中文
AI Engineer World's Fair

A reproducible experiment proves it: your character-AI evals can't catch a Hamilton who has read his own Broadway musical

The Miranda Hypothesis: How Hamilton Poisoned Persona Evals - Jacob E. Thomas, Results Gen · Jacob E. Thomas

58 min
EvalsContextAI ProductAgent

58 min total·Actually worth watching closely: ~25 min·3 must-watch clips

Orange = the 25 minutes worth watchingFor the rest, the guide is enough
Segment guide · 8 segments
  1. 0:05 8:17Watch

    Opening demo and the three paradigm stages

    Opens by showing the deployed scale of role-playing language agents and the open-source Companion framework, then asks Lincoln about war powers live and gets a plausible-sounding answer; from there it maps the field's three existing paradigm stages and states the central thesis.

    Hold on to "inherent executive authority" in that Lincoln answer — it sounds like Lincoln, but the whole talk will prove it is a twentieth-century construction.

    The live instantiation demo starts at 148s and is the most important setup in the talk; you need to see that fluent, plausible output with your own eyes for the later teardown to land.▶ Jump to 0:05
    Speaker · Jacob E. Thomas
  2. 8:22 17:49Watch

    The mask and the mirror: the Miranda hypothesis arrives

    Argues that "convincing" and "faithful" are separate properties and that current instruments (InCharacter's 80.7%) measure only the first; places the model's Hamilton beside the documentary Hamilton and introduces the Miranda hypothesis: the culturally dominant representation systematically overwrites the primary record.

    A Hamilton with a perfect personality-consistency score can simultaneously have read his own Broadway musical — and the evals are structurally unable to detect that dominant failure.

    There is something worth watching here: the side-by-side at 675s turns on the specific wording of the two texts (emotional narrative arc vs. legalistic syntax), which you can't catch by ear alone.▶ Jump to 8:22
    Speaker · Jacob E. Thomas
  3. 17:55 21:59Listen

    Evidence and mechanism: why alignment makes it worse

    Uses the Schuyler Mansion interpretive staff's work of un-teaching the musical as human-scale evidence for how corpus imbalance overwhelms the record; then argues that preference optimization rewards mythologized outputs because raters were shaped by the same cultural narratives, and that timelocked corpora don't rescue it either.

    Alignment is not a fix but algorithmic sycophancy: the repair has to happen at the encounter, not at the corpus substrate.

    This stretch is spoken argument and anecdotal evidence with no visual dependency — fine to listen to on a commute.▶ Jump to 17:55
    Speaker · Jacob E. Thomas
  4. 22:06 33:38Listen

    The fourth paradigm: epistemic simulation and its architectural choice

    Proposes epistemic simulation: the persona is a configuration rather than a property of the model, composed of a prompt, primary documents, a temporal anchor, an off-the-shelf model, and a human holding interpretive custody; then delivers the counterintuitive conclusion that for persona fidelity, fine-tuning is worse than context injection.

    "You don't train a persona, you assemble an encounter and keep the receipts" — the context window is auditable, revertible, and a kitchen-table capability available to anyone.

    None of the three sections depend on visuals; this is the densest stretch of argument in the talk and rewards focused listening. Concept-heavy, so consider replaying at speed.▶ Jump to 22:06
    Speaker · Jacob E. Thomas
  5. 33:46 40:39Skim

    The prism experiment: a pre-registered four-moment Lincoln design

    Uses the prism as the conceptual model (white light = the composite persona, prism = corpus plus temporal anchor, spectrum = several selves held apart), then lays out the matrix of four historical moments x three seeding conditions (primary sources / biography / bare model).

    C3, the bare model, is white light with no prism in the path — the control that establishes the Miranda distortion baseline; every prediction was pre-registered and time-stamped before data collection.

    The prism diagram and experimental matrix are slides — a glance at the matrix structure is enough to follow along; no need to watch every second.▶ Jump to 33:46
    Speaker · Jacob E. Thomas
  6. 40:43 47:36Watch

    The rubric and one cell dissected (the talk's high point)

    The rubric deliberately drops "does it sound like him" and weights anachronism detection at 40%; the talk then closes the loop on the opening demo, identifying that Lincoln answer as the output of cell C3 x 1847 and scoring it point by point against Lincoln's actual 1848 letter to Herndon.

    Plain but faithful must beat fluent but anachronistic — making voice a scoring axis would reward the exact error the instrument exists to catch.

    Strong visual dependency here: the Spielberg Lincoln clip, the text of the real letter, and the scored cell with observed results beside pre-registered predictions — the screen itself is the chain of evidence.▶ Jump to 40:43
    Speaker · Jacob E. Thomas
  7. 47:40 54:50Skim

    The six-step protocol and engineering the expert loop

    Lays out the reproducible six-step evaluation protocol and argues that the domain expert in the loop is a technical requirement, not a courtesy; the expert cost falls at build time and gate time (diagnostic questions, weighted rubric, sealed gold-standard vignettes), not at runtime.

    Fidelity is a relation between the output and a documentary record, which output-only automated metrics structurally cannot adjudicate — the persona must clear the expert gate before shipping and again after any base-model swap.

    From 3083s the speaker walks the slides through the practical implementation; the content is mostly bulleted checklists, so read the six-step protocol slide and listen to the rest.▶ Jump to 47:40
    Speaker · Jacob E. Thomas
  8. 54:59 58:13Listen

    Where it started and where it ends: Miranda distortion in a hospital room

    Reveals the personal origin of the whole project — an experience of Miranda distortion in a hospital room — and closes with an invitation: the six-step protocol is published in full with the paper, so any team with a frontier model can run their own figure in parallel.

    This is not one lab's benchmark but an open instrument — run it, with your own figure.

    An emotional spoken close with no visual dependency, but worth hearing in full to understand where the name "Miranda hypothesis" comes from.▶ Jump to 54:59
    Speaker · Jacob E. Thomas