中文
Figma Config

Designers don't need to learn ML: writing the North Star and working through disagreement, the craft you already have, is what it takes to lead AI evals.

Own the loop: how designers and researchers win with AI evals ft. Setor Zilevu (Figma) | Config 2026 · Rachel

19 min
EvalsAI Product

18 min total·Actually worth watching closely: ~6 min·3 must-watch clips

Orange = the 6 minutes worth watchingFor the rest, the guide is enough
Segment guide · 7 segments
  1. 0:11 2:20Listen

    Opening: why a designer is giving the evals talk

    Rachel and Megan host the Mezzanine stage and introduce Setor Zilevu, setting up the premise — AI evals are more than a raw technical tool, and they're not reserved for machine learning scientists.

    This isn't an ML lecture; it's a skills-transfer session for design and research practitioners.

    Pure intro and speaker bio, nothing on screen — just settle into the context.▶ Jump to 0:11
    Speaker · Rachel / Megan
  2. 2:20 5:00Watch

    John's story: passed every test, still wrong

    Setor walks through the cyberhuman system he helped build for stroke survivors — high-quality therapist feedback at home using interactive apps and wearable sensors — and what happened when it met John in the pilot at Emory University.

    An accurate model isn't a correct product; the definition of good has to match the domain expert and the real user.

    The rehab app is on screen twice (153s and 189s, both flagged visual moments) — you need to see it to understand what went wrong.▶ Jump to 2:20
    Speaker · Setor Zilevu
  3. 5:00 7:12Watch

    With no criteria, judgment is guesswork

    From the Michael Scott meme to a live comparison of two designs generated from the same prompt, showing there's no basis for calling one AI output better when no criteria were defined in advance.

    Answer what good means before you evaluate anything — otherwise every score is a guess.

    Visual moments cluster here (346s, 369s, 403s); the two-design comparison is live audience interaction you have to watch.▶ Jump to 5:00
    Speaker · Setor Zilevu
  4. 7:12 10:55Listen

    Demystifying evals: two myths and the loop

    Busts the myths that evals are strictly an engineering problem and that good only needs defining once, notes that human expectations keep evolving, then uses the Edna analogy to lay out the method: define intent, set success criteria, measure, learn, repeat.

    AI output is probabilistic, so quality can't be verified at the end — it has to be directed continuously.

    Mostly spoken argument you follow by ear; the one stretch worth looking up for, Edna, is already listed as a must-watch clip.▶ Jump to 7:12
    Speaker · Setor Zilevu
  5. 10:55 14:13Skim

    Three layers: human, auto, UX

    Explains human evals (rubric plus grader calibration), auto evals (an LLM taught to apply the calibrated rubric at scale) and UX evals (restoring the real context that vacuum scoring misses).

    If your graders aren't calibrated, auto evals scale noise rather than judgment; UX evals cover the blind spot the first two layers share.

    Framework slides — glance at the three-layer structure and listen closely for the calibration and blind-spot arguments.▶ Jump to 10:55
    Speaker · Setor Zilevu
  6. 14:13 16:40Listen

    Figma in practice: the failure-mode flywheel

    Figma's own experience: triage failure modes by frequency and severity, hand them to engineering to fix one by one, and watch the scores climb until leadership takes notice.

    Once quality work is organized around failure modes, design and research recommendations go from indirect influence to direct strategic input.

    Narrative rather than demo — the value is in the case details and the causal chain, so give it your attention.▶ Jump to 14:13
    Speaker · Setor Zilevu
  7. 16:40 18:15Listen

    The four levers, and the call to own the loop

    Sums up the four levers a practitioner has over how a model behaves — system prompts, eval design, data sets and capabilities — and closes out the session.

    Your existing skills renamed are eval skills: defining the North Star becomes writing the rubric, synthesizing findings becomes triaging failure modes.

    A spoken call to action and skill-mapping list — no need to watch the screen, just note the four levers.▶ Jump to 16:40
    Speaker · Setor Zilevu