Meta tells it firsthand: how AI design output went from "looks about right" to 87% production-ready, and the whole method is yours to copy
Designed to be read: making a system AI can actually use ft. Harvey Whiting (Meta) | Config 2026 · Harvey Whiting
15 min total·Actually worth watching closely: ~2 min·2 must-watch clips
- 0:12 – 3:20Listen
AI slop: design that's fast but wrong
Opens with the problem: most AI-generated design looks plausible, but the moment you put it next to a real design system it falls apart — failing in three consistent patterns: low fidelity, low repeatability, and context collapse under complexity.
Low repeatability — the same prompt giving a different result each time — erodes designers' trust more than any single bad output
This stretch lays out the problem in spoken argument with no key visuals flagged; listening on a commute costs you nothing▶ Jump to 0:12Speaker · Harvey Whiting - 3:20 – 5:30Listen
Three root causes and one verdict
Breaks down why the agent was dropping the ball: models trained on the open web don't know proprietary systems and don't know what they don't know; the design system is too complex for people to hold in full; and expecting one-shot output was never reasonable.
"Speed without quality is just expensive noise" — designers won't use a tool where speed comes at the cost of quality
Pure argument, dense with quotable lines but not dependent on visuals — good to listen to▶ Jump to 3:20Speaker · Harvey Whiting - 5:30 – 7:13Skim
Golden sets: giving AI the chance to iterate
If iteration is how designers reach quality, give AI the same opportunity: run the system through golden sets repeatedly, put a metric on every round of output, and set the pass bar at "at least 80% high fidelity".
Turns "is the AI output any good" from a gut feeling into a metric you can measure and iterate against
The evaluation method comes with slides; a glance at the metric definitions is enough, the value is in the narration▶ Jump to 5:30Speaker · Harvey Whiting - 7:13 – 10:25Skim
87%, then MCP
After iterative evaluation pushed the high-fidelity rate to 87%, Meta found manual feedback loops don't scale and turned to MCP to build a context delivery system: components, documentation, color, spacing, content standards and every other dimension captured in infrastructure for any AI tool to pull precisely, on demand.
The foundation is context — if your design system isn't structured in a way AI can consume, you're building on sand
The 87% results slide is the key visual moment flagged by the LLM (433s) and worth a pause; the rest is architecture diagrams you can skim▶ Jump to 7:13Speaker · Harvey Whiting - 10:29 – 13:00Skim
The third pillar: agents arrive
Context is knowledge, skills are knowing when and why to apply it; hand both to agents and you get a task handed over at night with results in the morning, agents exploring 10 directions in parallel, and a full feature design from about six words.
These tools are arriving faster than anyone expected; the teams with context and skills ready will be the fastest movers
633s is a visual moment flagged by the LLM — the agent workflow demo is worth watching; skim it alongside the narration▶ Jump to 10:29Speaker · Harvey Whiting - 13:00 – 15:29Listen
Three takeaways and the paradigm shift
Closes with three actionable takeaways — audit your context, embrace the feedback loop (re-prompting isn't rework, it's the workflow), get ready for agents — and names the paradigm shift: with context, skills and agents in place, design is becoming genuinely multimodal.
Treat AI like a new designer joining the team: give it the full background and problem definition, and treat iterating and correcting as the workflow itself
The action list and closing vision are mostly spoken with nothing being demoed; note down the three takeaways and you're set▶ Jump to 13:00Speaker · Harvey Whiting