Let the AI audit its own design work: two months of build and QA down to two hours
The metacognitive design loop ft. Jenny Au & Mike Green | Config 2026 · Mike Green
19 min total·Actually worth watching closely: ~7 min·3 must-watch clips
- 0:12 – 2:07Listen
The gatekeeper role is failing
The opening problem: design teams have long relied on manual gatekeeping to hold consistency, but once AI starts producing screens at volume, that line of defense breaks down.
The problem isn't that the AI draws badly — it's that the manual review step can't keep up with the output.
Mostly the speakers on stage laying out the background, with just a title slide on screen. Fine to listen to while doing something else.▶ Jump to 0:12 - 2:07 – 4:20Skim
4 billion combinations: the math doesn't work
They open up their design system and put numbers on the governance scale: 50-plus components, 200-plus style tokens, 100-plus rules, roughly 4 billion combinations — against a QA ratio of about 1 to 100.
Governance also swings on a pendulum of failure: too strict and the workflow slows to a halt, too loose and a thousand small inconsistencies pull the product slowly out of alignment.
The screen shows the design system's component panel and a few figures. A glance to register the order of magnitude is enough; there's nothing extra in listening word for word.▶ Jump to 2:07Speaker · Mike Green - 4:20 – 7:15Watch
The meta-cognitive loop: the AI checks itself before you see it
The core proposal: insert a self-check step after plan and build, where the AI corrects repeatedly against the design system's rules until every one passes, and each correction is solidified as a new rule by a commit.
The system gets smarter with use, because the fix for a mistake is deposited rather than re-explained every time.
The order and feedback paths between the loop's steps live entirely in that diagram — the arrows are the manual for this whole method. Audio alone falls apart here.▶ Jump to 4:20Speaker · Jenny Au - 7:15 – 10:10Listen
Two defenses: one for the code, one for the strategy
Introduces their own Core UI connection layer, wired straight to Figma live tokens and the latest component catalog as the source of truth, alongside the BRD agent and its three steps: a structured interview with the PO, translation into exact components, then explicit approval before handoff.
The worst moment to discover what you're building is wrong is after development has begun, so QA moves to the very front of the workflow.
Mostly an argument about what each defense covers and why the split falls that way; the screen is just bullets backing the words.▶ Jump to 7:15 - 10:10 – 12:29Watch
One component, two names
A live look at what surfaced while building the component mapping files: a site header goes by completely different names in the coded library and in the design file and component docs.
Consistent naming isn't aesthetic fussiness — it's the precondition for prompts working predictably, and this trap actually caught them.
The moment the two naming conventions sit side by side on screen is what lands, and the speaker is pointing right at it. Without the visual you barely feel how serious the problem is.▶ Jump to 10:10Speaker · Mike Green - 12:29 – 14:18Watch
The wreck: a table split in half
A failure case: the AI handled the header and the data rows as two independent layout problems, so the columns fell out of line and the grid broke.
An ungoverned agentic workflow doesn't just produce mistakes — it can drag overall consistency below where you started.
Exactly where the misalignment is and which column breaks only registers when you see that skewed table; no amount of narration fills that in.▶ Jump to 12:29Speaker · Mike Green - 14:18 – 19:26Listen
Two-hour delivery, and the four steps you can start tomorrow
The results and the reproducible path: two months of build and QA compressed to two hours with no loss of precision, via four steps — move QA up to the requirements phase, get design and code onto one shared language, let the loop self-check and self-correct, and mentor the AI so corrections become encoded rules.
Small teams need only three things to start: treat the design system as a source of truth you can code against, ask the AI why it used that particular spacing instead of telling it to fix the spacing, and build an immune system by encoding every correction.
The closing is mostly a verbal summary of the blueprint and the role shift, with nothing on screen worth watching. Catching those three things and the order of the four steps is enough.▶ Jump to 14:18