中文
Figma Config

Boston Dynamics from the inside: how generally capable hardware turns robotics into a software problem

Designing the product that can do anything ft. Brian Ringley | Config 2026 · Megan

21 min
AgentAI ProductEvals

21 min total·Actually worth watching closely: ~10 min·3 must-watch clips

Orange = the 10 minutes worth watchingFor the rest, the guide is enough
Segment guide · 6 segments
  1. 0:00 4:02Listen

    Why humanoids: an economics breakthrough, not a technology breakthrough

    After the Figma hosts introduce him, Brian states the core claim: invest once in generalized hardware, program everything with software, and you basically turn every application problem into a software problem. Atlas is going into Hyundai manufacturing facilities to fill edge cases where automation is of benefit but isn't economically viable with custom hardware.

    The value of the humanoid is in the economic model — generalized hardware turns a robotics problem into a software problem

    A purely spoken business argument with no key visuals — fine to listen to like a podcast▶ Jump to 0:00
    Speaker · Brian Ringley
  2. 4:02 7:16Watch

    Hardware trade-offs: look like a machine, not a person

    Three design decisions in a row: keep a face and directionality so workers sense they've been noticed and can anticipate the next movement; deliberately avoid looking human so nobody is tricked; and build the whole body from just two actuators, with arm and leg components nearly interchangeable and joints that rotate infinitely.

    Simplicity and symmetry hold down cost and protect reliability, and they also let the robot take the most efficient motion for the task and produce emergent behaviors

    Three visual moments cluster here (the face and directionality, the non-humanlike form, the invertible joints); these design decisions only land when you're looking at the robot itself▶ Jump to 4:02
    Speaker · Brian Ringley
  3. 7:16 11:44Listen

    The skill stack: reinforcement learning to walk, behavior cloning to work

    Mobility comes from reinforcement learning — give a motion example, simulate it many, many times, transfer to hardware, and 100 times per second the robot predicts its next actuator state. Dexterous manipulation comes from behavior cloning — an expert human demonstrates many, many times, and 30 times a second the policy picks its next moves from what it sees and the state of its body.

    The two training paths split cleanly — RL for mobility, behavior cloning for manipulation — but the latter has no semantic knowledge of what it's doing

    Mostly methodology over slides; the spoken content carries it, and understanding the split between the two paths is enough▶ Jump to 7:16
    Speaker · Brian Ringley
  4. 11:44 14:31Watch

    Adding the semantic layer: agents all the way down

    A behavior-cloned policy can act but doesn't know what it's doing, so visual language models associate images semantically with your language commands and visual language action models pass that on to the behavior policy. To collect more demonstrations the team gets pushed out of familiar web work into VR: the more the demonstrator forgets themselves, the higher the data quality.

    Robots have no internet-scale corpus, so behavior data can only be gathered one demonstration at a time by a person in a VR headset

    The VR teleoperated demonstration (around 795s) is worth pulling frames from — becoming the robot is far more vivid seen than heard▶ Jump to 11:44
    Speaker · Brian Ringley
  5. 14:31 18:46Watch

    An unprecedented UX problem and a zero-cost eval

    With tracking gloves on you can't press any buttons, and every tiny motion is controlling the robot — traditional input falls apart. Then comes the highlight of the talk: put a camera on your forehead, tell the agent it's the robot in the system prompt, wire it to a homemade web app, and you can test the agent's ability to receive a task instruction from home.

    Evaluating the upper-level agent needs neither a real robot nor a simulator — one person and one camera is enough

    Both visual moments are in this stretch (871s and 956s); the pretend-to-be-the-robot eval is the single thing most worth seeing with your own eyes▶ Jump to 14:31
    Speaker · Brian Ringley
  6. 18:46 21:03Listen

    The endgame: the factory itself is the robot

    A generalized platform needs 10,000 different software applications, and talking to your software is the way out at that scale. Agile humanoids fill the gaps, traditional industrial automation does the rest, and heterogeneous fleets are orchestrated by high-level factory systems organized by manufacturing execution systems.

    The humanoid is only a means to an end; the software-defined factory as a whole is the real robot

    The closing is a vision statement with no demo footage — the conclusion is all you need▶ Jump to 18:46
    Speaker · Brian Ringley