The used key is always bright.

Haozhe Jiang | Oct 7, 2026 min read

Reconstruct, Practice, Go Real: Guided Self-Improvement for Embodied Agents

Agentic robots have gone viral on Twitter since the release of GPT-6. We see them driving cars, peeling cucumbers, and folding towels. Behind this hype, a lot of problems remain. GPT is really slow and expensive, and seems to have issues with task coverage. It might take 2 minutes to fold a single towel, consumes a lot of API credits, and fail when you rotate the towel by 45 degrees. Some issues may be solved as we scale the model further, while some may not. But perhaps a much deeper question we should be asking is: what new capabilities are emerging in foundation models, how are they contributing to robotic systems, and which of them provide new dimensions of scaling?

Why We Do Agentic Robotics

Let us think about a simple robotics setup: bimanual grippers. A naive question I asked before I entered the field of robotics was: Why is bimanual manipulation not solved? Though my robotics friends laughed at me at the time, the question came from genuine confusion. Most tasks that we expect grippers to perform do not require complicated movements. To a large extent, we only need to figure out where the contact points should be, how firmly the grippers should grasp, and how the end effectors should move. These seem fairly easy to encode in Python scripts, so why bother with all the learning-based methods? There is some truth to this argument. The early days of manipulation policies could largely be described as writing scripts by hand to perform specific tasks, and I believe a well-written script could perform tasks really reliably. The point of engineering, however, is not only whether something is doable, but also how cheaply and quickly we can do it.

Writing scripts by hand to control robots, unfortunately, takes a long time, and the scripts do not generalize well. As foundation models have made code generation cheap for a while now (yes, two years is a long time in AI nowadays), this approach has been revived in Code-as-Policy (such as RATs ([1]) and ASPIRE ([2])). Code-as-Policy lets a foundation model take task instructions, combine perception and control primitives into executable programs, and send the programs to robots to execute. A naive Code-as-Policy approach has two drawbacks:

  • A lot of decisions are easy for a multimodal model to make, but hard to write as explicit logic. For instance, it is hard to determine through explicit logic whether the gripper has successfully grasped a corner of a towel.
  • A program usually does not work on the first attempt. We may encounter failures, each of which seems easy to fix, but together they make the policy awkward to maintain. For example, in towel folding, the policy could fail to grasp a corner because the gripper is not pressing hard enough, fail to lift the corners high enough to keep the towel from sliding too much, or fail to place the corners in the desired positions. We need to patch these small bugs over multiple iterations while maintaining the original functions, and end up in an awkward loop that still involves a lot of labor.

This suggests that we also need the capabilities of visual perception and efficient iteration. These capabilities, especially the latter, are exactly what multimodal agents bring us this year! I want to emphasize here that efficient iteration is fundamentally different from cheap generation. Generation is about writing good prompts, while iteration is about defining a good environment and a good goal.

At the other extreme, we could use Agent-as-Policy: keep the agent itself in the control loop, letting it observe the scene, decide what to do, execute, and observe again. The distinction between Code-as-Policy and Agent-as-Policy is whether we use the agent for open-loop control or closed-loop control. Most of the Twitter videos showing GPT-6 performing robotic tasks are closer to Agent-as-Policy. This makes it easier to handle situations that we did not anticipate. But asking a large model to reason through every movement can be slow and expensive.

We want to take advantage of both approaches and find the right balance. More concretely, we should maintain a library of reusable and generalizable skills alongside the basic perception and control primitives. At runtime, we then let an agent control the robots using these skills. At a higher level, agentic robotics is about how to orchestrate these capabilities and, most importantly at this time, identify stages of the development pipeline where efficient iteration can expand the robot’s capabilities.

The RPG Pipeline

In RPG, we propose a concrete development loop for agentic robotics in bimanual manipulation: use existing offline dataset to decide what skills to build, construct environments to practice them, and let agents improve the skills before taking them to the real robot.

One key insight is that years of research on bimanual manipulation have already given us offline datasets demonstrating most of what tasks we want robots to perform. Furthermore, language models already have prior knowledge about these tasks and the coding capability to turn the procedures into scripts. Together, they give us a starting point for building a skill library with broad coverage.

In RPG, we start from ABC, a teleoperation dataset that covers 193 tasks. A coding agent, the Constructor, reviews the videos and descriptions to identify useful skills and implement them using basic perception and control tools. For example, organizing sunglasses involves picking, placing, and closing a case. These operations can be reused in other tasks, so the demonstrations help us decide what skills to build.

The Constructor then build corresponding simulation environments to develop and test these skills. Though not explicitly enforced in RPG, efficient iteration can help with real-to-sim reconstruction as well. For example, a coding agent can build an object model, inspect the rendered result, and revise its geometry. Many ABC tasks involve complex scenes and several operations, so we are not trying to recreate each task as a digital twin. We extract reusable skills and focus the reconstruction effort on the objects and interactions needed to practice them.

With an initial skill library and a place to practice, self-improvement (Practice) can refine both the skills and the Runtime Agent’s instructions for using them. The Runtime Agent observes the scene and chooses skills and their arguments. During practice, we run it in simulation, diagnose failures, revise the skill code or system prompt, and test the changes. This is where the efficient iteration from the first section makes the development loop scalable. Many fixes are locally simple, but maintaining a growing codebase of shared skills and tracking how changes affect different tasks would be awkward for a human. Coding agents are good at this kind of work: inspecting code, making targeted edits, running tests, and integrating changes. Over 15 practice rounds, RPG adds 23 skills and makes 66 modifications to existing ones.

RPG constructs simulation practice tasks from demonstrations and improves shared skills through privileged execution, video analysis, and tested revisions.

Demonstrations guide practice task construction. Execution feedback and video analysis guide improvements to shared skills and runtime instructions.

During self-improvement, we find two designs particularly useful:

  • Privileged Agent: Comparing execution with and without simulator state helps us understand what needs fixing. We run a second agent from the same initial states, using the same model and skill library but giving it exact object positions and sizes. Its successes can guide improvements to the Runtime Agent’s perception or action selection, while failures shared by both agents can expose defects in their skills.
  • Video Analyzer: The multimodal capability we discussed in the first section helps turn execution feedback into concrete corrections. We compare the two agents’ executions with dataset videos when available, using frames, skill calls, and execution records to identify where things first went wrong. Understanding the physical outcome of the code gives the implementation agent a clearer idea of what to change.

The two designs turn out to work particularly well together. After five practice rounds, the full system reaches 78.2% success, compared with about 51% when either the Video Analyzer or the Privileged Agent is removed. Better failure diagnosis makes the coding agent’s iterations much more productive.

Over 15 practice rounds, success on held-out initializations of 22 simulated tasks rises from 28.6% after the first round to 95.0%. The foundation model’s weights stay fixed throughout. What improves is the code and instructions around it, and these improvements persist across executions. This gives us another dimension of scaling: more practice can turn the same model into a more capable robotic system.

Finally, in Go Real, we calibrate the system for the physical robot and freeze it for deployment. The resulting system achieves a higher success rate and executes tasks faster than the GPT-6 baseline, while costing significantly less. It succeeds in all 30 trials across the three physical tasks we tested in the real world.

Towel folding at 15× speed: RPG (Gemini 3.8 Flash, left) costs about $0.10 per run, compared with about $1.00 for CaP-Agent0 ([3]) (GPT-6 Astra Pro, right). RPG also achieves a higher success rate and executes tasks faster. CaP-Agent0 uses GPT to control the robot through basic perception and control primitives, without the skills developed during Practice. Footage from the project demos.

What Next?

Through RPG, we study how several capabilities of coding agents can contribute to robotic systems:

  • Cheap generation: Agents identify useful skills from demonstrations and implement them as code using existing perception and control tools.
  • Visual perception: Multimodal models make judgments that are awkward to encode as explicit logic, both during execution and when diagnosing failures.
  • Efficient iteration: Agents can help refine practice environments and maintain the Runtime Agent’s instructions and shared skill library through repeated edits and tests.

There is still a lot to improve. RPG takes significantly longer than a trained policy to perform the same task, and could be less robust to disturbance. One promising application is to use it as a cheap data engine for training VLAs, so the skills developed through practice can help train faster policies. We also need better ways to merge skill revisions as we practice more tasks and changes to shared code begin to interact.

However, RPG gives us promising ways to scale robot capabilities. Agents can iterate efficiently to reconstruct useful simulation environments and improve the runtime system, including its skill library and the Runtime Agent’s instructions. We expect the scaling principles established and tested through RPG to transfer beyond the particular tasks and hardware used here.

We have so far evaluated RPG on YAM workstations, and we expect the same development principles to be useful for other bimanual grippers. Extending it to dexterous manipulation and locomanipulation will require datasets with similar breadth. It would be interesting to explore how to use agents to make up for this gap.

Reference

[1] Junyi Zhang, Jiaxin Ge, Hanjun Yoo and others. Playful Agentic Robot Learning. arXiv preprint arXiv:2606.19419, 2026.

[2] Runyu Lu, Yubo Wu, Ethan Kou and others. ASPIRE: Agentic /Skills Discovery for Robotics. arXiv preprint arXiv:2607.00272, 2026.

[3] Max Fu, Justin Yu, Karim El-Refai and others. CaP-X: A Framework for Benchmarking and Improving Coding Agents for Robot Manipulation. arXiv preprint arXiv:2603.22435, 2026.