Robots should be able to apply their manipulation skills across changing tasks and environments. Vision-language-action (VLA) models provide these skills, while an agent can monitor their execution, diagnose a missed grasp, and guide recovery. Bringing the two together connects learned robot control with reasoning about what happened and how to proceed.
Recursive Harness Distillation (RHD) stores intervention experience in a playbook. A strong agent explores a frozen vision-language-action policy and writes guidance for a lighter agent. It then revises the playbook using the lighter agent’s execution feedback. The teacher, recipient, and robot policy all keep their model weights fixed.
With GPT-6 Astra as the teacher and GPT-5.6 Luna as the recipient, RHD reaches 66.7% success on SimplerEnv Bridge, exceeding Astra’s 54.2% without a playbook. On a physical Franka robot, the harness raises success from 37.3% to 64.0%.
1. Embodied intelligence
Embodied demonstrations show agents connecting visual reasoning to physical action: Astra guides object manipulation, and spatial puzzle systems coordinate motion around interlocked geometry. Robot execution requires this reasoning throughout a task, as objects move, grasps slip, and the scene changes.
Pairing an agent with a VLA brings complementary capabilities to execution. The VLA generates manipulation actions; the agent tracks progress and intervenes when recovery is needed. After an empty grasp, it can redirect the policy toward the object, adjust the approach, and check that the object moves with the fingers before continuing toward the destination.
Effective recovery draws on experience with the particular policy: which instructions change its approach, where to focus attention, and how action corrections affect the next movement. RHD transfers this experience from Astra to Luna through a playbook, then refines the guidance using Luna’s execution feedback.
2. A harness for VLA execution
The harness combines an intervention interface with a playbook. The interface exposes the frozen VLA’s instructions, visual attention, and proposed actions. The playbook describes when to use these controls and what to inspect after the robot moves.
Instruction editing
The agent rewrites the language instruction supplied to the VLA. It can identify the source object, request a different grasp, or specify the next movement. After an empty lift, for example, a revised instruction can redirect the policy toward a new approach to the object.
Attention editing
The agent changes attention over visual features during policy inference. It can emphasize a relevant image region, such as the contact area between the fingers and the target. The policy generates a new action proposal using this intervention while retaining its learned parameters.
Action editing
The agent inspects the generated action sequence and adjusts motion or gripper commands. It can compare alternative proposals before moving, select a useful prefix, and check the resulting observation. Instruction and attention edits influence candidate generation; action edits act directly on the proposal.
The playbook connects these controls to observable situations. A missed grasp calls for a change in approach and a check that the object follows the fingers. A motion correction calls for inspecting the next proposal to verify that it preserves the intended adjustment.
3. The distillation pipeline
RHD alternates between robot execution and playbook revision. Astra supplies the initial experience and proposes updates. Luna’s subsequent interactions determine how that guidance is refined.
Explore and distill
Astra interacts with the frozen VLA, tries interventions, and reflects on their physical outcomes. It writes an initial playbook that describes the situations it encountered, the adjustments it made, and the observations that confirmed progress. These procedures give Luna a starting point for operating the policy.
Execute with the playbook
Luna receives the task, the current observation, and the playbook. It chooses an intervention, inspects the VLA’s proposal, and executes a selected action prefix. The next image lets it assess object motion, grasp state, and progress toward the goal. Its observations, decisions, proposals, and executed actions form the feedback for the teacher.
Refine from recipient feedback
Astra reviews Luna’s execution and revises the relevant guidance. The overview illustrates one such revision: after an empty lift, compare the failed grasp geometry, refocus on target–finger contact, and verify that the object follows the gripper. Luna then runs with the revised playbook, producing the next round of feedback.
This recursive step changes the quality of the guidance. On the simulation tasks, Luna reaches 31.3% with the initial teacher playbook and 66.7% after refinement. The teacher learns which instructions the recipient can apply effectively through the same robot interface.
Rollouts across playbook revisions
The recordings below compare Luna’s execution with two playbook versions. In the carrot task, the revision focuses on changing the grasp geometry and verifying object motion. In the cube task, it emphasizes object–target alignment while maintaining a closed grasp.
Carrot manipulation: revise the grasp, preserve the correction, and verify object motion.
4. Robot manipulation results
Simulation with GR00T
We evaluate RHD on spoon placement, carrot placement, cube stacking, and eggplant placement in SimplerEnv Bridge. Luna operates the frozen GR00T-N1.7 policy through the harness.

The refined playbook raises Luna’s success from 43.8% to 66.7%, a gain of 22.9 percentage points. The recipient also exceeds Astra operating without a playbook. Giving Astra the same refined playbook raises its success to 79.2%.
| Agent | Playbook | Success |
|---|---|---|
| GR00T | — | 41.7 |
| Luna + GR00T | None | 43.8 |
| Astra + GR00T | None | 54.2 |
| Luna + GR00T | Refined | 66.7 |
| Astra + GR00T | Refined | 79.2 |
Real-world manipulation
For a physical Franka Panda, Astra adapts the playbook to a task-trained π0.5 policy. Luna uses it for cube-to-tray placement, cube stacking, and button pressing. The robot policy remains frozen while the harness guides execution.
The adapted playbook improves overall success from 37.3% to 64.0%. All three tasks improve, with cube stacking increasing from 12% to 52%.
| Method | Cube to tray | Cube stacking | Button pressing | Overall |
|---|---|---|---|---|
| π0.5 | 60 | 12 | 40 | 37.3 |
| Luna + π0.5No playbook | 0 | 0 | 0 | 0.0 |
| Luna + π0.5Refined playbook | 72 | 52 | 68 | 64.0 |
Deployment cost
With the refined playbook, Luna’s estimated inference cost averages $0.63 per episode, compared with $3.43 for Astra. This is approximately 5.5× lower deployment inference cost. Astra develops and refines the guidance; Luna consults the resulting playbook when operating the robot.
5. How the playbook changes execution
Refinement makes the playbook more specific about the physical evidence behind an intervention. Object identification, contact geometry, source departure, and the effect of the next action proposal become explicit checks in the recipient’s decisions.
Grasp verification and recovery
A closed gripper is one part of a grasp check. The playbook also directs the agent to observe whether the object leaves the surface and moves with the fingers. After an unsuccessful attempt, the next intervention changes the contact geometry and checks its effect in the returned image.

The spoon sequence uses a revised instruction to redirect the grasp. The eggplant sequence uses action control to check acquisition through a lift. In the carrot sequence, attention editing emphasizes the finger–object region. Each intervention is followed by a new observation that informs the next decision.
Contributions of the intervention controls
Instruction editing contributes the largest gain among the three controls. Removing it reduces success from 66.7% to 47.9%. Removing attention editing yields 64.6%, and removing action editing yields 62.5%. With all three available, the agent can choose its intervention according to the scene and the policy’s response.
Learning from the recipient’s decisions
The playbook is refined around Luna’s execution: which rule it follows, what action it chooses, and how the robot responds. Astra can then revise a vague recovery instruction into a procedure that names the relevant geometry, the intervention, and the evidence needed to proceed. That revised procedure remains available in subsequent tasks without changing model weights.