Recursive Harness DistillationAcross Agents for Robot Manipulation

Seungyeon Kim1 Junhoo Lee2 Minkyu Kim1 Baekseung Kim1 Nojun Kwak1
1Seoul National University2Korea Advanced Institute of Science and Technology

Robots should be able to apply their manipulation skills across changing tasks and environments. Vision-language-action (VLA) models provide these skills, while an agent can monitor their execution, diagnose a missed grasp, and guide recovery. Bringing the two together connects learned robot control with reasoning about what happened and how to proceed.

Recursive Harness Distillation (RHD) stores intervention experience in a playbook. A strong agent explores a frozen vision-language-action policy and writes guidance for a lighter agent. It then revises the playbook using the lighter agent’s execution feedback. The teacher, recipient, and robot policy all keep their model weights fixed.

With GPT-6 Astra as the teacher and GPT-5.6 Luna as the recipient, RHD reaches 66.7% success on SimplerEnv Bridge, exceeding Astra’s 54.2% without a playbook. On a physical Franka robot, the harness raises success from 37.3% to 64.0%.

1. Embodied intelligence

Block into bowl. Astra guides a robot arm through pick-and-place. RoboCurve ↗
Robot puzzles. Interlocked parts and ring threading with Dual ALOHA. Qineng Wang ↗

Embodied demonstrations show agents connecting visual reasoning to physical action: Astra guides object manipulation, and spatial puzzle systems coordinate motion around interlocked geometry. Robot execution requires this reasoning throughout a task, as objects move, grasps slip, and the scene changes.

Pairing an agent with a VLA brings complementary capabilities to execution. The VLA generates manipulation actions; the agent tracks progress and intervenes when recovery is needed. After an empty grasp, it can redirect the policy toward the object, adjust the approach, and check that the object moves with the fingers before continuing toward the destination.

Effective recovery draws on experience with the particular policy: which instructions change its approach, where to focus attention, and how action corrections affect the next movement. RHD transfers this experience from Astra to Luna through a playbook, then refines the guidance using Luna’s execution feedback.

2. A harness for VLA execution

The harness combines an intervention interface with a playbook. The interface exposes the frozen VLA’s instructions, visual attention, and proposed actions. The playbook describes when to use these controls and what to inspect after the robot moves.

Paper Figure 1: a strong agent interacts with a frozen VLA, revises a playbook, and refines it using a light agent’s execution experience
Recursive Harness Distillation. A strong agent refines the playbook through a light agent’s robot interactions. Figure 1 from the paper.

Instruction editing

The agent rewrites the language instruction supplied to the VLA. It can identify the source object, request a different grasp, or specify the next movement. After an empty lift, for example, a revised instruction can redirect the policy toward a new approach to the object.

Attention editing

The agent changes attention over visual features during policy inference. It can emphasize a relevant image region, such as the contact area between the fingers and the target. The policy generates a new action proposal using this intervention while retaining its learned parameters.

Action editing

The agent inspects the generated action sequence and adjusts motion or gripper commands. It can compare alternative proposals before moving, select a useful prefix, and check the resulting observation. Instruction and attention edits influence candidate generation; action edits act directly on the proposal.

The playbook connects these controls to observable situations. A missed grasp calls for a change in approach and a check that the object follows the fingers. A motion correction calls for inspecting the next proposal to verify that it preserves the intended adjustment.

3. The distillation pipeline

RHD alternates between robot execution and playbook revision. Astra supplies the initial experience and proposes updates. Luna’s subsequent interactions determine how that guidance is refined.

Explore and distill

Astra interacts with the frozen VLA, tries interventions, and reflects on their physical outcomes. It writes an initial playbook that describes the situations it encountered, the adjustments it made, and the observations that confirmed progress. These procedures give Luna a starting point for operating the policy.

Execute with the playbook

Luna receives the task, the current observation, and the playbook. It chooses an intervention, inspects the VLA’s proposal, and executes a selected action prefix. The next image lets it assess object motion, grasp state, and progress toward the goal. Its observations, decisions, proposals, and executed actions form the feedback for the teacher.

Refine from recipient feedback

Astra reviews Luna’s execution and revises the relevant guidance. The overview illustrates one such revision: after an empty lift, compare the failed grasp geometry, refocus on target–finger contact, and verify that the object follows the gripper. Luna then runs with the revised playbook, producing the next round of feedback.

This recursive step changes the quality of the guidance. On the simulation tasks, Luna reaches 31.3% with the initial teacher playbook and 66.7% after refinement. The teacher learns which instructions the recipient can apply effectively through the same robot interface.

Rollouts across playbook revisions

The recordings below compare Luna’s execution with two playbook versions. In the carrot task, the revision focuses on changing the grasp geometry and verifying object motion. In the cube task, it emphasizes object–target alignment while maintaining a closed grasp.

Luna + GR00T
Earlier playbookV15
Revised playbookV16

Carrot manipulation: revise the grasp, preserve the correction, and verify object motion.

4. Robot manipulation results

Simulation with GR00T

We evaluate RHD on spoon placement, carrot placement, cube stacking, and eggplant placement in SimplerEnv Bridge. Luna operates the frozen GR00T-N1.7 policy through the harness.

Real-world cube-to-tray, stacking, and button-pressing tasks, alongside the four SimplerEnv Bridge tasks
Real-world and simulation manipulation tasks. Figure 2 from the paper.

The refined playbook raises Luna’s success from 43.8% to 66.7%, a gain of 22.9 percentage points. The recipient also exceeds Astra operating without a playbook. Giving Astra the same refined playbook raises its success to 79.2%.

Simulation success (%)
AgentPlaybookSuccess
GR00T—41.7
Luna + GR00TNone43.8
Astra + GR00TNone54.2
Luna + GR00TRefined66.7
Astra + GR00TRefined79.2

Real-world manipulation

For a physical Franka Panda, Astra adapts the playbook to a task-trained π0.5 policy. Luna uses it for cube-to-tray placement, cube stacking, and button pressing. The robot policy remains frozen while the harness guides execution.

The adapted playbook improves overall success from 37.3% to 64.0%. All three tasks improve, with cube stacking increasing from 12% to 52%.

Real-world success (%)
MethodCube to trayCube stackingButton pressingOverall
Ï€0.560124037.3
Luna + π0.5No playbook0000.0
Luna + π0.5Refined playbook72526864.0

Deployment cost

With the refined playbook, Luna’s estimated inference cost averages $0.63 per episode, compared with $3.43 for Astra. This is approximately 5.5× lower deployment inference cost. Astra develops and refines the guidance; Luna consults the resulting playbook when operating the robot.

5. How the playbook changes execution

Refinement makes the playbook more specific about the physical evidence behind an intervention. Object identification, contact geometry, source departure, and the effect of the next action proposal become explicit checks in the recipient’s decisions.

Grasp verification and recovery

A closed gripper is one part of a grasp check. The playbook also directs the agent to observe whether the object leaves the surface and moves with the fingers. After an unsuccessful attempt, the next intervention changes the contact geometry and checks its effect in the returned image.

Paper Figure 7: instruction, action, and attention interventions during spoon, eggplant, and carrot manipulation
Interventions during execution. Spoon, eggplant, and carrot sequences illustrate how the agent applies instruction, action, and attention edits. Figure 7 from the paper.

The spoon sequence uses a revised instruction to redirect the grasp. The eggplant sequence uses action control to check acquisition through a lift. In the carrot sequence, attention editing emphasizes the finger–object region. Each intervention is followed by a new observation that informs the next decision.

Contributions of the intervention controls

Instruction editing contributes the largest gain among the three controls. Removing it reduces success from 66.7% to 47.9%. Removing attention editing yields 64.6%, and removing action editing yields 62.5%. With all three available, the agent can choose its intervention according to the scene and the policy’s response.

Learning from the recipient’s decisions

The playbook is refined around Luna’s execution: which rule it follows, what action it chooses, and how the robot responds. Astra can then revise a vague recovery instruction into a procedure that names the relevant geometry, the intervention, and the evidence needed to proceed. That revised procedure remains available in subsequent tasks without changing model weights.