BEE

Intervention-Adaptive Real-World Reinforcement Learning with Vision-Language-Action Models

Weihui Zhao1,2, Xiaohan Yan2,†, Zunian Wan2, Xuan Du2, Zhaozhan Chi2, Jianbo Mao2 Ruipu Wu2, Rushuai Yang2, Houlin Li2, Shukai Yang2, Jing Wu2, Yuxiang Yan2 Yongcheng Liu2, Chuankang Li2, Guanghui Ren2, Wei Shan2, Maoqing Yao2

1 South China University of Technology   2 AgiBot

† Corresponding author

PaperarXiv CodeComing Soon
Overview of BEE. A frozen VLA proposes an action chunk. A correction model turns human interventions into a per-dimension constraint, and a residual policy optimizes return subject to that constraint.
During training, a human correction constrains the residual on top of the frozen VLA chunk. At test time only the VLA and the residual run.
Method Constraint on the residual

The frozen VLA still proposes the chunk. A human correction limits each dimension and is not copied as the action.

Experiment 5 real, 1 sim

The paper has phone charging, snack hanging, cloth aligning, and bowl placing. Plug insertion and panel pressing were added later.

Result 91.2% on the paper tasks

Best of those four, against 57.5% for RLT and 42.1% for DSRL. The two later tasks are 100/100 and are not in this average.

Intervention 30.5% vs 49.0%

Corrections are used per dimension, not as a whole action to copy. Takeover is lower than RLT.

Abstract

Vision-language-action (VLA) models handle long-horizon manipulation, yet success hinges on a few precision-critical phases where millimeter-scale errors undo all prior progress. Online reinforcement learning (RL) can optimize exactly these actions, but free exploration is far too costly on real robots, which makes human corrections indispensable. Existing online RL methods for VLAs either cannot incorporate such corrections or fold them into undifferentiated supervision. Human corrections are not uniformly noisy: they are reliable along some action dimensions and variable along others.

We introduce BEE, an intervention-adaptive framework for real-world RL on a frozen VLA that lets the policy go BEyond Expert imitation. Human corrections are not actions to reproduce, but evidence about a constraint. A Correction Model predicts how a human would correct a given VLA proposal, and how consistent that correction is along each action dimension. This predicted consistency sets the per-dimension tightness of a constraint on policy optimization. Where corrections are consistent the policy stays close to the human; where they vary, the constraint relaxes.

We evaluate BEE on three real-world manipulation tasks and one LIBERO-Pro simulation task at a matched online-data budget. BEE attains the highest success rate on every task, 91.2% on average against 57.5% for RLT and 42.1% for DSRL, and the lowest human intervention rate on all real-world tasks.

Method

The VLA is fine-tuned on each task, then frozen, and still outputs the action chunk. BEE trains a residual added to that chunk. A human correction constrains the residual. It is not a label to imitate.

  1. 01 Keep the proposal

    The VLA is fine-tuned per task and then frozen. Its action chunk stays an explicit proposal.

  2. 02 Model the correction

    On a human takeover, the Correction Model predicts the residual relative to the VLA proposal and a variance on each action dimension. On an insertion, the downward correction is often consistent and the rotation is not. The constraint stays tight on the consistent part and relaxes on the rest.

  3. 03 Optimize the residual

    Add the residual to the VLA chunk. Training raises return, stays near the VLA proposal, and pulls harder on dimensions where the human correction is consistent.

Human Data Analysis

On phone charging, the correction model matches human corrections more closely than the VLA proposal, and its predicted spread follows the human spread on each joint.
On phone charging, the predicted correction is closer to the human than the VLA proposal on all seven intervened joints. The predicted spread follows the human spread, which varies by about 3× from the most consistent joint, q2, to the most variable, q7.
Constraint share across seven joints. A state-wise penalty loads the noisiest joint. The Mahalanobis penalty loads the most consistent joint.
A state-wise squared error gives the largest share to q5, one of the least consistent joints. Weighting by the predicted variance moves that share to q2, cuts the share of q5 by more than half, and leaves q7 with the smallest share.

Experiments

The four tasks with stills are the paper comparison. Plug insertion and dual-arm panel pressing were run later and are not in the 91.2% average. On those two, training converged within one hour, and each then succeeded on 100 of 100 tests.

Experiments Details

Phone charging

Dual-arm alignment and insertion

100%
Robot approaching the charger
Approach the charger
Gripper grasping the charger
Grasp the charger
Charger aligned with the socket
Align with the socket
Charger inserted into the socket
Insert the charger

Snack hanging

Deformable package onto a rack

85%
Robot approaching a snack package
Approach the snack
Gripper grasping the snack
Grasp the snack
Snack lifted toward the rack
Lift toward the rack
Snack hanging on the rack
Hang the snack

Cloth aligning

Grasp, transport, and seat the cloth

90%
Robot grasping cloth
Grasp the cloth
Cloth lifted off the table
Lift the cloth
Cloth moved over the frame
Move over the frame
Cloth aligned on the frame
Align onto the frame

Bowl placing

LIBERO-Pro spatial suite, task 9

90%
Simulated robot grasping a bowl
Grasp the bowl
Bowl held above the plate
Move above the plate
Bowl placed on the plate
Place the bowl

BEE Rollouts

Phone charging The left arm lifts the cable. The right arm grasps the charger, aligns it with the socket, and inserts it.
Snack hanging Grasp the deformable package, lift it to the rack, and hang it by the hole.
Plug insertion The arm aligns the plug with the socket and pushes it in.
Dual-arm panel pressing The two arms lower the panel onto the fixture and press it flat.
Cloth aligning Grasp the cloth, lift it over the target frame, and align it onto the frame.
Bowl placing Grasp the bowl, move it above the plate, and place it down. LIBERO-Pro spatial task 9.

Results

Every method starts from the same fine-tuned π0.5. Each real-robot task gets 90 episodes of robot time, and bowl placing gets 45. BEE and RLT use about 20 episodes as human corrections collected before online training, then 70 online episodes on the real tasks and 25 on bowl placing. DSRL and DAgger have no correction buffer, so those 20 episodes are spent online instead.

Average success BEE 91.2% RLT 57.5% DSRL 42.1%

Method Phone charging Snack hanging Cloth aligning Bowl placing Overall
SR ↑Int. ↓ SR ↑Int. ↓ SR ↑Int. ↓ SR ↑Int. ↓ SR ↑Int. ↓
Base policy 90.0 ± 5.0– 0.0 ± 0.0– 36.7 ± 7.6– 28.3 ± 7.6– 38.8–
RLT 58.3 ± 2.9†22.5 21.7 ± 5.862.1 73.3 ± 7.685.3 76.7 ± 17.626.2 57.549.0
DSRL 93.3 ± 2.9– 13.3 ± 2.9– 31.7 ± 2.9– 30.0 ± 5.0– 42.1–
BEE 100.0 ± 0.012.2 85.0 ± 0.017.0 90.0 ± 10.065.3 90.0 ± 13.227.4 91.230.5

SR is success rate (%). Int. is the share of control steps a person takes over. The operator steps in when the robot is about to fail in a way it cannot recover from, so corrections fall on the contact phase. † RLT at this budget. With about 3× the online data it reaches 96.7 ± 5.8 on phone charging. Base policy and DSRL do not train with online intervention.

DAgger, SiLRI, and HIL-SERL get the same robot time. On phone charging, DAgger spends the 20 seed episodes online instead and falls from the 90% base policy to 63.3 ± 5.8. SiLRI and HIL-SERL train a visuomotor policy from scratch. On snack hanging and bowl placing they stay at 0% success, with intervention on more than 90% of steps. A few dozen episodes is not enough to train those policies on a task that lasts about a minute.

Success rate versus online episodes for BEE and RLT on phone charging and snack hanging.
On phone charging, BEE reaches the base policy within a few episodes and then stays there. On snack hanging the base policy is 0%, and BEE rises inside the same budget. RLT is below BEE on both.
Ablation of success and intervention rate on phone charging and snack hanging.
On snack hanging, BEE is 85% success and 17% intervention. Without the VLA proposal in the correction model: 15%. Using the raw takeover as the target: 20% success and 55.5% intervention. Without the per-dimension constraint: 30% success and 62.8% intervention. Without the encoder: 60% success. The drop from removing the encoder is smaller than the other three.

Citation

@misc{zhao2026bee,
  title  = {BEE: Intervention-Adaptive Real-World Reinforcement Learning with Vision-Language-Action Models},
  author = {Weihui Zhao, Xiaohan Yan, Zunian Wan, Xuan Du, Zhaozhan Chi, Jianbo Mao, Ruipu Wu, Rushuai Yang, Houlin Li, Shukai Yang, Jing Wu, Yuxiang Yan, Yongcheng Liu, Chuankang Li, Guanghui Ren, Wei Shan, and Maoqing Yao},
  year   = {2026},
  eprint = {2609.27450},
  archivePrefix = {arXiv},
  url    = {https://arxiv.org/abs/2609.27450}
}