Phone charging
Dual-arm alignment and insertion
100%



BEE
Intervention-Adaptive Real-World Reinforcement Learning with Vision-Language-Action Models
1 South China University of Technology 2 AgiBot
† Corresponding author
The frozen VLA still proposes the chunk. A human correction limits each dimension and is not copied as the action.
The paper has phone charging, snack hanging, cloth aligning, and bowl placing. Plug insertion and panel pressing were added later.
Best of those four, against 57.5% for RLT and 42.1% for DSRL. The two later tasks are 100/100 and are not in this average.
Corrections are used per dimension, not as a whole action to copy. Takeover is lower than RLT.
Vision-language-action (VLA) models handle long-horizon manipulation, yet success hinges on a few precision-critical phases where millimeter-scale errors undo all prior progress. Online reinforcement learning (RL) can optimize exactly these actions, but free exploration is far too costly on real robots, which makes human corrections indispensable. Existing online RL methods for VLAs either cannot incorporate such corrections or fold them into undifferentiated supervision. Human corrections are not uniformly noisy: they are reliable along some action dimensions and variable along others.
We introduce BEE, an intervention-adaptive framework for real-world RL on a frozen VLA that lets the policy go BEyond Expert imitation. Human corrections are not actions to reproduce, but evidence about a constraint. A Correction Model predicts how a human would correct a given VLA proposal, and how consistent that correction is along each action dimension. This predicted consistency sets the per-dimension tightness of a constraint on policy optimization. Where corrections are consistent the policy stays close to the human; where they vary, the constraint relaxes.
We evaluate BEE on three real-world manipulation tasks and one LIBERO-Pro simulation task at a matched online-data budget. BEE attains the highest success rate on every task, 91.2% on average against 57.5% for RLT and 42.1% for DSRL, and the lowest human intervention rate on all real-world tasks.
The VLA is fine-tuned on each task, then frozen, and still outputs the action chunk. BEE trains a residual added to that chunk. A human correction constrains the residual. It is not a label to imitate.
The VLA is fine-tuned per task and then frozen. Its action chunk stays an explicit proposal.
On a human takeover, the Correction Model predicts the residual relative to the VLA proposal and a variance on each action dimension. On an insertion, the downward correction is often consistent and the rotation is not. The constraint stays tight on the consistent part and relaxes on the rest.
Add the residual to the VLA chunk. Training raises return, stays near the VLA proposal, and pulls harder on dimensions where the human correction is consistent.
The four tasks with stills are the paper comparison. Plug insertion and dual-arm panel pressing were run later and are not in the 91.2% average. On those two, training converged within one hour, and each then succeeded on 100 of 100 tests.
Dual-arm alignment and insertion
100%



Deformable package onto a rack
85%



Grasp, transport, and seat the cloth
90%



LIBERO-Pro spatial suite, task 9
90%


Every method starts from the same fine-tuned π0.5. Each real-robot task gets 90 episodes of robot time, and bowl placing gets 45. BEE and RLT use about 20 episodes as human corrections collected before online training, then 70 online episodes on the real tasks and 25 on bowl placing. DSRL and DAgger have no correction buffer, so those 20 episodes are spent online instead.
Average success BEE 91.2% RLT 57.5% DSRL 42.1%
| Method | Phone charging | Snack hanging | Cloth aligning | Bowl placing | Overall | |||||
|---|---|---|---|---|---|---|---|---|---|---|
| SR ↑ | Int. ↓ | SR ↑ | Int. ↓ | SR ↑ | Int. ↓ | SR ↑ | Int. ↓ | SR ↑ | Int. ↓ | |
| Base policy | 90.0 ± 5.0 | – | 0.0 ± 0.0 | – | 36.7 ± 7.6 | – | 28.3 ± 7.6 | – | 38.8 | – |
| RLT | 58.3 ± 2.9† | 22.5 | 21.7 ± 5.8 | 62.1 | 73.3 ± 7.6 | 85.3 | 76.7 ± 17.6 | 26.2 | 57.5 | 49.0 |
| DSRL | 93.3 ± 2.9 | – | 13.3 ± 2.9 | – | 31.7 ± 2.9 | – | 30.0 ± 5.0 | – | 42.1 | – |
| BEE | 100.0 ± 0.0 | 12.2 | 85.0 ± 0.0 | 17.0 | 90.0 ± 10.0 | 65.3 | 90.0 ± 13.2 | 27.4 | 91.2 | 30.5 |
SR is success rate (%). Int. is the share of control steps a person takes over. The operator steps in when the robot is about to fail in a way it cannot recover from, so corrections fall on the contact phase. † RLT at this budget. With about 3× the online data it reaches 96.7 ± 5.8 on phone charging. Base policy and DSRL do not train with online intervention.
DAgger, SiLRI, and HIL-SERL get the same robot time. On phone charging, DAgger spends the 20 seed episodes online instead and falls from the 90% base policy to 63.3 ± 5.8. SiLRI and HIL-SERL train a visuomotor policy from scratch. On snack hanging and bowl placing they stay at 0% success, with intervention on more than 90% of steps. A few dozen episodes is not enough to train those policies on a task that lasts about a minute.
@misc{zhao2026bee,
title = {BEE: Intervention-Adaptive Real-World Reinforcement Learning with Vision-Language-Action Models},
author = {Weihui Zhao, Xiaohan Yan, Zunian Wan, Xuan Du, Zhaozhan Chi, Jianbo Mao, Ruipu Wu, Rushuai Yang, Houlin Li, Shukai Yang, Jing Wu, Yuxiang Yan, Yongcheng Liu, Chuankang Li, Guanghui Ren, Wei Shan, and Maoqing Yao},
year = {2026},
eprint = {2609.27450},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2609.27450}
}