dice-rl-wam
DICE-RL post-training of a video-action world model on LIBERO-10
2026-09
Post-trained LingBot-VA, a video-action world model, on LIBERO-10 with DICE-RL: supervised fine-tuning on 300 demonstrations, then residual reinforcement learning against the frozen prior with an ensemble of ten critics.
- Model
- LingBot-VA, ~5B video-action world model
- Fine-tuning
- 300 demonstrations, 8× H100
- Reinforcement learning
- Residual MLP actor, ensemble of ten critics