Collecting and Scheduling Demonstrations for Multi-Step Precision Manipulation with VLAs

Published in IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), 2026

We present a complete pipeline for applying vision-language-action (VLA) models to multi-step precision manipulation: an affordable, leader-free teleoperation system for collecting high-quality demonstrations, and a systematic study of how those demonstrations should be scheduled during fine-tuning. A single operator controls both arms and the mobile base of a Mobile ALOHA platform from one wireless gamepad, and uses this system to collect 608 real-robot episodes for a tube pickup-and-insertion task. Fine-tuning π₀.₅ with nine recipes that vary the placement and quantity of failure-recovery (“correction”) demonstrations shows that placement is the dominant factor: the same 51 correction episodes yield full-task success from 0.00 to 0.70 depending only on where they enter the schedule, and the winning three-stage recipe is also the fastest.

Recommended citation: S. Yang. "Collecting and Scheduling Demonstrations for Multi-Step Precision Manipulation with VLAs. " IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS). 2026.