首页 /研究 /An Empirical Study on Stage-Information Interfaces for VLA Fine-Tuning
MANIPULATION

An Empirical Study on Stage-Information Interfaces for VLA Fine-Tuning

Yingwei Ji

发表年份
2026
访问权限
开放获取

摘要

One high-level instruction in long-horizon manipulation can cover several action stages. We use segmented action annotations as an intermediate representation between the full-task instruction and VLA action chunks. A progress module tracks the active stage, while the action policy receives stage information either as current-stage text or as a normalized ordinal stage index in robot state. We compare these interfaces with GR00T N1.6 on LIBERO-10 under direct fine-tuning and continuation fine-tuning from a full-task instruction baseline. Under direct fine-tuning, full-task instruction, current-stage text, and Ordinal Stage-State achieve mean success rates of 57.45%, 50.24%, and 54.36%, respectively, showing that explicit stage information does not automatically improve the policy. Under continuation, the corresponding means are 49.07%, 50.00%, and 53.75%, with Ordinal Stage-State exceeding both alternatives in all three paired runs. The observed benefit differs across interface representations and training arrangements.

关键词

cs.RO

相关论文

查看 MANIPULATION 分类全部论文