Flow-Based Vision-Language-Action Model Using Robot Velocity Fields
{shotakoba10267, koki.seno, ydaichi1207, komei.sugiura}@keio.jp
ACCV 2026
We will set the links as soon as possible.
At this moment, our code and additional report are provided as supplementary materials.
We focus on language-conditioned flow-based manipulation, where robot flows (robot velocity fields) serve as embodiment-agnostic, motion-centric representations for leveraging cross-embodiment data. This task is crucial because language-conditioned manipulation is essential for practical robotic systems, yet scaling robot foundation models remains bottlenecked by the labor-intensive collection of embodiment-specific data.
Existing methods either coarsely approximate robot flows with sparse keypoint displacements, or cannot handle language-conditioned manipulation. To address this limitation, we propose NarrativeFlow, which models robot flows as continuous velocity fields using a flow-matching formulation conditioned on language.
Moreover, we introduce an auxiliary objective that aligns the model's intermediate representations with task-relevant scene changes. Accordingly, NarrativeFlow captures task-relevant semantics, and generates robot flows that are physically consistent with real-world manipulation.
To validate NarrativeFlow, we have conducted experiments on standard datasets for language-conditioned manipulation. The experimental results show that NarrativeFlow outperforms representative baseline methods on standard evaluation metrics. Furthermore, through real-world experiments, we show that NarrativeFlow achieves higher success rates than baseline methods across multiple manipulation tasks.
We propose NarrativeFlow, a method for language-conditioned robot flow generation inspired by flow-based manipulation policies. The novelties of NarrativeFlow are as follows.
Architecture of NarrativeFlow. The vision-language fusion encoder integrates the embeddings of the language instruction and the initial image into the two subtask tokens: one for robot flow generation, and the other for the Narrative Delta loss. During training, the Narrative Delta loss aligns the corresponding subtask token with the narrative representations of task-relevant scene changes.
Qualitative comparison between NarrativeFlow and a baseline method (Im2Flow2Act). In each subfigure, the first row presents the language instruction, the initial image, and the ground-truth robot flow. The second and third rows show the predicted robot flows of the baseline method and the proposed method, respectively.
Quantitative comparison between the proposed method and baseline methods. The best scores for each metric are shown in bold.
Ablation study for subtask tokens and Narrative Delta loss. ✓ indicates the use of the Narrative Delta loss. The best scores for each metric are shown in bold.
Qualitative results of real-world experiments. The panels show successful executions of three manipulation tasks: (i) mobile close drawer, (ii) mobile bin picking, and (iii) mobile stack cup. In each panel, the top row shows the language instruction, the initial image, and the generated robot flow, while the bottom row shows the downstream manipulation conditioned on the robot flow.
Quantitative results of real-world experiments. For each method, we report success rates across 20 trials per task, along with the average rate over the three tasks. For the oracle method, the policy was conditioned on the corresponding ground-truth robot flows. The best scores among the proposed method and the baseline methods are shown in bold.
Coming soon.