الملخص
We focus on language-conditioned flow-based manipulation, where robot flows (robot velocity fields) serve as embodiment-agnostic, motion-centric representations for leveraging data collected from multiple robot platforms. This task is crucial because language-conditioned manipulation is essential for practical robotic systems, yet scaling robot foundation models remains limited by the labor-intensive collection of embodiment-specific data. Existing methods either coarsely approximate robot flows with sparse keypoint displacements, or cannot handle language-conditioned manipulation. To address this limitation, we propose NarrativeFlow, which models robot flows as continuous velocity fields using a flow-matching formulation conditioned on language. Accordingly, NarrativeFlow generates robot flows that are physically consistent with real-world manipulation. To validate NarrativeFlow, we have conducted experiments on standard datasets for language-conditioned manipulation. The experimental results show that NarrativeFlow outperforms representative baseline methods on standard evaluation metrics. Furthermore, through real-world experiments, we show that NarrativeFlow achieves higher success rates than baseline methods across multiple manipulation tasks. The project page is available at https://shota0520.github.io/NarrativeFlow-project-page/
الكلمات المفتاحية
الموضوع
بيانات النشر
- المجلة
- غير متاح
- وصول مفتوح
- وصول مفتوح أخضر
اقتبس هذه المقالة
APA 7
Kobayashi, S., Seno, K., Yashima, D., & Sugiura, K. (2026). NarrativeFlow: Flow-Based Vision-Language-Action Model Using Robot Velocity Fields. https://omanscience.com/ar/articles/narrativeflow-flow-based-vision-language-action-model-using-robot-velocity-fields
MLA 9
Kobayashi, Shota, et al. "NarrativeFlow: Flow-Based Vision-Language-Action Model Using Robot Velocity Fields." https://omanscience.com/ar/articles/narrativeflow-flow-based-vision-language-action-model-using-robot-velocity-fields.
شيكاغو (المؤلف–التاريخ)
Kobayashi, Shota, Koki Seno, Daichi Yashima, and Komei Sugiura. 2026. "NarrativeFlow: Flow-Based Vision-Language-Action Model Using Robot Velocity Fields." https://omanscience.com/ar/articles/narrativeflow-flow-based-vision-language-action-model-using-robot-velocity-fields.
هارفارد
Kobayashi, S., Seno, K., Yashima, D. and Sugiura, K. (2026) 'NarrativeFlow: Flow-Based Vision-Language-Action Model Using Robot Velocity Fields', Available at: https://omanscience.com/ar/articles/narrativeflow-flow-based-vision-language-action-model-using-robot-velocity-fields.
فانكوفر
Kobayashi S, Seno K, Yashima D, Sugiura K. NarrativeFlow: Flow-Based Vision-Language-Action Model Using Robot Velocity Fields. https://omanscience.com/ar/articles/narrativeflow-flow-based-vision-language-action-model-using-robot-velocity-fields
IEEE
S. Kobayashi, K. Seno, D. Yashima, and K. Sugiura, "NarrativeFlow: Flow-Based Vision-Language-Action Model Using Robot Velocity Fields," https://omanscience.com/ar/articles/narrativeflow-flow-based-vision-language-action-model-using-robot-velocity-fields.