الملخص

Assistance in collaborative manipulation is often initiated by user instructions, making high-level reasoning request-driven. In fluent human teamwork, however, partners often infer the next helpful step from the observed outcome of an action rather than waiting for instructions. Motivated by this, we investigate an event-driven formulation of proactive assistance, where human--object interaction outcomes initiate assistive reasoning without user-provided task specifications at inference time. To this end, we propose an event-driven framework that monitors workspace state changes with an event monitor and, upon event completion, extracts stabilized pre/post snapshots that characterize the resulting state transition. A frozen pretrained Vision-Language Model (VLM) then uses its semantic priors to infer the task context, decide whether assistance is appropriate, and, when needed, generate a sequence of assistive actions from the observed transition. To make outputs executable and verifiable, we restrict actions to a set of action primitives and reference objects via integer IDs.We evaluate the same framework across three distinct real world tabletop collaboration tasks without task-specific training or fine-tuning. The event-driven framework achieves performance comparable to variants given user instructions.

الكلمات المفتاحية

الموضوع

بيانات النشر

المجلة
غير متاح
وصول مفتوح
وصول مفتوح أخضر

اقتبس هذه المقالة

APA 7

Liu, F., Su, H., Chi, H., Geng, R., Ren, C., Liu, X., Xu, C., Ohsita, Y., & Zhang, L. (2026). Event-Driven Proactive Robot Assistance through Vision-Language Reasoning. https://omanscience.com/ar/articles/event-driven-proactive-robot-assistance-through-vision-language-reasoning

MLA 9

Liu, Fengkai, et al. "Event-Driven Proactive Robot Assistance through Vision-Language Reasoning." https://omanscience.com/ar/articles/event-driven-proactive-robot-assistance-through-vision-language-reasoning.

شيكاغو (المؤلف–التاريخ)

Liu, Fengkai, Hao Su, Haozhuang Chi, Rui Geng, Congzhi Ren, Xuqing Liu, Chenfei Xu, Yuichi Ohsita, and Liyun Zhang. 2026. "Event-Driven Proactive Robot Assistance through Vision-Language Reasoning." https://omanscience.com/ar/articles/event-driven-proactive-robot-assistance-through-vision-language-reasoning.

هارفارد

Liu, F., Su, H., Chi, H., Geng, R., Ren, C., Liu, X., Xu, C., Ohsita, Y. and Zhang, L. (2026) 'Event-Driven Proactive Robot Assistance through Vision-Language Reasoning', Available at: https://omanscience.com/ar/articles/event-driven-proactive-robot-assistance-through-vision-language-reasoning.

فانكوفر

Liu F, Su H, Chi H, Geng R, Ren C, Liu X, et al. Event-Driven Proactive Robot Assistance through Vision-Language Reasoning. https://omanscience.com/ar/articles/event-driven-proactive-robot-assistance-through-vision-language-reasoning

IEEE

F. Liu, H. Su, H. Chi, R. Geng, C. Ren, X. Liu, C. Xu, Y. Ohsita, and L. Zhang, "Event-Driven Proactive Robot Assistance through Vision-Language Reasoning," https://omanscience.com/ar/articles/event-driven-proactive-robot-assistance-through-vision-language-reasoning.