الباحثون

Xinnan Zhu

المنشورات 1

نسخة أولية وصول مفتوح

Exemplar2VQA: A Scalable Exemplar-Driven Visual Question Answering Generation Framework via Multi-Agent Coding

Jiayu Ying, Qijian Tian, Ruijie Xu وآخرون · 2026

Advancing spatial intelligence in Multimodal Large Language Models (MLLMs) is bottlenecked by the scarcity of complex, scalable 3D question-answer (QA) data. While manual annotation is labor-intensive, directly utilizing LLMs to synthesize these QA pairs often fails due to their inherent deficiencies in spatial and geo …

المؤلفون المشاركون