الباحثون

Farhan Tejani

المنشورات 1

نسخة أولية وصول مفتوح

Rethinking Latent Visual Reasoning: Grounding Latent Reasoning in Visual Evidence

Xi Xiao, Tianchen Zhao, Youngeun Kim وآخرون · 2026

Latent visual reasoning (LVR) enables multimodal large language models (MLLMs) to perform intermediate computation in continuous latent tokens rather than expressing every reasoning step in words. However, unlike textual CoT, latent reasoning is not directly observable, making it difficult to supervise what latent toke …

المؤلفون المشاركون