الباحثون

Ranjay Krishna

المنشورات 3

نسخة أولية وصول مفتوح

Why MLLMs Struggle to Count: Overcoming Individuation and Aggregation Bottlenecks with ConvStack

Liwei Che, Yihao Quan, Sen Fang وآخرون · 2026

Multimodal Large Language Models (MLLMs) consistently struggle with fine-grained visual counting, yet the underlying causes remain poorly understood. In this work, we present a mechanistic analysis of this failure mode, identifying two critical bottlenecks inherent to the global attention pipeline of MLLMs. First, we r …

نسخة أولية وصول مفتوح

OmniTaskonomy: When Does Visual Generation Improve Visual Understanding?

Jiaxin Ge, Yiming Qin, Ji Xie وآخرون · 2026

Training a model to generate visual content can encourage it to learn rich perceptual capabilities related to geometry, spatial relationships, and objectness; yet, its benefits for visual understanding remain unclear. We ask: when and how does visual generation supervision improve visual understanding? We study control …

المؤلفون المشاركون