الباحثون

Min-Hung Chen

المنشورات 6

نسخة أولية وصول مفتوح

Grounding What Shapes the Plan: Rethinking Groundedness for Physical Intelligence in Autonomous Driving

Minkyoung Cho, Zewei Zhou, Wenhao Ding وآخرون · 2026

Driving models increasingly ground reasoning in causal relations, spatial structure, perceptual evidence, and predicted futures. These advances make reasoning more faithful to the driving scene, but leave a fundamental question unresolved: what should groundedness mean when the model ultimately outputs an action? Corre …

نسخة أولية وصول مفتوح

VIEScore2: Unified Image Evaluation with Spatially Grounded Explanations

Xianda Du, Max Ku, Weiming Ren وآخرون · 2026

Existing synthetic image evaluators typically provide only a scalar quality score and do not identify the image regions that support it. We introduce VIEScore2, a unified evaluator for image generation and editing tasks with optional conditioning images. VIEScore2 represents an image as an N x N grid and jointly predic …

نسخة أولية وصول مفتوح

VETO: Video Efficient Token Optimization for Vision Language Models

Processing long videos with Vision-Language Models (VLMs) is bottlenecked by the quadratic cost of visual tokens, making long-form inference prohibitively expensive. While single-axis compression methods mitigate this, they hit a hard efficiency floor because they treat spatial and temporal redundancy independently. We …

نسخة أولية وصول مفتوح

World Editing: Intervening on Executable Worlds at Increasing Depth

Max Ku, Nok-Kan Law, Yu-Chien Tang وآخرون · 2026

Interactive world models are increasingly capable of generating environments and acting within them, yet deliberately editing an existing executable world remains underexplored. We formulate world editing as intervening on an existing world while preserving properties that should remain unchanged, and introduce interve …

نسخة أولية وصول مفتوح

ActionLens: Diagnosing Spatial-Temporal Binding Failures in Vision-Language Models

Video-capable vision-language models score above 80\% on popular benchmarks yet struggle with spatial-temporal binding: associating the right action with the right person at the right moment. We introduce ActionLens, a diagnostic benchmark of 6,701 multiple-choice video questions spanning five targeted diagnostics: tra …

نسخة أولية وصول مفتوح

TT-VidT: Decoupling the Temporal Axis for Efficient Motion-Centric Video Pretraining

Shih-Ying Yeh, Daniel Z. Kaplan, Xuehai Wang وآخرون · 2026

Comparisons in video self-supervised learning often evaluate complete training recipes rather than isolating the method itself: architecture, objective, data exposure, schedule, scale, and decoder capacity can all vary at once. This makes it hard to identify which choices yield motion-prioritized representations, whose …

المؤلفون المشاركون