الباحثون

Jingyan Jiang

المنشورات 5

نسخة أولية وصول مفتوح

RealtimeWAM: How Fast Can I Run My World Action Model?

Huanan Liu, Ye Li, Kangye Ji وآخرون · 2026

World Action Models (WAMs) combine visual dynamics modeling with action generation, but their high inference latency limits responsive robot control. Recent efforts accelerate inference by removing explicit future-video generation at test time, as in FastWAM, an approach that requires a specially tailored architectural …

نسخة أولية وصول مفتوح

ChunkTrust: Adapting Execution Horizons for Robot Policies with Action-Expert Evidence

Fanding Huang, Jingyan Jiang, Shifeng Bao وآخرون · 2026

Robot foundation policies predict action chunks, but how many actions to execute before replanning depends on the current task phase. We introduce ChunkTrust, which treats the execution horizon as a latent variable inferred from action-expert evidence rather than a fixed hyperparameter. Its training-free Action-aware H …

نسخة أولية وصول مفتوح

CurvSpec: Adaptive Multi-Curvature Learning for Partial Relevant Video Retrieval

Zhen Liu, Letian Li, Jinpeng Wang وآخرون · 2026 · 10.1145/3767308.3835167

Partially Relevant Video Retrieval (PRVR) seeks to retrieve untrim-med videos containing a moment that matches a text query, without temporal annotations. The relevant moment may last only seconds within a video spanning several minutes, creating an extremely low signal-to-noise ratio that makes PRVR more challenging t …

نسخة أولية وصول مفتوح

Seek Before You Move: Evidence Seeking for Progress Grounding in Vision-Language Navigation

Zhimin Wang, Meiyuan Zhu, Duo Wu وآخرون · 2026

Vision-Language Navigation (VLN) requires agents to continuously ground task progress from long-horizon instructions and partial egocentric observations. Existing VLM-based navigation agents typically reason only over available observations and may remain confident even when task-relevant evidence is missing. For examp …

نسخة أولية وصول مفتوح

FutureDuet: Decoupling Observation Access from Future Supervision in World Action Models

Jie Wu, Yuzhi Huang, Junqi Liu وآخرون · 2026

World Action Models (WAMs) augment robot action generation with future visual supervision. Existing WAMs commonly fuse main and wrist observations into one visual stream and train both with the same future-video objective, despite their different visual dynamics. A stable main camera reveals scene-level task evolution, …

المؤلفون المشاركون