الباحثون

Yu Tian

المنشورات 11

نسخة أولية وصول مفتوح

Beyond Visual Enhancement: Adaptive Multi-Context Steering to Mitigate LVLM Hallucinations

Shuran Ma, JiaLe Li, Yuxin Dong وآخرون · 2026

Hallucination remains a significant challenge in Large Vision-Language Models (LVLMs). Existing training-free methods generally mitigate hallucinations through contrastive decoding or visual enhancement, often increasing the relative influence of visual evidence during generation. This raises a fundamental question: Ca …

نسخة أولية وصول مفتوح

TERRA: Learning Transportable Latent Actions through Temporal Effect Representation and Relational Alignment

Latent actions supervise robot policies with action-like codes inferred from visual transitions, and their usefulness hinges on two questions: what a code keeps from a transition, and whether it still means the same thing when reused in a different initial state. The first is a tension in time: an endpoint difference d …

نسخة أولية وصول مفتوح

Keepsake: Selective Spatial Memory for Long-Horizon Video Generation

Long-horizon camera-controlled video generation relies on persistent memory to maintain scene consistency. Existing systems follow two strategies to achieve this consistency. Full-history approaches retain all generated observations, causing unbounded storage and retrieval costs. Selective-construction approaches reduc …

نسخة أولية وصول مفتوح

FedFit: Federated Fine-Tuning of LLMs via Vector-Bank Parameterization and Quantization

Hang Zou, Chao Zhang, Yuzhi Yang وآخرون · 2026

Federated Learning (FL) enables privacy-preserving fine-tuning of Large Language Models (LLMs), yet the massive communication overhead remains a critical bottleneck. Furthermore, applying Low-Rank Adaptation (LoRA) in FL faces a fundamental "aggregation dilemma" between the accurate Sum-of-Products (SoP) and the commun …

نسخة أولية وصول مفتوح

Mitigating Object Hallucination in Large Vision-Language Models via False Discovery Controlled Visual Data Splitting

Multiple object hallucination, where large vision-language models (LVLMs) generate objects not supported by the visual input, is a persistent challenge caused by visual uncertainty during decoding. Existing methods reduce hallucinations using contrastive signals, but they rely on heuristics and lack principled control …

نسخة أولية وصول مفتوح

PrivMeSA: Privacy-Aware Self-Evolving Multi-Agent System for Medicine via Local-Remote LLM Collaboration

Dannong Wang, Yuran Zhang, Bian Sun وآخرون · 2026

Clinical large language model (LLM) agents deployed locally can consult more capable remote models, but doing so risks exposing patient information. Privacy-conscious delegation places disclosure decisions with a local agent, yet removing explicit identifiers is insufficient: quasi-identifiers can accumulate across mul …

نسخة أولية وصول مفتوح

DeltaWAM: Change-Centric Visual Foresight via Delta Tokens for an Efficient World-Action Model

Tianyun Jiang, Wenrui Bao, Bingxin Xu وآخرون · 2026

World-Action Models (WAMs) offer visual foresight for robotic manipulation, but pixel-space models repeatedly reconstruct entire future scenes, incurring high computational cost and spatio-temporal redundancy. In physical manipulation, consecutive frames often share most of their visual context; the changes between the …

نسخة أولية وصول مفتوح

From Granular Revision Operations to Meaningful Revision Units: Evaluating LLMs for Revision Boundary Detection

Yu Tian, Andrew Potter, Katerina Christhilf وآخرون · 2026

Revision traces provide valuable evidence about students' writing processes, but their usefulness for learning analytics depends on how individual revisions are represented. Automated draft-comparison methods often produce granular edit operations that can fragment a single purposeful revision into multiple analytic un …

نسخة أولية وصول مفتوح

TelecomGPT-R1: Unified Post-Training for Reasoning Across Heterogeneous Telecom Tasks

Bohao Wang, Chenwei Wu, Hang Zou وآخرون · 2026

Large language models (LLMs) offer great potential to automate a broad range of telecom engineering tasks by reasoning over standards, network configurations, mathematical models, source code, and operational logs. However, existing telecom LLMs struggle to reliably reason across these diverse tasks and data types. Gen …

المؤلفون المشاركون