الباحثون

Yifan Wang

المنشورات 21

نسخة أولية وصول مفتوح

Autoregressive Retriever: Improving Query Understanding from Item Feedback for Universal Multimodal Retrieval

Jianfei Zhao, Yifan Wang, Feng Zhang وآخرون · 2026

Universal multimodal retrieval typically encodes a query once and ranks independently indexed items by embedding similarity. This design supports efficient search, but leaves the query representation unchanged even when retrieved items could help clarify the information need. We introduce the AutoRegressive Retriever ( …

نسخة أولية وصول مفتوح

Beyond Spatio-Temporal Priors: A Generalizable Approach for Dense Correspondence Matching

Luping Liu, Bingyi Kang, Yifan Wang وآخرون · 2026

Dense correspondence matching has historically been bounded by simplifying spatio-temporal priors, such as smooth motion and rigid geometry. While effective for classical tasks, these assumptions break down in image editing and reference-guided generation (IEG), where transformations can preserve visual identity while …

نسخة أولية وصول مفتوح

StoreBench: A Live-Commerce Environment for Evaluating and Training Autonomous Operator Agents

Reinforcement learning environments are now a primary lever for improving large language model (LLM) capabilities in post-training, yet most agentic benchmarks remain static: the world moves only when the agent acts, the reward is a terminal verdict, and the pass bar is set arbitrarily. We introduce StoreBench, a live- …

نسخة أولية وصول مفتوح

Arm-wise Compositional Generalization in Dual-Arm Vision-Language-Action Models

Zaibin Zhang, Binghao Ran, Yuhan Wu وآخرون · 2026

Generalization in multi-arm collaboration can be studied as composing familiar atomic skills in new ways across arms. However, existing evaluations offer limited insight into which training and architectural choices support this ability under different coordination requirements. We introduce \textbf{ACG-Bench}, a bench …

نسخة أولية وصول مفتوح

AgentPersonaBench: Benchmarking Persona-Driven User Simulation

Jintao Huang, Yifan Wang, Hongyu Shen وآخرون · 2026

We introduce AgentPersonaBench (APB), a benchmark evaluating whether persona conditioning faithfully steers downstream agent behavior. While language models are increasingly deployed for persona-driven user simulation, existing benchmarks primarily evaluate conversational styling or self-reports rather than authentic b …

نسخة أولية وصول مفتوح

MintEval: Do LLMs Implement the Trading Strategy You Asked For? A Behavioural-Equivalence Benchmark for Natural-Language-to-Strategy Code

Large language models are moving from producing trading signals to writing the code that executes them. The failure mode of the second role is silent: generated code runs, a backtest plots, yet the risk logic that the trader described is not the logic being executed. Existing code benchmarks test functional correctness …

نسخة أولية وصول مفتوح

Token-Level Video Reinforcement Learning

Yifan Wang, Gordon Guocheng Qian, Yanyu Li وآخرون · 2026

Reinforcement learning (RL) for video generation usually assigns one scalar reward to an entire sampled video. Yet a video is not uniformly flawed: some visual tokens may already satisfy the prompt, whereas others require correction. A scalar reward cannot localize errors, causing optimization to perturb satisfactory t …

نسخة أولية وصول مفتوح

NeurDuo-EEG: A Long-Sequence EEG Foundation Model with Persistent State and Explicit Memory

Yifan Wang, Haiping Liu, Yang Cui وآخرون · 2026

Electroencephalography (EEG) is recorded continuously over hours, with relevant dynamics spanning timescales from milliseconds to hours. Most EEG foundation models nevertheless process fixed windows independently, limiting their ability to capture information encoded in long-timescale dynamics. State-space architecture …

نسخة أولية وصول مفتوح

Targeting Pivotal Decisions for Credit Assignment in Agentic Reinforcement Learning

Group Relative Policy Optimization (GRPO) has become a promising approach for training large language model agents. However, its uniform assignment of trajectory-level advantages to all policy tokens fails to distinguish consequential decisions from less relevant ones, obscuring which intermediate decisions contributed …

نسخة أولية وصول مفتوح

MORSE: Multi-Context Ordering via Reverse Scoring for Evidence-Preserving Compression

Ke Wan, Yifan Wang, Liheng Lai وآخرون · 2026

Retrieval-augmented generation often relies on multiple retrieved contexts that contain substantial redundancy, motivating context compression to preserve useful information under limited input budgets. Likelihood-based compressors can account for cross-context redundancy through sequential scoring, but this makes evid …

نسخة أولية وصول مفتوح

CARE: Experience-Guided Atomic Corrective Execution for Vision-Language-Action Policies

Junlan Xiao, Junwei Jiang, Zaibin Zhang وآخرون · 2026

Vision-Language-Action (VLA) policies achieve strong performance in robotic manipulation but remain brittle once execution deviates from nominal trajectories. We propose CARE (Corrective Atomic Robotic Execution), a framework that improves recovery by learning from failures encountered during execution. Instead of gene …

نسخة أولية وصول مفتوح

CIBuzzBench: A Benchmark for Cross-Lingual Understanding of Chinese Internet Buzzwords

Yifan Wang, Junyu Lu, Qifan Wang وآخرون · 2026

Chinese social media has generated a vast and continually evolving lexicon of internet buzzwords whose meanings are often non-literal and deeply rooted in local cultural and pragmatic contexts. Existing research has primarily focused on interpreting these buzzwords within Chinese, leaving largely unexplored whether LLM …

نسخة أولية وصول مفتوح

VABench: Measuring Embodied Spatial Intelligence through Visual Demonstrations, Active Perception, and Metric Control

Zhongbo Zhang, Jiayi Jin, Yifan Wang وآخرون · 2026

Spatial intelligence requires more than describing object locations. Under incomplete observation, models must identify and acquire missing evidence, interpret it in a common spatial frame, and act on it. We introduce VA-Bench to evaluate the complete observe-reason-act-revise loop. General-purpose MLLMs learn procedur …

نسخة أولية وصول مفتوح

Learning Foresight without Explicit Trajectories for 3D Diffusion Policies

Zhongbo Zhang, Zaibin Zhang, Yifan Wang وآخرون · 2026

3D diffusion policies are strong at generating geometrically grounded actions from current observations, but successful manipulation requires not only knowing what motion is feasible now, but also anticipating where the interaction is heading. Existing policies largely leave such foresight to emerge implicitly from act …

نسخة أولية وصول مفتوح

TeleAntiFraud 2.0: A Refreshable, Profile-Grounded, and Audio-Based Benchmark for Telecom Fraud Detection

Huiyuan Liu, Zhiming Ma, Yanxing Liu وآخرون · 2026

Telecom fraud scripts evolve rapidly and are often designed to resemble routine service conversations, creating two key requirements for audio-based telecom-fraud evaluation. First, benchmarks must incorporate newly observed scam patterns without overwriting previously established test sets. Second, they must distingui …

نسخة أولية وصول مفتوح

FRAUDSkill: Structured Frozen-Weight Skill Optimization for Audio Anti-Fraud Detection

Chengxian Hu, Zhiming Ma, Mingjun Pan وآخرون · 2026

Large audio-language models have shown promise for anti-fraud detection by directly processing speech and reasoning over fraud-related evidence. Their deployment, however, requires predictions to follow a predefined label space and a structured decision protocol consisting of service-scenario identification, fraud dete …

المؤلفون المشاركون