الباحثون

Hao Li

المنشورات 25

نسخة أولية وصول مفتوح

Mine Odyssey: Benchmarking Spatial Agentic Intelligence in the Wild

Yuxuan Cao, Junlong Li, Hao Li وآخرون · 2026

Advances in foundation models are driving efforts to introduce agents to assist people in the physical world. Such agents require agentic spatial intelligence: exploring unfamiliar environments, updating spatial understanding through interaction, and adapting actions based on feedback to sustain progress toward a seque …

نسخة أولية وصول مفتوح

Cognition-Oriented Emotion Tracing from Causes to Consequences in Real-World Social Scenes

Hao Li, Jinye Zhang, Bobo Li وآخرون · 2026

Affective computing has progressed from categorical emotion recognition to open-ended affective analysis with large multimodal models. Yet affective science describes emotion as an unfolding process shaped by appraisal, regulation, and social interpretation, which remains underexplored computationally. We propose TRACE …

نسخة أولية وصول مفتوح

SIFT: Search Intent-to-Filter Transformer for Multi-Task Personalized Filter Ranking at Airbnb

Shashank Dabriwal, Tanya Piplani, Hao Li وآخرون · 2026

Search filters help guests navigate vast catalogs in two-sided marketplaces like Airbnb, and recommending the right filters can meaningfully lift booking conversion. Many such production filter-ranking systems, however, represent the guest through hand-engineered, pre-aggregated features generated by ETL pipelines. Thi …

نسخة أولية وصول مفتوح

Fluorescence-enhanced Whisker Array with Vision-based Deformation Analysis for Underwater Source Localization

Xiaochi Xie, Hao Li, Shixuan Zhao وآخرون · 2026

Deep-water biological observation is essential for understanding marine organisms and their interactions with the environment. However, conventional optical and acoustic approaches can introduce stimuli that alter animal behavior and bias biological observations. This paper proposes a fluorescence-enhanced whisker arra …

نسخة أولية وصول مفتوح

Lego-Like Stiffness Configuration of Planar Compliant Modules for Task-Specific Flexible Interfaces

Siyue Yao, Xiaochi Xie, Shixuan Zhao وآخرون · 2026

Compliant mechanisms provide compact and intrinsic structural compliance for regulating physical interactions between mechanisms and environments. However, different tasks demand distinct stiffness characteristics, often requiring task-specific optimization and redesign due to limited geometric design space and inheren …

نسخة أولية وصول مفتوح

SCOUT: Supply-Aware Cold-Start Proactive Query Suggestion for Travel Search

Hao Li, Shashank Reddy, Kedar Bellare وآخرون · 2026

Generative query suggestion, powered by Large Language Models (LLMs), has become increasingly popular in search and conversational systems to reduce user friction and guide intent formulation. Existing approaches align suggestions with user preferences (e.g., clicks or conversions). This works for open-ended applicatio …

نسخة أولية وصول مفتوح

PCLM: Small-target localization with frozen CLIP via prototype contrast and local magnification

Zhipeng Ye, Feng Jiang, Qiufeng Wang وآخرون · 2026

Small targets occupy few patches in a vision-language encoder, so spatial features often mix object appearance with surrounding content. We propose Prototype Contrast and Local Magnification (PCLM), a support-conditioned localization method that uses a frozen CLIP encoder. Five masked support images per class define fo …

نسخة أولية وصول مفتوح

Self-Reflection Fine-Tuning: Enhancing Agent Security against Prompt Injection Attacks from Failure Experience

Zixuan Wang, Hao Li, Fengyu Gao وآخرون · 2026

Large language model (LLM) agents are increasingly deployed in tool-augmented environments, but their reliance on external inputs makes them highly vulnerable to prompt injection attacks that can hijack task objectives. Existing safety alignment methods rely on static expert trajectories or preference optimization, lim …

نسخة أولية وصول مفتوح

LOCI: Spatial Linear Memory for Streaming World Models

Ji Xia, Tingting Liao, Xuezhi Liang وآخرون · 2026

When a camera revisits a previously observed region, a video world model should reproduce what was there before. This requires both remembering past observations and retrieving the right one for the current viewpoint. Key-value caches preserve visual detail but grow with video length; recurrent memory is compact but co …

نسخة أولية وصول مفتوح

Memorizon: Training World Models Beyond Their Context Window

Tingting Liao, Xuezhi Liang, Hao Li وآخرون · 2026

Streaming world models should render a place consistently across repeated visits. Directly supervising such revisits requires training samples that capture both visits, often spanning minutes. Yet dense attention over the full span incurs quadratic costs, making long-span supervision expensive. Memorizon breaks this co …

نسخة أولية وصول مفتوح

When Clipping Reverses Correction: Failure Dynamics of Pointwise Forward-KL On-Policy Self-Distillation

Di Huang, Hao Li, Yixin Chen وآخرون · 2026

On-policy self-distillation (OPSD) trains a student on its own generated responses using feedback from the same model conditioned on privileged information. On mathematical reasoning, the original OPSD study finds that stylistic tokens can dominate the training signal over math-related tokens, and that pointwise clippi …

نسخة أولية وصول مفتوح

False Frontiers: Diagnosing and Mitigating Co-Cheating in Self-Evolving Search Agents

Meijia Chen, Hao Li, Zheng Lu وآخرون · 2026

Self-evolving search agents build their own training curricula by jointly optimizing a proposer that generates questions and a solver that answers them. This closed loop introduces a failure mode we call co-cheating: the proposer and solver increasingly agree on shared errors, so internal reward improves without a matc …

نسخة أولية وصول مفتوح

GroundAnything: Reconciling Parallel Decoding with Precise Visual Grounding at Flash Speed

Qize Yu, Lianrui Fan, Bowen Ping وآخرون · 2026

Autoregressive (AR) grounding models serialize spatial predictions, introducing sequential latency and imposing a causal order on output tokens. We view grounding as visual evidence extraction: objects, locations, and spatial relations are jointly constrained by the image and query, yet their dependencies do not imply …

نسخة أولية وصول مفتوح

SEPAL: Separated Expert Pairs with Answer-Level Fusion for Reliable LLM Collaboration

Weijie Ren, Yanwen Zhang, Hao Li وآخرون · 2026

Multi-agent collaboration lets large language models (LLMs) improve question answering through deliberation and feedback. Yet shared discussion couples correction with exposure to the same mistakes, which can erode the diversity needed for voting. Self-consistency offers sampling diversity without feedback, while singl …

نسخة أولية وصول مفتوح

OPSRD: On-Policy Self-Role Distillation

Weijie Ren, Yanwen Zhang, Hao Li وآخرون · 2026

Role prompting elicits specialized behavior from large language models through an expert identity, offering a lightweight way to guide reasoning on demanding tasks. However, evaluating or distilling complete role-prompted answers can miss useful next-token preferences when the sampled solution remains incorrect. Transf …

نسخة أولية وصول مفتوح

The Teacher Is a Direction, Not a Destination: Extrapolating RL-Induced Representation Residuals in On-Policy Distillation

Hao Li, Meijia Chen, Weijie Ren وآخرون · 2026

On-policy distillation (OPD) trains a student to match the teacher's next-token distributions on the student's own trajectories and has yielded substantial empirical gains. Generalized variants allow the student to surpass the teacher by extrapolating an implicit reward in output space. The language-model head, however …

المؤلفون المشاركون