Authors

Hao Li

Publications 25

Preprint Open access

Fluorescence-enhanced Whisker Array with Vision-based Deformation Analysis for Underwater Source Localization

Xiaochi Xie, Hao Li, Shixuan Zhao et al. · 2026

Deep-water biological observation is essential for understanding marine organisms and their interactions with the environment. However, conventional optical and acoustic approaches can introduce stimuli that alter animal behavior and bias biological observations. This paper proposes a fluorescence-enhanced whisker arra …

Preprint Open access

Self-Reflection Fine-Tuning: Enhancing Agent Security against Prompt Injection Attacks from Failure Experience

Zixuan Wang, Hao Li, Fengyu Gao et al. · 2026

Large language model (LLM) agents are increasingly deployed in tool-augmented environments, but their reliance on external inputs makes them highly vulnerable to prompt injection attacks that can hijack task objectives. Existing safety alignment methods rely on static expert trajectories or preference optimization, lim …

Preprint Open access

LOCI: Spatial Linear Memory for Streaming World Models

Ji Xia, Tingting Liao, Xuezhi Liang et al. · 2026

When a camera revisits a previously observed region, a video world model should reproduce what was there before. This requires both remembering past observations and retrieving the right one for the current viewpoint. Key-value caches preserve visual detail but grow with video length; recurrent memory is compact but co …

Preprint Open access

Memorizon: Training World Models Beyond Their Context Window

Tingting Liao, Xuezhi Liang, Hao Li et al. · 2026

Streaming world models should render a place consistently across repeated visits. Directly supervising such revisits requires training samples that capture both visits, often spanning minutes. Yet dense attention over the full span incurs quadratic costs, making long-span supervision expensive. Memorizon breaks this co …

Preprint Open access

OPSRD: On-Policy Self-Role Distillation

Weijie Ren, Yanwen Zhang, Hao Li et al. · 2026

Role prompting elicits specialized behavior from large language models through an expert identity, offering a lightweight way to guide reasoning on demanding tasks. However, evaluating or distilling complete role-prompted answers can miss useful next-token preferences when the sampled solution remains incorrect. Transf …

Preprint Open access

The Teacher Is a Direction, Not a Destination: Extrapolating RL-Induced Representation Residuals in On-Policy Distillation

Hao Li, Meijia Chen, Weijie Ren et al. · 2026

On-policy distillation (OPD) trains a student to match the teacher's next-token distributions on the student's own trajectories and has yielded substantial empirical gains. Generalized variants allow the student to surpass the teacher by extrapolating an implicit reward in output space. The language-model head, however …

Co-authors