الباحثون

Jun Yu

المنشورات 5

نسخة أولية وصول مفتوح

Lifelong small-object navigation in changing object layouts: a benchmark and method

Jiagan Huang, Zikun Zhou, Zijian Ni وآخرون · 2026

Household robots need to continually navigate to different objects in the same environment, many of which are small and portable, such as tools and toys. Their small visual footprint and frequent occlusion make reliable observation difficult, and they may be moved by people without the robot observing the changes. We f …

نسخة أولية وصول مفتوح

Can Agent Harnesses and Inference Engines Hear Each Other? The HEAR Protocol for Agentic LLM Serving

Jiaqi Zhao, Haodong Chen, Jitai Hao وآخرون · 2026

LLM agents increasingly execute complex workflows involving multi-turn reasoning, tool use, and parallel agents. Efficient serving requires decisions that span two layers with complementary information: the agent harness understands workflow dependencies, context lifecycles, and execution objectives, whereas the infere …

نسخة أولية وصول مفتوح

Beyond Refusal Patterns: Safe-Role Internalization for Robust and Generalizable LLM Safety Alignment

Jinghao Pang, Jitai Hao, Qiang Huang وآخرون · 2026

Large Language Models (LLMs) have achieved remarkable capabilities but remain vulnerable to jailbreak attacks that elicit harmful or unsafe outputs. Existing safety alignment approaches, including Supervised Fine-Tuning (SFT) and Reinforcement Learning from Human Feedback (RLHF), often require substantial attack-specif …

نسخة أولية وصول مفتوح

SparseEngine: Sparse-First Inference Engine

Jitai Hao, Quansheng Gu, Qiang Huang وآخرون · 2026

Long-context LLM agents accumulate interaction histories that strain KV-cache memory and attention computation. Although sparse attention reduces these costs, heterogeneous cache representations and workflows hinder integration with existing inference engines, while prior sparse-serving abstractions support only specif …

نسخة أولية وصول مفتوح

MoRA: MoE Pruning via Router Bias Learning and Expert Approximation

Yushuai Sun, Zikun Zhou, Lin Gao وآخرون · 2026

Mixture-of-Experts (MoE) models enable parameter scaling with limited per-token computation by activating only a small subset of experts for each token, but deploying them still requires loading the complete expert pool into memory. Structured expert pruning can effectively reduce the memory usage by removing experts. …

المؤلفون المشاركون