الباحثون

Junfeng Fang

المنشورات 8

نسخة أولية وصول مفتوح

ReCal: Calibrating Structured Pruning for On-Policy Distillation Recovery

Houcheng Jiang, Mao Zheng, Mingyang Song وآخرون · 2026

Structured pruning reduces the deployment cost of reasoning language models, but the resulting capability degradation can hinder subsequent on-policy distillation (OPD) recovery. Because OPD relies on student-generated trajectories, pruning damage that persists after offline distillation can limit its effectiveness. We …

نسخة أولية وصول مفتوح

DNAlign: Dynamic Null-Space Safe Alignment for LLMs

Jisheng Dang, Yushuo Zhao, Dewei Liu وآخرون · 2026

Ensuring the safe and reliable deployment of large language models (LLMs) remains a fundamental challenge. Existing safety alignment approaches either incur high computational cost or unintentionally disrupt the model's core knowledge, leading to degraded fluency and factual accuracy on benign tasks. This reveals a per …

نسخة أولية وصول مفتوح

Super-Resolving Unseen Hyperspectral Sensors at Any Scale via Spatial Operators

Ji-Xuan He, Guohang Zhuang, Bo Junge وآخرون · 2026

Achieving cross-sensor generalization and arbitrary-scale reconstruction with a single model remains challenging in hyperspectral super-resolution (HSR). Although recent methods support arbitrary-scale reconstruction, applying them to new sensors or scales beyond the training range often requires additional data and co …

نسخة أولية وصول مفتوح

USA: Update-aware SAM for Cross-domain On-Policy Disitllation of Language Agents

Qiyong Zhong, Mao Zheng, Mingyang Song وآخرون · 2026

On-policy distillation instils multi-turn agentic reasoning through dense token-level supervision on the student's own trajectories, but a single domain saturates early, so further supervision has to be drawn from other domains. Multi-domain data mixing is the most direct way of incorporating them, at the cost of confl …

نسخة أولية وصول مفتوح

MAS-OPD: On-Policy Distillation for Multi-agent Systems

Qiyong Zhong, Mao Zheng, Mingyang Song وآخرون · 2026

Multi-agent systems (MAS) split a task across specialized roles and are promising on complex tasks, yet a prevailing approach relies on inference-time orchestration alone. General-purpose APIs are costly and hard to customize, while small models with role prompts rarely develop stable role competence or reliable collab …

نسخة أولية وصول مفتوح

Learning to Steer, Steering to See: Unveiling the Geometry of RLVR in Large Language Models via Trainable Vectors

Yuchen Cai, Ding Cao, Qixiang Yin وآخرون · 2026

Reinforcement learning (RL) has become a key paradigm for enhancing the reasoning of large language models, yet the high dimensionality of parameter updates makes its training dynamics hard to analyze. We study reinforcement learning with verifiable rewards (RLVR) and use vector steering to identify a low-dimensional e …

المؤلفون المشاركون