الباحثون

Zhaochun Ren

المنشورات 2

نسخة أولية وصول مفتوح

SearchJev: A Fast and Calibrated System-1 Model for Search Agents

Search agents repeatedly make short decisions about relevance, evidence sufficiency, and search actions. Using generative language models for these decisions introduces latency and unreliable confidence. We present SearchJev, a fast and calibrated System-1 model that separates search decisions from System-2 reasoning a …

نسخة أولية وصول مفتوح

Beyond Refusal Patterns: Safe-Role Internalization for Robust and Generalizable LLM Safety Alignment

Jinghao Pang, Jitai Hao, Qiang Huang وآخرون · 2026

Large Language Models (LLMs) have achieved remarkable capabilities but remain vulnerable to jailbreak attacks that elicit harmful or unsafe outputs. Existing safety alignment approaches, including Supervised Fine-Tuning (SFT) and Reinforcement Learning from Human Feedback (RLHF), often require substantial attack-specif …

المؤلفون المشاركون