الباحثون

Dongyub Jude Lee

المنشورات 3

نسخة أولية وصول مفتوح

Refusal Localizes, the Damage Relocates: Safety Layers Under Few-Sample Fine-Tuning

Jungseob Lee, Dongyub Jude Lee, Sugyeong Eo وآخرون · 2026

Fine-tuning adapts aligned large language models (LLMs) to downstream tasks, but a few dozen harmful examples can remove their refusal of harmful requests. Prior work localizes safety-related behavior to specific layers, directions, and tokens, suggesting targets for protection. We test whether successful localization …

نسخة أولية وصول مفتوح

Faster Block-Diffusion Serving with Distribution-Free Risk Guarantees

Jungseob Lee, Dongyub Jude Lee, Chanjun Park وآخرون · 2026

Block-diffusion language models are served at hand-picked operating points, such as acceptance thresholds, buffer depth, schedule, checkpoint and precision, and each point is chosen by its mean benchmark accuracy. However, a mean does not tell an operator how often a faster configuration fails on prompts that the slowe …

المؤلفون المشاركون