الباحثون

Kan Li

المنشورات 2

نسخة أولية وصول مفتوح

From Pareto to Preference: Personalized Test-Time Scaling via Amortized Agentic Policy Discovery

Xinglin Wang, Zishen Liu, Tong Zheng وآخرون · 2026

Test-time scaling (TTS) improves the reasoning capabilities of large language models by allocating additional inference computation. Existing approaches to improving TTS efficiency largely optimize accuracy against one resource dimension at a time, advancing either the accuracy--cost or accuracy--latency Pareto frontie …

نسخة أولية وصول مفتوح

Calibrate the Decisions That Change the Future: On-Policy Post-Training Quantization for Multimodal Large Language Models

Wenxiao Fan, Jingling Fu, Lichen Ma وآخرون · 2026

Post-training quantization (PTQ) lowers deployment cost for multimodal large language models, but calibration typically reconstructs fixed sequences with local objectives. This overlooks autoregressive feedback: a quantization-induced token change redirects the prefix and changes future states. Yet on-policy coverage a …

المؤلفون المشاركون