Authors

Youmin Chen

Publications 1

Preprint Open access

From Overloaded to Guaranteed: High-Throughput Multi-SLO Enforcement for LoRA-Assisted On-Premise LLM Deployment

Zeshen Zhang, Han Zhao, Weihao Cui et al. · 2026

As Large Language Models (LLMs) become essential in privacy-sensitive sectors like hospitals and government agencies, the on-premise LLM servers offer a cost-effective and secure alternative to public cloud services. However, these resource-constrained servers struggle to guarantee heterogeneous Service Level Objective …

Co-authors