الباحثون

Jeebak Mitra

المنشورات 1

نسخة أولية وصول مفتوح

Characterizing Parallelism Strategies in LLM Inference: Fundamental Compute-Communication Trade-offs

Large Language Model (LLM) inference has become the dominant workload in modern AI systems, requiring serving infrastructures to maximize throughput while meeting strict latency Service-Level Objectives (SLOs). Since state-of-the-art LLMs exceed the compute and memory capacity of a single GPU, inference is commonly dis …

المؤلفون المشاركون