Authors

Youngjin Cho

Publications 1

Preprint Open access

Characterizing High Bandwidth Flash for LLM Serving

Zack Yu, Chloe Wong, Coleman Hooper et al. · 2026

Large language model (LLM) serving requires substantial memory to store model weights and KV caches. As models grow larger and contexts become longer, memory capacity and bandwidth increasingly become bottlenecks for serving performance. Agentic workloads compound this pressure through repeated interactions over growin …

Co-authors