نسخة أولية وصول مفتوح
Evaluating Inference Compute for Generative AI: A Framework for Enterprise Workloads
LLM deployment is shifting from single-turn completion to agentic trajectories in which a model plans, calls tools, reads results and reasons at test time before acting. This inverts the economics of inference hardware: chat serving amortises weight reads across large batches, whereas agent trajectories are sequentiall …