Abstract

LLM deployment is shifting from single-turn completion to agentic trajectories in which a model plans, calls tools, reads results and reasons at test time before acting. This inverts the economics of inference hardware: chat serving amortises weight reads across large batches, whereas agent trajectories are sequentially dependent, run at effective batch one, and make per-token decode latency (TPOT) the dominant term in task completion time. Using a roofline analysis and a closed-form episode-latency model, we show why this regime favours accelerators that keep weights in on-die SRAM (Cerebras WSE-3/3T, Groq/NVIDIA LPU) or compiler-managed tiered memory (SambaNova SN40L/SN50), and why three vendor ecosystems converged in 2026 on disaggregated prefill/decode serving. We show that per-step reliability compounds exponentially in trajectory length-a 2% per-step failure rate erases a 2x decode advantage for a 20-step agent-so determinism and tail latency are first-order performance variables. We then propose a four-layer evaluation framework (silicon, serving system, agent episode, enterprise) with a metric set built on goodput at an agentic SLO and cost per successful episode, a six-axis benchmark protocol over six task families, a paired-bootstrap statistical design, an attestation protocol for vendor-run benchmarks, and TCO, availability and adoption-timing models with explicit break-even conditions. All performance figures are public and labelled by evidence class; we state seven falsifiable hypotheses and the experiments that test them, and argue that the most likely original result is that token-throughput rankings diverge from cost-per-successful-task rankings on long-horizon work.

Keywords

Subject

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Ali, A. R., Siddiqui, M. A., & Zahid, M. (2026). Evaluating Inference Compute for Generative AI: A Framework for Enterprise Workloads. https://omanscience.com/en/articles/evaluating-inference-compute-for-generative-ai-a-framework-for-enterprise-workloads

MLA 9

Ali, Abbas Raza, et al. "Evaluating Inference Compute for Generative AI: A Framework for Enterprise Workloads." https://omanscience.com/en/articles/evaluating-inference-compute-for-generative-ai-a-framework-for-enterprise-workloads.

Chicago (author–date)

Ali, Abbas Raza, Muhammad Ajmal Siddiqui, and Moona Zahid. 2026. "Evaluating Inference Compute for Generative AI: A Framework for Enterprise Workloads." https://omanscience.com/en/articles/evaluating-inference-compute-for-generative-ai-a-framework-for-enterprise-workloads.

Harvard

Ali, A. R., Siddiqui, M. A. and Zahid, M. (2026) 'Evaluating Inference Compute for Generative AI: A Framework for Enterprise Workloads', Available at: https://omanscience.com/en/articles/evaluating-inference-compute-for-generative-ai-a-framework-for-enterprise-workloads.

Vancouver

Ali AR, Siddiqui MA, Zahid M. Evaluating Inference Compute for Generative AI: A Framework for Enterprise Workloads. https://omanscience.com/en/articles/evaluating-inference-compute-for-generative-ai-a-framework-for-enterprise-workloads

IEEE

A. R. Ali, M. A. Siddiqui, and M. Zahid, "Evaluating Inference Compute for Generative AI: A Framework for Enterprise Workloads," https://omanscience.com/en/articles/evaluating-inference-compute-for-generative-ai-a-framework-for-enterprise-workloads.