Abstract

Industry applications often demand low-latency classification, yet current large language model (LLM) approaches remain poorly suited for latency-critical applications. Existing prompting and constrained decoding produce verbose, multi-token outputs that require expensive token-by-token generation, while encoder-based models achieve faster inference but sacrifice task flexibility. We propose Koa-action, a framework for low-latency atomic actions -- fast, single-step decisions such as classification, semantic endpointing, Boolean checks, and scoring -- formulated as constrained generation with single-token outputs. By introducing atomic label tokens and applying supervised fine-tuning, our method reduces classification to a deterministic one-step decoding problem. Across standard benchmarks, Koa-action delivers competitive accuracy with consistently low and stable latency. On a production intent-routing benchmark, Koa-action reaches 85.5% accuracy -- competitive with the strongest frontier models (Claude-4.8-Opus, Gemini-Pro-3.1) and ahead of GPT-5 and Gemini-2.5-Pro -- while answering in about half a second, several-fold faster than every frontier model (up to ~7.5x at the median) under identical serving conditions. Against the dedicated single-token system Jev/TypeSafe, Koa-action is competitive on accuracy and faster at the median, while also handling multimodal inputs and multi-label outputs that single-label text systems do not.

Keywords

Subject

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Dai, S., Pentyala, S. K., Liu, Y., Mehrotra, S., Banerjee, S., Zhu, J., Bi, B., Asur, S., & Mui, P. (2026). Koa-action: Fast and Consistent Structured Decision Making with Generative LLMs. https://omanscience.com/en/articles/koa-action-fast-and-consistent-structured-decision-making-with-generative-llms

MLA 9

Dai, Shenghong, et al. "Koa-action: Fast and Consistent Structured Decision Making with Generative LLMs." https://omanscience.com/en/articles/koa-action-fast-and-consistent-structured-decision-making-with-generative-llms.

Chicago (author–date)

Dai, Shenghong, Shiva Kumar Pentyala, Yingchi Liu, Shubham Mehrotra, Suman Banerjee, James Zhu, Bin Bi, Sitaram Asur, and Phil Mui. 2026. "Koa-action: Fast and Consistent Structured Decision Making with Generative LLMs." https://omanscience.com/en/articles/koa-action-fast-and-consistent-structured-decision-making-with-generative-llms.

Harvard

Dai, S., Pentyala, S. K., Liu, Y., Mehrotra, S., Banerjee, S., Zhu, J., Bi, B., Asur, S. and Mui, P. (2026) 'Koa-action: Fast and Consistent Structured Decision Making with Generative LLMs', Available at: https://omanscience.com/en/articles/koa-action-fast-and-consistent-structured-decision-making-with-generative-llms.

Vancouver

Dai S, Pentyala SK, Liu Y, Mehrotra S, Banerjee S, Zhu J, et al. Koa-action: Fast and Consistent Structured Decision Making with Generative LLMs. https://omanscience.com/en/articles/koa-action-fast-and-consistent-structured-decision-making-with-generative-llms

IEEE

S. Dai, S. K. Pentyala, Y. Liu, S. Mehrotra, S. Banerjee, J. Zhu, B. Bi, S. Asur, and P. Mui, "Koa-action: Fast and Consistent Structured Decision Making with Generative LLMs," https://omanscience.com/en/articles/koa-action-fast-and-consistent-structured-decision-making-with-generative-llms.