Abstract
User-written Triton kernels enable high-performance GPU computation within PyTorch, but their end-to-end latency can remain dominated by host-side orchestration, especially when device execution is short. Although torch.compile can generate native host wrappers for captured graphs, each invocation still passes through runtime-managed specialization lookup, guard evaluation, and preparation before reaching the wrapper. We present Trident, a compiler backend that removes this recurring overhead from the specialization cache-hit path. Trident introduces the Specialization Cache Module (SCM), which compiles guarded specialization selection, argument and execution-environment preparation, and host execution for multiple specializations into a single executable module. An invocation enters the SCM once, remains in compiled code when a specialization matches, and returns to Python only when a new specialization must be compiled. Built on Torch-MLIR, Trident lowers guards and host-side orchestration to native code while retaining calls to optimized runtime implementations of supported ATen operators. Our evalu- ation on two LLMs shows that Trident achieves up to a 1.47x speedup in model-level end-to-end latency over eager execution and up to 1.68x over torch.compile.
Keywords
Subject
Publication details
- Journal
- Not available
- Open access
- Green open access
Cite this article
APA 7
Liu, J., Liu, X., Zhang, S., Sun, W., Yang, R., Men, C., Lin, Y., & Li, S. (2026). Trident: Unifying Guarded Dispatch and Host Execution for PyTorch Triton Workloads. https://omanscience.com/en/articles/trident-unifying-guarded-dispatch-and-host-execution-for-pytorch-triton-workloads
MLA 9
Liu, Jinjie, et al. "Trident: Unifying Guarded Dispatch and Host Execution for PyTorch Triton Workloads." https://omanscience.com/en/articles/trident-unifying-guarded-dispatch-and-host-execution-for-pytorch-triton-workloads.
Chicago (author–date)
Liu, Jinjie, Xiaoyan Liu, Shuhan Zhang, Wenjia Sun, Ruilin Yang, Chunlei Men, Yonghua Lin, and Shaohua Li. 2026. "Trident: Unifying Guarded Dispatch and Host Execution for PyTorch Triton Workloads." https://omanscience.com/en/articles/trident-unifying-guarded-dispatch-and-host-execution-for-pytorch-triton-workloads.
Harvard
Liu, J., Liu, X., Zhang, S., Sun, W., Yang, R., Men, C., Lin, Y. and Li, S. (2026) 'Trident: Unifying Guarded Dispatch and Host Execution for PyTorch Triton Workloads', Available at: https://omanscience.com/en/articles/trident-unifying-guarded-dispatch-and-host-execution-for-pytorch-triton-workloads.
Vancouver
Liu J, Liu X, Zhang S, Sun W, Yang R, Men C, et al. Trident: Unifying Guarded Dispatch and Host Execution for PyTorch Triton Workloads. https://omanscience.com/en/articles/trident-unifying-guarded-dispatch-and-host-execution-for-pytorch-triton-workloads
IEEE
J. Liu, X. Liu, S. Zhang, W. Sun, R. Yang, C. Men, Y. Lin, and S. Li, "Trident: Unifying Guarded Dispatch and Host Execution for PyTorch Triton Workloads," https://omanscience.com/en/articles/trident-unifying-guarded-dispatch-and-host-execution-for-pytorch-triton-workloads.