Abstract
In-database predictive query processing increasingly applies Transformer-based models within relational pipelines. However, existing in-database inference typically exposes only tuple-level model inputs to the inference runtime, leaving relational predicates and metadata statistics invisible to neural execution planning. In this paper, we propose QCATS, a query context-aware transformer slicing framework that enables efficient sparse inference inside database systems. QCATS executes at query granularity: instead of routing individual tokens or tuples during inference, it uses query predicates and metadata statistics to pre-select context-aligned FFN slices before model execution. The framework comprises offline expert construction and lightweight query-level routing that dynamically selects experts during execution. QCATS further introduces system optimizations, including asynchronous CPU-GPU pipelines and routing-aware batching. Experiments on four predictive-query workloads with BERT-base and Qwen-0.6B show that QCATS achieves up to 4.42x latency reduction while preserving prediction accuracy comparable to dense baselines.
Keywords
Subject
Publication details
- Journal
- Not available
- Open access
- Green open access
Cite this article
APA 7
Li, Y., Xie, Z., Chen, K., & Shou, L. (2026). QCATS: Query Context-Aware Transformer Slicing for Efficient Predictive Query Processing. https://omanscience.com/en/articles/qcats-query-context-aware-transformer-slicing-for-efficient-predictive-query-processing
MLA 9
Li, Yueying, et al. "QCATS: Query Context-Aware Transformer Slicing for Efficient Predictive Query Processing." https://omanscience.com/en/articles/qcats-query-context-aware-transformer-slicing-for-efficient-predictive-query-processing.
Chicago (author–date)
Li, Yueying, Zhongle Xie, Ke Chen, and Lidan Shou. 2026. "QCATS: Query Context-Aware Transformer Slicing for Efficient Predictive Query Processing." https://omanscience.com/en/articles/qcats-query-context-aware-transformer-slicing-for-efficient-predictive-query-processing.
Harvard
Li, Y., Xie, Z., Chen, K. and Shou, L. (2026) 'QCATS: Query Context-Aware Transformer Slicing for Efficient Predictive Query Processing', Available at: https://omanscience.com/en/articles/qcats-query-context-aware-transformer-slicing-for-efficient-predictive-query-processing.
Vancouver
Li Y, Xie Z, Chen K, Shou L. QCATS: Query Context-Aware Transformer Slicing for Efficient Predictive Query Processing. https://omanscience.com/en/articles/qcats-query-context-aware-transformer-slicing-for-efficient-predictive-query-processing
IEEE
Y. Li, Z. Xie, K. Chen, and L. Shou, "QCATS: Query Context-Aware Transformer Slicing for Efficient Predictive Query Processing," https://omanscience.com/en/articles/qcats-query-context-aware-transformer-slicing-for-efficient-predictive-query-processing.