Abstract

In-database predictive query processing increasingly applies Transformer-based models within relational pipelines. However, existing in-database inference typically exposes only tuple-level model inputs to the inference runtime, leaving relational predicates and metadata statistics invisible to neural execution planning. In this paper, we propose QCATS, a query context-aware transformer slicing framework that enables efficient sparse inference inside database systems. QCATS executes at query granularity: instead of routing individual tokens or tuples during inference, it uses query predicates and metadata statistics to pre-select context-aligned FFN slices before model execution. The framework comprises offline expert construction and lightweight query-level routing that dynamically selects experts during execution. QCATS further introduces system optimizations, including asynchronous CPU-GPU pipelines and routing-aware batching. Experiments on four predictive-query workloads with BERT-base and Qwen-0.6B show that QCATS achieves up to 4.42x latency reduction while preserving prediction accuracy comparable to dense baselines.

Keywords

Subject

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Li, Y., Xie, Z., Chen, K., & Shou, L. (2026). QCATS: Query Context-Aware Transformer Slicing for Efficient Predictive Query Processing. https://omanscience.com/en/articles/qcats-query-context-aware-transformer-slicing-for-efficient-predictive-query-processing

MLA 9

Li, Yueying, et al. "QCATS: Query Context-Aware Transformer Slicing for Efficient Predictive Query Processing." https://omanscience.com/en/articles/qcats-query-context-aware-transformer-slicing-for-efficient-predictive-query-processing.

Chicago (author–date)

Li, Yueying, Zhongle Xie, Ke Chen, and Lidan Shou. 2026. "QCATS: Query Context-Aware Transformer Slicing for Efficient Predictive Query Processing." https://omanscience.com/en/articles/qcats-query-context-aware-transformer-slicing-for-efficient-predictive-query-processing.

Harvard

Li, Y., Xie, Z., Chen, K. and Shou, L. (2026) 'QCATS: Query Context-Aware Transformer Slicing for Efficient Predictive Query Processing', Available at: https://omanscience.com/en/articles/qcats-query-context-aware-transformer-slicing-for-efficient-predictive-query-processing.

Vancouver

Li Y, Xie Z, Chen K, Shou L. QCATS: Query Context-Aware Transformer Slicing for Efficient Predictive Query Processing. https://omanscience.com/en/articles/qcats-query-context-aware-transformer-slicing-for-efficient-predictive-query-processing

IEEE

Y. Li, Z. Xie, K. Chen, and L. Shou, "QCATS: Query Context-Aware Transformer Slicing for Efficient Predictive Query Processing," https://omanscience.com/en/articles/qcats-query-context-aware-transformer-slicing-for-efficient-predictive-query-processing.