Abstract

The increasing complexity and scale of modern cyber-attacks demand intelligent and computationally efficient Intrusion Detection Systems (IDS). However, designing effective data pre-processing pipelines traditionally involves substantial trial-and-error effort and repeated evaluation of alternative configurations. For large, high-dimensional network traffic data, this process can create a significant computational burden. In this work, we propose AutoDP-LLM, an automated pre-processing framework designed to reduce manual pipeline development and computational overhead. Specifically, AutoDP-LLM leverages Large Language Models (LLMs) to autonomously generate and validate executable data pre-processing pipelines. The framework combines deterministic host-side planning with LLM-based specialist agents to formulate data-processing strategies, synthesize executable code, and adaptively determine retained feature sets using semantic reasoning and training-derived statistical evidence, without requiring a predefined feature budget. Focusing on multiclass intrusion detection, we evaluate AutoDP-LLM on the UNSW-NB15 and NSL-KDD benchmark datasets using multiple downstream classifiers. Comparative experiments against conventional feature-selection methods show that AutoDP-LLM achieves competitive detection performance while automating the generation of compact and executable pre-processing pipelines. Component-level ablation experiments further demonstrate the complementary contributions of the semantic and statistical feature-reduction components. The repeated generation, validation, execution, and assessment of candidate pipelines are amenable to parallel execution, highlighting the potential of scalable computing environments, including high-performance computing (HPC) systems, to support automated IDS pipeline development.

Keywords

Subject

Publication details

DOI
10.1007/s11227-026-08894-8
Journal
Not available
Open access
Green open access

Cite this article

APA 7

Nguyen, B. P., Pham, G. K., Do, T. D., Trang, M. X., Le, M. T., Tran, X. N., Vu, H., Nguyen, T. C., Ngo, V. D., & Van Luong, T. (2026). AutoDP-LLM: Automating Data Pre-processing for Intrusion Detection Systems using Large Language Models. https://doi.org/10.1007/s11227-026-08894-8

MLA 9

Nguyen, Bao-Phong, et al. "AutoDP-LLM: Automating Data Pre-processing for Intrusion Detection Systems using Large Language Models." https://doi.org/10.1007/s11227-026-08894-8.

Chicago (author–date)

Nguyen, Bao-Phong, Gia-Khanh Pham, Thai-Duong Do, Mai Xuan Trang, Minh-Tuan Le, Xuan-Nam Tran, Huan Vu, Tien-Cuong Nguyen, Vu-Duc Ngo, and Thien Van Luong. 2026. "AutoDP-LLM: Automating Data Pre-processing for Intrusion Detection Systems using Large Language Models." https://doi.org/10.1007/s11227-026-08894-8.

Harvard

Nguyen, B. P., Pham, G. K., Do, T. D., Trang, M. X., Le, M. T., Tran, X. N., Vu, H., Nguyen, T. C., Ngo, V. D. and Van Luong, T. (2026) 'AutoDP-LLM: Automating Data Pre-processing for Intrusion Detection Systems using Large Language Models', doi:10.1007/s11227-026-08894-8.

Vancouver

Nguyen BP, Pham GK, Do TD, Trang MX, Le MT, Tran XN, et al. AutoDP-LLM: Automating Data Pre-processing for Intrusion Detection Systems using Large Language Models. doi:10.1007/s11227-026-08894-8

IEEE

B. P. Nguyen, G. K. Pham, T. D. Do, M. X. Trang, M. T. Le, X. N. Tran, H. Vu, T. C. Nguyen, V. D. Ngo, and T. Van Luong, "AutoDP-LLM: Automating Data Pre-processing for Intrusion Detection Systems using Large Language Models," doi: 10.1007/s11227-026-08894-8.