CyberSecLLM: A Hybrid Mamba–Transformer Foundation Model with ODE-Augmented State Dynamics for Zero-Shot Threat Intelligence

Roger Nick Anaedevha, Alexander Trofimov

Abstract


Contemporary intrusion detection systems rely on supervised paradigms that require extensive labeled data and rapidly become obsolete as adversaries evolve. This paper in- troduces CyberSecLLM, a family of domain-specific foundation models (3B, 7B, 13B parameters) pre-trained on six publicly available security datasets spanning cloud, industrial IoT, and enterprise environments. The architecture makes three contri- butions. First, it proposes ODE-augmented selective state space (ODE-S6) blocks, which embed neural ordinary differential equa- tion solvers with temporal adaptive batch normalization within Mamba layers to model continuous-time network state evolution, unifying the discrete selective scan with nonlinear continuous dynamics. Second, it interleaves these blocks with sparse cross- attention transformer layers grounded in a MITRE ATT&CK knowledge base and task-specialized mixture-of-experts routing. Third, the pre-training procedure combines five complementary objectives - masked flow modeling, contrastive attack discrim- ination, temporal forecasting, ATT&CK technique attribution, and a novel deep spatio-temporal point process loss that explicitly models inter-event attack timing via a neural conditional intensity function—balanced through learned uncertainty weights. Evalua- tion on CyberSecBench, an eleven-task benchmark incorporating three established LLM security suites (CTIBench, SecBench, CyberMetric), demonstrates that the 7B variant achieves 89.1% average zero-shot accuracy, a 20.5-point improvement over GPT- 4 and a 15.3-point gain over Foundation-Sec-8B. The architecture processes million-token network traces in 2.3 seconds on a single A100, a 19× speedup over quadratic transformers. All code, models, and the evaluation suite are released openly to support reproducibility.


Full Text:

PDF

References


R.N.Anaedevha,“CT-TGNN+:Continuous-timetemporalgraphneural networks for network intrusion detection,” TechRxiv Preprint, 2025. [Online]. Available: https://doi.org/10.36227/techrxiv.1375632

R.N.Anaedevha,“NeuralODEmodelsfordynamicalsystemrepresenta- tion learning,” 2024. [Online]. Available: https://github.com/rogerpanel/ Neural- ODE- models

W. Zheng, Y. Gao, Y. Sun, M. Boning, B. Yu, and T. Wong, “Improving Neural ODE training with temporal adaptive batch normalization,” in Proc. Adv. Neural Inf. Process. Syst. (NeurIPS), 2024.

A. Gu and T. Dao, “Mamba: Linear-time sequence modeling with selective state spaces,” in Proc. Int. Conf. Learn. Represent. (ICLR), 2024.

H. Mei and J. M. Eisner, “The Neural Hawkes Process: A neurally self- modulating multivariate point process,” in Proc. Adv. Neural Inf. Process. Syst. (NeurIPS), 2017, pp. 6754–6764.

S. Zuo, H. Jiang, Z. Li, T. Zhao, and H. Zha, “Transformer Hawkes Process,” in Proc. Int. Conf. Mach. Learn. (ICML), vol. 119, 2020, pp. 11692–11702.

M.Alametal.,“CTIBench:AbenchmarkforevaluatingLLMsincyber threat intelligence,” in Proc. NeurIPS Datasets Track, 2024.

Z. Jing et al., “SecBench: A comprehensive multi-dimensional benchmarking dataset for LLMs in cybersecurity,” arXiv preprint arXiv:2412.20787, 2024.

N.Tihanyietal.,“CyberMetric:Abenchmarkdatasetbasedonretrieval- augmented generation for evaluating LLMs in cybersecurity knowledge,” in Proc. IEEE CSR, 2024.

E. Aghaei, X. Niu, W. Shadid, and E. Al-Shaer, “SecureBERT: A domain-specific language model for cybersecurity,” in Proc. Security Privacy Workshops (SPW), 2023, pp. 285–296.

T.Michaeletal.,“SecGPT:AnexecutionisolationarchitectureforLLM- based systems,” arXiv preprint arXiv:2403.04960, 2024.

D. Liu et al., “SecurityLLM: Security-specific large language models,” arXiv preprint arXiv:2405.00893, 2024.

Y. Chen et al., “Foundation-Sec-8B: A cybersecurity-focused large language model,” arXiv preprint arXiv:2504.01838, 2025.

A. Gu, K. Goel, and C. Ré, “Efficiently modeling long sequences with structured state spaces,” in Proc. Int. Conf. Learn. Represent. (ICLR), 2022.

T. Dao and A. Gu, “Transformers are SSMs: Generalized models and efficient algorithms through structured state space duality,” in Proc. Int. Conf. Mach. Learn. (ICML), 2024.

R. T. Q. Chen, Y. Rubanova, J. Bettencourt, and D. K. Duvenaud, “Neu- ral ordinary differential equations,” in Proc. Adv. Neural Inf. Process. Syst. (NeurIPS), 2018, pp. 6571–6583.

X. Li, T. K. Wong, R. T. Q. Chen, and D. K. Duvenaud, “Scalable gradients for stochastic differential equations,” in Proc. AISTATS, 2020.

E. Dupont, A. Doucet, and Y. W. Teh, “Augmented neural ODEs,” in

Proc. NeurIPS, 2019, pp. 3140–3150.

P. Kidger, J. Morrill, J. Foster, and T. Lyons, “Neural controlled

differential equations for irregular time series,” in Proc. NeurIPS, 2020.

A. Baldwin, I. Gheyas, C. Ioannidis, D. Pym, and J. Williams, “Conta- gion in cyber security attacks,” J. Oper. Res. Soc., vol. 68, pp. 780–791, 2017.

C. Peng, M. Xu, S. Xu, and T. Hu, “Modeling and predicting extreme cyber attack rates via marked point processes,” J. Appl. Stat., vol. 44, no. 14, pp. 2534–2563, 2017.

P. Ferraro, C. King, and R. Shorten, “Identification and prediction of attacks to industrial control systems using temporal point processes,” J. Ambient Intell. Humaniz. Comput., vol. 14, pp. 14641–14653, 2022.

Q. Zhang, A. Lipani, O. Kirnap, and E. Yilmaz, “Self-Attentive Hawkes Process,” in Proc. Int. Conf. Mach. Learn. (ICML), vol. 119, 2020, pp. 11183–11193.

O.Lieberetal.,“Jamba:AhybridTransformer-Mambalanguagemodel,” arXiv preprint arXiv:2403.19887, 2024.

P. Glorioso et al., “Zamba: A compact 7B SSM hybrid model,” arXiv preprint arXiv:2405.18712, 2024.

M. Pióro et al., “MoE-Mamba: Efficient selective state space models with mixture of experts,” arXiv preprint arXiv:2401.04081, 2024.

W. Fedus, B. Zoph, and N. Shazeer, “Switch Transformers: Scaling to

trillion parameter models with simple and efficient sparsity,” J. Mach.

Learn. Res., vol. 23, no. 120, pp. 1–39, 2022.

D. Dai et al., “DeepSeekMoE: Towards ultimate expert specialization in

mixture-of-experts language models,” arXiv preprint arXiv:2401.06066,

X. Lin, G. Xiong, G. Gou, Z. Li, J. Shi, and J. Yu, “ET-BERT: A

contextualized datagram representation with pre-training transformers for encrypted traffic classification,” in Proc. ACM Web Conf., 2022, pp. 633–642.

X. Wang et al., “NetGPT: An AI-native network architecture for provi- sioning beyond personalized generative services,” IEEE Netw., 2024.

Z. Wang et al., “TrafficLLM: Enhancing large language models for net- work traffic analysis with generic traffic representation,” arXiv preprint arXiv:2504.04222, 2025.

A. Alsaheel et al., “CONTINUUM: Detecting APT attacks through spatial-temporal graph neural networks,” in Proc. NDSS, 2025.

M. Bhatt et al., “CyberSOCEval: Benchmarking LLMs capabilities for malware analysis and threat intelligence reasoning,” arXiv preprint arXiv:2509.20166, 2025.

Y. Yang et al., “Cybersecurity AI benchmark (CAIBench): A meta- benchmark for evaluating cybersecurity AI agents,” arXiv preprint arXiv:2510.24317, 2025.

D. Ruiz-Rodenas et al., “SynthCTI: LLM-driven synthetic CTI generation to enhance MITRE technique mapping,” arXiv preprint arXiv:2507.16852, 2025.

Y. Chen et al., “Knowledge-to-data: LLM-driven synthesis of struc- tured network traffic for testbed-free IDS evaluation,” arXiv preprint arXiv:2601.05022, 2025.

A. Gu, T. Dao, S. Ermon, A. Rudra, and C. Ré, “HiPPO: Recurrent memory with optimal polynomial projections,” in Proc. Adv. Neural Inf. Process. Syst. (NeurIPS), vol. 33, 2020, pp. 1474–1487.

A. G. Hawkes, “Spectra of some self-exciting and mutually exciting point processes,” Biometrika, vol. 58, no. 1, pp. 83–90, 1971.

N. Shazeer, “GLU Variants improve Transformer,” arXiv preprint arXiv:2002.05202, 2020.

A. Kendall, Y. Gal, and R. Cipolla, “Multi-task learning using uncer- tainty to weigh losses for scene geometry and semantics,” in Proc. IEEE Conf. Comput. Vis. Pattern Recognit. (CVPR), 2018, pp. 7482–7491.

I. Loshchilov and F. Hutter, “Decoupled weight decay regularization,” in Proc. Int. Conf. Learn. Represent. (ICLR), 2019.

S. Rajbhandari, J. Rasley, O. Rber, and Y. He, “ZeRO: Memory optimizations toward training trillion parameter models,” in Proc. Int. Conf. High Perform. Comput., 2020, pp. 1–16.

L. Xu, M. Skoularidou, A. Cuesta-Infante, and K. Veeramachaneni, “Modeling tabular data using conditional GAN,” in Proc. NeurIPS, 2019.

R. Hasani, M. Lechner, T. A. Wang, M. Elibol, and D. Rus, “Liquid

structural state-space models,” in Proc. ICLR, 2023.

J. Reha et al., “Anomaly detection in continuous-time temporal provenance graphs,” in Proc. NeurIPS Workshop TSALM, 2023.

J.Demšar,“Statistical comparisons of classifiers over multiple datasets,” J. Mach. Learn. Res., vol. 7, pp. 1–30, 2006.

Y. Rubanova, R. T. Q. Chen, and D. Duvenaud, “Latent ODEs for irregularly-sampled time series,” in Proc. NeurIPS, 2019.

E. De Brouwer, J. Simm, A. Arany, and Y. Moreau, “GRU-ODE-Bayes: Continuous modeling of sporadically-observed time series,” in Proc. NeurIPS, 2019.

P. Kidger, J. Morrill, J. Foster, and T. Lyons, “Neural controlled differential equations for irregular time series,” in Proc. NeurIPS, 2020, pp. 6696–6707.

J. Jia and A. R. Benson, “Neural jump stochastic differential equations,” in Proc. NeurIPS, 2019.

G. Fortino, C. Greco, A. Guzzo, and M. Ianni, “Neural network based temporal point processes for attack detection in industrial control systems,” in Proc. IEEE Int. Conf. Cyber Secur. Resilience (CSR), 2022.

P. Sun, S. Bagchi, et al., “Using Dirichlet marked Hawkes processes for insider threat detection,” ACM Digit. Threats: Res. Pract., vol. 3, no. 1, 2022.

W. W. Lo, S. Layeghy, M. Sarhan, M. Gallagher, and M. Portmann, “E-GraphSAGE: A graph neural network based intrusion detection system for IoT,” in Proc. IEEE/IFIP Netw. Oper. Manage. Symp. (NOMS), 2022.

E. Caville, W. W. Lo, S. Layeghy, and M. Portmann, “Anomal-E: A self-supervised network intrusion detection system based on graph neural networks,” Knowl.-Based Syst., vol. 258, p. 110030, 2022.

M. Sarhan, S. Layeghy, N. Moustafa, and M. Portmann, “NetFlow datasets for machine learning-based network intrusion detection systems,” in Proc. Big Data Technol. Appl. (BDTA), 2021.

R. Zhao et al., “Yet another traffic classifier: A masked autoencoder based traffic transformer with multi-level flow representation,” in Proc. AAAI Conf. Artif. Intell., 2023.

T. Wang, X. Xie, W. Wang, C. Wang, Y. Zhao, and Y. Cui, “NetMamba: Efficient network traffic classification via pre-training unidirectional Mamba,” in Proc. IEEE Int. Conf. Netw. Protocols (ICNP), 2024.

S. Guthula, R. Beltiukov, N. Battula, W. Guo, and A. Gupta, “netFound: Foundation model for network security,” arXiv preprint arXiv:2310.17025, 2023.

S. Hegselmann, A. Buendia, H. Lang, M. Agrawal, X. Jiang, and D. Sontag, “TabLLM: Few-shot classification of tabular data with large language models,” in Proc. Int. Conf. Artif. Intell. Statist. (AISTATS), PMLR vol. 206, 2023.

T. Dinh et al., “LIFT: Language-interfaced fine-tuning for non-language machine learning tasks,” in Proc. NeurIPS, 2022.

D. Lopez-Paz and M. Oquab, “Revisiting classifier two-sample tests,” in Proc. Int. Conf. Learn. Represent. (ICLR), 2017.

T. Brown et al., “Language models are few-shot learners,” in Proc. NeurIPS, 2020.


Refbacks

  • There are currently no refbacks.


Abava  Кибербезопасность ИТ-КОНГРЕСС ВМК МГУ 2026 СНЭ

ISSN: 2307-8162