Transformer-Based Multi-Signal Predictive Autoscaling for SLA-Aware Resource Management in Kubernetes-Orchestrated Cloud-Native Environments
DOI:
https://doi.org/10.63282/3050-9262.IJAIDSML-V5I2P127Keywords:
Kubernetes, Predictive Autoscaling, Transformer, Cloud-Native Systems, Service-Level Agreement, Service-Level Objective, Microservices, Resource Management, Observability, Time-Series Forecasting, Horizontal Pod Autoscaler, AIOpsAbstract
Kubernetes has become the dominant orchestration substrate for cloud-native applications, yet its native autoscaling mechanisms remain primarily reactive, threshold-driven, and limited in their ability to anticipate workload volatility before service-level agreement violations occur. Modern microservice systems exhibit non-linear interactions among request arrival rates, queueing delays, CPU saturation, memory pressure, network variability, pod cold-start latency, and downstream dependency bottlenecks. These characteristics make single-metric autoscaling policies insufficient for latency-sensitive workloads operating under strict service-level objectives. This paper proposes a Transformer-Based Multi-Signal Predictive Autoscaling framework for SLA-aware resource management in Kubernetes-orchestrated cloud-native environments. The proposed framework integrates heterogeneous observability signals, multi-horizon time-series forecasting, uncertainty-aware decision logic, and Kubernetes-native actuation to allocate resources before overload conditions materialize. Unlike conventional Horizontal Pod Autoscaler configurations that respond after resource utilization crosses predefined thresholds, the proposed approach forecasts near-future demand and performance risk using a Transformer encoder architecture designed to learn long-range dependencies, temporal seasonality, burst behavior, and cross-metric interactions. The framework translates predicted workload and latency risk into safe scaling actions through policy constraints that consider replica bounds, cooldown windows, pod readiness delays, cost budgets, and SLA violation probability. The paper develops the conceptual architecture, methodological workflow, evaluation metrics, and analytical discussion necessary for empirical implementation. The study argues that SLA-aware predictive autoscaling should be treated not merely as a forecasting task but as an integrated control problem involving observability quality, model calibration, decision governance, and runtime safety. The proposed model contributes to cloud resource management research by aligning deep temporal learning with Kubernetes operational semantics and by providing a structured pathway toward more reliable, efficient, and self-adaptive cloud-native platforms.
References
[1] Gunda, S. K., Yettapu, S. D. R., Bodakunti, S., & Bikki, S. B. (2023). Decision Intelligence Methodology for AI-Driven Agile Software Lifecycle Governance and Architecture-Centered Project Management. International Journal of Artificial Intelligence, Data Science, and Machine Learning, 4(1), 102-108. https://doi.org/10.63282/3050-9262.IJAIDSML-V4I1P112
[2] A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention Is All You Need,” in Advances in Neural Information Processing Systems 30, Long Beach, CA, USA, 2017, pp. 5998–6008. Available: https://papers.nips.cc/paper/7181-attention-is-all-you-need
[3] T.-T. Nguyen, Y.-J. Yeom, T. Kim, D.-H. Park, and S. Kim, “Horizontal Pod Autoscaling in Kubernetes for Elastic Container Orchestration,” Sensors, vol. 20, no. 16, Art. no. 4621, 2020. https://doi.org/10.3390/s20164621
[4] Gunda, S. K. G. (2023). The Future of Software Development and the Expanding Role of ML Models. International Journal of Emerging Research in Engineering and Technology, 4(2), 126-129. https://doi.org/10.63282/3050-922X.IJERET-V4I2P113
[5] T. Lorido-Botran, J. Miguel-Alonso, and J. A. Lozano, “A Review of Auto-scaling Techniques for Elastic Applications in Cloud Environments,” Journal of Grid Computing, vol. 12, no. 4, pp. 559–592, 2014. https://doi.org/10.1007/s10723-014-9314-7
[6] H. Zhou, S. Zhang, J. Peng, S. Zhang, J. Li, H. Xiong, and W. Zhang, “Informer: Beyond Efficient Transformer for Long Sequence Time-Series Forecasting,” in Proceedings of the AAAI Conference on Artificial Intelligence, vol. 35, no. 12, 2021, pp. 11106–11115. https://doi.org/10.1609/aaai.v35i12.17325
[7] Shim, S., Dhokariya, A., Doshi, D., Upadhye, S., Patwari, V., & Park, J. Y. (2023). Predictive auto-scaler for Kubernetes cloud. In 2023 IEEE 17th Annual Systems Conference (SysCon) (pp. 1–8). IEEE. https://doi.org/10.1109/SysCon53073.2023.10131106
[8] C. Qu, R. N. Calheiros, and R. Buyya, “Auto-scaling Web Applications in Clouds: A Taxonomy and Survey,” ACM Computing Surveys, vol. 51, no. 4, Art. no. 73, pp. 1–33, 2018. https://doi.org/10.1145/3148149
[9] B. Lim, S. Ö. Arık, N. Loeff, and T. Pfister, “Temporal Fusion Transformers for Interpretable Multi-horizon Time Series Forecasting,” International Journal of Forecasting, vol. 37, no. 4, pp. 1748–1764, 2021. https://doi.org/10.1016/j.ijforecast.2021.03.012
[10] Vu, D.-D., Tran, M.-N., & Kim, Y. (2022). Predictive hybrid autoscaling for containerized applications. IEEE Access, 10, 109768–109778. https://doi.org/10.1109/ACCESS.2022.3214985
[11] L. H. Phuc, L.-A. Phan, and T. Kim, “Traffic-Aware Horizontal Pod Autoscaler in Kubernetes-Based Edge Computing Infrastructure,” IEEE Access, vol. 10, pp. 18966–18977, 2022. https://doi.org/10.1109/ACCESS.2022.3150867
[12] T. Chen, R. Bahsoon, and X. Yao, “A Survey and Taxonomy of Self-Aware and Self-Adaptive Cloud Autoscaling Systems,” ACM Computing Surveys, vol. 51, no. 3, Art. no. 61, pp. 1–40, 2018. https://doi.org/10.1145/3190507
[13] H. Wu, J. Xu, J. Wang, and M. Long, “Autoformer: Decomposition Transformers with Auto-Correlation for Long-Term Series Forecasting,” in Advances in Neural Information Processing Systems, vol. 34, 2021, pp. 22419–22430. Available: https://arxiv.org/abs/2106.13008
[14] B. Burns, B. Grant, D. Oppenheimer, E. Brewer, and J. Wilkes, “Borg, Omega, and Kubernetes: Lessons Learned from Three Container-Management Systems over a Decade,” ACM Queue, vol. 14, no. 1, pp. 70–93, 2016. https://doi.org/10.1145/2898442.2898444
[15] D.-D. Vu, M.-N. Tran, and Y. Kim, “Predictive Hybrid Autoscaling for Containerized Applications,” IEEE Access, vol. 10, pp. 109768–109778, 2022. https://doi.org/10.1109/ACCESS.2022.3214985
[16] Y. Nie, N. H. Nguyen, P. Sinthong, and J. Kalagnanam, “A Time Series Is Worth 64 Words: Long-Term Forecasting with Transformers,” in Proceedings of the International Conference on Learning Representations, 2023. Available: https://arxiv.org/abs/2211.14730
[17] Y. Gan et al., “An Open-Source Benchmark Suite for Microservices and Their Hardware-Software Implications for Cloud and Edge Systems,” in Proceedings of the Twenty-Fourth International Conference on Architectural Support for Programming Languages and Operating Systems, Providence, RI, USA, 2019, pp. 3–18. https://doi.org/10.1145/3297858.3304013
[18] D. R. Augustyn, Ł. Wyciślik, and M. Sojka, “Tuning a Kubernetes Horizontal Pod Autoscaler for Meeting Performance and Load Demands in Cloud Deployments,” Applied Sciences, vol. 14, no. 2, Art. no. 646, 2024. https://doi.org/10.3390/app14020646
[19] O. Pozdniakova, D. Mažeika, and A. Cholomskis, “SLA-Adaptive Threshold Adjustment for a Kubernetes Horizontal Pod Autoscaler,” Electronics, vol. 13, no. 7, Art. no. 1242, pp. 1–28, 2024. https://doi.org/10.3390/electronics13071242
[20] Y. Garí, D. A. Monge, E. Pacini, C. Mateos, and C. G. Garino, “Reinforcement Learning-Based Application Autoscaling in the Cloud: A Survey,” Engineering Applications of Artificial Intelligence, vol. 102, Art. no. 104288, 2021. https://doi.org/10.1016/j.engappai.2021.104288
[21] M. Abdullah, W. Iqbal, A. Mahmood, F. Bukhari, and A. Erradi, “Predictive Autoscaling of Microservices Hosted in Fog Microdata Center,” IEEE Systems Journal, vol. 15, no. 1, pp. 1275–1286, 2021. https://doi.org/10.1109/JSYST.2020.2997518
[22] C. Zhu, B. Han, and Y. Zhao, “A Bi-Metric Autoscaling Approach for N-Tier Web Applications on Kubernetes,” Frontiers of Computer Science, vol. 16, no. 3, 2022. https://doi.org/10.1007/s11704-021-0118-1
[23] A. Verma, L. Pedrosa, M. R. Korupolu, D. Oppenheimer, E. Tune, and J. Wilkes, “Large-Scale Cluster Management at Google with Borg,” in Proceedings of the European Conference on Computer Systems, Bordeaux, France, 2015, pp. 1–17. https://doi.org/10.1145/2741948.2741964
[24] Z. Zhou, C. Zhang, L. Ma, J. Gu, H. Qian, Q. Wen, L. Sun, P. Li, and Z. Tang, “AHPA: Adaptive Horizontal Pod Autoscaling Systems on Alibaba Cloud Container Service for Kubernetes,” arXiv:2303.03640, 2023. Available: https://arxiv.org/abs/2303.03640
[25] S. Xie, J. Wang, B. Li, Z. Zhang, D. Li, and P. C. K. Hung, “PBScaler: A Bottleneck-Aware Autoscaling Framework for Microservice-Based Applications,” arXiv:2303.14620, 2023. Available: https://arxiv.org/abs/2303.14620










