Quantization and Compression Techniques for Model Deployment: Optimizing Neural Networks for Production
DOI:
https://doi.org/10.63282/3050-9262.IJAIDSML-V7I3P115Keywords:
Post-Training Quantization, INT8, KV-cache, Model Compression, CPU Inference, Mixed Precision, Operator Coverage, Inference Memory, Deployment Readiness, Neural Network OptimizationAbstract
Compression research reports compression ratios, and deployment engineering pays for something else. This paper reports a CPU benchmark study in which the binding constraint was neither model size nor latency but key-value cache memory, which grows linearly with sequence length and comes to dominate the inference memory budget well before weight storage does. Post-training INT8 quantization was applied to reduce that footprint, with latency treated as a secondary and unguaranteed benefit. The central finding concerns the gap between the compression ratio and the realized gain. Memory reduction followed the quantization ratio reliably. Runtime did not, because the software stack did not sustain low-precision execution end to end: several operators lacked native INT8 kernels, so tensors were repeatedly upcast to floating point and downcast back to INT8 between operations, and the conversion overhead consumed much of the arithmetic saving. The practical consequence is that a compression decision on CPU is governed by operator coverage across the execution path rather than by the nominal precision of the weights, and a technique that halves arithmetic width while doubling conversions can be a net loss. INT4 was considered and not pursued, on grounds of accuracy risk and the same coverage limitation in a more severe form. The paper also proposes a promotion process for moving a compressed model toward production against an uncompressed reference baseline. This is a benchmark feasibility study. No model was deployed to production or pilot, no quantitative result is reported, and Section IX states precisely what was not measured.
References
[1] B. Jacob, S. Kligys, B. Chen, M. Zhu, M. Tang, A. Howard, H. Adam, and D. Kalenichenko, "Quantization and Training of Neural Networks for Efficient Integer-Arithmetic-Only Inference," in Proc. IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2018, pp. 2704-2713. doi: 10.1109/CVPR.2018.00286
[2] R. Krishnamoorthi, "Quantizing deep convolutional networks for efficient inference: A whitepaper," arXiv:1806.08342, 2018. [Online]. Available: https://arxiv.org/abs/1806.08342
[3] M. Nagel, M. Fournarakis, R. A. Amjad, Y. Bondarenko, M. van Baalen, and T. Blankevoort, "A White Paper on Neural Network Quantization," arXiv:2106.08295, 2021. [Online]. Available: https://arxiv.org/abs/2106.08295
[4] A. Gholami, S. Kim, Z. Dong, Z. Yao, M. W. Mahoney, and K. Keutzer, "A Survey of Quantization Methods for Efficient Neural Network Inference," arXiv:2103.13630, 2021. [Online]. Available: https://arxiv.org/abs/2103.13630
[5] S. Han, H. Mao, and W. J. Dally, "Deep Compression: Compressing Deep Neural Networks with Pruning, Trained Quantization and Huffman Coding," arXiv:1510.00149, 2015. [Online]. Available: https://arxiv.org/abs/1510.00149
[6] J. Frankle and M. Carbin, "The Lottery Ticket Hypothesis: Finding Sparse, Trainable Neural Networks," arXiv:1803.03635, 2018. [Online]. Available: https://arxiv.org/abs/1803.03635
[7] G. Hinton, O. Vinyals, and J. Dean, "Distilling the Knowledge in a Neural Network," arXiv:1503.02531, 2015. [Online]. Available: https://arxiv.org/abs/1503.02531
[8] T. Dettmers, M. Lewis, Y. Belkada, and L. Zettlemoyer, "LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale," arXiv:2208.07339, 2022. [Online]. Available: https://arxiv.org/abs/2208.07339
[9] E. Frantar, S. Ashkboos, T. Hoefler, and D. Alistarh, "GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers," arXiv:2210.17323, 2022. [Online]. Available: https://arxiv.org/abs/2210.17323
[10] G. Xiao, J. Lin, M. Seznec, H. Wu, J. Demouth, and S. Han, "SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language Models," arXiv:2211.10438, 2022. [Online]. Available: https://arxiv.org/abs/2211.10438
[11] C. Hooper, S. Kim, H. Mohammadzadeh, M. W. Mahoney, Y. S. Shao, K. Keutzer, and A. Gholami, "KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization," arXiv:2401.18079, 2024. [Online]. Available: https://arxiv.org/abs/2401.18079
[12] Z. Liu, J. Yuan, H. Jin, S. Zhong, Z. Xu, V. Braverman, B. Chen, and X. Hu, "KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache," arXiv:2402.02750, 2024. [Online]. Available: https://arxiv.org/abs/2402.02750
[13] R. Pope, S. Douglas, A. Chowdhery, J. Devlin, J. Bradbury, A. Levskaya, J. Heek, K. Xiao, S. Agrawal, and J. Dean, "Efficiently Scaling Transformer Inference," arXiv:2211.05102, 2022. [Online]. Available: https://arxiv.org/abs/2211.05102
[14] R. Srinivasaraghavan, "Intelligent Inference Endpoint Scheduling for Heterogeneous Large Language Model Serving," International Journal of Computer Information Systems and Industrial Management Applications, 2026. doi: 10.70917/ijcisim-2026-4862.










