Benchmarking Hybrid On-Device and Cloud AI Inference for Android and Embedded Devices: Latency, Memory, Energy, Privacy, and Accuracy Trade-offs

Authors

  • Srikanth Puram Independent Researcher, Mobile and Embedded Software, Architecture Novi, Michigan, USA. Author

DOI:

https://doi.org/10.63282/3050-9262.IJAIDSML-V6I2P126

Keywords:

Android AI, Embedded AI, On-Device Inference, Cloud Inference, Hybrid AI, LiteRT, TensorFlow Lite, Gemini Nano, Gemma, ONNX Runtime Mobile, Latency Benchmarking, Edge Intelligence

Abstract

Android phones, automotive displays, wearables, home devices, retail terminals, industrial gateways, and other embedded gadgets increasingly run artificial intelligence workloads that were previously hosted almost exclusively in the cloud. On-device inference improves privacy, offline behavior, and round-trip latency, but it is constrained by model size, memory, battery, thermal limits, accelerator availability, and fragmented hardware. Cloud inference supports larger models and centralized updates, but it introduces network latency, availability risk, data-transfer cost, privacy exposure, and service dependency. This paper proposes a benchmarking and decision framework for hybrid on-device and cloud AI inference on Android and embedded devices. The framework evaluates five deployment modes: fully on-device inference, cloud-only inference, local prefiltering with cloud escalation, confidence-based hybrid routing, and cloud fallback after local failure. It defines workload classes, device classes, network profiles, privacy levels, and metrics including p50/p95 latency, time-to-first-token, throughput, peak RAM, model size, energy per inference, quantization accuracy loss, network round-trip delay, and privacy exposure. The paper also presents a routing algorithm that selects the inference location based on latency budget, privacy class, model memory footprint, battery state, network quality, and confidence threshold. The proposed framework is aligned with the technology landscape available by April 2025, including LiteRT/TensorFlow Lite foundations, Android on-device AI APIs, Gemini Nano experimental access, Gemma 3 class open models, ONNX Runtime Mobile, and MLPerf-style inference benchmarking.

References

[1] Google, “TensorFlow Lite is now LiteRT,” Google Developers Blog, Sep. 2024.

[2] Google, “LiteRT: High-Performance On-Device Machine Learning,” Google AI Edge Documentation, 2024.

[3] Android Developers Blog, “Gemini Nano experimental access available on Android,” Oct. 2024.

[4] Google AI for Developers, “Gemma releases: Gemma 3 model card and release notes,” Mar. 2025.

[5] ONNX Runtime, “ONNX Runtime Mobile Documentation,” Microsoft, 2024.

[6] Android Developers, “Neural Networks API,” Android NDK Documentation, 2024.

[7] R. David et al., “TensorFlow Lite Micro: Embedded Machine Learning on TinyML Systems,” arXiv:2010.08678, 2020.

[8] C. Banbury et al., “MLPerf Tiny Benchmark,” arXiv:2106.07597, 2021.

[9] V. J. Reddi et al., “MLPerf Inference Benchmark,” arXiv:1911.02549, 2019.

[10] J. Lee et al., “On-Device Neural Net Inference with Mobile GPUs,” arXiv:1907.01989, 2019.

[11] M. Tan and Q. Le, “EfficientNet: Rethinking Model Scaling for Convolutional Neural Networks,” ICML, 2019.

[12] A. Howard et al., “Searching for MobileNetV3,” ICCV, 2019.

[13] V. Sanh, L. Debut, J. Chaumond, and T. Wolf, “DistilBERT, a distilled version of BERT,” NeurIPS Workshop, 2019.

[14] Z. Sun et al., “MobileBERT: a Compact Task-Agnostic BERT for Resource-Limited Devices,” ACL, 2020.

[15] National Institute of Standards and Technology, “Artificial Intelligence Risk Management Framework (AI RMF 1.0),” NIST AI 100-1, Jan. 2023.

[16] OWASP Foundation, “Mobile Application Security Verification Standard,” 2024.

[17] E. Rescorla, “The Transport Layer Security (TLS) Protocol Version 1.3,” RFC 8446, Aug. 2018.

[18] M. Abadi et al., “TensorFlow: A System for Large-Scale Machine Learning,” OSDI, 2016.

[19] A. Paszke et al., “PyTorch: An Imperative Style, High-Performance Deep Learning Library,” NeurIPS, 2019.

[20] National Institute of Standards and Technology, “Secure Software Development Framework (SSDF) Version 1.1,” NIST SP 800-218, Feb. 2022.

Published

2025-06-10

Issue

Section

Articles

How to Cite

1.
Puram S. Benchmarking Hybrid On-Device and Cloud AI Inference for Android and Embedded Devices: Latency, Memory, Energy, Privacy, and Accuracy Trade-offs. IJAIDSML [Internet]. 2025 Jun. 10 [cited 2026 Jul. 24];6(2):229-33. Available from: https://ijaidsml.org/index.php/ijaidsml/article/view/621