Hybrid LLM and Knowledge Graph Framework for Explainable Healthcare Claims Adjudication

Authors

  • Selvakumar Kalyanasundaram Independent Researcher, Texas, USA. Author

DOI:

https://doi.org/10.63282/3050-9262.IJAIDSML-V7I3P113

Keywords:

Claims Adjudication, Explainable Artificial Intelligence (XAI), Healthcare Informatics, Knowledge Graphs, Large Language Models (Llms), Retrieval-Augmented Generation (RAG), Revenue Cycle Management

Abstract

Healthcare claims adjudication remains one of the most complex, high-stakes, and error-prone processes in modern health insurance operations. Manual review is costly, inconsistent, and inherently opaque. Automated rule-based engines improve throughput but lack adaptability and generate minimal audit-ready reasoning. This paper presents KLARA (Knowledge-augmented LLM Adjudication and Reasoning Architecture), a novel hybrid framework that tightly integrates large language models (LLMs) with a medical knowledge graph (MedKG) to adjudicate insurance claims with high accuracy, clinical fidelity, and full explainability. KLARA encodes clinical coding standards (ICD-11, CPT-4, HCPCS Level II), payer-specific benefit policies, and clinical evidence pathways as a richly interconnected graph. At inference time, a retrieval-augmented generation (RAG) pipeline grounds the LLM’s adjudication decisions in verified, authoritative graph nodes, preventing hallucination while preserving the model’s natural-language reasoning capacity. A chain-of-thought (CoT) prompting strategy elicits structured, step-by-step justification for every adjudication decision. Extensive experiments on three benchmarked datasets CMS synthetic claims, a de-identified commercial payer corpus (N = 124,000), and a publicly released Medicare Part B sample demonstrate that KLARA achieves 94.7% overall adjudication accuracy, a 38% reduction in false denial rate relative to the best competing baseline, and an explainability fidelity score of 0.91 on the SHAP-aligned XAI metric. Clinician reviewers rated 89.4% of KLARA’s natural-language justifications as “clinically sufficient.” These results establish KLARA as a significant advance toward trustworthy, auditable AI in healthcare revenue cycle management.

References

[1] CMS Office of the Actuary, “National Health Expenditure Accounts: Methodology Paper, 2023,” Centers for Medicare & Medicaid Services, Baltimore, MD, Tech. Rep., 2024.

[2] R. L. Himmelstein, M. Collins, and D. U. Woolhandler, “Health insurance claim denials and mortality: A retrospective cohort analysis,” JAMA Internal Medicine, vol. 183, no. 9, pp. 912-921, 2023.

[3] Office of Inspector General, “Medicare and Medicaid Programs: Improper Payments Report FY 2023,” U.S. Dept. of Health and Human Services, Washington, D.C., 2024.

[4] Council for Affordable Quality Healthcare (CAQH), “Index: Closing the Gap,” CAQH, Washington, D.C., Tech. Rep., 2023.

[5] E. J. Topol, “High-performance medicine: The convergence of human and artificial intelligence,” Nature Medicine, vol. 25, no. 1, pp. 44-56, 2019.

[6] A. Rajpurkar, E. Chen, O. Banerjee, and E. J. Topol, “AI in health and medicine,” Nature Medicine, vol. 28, no. 1, pp. 31-38, 2022.

[7] H. Chen, A. Sultan, Y. Tian, M. Chen, and S. Skiena, “Fast and accurate network embeddings via very sparse random projection,” in Proc. ACM CIKM, 2019, pp. 399-408.

[8] S. Bhatt, R. Das, and V. Kulkarni, “Scalability and maintainability challenges in insurance rule engines,” IEEE Trans. Services Computing, vol. 15, no. 3, pp. 1204-1216, 2022.

[9] M. Johnson and A. Patel, “Rule engine error propagation in multi-morbid claims adjudication,” Journal of Medical Systems, vol. 46, no. 8, p. 52, 2022.

[10] S. Bauder, T. M. Khoshgoftaar, and A. Richter, “Medicare fraud detection using machine learning methods,” in Proc. IEEE ICMLA, 2017, pp. 858-865.

[11] P. Hernandez, D. Kim, and J. Lee, “Graph neural networks for healthcare fraud detection,” IEEE Trans. Neural Networks Learning Syst., vol. 34, no. 7, pp. 3712-3726, 2023.

[12] T. Wang, R. Goyal, and S. Chaudhary, “Predicting prior authorization denials using ensemble learning on administrative claims,” Health Informatics Journal, vol. 29, no. 1, 2023.

[13] Z. Obermeyer, B. Powers, C. Vogeli, and S. Mullainathan, “Dissecting racial bias in an algorithm used to manage the health of populations,” Science, vol. 366, no. 6464, pp. 447-453, 2019.

[14] K. Choi, C. T. Bahadori, J. Searles, J. Coffman, M. Thompson, J. Bost, J. Tejedor-Sojo, and J. Sun, “Medical concept representation learning from electronic health records,” arXiv preprint arXiv:1602.03686, 2016.

[15] J. Lee, W. Yoon, S. Kim, D. Kim, S. Kim, C. Ho So, and J. Kang, “BioBERT: A pre-trained biomedical language representation model for biomedical text mining,” Bioinformatics, vol. 36, no. 4, pp. 1234-1240, 2020.

[16] E. Alsentzer, J. Murphy, W. Boag, W.-H. Weng, D. Jin, T. Naumann, and M. McDermott, “Publicly available clinical BERT embeddings,” in Proc. NAACL Workshop on Clin. NLP, 2019.

[17] X. Yang, Y. Chen, N. Peng, and M. Guo, “GatorTron: A large clinical language model to unlock patient information from unstructured electronic health records,” npj Digital Medicine, vol. 5, no. 1, p. 185, 2022.

[18] K. Singhal et al., “Towards expert-level medical question answering with large language models,” Nature Medicine, vol. 29, pp. 1958-1966, 2023.

[19] P. Nori, N. King, S. M. McKinney, D. Carignan, and E. Horvitz, “Capabilities of GPT-4 on medical challenge problems,” arXiv preprint arXiv:2303.13375, 2023.

[20] R. Chen, A. Bhattacharya, and D. Greenwald, “GPT-4 performance on prior authorization request classification,” npj Digital Medicine, vol. 7, no. 1, pp. 1-8, 2024.

[21] T. Nakamura, L. Park, and S. Rao, “LLM-generated denial letter quality in healthcare claims: A factual grounding audit,” JAMIA, vol. 31, no. 3, pp. 614-622, 2024.

[22] O. Bodenreider, “The Unified Medical Language System (UMLS): Integrating biomedical terminology,” Nucleic Acids Research, vol. 32, pp. D267-D270, 2004.

[23] The SNOMED International, “SNOMED CT: SNOMED Clinical Terms User Guide,” SNOMED International, London, Tech. Rep., 2024.

[24] O. Chandak, K. Huang, and M. Zitnik, “Building a knowledge graph to enable precision medicine,” Scientific Data, vol. 10, no. 1, p. 67, 2023.

[25] V. Ioannidis, X. Song, S. Manchanda, N. Li, X. Pan, D. Zheng, X. Ning, X. Zeng, and G. Karypis, “DRKG - Drug Repurposing Knowledge Graph,” arXiv preprint arXiv:2010.09124, 2020.

[26] B. Smith et al., “The OBO Foundry: Coordinated evolution of ontologies to support biomedical data integration,” Nature Biotechnology, vol. 25, no. 11, pp. 1251-1255, 2007.

[27] M. A. Rotmensch, Y. Halpern, A. Tlimat, S. Horng, and D. Sontag, “Learning a health knowledge graph from electronic medical records,” Scientific Reports, vol. 7, no. 1, pp. 1-11, 2017.

[28] Y. Zhang, R. Chen, J. Tang, W. S. Stewart, and J. Sun, “KDDI: Drug-drug interaction prediction based on knowledge graph embedding,” in Proc. IEEE ICDM, 2017, pp. 1105-1110.

[29] M. Zitnik, F. Li, A. Leskovec, and J. Leskovec, “Curating a COVID-19 data repository and forecasting county-level death counts,” Harvard Data Science Review, 2020.

[30] P. Lewis et al., “Retrieval-augmented generation for knowledge-intensive NLP tasks,” in Proc. NeurIPS, vol. 33, 2020, pp. 9459-9474.

[31] M. Yasunaga, H. Ren, A. Bosselut, P. Liang, and J. Leskovec, “QA-GNN: Reasoning with language models and knowledge graphs for question answering,” in Proc. NAACL, 2021, pp. 535-546.

[32] Y. Luo, S. Li, and T. Chen, “DRAGON: Deep bidirectional language-knowledge graph pretraining,” in Proc. NeurIPS, vol. 35, 2022.

[33] S. Wachter, B. Mittelstadt, and C. Russell, “Counterfactual explanations without opening the black box: Automated decisions and the GDPR,” Harvard Journal of Law & Technology, vol. 31, no. 2, pp. 841-887, 2018.

[34] CMS, “Medicare and Medicaid Programs: Prior Authorization Process and Interoperability,” Federal Register, vol. 89, no. 9, pp. 8758-8954, Jan. 2024.

[35] S. M. Lundberg and S.-I. Lee, “A unified approach to interpreting model predictions,” in Proc. NeurIPS, vol. 30, 2017.

[36] M. T. Ribeiro, S. Singh, and C. Guestrin, “Why should I trust you?: Explaining the predictions of any classifier,” in Proc. ACM SIGKDD, 2016, pp. 1135-1144.

[37] M. Sundararajan, A. Taly, and Q. Yan, “Axiomatic attribution for deep networks,” in Proc. ICML, vol. 70, 2017, pp. 3319-3328.

[38] M. Yasunaga, A. Leskovec, and P. Liang, “LinkBERT: Pretraining language models with document links,” in Proc. ACL, 2022, pp. 8003-8016.

Published

2026-08-10

Issue

Section

Articles

How to Cite

1.
Kalyanasundaram S. Hybrid LLM and Knowledge Graph Framework for Explainable Healthcare Claims Adjudication. IJAIDSML [Internet]. 2026 Aug. 10 [cited 2026 Sep. 14];7(3):112-9. Available from: https://ijaidsml.org/index.php/ijaidsml/article/view/646