ML-Powered Data Catalogs and Metadata Inference for Self-Describing Big Data Ecosystems

Authors

  • Sivadeep Katangoori Solutions Architect at Metanoia Solutions Inc., USA. Author

DOI:

https://doi.org/10.63282/3050-9262.IJAIDSML-V6I2P127

Keywords:

Ml-Powered Data Catalogs, Metadata Inference, Big Data Ecosystems, Self-Describing Data, Schema Detection, Data Discovery, Data Governance, Machine Learning, Data Lineage, Semantic Tagging, Automated Metadata Management

Abstract

In today's fast-changing data world, businesses have too much data that is too different and spread out. Standard methods for managing data don't always work well for finding, understanding, and applying data. Modern data catalogues are more than simply static lists of items; they are also dynamic places to store information that provide structure, meaning, and easy access to large data ecosystems. These catalogues can really shine when ML is applied to figure things out smartly, even when the data sets are too messy, not well-documented, or not well-organized. Data catalogues that utilize ML look at these current data trends, user behaviors & contextual signals to automatically tag, categorize & link their data assets. This makes the data self-descriptive and makes governance easier. This solves a big problem: creating information by hand is hard work, prone to their mistakes, and typically doesn't work well with different teams & also technologies. Our methodology utilizes both supervised and unsupervised ML techniques to infer schema details, data lineage, semantic links, and use contexts, therefore reducing dependence on human input while improving their data discoverability & the dependability. This paper suggests a scalable architecture that integrates ML models directly into the cataloguing process, enabling actual time metadata enhancement as data is ingested into the system. Additionally, we have a feedback mechanism that constantly improves the quality of inferences via user interactions & validations. Initial findings indicate significant improvements in data discoverability, accelerated onboarding for data users, and improved compliance readiness, resulting from enhanced metadata completeness and traceability. This work shows how data catalogues that use machine learning can change from simple documentation tools into these smart systems that make self-service analytics easier, promote data democracy & support truly self-describing big data platforms. This makes organizations more agile and focused on getting insights.

References

[1] Chard, K., D'Arcy, M., Heavner, B., Foster, I., Kesselman, C., Madduri, R., ... & Toga, A. (2016, December). I'll take that to go: Big data bags and minimal identifiers for exchange of large, complex datasets. In 2016 Ieee international conference on big data (big data) (pp. 319-328). IEEE.

[2] Nagorny, K., Scholze, S., Ruhl, M., & Colombo, A. W. (2018, May). Semantical support for a CPS data marketplace to prepare Big Data analytics in smart manufacturing environments. In 2018 IEEE Industrial Cyber-Physical Systems (ICPS) (pp. 206-211). IEEE.

[3] Suryadevara, S. S. K., & Shaik, K. (2023). Real-Time Anomaly Detection and Attack Mitigation for Cloud-Based Content Delivery Paths Using AI. International Journal of Emerging Research in Engineering and Technology, 4(1), 175-185. https://doi.org/10.63282/3050-922X.IJERET-V4I1P119

[4] Kumar Doodala, A. N. (2024). Validating UX consistency Across Omnichannel Platform. American International Journal of Computer Science and Technology, 6(6), 87-97. https://doi.org/10.63282/3117-5481/AIJCST-V6I6P109

[5] Lokers, R., Van Randen, Y., Knapen, R., Gaubitzer, S., Zudin, S., & Janssen, S. (2015, September). Improving access to big data in agriculture and forestry using semantic technologies. In Research Conference on Metadata and Semantics Research (pp. 369-380). Cham: Springer International Publishing.

[6] Vppalapati, M. (2024). Power-Bound Storage Design: Architecting Systems for Electrical Scarcity. International Journal of AI, BigData, Computational and Management Studies, 5(1), 208-217. https://doi.org/10.63282/3050-9416.IJAIBDCMS-V5I1P121

[7] Khalifa, S., Elshater, Y., Sundaravarathan, K., Bhat, A., Martin, P., Imam, F., ... & Statchuk, C. (2016). The six pillars for building big data analytics ecosystems. ACM Computing Surveys (CSUR), 49(2), 1-36.

[8] Parakala, A. (2023). Citizen-Facing Automation: Chatbots and Self-Service in Public Services. International Journal of AI, BigData, Computational and Management Studies, 4(4), 108-118. https://doi.org/10.63282/3050-9416.IJAIBDCMS-V4I4P112

[9] Gaddam, R. R., & Krishna, K. (2023). KFP v2 Artifact-Centric ML Pipeline Governance. International Journal of Artificial Intelligence, Data Science, and Machine Learning, 4(2), 142-153. https://doi.org/10.63282/3050-9262.IJAIDSML-V4I2P116

[10] Eynard-Bontemps, G., Abernathey, R., Hamman, J., Ponte, A., & Rath, W. (2019, January). The PANGEO Big Data Ecosystem and its use at CNES. In Big Data from Space (BiDS'19).... Turning Data into insights... 19-21 fébruary 2019, Munich, Germany.

[11] Muppaneni, R. K. (2022). Data Privacy in the Age of AI: How Dynamics 365 Handles Regulatory Challenges. International Journal of Artificial Intelligence, Data Science, and Machine Learning, 3(4), 159-170. https://doi.org/10.63282/3050-9262.IJAIDSML-V3I4P117

[12] Kumar Doodala, A. N. (2024). Service Virtualization for API-First development: A shift-Left Testing Strategy. American International Journal of Computer Science and Technology, 6(4), 50-58. https://doi.org/10.63282/3117-5481/AIJCST-V6I4P105

[13] Ghiringhelli, L. M., Carbogno, C., Levchenko, S., Mohamed, F., Huhs, G., Lüders, M., ... & Scheffler, M. (2017). Towards efficient data exchange and sharing for big-data driven materials science: metadata and data formats. npj computational materials, 3(1), 46.

[14] Suryadevara, S. S. K., & Nakirikanti, S. (2023). Privacy-Preserving Personalization Using Federated Learning in AEM . International Journal of AI, BigData, Computational and Management Studies, 4(4), 190-199. https://doi.org/10.63282/3050-9416.IJAIBDCMS-V4I4P119

[15] Srigadde, B. R., & Devaraju, J. M. (2024). Building a Reusable AI Connection Utility Class. International Journal of Emerging Research in Engineering and Technology, 5(2), 188-200. https://doi.org/10.63282/3050-922X.IJERET-V5I2P119

[16] Berman, J. J. (2013). Principles of big data: preparing, sharing, and analyzing complex information. Newnes.

[17] Shiramalla, R. (2024). Secure Multi-Cloud API Orchestration between Salesforce, Oracle CPQ, and Azure. American International Journal of Computer Science and Technology, 6(3), 102-113. https://doi.org/10.63282/3117-5481/AIJCST-V6I3P108

[18] Vppalapati, M. (2024). Cooling Domains as First-Class Failure Boundaries in Storage Architecture. American International Journal of Computer Science and Technology, 6(2), 96-106. https://doi.org/10.63282/3117-5481/AIJCST-V6I2P110

[19] Curry, E. (2020). Real-time linked dataspaces: Enabling data ecosystems for intelligent systems. Springer Nature.

[20] Takkalapally, D. (2024). ShiftLeft-AI: Machine Learning Framework for Proactive Performance Assurance in CI/CD Pipelines. International Journal of Artificial Intelligence, Data Science, and Machine Learning, 5(4), 285-296. https://doi.org/10.63282/3050-9262.IJAIDSML-V5I4P126

[21] Verma, D., Overton, J., Wright, B., Purcell, J., Santhar, S., Das, A., ... & Prasad, S. (2022, August). Self-describing digital assets and their applications in an integrated science and engineering ecosystem. In Smoky Mountains Computational Sciences and Engineering Conference (pp. 274-287). Cham: Springer Nature Switzerland.

[22] Allenki, S. S. (2023). Reducing Security Vulnerabilities with Encryption, IAM, and Regular Audits. International Journal of Emerging Trends in Computer Science and Information Technology, 4(1), 265-275. https://doi.org/10.63282/3050-9246.IJETCSIT-V4I1P127

[23] Srigadde, B. R. (2024). Agents, LLMs, and Salesforce with Multi-Cloud Provider (MCP). International Journal of Artificial Intelligence, Data Science, and Machine Learning, 5(3), 277-288. https://doi.org/10.63282/3050-9262.IJAIDSML-V5I3P127

[24] Amirian, P., van Loggerenberg, F., & Lang, T. (2017). Big data and big data technologies. In Big Data in Healthcare: Extracting Knowledge from Point-of-Care Machines (pp. 39-58). Cham: Springer International Publishing.

[25] Shiramalla, R. (2023). Optimizing Cross-Platform Enterprise Integrations Using Workato: A Case Study of Salesforce and Oracle SaaS Applications. International Journal of Emerging Trends in Computer Science and Information Technology, 4(1), 232-243. https://doi.org/10.63282/3050-9246.IJETCSIT-V4I1P124

[26] Muppaneni, K. (2024). Progressive Web Apps: Offline UX Benchmarking. International Journal of Emerging Trends in Computer Science and Information Technology, 5(2), 174-183. https://doi.org/10.63282/3050-9246.IJETCSIT-V5I2P119

[27] Nagarjuna, D. N., & Yogesh, N. (2015). A Survey on Hadoop Architecture & its Ecosystem to Process Big Data-Real World Hadoop Use Cases.

[28] Takkalapally, D., & Takkellapally, M. R. (2024). AI-SynPerf: Synthetic Data Intelligence Framework for 5G Mobile Performance Simulation. International Journal of Emerging Trends in Computer Science and Information Technology, 5(1), 182-194. https://doi.org/10.63282/3050-9246.IJETCSIT-V5I1P118

[29] Gaddam, R. R. (2023). Progressive Delivery for Models with Quality KPIs. American International Journal of Computer Science and Technology, 5(4), 33-47. https://doi.org/10.63282/3117-5481/AIJCST-V5I4P104

[30] Niu, C., Zhang, W., Byna, S., & Chen, Y. (2023, December). PSQS: Parallel Semantic Querying Service for Self-describing File Formats. In 2023 IEEE International Conference on Big Data (BigData) (pp. 536-541). IEEE.

[31] Muppaneni, R. K. (2022). From Legacy ERP to Cloud-First: A Transformation Story with Dynamics 365. International Journal of Emerging Research in Engineering and Technology, 3(4), 153-164. https://doi.org/10.63282/3050-922X.IJERET-V3I4P117

[32] Parakala, A. (2023). Vendor Highlights – IoT, AI, and Process Mining. International Journal of Emerging Trends in Computer Science and Information Technology, 4(4), 135-146. https://doi.org/10.63282/3050-9246.IJETCSIT-V4I4P115

[33] Mehta, S., & Mehta, V. (2016). Hadoop ecosystem: An introduction. International Journal of Science and Research (IJSR), 5(6), 557-562.

[34] Muppaneni, K., & Palem, V. (2024). Micro-Frontend Design Patterns for Multi-Framework Applications. International Journal of Emerging Research in Engineering and Technology, 5(3), 181-190. https://doi.org/10.63282/3050-922X.IJERET-V5I3P120

[35] Allenki, S. S. (2023). Applying Cloud Security Best Practices in Regulated Environments. American International Journal of Computer Science and Technology, 5(3), 48-60. https://doi.org/10.63282/3117-5481/AIJCST-V5I3P105

[36] Zender, C. S. (2008). Analysis of self-describing gridded geoscience data with netCDF Operators (NCO). Environmental Modelling & Software, 23(10-11), 1338-1342.

[37] Lukowski, M., Prokhorenkov, A., & Grossman, R. L. (2023). Towards self-describing and FAIR bulk formats for biomedical data. PLOS Computational Biology, 19(3), e1010944.

[38] Taluri, R. (2024). A Hybrid Data Engineering and Generative AI Architecture for Intelligent Data Governance, Metadata Management, and Automated Data Quality Assessment on AWS. International Journal of Artificial Intelligence, Data Science, and Machine Learning, 5(2), 230-240. https://doi.org/10.63282/3050-9262.IJAIDSML-V5I2P126

Published

2025-06-11

Issue

Section

Articles

How to Cite

1.
Katangoori S. ML-Powered Data Catalogs and Metadata Inference for Self-Describing Big Data Ecosystems. IJAIDSML [Internet]. 2025 Jun. 11 [cited 2026 Jul. 25];6(2):234-43. Available from: https://ijaidsml.org/index.php/ijaidsml/article/view/631