Independent Researcher, Dallas, TX, USA
Official Article Landing Page
A Lightweight Explainable Framework for Schema Matching and Data Fusion in Low-Code Data Engineering Pipelines
UnivColl International Multidisciplinary Research Journal International Open Access, Peer-reviewed, Refereed Journal
Google Scholar indexing depends on Google Scholar’s crawl schedule. DOI verification opens the official DOI resolver.
Official DOI Landing Verification
This page is the official article landing page for the DOI and published Version of Record.
Authors & Affiliations
Abstract
Low-code data engineering environments increasingly support rapid data ingestion, transformation, and integration, yet schema matching remains a persistent challenge when heterogeneous datasets use inconsistent naming conventions, limited metadata, and weak documentation. This study proposes a lightweight explainable framework for schema matching and data fusion in low-code data engineering pipelines. The framework integrates sentence embedding similarity, fuzzy string matching, keyword-based heuristics, and data-type compatibility validation to identify candidate correspondences between independently structured tabular schemas without requiring labeled training data, predefined ontologies, or extensive manual mapping. The approach is evaluated using publicly available Airbnb Listings and Reviews datasets, where the Listings file contains property-level attributes and the Reviews file contains review-event attributes, creating a practical case for schema alignment across related but differently structured data sources. In addition to hybrid similarity scoring, the framework includes an optional large language model based justification layer that generates natural-language explanations for predicted column matches, improving interpretability for analysts, data engineers, and low-code users. The experimental results show that a moderate similarity threshold of 0.65 achieved a precision of 0.50, recall of 0.33, and F1-score of 0.40, while a stricter threshold of 0.80 achieved a precision of 1.00, recall of 0.25, and F1-score of 0.40. These findings demonstrate the trade-off between broader mapping coverage and high-confidence matching, while confirming the value of explainable scoring for low-supervision schema alignment. The proposed framework contributes a domain-agnostic and low-footprint method for preparing heterogeneous data sources for downstream data fusion, ETL automation, and analytical workflows in low-code data engineering settings.
Keywords & Indexing Terms
Download Center
Indexing & Verification
Article Metrics
Metrics are indicative and may update periodically after indexing, downloads and citation tracking are enabled.
Article Integrity & Transparency
License & Copyright
Copyright © 2025 UnivColl International Multidisciplinary Research Journal. This work is licensed under a Creative Commons Attribution 4.0 International License . Authors retain copyright and grant the journal the right of first publication.
How to Cite
Manoj Parasa (2025). A Lightweight Explainable Framework for Schema Matching and Data Fusion in Low-Code Data Engineering Pipelines. UnivColl International Multidisciplinary Research Journal, 1(3), 1-25. https://doi.org/10.65919/uimrj.2025.v1i3001
PDF Preview
If the PDF preview does not load on your device, open the PDF directly or use the Download PDF button. Open PDF in new tab
Declarations
Funding
No external funding information has been declared unless stated in the published PDF.
Conflict of Interest
The authors declare no conflict of interest unless otherwise stated in the article.
Ethical Approval
Ethical approval status is as per the article and journal policy.
Data Availability
Data availability is as declared by the author(s) in the published article.
Author Contributions
Author contributions are recorded as per submitted manuscript and editorial records.
AI-use Declaration
AI-use declaration is governed by journal policy and author disclosure.
Publisher's Note
- âś“ The views, opinions and conclusions expressed in this article are solely those of the author(s).
- âś“ Publication of this article does not imply endorsement by UIMRJ, the editorial board or the publisher.
- âś“ Responsibility for the accuracy, originality and integrity of the work remains with the author(s).
- âś“ Readers are encouraged to independently evaluate and verify the information before application or citation.
- âś“ UIMRJ and UnivColl Publications shall not be held liable for any consequences arising from the use of the published content.
References
Showing first 3 references. Click “Show All References” to view complete list.
- [1] Rahm, E., & Bernstein, P. A. (2001). A survey of approaches to automatic schema matching. The VLDB Journal, 10(4), 334–350. https://doi.org/10.1007/s007780100057
- [2] Doan, A., Domingos, P., & Halevy, A. Y. (2001). Reconciling schemas of disparate data sources: A machine-learning approach. Proceedings of the ACM SIGMOD International Conference on Management of Data, 509–520. https://doi.org/10.1145/375663.375731
- [3] Melnik, S., Garcia-Molina, H., & Rahm, E. (2002). Similarity flooding: A versatile graph matching algorithm and its application to schema matching. Proceedings of the 18th International Conference on Data Engineering, 117–128. https://doi.org/10.1109/ICDE.2002.994702
- [4] Giunchiglia, F., & Shvaiko, P. (2004). Semantic matching. The Knowledge Engineering Review, 18(3), 265–280. https://doi.org/10.1017/S0269888904000074
- [5] Aumueller, D., Do, H. H., Massmann, S., & Rahm, E. (2005). Schema and ontology matching with COMA++. Proceedings of the ACM SIGMOD International Conference on Management of Data, 906–908. https://doi.org/10.1145/1066157.1066283
- [6] Bilke, A., & Naumann, F. (2005). Schema matching using duplicates. Proceedings of the 21st International Conference on Data Engineering, 69–80. https://doi.org/10.1109/ICDE.2005.126
- [7] Halevy, A. Y., Ashish, N., Bitton, D., Carey, M. J., Draper, D., Pollock, J., Rosenthal, A., & Sikka, V. (2005). Enterprise information integration: Successes, challenges and controversies. Proceedings of the ACM SIGMOD International Conference on Management of Data, 778–787. https://doi.org/10.1145/1066157.1066246
- [8] Do, H. H., & Rahm, E. (2007). Matching large schemas: Approaches and evaluation. Information Systems, 32(6), 857–885. https://doi.org/10.1016/j.is.2006.09.002
- [9] Elmagarmid, A. K., Ipeirotis, P. G., & Verykios, V. S. (2007). Duplicate record detection: A survey. IEEE Transactions on Knowledge and Data Engineering, 19(1), 1–16. https://doi.org/10.1109/TKDE.2007.250581
- [10] Bleiholder, J., & Naumann, F. (2008). Data fusion. ACM Computing Surveys, 41(1), 1–41. https://doi.org/10.1145/1456650.1456651
- [11] Bernstein, P. A., Madhavan, J., & Rahm, E. (2011). Generic schema matching, ten years later. Proceedings of the VLDB Endowment, 4(11), 695–701. https://doi.org/10.14778/3402707.3402710
- [12] Shvaiko, P., & Euzenat, J. (2013). Ontology matching: State of the art and future challenges. IEEE Transactions on Knowledge and Data Engineering, 25(1), 158–176. https://doi.org/10.1109/TKDE.2011.253
- [13] Cer, D., Yang, Y., Kong, S. Y., Hua, N., Limtiaco, N., St. John, R., Constant, N., Guajardo-Cespedes, M., Yuan, S., Tar, C., Strope, B., & Kurzweil, R. (2018). Universal Sentence Encoder for English. Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, 169–174. https://doi.org/10.18653/v1/D18-2029
- [14] Konda, P., Das, S., Suganthan, P. G. C., Doan, A., Ardalan, A., Ballard, J. R., Li, H., Panahi, F., Zhang, H., Naughton, J. F., Prasad, S., Krishnan, G., Deep, R., & Raghavendra, V. (2016). Magellan: Toward building entity matching management systems over data science stacks. Proceedings of the VLDB Endowment, 9(13), 1581–1584. https://doi.org/10.14778/3007263.3007314
- [15] Mudgal, S., Li, H., Rekatsinas, T., Doan, A., Park, Y., Krishnan, G., Deep, R., Arcaute, E., & Raghavendra, V. (2018). Deep learning for entity matching: A design space exploration. Proceedings of the ACM SIGMOD International Conference on Management of Data, 19–34. https://doi.org/10.1145/3183713.3196926
- [16] Guidotti, R., Monreale, A., Ruggieri, S., Turini, F., Giannotti, F., & Pedreschi, D. (2018). A survey of methods for explaining black box models. ACM Computing Surveys, 51(5), 1–42. https://doi.org/10.1145/3236009
- [17] Reimers, N., & Gurevych, I. (2019). Sentence-BERT: Sentence embeddings using Siamese BERT-networks. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing, 3982–3992. https://doi.org/10.18653/v1/D19-1410
- [18] Sahay, A., Indamutsa, A., Di Ruscio, D., & Pierantonio, A. (2020). Supporting the understanding and comparison of low-code development platforms. Proceedings of the 46th Euromicro Conference on Software Engineering and Advanced Applications, 171–178. https://doi.org/10.1109/SEAA51224.2020.00036
- [19] Barredo Arrieta, A., DĂaz-RodrĂguez, N., Del Ser, J., Bennetot, A., Tabik, S., Barbado, A., GarcĂa, S., Gil-LĂłpez, S., Molina, D., Benjamins, R., Chatila, R., & Herrera, F. (2020). Explainable artificial intelligence: Concepts, taxonomies, opportunities and challenges toward responsible AI. Information Fusion, 58, 82–115. https://doi.org/10.1016/j.inffus.2019.12.012
- [20] Luo, Y., Liang, P., Wang, C., Shahin, M., & Zhan, J. (2021). Characteristics and challenges of low-code development: The practitioners’ perspective. Proceedings of the ACM/IEEE International Symposium on Empirical Software Engineering and Measurement, 1–11. https://doi.org/10.1145/3475716.347578
Related Articles
Publish Your Research With UIMRJ
Submit your original research article, review paper, case study or conceptual paper to an international open access, peer-reviewed and refereed journal.