Open Access Peer Reviewed COPE Aligned DOI Registered Version of Record

Official Article Landing Page

A Lightweight Explainable Framework for Schema Matching and Data Fusion in Low-Code Data Engineering Pipelines

UnivColl International Multidisciplinary Research Journal International Open Access, Peer-reviewed, Refereed Journal

Google Scholar indexing depends on Google Scholar’s crawl schedule. Verify DOI opens the external DOI resolver.

Abstract

Low-code data engineering environments increasingly support rapid data ingestion, transformation, and integration, yet schema matching remains a persistent challenge when heterogeneous datasets use inconsistent naming conventions, limited metadata, and weak documentation. This study proposes a lightweight explainable framework for schema matching and data fusion in low-code data engineering pipelines. The framework integrates sentence embedding similarity, fuzzy string matching, keyword-based heuristics, and data-type compatibility validation to identify candidate correspondences between independently structured tabular schemas without requiring labeled training data, predefined ontologies, or extensive manual mapping. The approach is evaluated using publicly available Airbnb Listings and Reviews datasets, where the Listings file contains property-level attributes and the Reviews file contains review-event attributes, creating a practical case for schema alignment across related but differently structured data sources. In addition to hybrid similarity scoring, the framework includes an optional large language model based justification layer that generates natural-language explanations for predicted column matches, improving interpretability for analysts, data engineers, and low-code users. The experimental results show that a moderate similarity threshold of 0.65 achieved a precision of 0.50, recall of 0.33, and F1-score of 0.40, while a stricter threshold of 0.80 achieved a precision of 1.00, recall of 0.25, and F1-score of 0.40. These findings demonstrate the trade-off between broader mapping coverage and high-confidence matching, while confirming the value of explainable scoring for low-supervision schema alignment. The proposed framework contributes a domain-agnostic and low-footprint method for preparing heterogeneous data sources for downstream data fusion, ETL automation, and analytical workflows in low-code data engineering settings.

Official DOI Landing Verification

This is the official published Version of Record. DOI status: Registered.

DOI StatusRegistered
PDF StatusAvailable
LicenseCC BY-NC 4.0
MetadataMachine-readable
VersionVersion of Record
Article IntegrityVerified
DOI Record StatusRegistered
Google Scholar StatusMetadata Ready
Peer ReviewCompleted

Article Metadata

Article IDUIMRJ-V1I3-024
Article TypeResearch Article
Volume1
Issue3
Pages1-25
Published31 December 2025
LanguageEnglish
ISSN3108-1460
DOI Prefix10.65919
PublisherUnivColl Publications
Access ModelOpen Access
Peer ReviewDouble-Blind
LicenseCC BY-NC 4.0
Article DOI10.65919/uimrj.2025.v1i3001

Authors & Affiliations

* Corresponding Author. Author contact details are available in the published PDF article as per journal policy.

Received 07 November 2025
Accepted 10 December 2025
Published 31 December 2025

Keywords & Indexing Terms

Download Center

Indexing & Verification

Article Metrics

Article Views59
PDF Downloads9
Total Pages25

Metrics are indicative and may update periodically after indexing, downloads and citation tracking are enabled.

Article Integrity & Transparency

Peer ReviewDouble Blind
Plagiarism CheckCompleted
COPE ComplianceFollowed
Correction StatusNo correction issued
Retraction StatusNot retracted
VersionVersion of Record

License & Copyright

CC BY-NC 4.0

© 2025 The Author(s). The author(s) retain copyright and grant the journal the right of first publication. This work is licensed under the Creative Commons Attribution-NonCommercial 4.0 International License .

How to Cite

Manoj Parasa (2025). A Lightweight Explainable Framework for Schema Matching and Data Fusion in Low-Code Data Engineering Pipelines. UnivColl International Multidisciplinary Research Journal, 1(3), 1-25. https://doi.org/10.65919/uimrj.2025.v1i3001

PDF Preview

If the PDF preview does not load on your device, open the PDF directly or use the Download PDF button. Open PDF in new tab

Declarations

Funding

No external funding information has been declared unless stated in the published PDF.

Conflict of Interest

The authors declare no conflict of interest unless otherwise stated in the article.

Ethical Approval

Ethical approval status is as per the article and journal policy.

Data Availability

Data availability is as declared by the author(s) in the published article.

Author Contributions

Author contributions are recorded as per submitted manuscript and editorial records.

AI-use Declaration

AI-use declaration is governed by journal policy and author disclosure.

Publisher's Note

  • ✓ The views, opinions and conclusions expressed in this article are solely those of the author(s).
  • ✓ Publication of this article does not imply endorsement by UIMRJ, the editorial board or the publisher.
  • ✓ Responsibility for the accuracy, originality and integrity of the work remains with the author(s).
  • ✓ Readers are encouraged to independently evaluate and verify the information before application or citation.
  • ✓ UIMRJ and UnivColl Publications shall not be held liable for any consequences arising from the use of the published content.

References

Showing first 3 references. Click “Show All References” to view complete list.

  1. [1] Rahm, E., & Bernstein, P. A. (2001). A survey of approaches to automatic schema matching. The VLDB Journal, 10(4), 334–350. https://doi.org/10.1007/s007780100057
  2. [2] Doan, A., Domingos, P., & Halevy, A. Y. (2001). Reconciling schemas of disparate data sources: A machine-learning approach. Proceedings of the ACM SIGMOD International Conference on Management of Data, 509–520. https://doi.org/10.1145/375663.375731
  3. [3] Melnik, S., Garcia-Molina, H., & Rahm, E. (2002). Similarity flooding: A versatile graph matching algorithm and its application to schema matching. Proceedings of the 18th International Conference on Data Engineering, 117–128. https://doi.org/10.1109/ICDE.2002.994702
  4. [4] Giunchiglia, F., & Shvaiko, P. (2004). Semantic matching. The Knowledge Engineering Review, 18(3), 265–280. https://doi.org/10.1017/S0269888904000074
  5. [5] Aumueller, D., Do, H. H., Massmann, S., & Rahm, E. (2005). Schema and ontology matching with COMA++. Proceedings of the ACM SIGMOD International Conference on Management of Data, 906–908. https://doi.org/10.1145/1066157.1066283
  6. [6] Bilke, A., & Naumann, F. (2005). Schema matching using duplicates. Proceedings of the 21st International Conference on Data Engineering, 69–80. https://doi.org/10.1109/ICDE.2005.126
  7. [7] Halevy, A. Y., Ashish, N., Bitton, D., Carey, M. J., Draper, D., Pollock, J., Rosenthal, A., & Sikka, V. (2005). Enterprise information integration: Successes, challenges and controversies. Proceedings of the ACM SIGMOD International Conference on Management of Data, 778–787. https://doi.org/10.1145/1066157.1066246
  8. [8] Do, H. H., & Rahm, E. (2007). Matching large schemas: Approaches and evaluation. Information Systems, 32(6), 857–885. https://doi.org/10.1016/j.is.2006.09.002
  9. [9] Elmagarmid, A. K., Ipeirotis, P. G., & Verykios, V. S. (2007). Duplicate record detection: A survey. IEEE Transactions on Knowledge and Data Engineering, 19(1), 1–16. https://doi.org/10.1109/TKDE.2007.250581
  10. [10] Bleiholder, J., & Naumann, F. (2008). Data fusion. ACM Computing Surveys, 41(1), 1–41. https://doi.org/10.1145/1456650.1456651
  11. [11] Bernstein, P. A., Madhavan, J., & Rahm, E. (2011). Generic schema matching, ten years later. Proceedings of the VLDB Endowment, 4(11), 695–701. https://doi.org/10.14778/3402707.3402710
  12. [12] Shvaiko, P., & Euzenat, J. (2013). Ontology matching: State of the art and future challenges. IEEE Transactions on Knowledge and Data Engineering, 25(1), 158–176. https://doi.org/10.1109/TKDE.2011.253
  13. [13] Cer, D., Yang, Y., Kong, S. Y., Hua, N., Limtiaco, N., St. John, R., Constant, N., Guajardo-Cespedes, M., Yuan, S., Tar, C., Strope, B., & Kurzweil, R. (2018). Universal Sentence Encoder for English. Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, 169–174. https://doi.org/10.18653/v1/D18-2029
  14. [14] Konda, P., Das, S., Suganthan, P. G. C., Doan, A., Ardalan, A., Ballard, J. R., Li, H., Panahi, F., Zhang, H., Naughton, J. F., Prasad, S., Krishnan, G., Deep, R., & Raghavendra, V. (2016). Magellan: Toward building entity matching management systems over data science stacks. Proceedings of the VLDB Endowment, 9(13), 1581–1584. https://doi.org/10.14778/3007263.3007314
  15. [15] Mudgal, S., Li, H., Rekatsinas, T., Doan, A., Park, Y., Krishnan, G., Deep, R., Arcaute, E., & Raghavendra, V. (2018). Deep learning for entity matching: A design space exploration. Proceedings of the ACM SIGMOD International Conference on Management of Data, 19–34. https://doi.org/10.1145/3183713.3196926
  16. [16] Guidotti, R., Monreale, A., Ruggieri, S., Turini, F., Giannotti, F., & Pedreschi, D. (2018). A survey of methods for explaining black box models. ACM Computing Surveys, 51(5), 1–42. https://doi.org/10.1145/3236009
  17. [17] Reimers, N., & Gurevych, I. (2019). Sentence-BERT: Sentence embeddings using Siamese BERT-networks. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing, 3982–3992. https://doi.org/10.18653/v1/D19-1410
  18. [18] Sahay, A., Indamutsa, A., Di Ruscio, D., & Pierantonio, A. (2020). Supporting the understanding and comparison of low-code development platforms. Proceedings of the 46th Euromicro Conference on Software Engineering and Advanced Applications, 171–178. https://doi.org/10.1109/SEAA51224.2020.00036
  19. [19] Barredo Arrieta, A., Díaz-Rodríguez, N., Del Ser, J., Bennetot, A., Tabik, S., Barbado, A., García, S., Gil-López, S., Molina, D., Benjamins, R., Chatila, R., & Herrera, F. (2020). Explainable artificial intelligence: Concepts, taxonomies, opportunities and challenges toward responsible AI. Information Fusion, 58, 82–115. https://doi.org/10.1016/j.inffus.2019.12.012
  20. [20] Luo, Y., Liang, P., Wang, C., Shahin, M., & Zhan, J. (2021). Characteristics and challenges of low-code development: The practitioners’ perspective. Proceedings of the ACM/IEEE International Symposium on Empirical Software Engineering and Measurement, 1–11. https://doi.org/10.1145/3475716.347578

Related Articles

Publish Your Research With UIMRJ

Submit your original research article, review paper, case study or conceptual paper to an international open access, peer-reviewed and refereed journal.

Article Transparency Report

This transparency report provides information regarding peer review, editorial evaluation, publication ethics, metadata verification, DOI registration, article integrity, licensing and scholarly publishing standards followed by UIMRJ.

Conflict of InterestAs Declared by Author(s)
Funding SourceAs Declared in Article
AI Usage DeclarationAs Per Journal Policy
Ethical ComplianceVerified
Peer Review ModelDouble-Blind Peer Review
Editorial ScreeningCompleted
Plagiarism ScreeningCompleted
Research Integrity CheckVerified
DOI StatusRegistered
Metadata QualityMachine-Readable Metadata
Google Scholar CompatibilityMetadata Compatible
Schema.org MetadataImplemented
Dublin Core MetadataImplemented
Version StatusVersion of Record
Correction StatusNo Correction Issued
Retraction StatusNot Retracted
Access ModelOpen Access
LicenseCC BY-NC 4.0
Publication EthicsCOPE Aligned
Chat with Editor