A Unified Feedback-Driven Data Preparation Framework Using Persistent Feedback Artefacts
DOI:
https://doi.org/10.70882/noun-ijcea.2026.1148Keywords:
Adaptive Data Preparation, Context-Aware Retrieval, Feedback Reuse, Human-in-the-Loop, Persistent Feedback ArtefactsAbstract
Data preparation often requires analysts to make corrections that depend on the meaning and intended use of the data, not only its format. In many preparation tools, those corrections remain in the current file, script, or processing session and are unavailable when a similar problem appears in a later dataset. This study develops a Unified Feedback-Driven Data Preparation framework that incorporates the Persistent Feedback Artefact Model (PeFAM) into an end-to-end preparation workflow. The framework links dataset ingestion, staging, validation, transformation, feedback capture, persistent storage, analyst review, selective reuse, and export. A web prototype was implemented in PHP within a local XAMPP environment. CSV files were used for input and output, while MySQL records stored dataset settings, current-session feedback, and persistent feedback artefacts. The prototype retained validated corrections together with their original values and available contextual information. During a later task, it retrieved an artefact when the stored original value appeared in the current dataset and the affected attribute matched by exact column name or configured semantic role. The analyst then decided whether to select the suggestion. Selected items were copied to the active feedback store for use during transformation, while unselected items were not applied. The demonstration shows that these functions can operate within one preparation environment. The study does not evaluate efficiency, accuracy, scalability, or usability; those outcomes require a separate controlled evaluation. Its contribution is the operational connection between persistent artefact management and the core stages of data preparation.
References
Abdulkadir, S., Odion, P. O., Chinyio, D. T., Saidu, I. R., & Ahmad, M. A. (2026a). Feedback-driven automation in data preparation: A systematic literature review. Science World Journal (SWJ), 21(1), 290-303. https://doi.org/10.4314/swj.v21i1.39
Abdulkadir, S., Odion, P. O., Chinyio, D. T., Saidu, I. R., & Ahmad, M. A. (2026b). Formalising feedback as reusable knowledge in data preparation systems: A persistent feedback artefact model. FUDMA Journal of Sciences (FJS). Accepted for publication.
Azeroual, O. (2020). Data wrangling in database systems: Purging of dirty data. Data, 5(2), Article 50. https://doi.org/10.3390/data5020050
Cormier, K., Gagnier, K., Padron-Uy, J., Sareen, D., Parihar, A., Khmelevsky, Y., Hains, G., & Wong, A. (2025). Data extraction, transformation, and loading (ETL) process automation and data warehouse implementation for algorithmic trading machine learning modelling. In 2025 IEEE International Systems Conference (SysCon) (pp. 1–8). IEEE. https://doi.org/10.1109/syscon64521.2025.11014812
Fernandes, A. A. A., Koehler, M., Konstantinou, N., Pankin, P., Paton, N. W., & Sakellariou, R. (2023). Data preparation: A technological perspective and review. SN Computer Science, 4(4), Article 425. https://doi.org/10.1007/s42979-023-01828-8
Kasica, S., Berret, C., & Munzner, T. (2020). Table Scraps: An actionable framework for multi-table data wrangling from an artifact study of computational journalism. IEEE Transactions on Visualization and Computer Graphics, 27(2), 957–966. https://doi.org/10.1109/tvcg.2020.3030462
Konstantinou, N., Abel, E., Bellomarini, L., Bogatu, A., Civili, C., Irfanie, E., Koehler, M., Mazilu, L., Sallinger, E., Fernandes, A. A. A., Gottlob, G., Keane, J. A., & Paton, N. W. (2019). VADA: An architecture for end user informed data preparation. Journal of Big Data, 6(74), 1-32. https://doi.org/10.1186/s40537-019-0237-9
Konstantinou, N., & Paton, N. W. (2020). Feedback driven improvement of data preparation pipelines. Information Systems, 92, Article 101480. https://doi.org/10.1016/j.is.2019.101480
Liu, L., Hasegawa, S., Sampat, S. K., Xenochristou, M., Chen, W., Kato, T., Kakibuchi, T., & Asai, T. (2024). AutoDW: Automatic data wrangling leveraging large language models. In Proceedings of the 39th IEEE/ACM International Conference on Automated Software Engineering (pp. 2041–2052). ACM. https://doi.org/10.1145/3691620.3695267
Mosqueira-Rey, E., Hernández-Pereira, E., Alonso-Ríos, D., Bobes-Bascarán, J., & Fernández-Leal, Á. (2022). Human-in-the-loop machine learning: a state of the art. Artificial Intelligence Review, 56(4), 3005–3054. https://doi.org/10.1007/s10462-022-10246-w
Narayan, A., Chami, I., Orr, L., & Ré, C. (2022). Can foundation models wrangle your data? Proceedings of the VLDB Endowment, 16(4), 738-746. https://doi.org/10.14778/3574245.357425
Njeri, N. R. (2022). Data preparation for machine learning modelling. International Journal of Computer Applications Technology and Research, 11(06), 231–235. https://doi.org/10.7753/ijcatr1106.1008
Paton, N. (2019). Automating data preparation: Can we? Should we? Must we? Research Explorer, the University of Manchester. https://research.manchester.ac.uk/en/publications/automating-datapreparation-can-we-should-we-must-we/
Rezig, E. K., Ouzzani, M., Elmagarmid, A. K., Aref, W. G., & Stonebraker, M. (2019). Towards an end-to-end human-centric data cleaning framework. In Proceedings of the Workshop on Human-In-the-Loop Data Analytics (pp. 1–7). ACM. https://doi.org/10.1145/3328519.3329133
Wojciechowski, A. (2018). ETL workflow reparation by means of case-based reasoning. Information Systems Frontiers, 20, 21-43. https://doi.org/10.1007/s10796-016-9732-0
Wu, X., Xiao, L., Sun, Y., Zhang, J., Ma, T., & He, L. (2023). A survey of human-in-the-loop for machine learning. Future Generation Computer Systems, 135, 364–381. https://doi.org/10.1016/j.future.2022.05.014
Downloads
Published
Issue
Section
License
Copyright (c) 2026 Saadatu Abdulkadir, Philip O. Odion, Darius T. Chinyio, Isah R. Saidu, Muhammad A. Ahmad (Author)

This work is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License.

