Step-by-Step Data Cleaning Recommendations to Improve ML Prediction Accuracy

March 14, 2025 · Declared Dead · 🏛 International Conference on Extending Database Technology

"No code URL or promise found in abstract"

Evidence collected by the PWNC Scanner

Authors Sedir Mohammed, Felix Naumann, Hazar Harmouch arXiv ID 2503.11366 Category cs.DB: Databases Citations 4 Venue International Conference on Extending Database Technology Last Checked 4 months ago

Abstract

Data quality is crucial in machine learning (ML) applications, as errors in the data can significantly impact the prediction accuracy of the underlying ML model. Therefore, data cleaning is an integral component of any ML pipeline. However, in practical scenarios, data cleaning incurs significant costs, as it often involves domain experts for configuring and executing the cleaning process. Thus, efficient resource allocation during data cleaning can enhance ML prediction accuracy while controlling expenses. This paper presents COMET, a system designed to optimize data cleaning efforts for ML tasks. COMET gives step-by-step recommendations on which feature to clean next, maximizing the efficiency of data cleaning under resource constraints. We evaluated COMET across various datasets, ML algorithms, and data error types, demonstrating its robustness and adaptability. Our results show that COMET consistently outperforms feature importance-based, random, and another well-known cleaning method, achieving up to 52 and on average 5 percentage points higher ML prediction accuracy than the proposed baselines.

📄 View on arXiv 🌐 View on ar5iv 📑 PDF 🎉 Report Code Found

Community Contributions

Found the code? Know the venue? Think something is wrong? Let us know!

📜 Similar Papers

In the same crypt — Databases

R.I.P. 👻 Ghosted

The Case for Learned Index Structures

Tim Kraska, Alex Beutel, ... (+3 more)

cs.DB 🏛 SIGMOD 📚 1.2K cites 8 years ago

R.I.P. 👻 Ghosted

Untangling Blockchain: A Data Processing View of Blockchain Systems

Tien Tuan Anh Dinh, Rui Liu, ... (+4 more)

cs.DB 🏛 IEEE TKDE 📚 997 cites 8 years ago

R.I.P. 👻 Ghosted

Converting Static Image Datasets to Spiking Neuromorphic Datasets Using Saccades

Garrick Orchard, Ajinkya Jayawant, ... (+2 more)

cs.DB 🏛 Frontiers in Neuroscience 📚 905 cites 10 years ago

R.I.P. 👻 Ghosted

BLOCKBENCH: A Framework for Analyzing Private Blockchains

Tien Tuan Anh Dinh, Ji Wang, ... (+4 more)

cs.DB 🏛 SIGMOD 📚 872 cites 9 years ago

R.I.P. 👻 Ghosted

Data Synthesis based on Generative Adversarial Networks

Noseong Park, Mahmoud Mohammadi, ... (+4 more)

cs.DB 🏛 VLDB 📚 568 cites 8 years ago

R.I.P. 👻 Ghosted

HoloClean: Holistic Data Repairs with Probabilistic Inference

Theodoros Rekatsinas, Xu Chu, ... (+2 more)

cs.DB 🏛 VLDB 📚 544 cites 9 years ago

Died the same way — 👻 Ghosted

R.I.P. 👻 Ghosted

Federated Learning: Strategies for Improving Communication Efficiency

Jakub Konečný, H. Brendan McMahan, ... (+4 more)

cs.LG 🏛 arXiv 📚 5.2K cites 9 years ago

R.I.P. 👻 Ghosted

In-Datacenter Performance Analysis of a Tensor Processing Unit

Norman P. Jouppi, Cliff Young, ... (+73 more)

cs.AR 🏛 ISCA 📚 5.1K cites 9 years ago

R.I.P. 👻 Ghosted

Deep Convolutional Neural Networks for Computer-Aided Detection: CNN Architectures, Dataset Characteristics and Transfer Learning

Hoo-Chang Shin, Holger R. Roth, ... (+7 more)

cs.CV 🏛 IEEE TMI 📚 4.9K cites 10 years ago

R.I.P. 👻 Ghosted

Explanation in Artificial Intelligence: Insights from the Social Sciences

Tim Miller

cs.AI 🏛 AI 📚 4.9K cites 9 years ago