Clausthal University of TechnologyPrior LabsSiemens AIUniversity of MannheimUniversity of RostockJun 1, 2026arXiv:2606.02384

TabPrep: Closing the Feature Engineering Gap in Tabular Benchmarks

Andrej Tschalzev, Nick Erickson, Yuyang Wang, Huzefa Rangwala, Stefan Lüdtke, Heiner Stuckenschmidt, Christian Bartelt

AI Summary

This paper introduces TabPrep, a lightweight preprocessing pipeline that enhances feature engineering in tabular machine learning by addressing specific structural data patterns. The authors demonstrate that many popular model architectures have predictable limitations when these patterns are not accounted for, and they show that systematic feature engineering can significantly boost performance. By integrating TabPrep into model training across various benchmarks, the study reveals consistent performance improvements, often exceeding those achieved through model-centric innovations alone.

Key Contribution

Systematic feature engineering can outperform advanced model architectures, closing a critical evaluation gap in tabular benchmarks.

Abstract

Progress in tabular machine learning has largely focused on increasingly sophisticated model architectures. At the same time, feature engineering remains a critical yet underexplored component of real-world modeling pipelines that is entirely absent from modern benchmarks, which creates an unquantified evaluation gap. In this work, we introduce TabPrep, a lightweight preprocessing pipeline composed of feature generators that are carefully designed to target three specific structural data patterns. We show that many widely used model classes exhibit predictable blind spots to these patterns and that systematic feature engineering alone can establish new peak performance. Across the TabArena benchmark, integrating TabPrep into model training and tuning consistently improves performance for tree-based, neural, linear, and foundation models, often surpassing gains achieved by model-centric innovations alone. TabPrep outperforms previous automated feature engineering approaches in performance, efficiency, and applicability across datasets, enabling integration into large-scale benchmarks. By releasing TabPrep (see https://github.com/atschalz/tabprep), we enable researchers to integrate feature engineering into their benchmarking setup, filling a longstanding gap in tabular evaluations.

Data Curation & Synthetic Data Eval Frameworks & Benchmarks

Citation Metrics

Citations0

Influential citations0

References0

Year2026

VenueN/A

Related Papers

Finding related papers...

Search

TabPrep: Closing the Feature Engineering Gap in Tabular Benchmarks

Related Papers