In the world of machine learning, gradient boosting machines like XGBoost have long been the go-to solution for tabular data. Their ability to handle structured datasets with speed and accuracy has made them a staple in competitions and industry applications. But a recent benchmark has sent ripples through the community: two relatively new models, TabPFN and TabICL, have reportedly outperformed a heavily tuned XGBoost on all 14 datasets tested. The twist? These models require no training on the target dataset, making their victory even more striking.
The Benchmark That Shook the Boosting World
The benchmark, which has been circulating among machine learning practitioners, pitted TabPFN and TabICL against a meticulously tuned XGBoost across 14 diverse tabular datasets. The results were unanimous: both foundation models achieved better performance on every single dataset. This isn't just a minor improvement; in some cases, the margin was significant enough to question the dominance of gradient boosting in tabular tasks.
What makes this outcome particularly intriguing is that XGBoost was not a slouch. It was tuned using state-of-the-art hyperparameter optimization, likely involving extensive grid searches or Bayesian optimization. Yet, despite this effort, the no-training models came out on top. This raises important questions about the future of tabular machine learning and the role of pre-trained foundation models.
What Are TabPFN and TabICL?
TabPFN, short for Tabular Prior-Data Fitted Network, is a transformer-based model that was pre-trained on a large corpus of synthetic tabular datasets. The idea is that by exposing the model to a wide variety of data-generating processes, it learns to perform Bayesian inference in context. In practice, this means you can feed it a training set and a test set, and it will predict the test labels in a single forward pass, without any gradient updates. This is akin to how large language models perform few-shot learning.
TabICL, on the other hand, is a more recent entrant that builds on similar principles but introduces innovations to handle larger datasets and improve scalability. It also leverages in-context learning, where the model conditions on the provided training data to make predictions. Both models are part of a growing trend of foundation models for structured data, aiming to bring the success of pre-training from vision and language to the tabular domain.
Why This Matters: The End of Hyperparameter Tuning?
For years, the recipe for tabular data success has been: clean your data, engineer features, and then spend hours tuning your XGBoost or LightGBM model. This benchmark suggests that this paradigm might be shifting. If a model can achieve superior performance without any dataset-specific training, it could drastically reduce the time and expertise required to build high-performing models.
Moreover, the fact that TabPFN and TabICL won on all 14 datasets indicates robustness across different types of tabular problems, from small to medium-sized datasets. This consistency is a strong signal that these models are not just flukes but represent a genuine leap forward.
The Caveats: Where XGBoost Still Shines
Before we declare the death of gradient boosting, it's important to note the limitations of this benchmark. The datasets used are likely small to medium in size, as TabPFN and TabICL currently have constraints on the number of samples and features they can handle efficiently. XGBoost, with its scalability and ability to handle large-scale data, remains the tool of choice for many real-world applications where datasets have millions of rows.
Additionally, the benchmark focused on predictive performance, but other factors such as interpretability, training time, and inference latency also matter. XGBoost offers feature importance scores and is well-understood, while TabPFN and TabICL are more black-box. For regulated industries, this could be a barrier.
Implications for Practitioners and Researchers
For data scientists, this benchmark is a wake-up call to explore these new models. They are available as open-source libraries and can be easily integrated into existing workflows. For researchers, it highlights the potential of in-context learning for tabular data and encourages further investigation into scaling these models to larger datasets.
The victory of TabPFN and TabICL also underscores a broader trend: the convergence of deep learning and tabular data. While neural networks have historically struggled with tabular data compared to tree-based methods, foundation models like these are closing the gap.
Frequently Asked Questions
What exactly does "no training" mean for TabPFN and TabICL?
It means the models do not update their weights on the new dataset. Instead, they use the provided training examples as context to make predictions, similar to how a large language model answers a question based on a prompt. This process is called in-context learning.
Can TabPFN and TabICL handle large datasets like XGBoost?
Currently, they are best suited for small to medium-sized datasets, typically up to a few thousand samples and features. For very large datasets, XGBoost and other gradient boosting methods still have the advantage in terms of scalability and memory efficiency.
Are these models easy to use for someone without deep learning expertise?
Yes, libraries like TabPFN provide a scikit-learn compatible API, so you can use them with just a few lines of code. You don't need to understand the inner workings to get started.
Will this replace XGBoost entirely?
Not likely in the near term. XGBoost remains a robust, scalable, and interpretable choice for many applications. However, for smaller datasets where top accuracy is critical, TabPFN and TabICL are strong contenders that can save significant tuning time.
Where can I try these models?
Both TabPFN and TabICL are available as open-source projects on GitHub. You can install them via pip and experiment with your own datasets to see how they perform.

