How TabPFN covers wide tables
Each estimator sees up to 768 features. When a table is wider, TabPFN distributes columns across estimators so that every column is seen by some estimator, and averages the estimators’ predictions. With the default 8 estimators this covers about 6,000 columns. Beyond that, raisen_estimators so more columns are covered, or select features yourself.
When to select features
- Many noisy columns. Irrelevant features dilute attention. If you know that most columns carry no signal, filtering them helps both accuracy and speed.
- Beyond the recommended width. Above about 6,000 columns, either raise
n_estimatorsor select features so each estimator sees the columns that matter. - Latency budgets. Fewer columns mean faster fits and predictions, and a smaller KV cache.
Approaches
Greedy feature selection removes features individually and checks performance. This works particularly well on smaller data with low computational cost. Mutual information filtering ranks features by mutual information with the target and keeps the top k:FEATURE_SUBSAMPLING_METHOD to "gini_feature_importance":