StandardScaler or MinMaxScaler, imputing missing values, one-hot encoding categoricals, vectorizing text by hand, or extracting calendar features from datetime columns yourself.
Escalation path
When the default TabPFN does not meet your needs, try these approaches in roughly this order:1
Check column types
Declare categorical columns, especially integer-coded ones, through
categorical_features_indices or pandas category dtype. Keep datetime columns as datetime dtype and free text as pandas string dtype. For grouped data, include the group identifier as a categorical column. Column typing is the cheapest change and often the largest gain. See Preprocessing.2
Feature engineering
Add domain features TabPFN cannot derive from raw columns: ratios, interactions, group aggregations, and external signals. See Feature Engineering.
3
Row subsampling
On tables above the row limit, or with one dominant class or target value, control which rows each estimator sees with
SUBSAMPLE_SAMPLES and SAMPLE_SUBSAMPLING_METHOD. See Row subsampling.4
Metric tuning
Use
eval_metric and tuning_config to optimize for your specific evaluation metric, and check calibration on skewed targets. See Model Parameters.5
Feature selection
Beyond several thousand columns, or with many noisy features, try filtering to the most informative ones. See Feature Selection.
6
Preprocessing transforms
Experiment with different
PREPROCESS_TRANSFORMS and target transforms. The preprocessing guide explains what each estimator does and provides a practical tuning order.7
Thinking mode and fine-tuning
On the API, Thinking mode spends more compute at fit time and is the strongest option for grouped and time-ordered data. Locally, fine-tune the pretrained model when you have a specialized domain or distribution shift.
Guides
Feature Engineering
Encode domain knowledge into features that TabPFN cannot learn from raw columns alone.
Feature Selection
When to reduce feature count on very wide tables, and how TabPFN covers features itself.
Preprocessing
Configure column typing, text and date handling, row subsampling, and per-estimator transforms.
Model Parameters
Tune softmax temperature, metric optimization, estimator count, and class imbalance handling.
Related
Thinking mode
Inference-time compute scaling on the API.
Fine-tuning
Adapt TabPFN’s pretrained weights to your domain.