Skip to main content

softmax_temperature

Controls prediction sharpness for both TabPFNClassifier and TabPFNRegressor. It is applied to the model’s logits at prediction time.
  • Lower values such as 0.8 give sharper, more confident predictions.
  • Higher values such as 1.2 give softer predictions, useful when probability calibration matters.
  • 1.0 applies no change.
The default is "auto", which uses the temperature stored in the checkpoint: 1.0 for TabPFN-3.5 and TabPFN-3.5-Fast, 0.9 for TabPFN-3 and earlier. The current value was chosen with skewed and zero-heavy regression targets in mind, where a sharper temperature biases the predicted mean. Do not carry a hard-coded softmax_temperature=0.9 over from TabPFN-3 code.
With tuning_config={"calibrate_temperature": True}, the temperature is tuned on a holdout and overrides this value. For regression, eval_metric chooses whether the calibration optimizes "nll" or "crps".

Metric tuning

For metrics that are sensitive to decision thresholds such as F1 or balanced accuracy, use the built-in metric tuning:
Threshold tuning cannot optimize roc_auc or log_loss, and combining them with tune_decision_thresholds=True raises an error. Use calibrate_temperature=True for log loss.

Handling imbalanced data

TabPFN is robust to imbalance out of the box. Three options change how it treats the minority class:
  • balance_probabilities=True reweights predicted probabilities so each class counts equally. Use it when your metric weights classes equally, such as balanced accuracy or balanced log loss.
  • eval_metric="balanced_accuracy" with threshold tuning gives more control over the operating point.
  • Majority downsampling keeps every minority row in every estimator’s context and downsamples only the majority class. This applies when the table exceeds the row limit or when you set SUBSAMPLE_SAMPLES yourself. See Row subsampling.
Every one of these options trades one metric against another. Balancing and majority downsampling tend to improve ranking metrics such as AUC and worsen log loss, because the model sees a shifted class prior. Pick the metric you are judged by and test both settings.

n_estimators

n_estimators controls the number of ensemble members. Each estimator uses a different preprocessing configuration and a different feature order, which makes the average more robust. The default is "auto", which uses the count stored in the checkpoint: 8 for TabPFN-3.5 and TabPFN-3, 4 for TabPFN-3.5-Fast. A checkpoint that declares a count is run with exactly that count. Older checkpoints without a stored count start at 8 and are raised automatically on wide tables so that every feature is covered, up to a cap of 32. Pass an explicit integer to fix the budget or to trade speed against accuracy:
An explicit value is never auto-scaled. TabPFN warns at fit time if it is too small for every feature to be covered. On wide tables, more estimators cover more columns; on narrow tables, more estimators give diminishing returns.
The auto_scale_n_estimators argument is deprecated and will be removed in a future release. Pass an explicit n_estimators instead of setting auto_scale_n_estimators=False.