InferenceConfig
</> View source ↗TabPFNClassifier
and TabPFNRegressor interfaces. The options in this class are more advanced and
not expected to be changed by the (standard) user.
Several of the preprocessing options are supported by our code for efficiency
reasons (to avoid loading TabPFN multiple times). However, these can also be
applied outside of the model interface.
This class must be serializable as it is peristed in the model checkpoints.
Do not edit the default values in this class, as this can affect the backwards
compatibility of the model checkpoints. Instead, edit get_default().
Fields
| Name | Type | Default | Description |
|---|---|---|---|
PREPROCESS_ | list[Preprocessor | Required | The preprocessing applied to the data before passing it to TabPFN. See PreprocessorConfig for options and more details. If multiple PreprocessorConfig are provided, they are (repeatedly) applied across different estimators.By default, for classification, two preprocessors are applied: 1. Uses the original input data, all features transformed with a quantile scaler, and the first n-many components of SVD transformer (whereby n is a fract of on the number of features or samples). Categorical features are ordinal encoded but all categories with less than 10 features are ignored. 2. Uses the original input data, with categorical features as ordinal encoded. By default, for regression, two preprocessor are applied: 1. The same as for classification, with a minimal different quantile scaler. 2. The original input data power transformed and categories onehot encoded. |
MAX_ | int | 30 | The maximum number of unique values for a feature to be considered categorical. Otherwise, it is considered numerical. |
MIN_ | int | 4 | The minimum number of unique values for a feature to be considered numerical. Otherwise, it is considered categorical. |
MIN_ | int | 100 | The minimum number of samples in the data to run our infer which features might be categorical. |
MIN_ | int | 30 | Number of distinct values above which a string column is read as text rather than as a category. Only an undeclared column is subject to it: one listed in categorical_features_indices, or holding pandas’ category dtype, is a category at any cardinality. A separate decision from MAX_UNIQUE_FOR_CATEGORICAL_FEATURES, which governs numerical-vs-categorical: that one describes when a number is few enough to be a category, this one describes when a string is varied enough to be text rather than a category, and there is no reason the two should move together.Text is expanded into numeric features with TRANSFORM_TEXT; off, it is ordinal-encoded as a high-cardinality category and fit warns about it. |
SOFTMAX_ | float | DEFAULT_SOFTMAX_TEMPERATURE | The temperature applied to the model’s logits at predict time. Lower values make the predictions more confident, 1.0 is a no-op.The default is the value that shipped as the softmax_temperature argument of TabPFNClassifier/TabPFNRegressor before this became a config field, so checkpoints that predate the field (which is all of them up to and including the ones released with v8.5.0) keep their original behavior. Newer checkpoints are expected to store their own value and must do so explicitly.Setting this here overrides the checkpoint for every model in the ensemble, as does TabPFNClassifier(softmax_temperature=...); naming a temperature both ways at once is rejected. With neither, the value comes from the checkpoint, and an ensemble whose checkpoints declare different temperatures is rejected too. |
N_ | int | Literal[“auto”] | "auto" | How many estimators to run when the user leaves n_estimators="auto".An estimator is one forward pass over a differently preprocessed view of the data; more of them costs proportionally more compute. This means exactly what the n_estimators argument of the estimators means, and is used in its place when the user names no count:- If an int, that many estimators run, and feature-coverage scaling never raises it — a checkpoint that asks for a count gets that count, the same guarantee a user passing one gets. - If "auto" (the default), DEFAULT_N_ESTIMATORS estimators run, raised on wide tables so every feature is seen by some estimator (see scale_n_estimators_for_feature_coverage).The default leaves the decision where it was before this field existed, so checkpoints that predate it keep their original behavior. Newer checkpoints are expected to store a count of their own. Setting this here overrides the checkpoint, as does TabPFNClassifier(n_estimators=...); naming a count both ways at once is rejected. With neither, the value comes from the checkpoint, and an ensemble whose checkpoints declare different counts is rejected too. |
TRANSFORM_ | bool | False | Whether a column holding a genuine datetime dtype (datetime64, tz-aware, or period) is expanded into calendar features via skrub.DatetimeEncoder. Off, such a column is refused with an error naming it: cast or expand it yourself first. Only a real datetime dtype counts: a string column that merely looks like a date (e.g. “2020-01-01”) is read as a plain category or text either way.On, the same columns have to hold datetimes at predict, in a DataFrame, and none of them may be listed in categorical_features_indices; each of these is refused with an error saying so rather than guessed at. The fine-tuning estimators do not run this conversion, whatever inference_config they are handed, so a datetime column has to be converted before fine-tuning. |
TRANSFORM_ | bool | False | Whether a text column, a pandas string or pyarrow string column with more than MIN_CARDINALITY_FOR_TEXT distinct values, is expanded into TEXT_N_COMPONENTS numeric features via skrub.StringEncoder (tf-idf over character n-grams, truncated SVD). Off, such a column is ordinal-encoded as a high-cardinality category and fit warns about it. An object column is never expanded. Not run by the fine-tuning estimators. |
TEXT_ | int | 30 | Features a text column is expanded into with TRANSFORM_TEXT: the leading components of a truncated SVD over its tf-idf matrix. Fewer when the column has fewer character n-grams than that. |
OUTLIER_ | float | None | Literal[“auto”] | "auto" | The number of standard deviations from the mean to consider a sample an outlier. - If None, no outliers are removed.- If float, the number of standard deviations from the mean to consider a sample an outlier. - If “auto”, the OUTLIER_REMOVAL_STD is automatically determined. -> 12.0 for classification and None for regression. |
FEATURE_ | Literal[“shuffle”, “rotate”] | None | "shuffle" | The method used to shift features during preprocessing for ensembling to emulate the effect of invariance to feature position. Without ensembling, TabPFN is not invariant to feature position due to using a transformer. Moreover, shifting features can have a positive effect on the model’s performance. The options are: - If “shuffle”, the features are shuffled. - If “rotate”, the features are rotated (think of a ring). - If None, no feature shifting is done. |
CLASS_ | Literal[“rotate”, “shuffle”] | None | "shuffle" | The method used to shift classes during preprocessing for ensembling to emulate the effect of invariance to class order. Without ensembling, TabPFN is not invariant to class order due to using a transformer. Shifting classes can have a positive effect on the model’s performance. The options are: - If “shuffle”, the classes are shuffled. - If “rotate”, the classes are rotated (think of a ring). - If None, no class shifting is done. |
FINGERPRINT_ | bool | True | Whether to add a fingerprint feature to the data. The added feature is a hash of the row, counting up for duplicates. This helps TabPFN to distinguish between duplicated data points in the input data. Otherwise, duplicates would be less obvious during attention. This is expected to improve prediction performance and help with stability if the data has many sample duplicates. |
POLYNOMIAL_ | Literal[“no”, “all”] | int | "no" | The number of 2 factor polynomial features to generate and add to the original data before passing the data to TabPFN. The polynomial features are generated by multiplying the original features together, e.g., this might add a feature x1*x2 to the features, if x1 and x2 are features. In total, this can add up O(n^2) many features. Adding polynomial features can improve predictive performance by exploiting simple feature engineering.- If “no”, no polynomial features are added. - If “all”, all possible polynomial features are added. - If an int, determines the maximal number of polynomial features to add to the original data. |
SUBSAMPLE_ | int | float | list | None | None | Subsample the input data sample/row-wise before performing any preprocessing and the TabPFN forward pass. - If None, no subsampling is done.- If an int, the number of samples to subsample (or oversample if SUBSAMPLE_SAMPLES is larger than the number of samples).- If a float, the percentage of samples to subsample. - If a list arrays of indices, the indices to subsample for each estimator. If the length of the outer list is less than the number of estimators, the indices are repeated for the remaining estimators. |
SAMPLE_ | Literal[“auto”, “balanced”, “stratified”, “majority_downsample”] | "auto" | How rows are drawn for each estimator when SUBSAMPLE_SAMPLES is an int or float. Ignored when SUBSAMPLE_SAMPLES is None or a list of explicit indices.- “balanced”: Round-robin sampling from a shared shuffled pool of all rows so each row appears approximately equally often across estimators. Ignores the class labels. - “stratified”: Preserves the class proportions of the training data in every subsample while guaranteeing at least one row per class. Classification only. - “majority_downsample”: Groups rows by exact target value, keeps every row outside the single most frequent group, and fills the remaining budget from that majority group. This mode is designed for datasets with one dominant target value. For binary classification this keeps the whole minority class and fills up with majority rows. For regression it targets zero-inflated or spiky targets: the repeated value is downsampled while all other values are kept. SUBSAMPLE_SAMPLES must exceed the number of non-majority rows so at least one majority row remains. If there is no unique most frequent target value, a warning is emitted and the method falls back to “stratified” for classification or “balanced” for regression. Downsampling the majority shifts the target prior that the model sees: the majority value is underrepresented in every context relative to the training data. Predicted probabilities and regression means inherit that shift, so the predicted level typically needs a correction, for example rescaling regression predictions to the training mean. Rankings are unaffected. For classification, a warning is emitted if subsampling makes the original majority class smaller than another class.- “auto”: “stratified” for classification and “balanced” for regression. |
ENABLE_ | bool | False | Move quantile transform, SVD feature generation, and feature shuffling to GPU / torch. When True, these operations run on the same device as the model, which can be significantly faster for large datasets (>10 k rows). When False (default), all preprocessing runs on CPU / sklearn as before.Only quantile_uni* transforms are accelerated (the torch quantile transformer only supports uniform output). Other transforms stay on CPU regardless of this flag. SVD and shuffle always move to GPU / torch when this flag is set. |
FEATURE_ | Literal[“balanced”, “random”, “constant_and_balanced”, “gini_feature_importance”, “auto”] | "balanced" | The method used to subsample features when the dataset has more features than max_features_per_estimator. The options are: - “random”: Each estimator independently draws a random subset of features. - “balanced”: Round-robin sampling from a shared shuffled pool so each feature appears approximately equally across estimators. - “constant_and_balanced”: Always include the first N features (see FEATURE_SUBSAMPLING_CONSTANT_FEATURE_COUNT), then use balanced subsampling for the rest.- “gini_feature_importance”: Use LightGBM gain importance to rank features. Always include the top-K most important features (see FEATURE_SUBSAMPLING_IMPORTANCE_TOP_K_COUNT), fill the rest via balanced round-robin sampling from the remaining features.- “auto”: Automatically selects the method based on dataset size and whether feature subsampling is needed. Uses “gini_feature_importance” when n_samples > AUTO_FEATURE_SUBSAMPLING_IMPORTANCE_MIN_SAMPLES(=100_000) and subsampling is required (importance scoring is more accurate on larger datasets), otherwise falls back to “balanced”. |
FEATURE_ | int | 50 | The number of leading features that are always included when using the ‘constant_and_balanced’ feature subsampling method. Only used when FEATURE_SUBSAMPLING_METHOD is ‘constant_and_balanced’. |
FEATURE_ | int | float | Literal[“auto”] | "auto" | Number of top important features always included per estimator when FEATURE_SUBSAMPLING_METHOD is an importance-based method. The remaining budget up to max_features_per_estimator is filled randomly from the remaining features.- If an int, that many features are always included. - If a float in (0, 1], resolved as ceil(value * n_total_features). - If “auto”, uses top-k=AUTO_FEATURE_SUBSAMPLING_TOP_K(=150) when n_features > AUTO_FEATURE_SUBSAMPLING_TOP_K_MIN_FEATURES(=200); otherwise no importance filtering is done. |
REGRESSION_ | tuple[str | None, …] | (None, "safepower") | The preprocessing applied to the target variable before passing it to TabPFN for regression. This can be understood as scaling the target variable to better predict it. The preprocessors should be passed as a tuple/list and are then (repeatedly) used by the estimators in the ensembles. By default, we use no preprocessing and a power transformation (if we have more than one estimator). The options are: - None: no preprocessing is done.- One of the options from tabpfn.preprocessing.get_all_reshape_feature_distribution_preprocessors() |
USE_ | bool | False | Whether to round the probabilities to float 16 to match the precision of scikit-learn. This can help with reproducibility and compatibility with scikit-learn but is not recommended for general use. This is not exposed to the user or as a hyperparameter. To improve reproducibility,set ._sklearn_16_decimal_precision = True before calling .predict() or .predict_proba(). |
MAX_ | int | 10 | The number of classes seen during pretraining for classification. If the number of classes is larger than this number, TabPFN requires an additional step to predict for more than classes. |
MAX_ | int | 500 | The number of features that the pretraining was intended for. If the number of features is larger than this number, you may see degraded performance. Note, this is not the number of features seen by the model during pretraining but also accounts for expected generalization (i.e., length extrapolation). |
MAX_ | int | 10000 | The number of samples that the pretraining was intended for. If the number of samples is larger than this number, you may see degraded performance. Note, this is not the number of samples seen by the model during pretraining but also accounts for expected generalization (i.e., length extrapolation). |
MAX_ | int | 1000 | The number of samples above which CPU inference is disallowed by default due to slow performance. Raise via ignore_pretraining_limits or the TABPFN_ALLOW_CPU_LARGE_DATASET setting. |
FIX_ | bool | True | Whether to repair any borders of the bar distribution in regression that are NaN after the transformation. This can happen due to multiple reasons and should in general always be done. |
PASSTHROUGH_ | bool | False | Whether to pass infinite values through to the model instead of rejecting them. When True, +/-inf are temporarily replaced with NaN for preprocessing and restored afterwards; when False, infinities are rejected at input validation. |
InferenceConfig.override_with_user_input_and_resolve_auto
</> View source ↗user_config overwritten.
InferenceConfig.override_with_user_input_and_resolve_auto(
user_config: dict | InferenceConfig | None,
) -> InferenceConfig
| Parameter | Type | Default | Description |
|---|---|---|---|
user_ | dict | Inference | Required | Config provided by the user at inference time. If a dictionary, then the keys must match attributes of InferenceConfig and will be used to override these attributes. If an InferenceConfig object, then the whole config is overridden with the values from the user config. Deprecated. If None, then a copy of this config is returned with no fields changed. |
| Type | Description |
|---|---|
Inference | — |
InferenceConfig.equals_ignoring_overridable_fields
</> View source ↗other agree on every non-overridable field.
A mismatch in one of OVERRIDABLE_FIELDS between the checkpoints of one
ensemble gets its own error (see
raise_if_checkpoints_disagree_on_overridable_fields), since the user can
resolve it by naming a value; any other mismatch is unfixable.
InferenceConfig.equals_ignoring_overridable_fields(
other: InferenceConfig,
) -> bool
| Parameter | Type | Default | Description |
|---|---|---|---|
other | Inference | Required | — |
| Type | Description |
|---|---|
bool | — |
InferenceConfig.get_resolved_outlier_removal_std
</> View source ↗InferenceConfig.get_resolved_outlier_removal_std(
estimator_type: Literal["regressor", "classifier"],
) -> float | None
| Parameter | Type | Default | Description |
|---|---|---|---|
estimator_ | Literal[“regressor”, “classifier”] | Required | — |
| Type | Description |
|---|---|
float | None | — |
InferenceConfig.get_default
</> View source ↗InferenceConfig.get_default(
task_type: TaskType,
model_version: ModelVersion,
) -> InferenceConfig
Type aliases
Type aliases
TaskType = Literal["multiclass", "regression"]
| Parameter | Type | Default | Description |
|---|---|---|---|
task_ | Task | Required | — |
model_ | Model | Required | — |
| Type | Description |
|---|---|
Inference | — |