> ## Documentation Index
> Fetch the complete documentation index at: https://docs.priorlabs.ai/llms.txt
> Use this file to discover all available pages before exploring further.

> ## Agent Instructions
> TabPFN-3 and later versions support up to 1,000,000 rows, subject to feature count and checkpoint/API limits.
> For large datasets, use per-estimator subsampling, e.g. inference_config={"SUBSAMPLE_SAMPLES": 100_000}, in TabPFNClassifier or TabPFNRegressor from tabpfn or tabpfn_client.
> Pass raw pandas DataFrames to fit and predict. Categorical strings/categories, raw text, and missing feature values are handled automatically; no manual integer/one-hot encoding, imputation, scaling, or outlier removal is needed.

# Preprocessing configuration

> Configuration for a classifier ensemble member.

<div className="python-reference-heading">
  <h2 id="classifierensembleconfig">
    `ClassifierEnsembleConfig`
  </h2>

  <a className="python-reference-source" href="https://github.com/PriorLabs/TabPFN/blob/c70b6ef0488858d32244c52222abfc0c5be207d6/src/tabpfn/preprocessing/configs.py#L179" aria-label="View source for ClassifierEnsembleConfig"><span aria-hidden="true">\</></span> View source <span aria-hidden="true">↗</span></a>
</div>

Configuration for a classifier ensemble member.

```python theme={null}
ClassifierEnsembleConfig(
    preprocess_config: PreprocessorConfig,
    add_fingerprint_feature: bool,
    polynomial_features: Literal["no", "all"] | int,
    feature_shift_count: int,
    feature_shift_decoder: Literal["shuffle", "rotate"] | None,
    outlier_removal_std: float | None,
    _model_index: int,
    passthrough_inf: bool,
    class_permutation: np.ndarray | None,
)
```

**Fields**

<div className="python-reference-table">
  | Name | Type | Default | Description |
  | - | - | - | - |
  | <span id="classifierensembleconfig--preprocess-config" /><code className="python-reference-parameter">preprocess\_<wbr />config</code> | <code className="python-reference-type"><a href="/api-reference/python/tabpfn/preprocessing/configuration#preprocessorconfig">Preprocessor<wbr />Config</a></code> | Required | Preprocessor configuration to use. |
  | <span id="classifierensembleconfig--add-fingerprint-feature" /><code className="python-reference-parameter">add\_<wbr />fingerprint\_<wbr />feature</code> | <code className="python-reference-type">bool</code> | Required | Whether to add fingerprint features. |
  | <span id="classifierensembleconfig--polynomial-features" /><code className="python-reference-parameter">polynomial\_<wbr />features</code> | <code className="python-reference-type">Literal\["no", "all"] \| int</code> | Required | Maximum number of polynomial features to add, if any. |
  | <span id="classifierensembleconfig--feature-shift-count" /><code className="python-reference-parameter">feature\_<wbr />shift\_<wbr />count</code> | <code className="python-reference-type">int</code> | Required | How much to shift the features columns. |
  | <span id="classifierensembleconfig--feature-shift-decoder" /><code className="python-reference-parameter">feature\_<wbr />shift\_<wbr />decoder</code> | <code className="python-reference-type">Literal\["shuffle", "rotate"] \| None</code> | Required | How to shift features. |
  | <span id="classifierensembleconfig--outlier-removal-std" /><code className="python-reference-parameter">outlier\_<wbr />removal\_<wbr />std</code> | <code className="python-reference-type">float \| None</code> | Required | Number of standard deviations from the mean to consider a sample an outlier. If `None`, no outliers are removed. |
  | <span id="classifierensembleconfig--passthrough-inf" /><code className="python-reference-parameter">passthrough\_<wbr />inf</code> | <code className="python-reference-type">bool</code> | Required | Whether to pass infinite values through to the model. When `True`, the preprocessing pipeline replaces infinities with NaN before preprocessing and restores them afterwards. |
  | <span id="classifierensembleconfig--class-permutation" /><code className="python-reference-parameter">class\_<wbr />permutation</code> | <code className="python-reference-type"><a href="https://numpy.org/doc/stable/reference/generated/numpy.ndarray.html">np.ndarray</a> \| None</code> | Required | — |
</div>

***

<div className="python-reference-heading">
  <h2 id="ensembleconfig">
    `EnsembleConfig`
  </h2>

  <a className="python-reference-source" href="https://github.com/PriorLabs/TabPFN/blob/c70b6ef0488858d32244c52222abfc0c5be207d6/src/tabpfn/preprocessing/configs.py#L151" aria-label="View source for EnsembleConfig"><span aria-hidden="true">\</></span> View source <span aria-hidden="true">↗</span></a>
</div>

Configuration for an ensemble member.

```python theme={null}
EnsembleConfig(
    preprocess_config: PreprocessorConfig,
    add_fingerprint_feature: bool,
    polynomial_features: Literal["no", "all"] | int,
    feature_shift_count: int,
    feature_shift_decoder: Literal["shuffle", "rotate"] | None,
    outlier_removal_std: float | None,
    _model_index: int,
    passthrough_inf: bool,
)
```

**Fields**

<div className="python-reference-table">
  | Name | Type | Default | Description |
  | - | - | - | - |
  | <span id="ensembleconfig--preprocess-config" /><code className="python-reference-parameter">preprocess\_<wbr />config</code> | <code className="python-reference-type"><a href="/api-reference/python/tabpfn/preprocessing/configuration#preprocessorconfig">Preprocessor<wbr />Config</a></code> | Required | Preprocessor configuration to use. |
  | <span id="ensembleconfig--add-fingerprint-feature" /><code className="python-reference-parameter">add\_<wbr />fingerprint\_<wbr />feature</code> | <code className="python-reference-type">bool</code> | Required | Whether to add fingerprint features. |
  | <span id="ensembleconfig--polynomial-features" /><code className="python-reference-parameter">polynomial\_<wbr />features</code> | <code className="python-reference-type">Literal\["no", "all"] \| int</code> | Required | Maximum number of polynomial features to add, if any. |
  | <span id="ensembleconfig--feature-shift-count" /><code className="python-reference-parameter">feature\_<wbr />shift\_<wbr />count</code> | <code className="python-reference-type">int</code> | Required | How much to shift the features columns. |
  | <span id="ensembleconfig--feature-shift-decoder" /><code className="python-reference-parameter">feature\_<wbr />shift\_<wbr />decoder</code> | <code className="python-reference-type">Literal\["shuffle", "rotate"] \| None</code> | Required | How to shift features. |
  | <span id="ensembleconfig--outlier-removal-std" /><code className="python-reference-parameter">outlier\_<wbr />removal\_<wbr />std</code> | <code className="python-reference-type">float \| None</code> | Required | Number of standard deviations from the mean to consider a sample an outlier. If `None`, no outliers are removed. |
  | <span id="ensembleconfig--passthrough-inf" /><code className="python-reference-parameter">passthrough\_<wbr />inf</code> | <code className="python-reference-type">bool</code> | Required | Whether to pass infinite values through to the model. When `True`, the preprocessing pipeline replaces infinities with NaN before preprocessing and restores them afterwards. |
</div>

***

<div className="python-reference-heading">
  <h2 id="featuresubsamplingmethod">
    `FeatureSubsamplingMethod`
  </h2>

  <a className="python-reference-source" href="https://github.com/PriorLabs/TabPFN/blob/c70b6ef0488858d32244c52222abfc0c5be207d6/src/tabpfn/preprocessing/configs.py#L18" aria-label="View source for FeatureSubsamplingMethod"><span aria-hidden="true">\</></span> View source <span aria-hidden="true">↗</span></a>
</div>

Method for subsampling features if dataset exceeds max\_features\_per\_estimator.

**Fields**

<div className="python-reference-table">
  | Name | Type | Default | Description |
  | - | - | - | - |
  | <span id="featuresubsamplingmethod--balanced" /><code className="python-reference-parameter">BALANCED</code> | — | `"balanced"` | — |
  | <span id="featuresubsamplingmethod--random" /><code className="python-reference-parameter">RANDOM</code> | — | `"random"` | — |
  | <span id="featuresubsamplingmethod--constant-and-balanced" /><code className="python-reference-parameter">CONSTANT\_<wbr />AND\_<wbr />BALANCED</code> | — | `"constant_and_balanced"` | — |
  | <span id="featuresubsamplingmethod--gini-feature-importance" /><code className="python-reference-parameter">GINI\_<wbr />FEATURE\_<wbr />IMPORTANCE</code> | — | `"gini_feature_importance"` | — |
  | <span id="featuresubsamplingmethod--auto" /><code className="python-reference-parameter">AUTO</code> | — | `"auto"` | — |
</div>

***

<div className="python-reference-heading">
  <h2 id="preprocessorconfig">
    `PreprocessorConfig`
  </h2>

  <a className="python-reference-source" href="https://github.com/PriorLabs/TabPFN/blob/c70b6ef0488858d32244c52222abfc0c5be207d6/src/tabpfn/preprocessing/configs.py#L37" aria-label="View source for PreprocessorConfig"><span aria-hidden="true">\</></span> View source <span aria-hidden="true">↗</span></a>
</div>

Configuration for data preprocessing.

```python theme={null}
PreprocessorConfig(
    name: Literal["power", "safepower", "power_box", "safepower_box", "quantile_uni_coarse", "quantile_norm_coarse", "quantile_uni", "quantile_norm", "quantile_uni_fine", "quantile_norm_fine", "quantile_uni_extrapolate", "squashing_scaler_default", "squashing_scaler_max10", "robust", "kdi", "none", "kdi_random_alpha", "kdi_uni", "kdi_random_alpha_uni", "adaptive", "norm_and_kdi", "kdi_alpha_0.3_uni", "kdi_alpha_0.5_uni", "kdi_alpha_0.8_uni", "kdi_alpha_1.0_uni", "kdi_alpha_1.2_uni", "kdi_alpha_1.5_uni", "kdi_alpha_2.0_uni", "kdi_alpha_3.0_uni", "kdi_alpha_5.0_uni", "kdi_alpha_0.3", "kdi_alpha_0.5", "kdi_alpha_0.8", "kdi_alpha_1.0", "kdi_alpha_1.2", "kdi_alpha_1.5", "kdi_alpha_2.0", "kdi_alpha_3.0", "kdi_alpha_5.0"],
    categorical_name: Literal["none", "numeric", "onehot", "ordinal", "ordinal_shuffled", "ordinal_very_common_categories_shuffled"] = "none",
    append_original: bool | Literal["auto"] = False,
    max_features_per_estimator: int = 500,
    global_transformer_name: Literal["svd", "svd_quarter_components"] | None = None,
    max_onehot_cardinality: int | None = None,
    differentiable: bool = False,
)
```

**Fields**

<div className="python-reference-table">
  | Name | Type | Default | Description |
  | - | - | - | - |
  | <span id="preprocessorconfig--name" /><code className="python-reference-parameter">name</code> | <code className="python-reference-type">Literal\["power", "safepower", "power\_box", "safepower\_box", "quantile\_uni\_coarse", "quantile\_norm\_coarse", "quantile\_uni", "quantile\_norm", "quantile\_uni\_fine", "quantile\_norm\_fine", "quantile\_uni\_extrapolate", "squashing\_scaler\_default", "squashing\_scaler\_max10", "robust", "kdi", "none", "kdi\_random\_alpha", "kdi\_uni", "kdi\_random\_alpha\_uni", "adaptive", "norm\_and\_kdi", "kdi\_alpha\_0.3\_uni", "kdi\_alpha\_0.5\_uni", "kdi\_alpha\_0.8\_uni", "kdi\_alpha\_1.0\_uni", "kdi\_alpha\_1.2\_uni", "kdi\_alpha\_1.5\_uni", "kdi\_alpha\_2.0\_uni", "kdi\_alpha\_3.0\_uni", "kdi\_alpha\_5.0\_uni", "kdi\_alpha\_0.3", "kdi\_alpha\_0.5", "kdi\_alpha\_0.8", "kdi\_alpha\_1.0", "kdi\_alpha\_1.2", "kdi\_alpha\_1.5", "kdi\_alpha\_2.0", "kdi\_alpha\_3.0", "kdi\_alpha\_5.0"]</code> | Required | Name of the preprocessor. |
  | <span id="preprocessorconfig--categorical-name" /><code className="python-reference-parameter">categorical\_<wbr />name</code> | <code className="python-reference-type">Literal\["none", "numeric", "onehot", "ordinal", "ordinal\_shuffled", "ordinal\_very\_common\_categories\_shuffled"]</code> | `"none"` | Name of the categorical encoding method. Options: "none", "numeric", "onehot", "ordinal", "ordinal\_shuffled", "none". |
  | <span id="preprocessorconfig--append-original" /><code className="python-reference-parameter">append\_<wbr />original</code> | <code className="python-reference-type">bool \| Literal\["auto"]</code> | `False` | — |
  | <span id="preprocessorconfig--max-features-per-estimator" /><code className="python-reference-parameter">max\_<wbr />features\_<wbr />per\_<wbr />estimator</code> | <code className="python-reference-type">int</code> | `500` | Maximum number of features per estimator. In case the dataset has more features than this, the features are subsampled for each estimator independently. If append to original is set to `True` we can still have more features. |
  | <span id="preprocessorconfig--global-transformer-name" /><code className="python-reference-parameter">global\_<wbr />transformer\_<wbr />name</code> | <code className="python-reference-type">Literal\["svd", "svd\_quarter\_components"] \| None</code> | `None` | Name of the global transformer to use. |
  | <span id="preprocessorconfig--max-onehot-cardinality" /><code className="python-reference-parameter">max\_<wbr />onehot\_<wbr />cardinality</code> | <code className="python-reference-type">int \| None</code> | `None` | Maximum number of unique values a categorical feature can have to be one-hot encoded. Features with higher cardinality are passed through unchanged to ordinal encoding. If `None`, all categorical features are one-hot encoded. |
  | <span id="preprocessorconfig--differentiable" /><code className="python-reference-parameter">differentiable</code> | <code className="python-reference-type">bool</code> | `False` | — |
</div>

***

<div className="python-reference-heading">
  <h2 id="regressorensembleconfig">
    `RegressorEnsembleConfig`
  </h2>

  <a className="python-reference-source" href="https://github.com/PriorLabs/TabPFN/blob/c70b6ef0488858d32244c52222abfc0c5be207d6/src/tabpfn/preprocessing/configs.py#L186" aria-label="View source for RegressorEnsembleConfig"><span aria-hidden="true">\</></span> View source <span aria-hidden="true">↗</span></a>
</div>

Configuration for a regression ensemble member.

```python theme={null}
RegressorEnsembleConfig(
    preprocess_config: PreprocessorConfig,
    add_fingerprint_feature: bool,
    polynomial_features: Literal["no", "all"] | int,
    feature_shift_count: int,
    feature_shift_decoder: Literal["shuffle", "rotate"] | None,
    outlier_removal_std: float | None,
    _model_index: int,
    passthrough_inf: bool,
    target_transform: TransformerMixin | Pipeline | None,
)
```

**Fields**

<div className="python-reference-table">
  | Name | Type | Default | Description |
  | - | - | - | - |
  | <span id="regressorensembleconfig--preprocess-config" /><code className="python-reference-parameter">preprocess\_<wbr />config</code> | <code className="python-reference-type"><a href="/api-reference/python/tabpfn/preprocessing/configuration#preprocessorconfig">Preprocessor<wbr />Config</a></code> | Required | Preprocessor configuration to use. |
  | <span id="regressorensembleconfig--add-fingerprint-feature" /><code className="python-reference-parameter">add\_<wbr />fingerprint\_<wbr />feature</code> | <code className="python-reference-type">bool</code> | Required | Whether to add fingerprint features. |
  | <span id="regressorensembleconfig--polynomial-features" /><code className="python-reference-parameter">polynomial\_<wbr />features</code> | <code className="python-reference-type">Literal\["no", "all"] \| int</code> | Required | Maximum number of polynomial features to add, if any. |
  | <span id="regressorensembleconfig--feature-shift-count" /><code className="python-reference-parameter">feature\_<wbr />shift\_<wbr />count</code> | <code className="python-reference-type">int</code> | Required | How much to shift the features columns. |
  | <span id="regressorensembleconfig--feature-shift-decoder" /><code className="python-reference-parameter">feature\_<wbr />shift\_<wbr />decoder</code> | <code className="python-reference-type">Literal\["shuffle", "rotate"] \| None</code> | Required | How to shift features. |
  | <span id="regressorensembleconfig--outlier-removal-std" /><code className="python-reference-parameter">outlier\_<wbr />removal\_<wbr />std</code> | <code className="python-reference-type">float \| None</code> | Required | Number of standard deviations from the mean to consider a sample an outlier. If `None`, no outliers are removed. |
  | <span id="regressorensembleconfig--passthrough-inf" /><code className="python-reference-parameter">passthrough\_<wbr />inf</code> | <code className="python-reference-type">bool</code> | Required | Whether to pass infinite values through to the model. When `True`, the preprocessing pipeline replaces infinities with NaN before preprocessing and restores them afterwards. |
  | <span id="regressorensembleconfig--target-transform" /><code className="python-reference-parameter">target\_<wbr />transform</code> | <code className="python-reference-type">Transformer<wbr />Mixin \| Pipeline \| None</code> | Required | — |
</div>

***

<div className="python-reference-heading">
  <h2 id="samplesubsamplingmethod">
    `SampleSubsamplingMethod`
  </h2>

  <a className="python-reference-source" href="https://github.com/PriorLabs/TabPFN/blob/c70b6ef0488858d32244c52222abfc0c5be207d6/src/tabpfn/preprocessing/configs.py#L28" aria-label="View source for SampleSubsamplingMethod"><span aria-hidden="true">\</></span> View source <span aria-hidden="true">↗</span></a>
</div>

Method for subsampling rows per estimator when SUBSAMPLE\_SAMPLES is set.

**Fields**

<div className="python-reference-table">
  | Name | Type | Default | Description |
  | - | - | - | - |
  | <span id="samplesubsamplingmethod--auto" /><code className="python-reference-parameter">AUTO</code> | — | `"auto"` | — |
  | <span id="samplesubsamplingmethod--balanced" /><code className="python-reference-parameter">BALANCED</code> | — | `"balanced"` | — |
  | <span id="samplesubsamplingmethod--stratified" /><code className="python-reference-parameter">STRATIFIED</code> | — | `"stratified"` | — |
  | <span id="samplesubsamplingmethod--majority-downsample" /><code className="python-reference-parameter">MAJORITY\_<wbr />DOWNSAMPLE</code> | — | `"majority_downsample"` | — |
</div>


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.