> ## Documentation Index
> Fetch the complete documentation index at: https://docs.priorlabs.ai/llms.txt
> Use this file to discover all available pages before exploring further.

> ## Agent Instructions
> TabPFN-3 and later versions support up to 1,000,000 rows, subject to feature count and checkpoint/API limits.
> For large datasets, use per-estimator subsampling, e.g. inference_config={"SUBSAMPLE_SAMPLES": 100_000}, in TabPFNClassifier or TabPFNRegressor from tabpfn or tabpfn_client.
> Pass raw pandas DataFrames to fit and predict. Categorical strings/categories, raw text, and missing feature values are handled automatically; no manual integer/one-hot encoding, imputation, scaling, or outlier removal is needed.

# Preprocessing pipelines

> Modality-aware preprocessing pipeline that handles column slicing.

<div className="python-reference-heading">
  <h2 id="preprocessingpipeline">
    `PreprocessingPipeline`
  </h2>

  <a className="python-reference-source" href="https://github.com/PriorLabs/TabPFN/blob/c70b6ef0488858d32244c52222abfc0c5be207d6/src/tabpfn/preprocessing/pipeline_interface.py#L329" aria-label="View source for PreprocessingPipeline"><span aria-hidden="true">\</></span> View source <span aria-hidden="true">↗</span></a>
</div>

Modality-aware preprocessing pipeline that handles column slicing.

This pipeline applies a sequence of preprocessing steps to data,
where each step can be registered to target specific feature modalities.
The pipeline handles slicing columns based on registered modalities,
passing only relevant columns to each step, reassembling data after each
step, and tracking feature schema updates.

For backwards compatibility, steps can be registered as (step, modalities)
tuples where the step receives only columns matching the specified modalities,
or as bare steps that receive all columns. In the latter case, the column modalities
need to be tracked by the step itself. Going forward, only the former format will be
supported.

Initialize the pipeline with preprocessing steps.

```python theme={null}
PreprocessingPipeline(
    steps: list[PreprocessingStep | StepWithModalities],
)
```

**Parameters**

<div className="python-reference-table">
  | Parameter | Type | Default | Description |
  | - | - | - | - |
  | <span id="preprocessingpipeline--steps" /><code className="python-reference-parameter">steps</code> | <code className="python-reference-type">list\[<a href="https://github.com/PriorLabs/TabPFN/blob/c70b6ef0488858d32244c52222abfc0c5be207d6/src/tabpfn/preprocessing/pipeline_interface.py#L153">Preprocessing<wbr />Step</a> \| <a href="https://github.com/PriorLabs/TabPFN/blob/c70b6ef0488858d32244c52222abfc0c5be207d6/src/tabpfn/preprocessing/pipeline_interface.py#L17">Step<wbr />With<wbr />Modalities</a>]</code> | Required | List of preprocessing steps. Each can be a [`PreprocessingStep`](https://github.com/PriorLabs/TabPFN/blob/c70b6ef0488858d32244c52222abfc0c5be207d6/src/tabpfn/preprocessing/pipeline_interface.py#L153) (receives all columns) or a tuple of ([`PreprocessingStep`](https://github.com/PriorLabs/TabPFN/blob/c70b6ef0488858d32244c52222abfc0c5be207d6/src/tabpfn/preprocessing/pipeline_interface.py#L153), set\[FeatureModality]) where the step receives only columns matching the specified modalities. |
</div>

***

<div className="python-reference-heading">
  <h2 id="preprocessingpipeline-fit-transform">
    `PreprocessingPipeline.fit_transform`
  </h2>

  <a className="python-reference-source" href="https://github.com/PriorLabs/TabPFN/blob/c70b6ef0488858d32244c52222abfc0c5be207d6/src/tabpfn/preprocessing/pipeline_interface.py#L375" aria-label="View source for PreprocessingPipeline.fit_transform"><span aria-hidden="true">\</></span> View source <span aria-hidden="true">↗</span></a>
</div>

Fit and transform the data using the pipeline.

```python theme={null}
PreprocessingPipeline.fit_transform(
    X: np.ndarray | torch.Tensor,
    feature_schema: FeatureSchema,
) -> PreprocessingPipelineResult
```

**Parameters**

<div className="python-reference-table">
  | Parameter | Type | Default | Description |
  | - | - | - | - |
  | <span id="preprocessingpipeline-fit-transform--x" /><code className="python-reference-parameter">X</code> | <code className="python-reference-type"><a href="https://numpy.org/doc/stable/reference/generated/numpy.ndarray.html">np.ndarray</a> \| torch.Tensor</code> | Required | 2d array of shape (n\_samples, n\_features). |
  | <span id="preprocessingpipeline-fit-transform--feature-schema" /><code className="python-reference-parameter">feature\_<wbr />schema</code> | <code className="python-reference-type"><a href="https://github.com/PriorLabs/TabPFN/blob/c70b6ef0488858d32244c52222abfc0c5be207d6/src/tabpfn/preprocessing/datamodel.py#L145">Feature<wbr />Schema</a></code> | Required | feature schema. |
</div>

**Returns**

<div className="python-reference-table python-reference-returns">
  | Type | Description |
  | - | - |
  | <code className="python-reference-type"><a href="/api-reference/python/tabpfn/preprocessing/pipeline#preprocessingpipelineresult">Preprocessing<wbr />Pipeline<wbr />Result</a></code> | [`PreprocessingPipelineResult`](/api-reference/python/tabpfn/preprocessing/pipeline#preprocessingpipelineresult) with transformed data and updated schema. |
</div>

***

<div className="python-reference-heading">
  <h2 id="preprocessingpipeline-transform">
    `PreprocessingPipeline.transform`
  </h2>

  <a className="python-reference-source" href="https://github.com/PriorLabs/TabPFN/blob/c70b6ef0488858d32244c52222abfc0c5be207d6/src/tabpfn/preprocessing/pipeline_interface.py#L394" aria-label="View source for PreprocessingPipeline.transform"><span aria-hidden="true">\</></span> View source <span aria-hidden="true">↗</span></a>
</div>

Transform the data using the fitted pipeline.

```python theme={null}
PreprocessingPipeline.transform(
    X: np.ndarray | torch.Tensor,
) -> PreprocessingPipelineResult
```

**Parameters**

<div className="python-reference-table">
  | Parameter | Type | Default | Description |
  | - | - | - | - |
  | <span id="preprocessingpipeline-transform--x" /><code className="python-reference-parameter">X</code> | <code className="python-reference-type"><a href="https://numpy.org/doc/stable/reference/generated/numpy.ndarray.html">np.ndarray</a> \| torch.Tensor</code> | Required | 2d array of shape (n\_samples, n\_features). |
</div>

**Returns**

<div className="python-reference-table python-reference-returns">
  | Type | Description |
  | - | - |
  | <code className="python-reference-type"><a href="/api-reference/python/tabpfn/preprocessing/pipeline#preprocessingpipelineresult">Preprocessing<wbr />Pipeline<wbr />Result</a></code> | [`PreprocessingPipelineResult`](/api-reference/python/tabpfn/preprocessing/pipeline#preprocessingpipelineresult) with transformed data and feature schema. |
</div>

***

<div className="python-reference-heading">
  <h2 id="preprocessingpipeline-num-added-features">
    `PreprocessingPipeline.num_added_features`
  </h2>

  <a className="python-reference-source" href="https://github.com/PriorLabs/TabPFN/blob/c70b6ef0488858d32244c52222abfc0c5be207d6/src/tabpfn/preprocessing/pipeline_interface.py#L413" aria-label="View source for PreprocessingPipeline.num_added_features"><span aria-hidden="true">\</></span> View source <span aria-hidden="true">↗</span></a>
</div>

Return the number of added features.

Threads an evolving feature schema through the steps so that each step
sees the feature count that includes columns added by prior steps.

```python theme={null}
PreprocessingPipeline.num_added_features(
    n_samples: int,
    feature_schema: FeatureSchema,
) -> int
```

**Parameters**

<div className="python-reference-table">
  | Parameter | Type | Default | Description |
  | - | - | - | - |
  | <span id="preprocessingpipeline-num-added-features--n-samples" /><code className="python-reference-parameter">n\_<wbr />samples</code> | <code className="python-reference-type">int</code> | Required | — |
  | <span id="preprocessingpipeline-num-added-features--feature-schema" /><code className="python-reference-parameter">feature\_<wbr />schema</code> | <code className="python-reference-type"><a href="https://github.com/PriorLabs/TabPFN/blob/c70b6ef0488858d32244c52222abfc0c5be207d6/src/tabpfn/preprocessing/datamodel.py#L145">Feature<wbr />Schema</a></code> | Required | — |
</div>

**Returns**

<div className="python-reference-table python-reference-returns">
  | Type | Description |
  | - | - |
  | <code className="python-reference-type">int</code> | — |
</div>

***

<div className="python-reference-heading">
  <h2 id="preprocessingpipeline-has-data-dependent-feature-expansion">
    `PreprocessingPipeline.has_data_dependent_feature_expansion`
  </h2>

  <a className="python-reference-source" href="https://github.com/PriorLabs/TabPFN/blob/c70b6ef0488858d32244c52222abfc0c5be207d6/src/tabpfn/preprocessing/pipeline_interface.py#L430" aria-label="View source for PreprocessingPipeline.has_data_dependent_feature_expansion"><span aria-hidden="true">\</></span> View source <span aria-hidden="true">↗</span></a>
</div>

Return `True` if any step has data-dependent feature expansion.

```python theme={null}
PreprocessingPipeline.has_data_dependent_feature_expansion() -> bool
```

**Returns**

<div className="python-reference-table python-reference-returns">
  | Type | Description |
  | - | - |
  | <code className="python-reference-type">bool</code> | — |
</div>

***

<div className="python-reference-heading">
  <h2 id="clean-data">
    `clean_data`
  </h2>

  <a className="python-reference-source" href="https://github.com/PriorLabs/TabPFN/blob/c70b6ef0488858d32244c52222abfc0c5be207d6/src/tabpfn/preprocessing/clean.py#L168" aria-label="View source for clean_data"><span aria-hidden="true">\</></span> View source <span aria-hidden="true">↗</span></a>
</div>

Clean the data by converting dtypes and ordinally encoding categorical columns.

```python theme={null}
clean_data(
    X: np.ndarray,
    feature_schema: FeatureSchema,
    *,
    passthrough_inf: bool = False,
) -> tuple[np.ndarray, OrderPreservingColumnTransformer, FeatureSchema]
```

**Parameters**

<div className="python-reference-table">
  | Parameter | Type | Default | Description |
  | - | - | - | - |
  | <span id="clean-data--x" /><code className="python-reference-parameter">X</code> | <code className="python-reference-type"><a href="https://numpy.org/doc/stable/reference/generated/numpy.ndarray.html">np.ndarray</a></code> | Required | The data to clean. |
  | <span id="clean-data--feature-schema" /><code className="python-reference-parameter">feature\_<wbr />schema</code> | <code className="python-reference-type"><a href="https://github.com/PriorLabs/TabPFN/blob/c70b6ef0488858d32244c52222abfc0c5be207d6/src/tabpfn/preprocessing/datamodel.py#L145">Feature<wbr />Schema</a></code> | Required | The feature schema corresponding to the data. |
  | <span id="clean-data--passthrough-inf" /><code className="python-reference-parameter">passthrough\_<wbr />inf</code> | <code className="python-reference-type">bool</code> | `False` | If `True`, +/-inf values are carried through the ordinal encoding stage unchanged instead of crashing it (see `process_text_na_dataframe`). |
</div>

**Returns**

<div className="python-reference-table python-reference-returns">
  | Type | Description |
  | - | - |
  | <code className="python-reference-type">tuple\[<a href="https://numpy.org/doc/stable/reference/generated/numpy.ndarray.html">np.ndarray</a>, <a href="https://github.com/PriorLabs/TabPFN/blob/c70b6ef0488858d32244c52222abfc0c5be207d6/src/tabpfn/preprocessing/steps/preprocessing_helpers.py#L471">Order<wbr />Preserving<wbr />Column<wbr />Transformer</a>, <a href="https://github.com/PriorLabs/TabPFN/blob/c70b6ef0488858d32244c52222abfc0c5be207d6/src/tabpfn/preprocessing/datamodel.py#L145">Feature<wbr />Schema</a>]</code> | A tuple containing the cleaned data, the ordinal encoder, and the inferred feature modalities. |
</div>

***

<div className="python-reference-heading">
  <h2 id="fit-preprocessing">
    `fit_preprocessing`
  </h2>

  <a className="python-reference-source" href="https://github.com/PriorLabs/TabPFN/blob/c70b6ef0488858d32244c52222abfc0c5be207d6/src/tabpfn/preprocessing/transform.py#L115" aria-label="View source for fit_preprocessing"><span aria-hidden="true">\</></span> View source <span aria-hidden="true">↗</span></a>
</div>

Fit preprocessing pipelines in parallel.

```python theme={null}
fit_preprocessing(
    configs: Sequence[EnsembleConfig],
    X_train: np.ndarray,
    y_train: np.ndarray,
    *,
    feature_schema: FeatureSchema,
    n_preprocessing_jobs: int,
    parallel_mode: Literal["block", "as-ready", "in-order"],
    pipelines: Sequence[PreprocessingPipeline],
    subsample_feature_indices: list[np.ndarray | None] | None = None,
    subsample_row_indices: list[np.ndarray] | None = None,
) -> Iterator[tuple[int, EnsembleConfig, PreprocessingPipeline, np.ndarray, np.ndarray, FeatureSchema]]
```

**Parameters**

<div className="python-reference-table">
  | Parameter | Type | Default | Description |
  | - | - | - | - |
  | <span id="fit-preprocessing--configs" /><code className="python-reference-parameter">configs</code> | <code className="python-reference-type">Sequence\[<a href="/api-reference/python/tabpfn/preprocessing/configuration#ensembleconfig">Ensemble<wbr />Config</a>]</code> | Required | List of ensemble configurations. |
  | <span id="fit-preprocessing--x-train" /><code className="python-reference-parameter">X\_<wbr />train</code> | <code className="python-reference-type"><a href="https://numpy.org/doc/stable/reference/generated/numpy.ndarray.html">np.ndarray</a></code> | Required | Training data. |
  | <span id="fit-preprocessing--y-train" /><code className="python-reference-parameter">y\_<wbr />train</code> | <code className="python-reference-type"><a href="https://numpy.org/doc/stable/reference/generated/numpy.ndarray.html">np.ndarray</a></code> | Required | Training target. |
  | <span id="fit-preprocessing--feature-schema" /><code className="python-reference-parameter">feature\_<wbr />schema</code> | <code className="python-reference-type"><a href="https://github.com/PriorLabs/TabPFN/blob/c70b6ef0488858d32244c52222abfc0c5be207d6/src/tabpfn/preprocessing/datamodel.py#L145">Feature<wbr />Schema</a></code> | Required | feature schema. |
  | <span id="fit-preprocessing--n-preprocessing-jobs" /><code className="python-reference-parameter">n\_<wbr />preprocessing\_<wbr />jobs</code> | <code className="python-reference-type">int</code> | Required | Number of worker processes to use. If `1`, then the preprocessing is performed in the current process. This     avoids multiprocessing overheads, but may not be able to full saturate     the CPU. Note that the preprocessing itself will parallelise over     multiple cores, so one job is often enough. If `>1`, then different estimators are dispatched to different proceses,     which allows more parallelism but incurs some overhead. If `-1`, then creates as many workers as CPU cores. As each worker itself     uses multiple cores, this is likely too many. It is best to select this value by benchmarking. |
  | <span id="fit-preprocessing--parallel-mode" /><code className="python-reference-parameter">parallel\_<wbr />mode</code> | <code className="python-reference-type">Literal\["block", "as-ready", "in-order"]</code> | Required | Parallel mode to use.<br /><br />\* `"block"`: Blocks until all workers are done. Returns in order.<br />\* `"as-ready"`: Returns results as they are ready. Any order.<br />\* `"in-order"`: Returns results in order, blocking only in the order that     needs to be returned in. |
  | <span id="fit-preprocessing--pipelines" /><code className="python-reference-parameter">pipelines</code> | <code className="python-reference-type">Sequence\[<a href="/api-reference/python/tabpfn/preprocessing/pipeline#preprocessingpipeline">Preprocessing<wbr />Pipeline</a>]</code> | Required | Preprocessing pipelines, one per configuration. |
  | <span id="fit-preprocessing--subsample-feature-indices" /><code className="python-reference-parameter">subsample\_<wbr />feature\_<wbr />indices</code> | <code className="python-reference-type">list\[<a href="https://numpy.org/doc/stable/reference/generated/numpy.ndarray.html">np.ndarray</a> \| None] \| None</code> | `None` | Indices of features to subsample. If not provided, no features are subsampled. |
  | <span id="fit-preprocessing--subsample-row-indices" /><code className="python-reference-parameter">subsample\_<wbr />row\_<wbr />indices</code> | <code className="python-reference-type">list\[<a href="https://numpy.org/doc/stable/reference/generated/numpy.ndarray.html">np.ndarray</a>] \| None</code> | `None` | Indices of rows to subsample per estimator. If not provided, no row subsampling is done. |
</div>

**Returns**

<div className="python-reference-table python-reference-returns">
  | Type | Description |
  | - | - |
  | <code className="python-reference-type">Iterator\[tuple\[int, <a href="/api-reference/python/tabpfn/preprocessing/configuration#ensembleconfig">Ensemble<wbr />Config</a>, <a href="/api-reference/python/tabpfn/preprocessing/pipeline#preprocessingpipeline">Preprocessing<wbr />Pipeline</a>, <a href="https://numpy.org/doc/stable/reference/generated/numpy.ndarray.html">np.ndarray</a>, <a href="https://numpy.org/doc/stable/reference/generated/numpy.ndarray.html">np.ndarray</a>, <a href="https://github.com/PriorLabs/TabPFN/blob/c70b6ef0488858d32244c52222abfc0c5be207d6/src/tabpfn/preprocessing/datamodel.py#L145">Feature<wbr />Schema</a>]]</code> | Iterator of tuples containing the config index, ensemble configuration, the fitted preprocessing pipeline, the transformed training data, the transformed target, and the feature schema. |
</div>

***

<div className="python-reference-heading">
  <h2 id="generate-classification-ensemble-configs">
    `generate_classification_ensemble_configs`
  </h2>

  <a className="python-reference-source" href="https://github.com/PriorLabs/TabPFN/blob/c70b6ef0488858d32244c52222abfc0c5be207d6/src/tabpfn/preprocessing/ensemble.py#L1279" aria-label="View source for generate_classification_ensemble_configs"><span aria-hidden="true">\</></span> View source <span aria-hidden="true">↗</span></a>
</div>

Generate ensemble configurations for classification.

```python theme={null}
generate_classification_ensemble_configs(
    *,
    num_estimators: int,
    add_fingerprint_feature: bool,
    polynomial_features: Literal["no", "all"] | int,
    feature_shift_decoder: Literal["shuffle", "rotate"] | None,
    preprocessor_configs: Sequence[PreprocessorConfig],
    class_shift_method: Literal["rotate", "shuffle"] | None,
    n_classes: int,
    random_state: int | np.random.Generator | None,
    num_models: int,
    outlier_removal_std: float | None,
    passthrough_inf: bool = False,
) -> list[ClassifierEnsembleConfig]
```

**Parameters**

<div className="python-reference-table">
  | Parameter | Type | Default | Description |
  | - | - | - | - |
  | <span id="generate-classification-ensemble-configs--num-estimators" /><code className="python-reference-parameter">num\_<wbr />estimators</code> | <code className="python-reference-type">int</code> | Required | Number of ensemble configurations to generate. |
  | <span id="generate-classification-ensemble-configs--add-fingerprint-feature" /><code className="python-reference-parameter">add\_<wbr />fingerprint\_<wbr />feature</code> | <code className="python-reference-type">bool</code> | Required | Whether to add fingerprint features. |
  | <span id="generate-classification-ensemble-configs--polynomial-features" /><code className="python-reference-parameter">polynomial\_<wbr />features</code> | <code className="python-reference-type">Literal\["no", "all"] \| int</code> | Required | Maximum number of polynomial features to add, if any. |
  | <span id="generate-classification-ensemble-configs--feature-shift-decoder" /><code className="python-reference-parameter">feature\_<wbr />shift\_<wbr />decoder</code> | <code className="python-reference-type">Literal\["shuffle", "rotate"] \| None</code> | Required | How shift features |
  | <span id="generate-classification-ensemble-configs--preprocessor-configs" /><code className="python-reference-parameter">preprocessor\_<wbr />configs</code> | <code className="python-reference-type">Sequence\[<a href="/api-reference/python/tabpfn/preprocessing/configuration#preprocessorconfig">Preprocessor<wbr />Config</a>]</code> | Required | Preprocessor configurations to use on the data. |
  | <span id="generate-classification-ensemble-configs--class-shift-method" /><code className="python-reference-parameter">class\_<wbr />shift\_<wbr />method</code> | <code className="python-reference-type">Literal\["rotate", "shuffle"] \| None</code> | Required | How to shift classes for classpermutation. |
  | <span id="generate-classification-ensemble-configs--n-classes" /><code className="python-reference-parameter">n\_<wbr />classes</code> | <code className="python-reference-type">int</code> | Required | Number of classes. |
  | <span id="generate-classification-ensemble-configs--random-state" /><code className="python-reference-parameter">random\_<wbr />state</code> | <code className="python-reference-type">int \| np.random.Generator \| None</code> | Required | Random number generator. |
  | <span id="generate-classification-ensemble-configs--num-models" /><code className="python-reference-parameter">num\_<wbr />models</code> | <code className="python-reference-type">int</code> | Required | Number of models to use. |
  | <span id="generate-classification-ensemble-configs--outlier-removal-std" /><code className="python-reference-parameter">outlier\_<wbr />removal\_<wbr />std</code> | <code className="python-reference-type">float \| None</code> | Required | The standard deviation to remove outliers. |
  | <span id="generate-classification-ensemble-configs--passthrough-inf" /><code className="python-reference-parameter">passthrough\_<wbr />inf</code> | <code className="python-reference-type">bool</code> | `False` | Whether to pass infinite values through to the model. |
</div>

**Returns**

<div className="python-reference-table python-reference-returns">
  | Type | Description |
  | - | - |
  | <code className="python-reference-type">list\[<a href="/api-reference/python/tabpfn/preprocessing/configuration#classifierensembleconfig">Classifier<wbr />Ensemble<wbr />Config</a>]</code> | List of ensemble configurations. |
</div>

***

<div className="python-reference-heading">
  <h2 id="generate-regression-ensemble-configs">
    `generate_regression_ensemble_configs`
  </h2>

  <a className="python-reference-source" href="https://github.com/PriorLabs/TabPFN/blob/c70b6ef0488858d32244c52222abfc0c5be207d6/src/tabpfn/preprocessing/ensemble.py#L1358" aria-label="View source for generate_regression_ensemble_configs"><span aria-hidden="true">\</></span> View source <span aria-hidden="true">↗</span></a>
</div>

Generate ensemble configurations for regression.

```python theme={null}
generate_regression_ensemble_configs(
    *,
    num_estimators: int,
    add_fingerprint_feature: bool,
    polynomial_features: Literal["no", "all"] | int,
    feature_shift_decoder: Literal["shuffle", "rotate"] | None,
    preprocessor_configs: Sequence[PreprocessorConfig],
    target_transforms: Sequence[TransformerMixin | Pipeline | None],
    random_state: int | np.random.Generator | None,
    num_models: int,
    outlier_removal_std: float | None,
    passthrough_inf: bool = False,
) -> list[RegressorEnsembleConfig]
```

**Parameters**

<div className="python-reference-table">
  | Parameter | Type | Default | Description |
  | - | - | - | - |
  | <span id="generate-regression-ensemble-configs--num-estimators" /><code className="python-reference-parameter">num\_<wbr />estimators</code> | <code className="python-reference-type">int</code> | Required | Number of ensemble configurations to generate. |
  | <span id="generate-regression-ensemble-configs--add-fingerprint-feature" /><code className="python-reference-parameter">add\_<wbr />fingerprint\_<wbr />feature</code> | <code className="python-reference-type">bool</code> | Required | Whether to add fingerprint features. |
  | <span id="generate-regression-ensemble-configs--polynomial-features" /><code className="python-reference-parameter">polynomial\_<wbr />features</code> | <code className="python-reference-type">Literal\["no", "all"] \| int</code> | Required | Maximum number of polynomial features to add, if any. |
  | <span id="generate-regression-ensemble-configs--feature-shift-decoder" /><code className="python-reference-parameter">feature\_<wbr />shift\_<wbr />decoder</code> | <code className="python-reference-type">Literal\["shuffle", "rotate"] \| None</code> | Required | How shift features |
  | <span id="generate-regression-ensemble-configs--preprocessor-configs" /><code className="python-reference-parameter">preprocessor\_<wbr />configs</code> | <code className="python-reference-type">Sequence\[<a href="/api-reference/python/tabpfn/preprocessing/configuration#preprocessorconfig">Preprocessor<wbr />Config</a>]</code> | Required | Preprocessor configurations to use on the data. |
  | <span id="generate-regression-ensemble-configs--target-transforms" /><code className="python-reference-parameter">target\_<wbr />transforms</code> | <code className="python-reference-type">Sequence\[Transformer<wbr />Mixin \| Pipeline \| None]</code> | Required | Target transformations to apply. |
  | <span id="generate-regression-ensemble-configs--random-state" /><code className="python-reference-parameter">random\_<wbr />state</code> | <code className="python-reference-type">int \| np.random.Generator \| None</code> | Required | Random number generator. |
  | <span id="generate-regression-ensemble-configs--num-models" /><code className="python-reference-parameter">num\_<wbr />models</code> | <code className="python-reference-type">int</code> | Required | Number of models to use. |
  | <span id="generate-regression-ensemble-configs--outlier-removal-std" /><code className="python-reference-parameter">outlier\_<wbr />removal\_<wbr />std</code> | <code className="python-reference-type">float \| None</code> | Required | The standard deviation to remove outliers. |
  | <span id="generate-regression-ensemble-configs--passthrough-inf" /><code className="python-reference-parameter">passthrough\_<wbr />inf</code> | <code className="python-reference-type">bool</code> | `False` | Whether to pass infinite values through to the model. |
</div>

**Returns**

<div className="python-reference-table python-reference-returns">
  | Type | Description |
  | - | - |
  | <code className="python-reference-type">list\[<a href="/api-reference/python/tabpfn/preprocessing/configuration#regressorensembleconfig">Regressor<wbr />Ensemble<wbr />Config</a>]</code> | List of ensemble configurations. |
</div>

***

<div className="python-reference-heading">
  <h2 id="preprocessingpipelineresult">
    `PreprocessingPipelineResult`
  </h2>

  <a className="python-reference-source" href="https://github.com/PriorLabs/TabPFN/blob/c70b6ef0488858d32244c52222abfc0c5be207d6/src/tabpfn/preprocessing/pipeline_interface.py#L140" aria-label="View source for PreprocessingPipelineResult"><span aria-hidden="true">\</></span> View source <span aria-hidden="true">↗</span></a>
</div>

Result from the preprocessing pipeline.

```python theme={null}
PreprocessingPipelineResult(
    X: np.ndarray | torch.Tensor,
    feature_schema: FeatureSchema,
)
```

**Fields**

<div className="python-reference-table">
  | Name | Type | Default | Description |
  | - | - | - | - |
  | <span id="preprocessingpipelineresult--x" /><code className="python-reference-parameter">X</code> | <code className="python-reference-type"><a href="https://numpy.org/doc/stable/reference/generated/numpy.ndarray.html">np.ndarray</a> \| torch.Tensor</code> | Required | The transformed array. |
  | <span id="preprocessingpipelineresult--feature-schema" /><code className="python-reference-parameter">feature\_<wbr />schema</code> | <code className="python-reference-type"><a href="https://github.com/PriorLabs/TabPFN/blob/c70b6ef0488858d32244c52222abfc0c5be207d6/src/tabpfn/preprocessing/datamodel.py#L145">Feature<wbr />Schema</a></code> | Required | Updated feature schema (may have new columns added). |
</div>

***

<div className="python-reference-heading">
  <h2 id="preprocessingstepresult">
    `PreprocessingStepResult`
  </h2>

  <a className="python-reference-source" href="https://github.com/PriorLabs/TabPFN/blob/c70b6ef0488858d32244c52222abfc0c5be207d6/src/tabpfn/preprocessing/pipeline_interface.py#L110" aria-label="View source for PreprocessingStepResult"><span aria-hidden="true">\</></span> View source <span aria-hidden="true">↗</span></a>
</div>

Result of a feature preprocessing step.

```python theme={null}
PreprocessingStepResult(
    X: np.ndarray | torch.Tensor,
    feature_schema: FeatureSchema,
    X_added: np.ndarray | torch.Tensor | None = None,
    modality_added: FeatureModality | None = None,
)
```

**Fields**

<div className="python-reference-table">
  | Name | Type | Default | Description |
  | - | - | - | - |
  | <span id="preprocessingstepresult--x" /><code className="python-reference-parameter">X</code> | <code className="python-reference-type"><a href="https://numpy.org/doc/stable/reference/generated/numpy.ndarray.html">np.ndarray</a> \| torch.Tensor</code> | Required | Transformed array. For steps registered with specific modalities, this is only the transformed columns (not the full array). The shape should match the input shape unless columns are removed. |
  | <span id="preprocessingstepresult--feature-schema" /><code className="python-reference-parameter">feature\_<wbr />schema</code> | <code className="python-reference-type"><a href="https://github.com/PriorLabs/TabPFN/blob/c70b6ef0488858d32244c52222abfc0c5be207d6/src/tabpfn/preprocessing/datamodel.py#L145">Feature<wbr />Schema</a></code> | Required | Feature schema for the columns this step processed. Contains 0-based indices relative to the step's input. Should NOT include added\_columns - the pipeline handles that. |
  | <span id="preprocessingstepresult--x-added" /><code className="python-reference-parameter">X\_<wbr />added</code> | <code className="python-reference-type"><a href="https://numpy.org/doc/stable/reference/generated/numpy.ndarray.html">np.ndarray</a> \| torch.Tensor \| None</code> | `None` | Optional new features to append (e.g., fingerprint features). These are handled by the pipeline, which concatenates them and updates the schema accordingly. Steps should NOT concatenate these internally. |
  | <span id="preprocessingstepresult--modality-added" /><code className="python-reference-parameter">modality\_<wbr />added</code> | <code className="python-reference-type"><a href="https://github.com/PriorLabs/TabPFN/blob/c70b6ef0488858d32244c52222abfc0c5be207d6/src/tabpfn/preprocessing/datamodel.py#L90">Feature<wbr />Modality</a> \| None</code> | `None` | Modality for the added features. Required if [`X_added`](/api-reference/python/tabpfn/preprocessing/pipeline#preprocessingstepresult--x-added) is provided. |
</div>


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.