> ## Documentation Index
> Fetch the complete documentation index at: https://docs.priorlabs.ai/llms.txt
> Use this file to discover all available pages before exploring further.

> ## Agent Instructions
> TabPFN-3 and later versions support up to 1,000,000 rows, subject to feature count and checkpoint/API limits.
> For large datasets, use per-estimator subsampling, e.g. inference_config={"SUBSAMPLE_SAMPLES": 100_000}, in TabPFNClassifier or TabPFNRegressor from tabpfn or tabpfn_client.
> Pass raw pandas DataFrames to fit and predict. Categorical strings/categories, raw text, and missing feature values are handled automatically; no manual integer/one-hot encoding, imputation, scaling, or outlier removal is needed.

# Imputation and data generation

> Impute missing values in the input data using the fitted TabPFN models.

<Info>
  Looking for usage documentation? Check out [Data generation](/capabilities/data-generation).
</Info>

<div className="python-reference-heading">
  <h2 id="unsupervised-unsupervised-tabpfnunsupervisedmodel-generate-synthetic-data">
    `TabPFNUnsupervisedModel.generate_synthetic_data`
  </h2>

  <a className="python-reference-source" href="https://github.com/PriorLabs/tabpfn-extensions/blob/840c15a1848a986b39c85bc17efc61e0e377f983/src/tabpfn_extensions/unsupervised/unsupervised.py#L896" aria-label="View source for TabPFNUnsupervisedModel.generate_synthetic_data"><span aria-hidden="true">\</></span> View source <span aria-hidden="true">↗</span></a>
</div>

Generate synthetic tabular data samples using the fitted TabPFN models.

This method uses imputation to create synthetic data, starting with a matrix of NaN
values and filling in each feature sequentially. Samples are generated feature by
feature in a single pass, with each feature conditioned on previously generated features.

```python theme={null}
TabPFNUnsupervisedModel.generate_synthetic_data(
    n_samples: int = 100,
    t: float = 1.0,
    n_permutations: int = 3,
    dag: dict[int, list[int]] | None = None,
) -> torch.Tensor
```

**Parameters**

<div className="python-reference-table">
  | Parameter | Type | Default | Description |
  | - | - | - | - |
  | <span id="unsupervised-unsupervised-tabpfnunsupervisedmodel-generate-synthetic-data--n-samples" /><code className="python-reference-parameter">n\_<wbr />samples</code> | <code className="python-reference-type">int</code> | `100` | int, default=100 Number of synthetic samples to generate |
  | <span id="unsupervised-unsupervised-tabpfnunsupervisedmodel-generate-synthetic-data--t" /><code className="python-reference-parameter">t</code> | <code className="python-reference-type">float</code> | `1.0` | float, default=1.0 Temperature parameter for sampling. Controls randomness:<br />- Higher values (e.g., 1.0) produce more diverse samples<br />- Lower values (e.g., 0.1) produce more deterministic samples |
  | <span id="unsupervised-unsupervised-tabpfnunsupervisedmodel-generate-synthetic-data--n-permutations" /><code className="python-reference-parameter">n\_<wbr />permutations</code> | <code className="python-reference-type">int</code> | `3` | int, default=3 Number of feature permutations to use for generation More permutations may provide more robust results but increase computation time |
  | <span id="unsupervised-unsupervised-tabpfnunsupervisedmodel-generate-synthetic-data--dag" /><code className="python-reference-parameter">dag</code> | <code className="python-reference-type">dict\[int, list\[int]] \| None</code> | `None` | dict\[int, list\[int]] \| `None`, default=`None` Optional Directed Acyclic Graph mapping each column index to the list of column indices it depends on. When provided, columns are generated in topological order and each column is conditioned on its DAG parents only. Every column must appear as a key; map a column to an empty list to sample it marginally. Useful for causally-informed synthesis. |
</div>

**Returns**

<div className="python-reference-table python-reference-returns">
  | Type | Description |
  | - | - |
  | <code className="python-reference-type">torch.Tensor</code> | torch.Tensor:     Generated synthetic data of shape (`n_samples`, n\_features) |
</div>

**Raises**

`AssertionError`

If the model is not fitted (self.X\_ does not exist)

`ValueError`

If `dag` contains a cycle or does not specify every feature

***

<div className="python-reference-heading">
  <h2 id="unsupervised-unsupervised-tabpfnunsupervisedmodel-impute">
    `TabPFNUnsupervisedModel.impute`
  </h2>

  <a className="python-reference-source" href="https://github.com/PriorLabs/tabpfn-extensions/blob/840c15a1848a986b39c85bc17efc61e0e377f983/src/tabpfn_extensions/unsupervised/unsupervised.py#L660" aria-label="View source for TabPFNUnsupervisedModel.impute"><span aria-hidden="true">\</></span> View source <span aria-hidden="true">↗</span></a>
</div>

Impute missing values in the input data using the fitted TabPFN models.

This method fills missing values (np.nan) in the input data by predicting
each missing value based on the observed values in the same sample. The
imputation uses multiple random feature permutations to improve robustness.

```python theme={null}
TabPFNUnsupervisedModel.impute(
    X: torch.Tensor | np.ndarray | pd.DataFrame,
    t: float = 1e-09,
    n_permutations: int = 10,
    dag: dict[int, list[int]] | None = None,
) -> torch.Tensor
```

**Parameters**

<div className="python-reference-table">
  | Parameter | Type | Default | Description |
  | - | - | - | - |
  | <span id="unsupervised-unsupervised-tabpfnunsupervisedmodel-impute--x" /><code className="python-reference-parameter">X</code> | <code className="python-reference-type">torch.Tensor \| <a href="https://numpy.org/doc/stable/reference/generated/numpy.ndarray.html">np.ndarray</a> \| <a href="https://pandas.pydata.org/docs/reference/api/pandas.DataFrame.html">pd.Data<wbr />Frame</a></code> | Required | Union\[torch.Tensor, [`np.ndarray`](https://numpy.org/doc/stable/reference/generated/numpy.ndarray.html), [`pd.DataFrame`](https://pandas.pydata.org/docs/reference/api/pandas.DataFrame.html)] Input data of shape (n\_samples, n\_features) with missing values encoded as np.nan. |
  | <span id="unsupervised-unsupervised-tabpfnunsupervisedmodel-impute--t" /><code className="python-reference-parameter">t</code> | <code className="python-reference-type">float</code> | `1e-09` | float, default=0.000000001 Temperature for sampling from the imputation distribution. Lower values result in more deterministic imputations, while higher values introduce more randomness. |
  | <span id="unsupervised-unsupervised-tabpfnunsupervisedmodel-impute--n-permutations" /><code className="python-reference-parameter">n\_<wbr />permutations</code> | <code className="python-reference-type">int</code> | `10` | int, default=10 Number of random feature permutations to use for imputation. Higher values may improve robustness but increase computation time. |
  | <span id="unsupervised-unsupervised-tabpfnunsupervisedmodel-impute--dag" /><code className="python-reference-parameter">dag</code> | <code className="python-reference-type">dict\[int, list\[int]] \| None</code> | `None` | dict\[int, list\[int]] \| `None`, default=`None` Optional Directed Acyclic Graph mapping each column index to the list of column indices it depends on. When provided, columns are imputed in topological order and each column is conditioned on its DAG parents instead of all other features. Every column must appear as a key; map a column to an empty list to impute it without conditioning. Useful for causally-informed imputation. |
</div>

**Returns**

<div className="python-reference-table python-reference-returns">
  | Type | Description |
  | - | - |
  | <code className="python-reference-type">torch.Tensor</code> | torch.Tensor     Imputed data with missing values replaced, of shape (n\_samples, n\_features). |
</div>

**Note**

The model must be fitted with training data before calling this method.

***

<div className="python-reference-heading">
  <h2 id="unsupervised-unsupervised-tabpfnunsupervisedmodel-impute-underscore">
    `TabPFNUnsupervisedModel.impute_`
  </h2>

  <a className="python-reference-source" href="https://github.com/PriorLabs/tabpfn-extensions/blob/840c15a1848a986b39c85bc17efc61e0e377f983/src/tabpfn_extensions/unsupervised/unsupervised.py#L320" aria-label="View source for TabPFNUnsupervisedModel.impute_"><span aria-hidden="true">\</></span> View source <span aria-hidden="true">↗</span></a>
</div>

Impute missing values (np.nan) in `X` by sampling all cells independently from the trained models.

```python theme={null}
TabPFNUnsupervisedModel.impute_(
    X: torch.Tensor,
    t: float = 1e-09,
    n_permutations: int = 10,
    condition_on_all_features: bool = True,
    dag: dict[int, list[int]] | None = None,
    fast_mode: bool = False,
) -> torch.Tensor
```

**Parameters**

<div className="python-reference-table">
  | Parameter | Type | Default | Description |
  | - | - | - | - |
  | <span id="unsupervised-unsupervised-tabpfnunsupervisedmodel-impute-underscore--x" /><code className="python-reference-parameter">X</code> | <code className="python-reference-type">torch.Tensor</code> | Required | torch.Tensor Input data of shape (n\_samples, n\_features) with missing values encoded as np.nan |
  | <span id="unsupervised-unsupervised-tabpfnunsupervisedmodel-impute-underscore--t" /><code className="python-reference-parameter">t</code> | <code className="python-reference-type">float</code> | `1e-09` | float, default=0.000000001 Temperature for sampling from the imputation distribution, lower values are more deterministic |
  | <span id="unsupervised-unsupervised-tabpfnunsupervisedmodel-impute-underscore--n-permutations" /><code className="python-reference-parameter">n\_<wbr />permutations</code> | <code className="python-reference-type">int</code> | `10` | int, default=10 Number of permutations to use for imputation |
  | <span id="unsupervised-unsupervised-tabpfnunsupervisedmodel-impute-underscore--condition-on-all-features" /><code className="python-reference-parameter">condition\_<wbr />on\_<wbr />all\_<wbr />features</code> | <code className="python-reference-type">bool</code> | `True` | bool, default=`True` Whether to condition on all other features (`True`) or only previous features (`False`) |
  | <span id="unsupervised-unsupervised-tabpfnunsupervisedmodel-impute-underscore--dag" /><code className="python-reference-parameter">dag</code> | <code className="python-reference-type">dict\[int, list\[int]] \| None</code> | `None` | dict\[int, list\[int]] \| `None`, default=`None` Optional Directed Acyclic Graph mapping each column index to its list of parent column indices (i.e. the features it depends on). When provided, columns are imputed in topological order and each column is conditioned on exactly its DAG parents. Mutually exclusive with `condition_on_all_features=True`. Every feature must appear as a key; map a feature to an empty list to impute it without conditioning. |
  | <span id="unsupervised-unsupervised-tabpfnunsupervisedmodel-impute-underscore--fast-mode" /><code className="python-reference-parameter">fast\_<wbr />mode</code> | <code className="python-reference-type">bool</code> | `False` | bool, default=`False` Whether to use faster settings for testing |
</div>

**Returns**

<div className="python-reference-table python-reference-returns">
  | Type | Description |
  | - | - |
  | <code className="python-reference-type">torch.Tensor</code> | torch.Tensor: Imputed data with missing values replaced |
</div>

**Raises**

`ValueError`

If `dag` is combined with `condition_on_all_features=True`,
contains a cycle, or does not specify every feature.

***

<div className="python-reference-heading">
  <h2 id="unsupervised-unsupervised-tabpfnunsupervisedmodel-impute-single-permutation">
    `TabPFNUnsupervisedModel.impute_single_permutation_`
  </h2>

  <a className="python-reference-source" href="https://github.com/PriorLabs/tabpfn-extensions/blob/840c15a1848a986b39c85bc17efc61e0e377f983/src/tabpfn_extensions/unsupervised/unsupervised.py#L443" aria-label="View source for TabPFNUnsupervisedModel.impute_single_permutation_"><span aria-hidden="true">\</></span> View source <span aria-hidden="true">↗</span></a>
</div>

Impute missing values (np.nan) in `X` by sampling all cells independently from the trained models.

```python theme={null}
TabPFNUnsupervisedModel.impute_single_permutation_(
    X: torch.Tensor,
    feature_permutation: list[int] | tuple[int, ...],
    t: float = 1e-09,
    condition_on_all_features: bool = True,
) -> tuple[torch.Tensor, dict[str, torch.Tensor]]
```

**Parameters**

<div className="python-reference-table">
  | Parameter | Type | Default | Description |
  | - | - | - | - |
  | <span id="unsupervised-unsupervised-tabpfnunsupervisedmodel-impute-single-permutation--x" /><code className="python-reference-parameter">X</code> | <code className="python-reference-type">torch.Tensor</code> | Required | Input data of the shape (num\_examples, num\_features) with missing values encoded as np.nan |
  | <span id="unsupervised-unsupervised-tabpfnunsupervisedmodel-impute-single-permutation--feature-permutation" /><code className="python-reference-parameter">feature\_<wbr />permutation</code> | <code className="python-reference-type">list\[int] \| tuple\[int, ...]</code> | Required | — |
  | <span id="unsupervised-unsupervised-tabpfnunsupervisedmodel-impute-single-permutation--t" /><code className="python-reference-parameter">t</code> | <code className="python-reference-type">float</code> | `1e-09` | Temperature for sampling from the imputation distribution, lower values are more deterministic |
  | <span id="unsupervised-unsupervised-tabpfnunsupervisedmodel-impute-single-permutation--condition-on-all-features" /><code className="python-reference-parameter">condition\_<wbr />on\_<wbr />all\_<wbr />features</code> | <code className="python-reference-type">bool</code> | `True` | — |
</div>

**Returns**

<div className="python-reference-table python-reference-returns">
  | Type | Description |
  | - | - |
  | <code className="python-reference-type">tuple\[torch.Tensor, dict\[str, torch.Tensor]]</code> | Imputed data, with missing values replaced |
</div>

***

<div className="python-reference-heading">
  <h2 id="unsupervised-unsupervised-tabpfnunsupervisedmodel-sample-from-model-prediction">
    `TabPFNUnsupervisedModel.sample_from_model_prediction_`
  </h2>

  <a className="python-reference-source" href="https://github.com/PriorLabs/tabpfn-extensions/blob/840c15a1848a986b39c85bc17efc61e0e377f983/src/tabpfn_extensions/unsupervised/unsupervised.py#L496" aria-label="View source for TabPFNUnsupervisedModel.sample_from_model_prediction_"><span aria-hidden="true">\</></span> View source <span aria-hidden="true">↗</span></a>
</div>

Sample values from a model's prediction distribution.

```python theme={null}
TabPFNUnsupervisedModel.sample_from_model_prediction_(
    column_idx: int,
    X_fit: torch.Tensor,
    model: Any,
    X_predict: torch.Tensor,
    t: float,
) -> tuple[dict[str, Any] | torch.Tensor, torch.Tensor]
```

**Parameters**

<div className="python-reference-table">
  | Parameter | Type | Default | Description |
  | - | - | - | - |
  | <span id="unsupervised-unsupervised-tabpfnunsupervisedmodel-sample-from-model-prediction--column-idx" /><code className="python-reference-parameter">column\_<wbr />idx</code> | <code className="python-reference-type">int</code> | Required | Index of the column being predicted |
  | <span id="unsupervised-unsupervised-tabpfnunsupervisedmodel-sample-from-model-prediction--x-fit" /><code className="python-reference-parameter">X\_<wbr />fit</code> | <code className="python-reference-type">torch.Tensor</code> | Required | Training data used to determine feature type |
  | <span id="unsupervised-unsupervised-tabpfnunsupervisedmodel-sample-from-model-prediction--model" /><code className="python-reference-parameter">model</code> | <code className="python-reference-type">Any</code> | Required | The trained model (classifier or regressor) |
  | <span id="unsupervised-unsupervised-tabpfnunsupervisedmodel-sample-from-model-prediction--x-predict" /><code className="python-reference-parameter">X\_<wbr />predict</code> | <code className="python-reference-type">torch.Tensor</code> | Required | Input data for prediction |
  | <span id="unsupervised-unsupervised-tabpfnunsupervisedmodel-sample-from-model-prediction--t" /><code className="python-reference-parameter">t</code> | <code className="python-reference-type">float</code> | Required | Temperature parameter for sampling (lower values = more deterministic) |
</div>

**Returns**

<div className="python-reference-table python-reference-returns">
  | Type | Description |
  | - | - |
  | <code className="python-reference-type">tuple\[dict\[str, Any] \| torch.Tensor, torch.Tensor]</code> | tuple containing:<br />    - The raw prediction output (dictionary for regressors, tensor for classifiers)<br />    - The sampled values as a tensor |
</div>

***

<div className="python-reference-heading">
  <h2 id="unsupervised-simple-impute-impute-column">
    `impute_column`
  </h2>

  <a className="python-reference-source" href="https://github.com/PriorLabs/tabpfn-extensions/blob/840c15a1848a986b39c85bc17efc61e0e377f983/src/tabpfn_extensions/unsupervised/simple_impute.py#L46" aria-label="View source for impute_column"><span aria-hidden="true">\</></span> View source <span aria-hidden="true">↗</span></a>
</div>

Impute column `col` of `X` (`np.nan` = missing) in place.

Uses every other column as features (feature NaNs left in — TabPFN handles
them), fits `model` on the rows where `col` is observed, and predicts the
rows where it is missing.

```python theme={null}
impute_column(
    X: np.ndarray,
    col: int,
    model: object,
    *,
    classification: bool = False,
) -> np.ndarray
```

**Parameters**

<div className="python-reference-table">
  | Parameter | Type | Default | Description |
  | - | - | - | - |
  | <span id="unsupervised-simple-impute-impute-column--x" /><code className="python-reference-parameter">X</code> | <code className="python-reference-type"><a href="https://numpy.org/doc/stable/reference/generated/numpy.ndarray.html">np.ndarray</a></code> | Required | Data of shape `(n_samples, n_features)` with missing values as `np.nan`. |
  | <span id="unsupervised-simple-impute-impute-column--col" /><code className="python-reference-parameter">col</code> | <code className="python-reference-type">int</code> | Required | Index of the column to impute. |
  | <span id="unsupervised-simple-impute-impute-column--model" /><code className="python-reference-parameter">model</code> | <code className="python-reference-type">object</code> | Required | A fitted-on-call TabPFN estimator (regressor, or classifier when `classification=True`). |
  | <span id="unsupervised-simple-impute-impute-column--classification" /><code className="python-reference-parameter">classification</code> | <code className="python-reference-type">bool</code> | `False` | If `True`, treat the target as categorical (fit on integer labels and predict class labels). |
</div>

**Returns**

<div className="python-reference-table python-reference-returns">
  | Type | Description |
  | - | - |
  | <code className="python-reference-type"><a href="https://numpy.org/doc/stable/reference/generated/numpy.ndarray.html">np.ndarray</a></code> | [`np.ndarray`](https://numpy.org/doc/stable/reference/generated/numpy.ndarray.html): `X` with column `col` filled (mutated in place and returned). |
</div>

***

<div className="python-reference-heading">
  <h2 id="unsupervised-simple-impute-simple-impute">
    `simple_impute`
  </h2>

  <a className="python-reference-source" href="https://github.com/PriorLabs/tabpfn-extensions/blob/840c15a1848a986b39c85bc17efc61e0e377f983/src/tabpfn_extensions/unsupervised/simple_impute.py#L87" aria-label="View source for simple_impute"><span aria-hidden="true">\</></span> View source <span aria-hidden="true">↗</span></a>
</div>

Impute all missing values (`np.nan`) in `X`, one column at a time.

For each column that contains missing values, fit a TabPFN model on the rows
where that column is observed and predict the missing rows, using all other
columns as features. Numerical columns use a regressor; categorical columns
use a classifier.

```python theme={null}
simple_impute(
    X: np.ndarray,
    tabpfn_clf: object | None = None,
    tabpfn_reg: object | None = None,
    categorical_features: list[int] | None = None,
    use_imputed: bool = False,
) -> np.ndarray
```

**Parameters**

<div className="python-reference-table">
  | Parameter | Type | Default | Description |
  | - | - | - | - |
  | <span id="unsupervised-simple-impute-simple-impute--x" /><code className="python-reference-parameter">X</code> | <code className="python-reference-type"><a href="https://numpy.org/doc/stable/reference/generated/numpy.ndarray.html">np.ndarray</a></code> | Required | Data of shape `(n_samples, n_features)` with missing values encoded as `np.nan`. Converted to a float array (a copy is made; the input is not modified). |
  | <span id="unsupervised-simple-impute-simple-impute--tabpfn-clf" /><code className="python-reference-parameter">tabpfn\_<wbr />clf</code> | <code className="python-reference-type">object \| None</code> | `None` | TabPFN classifier used for categorical columns. Defaults to `TabPFNClassifier()` when needed. |
  | <span id="unsupervised-simple-impute-simple-impute--tabpfn-reg" /><code className="python-reference-parameter">tabpfn\_<wbr />reg</code> | <code className="python-reference-type">object \| None</code> | `None` | TabPFN regressor used for numerical columns. Defaults to `TabPFNRegressor()` when needed. |
  | <span id="unsupervised-simple-impute-simple-impute--categorical-features" /><code className="python-reference-parameter">categorical\_<wbr />features</code> | <code className="python-reference-type">list\[int] \| None</code> | `None` | Indices of categorical columns. If `None`, they are inferred from the data. |
  | <span id="unsupervised-simple-impute-simple-impute--use-imputed" /><code className="python-reference-parameter">use\_<wbr />imputed</code> | <code className="python-reference-type">bool</code> | `False` | How to treat already-imputed columns when imputing later ones.<br /><br />\* `False` (default): every column is imputed using only the   *original* observed data as features, so the result does not depend   on column order.<br />\* `True`: columns are filled left to right in place, so a column   imputed earlier is used as a feature for later columns (chained;   order-dependent). |
</div>

**Returns**

<div className="python-reference-table python-reference-returns">
  | Type | Description |
  | - | - |
  | <code className="python-reference-type"><a href="https://numpy.org/doc/stable/reference/generated/numpy.ndarray.html">np.ndarray</a></code> | [`np.ndarray`](https://numpy.org/doc/stable/reference/generated/numpy.ndarray.html): A new array with all missing values imputed. |
</div>


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.