> ## Documentation Index
> Fetch the complete documentation index at: https://docs.priorlabs.ai/llms.txt
> Use this file to discover all available pages before exploring further.

> ## Agent Instructions
> TabPFN-3 and later versions support up to 1,000,000 rows, subject to feature count and checkpoint/API limits.
> For large datasets, use per-estimator subsampling, e.g. inference_config={"SUBSAMPLE_SAMPLES": 100_000}, in TabPFNClassifier or TabPFNRegressor from tabpfn or tabpfn_client.
> Pass raw pandas DataFrames to fit and predict. Categorical strings/categories, raw text, and missing feature values are handled automatically; no manual integer/one-hot encoding, imputation, scaling, or outlier removal is needed.

# Feature selection

> Select features with cross-validation and inspect the selection results.

<Info>
  Looking for usage documentation? Check out [Interpretability](/capabilities/interpretability).
</Info>

<div className="python-reference-heading">
  <h2 id="feature-selection">
    `feature_selection.feature_selection`
  </h2>

  <a className="python-reference-source" href="https://github.com/PriorLabs/tabpfn-extensions/blob/840c15a1848a986b39c85bc17efc61e0e377f983/src/tabpfn_extensions/interpretability/feature_selection.py#L104" aria-label="View source for feature_selection.feature_selection"><span aria-hidden="true">\</></span> View source <span aria-hidden="true">↗</span></a>
</div>

Sequential feature selection wrapper around scikit-learn's SFS.

Picks a subset of features that work well for `estimator` by
repeatedly fitting `estimator` on candidate subsets and keeping the
one that maximizes a cross-validated score. Forwards the relevant
[`SequentialFeatureSelector`](https://scikit-learn.org/stable/modules/generated/sklearn.feature_selection.SequentialFeatureSelector.html) hyperparameters; always computes the
baseline (all-features) and selected (subset-only) CV scores so they
are available on the returned object regardless of `verbose`.

Note that while we expose feature selection here, **TabPFN is very
robust to noisy / uninformative features** in its native
in-context-learning regime, so the *accuracy gain* from running this
selector is often marginal — on real datasets the all-features
baseline and the selected subset typically score within
cross-validation noise of each other. The value of running selection
on TabPFN is usually more in terms of interpretability.
Note that other interpretability methods, such as SHAP, are also supported
and are generally much faster because they can use the KV cache.

Sequential feature selection is expensive: forward selection with
`n_features_to_select=k` on `d` features uses on the order of
`cv * sum_{i=0..k-1} (d - i)` model fits, plus 2 more for the
baseline / selected CV scores. `n_jobs` parallelizes
candidate-feature evaluation within each round; pass `-1` for all
cores. Note that v3's KV cache does *not* help here — every candidate
has a different `X_train` so the cache invalidates between fits.

```python theme={null}
feature_selection.feature_selection(
    estimator: BaseEstimator,
    X: np.ndarray,
    y: np.ndarray,
    n_features_to_select: int | float | str,
    feature_names: list[str] | None = None,
    *,
    cv: int | BaseCrossValidator | Iterable = 5,
    scoring: str | Callable | None = None,
    direction: str = "forward",
    n_jobs: int | None = None,
    tol: float | None = None,
    verbose: bool = True,
    **kwargs: Any,
) -> FeatureSelectionResult
```

**Parameters**

<div className="python-reference-table">
  | Parameter | Type | Default | Description |
  | - | - | - | - |
  | <span id="feature-selection-estimator" /><code className="python-reference-parameter">estimator</code> | <code className="python-reference-type">Base<wbr />Estimator</code> | Required | The model to use for feature selection. |
  | <span id="feature-selection-x" /><code className="python-reference-parameter">X</code> | <code className="python-reference-type"><a href="https://numpy.org/doc/stable/reference/generated/numpy.ndarray.html">np.ndarray</a></code> | Required | Input features, shape `(n_samples, n_features)`. |
  | <span id="feature-selection-y" /><code className="python-reference-parameter">y</code> | <code className="python-reference-type"><a href="https://numpy.org/doc/stable/reference/generated/numpy.ndarray.html">np.ndarray</a></code> | Required | Target values, shape `(n_samples,)`. |
  | <span id="feature-selection-n-features-to-select" /><code className="python-reference-parameter">n\_<wbr />features\_<wbr />to\_<wbr />select</code> | <code className="python-reference-type">int \| float \| str</code> | Required | Number of features to keep. `int` for an absolute count, `float` for a fraction of the total, or `"auto"` to let `tol` decide (requires `tol`). |
  | <span id="feature-selection-feature-names" /><code className="python-reference-parameter">feature\_<wbr />names</code> | <code className="python-reference-type">list\[str] \| None</code> | `None` | Optional list of feature names. When provided, the returned [`FeatureSelectionResult`](/api-reference/python/tabpfn-extensions/interpretability/feature-selection#featureselectionresult) carries the selected names under `selected_names`. |
  | <span id="feature-selection-cv" /><code className="python-reference-parameter">cv</code> | <code className="python-reference-type">int \| Base<wbr />Cross<wbr />Validator \| Iterable</code> | `5` | Cross-validation folds — `int` (k-fold), CV generator, or iterable of splits. Default `5`. |
  | <span id="feature-selection-scoring" /><code className="python-reference-parameter">scoring</code> | <code className="python-reference-type">str \| Callable \| None</code> | `None` | Metric to maximize. `str` (e.g. `"roc_auc"`, `"neg_log_loss"`, `"r2"`) or a callable. Default `None` uses sklearn's per-estimator default (`accuracy` for classifiers, `r2` for regressors). |
  | <span id="feature-selection-direction" /><code className="python-reference-parameter">direction</code> | <code className="python-reference-type">str</code> | `"forward"` | `"forward"` (start empty, add features) or `"backward"` (start full, remove features). Backward is much more expensive but sometimes preferred when features are redundant. |
  | <span id="feature-selection-n-jobs" /><code className="python-reference-parameter">n\_<wbr />jobs</code> | <code className="python-reference-type">int \| None</code> | `None` | Parallelism over candidate features in each round. `-1` uses all cores. Default `None` is single-threaded. |
  | <span id="feature-selection-tol" /><code className="python-reference-parameter">tol</code> | <code className="python-reference-type">float \| None</code> | `None` | Stop condition for auto selection — only used when `n_features_to_select="auto"`. Forward selection stops when adding a feature improves the CV score by less than `tol`. |
  | <span id="feature-selection-verbose" /><code className="python-reference-parameter">verbose</code> | <code className="python-reference-type">bool</code> | `True` | When `True` (default), print the pre- and post-selection CV scores and the names of the selected features. The scores are computed and returned regardless. |
  | <span id="feature-selection-kwargs" /><code className="python-reference-parameter">\*\*kwargs</code> | <code className="python-reference-type">Any</code> | — | Forwarded to [`SequentialFeatureSelector`](https://scikit-learn.org/stable/modules/generated/sklearn.feature_selection.SequentialFeatureSelector.html) for forward compatibility with future sklearn options. |
</div>

**Returns**

<div className="python-reference-table python-reference-returns">
  | Type | Description |
  | - | - |
  | <code className="python-reference-type"><a href="/api-reference/python/tabpfn-extensions/interpretability/feature-selection#featureselectionresult">Feature<wbr />Selection<wbr />Result</a></code> | Selection results containing the fitted selector, support mask, selected feature indices and names, and baseline and selected-feature cross-validation scores. |
</div>

***

<div className="python-reference-heading">
  <h2 id="featureselectionresult">
    `feature_selection.FeatureSelectionResult`
  </h2>

  <a className="python-reference-source" href="https://github.com/PriorLabs/tabpfn-extensions/blob/840c15a1848a986b39c85bc17efc61e0e377f983/src/tabpfn_extensions/interpretability/feature_selection.py#L68" aria-label="View source for feature_selection.FeatureSelectionResult"><span aria-hidden="true">\</></span> View source <span aria-hidden="true">↗</span></a>
</div>

Result of running [`feature_selection`](/api-reference/python/tabpfn-extensions/interpretability/feature-selection#feature-selection).

**Attributes**

| Attribute | Type | Description |
| - | - | - |
| `selector` | <code className="python-reference-type"><a href="https://scikit-learn.org/stable/modules/generated/sklearn.feature_selection.SequentialFeatureSelector.html">Sequential<wbr />Feature<wbr />Selector</a></code> | The underlying fitted [`SequentialFeatureSelector`](https://scikit-learn.org/stable/modules/generated/sklearn.feature_selection.SequentialFeatureSelector.html). Use it for `.transform(X)` to project to the selected columns, or for any sklearn-style downstream work. |
| `support_mask` | <code className="python-reference-type"><a href="https://numpy.org/doc/stable/reference/generated/numpy.ndarray.html">np.ndarray</a></code> | Boolean array of shape `(n_features,)` — `True` for the columns SFS picked. |
| `selected_indices` | <code className="python-reference-type">list\[int]</code> | Integer indices of the selected columns, in ascending order. |
| `selected_names` | <code className="python-reference-type">list\[str] \| None</code> | Selected feature names, in the same order as `selected_indices`. `None` iff `feature_names` wasn't passed. |
| `baseline_score_mean` | <code className="python-reference-type">float</code> | Mean cross-validated score of `estimator` on **all** features, using the same `cv` and `scoring` as the selection step. |
| `baseline_score_std` | <code className="python-reference-type">float</code> | Standard deviation across CV folds for the baseline score. |
| `selected_score_mean` | <code className="python-reference-type">float</code> | Mean cross-validated score of `estimator` on the **selected** subset of features. |
| `selected_score_std` | <code className="python-reference-type">float</code> | Standard deviation across CV folds for the selected-subset score. |


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.