Looking for usage documentation? Check out Interpretability.
feature_selection.feature_selection
View source estimator by
repeatedly fitting estimator on candidate subsets and keeping the
one that maximizes a cross-validated score. Forwards the relevant
SequentialFeatureSelector hyperparameters; always computes the
baseline (all-features) and selected (subset-only) CV scores so they
are available on the returned object regardless of verbose.
Note that while we expose feature selection here, TabPFN is very
robust to noisy / uninformative features in its native
in-context-learning regime, so the accuracy gain from running this
selector is often marginal — on real datasets the all-features
baseline and the selected subset typically score within
cross-validation noise of each other. The value of running selection
on TabPFN is usually more in terms of interpretability.
Note that other interpretability methods, such as SHAP, are also supported
and are generally much faster because they can use the KV cache.
Sequential feature selection is expensive: forward selection with
n_features_to_select=k on d features uses on the order of
cv * sum_{i=0..k-1} (d - i) model fits, plus 2 more for the
baseline / selected CV scores. n_jobs parallelizes
candidate-feature evaluation within each round; pass -1 for all
cores. Note that v3’s KV cache does not help here — every candidate
has a different X_train so the cache invalidates between fits.
feature_selection.FeatureSelectionResult
View source feature_selection.
Attributes