Skip to main content
Looking for usage documentation? Check out Interpretability.

feature_selection.feature_selection

View source
Sequential feature selection wrapper around scikit-learn’s SFS. Picks a subset of features that work well for estimator by repeatedly fitting estimator on candidate subsets and keeping the one that maximizes a cross-validated score. Forwards the relevant SequentialFeatureSelector hyperparameters; always computes the baseline (all-features) and selected (subset-only) CV scores so they are available on the returned object regardless of verbose. Note that while we expose feature selection here, TabPFN is very robust to noisy / uninformative features in its native in-context-learning regime, so the accuracy gain from running this selector is often marginal — on real datasets the all-features baseline and the selected subset typically score within cross-validation noise of each other. The value of running selection on TabPFN is usually more in terms of interpretability. Note that other interpretability methods, such as SHAP, are also supported and are generally much faster because they can use the KV cache. Sequential feature selection is expensive: forward selection with n_features_to_select=k on d features uses on the order of cv * sum_{i=0..k-1} (d - i) model fits, plus 2 more for the baseline / selected CV scores. n_jobs parallelizes candidate-feature evaluation within each round; pass -1 for all cores. Note that v3’s KV cache does not help here — every candidate has a different X_train so the cache invalidates between fits.
Parameters
Returns

feature_selection.FeatureSelectionResult

View source
Result of running feature_selection. Attributes