Skip to main content
Looking for usage documentation? Check out Embeddings.

Installation

TabPFNEmbedding

View source
scikit-learn style transformer that extracts TabPFN embeddings. When n_fold >= 2, fit produces out-of-fold (OOF) embeddings for the training data — the robust variant from “A Closer Look at TabPFN v2: Strength, Limitation, and Extension” (https://arxiv.org/abs/2502.17361) — and then refits a single model on the full training set for use on unseen data. The OOF embeddings are stored on train_embeddings_ and returned by fit_transform. transform(X) ALWAYS uses the final, full-data model — it does NOT return cached OOF embeddings, even when X happens to equal the training set. For OOF embeddings call fit_transform (or read train_embeddings_). Note on output shape: transform returns a 3D array of shape (n_estimators, n_samples, embed_dim). It is not a drop-in input for sklearn.pipeline.Pipeline / ColumnTransformer — those expect 2D output. Pick an ensemble member (embeds[0]) or aggregate across axis=0 before passing to a downstream 2D estimator.
Parameters
Attributes Examples

TabPFNEmbedding.fit

View source

TabPFNEmbedding.fit_transform

View source
Fit and return embeddings for the training data. For n_fold >= 2 these are out-of-fold embeddings. For n_fold == 0 they come from the single full-data model.
Parameters
Returns

TabPFNEmbedding.get_embeddings

View source
DEPRECATED. Use fit_transform (OOF) or transform (unseen).
Parameters
Returns

TabPFNEmbedding.transform

View source
Embed unseen data X using the full-data model. Use this method when you have new, held-out data that was not part of training. It always runs inference through model_ (trained on the full training set) and never returns cached embeddings. If you want embeddings for the training data, prefer fit_transform(X_train, y_train), which yields out-of-fold embeddings for n_fold >= 2 (avoiding label leakage) or reads train_embeddings_ after a fit call.
Parameters
Returns