Looking for usage documentation? Check out Embeddings.
Installation
TabPFNEmbedding
View source n_fold >= 2, fit produces out-of-fold (OOF) embeddings for the
training data — the robust variant from “A Closer Look at TabPFN v2:
Strength, Limitation, and Extension” (https://arxiv.org/abs/2502.17361) —
and then refits a single model on the full training set for use on unseen
data. The OOF embeddings are stored on train_embeddings_ and returned
by fit_transform.
transform(X) ALWAYS uses the final, full-data model — it does NOT
return cached OOF embeddings, even when X happens to equal the
training set. For OOF embeddings call fit_transform (or read
train_embeddings_).
Note on output shape: transform returns a 3D array of shape
(n_estimators, n_samples, embed_dim). It is not a drop-in input for
sklearn.pipeline.Pipeline / ColumnTransformer — those expect 2D
output. Pick an ensemble member (embeds[0]) or aggregate across
axis=0 before passing to a downstream 2D estimator.
Examples
TabPFNEmbedding.fit
View source TabPFNEmbedding.fit_transform
View source n_fold >= 2 these are out-of-fold embeddings. For
n_fold == 0 they come from the single full-data model.
TabPFNEmbedding.get_embeddings
View source fit_transform (OOF) or transform (unseen).
TabPFNEmbedding.transform
View source X using the full-data model.
Use this method when you have new, held-out data that was not part of
training. It always runs inference through model_ (trained on the
full training set) and never returns cached embeddings.
If you want embeddings for the training data, prefer
fit_transform(X_train, y_train), which yields out-of-fold
embeddings for n_fold >= 2 (avoiding label leakage) or reads
train_embeddings_ after a fit call.