Skip to main content
Feature engineering is one of the most impactful ways to improve TabPFN’s performance. The goal is to encode domain knowledge that TabPFN cannot learn from raw columns alone. The encoding work that used to be feature engineering, such as vectorizing text, expanding dates, or reducing high-cardinality columns, is handled by the model. What remains is the domain knowledge.

Domain-specific features

Create features that capture known relationships in your data:
  • Ratios: price / area, revenue / headcount
  • Interactions: weight / (height ** 2) (BMI), voltage * current (power)
  • Group aggregations: mean, count, or standard deviation of a numeric column grouped by a categorical, for example average spend per customer segment
  • External signals: holidays, weather, prices, or other context that is not in the table

Group identifiers

When rows belong to customers, policies, sites, or machines, keep the identifier in the table and declare it as categorical. TabPFN uses categorical columns at any cardinality, and the group column lets it separate what belongs to the group from what generalizes to groups unseen at training time.
Do not replace the identifier with the fingerprint feature or hash it into a number. For the strongest results on grouped and time-ordered data, use Thinking mode on the API.

Datetime features

TabPFN expands genuine datetime columns into calendar features itself: year, day of year, seconds since the epoch, and cyclical month, day, and weekday pairs, plus time of day when the column carries one. Keep the column as a datetime dtype and pass it as is.
Manual extraction is still useful for signals the automatic expansion does not produce, and as an optional experiment when a specific calendar component matters more than the default set:
The automatic expansion is controlled by the TRANSFORM_DATES setting, which is on for current checkpoints and off for TabPFN-3 and older, where a raw datetime column is refused. The preprocessing guide covers the setting and the older checkpoints.
The TabPFN API processes dates with its own pipeline. Manual extraction is optional there as well.

Text and string features

Start by passing text columns directly alongside numerical and categorical features. The local tabpfn package expands text into numeric features; the Plus and Thinking models on the API add proprietary text processing for stronger results on text-rich datasets.
  • Free text: Pass descriptions, reviews, or notes as pandas string dtype. Manual vectorization is not required.
  • Identifiers: A string column that is an ID rather than free text, such as a merchant name or SKU, works better declared as categorical. Otherwise, above 30 distinct values it is treated as text.
  • Categorical labels: Keep columns such as product category or subscription plan alongside the text.
  • Optional feature engineering: Compare CountVectorizer, TfidfVectorizer, or domain-specific text features against the raw-text baseline on held-out data. Fit any learned preprocessing on the training split only.
See Text features for a mixed-data example and Text features in preprocessing for how the package decides which columns are text.