Domain-specific features
Create features that capture known relationships in your data:- Ratios:
price / area,revenue / headcount - Interactions:
weight / (height ** 2)(BMI),voltage * current(power) - Group aggregations: mean, count, or standard deviation of a numeric column grouped by a categorical, for example average spend per customer segment
- External signals: holidays, weather, prices, or other context that is not in the table
Group identifiers
When rows belong to customers, policies, sites, or machines, keep the identifier in the table and declare it as categorical. TabPFN uses categorical columns at any cardinality, and the group column lets it separate what belongs to the group from what generalizes to groups unseen at training time.Datetime features
TabPFN expands genuine datetime columns into calendar features itself: year, day of year, seconds since the epoch, and cyclical month, day, and weekday pairs, plus time of day when the column carries one. Keep the column as a datetime dtype and pass it as is.TRANSFORM_DATES setting, which is on for current checkpoints and off for TabPFN-3 and older, where a raw datetime column is refused. The preprocessing guide covers the setting and the older checkpoints.
The TabPFN API processes dates with its own pipeline. Manual extraction is optional there as well.
Text and string features
Start by passing text columns directly alongside numerical and categorical features. The localtabpfn package expands text into numeric features; the Plus and Thinking models on the API add proprietary text processing for stronger results on text-rich datasets.
- Free text: Pass descriptions, reviews, or notes as pandas
stringdtype. Manual vectorization is not required. - Identifiers: A string column that is an ID rather than free text, such as a merchant name or SKU, works better declared as categorical. Otherwise, above 30 distinct values it is treated as text.
- Categorical labels: Keep columns such as product category or subscription plan alongside the text.
- Optional feature engineering: Compare
CountVectorizer,TfidfVectorizer, or domain-specific text features against the raw-text baseline on held-out data. Fit any learned preprocessing on the training split only.