Local package and API
Use a package release with TabPFN-3.5 support. See the quickstart for installation, authentication, and model selection.
Plus text limits
TabPFN-3.5-Plus text preprocessing supports at most 10 million tokens in total and 2,500 characters per text value. These limits apply in addition to the model’s row and column limits.Pass a mixed table
Keep text in its original columns. You do not need to manually vectorize it before passing it to TabPFN-3.5. For a churn classification task, a customer dataset might contain:
Load your data and split it into training and test sets:
string dtype so they can be detected as text, including on pandas 2. Declaring plan as category keeps it categorical regardless of how many distinct values it contains.
Select TabPFN-3.5 locally or TabPFN-3.5-Plus through the API:
TabPFNRegressor from the same package. The feature table can still contain text, categorical labels, and numerical values.
Preparing text columns
- Keep text features alongside the other relevant columns, so the model can use both the text and the structured data.
- Use the same feature columns when fitting and predicting.
- Include only text available at prediction time. For example, notes written after a customer churns should not be inputs to a prediction of that churn.
- Start with raw text. Evaluate manual vectorization or domain-specific text features against that baseline on held-out data, and fit learned preprocessing only on the training split.