Skip to main content

Installation

Get estimator tags in a consistent format across different sklearn versions. This function provides compatibility between sklearn versions before and after 1.6. It returns either a Tags object (sklearn >= 1.6) or a converted Tags object from the dictionary format (sklearn < 1.6) containing metadata about the estimator’s requirements and capabilities.
Parameters
Returns

ClientTabPFNClassifier

View source

ClientTabPFNClassifier.get_params

View source
Return parameters for this estimator.
Parameters
Returns

ClientTabPFNRegressor

View source

ClientTabPFNRegressor.get_params

View source
Return parameters for this estimator.
Parameters
Returns

FakeTorchDevice

View source
Fake used to represent torch.device used when PyTorch is not installed.
Parameters

TabPFNEstimator

View source

TabPFNEstimator.fit

View source

TabPFNEstimator.predict

View source

get_max_num_classes

View source
Infer the max number of classes a TabPFN estimator can predict in one fit. This is the single source of truth for the TabPFN class-count limit across tabpfn-extensions, so a fix here propagates everywhere (unsupervised classifier/regressor routing, the many-class output-coding wrapper, …). The value is read from the model’s inference config (get_inference_config().MAX_NUMBER_OF_CLASSES), which TabPFN exposes as of v8.0.0 (the minimum version this package depends on).
Parameters
Returns

get_tabpfn_models

View source
Get the TabPFN model classes for the selected backend. USE_TABPFN_LOCAL selects the backend; the function does not silently fall back to the other one:
  1. USE_TABPFN_LOCAL is True -> the standard tabpfn package
  2. USE_TABPFN_LOCAL is False -> the tabpfn-client API backend
If the selected backend is not installed, an ImportError is raised naming that backend, rather than quietly using the other one.
Returns
Raises ImportError If the selected TabPFN backend is not installed

infer_categorical_features

View source
Infer which columns are categorical features (constraint (a) only). This answers a data question — “is this column categorical?” — and is deliberately independent of any model constraint. Whether a categorical column has few enough levels for a TabPFN classifier to predict it (constraint (b)) is a separate concern; derive that limit with get_max_num_classes and apply it at the point of use. A column is treated as categorical if any of these hold:
  1. It is in the caller-provided categorical_features list.
  2. It has a string/object/category dtype (pandas DataFrame).
  3. It contains string values (numpy object array).
  4. It is low-cardinality: at most MAX_UNIQUE_VALUES_FOR_CATEGORICAL unique values, with more than MIN_SAMPLES_PER_CATEGORY samples per unique value on average, to avoid mislabelling columns that only look low-cardinality because the sample is too thin per level.
Parameters
Returns

Where TabPFN itself runs: a CPU stand-in when the client serves the model.
Parameters
Returns

infer_torch_device

View source
Where torch work of this process runs, with or without the local tabpfn. With tabpfn installed this is TabPFN’s own reading of device. Without it, the same rule on torch directly: for "auto", CUDA, else MPS, else the CPU, minus what TABPFN_EXCLUDE_DEVICES names; anything else is parsed as a torch device, the first of several.
Parameters
Returns

Check if an estimator is a TabPFN model.
Parameters
Returns

Cartesian product of a dictionary of lists. This function takes a dictionary where each value is a list, and returns an iterator over dictionaries where each key is mapped to one element from the corresponding list.
Parameters
Returns
Example

Apply softmax function to convert logits to probabilities.
Parameters
Returns

warn_if_no_kv_cache

View source
Warn if a TabPFN model isn’t configured to use the KV cache. The KV cache (improved with TabPFN-3) caches the encoder pass over the training set so that repeated predicts against the same fitted model don’t re-encode the training data each time. Extensions that issue many predicts per fit (e.g. imputation-based SHAP, certain feature-selection or HPO routines) benefit from it — without the cache, the encoder pass over the training set runs on every predict and these extensions can be 10-100x slower than necessary. What has to hold depends on the backend. An endpoint-backed estimator (self-hosted container, SageMaker, Foundry) needs use_kv_cache=True, which is the only condition. A local model needs both:
  1. model was constructed with fit_mode="fit_with_cache" (a constructor argument, must be set BEFORE .fit()).
  2. model.executor_.keep_cache_on_device is True (set AFTER .fit(); usually the default but worth setting explicitly).
This helper warns if either is missing, but does not raise — users may have intentional reasons (e.g. memory).
Parameters
Returns