Looking for usage documentation? Check out Regression.
TabPFNRegressor
</> View source ↗create_default_for_version()
instead. You can also use model_path to specify a particular model.
TabPFNRegressor(
*,
n_estimators: int | Literal["auto"] = "auto",
auto_scale_n_estimators: bool = True,
categorical_features_indices: Sequence[int] | None = None,
softmax_temperature: float | Literal["auto"] = "auto",
average_before_softmax: bool = False,
model_path: str | Path | list[str] | list[Path] | Literal["auto"] | RegressorModelSpecs | list[RegressorModelSpecs] = "auto",
device: DevicesSpecification = "auto",
ignore_pretraining_limits: bool = False,
inference_precision: _dtype | Literal["autocast", "auto"] = "auto",
fit_mode: Literal["low_memory", "fit_preprocessors", "fit_with_cache", "batched"] = "fit_preprocessors",
memory_saving_mode: MemorySavingMode = "auto",
keep_cache_on_device: bool = True,
kv_cache_precision: Literal["auto", "int8", "fp8"] | None = None,
random_state: int | np.random.RandomState | np.random.Generator | None = 0,
n_jobs: Annotated[int | None, deprecated("Use n_preprocessing_jobs")] = None,
n_preprocessing_jobs: int = 1,
inference_config: dict | InferenceConfig | None = None,
differentiable_input: bool = False,
eval_metric: str | RegressorEvalMetrics | None = None,
tuning_config: dict | RegressorTuningConfig | None = None,
show_progress_bar: bool = False,
)
Type aliases
Type aliases
DevicesSpecification = torch.device | str | Sequence[torch.device | str] | Literal["auto"]
MemorySavingMode = bool | Literal["auto"] | float | int
| Parameter | Type | Default | Description |
|---|---|---|---|
n_ | int | Literal[“auto”] | "auto" | The number of estimators in the TabPFN ensemble. We aggregate the predictions of n_estimators-many forward passes of TabPFN. Each forward pass has (slightly) different input data. Think of this as an ensemble of n_estimators-many “prompts” of the input data. With the default "auto", the count comes from the checkpoint (InferenceConfig.N_ESTIMATORS), which is itself "auto" unless the checkpoint names a count. "auto" means DEFAULT_N_ESTIMATORS, raised on wide datasets so every feature is seen by some estimator (i.e. when the data has more than max_features_per_estimator features per estimator), to the smallest value that lets every feature appear in at least one ensemble member, emitting a warning when it does so. That auto-scaled value is capped at MAX_AUTO_SCALED_N_ESTIMATORS; beyond that some features may never be sampled unless you raise n_estimators yourself. An explicit integer — yours or the checkpoint’s — is never overridden: if it is too small to cover every feature, a warning is emitted at fit time and the value is used as given. Your integer cannot be combined with an N_ESTIMATORS in inference_config, which is the other way of naming a count. |
auto_ | bool | True | Deprecated, removed in v9 — pass an explicit n_estimators instead. Only applies when n_estimators="auto", where False keeps the auto value at DEFAULT_N_ESTIMATORS rather than raising it for feature coverage, exactly what passing that count as n_estimators does. Passing False emits a FutureWarning at fit time. |
categorical_ | Sequence[int] | None | None | The indices of the columns that are suggested to be treated as categorical. If None, the model will infer the categorical columns. A column with pandas’ category dtype counts as listed here. A string column declared this way is read as categorical whatever its cardinality, never as text; for a numeric one, we might ignore the suggestion to better fit the data seen during pre-training.!!! note The indices are 0-based and should represent the data passed to .fit(). If the data changes between the initializations of the model and the .fit(), consider setting the .categorical_features_indices attribute after the model was initialized and before .fit(). |
softmax_ | float | Literal[“auto”] | "auto" | The temperature for the softmax function. This is used to control the confidence of the model’s predictions. Lower values make the model’s predictions more confident. This is only applied when predicting during a post-processing step. Set softmax_temperature=1.0 for no effect.If "auto" (the default), the temperature is taken from the checkpoint (InferenceConfig.SOFTMAX_TEMPERATURE), which is 0.9 for every checkpoint released up to and including v8.5.0. Passing a float overrides the checkpoint for every model in the ensemble; it cannot be combined with a SOFTMAX_TEMPERATURE in inference_config, which is the other way of naming one. |
average_ | bool | False | Only used if n_estimators > 1. Whether to average the predictions of the estimators before applying the softmax function. This can help to improve predictive performance when there are many classes or when calibrating the model’s confidence. This is only applied when predicting during a post-processing.- If True, the predictions are averaged before applying the softmax function. Thus, we average the logits of TabPFN and then apply the softmax.- If False, the softmax function is applied to each set of logits. Then, we average the resulting probabilities of each forward pass. |
model_ | str | Path | list[str] | list[Path] | Literal[“auto”] | Regressor | "auto" | The path to the TabPFN model file, i.e., the pre-trained weights. - If "auto", the model will be downloaded upon first use. This defaults to your system cache directory, but can be overwritten with the use of an environment variable TABPFN_MODEL_CACHE_DIR.- If a path or a string of a path, the model will be loaded from the user-specified location if available, otherwise it will be downloaded to this location. Details on available checkpoints are available in the repository README. |
device | Devices | "auto" | The device(s) to use for inference. See the documentation of .to(). |
ignore_ | bool | False | Whether to ignore the pre-training limits of the model. The TabPFN models have been pre-trained on a specific range of input data. If the input data is outside of this range, the model may not perform well. You may ignore our limits to use the model on data outside the pre-training range. - If True, the model will not raise an error if the input data is outside the pre-training range. Also suppresses error when using the model with a large dataset on CPU.- If False, you can use the model outside the pre-training range, but the model could perform worse.!!! note For version 2.5, the pre-training limits are: - 50_000 samples/rows - 2_000 features/columns (Note that for more than 500 features we subsample 500 features per estimator. It is therefore important to use a sufficiently large number of n_estimators.) |
inference_ | _dtype | Literal[“autocast”, “auto”] | "auto" | The precision to use for inference. This can dramatically affect the speed and reproducibility of the inference. Higher precision can lead to better reproducibility but at the cost of speed. By default, we optimize for speed and use torch’s mixed-precision autocast. The options are: - If torch.dtype, we force precision of the model and data to be the specified torch.dtype during inference. This can is particularly useful for reproducibility. Here, we do not use mixed-precision.- If "autocast", enable PyTorch’s mixed-precision autocast. Ensure that your device is compatible with mixed-precision.- If "auto", we determine whether to use autocast or not depending on the device type. |
fit_ | Literal[“low_memory”, “fit_preprocessors”, “fit_with_cache”, “batched”] | "fit_preprocessors" | Determine how the TabPFN model is “fitted”. The mode determines how the data is preprocessed and cached for inference. This is unique to an in-context learning foundation model like TabPFN, as the “fitting” is technically the forward pass of the model. The options are: - If "low_memory", the data is preprocessed on-demand during inference when calling .predict() or .predict_proba(). This is the most memory-efficient mode but can be slower for large datasets because the data is (repeatedly) preprocessed on-the-fly. Ideal with low GPU memory and/or a single call to .fit() and .predict().- If "fit_preprocessors", the data is preprocessed and cached once during the .fit() call. During inference, the cached preprocessing (of the training data) is used instead of re-computing it. Ideal with low GPU memory and multiple calls to .predict() with the same training data.- If "fit_with_cache", the data is preprocessed and cached once during the .fit() call like in fit_preprocessors. Moreover, the transformer key-value cache is also initialized, allowing for much faster inference on the same data at a large cost of memory. Ideal with very high GPU memory and multiple calls to .predict() with the same training data.- If "batched", the already pre-processed data is iterated over in batches. This can only be done after the data has been preprocessed with the get_preprocessed_datasets function. This is primarily used only for inference with the InferenceEngineBatchedNoPreprocessing class in Fine-Tuning. The fit_from_preprocessed() function sets this attribute internally. |
memory_ | Memory | "auto" | Enable GPU/CPU memory saving mode. This can both avoid out-of-memory errors and improve fit+predict speed by reducing memory pressure. It saves memory by automatically batching certain model computations within TabPFN. - If “auto”: memory saving mode is enabled/disabled automatically based on a heuristic - If True/False: memory saving mode is forced enabled/disabled.If speed is important to your application, you may wish to manually tune this option by comparing the time taken for fit+predict with it set to False and True.!!! warning This does not batch the original input data. We still recommend to batch the test set as necessary if you run out of memory. |
keep_ | bool | True | Only relevant when fit_mode="fit_with_cache". If True (default), the key-value cache is kept on the inference device (e.g. GPU). Uses more device memory but gives lower latency. If False, the cache is stored on CPU. |
kv_ | Literal[“auto”, “int8”, “fp8”] | None | None | Only relevant when fit_mode="fit_with_cache". Resolved against what the model architecture supports. None (default) picks the architecture default ("int8" when it can quantize, e.g. TabPFN-3, else "auto"); "int8" quantizes the key-value cache to save memory; "fp8" stores it as 8-bit floats instead (same size, float rounding semantics; not supported on MPS); "auto" keeps the computed dtype. Requesting a quantized precision on an architecture that cannot quantize warns and falls back to "auto". |
random_ | int | np.random.Random | 0 | Controls the randomness of the model. Pass an int for reproducible results and see the scikit-learn glossary for more information. If None, the randomness is determined by the system when calling .fit().!!! warning We depart from the usual scikit-learn behavior in that by default we provide a fixed seed of 0.!!! note Even if a seed is passed, we cannot always guarantee reproducibility due to PyTorch’s non-deterministic operations and general numerical instability. To get the most reproducible results across hardware, we recommend using a higher precision as well (at the cost of a much higher inference time). Likewise, for scikit-learn, consider passing USE_SKLEARN_16_DECIMAL_PRECISION=True as kwarg. |
n_ | int | None | None | Deprecated, use n_preprocessing_jobs instead. This parameter never had any effect. |
n_ | int | 1 | The number of worker processes to use for the preprocessing. If 1, the preprocessing will be performed in the current process, parallelised across multiple CPU cores. If >1 and n_estimators > 1, then different estimators will be dispatched to different processes.We strongly recommend setting this to 1, which has the lowest overhead and can often fully utilise the CPU. Values >1 can help if you have lots of CPU cores available, but can also be slower. |
inference_ | dict | InferenceInferenceConfig options | None | For advanced users, additional advanced arguments that adjust the behavior of the model interface. See tabpfn.inference_config.InferenceConfig for details and options.- If None, the default InferenceConfig is used.- If dict, the key-value pairs are used to update the default InferenceConfig. Raises an error if an unknown key is passed.- If InferenceConfig, the object replaces the checkpoint’s config as a whole, so any field not set on it takes a class default rather than the value the checkpoint declares. Deprecated. |
differentiable_ | bool | False | If true, preprocessing attempts to be end-to-end differentiable. Less relevant for standard regression fine-tuning compared to prompt-tuning. |
eval_ | str | Regressor | None | Metric by which predictions will be evaluated on test data for temperature calibration. For currently supported metrics, see tabpfn.inference_tuning.RegressorEvalMetrics. |
tuning_ | dict | Regressor | None | The settings to use to tune the model’s predictions for the specified eval_metric. See tabpfn.inference_tuning.RegressorTuningConfig for details and options. |
show_ | bool | False | Whether to show a progress bar during inference. Defaults to False. |
- TabPFN-3 and later versions support up to 1,000,000 rows, subject to feature count, checkpoint limits, and memory.
- For large datasets or limited memory, use per-estimator subsampling,
e.g.
inference_config={"SUBSAMPLE_SAMPLES": 50_000}. - Pass raw pandas DataFrames to
fitandpredict. Categorical strings/categories and missing feature values are handled automatically; no manual integer/one-hot encoding, imputation, scaling, or outlier removal is needed.
| Attribute | Type | Description |
|---|---|---|
configs_ | list[Architecture | The configurations of the loaded models to be used for inference. The concrete type of these configs is defined by the architectures in use and should be inspected at runtime, but they will be subclasses of ArchitectureConfig. |
models_ | list[Architecture] | The loaded models to be used for inference. The models can be different PyTorch modules, but will be subclasses of Architecture. |
inference_config_ | Inference | Additional configuration of inference for expert users. |
devices_ | tuple[torch.device, …] | The devices determined to be used. The devices are determined based on the device argument to the constructor, and the devices available on the system. See the constructor documentation for details. |
feature_names_in_ | npt.ND | The feature names of the input data. May not be set if the input data does not have feature names, such as with a numpy array. |
n_features_in_ | int | The number of features in the input data used during fit(). |
n_train_samples_ | int | The number of training samples used during fit(). |
inferred_feature_schema_ | Feature | The inferred feature schema. This contains the feature modalities per column, using heuristics and user-provided indices for categorical features. |
n_outputs_ | Literal[1] | The number of outputs the model supports. Only 1 for now |
znorm_space_bardist_ | Full | The bar distribution of the target variable, used by the model. This is the bar distribution in the normalized target space. |
raw_space_bardist_ | Full | The bar distribution in the raw target space, used for computing the predictions. |
use_autocast_ | bool | Whether torch’s autocast should be used. |
forced_inference_dtype_ | _dtype | None | The forced inference dtype for the model based on inference_precision. |
executor_ | Inference | The inference engine used to make predictions. |
ordinal_encoder_ | Order | The column transformer used to preprocess categorical data to be numeric. |
date_transformer_ | Date | The transformer that converted every temporal column before validation. |
text_transformer_ | Text | The transformer that expanded every text column before validation. |
categorical_features_indices_ | list[int] | None | Declared categorical column positions after date/text expansion, including columns declared through pandas category dtype. Expanded source columns are removed and their generated features appended, so these positions can differ from those in the original fit input. |
eval_metric_ | Regressor | The validated evaluation metric to optimize for during prediction. |
ensemble_softmax_temperature_ | float | The temperature applied to the aggregated ensemble distribution at predict time, after the per-estimator softmax_temperature. This is 1.0, a no-op, when no temperature calibration is done. |
softmax_temperature_ | float | The resolved per-estimator softmax_temperature, i.e. the one the checkpoint declares unless it was overridden. |
TabPFNRegressor.create_default_for_version
</> View source ↗TabPFNRegressor.create_default_for_version(
version: ModelVersion,
**overrides,
) -> Self
| Parameter | Type | Default | Description |
|---|---|---|---|
version | Model | Required | — |
**overrides | —TabPFNRegressor options | — | — |
| Type | Description |
|---|---|
Self | — |
TabPFNRegressor.estimator_type
</> View source ↗| Type | Description |
|---|---|
Literal[“regressor”] | — |
TabPFNRegressor.model_
</> View source ↗| Type | Description |
|---|---|
Architecture | — |
TabPFNRegressor.norm_bardist_
</> View source ↗raw_space_bardist_ instead.
This attribute will be removed in a future version.
Returns
| Type | Description |
|---|---|
Full | — |
TabPFNRegressor.bardist_
</> View source ↗znorm_space_bardist_ instead.
This attribute will be removed in a future version.
Returns
| Type | Description |
|---|---|
Full | — |
TabPFNRegressor.get_inference_config
</> View source ↗fit(). Any inference_config override
passed to the constructor is considered.
TabPFNRegressor.get_inference_config() -> InferenceConfig
| Type | Description |
|---|---|
Inference | A deep copy of the active inference config. |
Scikit-learn methods
Seeget_params and set_params.