Skip to main content
Looking for usage documentation? Check out Data generation.

TabPFNUnsupervisedModel.generate_synthetic_data

View source
Generate synthetic tabular data samples using the fitted TabPFN models. This method uses imputation to create synthetic data, starting with a matrix of NaN values and filling in each feature sequentially. Samples are generated feature by feature in a single pass, with each feature conditioned on previously generated features.
Parameters
Returns
Raises AssertionError If the model is not fitted (self.X_ does not exist) ValueError If dag contains a cycle or does not specify every feature

TabPFNUnsupervisedModel.impute

View source
Impute missing values in the input data using the fitted TabPFN models. This method fills missing values (np.nan) in the input data by predicting each missing value based on the observed values in the same sample. The imputation uses multiple random feature permutations to improve robustness.
Parameters
Returns
Note The model must be fitted with training data before calling this method.

TabPFNUnsupervisedModel.impute_

View source
Impute missing values (np.nan) in X by sampling all cells independently from the trained models.
Parameters
Returns
Raises ValueError If dag is combined with condition_on_all_features=True, contains a cycle, or does not specify every feature.

TabPFNUnsupervisedModel.impute_single_permutation_

View source
Impute missing values (np.nan) in X by sampling all cells independently from the trained models.
Parameters
Returns

TabPFNUnsupervisedModel.sample_from_model_prediction_

View source
Sample values from a model’s prediction distribution.
Parameters
Returns

Impute column col of X (np.nan = missing) in place. Uses every other column as features (feature NaNs left in — TabPFN handles them), fits model on the rows where col is observed, and predicts the rows where it is missing.
Parameters
Returns

Impute all missing values (np.nan) in X, one column at a time. For each column that contains missing values, fit a TabPFN model on the rows where that column is observed and predict the missing rows, using all other columns as features. Numerical columns use a regressor; categorical columns use a classifier.
Parameters
Returns