> ## Documentation Index
> Fetch the complete documentation index at: https://docs.priorlabs.ai/llms.txt
> Use this file to discover all available pages before exploring further.

> ## Agent Instructions
> TabPFN-3 and later versions support up to 1,000,000 rows, subject to feature count and checkpoint/API limits.
> For large datasets, use per-estimator subsampling, e.g. inference_config={"SUBSAMPLE_SAMPLES": 100_000}, in TabPFNClassifier or TabPFNRegressor from tabpfn or tabpfn_client.
> Pass raw pandas DataFrames to fit and predict. Categorical strings/categories, raw text, and missing feature values are handled automatically; no manual integer/one-hot encoding, imputation, scaling, or outlier removal is needed.

# Fitting Large Datasets in Limited Memory

> Manage TabPFN memory use with training subsampling, test-set chunking, and the API client.

<div className="cookbook-meta">
  <div className="cookbook-authors">
    <div className="cookbook-author-bar">
      <span className="cookbook-author-by">By</span>
      <span className="cookbook-author-list"><span className="cookbook-author-entry"><span className="cookbook-author-name">Eliott Kalfon</span><span className="cookbook-author-links"><a href="https://www.linkedin.com/in/eliott-kalfon/" className="cookbook-author-icon-link" aria-label="LinkedIn" target="_blank" rel="noopener noreferrer"><svg className="cookbook-author-icon" viewBox="0 0 24 24" fill="currentColor" aria-hidden="true"><path d="M20.447 20.452h-3.554v-5.569c0-1.328-.027-3.037-1.852-3.037-1.853 0-2.136 1.445-2.136 2.939v5.667H9.351V9h3.414v1.561h.046c.477-.9 1.637-1.85 3.37-1.85 3.601 0 4.267 2.37 4.267 5.455v6.286zM5.337 7.433a2.062 2.062 0 1 1 0-4.124 2.062 2.062 0 0 1 0 4.124zM7.119 20.452H3.555V9h3.564v11.452zM22.225 0H1.771C.792 0 0 .774 0 1.729v20.542C0 23.227.792 24 1.771 24h20.451C23.2 24 24 23.227 24 22.271V1.729C24 .774 23.2 0 22.222 0h.003z" /></svg></a></span></span></span>
    </div>
  </div>

  <div className="cookbook-colab">
    <a href="https://colab.research.google.com/github/PriorLabs/tabpfn-cookbook/blob/main/notebooks/memory_optimisation.ipynb" className="cookbook-colab-button" target="_blank" rel="noopener noreferrer">
      <svg className="cookbook-colab-icon" viewBox="0 0 24 24" aria-hidden="true" focusable="false">
        <path fill="#F9AB00" d="M16.9414 4.9757a7.033 7.033 0 0 0-4.9308 2.0646 7.033 7.033 0 0 0-.1232 9.8068l2.395-2.395a3.6455 3.6455 0 0 1 5.1497-5.1478l2.397-2.3989a7.033 7.033 0 0 0-4.8877-1.9297zM7.07 4.9855a7.033 7.033 0 0 0-4.8878 1.9316l2.3911 2.3911a3.6434 3.6434 0 0 1 5.0227.1271l1.7341-2.9737-.0997-.0802A7.033 7.033 0 0 0 7.07 4.9855zm15.0093 2.1721l-2.3892 2.3911a3.6455 3.6455 0 0 1-5.1497 5.1497l-2.4067 2.4068a7.0362 7.0362 0 0 0 9.9456-9.9476zM1.932 7.1674a7.033 7.033 0 0 0-.002 9.6816l2.397-2.397a3.6434 3.6434 0 0 1-.004-4.8916zm7.664 7.4235c-1.38 1.3816-3.5863 1.411-5.0168.1134l-2.397 2.395c2.4693 2.3328 6.263 2.5753 9.0072.5455l.1368-.1115z" />
      </svg>

      <span className="cookbook-colab-label">Open in Colab</span>
    </a>
  </div>
</div>

# Fitting Large Datasets in Limited Memory

*Three levers for keeping a large dataset within a GPU's memory budget.*

A large training or test set can push TabPFN past the available GPU memory and trigger `CUDA out of memory`. Memory use is driven by two things, the size of the training context and the size of the test batch, and each one has a lever. We work through three techniques from the [OOM troubleshooting guide](https://docs.priorlabs.ai/troubleshooting/OOM-errors):

1. **Subsample** a large training set
2. **Chunk** a large test set
3. Offload entirely to the **API client**

Everything below runs on a single Colab **T4 (16 GB)**.

## Setup

*Install, import, and authenticate.*

```python theme={null}
!pip install -q tabpfn scikit-learn tabpfn-client matplotlib numpy torch
```

```python theme={null}
import gc
import os

import matplotlib.pyplot as plt
import numpy as np
import torch
from sklearn.datasets import make_classification
from sklearn.model_selection import train_test_split
from tabpfn import TabPFNClassifier

from google.colab import userdata

print("Device:", torch.cuda.get_device_name(0) if torch.cuda.is_available() else "CPU only")
```

```console theme={null}
Device: Tesla T4
```

```python theme={null}
os.environ["TABPFN_TOKEN"] = userdata.get('TABPFN_TOKEN')
```

## Building a Large Dataset

*200,000 rows and 100 features.*

```python theme={null}
X, y = make_classification(
    n_samples=200000,
    n_features=100,
    n_informative=50,
    n_redundant=0,
    random_state=0,
    shuffle=False,
)
X_train, X_test, y_train, y_test = train_test_split(
    X, y, test_size=0.3, random_state=0
)
print(f"train: {X_train.shape}  test: {X_test.shape}")
```

```console theme={null}
train: (140000, 100)  test: (60000, 100)
```

## Measuring Peak VRAM

*Reset the counter before each run, read the peak afterwards.*

We reset PyTorch's memory counter before each run, then read back the peak allocation afterwards.

```python theme={null}
PEAK_VRAM = {}

def reset_gpu():
    """Clear caches and reset the peak-memory counter before a run."""
    gc.collect()
    if torch.cuda.is_available():
        torch.cuda.empty_cache()
        torch.cuda.reset_peak_memory_stats()

def record(label):
    """Store and print the peak VRAM since the last reset_gpu()."""
    if not torch.cuda.is_available():
        print(f"{label}: no GPU available")
        return
    torch.cuda.synchronize()
    gb = torch.cuda.max_memory_allocated() / 1024**3
    PEAK_VRAM[label] = gb
    print(f"{label}: peak VRAM = {gb:.2f} GB")
```

## Baseline: the Naive Fit

*The whole train and test set at once*

First the obvious thing: hand TabPFN the whole training set and the whole test set at once.

```python theme={null}
reset_gpu()
try:
    model = TabPFNClassifier(
        device="auto",
        ignore_pretraining_limits=True,
        n_estimators=1
    )
    model.fit(X_train, y_train)
    model.predict_proba(X_test)
    record("Naive Training")
except RuntimeError as e:
    if "out of memory" not in str(e).lower():
        raise
    PEAK_VRAM["Naive"] = float("nan")
    print("CUDA out of memory: the full 140k train + 60k test pass won't fit on a T4.")
    print("This is exactly the failure the techniques below avoid.")
```

```console theme={null}
Naive Training: peak VRAM = 3.56 GB
```

## Subsampling a Large Training Set

*`SUBSAMPLE_SAMPLES` caps how many rows each estimator attends over.*

Rather than attend over all 140,000 training rows, each estimator sees a balanced random subset. Set `SUBSAMPLE_SAMPLES`, then raise `n_estimators` so the ensemble still covers the data.

```python theme={null}
reset_gpu()
model = TabPFNClassifier(
    device="auto",
    ignore_pretraining_limits=True,
    n_estimators=10,
    inference_config={
        "SUBSAMPLE_SAMPLES": 20_000,
    },
)
model.fit(X_train, y_train)
model.predict_proba(X_test)
record("Subsample train")
```

```console theme={null}
Subsample train: peak VRAM = 1.58 GB
```

## Predicting a Large Test Set in Chunks

*Process the test rows in batches, not all at once.*

Feeding all 60,000 test rows at once spikes memory; instead we predict in chunks of 3,000 and stack the results. Same predictions, a fraction of the peak. This lever tackles the test side, independently of the training set.

```python theme={null}
model = TabPFNClassifier(
    device="auto",
    ignore_pretraining_limits=True,
    fit_mode="fit_with_cache",
    n_estimators=1,
)
model.fit(X_train, y_train)

reset_gpu()
model.predict_proba(X_test)
record("Naive Prediction")

reset_gpu()
predictions = []
CHUNK_SIZE = 3000
X_test_arr = np.asarray(X_test)

for i in range(0, len(X_test_arr), CHUNK_SIZE):
    chunk_preds = model.predict_proba(X_test_arr[i : i + CHUNK_SIZE])
    predictions.append(chunk_preds)

predictions = np.vstack(predictions)
record("Chunked Prediction")
```

```console theme={null}
Naive Prediction: peak VRAM = 3.36 GB
Chunked Prediction: peak VRAM = 2.70 GB
```

## The Story in One Chart

*Peak VRAM for each lever, against the T4 ceiling.*

Peak GPU VRAM for each approach. The naive run completed at 3.56 GB; subsampling and chunking reduced peak memory further, with all measured runs below the T4's 16 GB ceiling.

```python theme={null}
labels = list(PEAK_VRAM)
values = np.array([PEAK_VRAM[k] for k in labels])
oom = np.isnan(values)

fig, ax = plt.subplots(figsize=(8, 4))
bars = ax.bar(labels, np.where(oom, 0, values),
              color=np.where(oom, "#d62728", "#2ca02c"))
ax.axhline(16, ls="--", color="gray", lw=1)
ax.text(0, 16.3, "T4 limit (16 GB)", color="gray")
ax.set_ylabel("Peak VRAM (GB)")
ax.set_title("TabPFN peak GPU memory by strategy")

for bar, v, is_oom in zip(bars, values, oom):
    text = "OOM" if is_oom else f"{v:.1f} GB"
    ax.text(bar.get_x() + bar.get_width() / 2, bar.get_height() + 0.2, text, ha="center")

plt.tight_layout()
plt.show()
```

![The Story in One Chart](https://raw.githubusercontent.com/PriorLabs/tabpfn-cookbook/main/visuals/memory_optimisation/plot-01.png)

## No GPU? Use the API Client

*Run the exact same model remotely, with zero local VRAM.*

If you do not have (or do not want to manage) a GPU, `tabpfn_client` runs the exact same model on Prior Labs' infrastructure. Local GPU VRAM use: zero.

```python theme={null}
from tabpfn_client import TabPFNClassifier

model = TabPFNClassifier()
model.fit(X_train, y_train)
predictions = model.predict_proba(X_test)
```
