> ## Documentation Index
> Fetch the complete documentation index at: https://docs.priorlabs.ai/llms.txt
> Use this file to discover all available pages before exploring further.

> ## Agent Instructions
> TabPFN-3 and later versions support up to 1,000,000 rows, subject to feature count and checkpoint/API limits.
> For large datasets, use per-estimator subsampling, e.g. inference_config={"SUBSAMPLE_SAMPLES": 100_000}, in TabPFNClassifier or TabPFNRegressor from tabpfn or tabpfn_client.
> Pass raw pandas DataFrames to fit and predict. Categorical strings/categories, raw text, and missing feature values are handled automatically; no manual integer/one-hot encoding, imputation, scaling, or outlier removal is needed.

# Metering

> How TabPFN API predictions and Thinking fits consume tokens, share usage budgets, and are charged.

API predictions and [Thinking fits](/capabilities/thinking-mode) consume **tokens** from the same daily and monthly budgets. There is no separate monthly allowance for the number of Thinking fits. Your usage depends on the computation needed for the model, dataset, and settings you choose.

<Note>
  Usage quotas are separate from per-minute and per-hour request limits. See [API rate limits](/api/rate-limits) for the default limits on uploads, fits, and predictions.
</Note>

***

## Token cost

Tokens measure computation. They are usage units, not a count of words or characters in your data.

For **TabPFN-3.5**, token usage depends on the selected model and the work needed for your request:

* **Dataset size:** The training rows, rows to predict, and feature columns all contribute to the workload. Rows contribute more than columns, and column-row interaction also contributes to the end token charge.
* **Number of estimators:** `n_estimators` controls how many model passes contribute to the prediction. More estimators use more tokens and can improve prediction quality.
* **Type of work:** Thinking fits and predictions require additional computation. Successful KV cache reuse lowers the prediction charge.
* **Minimum charge:** Each billable TabPFN-3 or TabPFN-3.5 operation has a minimum charge of **10,000 tokens** to cover request overhead.

**TabPFN v2.x** charges in proportion to the combined training and prediction rows, the number of columns, and the number of estimators. It uses **8 estimators** when not specified, with a minimum charge of **5,000 tokens** per prediction request.

## What consumes tokens

<div className="metering-operations">
  | **Operation** | **How usage is charged** |
  | - | - |
  | Uploads and standard fits | No separate token charge. |
  | Standard predictions | Charged for each prediction request, based on the model, dataset, and estimator count. |
  | Thinking fits | Charged from the same token budgets as predictions. Higher Thinking effort uses more tokens. |
  | Thinking predictions | Charged for the additional prediction compute of the fitted Thinking model. |
  | Predictions using KV cache | A lower charge applies when the cache is reused. If the request falls back to standard inference, the standard prediction charge applies. |
</div>

The minimum charge still applies to cached predictions. A lower token charge does not change the [request rate limits](/api/rate-limits).

## Estimate and monitor usage

With `tabpfn-client`, estimate a prediction's token cost before running it. Given your training and test feature tables, `X_train` and `X_test`:

```python theme={null}
from tabpfn_client import estimate_cost

estimate = estimate_cost(
    X_train,
    X_test,
    model_version="v3.5",
    n_estimators=8,
)
print(estimate.estimated_cost)
```

The estimate sends only dataset dimensions and your settings; it does not upload the data or consume tokens. Use the same model and estimator settings as your planned prediction. The final charge can differ if the fitted model uses different settings or a cached prediction falls back to standard inference.

Check your usage and account limits on the [Usage page](https://platform.priorlabs.ai/account/usage), or get a usage summary with `tabpfn-client`:

```python theme={null}
from tabpfn_client import get_api_usage

print(get_api_usage())
```

***

## Usage pools

Predictions and Thinking fits share a monthly token budget and a daily cap. Every charged operation counts toward both. There is no separate Thinking fit counter.

| Pool | Default limit | Reset schedule |
| - | - | - |
| Daily tokens | 5,000,000 | Midnight UTC |
| Monthly tokens | 20,000,000 | 1st of each month, midnight UTC |

Your account may have custom limits; check the Usage page for the values that apply to you. If a request's estimated cost exceeds either remaining budget, the API returns **HTTP 429** with a message indicating which limit was hit and when it resets. This applies to predictions using existing fitted models as well as new Thinking fits.

A warning email is sent when either daily or monthly usage crosses 70% of the limit.

***

## When tokens are charged

The API reserves the estimated token cost before starting a billable request. This reservation counts toward both budgets while the work is running. Afterward, the charge is adjusted using the actual model settings and execution details, and any unused reservation is returned.

For a Thinking fit on a newly uploaded dataset, the initial reservation may use file size until the actual row and column counts are known. The final charge is adjusted once those details are available.

If the API confirms that computation did not start, the reservation is refunded. A failed or timed-out request can still consume tokens if work started or its execution status is uncertain.

***

## HTTP error codes

| Status | Condition |
| - | - |
| 429 | Insufficient daily or monthly tokens, or a request rate limit exceeded. The response identifies the limit and when it resets. |
| 401 / 403 | Authentication failure — missing, invalid, or expired token. |
| 422 | Validation error — dataset exceeds model limits, invalid parameters, or upload issues. |

***

## Higher limits

Request higher token budgets through the [Usage page](https://platform.priorlabs.ai/account/usage). Include your expected daily and monthly usage and the types of workloads you plan to run.

***

<CardGroup cols={2}>
  <Card title="API rate limits" icon="gauge-high" href="/api/rate-limits">
    Per-minute and per-hour limits for uploads, fits, and predictions.
  </Card>

  <Card title="Thinking mode" icon="brain" href="/capabilities/thinking-mode">
    Configure fit-time optimization and understand thinking parameters.
  </Card>

  <Card title="REST quickstart" icon="rocket" href="/api/rest-quickstart">
    Full upload → fit → predict walkthrough.
  </Card>

  <Card title="TabPFN-3 changelog" icon="star" href="/changelog/tabpfn-3">
    What's new in v3 — scale, capabilities, and migration.
  </Card>

  <Card title="Security" icon="shield" href="/api/security">
    Encryption, data isolation, and access controls.
  </Card>
</CardGroup>


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.