> For the complete documentation index, see [llms.txt](https://docs.everesteer.ai/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://docs.everesteer.ai/data/benchmark-models.md).

# Benchmark models

The benchmark prediction files, the designated benchmark, and what they are for.

Benchmark model predictions are downloadable parquet files aligned to each split. They serve two purposes:

1. The **designated benchmark** is the reference series that the AIMC term orthogonalises against.
2. The prediction files let you estimate correlation to the benchmark, orthogonalise your signal, and build ensembles.

## The designated benchmark

The designated benchmark is a single column named `v1_sherpa` in the benchmark prediction files below. It is a float prediction between 0 and 1 for every row in the dataset. Atlas publishes no separate designated-benchmark file.

The AIMC term is computed as your prediction's covariance after orthogonalising against this series. The designated benchmark is the same reference in every lane: the tournament, the partial scores, and the resolved scores. A round without a designated benchmark does not settle.

{% hint style="warning" %}
Copying the designated benchmark as your prediction scores zero on AIMC. The orthogonalisation strips the shared component. To score on AIMC, your prediction must differ from the benchmark in a way that correlates with the target.
{% endhint %}

## Benchmark prediction files

The benchmark predictions are served as parquet files matching each split:

| File                                      | What it contains                         |
| ----------------------------------------- | ---------------------------------------- |
| `eiq_train_benchmark_models.parquet`      | Predictions over the `train` split.      |
| `eiq_validation_benchmark_models.parquet` | Predictions over the `validation` split. |
| `eiq_live_benchmark_models.parquet`       | Predictions for the current round.       |

Each file is indexed by `id` and carries an `exped` column and one column per benchmark model. The `v1_sherpa` column is the designated benchmark; on Atlas it is the only benchmark column.

## How to download

{% tabs %}
{% tab title="Python SDK" %}

```python
# Download the benchmark predictions for a split
val_bench = pd.read_parquet(
    client.download_benchmark(universe="futures", split="validation")
)
benchmark = val_bench["v1_sherpa"]  # the designated benchmark, indexed by id
```

`download_benchmark` resolves the `validation_benchmark_models` / `live_benchmark_models` / `train_benchmark_models` file. Event-scoped keys can only download `train_benchmark_models` while an event is running; `validation_benchmark_models` and `live_benchmark_models` are withheld to prevent cloning the event benchmark's board score.
{% endtab %}

{% tab title="MCP" %}
The `download_benchmark` tool downloads the benchmark predictions for a split, returning a file path. `get_benchmarks` returns the benchmark catalogue with per-slice validation metrics and the list of published models.
{% endtab %}

{% tab title="curl" %}

```bash
curl -L -H "X-API-Key: $EIQ_API_KEY" \
  https://api.everesteer.ai/api/v1/futures/data/validation_benchmark_models \
  -o validation_benchmark_models.parquet
```

{% endtab %}
{% endtabs %}

## How to use them offline

You can develop against the benchmark files without submitting:

* **Compute your correlation to the benchmark.** A high correlation means you are tracking the reference; the AIMC term will return little.
* **Orthogonalise your prediction.** Subtract the benchmark-projected component from your signal to produce a residual that AIMC can score.
* **Ensemble with the benchmark.** Blend your prediction with the benchmark or with other reference models.

The `scoring` module in the SDK provides `aimc20(predictions, ai_model, target)` for offline orthogonalisation experiments.

## The benchmarks view

The benchmark catalogue is available as `GET /api/v1/benchmarks/board` (the MCP tool `get_benchmarks` returns the same data). It lists every published benchmark model with its mean CORR, mean AIMC and Sharpe across the scored rounds. The event leaderboard also has a benchmarks view that shows how the designated benchmark, and any reference models, perform on the current event board.


---

# Agent Instructions
This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com.

## Querying This Documentation
If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question.

Perform an HTTP GET request on the current page URL with the `ask` query parameter, and the optional `goal` query parameter:

```
GET https://docs.everesteer.ai/data/benchmark-models.md?ask=<question>&goal=<endgoal>
```

`ask` is the immediate question: it should be specific, self-contained, and written in natural language.
`goal` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal.

The response will contain a direct answer to the question and relevant excerpts and sources from the documentation.

Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections.
