> For the complete documentation index, see [llms.txt](https://docs.everesteer.ai/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://docs.everesteer.ai/data/datasets.md).

# Datasets

What you can download, which split tokens exist, and how fresh the files are.

Atlas is served as one obfuscated feature matrix with rank-based targets, split three ways, plus benchmark predictions, example predictions and two JSON metadata files. Every file is versioned with the dataset.

## How to download

{% tabs %}
{% tab title="Python SDK" %}

```python
import os
from everestapi import EverestAPI

client = EverestAPI(api_key=os.environ["EIQ_API_KEY"])
train_path      = client.download_dataset(universe="futures", split="train")
validation_path = client.download_dataset(universe="futures", split="validation")
live_path       = client.download_dataset(universe="futures", split="live")

import pandas as pd
train = pd.read_parquet(train_path)
```

`download_dataset` returns a local file path, not a DataFrame. Read it with `pandas.read_parquet`.
{% endtab %}

{% tab title="MCP" %}
The `download_dataset` tool downloads a split and writes it to a local file, then returns the path. Call `get_dataset_schema` first to see the live catalogue: the feature sets, the target list and which splits are labeled.
{% endtab %}

{% tab title="curl" %}

```bash
curl -L -H "X-API-Key: $EIQ_API_KEY" \
  https://api.everesteer.ai/api/v1/futures/data/train \
  -o train.parquet
```

Large downloads may redirect to a short-lived signed URL. The SDK follows these redirects automatically; with curl, pass `-L` to follow them.
{% endtab %}
{% endtabs %}

## The split tokens

The futures download route accepts these tokens:

| Token                         | File                                      |
| ----------------------------- | ----------------------------------------- |
| `train`                       | `eiq_train.parquet`                       |
| `validation`                  | `eiq_validation.parquet`                  |
| `live`                        | `eiq_live.parquet`                        |
| `train_benchmark_models`      | `eiq_train_benchmark_models.parquet`      |
| `validation_benchmark_models` | `eiq_validation_benchmark_models.parquet` |
| `live_benchmark_models`       | `eiq_live_benchmark_models.parquet`       |
| `validation_example_preds`    | `eiq_validation_example_preds.parquet`    |
| `live_example_preds`          | `eiq_live_example_preds.parquet`          |
| `features.json`               | `eiq_features.json`                       |
| `metadata.json`               | `eiq_metadata.json`                       |

Each token maps to a single file. Atlas publishes no separate designated-benchmark file: the designated benchmark is the `v1_sherpa` column of the benchmark files. An event-scoped key is served the event tree and can fetch `train`, `validation`, `live`, `train_benchmark_models`, `validation_example_preds`, `features.json` and `metadata.json`. The validation and live benchmark files and the live example predictions return 404 for it while an event runs, so that nobody can clone the event benchmark's board score. The file table with what is inside each one is on the [All dataset files](/data/all-dataset-files.md) page.

## The versioned download path

Full-scope keys can also download through the explicit-version path:

```bash
curl -L -H "X-API-Key: $EIQ_API_KEY" \
  https://api.everesteer.ai/api/v1/data/download/{version}/futures/{split} \
  -o train.parquet
```

`GET /api/v1/data/versions` returns the current version and the universes it covers. Pass the returned version string in the path. A mismatched version is rejected with 404 and the error names the current value. This path is for full-scope keys; event-scoped keys are served the event tree through the futures path instead.

For benchmark predictions there is a dedicated SDK method, `download_benchmark`, documented on the [Benchmark models](/data/benchmark-models.md) page.

## Freshness

The three core splits are static between rounds. The `live` artifact updates when a new round opens, and `eiq_features.json` and `eiq_metadata.json` update with it. Re-download `live` for every round: each round's live split is a new id namespace, and a stale download carries ids that no longer join the open round.

| Condition                                           | Behaviour                                                                           |
| --------------------------------------------------- | ----------------------------------------------------------------------------------- |
| No round is open                                    | `live` returns 404.                                                                 |
| The next round is being republished (event cadence) | `live` returns 409 while the republish is in progress, with a `Retry-After` header. |
| Unknown split token                                 | 400, naming the valid set.                                                          |

Between rounds on the tournament lane, `live` returns 404 until the next round opens. The 409 refusal is the event cadence lane, while the next sealed round is being published.

## The two JSON files

* `eiq_features.json` carries the feature sets and the target list for the active dataset version. Its top-level keys are `feature_sets` and `targets`. Read the live catalogue from here or from `get_dataset_schema`; do not copy a static list into your code.
* `eiq_metadata.json` carries the dataset version, the frequency and basis of expeds, the split boundaries, the feature and target counts, the primary target, and the feature and target encodings. Point at this file rather than quoting numbers on these pages.

See [All dataset files](/data/all-dataset-files.md) for the shapes of both files.


---

# Agent Instructions
This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com.

## Querying This Documentation
If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question.

Perform an HTTP GET request on the current page URL with the `ask` query parameter, and the optional `goal` query parameter:

```
GET https://docs.everesteer.ai/data/datasets.md?ask=<question>&goal=<endgoal>
```

`ask` is the immediate question: it should be specific, self-contained, and written in natural language.
`goal` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal.

The response will contain a direct answer to the question and relevant excerpts and sources from the documentation.

Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections.
