> For the complete documentation index, see [llms.txt](https://docs.everesteer.ai/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://docs.everesteer.ai/data/all-dataset-files.md).

# All dataset files

Every downloadable artifact in the Atlas tree, with what is inside each.

The dataset version is published in `eiq_metadata.json`. Files refresh per round with the open Himalayas round. The `live` artifacts update when a new round opens.

{% tabs %}
{% tab title="Core splits" %}

| File                     | Description                                                                                                                        |
| ------------------------ | ---------------------------------------------------------------------------------------------------------------------------------- |
| `eiq_train.parquet`      | Training split. Obfuscated features, labeled targets, full history. All structural columns present.                                |
| `eiq_validation.parquet` | Validation split. Obfuscated features, target columns present but every value is NaN (blanked). The day-0 practice board.          |
| `eiq_live.parquet`       | Live split. Obfuscated features for the current round. Target columns present but every value is NaN. New id namespace each round. |

All three share the same column layout. See [Column definitions](/data/column-definitions.md) for the full schema.
{% endtab %}

{% tab title="Benchmarks" %}

| File                                      | Description                                                                                                                     |
| ----------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------- |
| `eiq_train_benchmark_models.parquet`      | Benchmark-model predictions over the `train` split. Indexed by `id`, with an `exped` column and one column per benchmark model. |
| `eiq_validation_benchmark_models.parquet` | Benchmark-model predictions over the `validation` split. Same column layout.                                                    |
| `eiq_live_benchmark_models.parquet`       | Benchmark-model predictions for the current round. Same column layout.                                                          |

The benchmark columns are float predictions between 0 and 1. The designated benchmark column is `v1_sherpa`, the series the AIMC term orthogonalises against; on Atlas it is the only benchmark column. Atlas publishes no separate designated-benchmark file. See [Benchmark models](/data/benchmark-models.md) for how to use them.
{% endtab %}

{% tab title="Examples" %}

| File                                   | Description                                                                                     |
| -------------------------------------- | ----------------------------------------------------------------------------------------------- |
| `eiq_validation_example_preds.parquet` | Example predictions over the `validation` split. Indexed by `id`, with one `prediction` column. |
| `eiq_live_example_preds.parquet`       | Example predictions for the current round. Indexed by `id`, with one `prediction` column.       |

On Atlas they carry the designated benchmark's predictions. Use them to confirm your pipeline loads and writes the expected shape. Do not submit them as your own predictions.
{% endtab %}

{% tab title="Metadata" %}

| File                | Description                                                                                                                                                                                                                                                                                                                                                                                                                        |
| ------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `eiq_features.json` | Feature metadata. Top-level keys: `feature_sets` (a dict of set name to list of feature column names) and `targets` (the full list of target column names including the primary, the alias and every auxiliary).                                                                                                                                                                                                                   |
| `eiq_metadata.json` | Build metadata. Top-level keys include `version` (dataset version string), `exped_frequency`, `exped_basis` (what one exped means), `feature_count`, `target_count`, `primary_target`, `platform_target_alias` (the graded column), `feature_encoding`, `target_encoding`, `splits` (the split boundaries) and plain-language notes on obfuscation, targets and freshness. The key set differs between builds: read what is there. |

Read the scale figures from `eiq_metadata.json` rather than copying them into your code. The feature set catalogue in `eiq_features.json` is the live source of truth; do not copy a static list.
{% endtab %}
{% endtabs %}

## Download these files

See the [Datasets](/data/datasets.md) page for the split tokens, the SDK methods and the curl commands.


---

# Agent Instructions
This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com.

## Querying This Documentation
If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question.

Perform an HTTP GET request on the current page URL with the `ask` query parameter, and the optional `goal` query parameter:

```
GET https://docs.everesteer.ai/data/all-dataset-files.md?ask=<question>&goal=<endgoal>
```

`ask` is the immediate question: it should be specific, self-contained, and written in natural language.
`goal` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal.

The response will contain a direct answer to the question and relevant excerpts and sources from the documentation.

Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections.
