> For the complete documentation index, see [llms.txt](https://docs.everesteer.ai/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://docs.everesteer.ai/research-guide/evaluation-protocol.md).

# Evaluation protocol

Split by exped, leave a gap, score per exped, reproduce the round score.

The platform scores your predictions per exped (cross-section), then averages across the round. An honest offline evaluation follows the same shape.

## Split by exped, never by row

The time column is `exped` (values like `exped_0850`). Each exped is one forward-looking target. Fitting on some rows of an exped and testing on other rows of the same exped uses the same future information for both training and evaluation. That overstates your score.

Always split by exped: fit on earlier expeds and hold out a later block.

## Leave a gap

The primary target is forward-looking over a fixed horizon. On the tournament dataset the primary target `target_everest` is a 20-day forward return, but no target name carries its horizon, so read the horizon from the dataset's own documentation rather than from the column name. The label of exped N and the label of exped N+1 overlap in time: the target for N draws from forward returns starting at N, and the target for N+1 draws from forward returns starting at N+1. The two share 19 of 20 days.

Leave a gap at least as long as the primary target's horizon between your last fit exped and your first hold-out exped. A gap of 20 expeds ensures no hold-out label overlaps your fit window. Without it, your model has seen information from the hold-out period even though the rows are separate.

## Walk-forward folds

For a more robust estimate, use walk-forward folds: several train/hold-out splits where the train window grows and the hold-out window slides forward, each with a gap. The platform's hosted training does this internally with an exped-purged, embargoed cross-validation.

## Score per exped, then average

Compute the term (CORR, AIMC, NCORR) on each exped separately, then average across expeds. This matches how the server scores your submission.

```python
from everestapi import scoring

# target_col is the column your dataset is graded on. Read it once from
# client.get_dataset_schema()["primary_target"], never from a remembered name.
corrs = []
for exped in hold_out["exped"].unique():
    mask = hold_out["exped"] == exped
    c = scoring.corr20(predictions[mask], hold_out.loc[mask, target_col])
    corrs.append(c)
mean_corr = sum(corrs) / len(corrs)
```

## Use the offline toolkit

The `everestapi[scoring]` extra ships the scoring functions. Install it:

```bash
pip install "everesteer-api[scoring]"
```

The `everestapi[scoring]` toolkit still implements the earlier CORR and NCORR definitions (a power-transformed Pearson correlation, and a correlation after neutralising against the core features), so its `corr20` and `ncorr` do not reproduce the server's current terms. `aimc20` matches the server kernel. Official numbers are server-side.

See the [offline scoring toolkit](/for-developers/offline-scoring-toolkit.md) page for full function signatures.

## Reproduce a board's ordering

The round score is `b * arctan((CORR + AIMC + NCORR) / b)`, bounded per round. To reproduce it:

1. Compute each term per exped, then average.
2. Read the scale `b` from `GET /api/v1/scoring` (SDK: `explain_scoring`), and apply the arctan to the sum.
3. Payout is stake times that score.

The `rank_metric` field on every leaderboard response tells you what that board was ordered by: `round_score` (scored boards), `corr20` (fallback before scoring), `final_corr20` (held-out final window), or `mean_payout` (tournament agent board).

## The validation split is a practice board

The `validation` split has blank targets. It is scored server-side and is useful as a dry run. It is not a hold-out you can evaluate on yourself. The platform's hosted training reports CV numbers from the labeled `train` split only; the validation split is never scored in-job.

## Degenerate cross-sections

On an exped with too few overlapping ids or no variance, a term scores 0.0 and enters the average as a zero. A term is reported as null only when no exped could produce it at all (for example, none of the neutralisation features present).


---

# Agent Instructions
This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com.

## Querying This Documentation
If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question.

Perform an HTTP GET request on the current page URL with the `ask` query parameter, and the optional `goal` query parameter:

```
GET https://docs.everesteer.ai/research-guide/evaluation-protocol.md?ask=<question>&goal=<endgoal>
```

`ask` is the immediate question: it should be specific, self-contained, and written in natural language.
`goal` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal.

The response will contain a direct answer to the question and relevant excerpts and sources from the documentation.

Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections.
