> For the complete documentation index, see [llms.txt](https://docs.everesteer.ai/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://docs.everesteer.ai/for-developers/offline-scoring-toolkit.md).

# Offline scoring toolkit

Score a hold-out offline with the platform's metric definitions.

The `everestapi[scoring]` extra provides functions that implement close approximations of the platform's CORR20, AIMC and NCORR transforms. Use them to score a hold-out you carve from the `train` split before you spend a submission. The official numbers are computed server-side; the toolkit is for relative comparisons and tuning.

```bash
pip install "everesteer-api[scoring]"
```

## Functions

Every function accepts pandas Series or numpy arrays.

### `corr20(predictions, target)`

Rank-gaussianize the predictions, centre the target, apply the signed power transform, then Pearson correlation. Returns a float; 0.0 on a degenerate input.

### `aimc20(predictions, ai_model, target)`

Rank-gaussianize both the predictions and the benchmark series, remove the component of the predictions that lies along the benchmark, then take the covariance of the residual with the centred target. Returns a float.

### `ncorr(predictions, features, target)`

Neutralise the rank-gaussianized predictions against the feature columns you pass, then score the residual with the CORR kernel. Returns a float.

### `feature_exposure(predictions, features)`

Maximum absolute correlation between the predictions and any single feature column. Returns a float.

### `payout(corr20, aimc, *, corr_weight, aimc_weight, payout_cap, ncorr=0.0, ncorr_weight=0.0, stake=1.0, payout_factor=1.0, stake_return_amplitude=None)`

Apply the payout formula. Call `explain_scoring` for the live weights and clip. Everything after the two positional terms must be passed by keyword; `stake_return_amplitude` is the bounded-return knob a money event reports.

### `score(predictions, target, *, ai_model=None, features=None, corr_weight=None, aimc_weight=None, ncorr_weight=None, payout_cap=None)`

Compute every metric the inputs allow in one call: always `corr20`; `aimc20` when `ai_model` is given; `ncorr` and `feature_exposure` when `features` is given; a `payout` when the weights are given. Returns a dict.

## Correct evaluation protocol

Score a hold-out you carve from the labeled `train` split. Do not score against the `validation` or `live` splits: their target columns are blank (NaN).

```python
import pandas as pd
from everestapi import scoring
from everestapi import EverestAPI

client = EverestAPI()

# Download the training split
train_path = client.download_dataset(split="train")
train = pd.read_parquet(train_path)

# The graded column is named by the dataset, so read it rather than typing it.
target_col = client.get_dataset_schema()["primary_target"]

# Split by exped: fit on the earlier block, hold out the latest block,
# and leave a gap at least as long as the target horizon between them.
expeds = sorted(train["exped"].unique())
GAP, HOLDOUT = 20, 60
fit_expeds = expeds[: -(GAP + HOLDOUT)]
holdout_expeds = expeds[-HOLDOUT:]
training = train[train["exped"].isin(fit_expeds)].dropna(subset=[target_col])
holdout = train[train["exped"].isin(holdout_expeds)].dropna(subset=[target_col])

feature_cols = [c for c in train.columns if c.startswith("feature_")]
my_model.fit(training[feature_cols], training[target_col])
holdout = holdout.assign(prediction=my_model.predict(holdout[feature_cols]))

# Score per exped, then average, the way the server does
corrs = [
    scoring.corr20(g["prediction"], g[target_col])
    for _, g in holdout.groupby("exped")
]
mean_corr = sum(corrs) / len(corrs)

# For AIMC, align the designated benchmark on id from the train benchmark file
bench = pd.read_parquet(client.download_benchmark(split="train"))
ref = (bench.set_index("id") if "id" in bench.columns else bench)["v1_sherpa"]
ids = holdout["id"] if "id" in holdout.columns else holdout.index
benchmark = ref.reindex(ids).to_numpy()

aimcs = []
for _, g in holdout.groupby("exped"):
    g_ids = g["id"] if "id" in g.columns else g.index
    aimcs.append(scoring.aimc20(g["prediction"], ref.reindex(g_ids), g[target_col]))
mean_aimc = sum(aimcs) / len(aimcs)
```

## Live payout weights

The `payout()` function requires weights as explicit keyword arguments. Read them from the live API rather than hardcoding them.

```python
weights = client.explain_scoring()
corr_weight = weights["weights"]["corr"]
aimc_weight = weights["weights"]["aimc"]
ncorr_weight = weights["weights"]["ncorr"]
payout_cap = weights["weights"]["payout_cap"]

payout_value = scoring.payout(
    corr=corr_value,
    aimc=aimc_value,
    corr_weight=corr_weight,
    aimc_weight=aimc_weight,
    payout_cap=payout_cap,
    ncorr=ncorr_value,
    ncorr_weight=ncorr_weight,
)
```

## Approximation caveat

The official numbers are computed server-side, and the toolkit does not yet match two of them:

* **CORR20.** The toolkit computes a power-transformed Pearson correlation. The server's CORR is the covariance of the rank-gaussianized predictions with the centred target.
* **NCORR.** The toolkit neutralises against the core features and correlates the residual. The server's NCORR is the AIMC kernel with the equal-weight core-feature average in place of the benchmark, so `aimc20(predictions, core_features.mean(axis=1), target)` reproduces it, except on an exped where that average lost (below).
* **AIMC20.** The SDK matches the server kernel exactly, except on an exped where the benchmark lost (below).
* **A reference's losing exped.** Where the benchmark's own covariance with the target is negative, the server's AIMC is 0; where the core-feature average's is, the server's NCORR is 0. The toolkit does not apply this.

Use the toolkit for tuning and relative comparisons. The platform's leaderboard is the final authority.

## See also

* [scoring/definitions.md](/scoring/definitions.md) for the full metric definitions.
* [scoring/payout-and-payout-factor.md](/scoring/payout-and-payout-factor.md) for how a score becomes money.


---

# Agent Instructions
This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com.

## Querying This Documentation
If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question.

Perform an HTTP GET request on the current page URL with the `ask` query parameter, and the optional `goal` query parameter:

```
GET https://docs.everesteer.ai/for-developers/offline-scoring-toolkit.md?ask=<question>&goal=<endgoal>
```

`ask` is the immediate question: it should be specific, self-contained, and written in natural language.
`goal` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal.

The response will contain a direct answer to the question and relevant excerpts and sources from the documentation.

Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections.
