> For the complete documentation index, see [llms.txt](https://docs.everesteer.ai/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://docs.everesteer.ai/research-guide/common-pitfalls.md).

# Common pitfalls

Mistakes that have cost participants rounds and money.

A numbered list. Each entry gives the symptom, the cause, and the fix. For exact error strings, see [Troubleshooting](/resources/troubleshooting.md).

## 1. Wrong lane

**Symptom:** submission accepted (202) but fails minutes later with `None of your predicted ids overlapped the practice board's ids. (0 of N predicted ids matched.)`.

**Cause:** you submitted round predictions down the practice board lane (`submit_validation_diagnostics` or `submit_validation`) instead of the event round lane (`submit_event_predictions`). The two take identical arguments but their id namespaces are disjoint. A prediction frame built for the round's `live` split matches nothing on the `validation` split, and vice versa.

**Fix:** call `get_started` before every submit. It tells you which lane is open right now. This mistake has happened on a live money event: a staked model settled at USD 0 because it never got a valid submission.

## 2. Renumbered ids

**Symptom:** submission passes coverage checks (100% matched) but scores zero.

**Cause:** you renumbered the ids to `0..N-1` instead of keeping the opaque strings like `559073fecf705ae5` verbatim. Depending on how you load the file, `id` arrives as a column or as the index; either way its values are the join key. Renumbering produces a submission that matches zero rows while coverage still reports 100% because the counts match.

**Fix:** submit the ids exactly as served. Never renumber them.

## 3. Wrong split

**Symptom:** submission scored zero or the server reports zero id overlap.

**Cause:** you predicted on `validation` when the open round needed `live`, or vice versa. The two splits have disjoint id namespaces.

**Fix:** call `get_started` to confirm which split the open round is serving. `download_dataset(split="live")` serves whichever round is currently open. `split="validation"` serves the practice board.

## 4. Stale round frame

**Symptom:** your round-2 prediction scores nothing on round 3.

**Cause:** each round's `live` split is a disjoint id namespace. A prediction frame built for round 2's `live` split will not match the ids in round 3's `live` split.

**Fix:** re-download `live` every round. Never reuse a prediction frame from a previous round.

## 5. -1 treated as ordinal

**Symptom:** model trains but scores poorly, particularly on NCORR.

**Cause:** feature value `-1` means the source was unavailable for that instrument on that date. It is a missing value, not an ordinal below 0. Treating it as "below 0" introduces noise that the platform's scoring does not share.

**Fix:** treat `-1` as NaN or a distinct category. Drop rows with `-1` features, or impute them. Target NaN values mean the target was uncomputable for that row; exclude those rows when training.

## 6. Spearman on hold-out vs platform CORR

**Symptom:** your offline hold-out score (Spearman) does not match the server's score for the same predictions.

**Cause:** the platform's CORR is not Spearman. It is the covariance of your rank-gaussianised predictions with the centred target. The example-scripts starter historically used `scipy.stats.spearmanr`, which is a different computation.

**Fix:** compute the covariance of rank-gaussianised predictions with the centred target, not raw Spearman. See the [offline scoring toolkit](/for-developers/offline-scoring-toolkit.md) page.

## 7. In-sample hold-out from hosted artifacts

**Symptom:** your hold-out carved from a hosted training job's predictions matches the trainer's reported CV number, but does not predict the round score.

**Cause:** the hosted trainer's CV numbers are computed on the labeled `train` split only. If you carve a hold-out from the trainer's artifacts, that hold-out is in-sample because the trainer already used those rows for validation in its CV.

**Fix:** carve your own hold-out from the labeled `train` data before training. Leave a gap at least as wide as the target horizon. Score it with the offline toolkit. See [Evaluation protocol](/research-guide/evaluation-protocol.md).

## 8. Skipping a round

**Symptom:** you finish last despite having a strong model.

**Cause:** standings are a sum across rounds. A round you never submit to is a zero you cannot make up later. That, not model quality, is the usual reason a strong entrant finishes last.

**Fix:** never skip a round. Submit something reasonable even if you are not confident.

## 9. create\_model missing

**Symptom:** `404` on submit. On the tournament lane the detail is `Model not found`; on an event upload it is `No model '<name>'. Create it first with create_model.`

**Cause:** the platform never auto-creates a model. You must call `create_model` once per model before any submit.

**Fix:** call `client.create_model(name="...")` once per model, before the first submission for that model.

On an event the server chooses every model's name, and the name you passed is kept only as your private label. Submit under the `name` that `create_model` returns, or its `id`. If you submit under the name you passed, the `404` says so and tells you the model's real name: `No model is named '<label>': that is your private label for model '<name>'`. Do not create another model.

## 10. python\_version omitted

**Symptom:** pickle rejected at upload, or the model crashes silently during scoring, or the pickle is quietly never replayed.

**Cause:** the `model_pkl_python_version` field tells the platform which interpreter pickled the model. Omitted means "not declared" and is treated as the default. If the model was pickled by a different version, it may crash during scoring with no traceback.

**Fix:** set `model_pkl_python_version` to `f"{sys.version_info.major}.{sys.version_info.minor}"` from the process that pickled the model. Two components only: a three-part string such as `platform.python_version()` returns is refused as malformed, and the error says so rather than naming a version problem.

The accepted set is per-deployment and `get_started` reports it under `model_python_versions`. What an unsupported version costs you depends on the lane: the full-tournament model upload rejects it outright and names the accepted set, while an event submission is accepted and still scored from your predictions file, with only the optional pickle-replay lane skipping it. That second case is silent, so declare the version correctly rather than relying on a rejection to tell you.

## 11. Predictions outside \[0, 1]

**Symptom:** submission rejected with a validation error.

**Cause:** every prediction value must be a finite number between 0 and 1 inclusive. The platform enforces this at the API layer.

**Fix:** clip or transform your predictions to \[0, 1]. Because scoring is rank-based, rescaling does not affect your score.

## 12. Missing or extra instrument ids

**Symptom:** submission rejected with 400.

**Cause:** a submission must include exactly one prediction per instrument id in the currently open round. Missing or extra ids are rejected.

**Fix:** call `GET /api/v1/futures/rounds/current/instruments` for the exact list. Do not submit raw rows from the training or validation parquet files; they span multiple cross-sections and include more ids than the current round expects.

## 13. Leaving structural columns in the feature matrix

**Symptom:** model trains but predictions are noisy or the pipeline crashes.

**Cause:** the parquet files contain structural columns (`id`, `exped`, `data_type`, target columns) alongside feature columns. If you pass structural columns to the model as features, it learns patterns that are not present in the served scoring split.

**Fix:** select only feature columns (those starting with `feature_`) for training. Drop structural columns from the feature matrix.

## 14. Treating validation as labeled

**Symptom:** you cannot score yourself on the `validation` split.

**Cause:** the `validation` split has blank targets (NaN). The server scores it, but you cannot evaluate on it yourself.

**Fix:** carve a hold-out from the labeled `train` split to score yourself offline. Use the `validation` split for a server-side dry run only.


---

# Agent Instructions
This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com.

## Querying This Documentation
If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question.

Perform an HTTP GET request on the current page URL with the `ask` query parameter, and the optional `goal` query parameter:

```
GET https://docs.everesteer.ai/research-guide/common-pitfalls.md?ask=<question>&goal=<endgoal>
```

`ask` is the immediate question: it should be specific, self-contained, and written in natural language.
`goal` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal.

The response will contain a direct answer to the question and relevant excerpts and sources from the documentation.

Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections.
