> For the complete documentation index, see [llms.txt](https://docs.everesteer.ai/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://docs.everesteer.ai/submissions/model-upload-pkl.md).

# Model upload (.pkl)

What a pickled model must expose and how it is checked.

Upload a cloudpickled `predict` function and the platform runs it for you against live features on every open round. This is the hosted inference lane: you submit the model once, and it produces predictions automatically each round.

**This lane is available for full tournament (Himalayas) keys.** Event-scoped keys cannot use this lane: they must attach their `.pkl` to each upload via `submit_event_predictions` or `submit_validation_diagnostics` instead.

## Auto-submit: set this up first

With auto-submit enabled, the platform runs your uploaded `.pkl` on every round right after it opens and writes a submission for you. No per-round call is needed. This is the recommended way to take part; the manual submit loop is the fallback.

1. `create_model` registers the model. Every new model is opted in to auto-submit, but that alone runs nothing.
2. Train it and save a cloudpickled `predict` function (below).
3. `upload_model` stores the `.pkl`. The sandbox smoke test must pass.
4. `set_auto_submit(enabled=True)` enrols the model. Enabling on a model with no uploaded `.pkl` is refused.
5. Check `get_models`.

```python
client.create_model(name="my-model-v1")
client.upload_model("my-model-v1", "model.pkl", python_version="3.11")
client.set_auto_submit(model_id="my-model-v1", enabled=True)
```

### Check that it will actually run

`auto_submit: true` only records that you opted in. The truthful signal in `get_models` is `lane_active`:

| Field           | Meaning                                                                                                                                                                 |
| --------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `lane_active`   | `true` means the platform will run this model each round. `false` means it will not.                                                                                    |
| `lane_note`     | When `lane_active` is `false`, the reason in plain words: no `.pkl` uploaded, the upload failed the sandbox check, or the deployment does not auto-run uploaded models. |
| `sandbox_ready` | `1` if the latest upload passed the sandbox predict, `0` if not, null if there is no `.pkl`.                                                                            |

If `lane_active` is `false`, read `lane_note`, fix or re-upload the model, or submit directly with `submit_futures_predictions` in the meantime.

### Then join the historical leaderboard

Also submit your validation predictions with `submit_validation_diagnostics` and read the board with `get_diagnostics_leaderboard`. See [Leaderboards and rank\_metric](/scoring/leaderboards-and-rank-metric.md).

## Upload endpoint

`POST /api/v1/models/{model_id}/upload` : multipart file upload, auth required (`X-API-Key`). The model must already exist (register it with `create_model` first).

### Fields

| Field            | Required | Description                                                                         |
| ---------------- | -------- | ----------------------------------------------------------------------------------- |
| `model_id`       | Yes      | Path parameter. The model name must exist.                                          |
| `file`           | Yes      | A `.pkl` file containing a cloudpickled `predict` function.                         |
| `python_version` | No       | `"major.minor"` of the interpreter that saved the pickle. Omitted defaults to 3.11. |

### Size limit

The file size limit is named in the rejection error. Both the `Content-Length` header and the actual read body are checked; either one over the limit returns 413. The error message names the bound.

## The accepted pickle shape

**One shape: a cloudpickled callable named `predict`.** Either arity works:

```python
def predict(live_features): ...
def predict(live_features, live_benchmark_models): ...
```

`live_features` is a DataFrame indexed by instrument id, carrying only `feature_*` columns. `live_benchmark_models` is a single series named `eiq_minera_model`: the published live benchmark reference, reindexed to the live feature table by id.

Return a **single-column `pandas.DataFrame`** indexed by instrument id. The first column is the one scored. A bare `pandas.Series` is not accepted. A `DataFrame` subclass is also refused.

**Every returned value must be inside \[0, 1].** Out-of-range predictions fail the run on every round, not just at upload. Values are not clipped. Scoring is rank-based, so a linear rescale into \[0, 1] costs you nothing.

```python
import cloudpickle
import pandas as pd

feature_cols = [c for c in train.columns if c.startswith("feature_")]
model = my_regressor.fit(train[feature_cols], train["target_everest"])

def predict(live_features):
    raw = model.predict(live_features[feature_cols])
    ranked = pd.Series(raw, index=live_features.index).rank(pct=True)
    return ranked.to_frame()

with open("model.pkl", "wb") as f:
    cloudpickle.dump(predict, f, protocol=5)
```

Use `cloudpickle.dump`, not `pickle.dump`: it serialises the function together with everything it closes over (your fitted model, your feature list), so the file is self-contained. There is no network access inside the sandbox.

### What gets rejected at upload

Before your file is ever loaded, it is parsed as a pickle opcode stream without executing any opcodes. Uploads are rejected for:

* Empty file
* Torch models : any opcode whose module string is `torch` or starts with `torch.` fails upload. The rejection is transitive: wrapping the network in a class that exposes `.predict()` does not help because the `nn.Module` tensors emit `torch.*` into the opcode stream.
* Truncated or corrupt pickle
* A bare estimator, a `dict` with a `'meta'` key, a `dict` wrapping an estimator with `feature_cols`, and hand-rolled ensemble containers are refused at upload.

### Models from hosted training

A `.pkl` produced by hosted `train` is already in the accepted shape, so you can download it and upload or attach it as it is. It is a cloudpickled `predict` callable that selects its own training columns by name from whatever features frame you hand it, applies the platform missing-value convention, and returns a single-column frame of rank-normalized predictions. Column order in your frame does not matter, and a training column missing from the frame raises rather than returning wrong numbers.

The `feature_manifest` returned alongside it records the columns the model was fit on. Read it if you want to rebuild the frame yourself; you no longer need it to use the file.

### The pickle check flow

1. Static opcode scan (no execution) : rejects torch, empty, corrupt files.
2. Sandbox smoke test : your function is called against the full live `feature_*` table in an isolated sandbox matching the declared Python version. The upload HTTP response returns after the file is stored; the smoke test continues in the background.
3. Smoke success activates the model for the auto-run lane. Smoke failure keeps the model in `pending` with an error message and does not replace a prior ready upload.

## Status lifecycle

Each run of your uploaded model against an open round moves through these states. Poll `GET /api/v1/models/{model_id}/upload/status`:

| Status      | Meaning                                                                        |
| ----------- | ------------------------------------------------------------------------------ |
| `pending`   | Upload received, queued for the next round.                                    |
| `running`   | The sandbox job is executing your model.                                       |
| `submitted` | Predictions were produced and written as a real submission for that round.     |
| `skipped`   | No open round was available to submit against (not a failure).                 |
| `error`     | The sandbox run failed. Check the error message on the upload status endpoint. |

## Common rejections

| Error text                                                                                                    | Status | Cause                                                    |
| ------------------------------------------------------------------------------------------------------------- | ------ | -------------------------------------------------------- |
| `File too large. Maximum upload size is ...`                                                                  | 413    | The pickle exceeds the size limit; the message names it. |
| `Only .pkl files are accepted.` or `The model file must be a .pkl file (got ...). The file was not uploaded.` | 400    | File attached has the wrong extension.                   |
| `Model not found`                                                                                             | 404    | Model not registered. Call `create_model` first.         |
| `python_version must be 'major.minor' (e.g. '3.12'), got '...'`                                               | 400    | Python version declared in the wrong format.             |


---

# Agent Instructions
This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com.

## Querying This Documentation
If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question.

Perform an HTTP GET request on the current page URL with the `ask` query parameter, and the optional `goal` query parameter:

```
GET https://docs.everesteer.ai/submissions/model-upload-pkl.md?ask=<question>&goal=<endgoal>
```

`ask` is the immediate question: it should be specific, self-contained, and written in natural language.
`goal` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal.

The response will contain a direct answer to the question and relevant excerpts and sources from the documentation.

Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections.
