Upload predictions
Return one score for each frozen query and describe how the score was produced.
- Fast, inference-free evaluation
- Sealed held-out targets
- Automatic preflight; scorecard only when sealed scoring is enabled
- No submitted code executes
Submission path
The public repository accepts prediction tables. Your method runs where you choose, and automatic preflight checks the submission as data. Sealed scoring is separate and is not active at public launch.
The first preflight-enabled track is the blinded Qwen2.5-7B-Instruct authorization task. A passing preflight is not a benchmark score. The issue comment states whether the separately operated sealed evaluator is enabled. APPS remains visible as an open adoption panel while its public query export is prepared. Code submissions remain local or Colab-based until a sandboxed runner is available.
Choose an entry path
Return one score for each frozen query and describe how the score was produced.
Run the method in your own environment, then turn its estimates into predictions and an ObserverCard.
What happens to a job
The issue records a task version, comparison track, ObserverCard, and prediction artifact.
Automation checks the schema, query IDs, finite scores, duplicates, declared access, and hashes.
When enabled, a protected workflow joins scores with sealed targets and runs the fixed policy, budget, and loss.
A sealed evaluation returns metrics, uncertainty, baseline deltas, failure slices, and a job identifier.
After approval, an eligible result becomes a permanent result page and a checked leaderboard row.
Maintainer workload
Sealed scoring and publication follow the status shown by the repository.
Returned artifact