Submission path

Bring scores. Get a task-specific answer.

The public repository accepts prediction tables. Your method runs where you choose, and automatic preflight checks the submission as data. Sealed scoring is separate and is not active at public launch.

preflight live

Automatic checks without arbitrary code execution

The first preflight-enabled track is the blinded Qwen2.5-7B-Instruct authorization task. A passing preflight is not a benchmark score. The issue comment states whether the separately operated sealed evaluator is enabled. APPS remains visible as an open adoption panel while its public query export is prepared. Code submissions remain local or Colab-based until a sandboxed runner is available.

Choose an entry path

Two ways to participate

local pathCustom implementation

Run an observer

Run the method in your own environment, then turn its estimates into predictions and an ObserverCard.

  • Supports activation, attribution, DLA, SAE, or custom methods
  • Model weights stay in your environment
  • Produces the same prediction and ObserverCard boundary

What happens to a job

Five mechanical stages

  1. Submitted

    The issue records a task version, comparison track, ObserverCard, and prediction artifact.

  2. Validated

    Automation checks the schema, query IDs, finite scores, duplicates, declared access, and hashes.

  3. Evaluated

    When enabled, a protected workflow joins scores with sealed targets and runs the fixed policy, budget, and loss.

  4. Returned

    A sealed evaluation returns metrics, uncertainty, baseline deltas, failure slices, and a job identifier.

  5. Published

    After approval, an eligible result becomes a permanent result page and a checked leaderboard row.

Maintainer workload

Routine preflight needs no manual review.

Sealed scoring and publication follow the status shown by the repository.

Pass preflight automatically when

  • The task and comparison track already exist
  • Only inert CSV and JSON data are submitted
  • All IDs, hashes, and scores validate, and an access tier is declared
  • The result identity is new and the workflow is reproducible

Review only when

  • The submission adds or executes code
  • It changes the task, controller, budget, targets, or loss
  • The author requests an implementation-verified badge
  • There is a security warning, identity collision, or appeal

Returned artifact

A sealed evaluation produces the same core evidence for every observer.