A task-relative observer workbench

Does your estimate improve the action that matters?

ObserverBench compares internal-state estimates under a fixed task, intervention rule, budget, and loss. It reports both how well an observer predicts and what happens when a controller acts on it.

Mechanistic Tomography develops the measurement framework; ObserverBench tests the resulting estimate through action.

Release status

The public repository automates preflight. Sealed scoring is a separate service.

Automatic at launch

  • Validate a prediction table and ObserverCard
  • Require an immutable GitHub commit URL
  • Check schema, IDs, size, values, and declared access
  • Store a short-lived inert preflight artifact
  • Report pass or fail without executing contributor code

Requires the sealed evaluator

  • Join predictions to evaluator-held targets
  • Return a redacted scorecard
  • Compare with the matching checked panel
  • Approve a result for leaderboard publication

This service is not active at public launch. The issue comment will state if it is enabled later. A preflight pass alone is not a score or rank.

One clear path

From observer to evidence

Visitors can prepare and validate a submission without a GPU or a conversation with the maintainers.

  1. 1

    Choose a task

    The task fixes what can be observed, what action follows, and what counts as loss.

  2. 2

    Run or upload

    Use frozen measurements, a Colab runner, or upload one risk or effect score per query.

  3. 3

    Pass preflight

    Public automation checks IDs, schema, provenance, access fields, and finite values.

  4. 4

    Receive a scorecard

    When sealed scoring is enabled, see prediction quality, action loss, uncertainty, cost, and failure slices together.

  5. 5

    Compare fairly

    An approved result gets a position only among rows using the same task version and comparison track.

Tangible outcomes

What a visitor takes away

01

A direct answer

Did this observer beat no action and the relevant existing methods on downstream loss?

02

A fair comparison

Where does it sit among observers with the same information, budget, controller, and target?

03

A failure map

Which targets, prompt families, action budgets, or operating conditions make it fail?

04

A durable artifact

A downloadable scorecard, provenance hashes, and a result page that another researcher can reproduce.

Task catalog

Four checked questions from current models

View all checked rows →
compression testSafety · Qwen2.5-7B

Does a sparse observer preserve the decision?

The SAE readout uses 127 fitted coefficients instead of 3,584 residual coordinates, but mean violations rise from 2.05 to 4.67. Compression without a better decision.

Published SAE integration · 1% attacks · 2% audit

interaction gainMechanisms · Qwen2.5-7B

Do pair terms improve held-out effects?

On the induction-copy task, an interaction-aware predictor lowers held-out MAE from 0.121 to 0.040 with the same 128-measurement budget.

Qwen2.5-7B-Instruct · frozen intervention table

Ranking rule

No global “best observer.”

A score is task-relative. ObserverBench can return an official rank only when the task version, observation boundary, action budget, controller, and loss match. Cross-track results remain visible as comparisons, and evaluator oracles remain unranked reference bounds.