A direct answer
Did this observer beat no action and the relevant existing methods on downstream loss?
A task-relative observer workbench
ObserverBench compares internal-state estimates under a fixed task, intervention rule, budget, and loss. It reports both how well an observer predicts and what happens when a controller acts on it.
Mechanistic Tomography develops the measurement framework; ObserverBench tests the resulting estimate through action.
Release status
This service is not active at public launch. The issue comment will state if it is enabled later. A preflight pass alone is not a score or rank.
One clear path
Visitors can prepare and validate a submission without a GPU or a conversation with the maintainers.
The task fixes what can be observed, what action follows, and what counts as loss.
Use frozen measurements, a Colab runner, or upload one risk or effect score per query.
Public automation checks IDs, schema, provenance, access fields, and finite values.
When sealed scoring is enabled, see prediction quality, action loss, uncertainty, cost, and failure slices together.
An approved result gets a position only among rows using the same task version and comparison track.
Tangible outcomes
Did this observer beat no action and the relevant existing methods on downstream loss?
Where does it sit among observers with the same information, budget, controller, and target?
Which targets, prompt families, action budgets, or operating conditions make it fail?
A downloadable scorecard, provenance hashes, and a result page that another researcher can reproduce.
Task catalog
On the frozen APPS task, the verdict-adjacent residual probe reaches AUROC 0.977 and 1.976 mean violations at 1% attacks and a 2% audit budget—the strongest deployable row.
No. The detailed prompted monitor beats the official Gemma Scope SAE probe on AUROC (0.893 vs. 0.870) but produces more violations (10.80 vs. 9.07).
The SAE readout uses 127 fitted coefficients instead of 3,584 residual coordinates, but mean violations rise from 2.05 to 4.67. Compression without a better decision.
On the induction-copy task, an interaction-aware predictor lowers held-out MAE from 0.121 to 0.040 with the same 128-measurement budget.
Ranking rule
A score is task-relative. ObserverBench can return an official rank only when the task version, observation boundary, action budget, controller, and loss match. Cross-track results remain visible as comparisons, and evaluator oracles remain unranked reference bounds.