A rigorous, open-source instrument that treats systemic awareness as the prerequisite for autonomy — the construct, the lineage, the dimensions, the engine, and the rigor behind SystemsBench.
SystemsBench marks a pivot in AI evaluation — beyond the industry's obsession with trivia recall and code generation. At its core is the Systems Thinking & Leverage Competence (STLC) construct: a model's capacity to identify and understand complex systems, predict their behavior through feedback and stocks, and devise high-leverage modifications that produce the intended effect.
In an era when we are rushing to grant AI control over hospitals, power grids, and financial markets, the absence of an STLC metric is a terrifying blind spot. Current benchmarks are caught in the Parameter Trap — rewarding models for fiddling with surface-level numbers while remaining blind to the structural domino effects of their actions. SystemsBench measures the missing axis of consequence.
Perception (I) — identify boundaries, structure, and feedback loops.
Dynamics (II) — reason about movement through time: stocks, flows, delays, nonlinearity, emergence.
Intervention (III) — find high-leverage points and design interventions that survive contact with the real world.
Running through all three strata: does the model know the limits of its own understanding, and refuse to impose false determinism on complex domains? A mind confidently wrong about a system is more dangerous than one that knows it is uncertain. Calibration decides how much any other score can be trusted.
SystemsBench stands on a deliberate lineage. Each theorist contributes a specific, load-bearing piece of the construct.
| Theorist | Specific contribution to SystemsBench |
|---|---|
| Donella Meadows | The 12-point leverage ladder (Parameters → Paradigms); the "Dutch Houses" information-flow paradigm. |
| Peter Senge | The Iceberg Model; system archetypes — Tragedy of the Commons, Fixes-that-Fail, the Antibiotic Spiral. |
| Jay Forrester | System Dynamics; stock-flow-feedback-delay simulation; the counterintuitive behavior of social systems. |
| Russell Ackoff | The distinction between dissolving (structural redesign) and solving (symptom patching). |
| Herbert Simon | Bounded Rationality — the framework for identifying structural causes over actor blame. |
| Sweeney & Sterman | Bathtub Dynamics — the empirical proof of the "Misperception Trap" in stock-flow inference. |
Drawing from Senge, we require models to look below the waterline. Reactive answers score low; structural and paradigm-level diagnosis scores high.
Following Ackoff: patching symptoms within the existing system is solving; redesigning the system so the problem becomes structurally impossible is dissolving.
Stop the oil leak in a combustion engine. The leak — and every future leak — remains possible. The structure is untouched.
An "oil leak" becomes mathematically and structurally impossible. The problem class is gone, not managed. This is the move SystemsBench rewards.
STLC is measured across five dimensions (A–E), each weighted by its importance to the final Systems Quotient. Leverage placement carries the most weight — it is the closest proxy for real-world consequence.
Can the model map stocks, flows, and feedback loops — with correct polarity — before it prescribes? Grounded in Forrester and Sweeney & Sterman's Bathtub Dynamics.
Does it locate causes in systemic structure rather than blaming individuals? Grounded in Senge's Iceberg and Simon's Bounded Rationality.
The "Holy Grail." Interventions are ranked on the Meadows 12-point hierarchy. High scores require structural leverage — e.g. restructuring information flows (the "Dutch Houses," where moving electricity meters from the basement to the hallway cut consumption ~30% with no price change).
Can it anticipate behavior-over-time — overshoot, oscillation, policy resistance? Grounded in Sterman's Business Dynamics and Senge's archetypes.
Calibration and domain matching (Complex vs. Complicated). Grounded in Snowden's Cynefin and Checkland's Soft Systems Methodology.
Where a model chooses to intervene is located on this hierarchy — from the lowest leverage (12) to the highest (1).
Models are forbidden from using "paradigm shift" (Level 1) as a hollow slogan. The score rewards the highest-feasible leverage: a well-justified Level 6 intervention scores higher than an unfeasible, unjustified Level 1 claim. Ambition without feasibility is not insight.
The Davara operator runs the nine-phase SenseRun loop so the benchmark compounds in quality — one bounded, reversible, gated, logged enhancement per run. Each phase carries a guardrail.
We reject the industry impulse to fill dashboards with fuzzy numbers. If a dimension's jury cannot be certified, the score is reported as UNCALIBRATED. AI judges must clear two mathematical gates against a human gold set before any score is trusted.
Ensures AI judges agree with the nuance of human expert labels — not just the easy cases.
Ensures the AI jury ranks models in the same relative order that human experts do.
The original safety gate required 2 of 3 judges to be within 0.5 points on a three-point scale (0, 0.5, 1.0). The engine proved this was a presence check, not an agreement check: by the Pigeonhole Principle, with only three possible scores it is impossible for three judges not to have at least two within a 0.5-point span. The rule was mathematically unfalsifiable — and was caught and removed before it could certify anything.
A private, expert-validated vault of L4 items, held offline and never published. If a model collapses on the Diamond Split while excelling on public items, we have confirmed memorization over intelligence.
Every run regenerates surface variables — industry, actors, numbers — from symbolic templates. Prompts are date-stamped past the training cutoff to expose the "contamination cliff," where a model leans on training data instead of live reasoning.
SystemsBench improves its rubrics and prunes low-discrimination items faster than labs can scrape or memorize them. The instrument evolves alongside the intelligence it measures.
In v0.8 the gold set is still empty: 299 of 300 items carry no gold reference, so the leverage-graded half of the benchmark stays dark and every jury sub-score reports UNCALIBRATED. Populating it with expert labels is the immediate hurdle to certified STLC scoring.
Move from text descriptions to live, mathematically rigorous environments (e.g. regional power grids). We will no longer grade a model's words — we will grade the behavior-over-time of the environment under its control.
A high SQ is impossible without calibration. Capability without failure honesty is a hazard, not an intelligence.
SystemsBench is the pre-flight check for civilization — ensuring an AI can see the curvature of the world before we grant it control.
Map the system. Find where to push.
Predict what happens. Make humanity proud.