Not what a model knows. Where it chooses to push. The open-source ruler for the one capability that decides whether an intelligence can be trusted near a living system: can it see what else moves when it acts? Nothing else measures it. So we built it, and we are giving it away.
Led by Ember Seoni · Developed & Maintained by Outlier.Systems · Operated by Davara · Built on the Donella Meadows lineage
No human marks this paper. Pick a wicked, self-reinforcing system below. A hypothetical frontier model must map its structure, find where to push, predict what happens — and reflect on its own reasoning. Then three deterministic scorers and a five-judge blind jury grade the trace, slowly, in front of you.
We have gotten extraordinary at measuring recall, code, and math, and those benchmarks are saturating: everyone clusters at the top and the numbers stop discriminating. The frontier moved. The rulers did not. None of them ask the question that actually decides whether an AI is safe to trust with a hospital, a grid, a market, or a nation:
SystemsBench measures the missing axis. It evaluates the capacity to map a system, trace its feedback, locate its leverage, anticipate emergence, and design interventions that survive contact with the real world. As AI is handed control of systems that do not forgive linear thinking, this becomes one of the most important benchmarks in the field.
Systems intelligence is not one skill. It is three strata — each a precondition for the next — and a fourth layer that runs through all of them. Ten capabilities are scored; one cross-cutting measure governs how much the others can be trusted. Together they form a model's Systems Profile.
A mind that is confidently wrong about a system is more dangerous than one that knows it is uncertain. Calibration runs through all three strata — and decides how much weight any other score deserves. Meadows' first instruction for complex systems was to stay humble, stay a learner.
A single number hides where an intelligence is strong and where it is dangerous. SystemsBench returns a profile — a shape across all ten capabilities — plotted against a human reference. It resolves to one figure: the Systems Quotient (SQ).
A model that reasons brilliantly but does not know when it is wrong cannot post a high SQ — the calibration term holds it honest. Capability without calibration is not intelligence; it is hazard.
For the formats that can be checked like arithmetic, a program computes the score. No judge, no bias, fully reproducible. Three executable lanes ship today, each with a self-test you can run in seconds.
Given the flows, infer the stock's path: the bathtub task that fools most humans. Exact numeric match.
Map the feedback structure. Each loop's polarity is recomputed as the product of its signed edges and checked against the reference. A diagram graded by computation, not opinion.
Predict the behavior over time after an intervention, matched against a reference trajectory. The shape is the proof.
A cross-family panel of judge models, never the candidate's own family, grades against an anchored rubric, reference-guided and swap-averaged, five raters blind to one another. Agreement is measured and published, Krippendorff's α and Kendall's τ, never assumed.
No judge's score counts until it is validated against human-graded answers (α ≥ 0.667 and τ ≥ 0.7). Until then the format reports UNCALIBRATED — not scored: never imputed, never averaged in, never silently zeroed. An unparseable reply yields PARSE_ERROR, not a deceptive zero. We would rather report a blank than a fake number.
# 1. Prove the instrument is sound: every reference round-trips, every trap is caught $ python3 engine/cld-score.py calibrate items/cld_oracle.json → 31/31 PASS $ python3 engine/dyn-score.py calibrate items/dyn_oracle.json → 34/34 PASS $ python3 engine/harness.py selftest → 61/61 PASS # 2. See the question a model would actually receive: scenario + the exact answer schema $ python3 engine/harness.py template DYN DYN-FISH-001 # 3. Grade a structured answer and read the dimension breakdown $ python3 engine/dyn-score.py score items/dyn_oracle.json DYN-FISH-001 resp.json
SystemsBench is not a static document — it is a compounding one. Four Markdown files form a research-and-innovation engine: the spec, the procedure, the memory, and the horizon. Each run makes them sharper.
Most benchmarks are static targets that quietly decay as models memorize them. SystemsBench applies its own discipline to itself. One command, Davara /SystemsBenchSenseRun, makes exactly one bounded, reversible, gated, logged improvement per run, then feeds the result back into the next. Ten runs logged so far. The benchmark compounds instead of rotting.
Twelve principles hold the benchmark up, in the order they were pulled. Each is a decision about where SystemsBench itself is held, and the first four carry most of the weight. The full set, with Meadows quoted directly, lives in FOUNDATIONS.md.
This is a research preview, and we would rather tell you exactly what is real than oversell it. Every number below is derived from the files by the repo's own audit script. If a README disagrees with the count, the README is wrong. That is the point.
We are handing intelligence the controls of living systems: grids, markets, hospitals, supply chains, a piece of itself. The one thing that decides whether that goes well is whether it can see what else moves when it pushes.
Nobody had measured it. So we built the ruler, made it honest by construction, taught it to improve itself, and gave it away. That is our contribution. Fork it. Break it. Make it better.
As models become agents that act on the world, that write, run, deploy, and decide, the valuable, hard-to-fake capability is no longer knowledge recall. It is knowing where to push. SystemsBench is the instrument that measures it.
Before an AI is trusted with a grid, a market, a hospital, or a supply chain, we should know whether it can predict what else moves when it intervenes. SystemsBench is that pre-flight check.
Recall and code are saturating. Systems intelligence is the open frontier — and the one that maps most directly onto real-world consequence. A benchmark here reshapes what labs optimize for.
Open-source, auditable, and self-improving. The rubric, item bank, and calibration data compound every run — a shared instrument the whole field can build on, not a closed leaderboard.
Map the system. Find where to push.
Predict what happens. Make humanity proud.