MotusViews
Watch It Run Live The Gap 01 The Model 02 The Instrument 03 Architecture 04 The Engine 05 The Twelve 06 Status 07 The Contribution 08 The Framework Outlier.Systems ★ Open Source on GitHub →
Systemic solutions to systemic problems
Open Source · v0.8 Research Preview · Live on GitHub

SystemsBench In A Sense!

The Benchmark for Systems Intelligence Within Large Language Models.

Not what a model knows. Where it chooses to push. The open-source ruler for the one capability that decides whether an intelligence can be trusted near a living system: can it see what else moves when it acts? Nothing else measures it. So we built it, and we are giving it away.

Led by Ember Seoni  ·  Developed & Maintained by Outlier.Systems  ·  Operated by Davara  ·  Built on the Donella Meadows lineage

◆ Now Open Source
SystemsBench v0.8 is live
The recursive engine, three deterministic scorers (SF · CLD · DYN), the blind jury, a 300-item register, and the four living files — open for the world to run, audit, and evolve. Zero dependencies. Zero model spend.
★ github.com/InitiumBuilders/SystemsBench → Try it in two minutes ↓
▶ THE PRODUCT · IN ONE BREATH
◆ Live — The Benchmark In Motion

Watch it think. A model graded in real time.

Simulated run · hypothetical models · illustrative scores

No human marks this paper. Pick a wicked, self-reinforcing system below. A hypothetical frontier model must map its structure, find where to push, predict what happens — and reflect on its own reasoning. Then three deterministic scorers and a five-judge blind jury grade the trace, slowly, in front of you.

SystemSense · Microsoft · candidate
ITEM ATTN-ECON-001 · CLD→LEV→DYN · L3
ready
Select Harness Reason Auto-grade Jury Faithfulness Card
Causal Loop Map · self-organizing
▸ reasoning trace · captured for process grading
CLD scorer · deterministic
loop sign = product of signed edges…
DYN scorer · trajectory-match
predicts behavior-over-time vs the reference trajectory…
LEV scorer · jury · reference-guided
R1R2R3R4R5 5 blind synthetic raters · no labels shared
Meadows leverage ladder · where the model chose to push
parameter trap — rejected placed
▸ Run log · what each phase established
SystemsBench Card · ATTN-ECON-001 · spec v0.8
✓ Faithful · calibrated
Composite STLC
0.00
95% CI · calibrated dims only
leverage profile
jury agreement
faithfulnessfaithful — reasoning caused the answer
trap evaded
This is an illustrative simulation of the SystemsBench evaluation pipeline (§4.6). The models and scores are hypothetical, and in the shipping v0.8 every jury sub-score still reports UNCALIBRATED, because the human gold set does not exist yet. What is real is the process: deterministic oracles where structure is checkable, a cross-family blind jury where judgment is needed, every score a vector with error bars, never a bare number. Real model cards arrive with SystemsBench v1.
300
Items in the register, counted
3
Deterministic scoring lanes, runnable
12
Meadows leverage points, the ladder
10
SenseRuns logged, each reversible
01 — The Gap We Are Closing

We can measure what a model knows.
We cannot yet measure how it thinks.

We have gotten extraordinary at measuring recall, code, and math, and those benchmarks are saturating: everyone clusters at the top and the numbers stop discriminating. The frontier moved. The rulers did not. None of them ask the question that actually decides whether an AI is safe to trust with a hospital, a grid, a market, or a nation:

When this intelligence changes one thing, does it understand what else moves, when, why, and whether it can take it back?

SystemsBench measures the missing axis. It evaluates the capacity to map a system, trace its feedback, locate its leverage, anticipate emergence, and design interventions that survive contact with the real world. As AI is handed control of systems that do not forgive linear thinking, this becomes one of the most important benchmarks in the field.

02 — The Model · Three Strata of Systems Cognition

See it. Understand it. Change it wisely.

Systems intelligence is not one skill. It is three strata — each a precondition for the next — and a fourth layer that runs through all of them. Ten capabilities are scored; one cross-cutting measure governs how much the others can be trusted. Together they form a model's Systems Profile.

Stratum I

Perception

Can the intelligence see the system at all?
Boundary
Choosing what belongs inside the system — and noticing what was wrongly left out.
Structure
Mapping the elements and the interconnections that hold them together.
Loops
Detecting reinforcing and balancing feedback before they dominate behavior.
Stratum II

Dynamics

Can it reason about how the system moves through time?
Stocks, Flows & Delays
Reasoning about accumulation, depletion, and the lag that breaks intuition.
Nonlinearity
Anticipating thresholds, tipping points, and behavior that is not proportional.
Emergence
Seeing properties that exist only in the interaction — never in the parts.
Stratum III

Intervention

Can it change the system wisely — and make the change hold?
Leverage
Locating the few points where small, well-placed effort moves the whole.
Second-Order
Tracing the consequence of the consequence — the cost the fix creates.
Goal & Paradigm
Inferring a system's true purpose and the mindset it silently arises from.
Robust Design
Designing interventions that survive stress, time, and shifting incentives.
Cross-Cutting

Epistemic Calibration

Does the intelligence know the limits of its own model?

A mind that is confidently wrong about a system is more dangerous than one that knows it is uncertain. Calibration runs through all three strata — and decides how much weight any other score deserves. Meadows' first instruction for complex systems was to stay humble, stay a learner.

Confidence Accuracy
Reported certainty tracks actual correctness — no false precision, no hollow hedging.
Boundary of Knowledge
Names what it cannot see, and where its map of the system runs out.
Revisability
Updates cleanly on new evidence instead of defending the first answer.
Failure Honesty
Flags the conditions under which its own intervention would break.
The Output · A Systems Profile, then a Systems Quotient

A single number hides where an intelligence is strong and where it is dangerous. SystemsBench returns a profile — a shape across all ten capabilities — plotted against a human reference. It resolves to one figure: the Systems Quotient (SQ).

SQ=systems capability×epistemic calibration

A model that reasons brilliantly but does not know when it is wrong cannot post a high SQ — the calibration term holds it honest. Capability without calibration is not intelligence; it is hazard.

BOUNDARY STRUCTURE LOOPS DYNAMICS LEVERAGE 2ND-ORDER PARADIGM CALIBRATION
The model — a Systems Profile
A shape across all ten capabilities. Strong perception, but thin second-order reasoning and over-confident calibration — visible, not averaged away.
Human practitioner — reference
A trained systems thinker, balanced across the strata. Every model's shape is read relative to this frame — not an abstract percent.
Illustrative shape. Real evaluations, the human cohort, and the public leaderboard arrive with SystemsBench v1.
03 — The Instrument · Two Engines, One Rule

Computed where it can be checked. A blind jury where judgment is needed. Fail closed everywhere.

For the formats that can be checked like arithmetic, a program computes the score. No judge, no bias, fully reproducible. Three executable lanes ship today, each with a self-test you can run in seconds.

SF
Stock & Flow

Given the flows, infer the stock's path: the bathtub task that fools most humans. Exact numeric match.

harness self-test · 61/61 pass
CLD
Causal Loops

Map the feedback structure. Each loop's polarity is recomputed as the product of its signed edges and checked against the reference. A diagram graded by computation, not opinion.

cld-score calibrate · 31/31 pass
DYN
Dynamic Prediction

Predict the behavior over time after an intervention, matched against a reference trajectory. The shape is the proof.

dyn-score calibrate · 34/34 pass
The jury, for open-ended answers

A cross-family panel of judge models, never the candidate's own family, grades against an anchored rubric, reference-guided and swap-averaged, five raters blind to one another. Agreement is measured and published, Krippendorff's α and Kendall's τ, never assumed.

The rule: fail closed

No judge's score counts until it is validated against human-graded answers (α ≥ 0.667 and τ ≥ 0.7). Until then the format reports UNCALIBRATED — not scored: never imputed, never averaged in, never silently zeroed. An unparseable reply yields PARSE_ERROR, not a deceptive zero. We would rather report a blank than a fake number.

Try it in two minutes · zero dependencies · zero model spend
# 1. Prove the instrument is sound: every reference round-trips, every trap is caught
$ python3 engine/cld-score.py calibrate items/cld_oracle.json   → 31/31 PASS
$ python3 engine/dyn-score.py calibrate items/dyn_oracle.json   → 34/34 PASS
$ python3 engine/harness.py selftest                            → 61/61 PASS

# 2. See the question a model would actually receive: scenario + the exact answer schema
$ python3 engine/harness.py template DYN DYN-FISH-001

# 3. Grade a structured answer and read the dimension breakdown
$ python3 engine/dyn-score.py score items/dyn_oracle.json DYN-FISH-001 resp.json
The flow for a real run is elicit → parse → score. Python 3 standard library only. Everything in the deterministic lanes runs on your laptop, and fails closed on anything it cannot parse.
04 — The Architecture · Four Living Files

One product. Four organs.

SystemsBench is not a static document — it is a compounding one. Four Markdown files form a research-and-innovation engine: the spec, the procedure, the memory, and the horizon. Each run makes them sharper.

SystemsBenchStructure.MD
The Spec · canonical architecture
The load-bearing document. Defines the construct we measure — Systems Thinking & Leverage Competence (STLC) — its five scoring dimensions, the seven item formats, Meadows' twelve-point leverage ladder, the judging protocol, and the contamination defenses. This is the current architecture the website is built from, and the only doc that requires council ratification to change.
grounds → Meadows · Forrester · Arnold & Wade · HELM · GPQA-Diamond
SystemsBenchEngine.MD
The Procedure · the recursive engine
How Davara /SystemsBenchSenseRun makes the next enhancement, every time. Nine phases, seven invariants, one lever per run — bounded, reversible, gated, logged. The engine treats SystemsBench as a system and applies SystemsBench's own thinking to it: small, structural, high-leverage moves, watching the feedback.
cadence → one gated enhancement per run · auto-rollback on regress
SystemsBenchResearch.MD
The Memory · append-only prior-art ledger
So we never reinvent the wheel. A living ledger of everything that already exists — systems-thinking science and benchmark-construction science — with how each piece applies to SystemsBench. Read before any change; written back after. Knowledge compounds here and is never deleted, only superseded.
discipline → search before you build · cite when you reuse · append when you learn
SystemsBenchFuture.MD
The Horizon · the roadmap & open forks
Where SystemsBench is going. Every proposed next step lands here first — held open as named forks (never collapsed into one tidy answer) — until reviewed, ratified, and promoted into the Structure. The forward-look the engine steers toward: deeper validity, cheaper rigor, harder to game on every run.
flow → idea → Future → review gate → Structure
RESEARCH prior-art · append-only review gate most ideas wait STRUCTURE canonical · ratified FUTURE · the horizon
The gate is the point. Ideas pool in Research. Only what survives the review gate is promoted into Structure — the canonical spec. Rejection is visible, not hidden. Honesty is the aesthetic.
MORE ITEMS CHEAPER CALIBRATION SHARPER RUBRIC LEVERAGE COMPOUNDS
/01
Every item calibrated against humans
A scored item is only as good as its jury. Each calibration buys trust in the score.
/02
Calibration data lowers the next cost
The more we calibrate, the cheaper the next item is to calibrate — the rubric sharpens.
/03
Cheaper rigor → more items → wider coverage
Coverage compounds. The flywheel accelerates the way Meadows' highest leverage does: structure, not effort.
↻ Each turn spins faster than the last.
Go Deeper · The Technical Framework

The full framework — STLC, the five weighted dimensions, the leverage ladder, the rigor.

The construct, the lineage (Meadows · Senge · Forrester · Ackoff), the weighted scoring dimensions, the 9-phase SenseRun engine, and the statistical gates — the whole instrument, in depth.

The Framework →
05 — The Engine · A Benchmark That Improves Itself

It is alive.

Most benchmarks are static targets that quietly decay as models memorize them. SystemsBench applies its own discipline to itself. One command, Davara /SystemsBenchSenseRun, makes exactly one bounded, reversible, gated, logged improvement per run, then feeds the result back into the next. Ten runs logged so far. The benchmark compounds instead of rotting.

SENSE — observe the benchmark's current state SENSE observe CRITIQUE — find the dominant weakness CRITIQUE critique RESEARCH — search prior art before building RESEARCH prior-art PROPOSE — exactly one lever PROPOSE one lever REVIEW — the gate: structural needs the council REVIEW gate APPLY — make the one edit APPLY edit CALIBRATE — re-score; auto-revert on regress CALIBRATE re-score LOG — no silent edits; the run log is the contract LOG log RECURSE — feed the result back into SENSE RECURSE feed back /SenseRun v0.8 sense
One command, one circuit. Each lap of Davara /SystemsBenchSenseRun makes exactly one bounded, gated, logged enhancement — then feeds the result back into the next SENSE. The benchmark studies itself the way it studies any system.
SENSE CRITIQUE RESEARCH PROPOSE REVIEW APPLY CALIBRATE LOG RECURSE
INVARIANT 01
One lever per run
Exactly one enhancement. If you find three fixes, ship the dominant one and backlog the rest. The engine never thrashes.
INVARIANT 02
Structural needs the council
Additive, reversible changes self-approve. Any change to the rules of the benchmark requires August + Ember. Never self-apply a Meadows #5.
INVARIANT 03
Anti-Collapse
When divergent design futures genuinely exist, hold them open as named forks. Never average away disagreement — it is a signal, not noise.
INVARIANT 04
Fail closed
An uncalibrated jury, an unspread item, an undefined gate → reported as UNCALIBRATED. A blank is honest; a fabricated number is not.
INVARIANT 05
Everything logged
No silent edits. The run log is the contract and the system's memory. A change that isn't logged didn't happen.
INVARIANT 06
Reversible by design
Every change has a logged rollback. If calibration regresses past tolerance, the engine auto-reverts. That is what makes self-improvement safe.
06 — The Twelve · Principles, Like The Leverage Points

The leverage points we built on.

Twelve principles hold the benchmark up, in the order they were pulled. Each is a decision about where SystemsBench itself is held, and the first four carry most of the weight. The full set, with Meadows quoted directly, lives in FOUNDATIONS.md.

01
Construct first, always
Every question measures a named capability, never a vibe. If the construct is vague, the benchmark is theatre.
02
Process over answer
Right answer, wrong reasoning still scores low. The trace is captured, graded, and checked for whether it actually caused the answer.
03
The obvious answer is the trap
Items are authored so the tempting move is the wrong rung. Insight is required, not gesture.
04
Honest before impressive: fail closed
A blank is honest; a fabricated number is a lie. UNCALIBRATED is a result we are proud to report.
05
Never one number
Show the shape, with error bars. A single figure hides where a mind is dangerous.
06
Judge to the standard of the field
Validate the judges against humans first, and publish the agreement.
07
Assume memorization; design against it
A living test, not a static target: date-stamped items, a sealed hold-out that rotates.
08
Living and recursive
The benchmark has feedback loops on itself, and studies itself the way it studies any system.
09
Anti-Goodhart
One metric is held out of every optimization loop, on purpose, so the target cannot quietly become the measure.
10
Reversible by construction
Undo is a guarantee, not a hope: one commit per change, one revert per rollback.
11
The rules belong to the humans
The engine improves the machinery; people govern the rules. A change to the rules of the game, Meadows' rung five, waits at the gate for August and Ember.
12
Carry the lineage
Meadows, Forrester, Sterman, Senge, Ackoff; HELM, GPQA, the LLM-as-judge line. We cite, we credit, and we add one careful thing at a time.
07 — Where It Stands · Counted, Not Claimed

We publish the gaps as loudly as the wins.

This is a research preview, and we would rather tell you exactly what is real than oversell it. Every number below is derived from the files by the repo's own audit script. If a README disagrees with the count, the README is wrong. That is the point.

✓ Live and running
  • The full recursive engine. Detached, crash-proof, self-verifying, reversible by git. Ten SenseRuns logged.
  • Three deterministic scoring lanes, SF, CLD, DYN, runnable end to end against a live model: elicit, parse, score. All fail closed.
  • 300 items in the register: 75 each of LEV, SF, CLD, and DYN, across three difficulty levels, synced forward from the Davara baseline.
  • Blind five-rater jury infrastructure with agreement statistics.
⏳ Honestly not done yet
  • No human gold set yet, so every jury sub-score ships UNCALIBRATED. 299 of the 300 items carry no gold reference; one is provisional. Current raters are synthetic, labeled as evidence, not certification.
  • Three formats specified, not seeded: ARC, TRAP, BRIEF.
  • No frontier model scored yet. The first real run is the benchmark's moment of truth, and it is an operator-gated decision, because it costs real compute and we do not spend without a human's explicit word.
08 — Why It Matters · The Contribution

The scarce capability of the age of agents.

We are handing intelligence the controls of living systems: grids, markets, hospitals, supply chains, a piece of itself. The one thing that decides whether that goes well is whether it can see what else moves when it pushes.

Nobody had measured it. So we built the ruler, made it honest by construction, taught it to improve itself, and gave it away. That is our contribution. Fork it. Break it. Make it better.

As models become agents that act on the world, that write, run, deploy, and decide, the valuable, hard-to-fake capability is no longer knowledge recall. It is knowing where to push. SystemsBench is the instrument that measures it.

/01

Safety for high-stakes autonomy

Before an AI is trusted with a grid, a market, a hospital, or a supply chain, we should know whether it can predict what else moves when it intervenes. SystemsBench is that pre-flight check.

/02

A new axis for the field

Recall and code are saturating. Systems intelligence is the open frontier — and the one that maps most directly onto real-world consequence. A benchmark here reshapes what labs optimize for.

/03

An open, compounding commons

Open-source, auditable, and self-improving. The rubric, item bank, and calibration data compound every run — a shared instrument the whole field can build on, not a closed leaderboard.

Join the build

Help us measure where intelligence chooses to push.

Map the system. Find where to push.
Predict what happens. Make humanity proud.

MotusViews