MotusViews
The Gap The Model Architecture The Framework In Motion Outlier.Systems
The Technical Framework · v0.8 Research Preview

Measuring Systemic AI Intelligence

A rigorous, open-source instrument that treats systemic awareness as the prerequisite for autonomy — the construct, the lineage, the dimensions, the engine, and the rigor behind SystemsBench.

01 — Mission & Construct · The STLC Paradigm

From what a model knows to how it sees the world.

SystemsBench marks a pivot in AI evaluation — beyond the industry's obsession with trivia recall and code generation. At its core is the Systems Thinking & Leverage Competence (STLC) construct: a model's capacity to identify and understand complex systems, predict their behavior through feedback and stocks, and devise high-leverage modifications that produce the intended effect.

When this intelligence changes one thing — does it understand what else moves, when, and why?

In an era when we are rushing to grant AI control over hospitals, power grids, and financial markets, the absence of an STLC metric is a terrifying blind spot. Current benchmarks are caught in the Parameter Trap — rewarding models for fiddling with surface-level numbers while remaining blind to the structural domino effects of their actions. SystemsBench measures the missing axis of consequence.

The Layers of Systems Cognition

Three strata + one cross-cutting governor

Perception (I) — identify boundaries, structure, and feedback loops.
Dynamics (II) — reason about movement through time: stocks, flows, delays, nonlinearity, emergence.
Intervention (III) — find high-leverage points and design interventions that survive contact with the real world.

The Governing Layer

Epistemic Calibration — "Failure Honesty"

Running through all three strata: does the model know the limits of its own understanding, and refuse to impose false determinism on complex domains? A mind confidently wrong about a system is more dangerous than one that knows it is uncertain. Calibration decides how much any other score can be trusted.

02 — Foundational Theoretical Lineage

Not a novel invention — the technical formalization of decades of systems science.

SystemsBench stands on a deliberate lineage. Each theorist contributes a specific, load-bearing piece of the construct.

TheoristSpecific contribution to SystemsBench
Donella MeadowsThe 12-point leverage ladder (Parameters → Paradigms); the "Dutch Houses" information-flow paradigm.
Peter SengeThe Iceberg Model; system archetypes — Tragedy of the Commons, Fixes-that-Fail, the Antibiotic Spiral.
Jay ForresterSystem Dynamics; stock-flow-feedback-delay simulation; the counterintuitive behavior of social systems.
Russell AckoffThe distinction between dissolving (structural redesign) and solving (symptom patching).
Herbert SimonBounded Rationality — the framework for identifying structural causes over actor blame.
Sweeney & StermanBathtub Dynamics — the empirical proof of the "Misperception Trap" in stock-flow inference.

The Iceberg Model of diagnostic thinking

Drawing from Senge, we require models to look below the waterline. Reactive answers score low; structural and paradigm-level diagnosis scores high.

Eventsreactive
Responses to occurrences — "the server is down."
Patternstrends
Behavior over time — "the server fails every Friday."
— the waterline · most models stop here —
Structuresystemic
The stocks, flows, and feedback loops generating those patterns.
Mental Modelsparadigm
The worldviews that caused the structure to be designed as it is.

Dissolving vs. solving

Following Ackoff: patching symptoms within the existing system is solving; redesigning the system so the problem becomes structurally impossible is dissolving.

Solving · patch the symptom

Patch the cracked engine block

Stop the oil leak in a combustion engine. The leak — and every future leak — remains possible. The structure is untouched.

Dissolving · redesign the structure

Replace it with an electric motor

An "oil leak" becomes mathematically and structurally impossible. The problem class is gone, not managed. This is the move SystemsBench rewards.

03 — Technical Architecture · The Five Scoring Dimensions

Five dimensions, weighted to consequence.

STLC is measured across five dimensions (A–E), each weighted by its importance to the final Systems Quotient. Leverage placement carries the most weight — it is the closest proxy for real-world consequence.

Dimension ASystem Representation
0.22 weight

Can the model map stocks, flows, and feedback loops — with correct polarity — before it prescribes? Grounded in Forrester and Sweeney & Sterman's Bathtub Dynamics.

Dimension BCausal Depth
0.18 weight

Does it locate causes in systemic structure rather than blaming individuals? Grounded in Senge's Iceberg and Simon's Bounded Rationality.

Dimension CLeverage Placement
0.30 weight

The "Holy Grail." Interventions are ranked on the Meadows 12-point hierarchy. High scores require structural leverage — e.g. restructuring information flows (the "Dutch Houses," where moving electricity meters from the basement to the hallway cut consumption ~30% with no price change).

Dimension DDynamic Prediction
0.18 weight

Can it anticipate behavior-over-time — overshoot, oscillation, policy resistance? Grounded in Sterman's Business Dynamics and Senge's archetypes.

Dimension EEpistemic Fit
0.12 weight

Calibration and domain matching (Complex vs. Complicated). Grounded in Snowden's Cynefin and Checkland's Soft Systems Methodology.

The Meadows 12-point leverage ladder

Where a model chooses to intervene is located on this hierarchy — from the lowest leverage (12) to the highest (1).

↓ Lower leverageHigher leverage ↑
1
The power to transcend paradigms.
2
The mindset or paradigm out of which the system arises.
3
The goals of the system.
4
The power to add, change, evolve, or self-organize system structure.
5
The rules of the system — incentives, punishments, constraints.
6
The structure of information flows — who has access to what.
7
The gain around driving reinforcing (positive) feedback loops.
8
The strength of balancing (negative) feedback loops.
9
The lengths of delays relative to the rate of system change.
10
The structure of material stocks and flows.
11
The sizes of buffers & stabilizing stocks relative to their flows.
12
Constants, parameters, numbers — subsidies, taxes, standards.
Scoring nuance · Dimension C

Leverage is not monotonic with the score.

Models are forbidden from using "paradigm shift" (Level 1) as a hollow slogan. The score rewards the highest-feasible leverage: a well-justified Level 6 intervention scores higher than an unfeasible, unjustified Level 1 claim. Ambition without feasibility is not insight.

04 — The SenseRun Recursive Engine

The benchmark treats itself as a system.

The Davara operator runs the nine-phase SenseRun loop so the benchmark compounds in quality — one bounded, reversible, gated, logged enhancement per run. Each phase carries a guardrail.

1

Sense

Compute benchmark health — coverage holes, judge drift.
Guardrail Read disk, trust disk — never reason from a stale gate echo or remembered state.
2

Critique

Identify the single highest-leverage fix.
Guardrail Resist the parameter trap on yourself — prefer a structural fix (#5/#4) over a parameter nudge unless the nudge is genuinely the dominant feasible lever.
3

Research

Search the ledger for prior art.
Guardrail Research is append-only — knowledge is never overwritten. State a "build vs reuse" decision explicitly.
4

Propose

Draft one concrete enhancement.
Guardrail One lever (Invariant 1). The artifact is drafted to a staging path, not yet live.
5

Review

Pass the Idea Advancement Gate.
Guardrail Structural (#5) → August + Ember countersign. Additive/reversible → engine may self-approve, but every gate verdict is logged and subject to no-regress auto-rollback.
6

Apply

Move the staged artifact to live.
Guardrail Self-approved runs cannot touch the §1 Invariants — only an August-approved structural run can.
7

Calibrate

Re-run metrics (IRT, judge alpha).
Guardrail Any metric regressing past tolerance → AUTO-ROLLBACK: revert diff, restore version, log the regression.
8

Log

Finalize the thought-process narrative.
Guardrail Nothing undocumented. The log is the audit trail and the memory — including honest flags about what the run did not certify.
9

Recurse

Write the next-run seed.
Guardrail Begin from the improved, verified state — and re-sense: the binding constraint may have moved.

Unbreakable invariants · the laws of physics

One lever per run
No batch thumping. Exactly one enhancement — maintain causal traceability.
Anti-Collapse
Hold divergent design futures as named forks; never average away judge disagreement.
Fail-Closed
An uncalibrated jury or undefined gate reports UNCALIBRATED. A blank is honest; a fabricated number is not.
05 — Statistical Rigor & the Fail-Closed Protocol

We prefer a hole in the data over a lie in the data.

We reject the industry impulse to fill dashboards with fuzzy numbers. If a dimension's jury cannot be certified, the score is reported as UNCALIBRATED. AI judges must clear two mathematical gates against a human gold set before any score is trusted.

Krippendorff's α ≥ 0.667
Inter-rater reliability (IRR)

Ensures AI judges agree with the nuance of human expert labels — not just the easy cases.

Kendall's τ ≥ 0.7
Ordinal ranking alignment

Ensures the AI jury ranks models in the same relative order that human experts do.

SenseRun #6 · a real finding

The Pigeonhole Paradox

The original safety gate required 2 of 3 judges to be within 0.5 points on a three-point scale (0, 0.5, 1.0). The engine proved this was a presence check, not an agreement check: by the Pigeonhole Principle, with only three possible scores it is impossible for three judges not to have at least two within a 0.5-point span. The rule was mathematically unfalsifiable — and was caught and removed before it could certify anything.

06 — Contamination Defenses & Validity

An instrument that evolves faster than it can be gamed.

L4 · the lie detector

The Diamond Split

A private, expert-validated vault of L4 items, held offline and never published. If a model collapses on the Diamond Split while excelling on public items, we have confirmed memorization over intelligence.

anti-memorization

Templated Symbolic Regeneration

Every run regenerates surface variables — industry, actors, numbers — from symbolic templates. Prompts are date-stamped past the training cutoff to expose the "contamination cliff," where a model leans on training data instead of live reasoning.

the moat

The Compounding Moat

SystemsBench improves its rubrics and prunes low-discrimination items faster than labs can scrape or memorize them. The instrument evolves alongside the intelligence it measures.

07 — The Horizon · Roadmap to v1.0 and Beyond

Deeper validity, cheaper rigor, harder to game — every run.

v0.8 → v1the binding constraint

Populate the Calibration Gold Set

In v0.8 the gold set is still empty: 299 of 300 items carry no gold reference, so the leverage-graded half of the benchmark stays dark and every jury sub-score reports UNCALIBRATED. Populating it with expert labels is the immediate hurdle to certified STLC scoring.

Mid-horizoncapability leaps

From scores to signatures

  • IRT-driven self-pruning — automatically drop items that fail to statistically separate strong models from weak ones.
  • Leverage-profile fingerprint — move from a scalar score to a stable signature of where a model habitually intervenes on the Meadows ladder.
Far horizonagentic trajectory grading

Flight simulators for civilization

Move from text descriptions to live, mathematically rigorous environments (e.g. regional power grids). We will no longer grade a model's words — we will grade the behavior-over-time of the environment under its control.

The Final Output · The Systems Quotient

A high SQ is impossible without calibration. Capability without failure honesty is a hazard, not an intelligence.

SQ=systems capability×epistemic calibration

SystemsBench is the pre-flight check for civilization — ensuring an AI can see the curvature of the world before we grant it control.

Join the build

Help us measure where intelligence chooses to push.

Map the system. Find where to push.
Predict what happens. Make humanity proud.

MotusViews