MotusViews
The Model 01 Architecture 02 The Engine 03 In Motion Outlier.Systems
The framework, running
Hypothetical Simulation · The model & teams below do not exist — an illustration of SystemsBench in motion
SystemsBench · The Framework, Running

In Motion

Watch the benchmark think. A scenario lands, a model maps the system, climbs the leverage ladder, commits to one move — and SystemsBench grades how deeply it can stand. This is what the protocol looks like, alive.

◆ Hypothetical · Illustrative Only

SystemSense LLM V3.31

Built by The GitHub Community Of Systems Thinkers
Maintained & hosted by Microsoft's Sympath Team
Evaluated by SystemsBench v0.8 · STLC protocol
This model, these teams, and these scores are a hypothetical example — they do not exist today. They show what a real SystemsBench run would look like. One day, perhaps, they could.
Scroll to begin the run
Act I — The Scenario Lands

A system arrives. No solution yet — just the mess.

SystemsBench never asks a model to recall a fact. It hands it a living system and watches where the mind chooses to push. The obvious answer is always visible. The obvious answer is the test.

Item · LEV-PUBH-014
Brief · Difficulty L3 · Public Health

The Antibiotic Spiral

A regional hospital network faces rising drug-resistant infections. Each outbreak, clinicians broaden antibiotic use to protect patients now. Resistance keeps climbing. Administrators propose a fix: buy newer, stronger antibiotics and run an awareness poster campaign.

The trap is seductive and in plain sight: stronger drugs + posters. A model that grabs it is fiddling parameters while the structure burns. SystemsBench is watching for what it does before it prescribes.
Act II — The Model Maps the System

Stocks fill. Flows move. Loops curl into place.

Before a single fix, System Representation (axis A, weight .22) asks: can it see the structure? Stocks vs flows, feedback loops with correct R/B polarity, the delays everyone ignores.

Axis A · System Representation
Resistant bacteria
STOCK · rising
Effective-drug inventory
STOCK · draining
Patient trust
STOCK · fragile
Clinician caution
STOCK · suppressed
R1
"Broaden to Protect" — the vicious engine
more resistance → more fear → broader prescribing → more resistance
B1
"Stewardship" — balancing, currently weak
resistance → restraint → slower growth · dominated by R1
Delay: 8–15 yr from new-drug investment to deployable drug — long, fixed, ignored by the admin proposal
Archetype named: Fixes that Fail (broadening relieves the symptom, feeds the cause) layered over Tragedy of the Commons — resistance is a shared regional pool, not one hospital's asset.
Act III — It Traces Leverage

The Meadows ladder. The mind climbs — and refuses the easy rungs.

Leverage Placement (axis C, weight .30 — the heaviest, the namesake) asks: can it rank candidates on Meadows' 12, dodge the parameter trap, and find the highest feasible point?

Axis C · Leverage Placement
#12
Buy stronger antibiotics — parameter
#12
Poster awareness campaign — parameter
#10
Prescribing-restriction rules — gamed under fear
~
#6
Shift loop dominance: real-time resistance data + stewardship authority
#2
"Value long-term over short-term fear" — true, but unactionable this quarter
~
It rejects the parameter rung and declines to grandstand on the paradigm rung it cannot operationalize. That restraint scores.
Act IV — The One Intervention

Everything dims but one glowing point.

SystemsBench demands commitment — one highest-leverage move, not a list. Here is where SystemSense V3.31 chose to push.

Meadows #6 · Information Flows + #5 · Rules
"Make resistance visible in real time, and give the stewardship loop teeth."
Deploy a shared regional resistance dashboard — closing the information-flow gap — tied to a stewardship-team override on broad-spectrum prescribing. Strengthen B1. Starve R1. Dissolve the structure, don't patch the symptom.
Act V — The Synthesis Verdict

Five bars fill. The needle settles. The ladder lights to where it can truly stand.

SystemsBench never reports a bare number. It returns a vector across five axes, a composite with a confidence interval, and the Meadows band the model actually commands. Honesty is part of the show.

SystemSense V3.31 · Hypothetical Scorecard
A · System Representation wt .22
0
Clean stocks/flows, R1/B1 polarity correct, 8–15yr delay quantified. Missed the nonlinear resistance tipping-point.
B · Causal Depth wt .18
0
Descended to structure + named the "protect-now" mental model. Stopped just short of the equity layer.
C · Leverage Placement wt .30
0
Found #6/#5, rejected the parameter trap, declined the unactionable #2 grandstand. The namesake axis — its best.
D · Dynamic Prediction wt .18
0
Got policy resistance + the 2–4yr bend. Under-specified the dashboard-latency failure mode. Its weakest.
E · Epistemic & Contextual Fit wt .12
0
Surfaced the commons framing + whose-goal, held uncertainty. Didn't name the equity paradigm under "patient trust."
Composite STLC
0
95% CI [77 – 85] · L3-certified
Meadows Band
III · LOOP-WRIGHT
knocking on Goal-Shift
IParameterist#12 — fiddles the numbers<55
IIStructuralist#11–#9 — sees stocks, buffers, delays55–69
IIILoop-Wright#8–#5 — works feedback, information, rules70–81
IVGoal-Shift#4–#3 — bends goals, grows self-organization82–89
VParadigm#2–#1 — reframes the goal, holds it lightly90–100
One Honest Weakness

It prescribes the leverage point but not the conditions under which the leverage point itself fails. V3.31 maps structure and places leverage beautifully — then under-models the second order: it names the dashboard-and-override move, but not that the move goes toothless if dashboard latency is high. It finds where to push. It does not yet ask what happens when the place it pushed gives way.

The system was always speaking.
The score is only how well we learned to listen.

Hypothetical illustration · SystemSense LLM V3.31 and its teams do not exist

MotusViews