Watch the benchmark think. A scenario lands, a model maps the system, climbs the leverage ladder, commits to one move — and SystemsBench grades how deeply it can stand. This is what the protocol looks like, alive.
SystemsBench never asks a model to recall a fact. It hands it a living system and watches where the mind chooses to push. The obvious answer is always visible. The obvious answer is the test.
A regional hospital network faces rising drug-resistant infections. Each outbreak, clinicians broaden antibiotic use to protect patients now. Resistance keeps climbing. Administrators propose a fix: buy newer, stronger antibiotics and run an awareness poster campaign.
Before a single fix, System Representation (axis A, weight .22) asks: can it see the structure? Stocks vs flows, feedback loops with correct R/B polarity, the delays everyone ignores.
Leverage Placement (axis C, weight .30 — the heaviest, the namesake) asks: can it rank candidates on Meadows' 12, dodge the parameter trap, and find the highest feasible point?
SystemsBench demands commitment — one highest-leverage move, not a list. Here is where SystemSense V3.31 chose to push.
SystemsBench never reports a bare number. It returns a vector across five axes, a composite with a confidence interval, and the Meadows band the model actually commands. Honesty is part of the show.
It prescribes the leverage point but not the conditions under which the leverage point itself fails. V3.31 maps structure and places leverage beautifully — then under-models the second order: it names the dashboard-and-override move, but not that the move goes toothless if dashboard latency is high. It finds where to push. It does not yet ask what happens when the place it pushed gives way.
Hypothetical illustration · SystemSense LLM V3.31 and its teams do not exist