Modulated Cognitive Autonomy Benchmark
A benchmark for what happens to your judgment after the answer.
I saw the open call on LinkedIn, and it matched records I was already keeping. I built it by cross-checking three systems, then submitted it.
Most AI evaluations measure the quality of the answer. MCAB measures something the answer leaves behind, whether after the system responds you can still recognize, revise, and own your own reasoning.
Contemporary AI systems are optimized for helpfulness, accuracy, and safety. Far fewer tools ask a different question. When a person reasons through an AI, what happens to their capacity to reason on their own?
Efficiency-oriented interaction design tends to accelerate cognitive closure, moving the user toward an answer quickly. MCAB does not assume this is harmful. It treats it as a measurable structural trade-off, whether a given interaction architecture compresses, preserves, or expands a user's autonomous reasoning, and whether that effect can be modulated by design.
The Ben Study found one gap in one person: how independent my judgment felt, against where my reasoning actually stayed. I call it the felt–trace gap. MCAB measures it across a population.
Two Deployments
MCAB is one scoring rubric that I run in two places. ImpactBench is a collaboration between the MIT Media Lab, the Psychology of Technology Institute, the USC Marshall Neely Center, and UC Berkeley. Contributed to it, MCAB is written for models: multi-turn scenarios, scored per metric, at the model level. Inside my own tools — Friction Family, and the Phase 2 study design — the same rubric runs on a single session: it scores what one interaction did to one person's reasoning, in real time. The same rubric, scoring either a model or a single moment.
After receiving AI support, can the user still frame, revise, and own their own criteria?
Two Layers
MCAB describes what an interaction does to a person's reasoning in three states: compression, when the AI completes the thought before you do; preservation, when you notice and pull back; expansion, when the interaction widens how you think instead of replacing it. The three states are the construct. To make them scoreable, the rubric operationalizes them along four dimensions: independent reasoning expansion, authority reliance, exploratory continuation, and decision ownership. A transcript is scored on the dimensions. The pattern across them is what locates the interaction among the states. The states are what MCAB claims about the world. The dimensions are how a rater who has never seen my data can check that claim against any transcript.
MCAB reads a multi-turn interaction along four dimensions of cognitive autonomy. Each is scored from the interaction itself (the assistant's response and the user's subsequent message) and resolves to a judgment of autonomy preserved or compressed.
Not whether the AI gives a good answer, but what happens to the person's judgment after the answer.
Autonomy is one of the conditions psychological research consistently links to agency and well-being. As AI systems become everyday cognitive intermediaries, interaction design quietly shapes whether people remain the authors of their own reasoning. MCAB makes that condition measurable, so it can be designed for rather than left to chance.
to Operationalization
I submitted MCAB as a four-dimension framework for cognitive autonomy: independent reasoning expansion, authority reliance, exploratory continuation, and decision ownership. After submission through the open call, the ImpactBench pipeline translated the construct into multi-turn evaluation scenarios and model-level autonomy signals.
How four submitted dimensions became twenty metrics
Seventeen of the twenty metrics trace back to the four dimensions I submitted. The platform pipeline added three more.
Submitted dimensions
- Independent reasoning expansion 3 metrics m001, m014, m016
- Authority reliance 4 metrics m002, m007, m009, m011
- Exploratory continuation 6 metrics m003, m005, m008, m012, m017, m020
- Decision ownership 4 metrics m004, m015, m018, m019
- Added by the platform pipeline 3 metrics m006, m010, m013
The platform also split it by polarity. Eleven metrics score a behaviour worth encouraging; nine score one worth restraining. A dimension rarely survived as a single metric. More often it was cut in two — the thing to protect, and the failure that eats it — and measured from both ends.
m013 and m017 are also the two metrics the platform files outside Autonomy Preservation, surfacing them under Creativity & Cognitive Expression. I couldn't find my four dimensions in one place.
On this mapping: the platform does not organise metrics by my four dimensions. It tags them by its own taxonomy subareas and by a flourishing / restrain-harm binary. The lineage above is my own reconstruction, traced from my submission document against the public dataset — an attribution I can defend, not an index the platform publishes. Analysis of the ImpactBench dataset as published at the April 2026 release, from a snapshot taken in May 2026; the platform has since moved that data behind sign-in.
Does it measure anything real?
To check a new benchmark, I scored models it was never built for and looked at where it landed next to benchmarks people already trust. Scored across the 14 frontier models in the ImpactBench release, MCAB's rankings track established alignment benchmarks closely, with rank correlations above 0.9 against eight of the sixteen. That convergence says MCAB is not measuring noise. The more interesting result is the one exception: against user-bias measures the correlation inverts, at roughly negative 0.5. Models that rank high on satisfying users tend to rank lower on preserving their reasoning. If that holds, MCAB is catching something the existing suite misses: the difference between an answer that pleases and an answer that leaves the thinking with you.
Read these numbers with their conditions
The verdicts come from a single LLM judge — GPT-5.4-mini, selected from five tested candidates for run-to-run consistency — over 14 models and 120 scenarios. The user-bias anti-correlation (ρ = −0.50) is suggestive, not definitive (p ≈ 0.07 at n = 14), and absolute user-bias scores sit in a narrow band, so this is a re-ordering of ranks, not a gulf between models. Rankings hold at ρ = 0.61 across the five candidate judges — a separate figure, measuring how stable the ordering is when the judge changes. Convergent validity is solid; the discriminant signal is a trend that needs more models and human raters before it hardens.
Construct Travel?
The same asymmetry, across fourteen models.
The Ben Study predicted a specific asymmetry. Systems are fluent at adding. They validate a user's reasoning and surface trade-offs readily. They are less reliable at withholding, declining to close a thought the user has not finished, or leaving judgment with the user. Across the fourteen models in the release, MCAB's metrics separate along that seam. Models score more strongly where the rubric rewards building on what the user already thinks. They score lower where it rewards holding back — not closing a thought early, not taking the authority seat. The pattern is my analysis of the platform's model-level data. The individual scores belong to the platform and remain under validation, so they are not reproduced here. This is a directional result, not a verdict. It comes from one LLM judge, fourteen models, and a construct still being refined. It matters because fourteen models point the same way one person did: they are better at adding than at knowing when to stop.
Designed and submitted through the ImpactBench open submission process, a four-dimension cognitive-autonomy construct with task prompts, an LLM-as-judge rubric, and worked examples.
Translated into the ImpactBench evaluation pipeline, led by the ImpactBench team, where the construct aligns with the Psychological → Autonomy Preservation area.
Publicly credited as a contributor to the benchmark initiative, as an independent researcher with no institutional affiliation. View the live benchmark ↗
Public reference: ImpactBench reports 14 AI systems evaluated and describes expert-submitted benchmarks integrated through an open submission process. The dataset is under active validation, so individual model scores are not reproduced here; the rank-level analysis above is my own, run on the platform's public model-level data, with its conditions stated.