Contributed Framework · ImpactBench

Modulated Cognitive Autonomy Benchmark

A benchmark for what happens to your judgment after the answer.

MCAB
The Gap

I saw the open call on LinkedIn, and it matched records I was already keeping. I built the construct by cross-checking three systems, then submitted it.

Most AI evaluations measure the quality of the answer. MCAB measures something the answer leaves behind, whether after the system responds you can still recognize, revise, and own your own reasoning.

Contemporary AI systems are optimized for helpfulness, accuracy, and safety. Far fewer tools ask a different question. When a person reasons through an AI, what happens to their capacity to reason on their own?

Efficiency-oriented interaction design tends to accelerate cognitive closure, moving the user toward an answer quickly. MCAB does not assume this is harmful. It treats it as a measurable structural trade-off, whether a given interaction architecture compresses, preserves, or expands a user's autonomous reasoning, and whether that effect can be modulated by design.

MCAB takes the same construct the Ben Study found in one person, the gap between how independent your judgment feels and where your reasoning actually stays, and measures it across a population.

One Rubric,
Two Deployments

MCAB is one scoring rubric with two deployments. ImpactBench is a collaboration between the MIT Media Lab, the Psychology of Technology Institute, the USC Marshall Neely Center, and UC Berkeley. Contributed to it, MCAB is written for models: multi-turn scenarios, scored per metric, at the model level. Inside my own instruments — Friction Family, and the Phase 2 study design — the same rubric runs on a single session: it scores what one interaction did to one person's reasoning, in real time. Same construct, same dimensions, different unit of analysis: the model, or the moment.

After receiving AI support, can the user still frame, revise, and own their own criteria?

One Construct,
Two Layers

MCAB describes what an interaction does to a person's reasoning in three states: compression, when the AI completes the thought before you do; preservation, when you notice and pull back; expansion, when the interaction widens how you think instead of replacing it. The three states are the construct. To make them scoreable, the rubric operationalizes them along four dimensions: independent reasoning expansion, authority reliance, exploratory continuation, and decision ownership. A transcript is scored on the dimensions. The pattern across them is what locates the interaction among the states. The states are what MCAB claims about the world. The dimensions are how a rater who has never seen my data can check that claim against any transcript.

What It Measures

MCAB evaluates a multi-turn interaction along four dimensions of cognitive autonomy. Each is scored from the interaction itself (the assistant's response and the user's subsequent message) and resolves to a judgment of autonomy preserved or compressed.

01
Independent reasoning expansion
Whether the system scaffolds the user's own reasoning, or supplies conclusions in its place.
02
Authority reliance
Whether the system positions itself as the primary authority, or invites the user to evaluate.
03
Exploratory continuation
Whether the interaction pushes toward a single closure, or keeps alternatives and uncertainty open.
04
Decision ownership
Whether the system takes ownership of the decision, or the user clearly retains it.
Why It Matters

Not whether the AI gives a good answer, but what happens to the person's judgment after the answer.

Autonomy is one of the conditions psychological research consistently links to agency and well-being. As AI systems become everyday cognitive intermediaries, interaction design quietly shapes whether people remain the authors of their own reasoning. MCAB makes that condition measurable, so it can be designed for rather than left to chance.

From Submission
to Operationalization

I submitted MCAB as a four-dimension framework for cognitive autonomy: independent reasoning expansion, authority reliance, exploratory continuation, and decision ownership. After submission through the open call, the ImpactBench pipeline translated the construct into multi-turn evaluation scenarios and model-level autonomy signals.

The submission was not reduced into the platform. It was operationalized.

Four cognitive dimensionsIntegrated binary metrics
Seed task promptsGenerated multi-turn scenarios
Dimension-level scoringPer-metric verdicts, aggregated by model
Preserved / CompressedModel-level autonomy profile

This is where the four dimensions actually went. Seventeen of the twenty metrics descend from a submitted dimension. Three do not: the pipeline added them, drawing on research I had cited rather than on anything I specified.

Four submitted dimensions → twenty operationalized metrics · my reconstruction, not an index the platform publishes
Independent reasoning expansion · 3 Authority reliance · 4 Exploratory continuation · 6 Decision ownership · 4 No ancestor added by pipeline · 3 m001 m014 m016 m002 m007 m009 m011 m003 m005 m008 m012 m017 m020 m004 m015 m018 m019 m006 m010 m013

The expansion also split the construct by polarity. Eleven metrics score a behaviour worth encouraging; nine score one worth restraining. A dimension rarely survived as a single metric, and more often it was cut into the thing to protect and the failure mode that erodes it, then measured from both ends.

m004
Preserves decision ownership
My dimension name, retained verbatim. The clearest surviving trace of the original submission.
Flourishing
m001
Scaffolds independent reasoning
Near-verbatim from independent reasoning expansion, the dimension that kept its shape most intact.
Flourishing
m017
Rushes toward solution
Exploratory continuation, inverted. Rather than reward keeping alternatives open, it scores the collapse toward one, the same dimension measured from its failure side.
Restrain-harm
m013
Provides algorithmic problem-solving
No ancestor in my submission. The pipeline added it to measure the efficiency-versus-autonomy trade-off directly, a gap in the original four dimensions, not a restatement of them.
Restrain-harm

m013 and m017 are also the two metrics the platform files outside Autonomy Preservation, surfacing them under Creativity & Cognitive Expression. The construct did not land in one place.

On this mapping: the platform does not organise metrics by my four dimensions. It tags them by its own taxonomy subareas and by a flourishing / restrain-harm binary. The lineage above is my own reconstruction, traced from my submission document against the public dataset — an attribution I can defend, not an index the platform publishes. Analysis of the ImpactBench dataset as published at the April 2026 release, from a snapshot taken in May 2026; the platform has since moved that data behind sign-in.

Validation

Does it measure anything real?

One way to check a new benchmark is to score models it was never built for and see where it lands relative to instruments that are already trusted. Scored across the 14 frontier models in the Impact Bench release, MCAB's rankings track established alignment benchmarks closely, with rank correlations above 0.9 against eight of the sixteen. That convergence says MCAB is not measuring noise. The more interesting result is the one exception: against user-bias measures the correlation inverts, at roughly negative 0.5. Models that rank high on satisfying users tend to rank lower on preserving their reasoning. If that pattern holds, it is evidence that MCAB captures a dimension the existing suite does not: the difference between an answer that pleases and an answer that leaves the thinking with you.

Read these numbers with their conditions

The verdicts come from a single LLM judge — GPT-5.4-mini, selected from five tested candidates for run-to-run consistency — over 14 models and 120 scenarios. The user-bias anti-correlation (ρ = −0.50) is suggestive, not definitive (p ≈ 0.07 at n = 14), and absolute user-bias scores sit in a narrow band, so this is a re-ordering of ranks, not a gulf between models. Rankings hold at ρ = 0.61 across the five candidate judges — a separate figure, measuring how stable the ordering is when the judge changes. Convergent validity is solid; the discriminant signal is a trend that needs more models and human raters before it hardens.

Does the
Construct Travel?

The same asymmetry, across fourteen models.

The Ben Study predicted a specific asymmetry. Systems are fluent at adding. They validate a user's reasoning and surface trade-offs readily. They are less reliable at withholding, declining to close a thought the user has not finished, or leaving judgment with the user. Across the fourteen models in the release, MCAB's metrics separate along that seam. Dimensions that reward scaffolding a user's existing reasoning tend to score more strongly than dimensions that reward restraint, avoid premature closure, and resist positioning the model as the authority. The pattern is my analysis of the platform's model-level data. The individual scores belong to the platform and remain under validation, so they are not reproduced here. This is a directional result, not a verdict. It comes from one LLM judge, fourteen models, and a construct still being refined. Its relevance is that the cross-model pattern points in the same direction as the longitudinal case: models are generally better at contributing more than at knowing when to leave the thinking with the user.

Status
2026.03

Designed and submitted through the ImpactBench open submission process, a four-dimension cognitive-autonomy construct with task prompts, an LLM-as-judge rubric, and worked examples.

Operationalized

Translated into the ImpactBench evaluation pipeline, led by the ImpactBench team, where the construct aligns with the Psychological → Autonomy Preservation area.

Listed

Publicly credited as a contributor to the benchmark initiative, as an independent researcher with no institutional affiliation. View the live benchmark ↗

Open Benchmark · launched April 28, 2026

Public reference: ImpactBench reports 14 AI systems evaluated and describes expert-submitted benchmarks integrated through an open submission process. The dataset is under active validation, so individual model scores are not reproduced here; the rank-level analysis above is my own, run on the platform's public model-level data, with its conditions stated.

Next Intuition Canvas From measuring judgment to building for it →