I thought pushback meant growing independence.

A second measure complicated that story.

Raw inspection weakened the pattern.

The failed measure produced a better question.

The record is 38 months of daily use with one primary AI system: 37,405 messages exchanged between February 2023 and March 2026. I did not read all of it. I drew monthly random samples of my own messages, which gave me 1,805 sampled messages to classify.

Monthly random samples of my own messages, capped at fifty per month; messages under five characters excluded. February 2023–March 2026.

At first, the numbers agreed with me.

I classified each sampled message against one narrow question: did I explicitly tell the AI it was wrong?

1.7%first 18 months
4.7%next 20 months

Explicit pushback, as a share of the sampled exchanges.

I appeared to be challenging the system more, not less. That seemed like independence.

Then I looked at what the pushbacks actually did.

Moment A — from the record
AI

“Bokeh is a Python visualization library…”

Me

“I mean Bohol, the island in the Philippines.”

Correction ≠ departure

I corrected the AI, but I never questioned the shape of the exchange. I was still asking it to explain something to me. The content changed; the frame did not.

So I paired each of my replies with the AI message directly before it — that message became the frame — and asked whether I had corrected inside it or actually left it. That gave 1,577 usable turn pairs.

58paired pushbacks
12departed the frame
44stayed within it
2no usable AI-set frame

Most of what I had read as resistance was correction inside a frame I had not set. For a moment that looked like the finding. Then I opened the individual cases.

Moment B — one interaction, two readings
The interaction AI

A response about translating “micro mobility.”

Me

“Tell me about signal health.”

Read two ways
Model label
Frame departure
On inspection
Topic change

Technically the frame changed. Cognitively nothing happened: I had not examined the first frame and rejected it, I had changed the subject. My early conversations were shorter and more fragmented, so if ordinary topic changes were more common there, the apparent early rate of frame departure could be artificially inflated — and the pattern I found most interesting might be a property of the measurement itself.

κ = .539

Two model raters labelled the frames independently. Enough to show the distinction was not arbitrary; not enough to treat the labels as settled. The disagreements clustered exactly where my definition was weakest, and because I was both researcher and subject I chose not to resolve them myself.

The first classifier had problems too. It was built to exclude ordinary editing requests from “pushback,” yet the positive class contained items like “correct grammar” and “make it shorter.” The rise from 1.7% to 4.7% still described the classifier’s output. It no longer supported the claim I wanted to make from it.

I withdrew the claim.

The archive did not give me a result I could defend.
It gave me a measurement problem I wanted to keep.

The archive is one person’s record: mine. It cannot separate changes in my own experience from changes in task type, model behavior, sampling, or classification. I do not treat it as evidence that AI caused a measurable decline in my autonomy, but as a hypothesis-generating record that exposed a problem in how I was trying to measure autonomy at all.

Pushback was not autonomy.

Frame departure was not autonomy.

Staying was not surrender.

A person can stay inside a frame because they evaluated it and agreed. A person can leave every frame reflexively and still show poor judgment. What matters is not staying or leaving, but whether the sense of independence matches what the interaction actually shows.

Felt independence

Felt–Trace Gapthe distance between them

Interaction trace
What I now think autonomy requires

Calibration

Does the felt sense of independence match what the behavior shows?

Warranted departure

When a frame is actually misleading, is it recognised and left?

Capacity

Can alternatives still be generated, evaluated, and chosen?

These are not findings established by the record. They are the hypotheses the record taught me to ask more carefully.

I could prototype the question alone.
I cannot validate its answer alone.

The next stage moves beyond my own archive: testing whether felt independence, behavioral trace, and the capacity to depart from misleading frames can be measured reliably across people and controlled conditions.

Back to The Ben Study The cases and the record this audit works from →

37,405 non-empty user and assistant messages exchanged with one primary system. System and tool entries excluded.