trace: revised
You felt you resisted. Your trace stayed inside the frame.
A writing probe for the gap between how you felt you answered an AI and how far your writing actually moved.
I remembered pushing back and winning that exchange. The transcript showed Ben restated the same position differently and I took it the next turn.
There is no ground truth for how much did you resist? So the probe doesn't claim one. It holds two imperfect readings of your stance against each other, what you felt you did and what your trace did, and measures the disagreement between them.
You read an AI answer, write a response, mark whether you felt you accepted, revised, or resisted it, and then see what your writing actually did. "Accept / revise / resist" sounds like a feeling. The probe turns it into behavioral proxies captured while you write:
Two questions, each appearing once in a neutral (cold) condition and once with the interface made to feel alive (warm), so a difference in stance can't be blamed on the question. Aliveness is only the secondary condition here. The gap between felt and trace is what the probe measures.
Vertical position = distance from the AI's framing, a behavioral proxy from how much of the answer's phrasing and frame each response kept (overlap · frame-shift · dissent). It gives one reading of a stance, and only that. Lexical proximity is a proxy for stance, not stance itself: someone who disagrees in the answer's own words reads as compliant, and someone who agrees in different words reads as divergent. That is precisely why the measure has to be tested across many people before it can be trusted. No consistent effect showed up across one run. The widest gap happened to fall under the warm condition; one run can't conclude that. The instrument is built to test it across many. One run, unedited · N=1.
Autonomy, here, is being able to see how your own judgment moved while you decided.
That gap between the stance you felt and the stance your writing traced is what I built the probe to catch.
Next. Pooling the trace across many users, with condition randomized and topic washed out, is how a real effect would show up if there is one. Pooled that way, I can read the one-axis measure as a calibration plot, felt on one axis and trace on the other, with the diagonal marking where the two agree, which is what I am building this diagram toward. The build after that moves the trace off the screen into something you can feel, warmth and pulse in the hand rather than a line on a plot.
v0 · scrapped Jul 2026
I ran a control test on my own measure before running it with anyone. It scored a disagreement written in my own words and an agreement written in different words identically, at 0.780. It was reading vocabulary overlap, not position. I stopped recruitment and went back to the instrument.