Skip to content
yash.

Is there a "persuasion" direction inside the model?

Probing and steering an open model on real r/ChangeMyView replies. Lab notes, not a tutorial.

2026-10-05experiment~1 mingithub ↗

Sample post with placeholder content. Replace with your real notes. It's here to show every building block a post can use.

The question

Replies that earn a Δ on CMV changed someone's mind. If persuasiveness is linearly represented, a simple probe on the residual stream should separate Δ from non-Δ replies, and adding that direction back should make generations more persuasive.

What I read

The three sources in the frontmatter show up at the bottom of this post and on the Library page automatically.

Key ideas, in my words

A steering vector is the difference between mean activations of the two groups at layer ℓ\ell:

vℓ=μℓ(Δ)−μℓ(¬Δ),h′=h+α vℓv_\ell = \mu_\ell(\Delta) - \mu_\ell(\neg\Delta), \qquad h' = h + \alpha \, v_\ell

What I built

# grab residual-stream activations for each reply
for reply in replies:
    _, cache = model.run_with_cache(reply.text)
    X.append(cache[f"blocks.{layer}.hook_resid_post"].mean(1))
    y.append(reply.got_delta)

What I found

layer 0probe accuracy by layerlayer 23
Fig 1. Probe accuracy by layer (illustrative placeholder values). Hover a bar.

I see your point, but the evidence you cite supports the opposite conclusion

Fig 2. Hover or tap a word to see attention arcs (illustrative weights).
[ Steer α and generate: Hugging Face Space goes here ]

What broke

What broke
Δ replies are longer on average. Is the probe just detecting length? Need a length-matched control.

What's next

Length-matched dataset, sweep more layers, and test whether steered generations actually read as more persuasive to humans.

Sources

Discussion

[ giscus comments: set site.giscus in src/lib/site.ts ]