Is there a "persuasion" direction inside the model?
Probing and steering an open model on real r/ChangeMyView replies. Lab notes, not a tutorial.
Sample post with placeholder content. Replace with your real notes. It's here to show every building block a post can use.
The question
Replies that earn a Δ on CMV changed someone's mind. If persuasiveness is linearly represented, a simple probe on the residual stream should separate Δ from non-Δ replies, and adding that direction back should make generations more persuasive.
What I read
The three sources in the frontmatter show up at the bottom of this post and on the Library page automatically.
Key ideas, in my words
A steering vector is the difference between mean activations of the two groups at layer :
What I built
# grab residual-stream activations for each reply
for reply in replies:
_, cache = model.run_with_cache(reply.text)
X.append(cache[f"blocks.{layer}.hook_resid_post"].mean(1))
y.append(reply.got_delta)What I found
I see your point, but the evidence you cite supports the opposite conclusion
What broke
What's next
Length-matched dataset, sweep more layers, and test whether steered generations actually read as more persuasive to humans.
Sources
- paperRepresentation Engineering: A Top-Down Approach to AI Transparency · Placeholder note: my one-line takeaway.
- paperSteering Language Models With Activation Engineering
- paperWinning Arguments: Interaction Dynamics and Persuasion Strategies in Good-faith Online Discussions