Full text coming soon.
The Alignment Problem Is a Values Problem, Not a Technical One
We keep framing alignment as an engineering challenge — as if the right loss function will eventually produce a machine that wants what we want. But we don't agree on what we want. The alignment problem is downstream of that disagreement, not upstream of it.
More in AI
The Model Doesn't Know It's Lying
There's a particular kind of unease that comes from watching a language model confabulate with total confidence. It's not deception in any meaningful sense — there's no intent, no self-awareness of the gap. That's what makes it stranger than lying.
Context Windows and the Illusion of Memory
A language model with a million-token context window doesn't remember anything. It processes. The distinction matters more than it seems — memory implies continuity, a self that persists. What we're building are very sophisticated present-tense machines.