A stub, and possibly the note that ties the rest together.
Most alignment writing starts at values. Does the system want what we want. Will it keep wanting that as it gets more capable.
Those are real questions. But there is a smaller one underneath that gets skipped, and it is the one I actually work on.
Before a system can be aligned to anything, it has to be right about what it is referring to.
A system perfectly aligned to human values, acting on the wrong person, does harm. Not because its values drifted. Because its reference did. The values were fine. It was talking about somebody else.
That failure mode is not exotic. It is the ordinary, daily failure of every identity system I have worked with, and it does not require any superintelligence to show up.
Why this might matter more later, not less
The usual assumption is that a more capable system makes fewer of these mistakes. Probably true on average.
But capability also means acting faster, on more things, with less human inspection. So the rate falls while the blast radius grows. I am not sure those cancel out, and I have not seen anyone try to work out which way it goes.
Things to work out
- Is reference failure a special case of misalignment, or a separate axis
- Does anyone measure identity error rates in deployed AI systems, and if not, why not
- Whether the singularity arguments assume reference is solved, and what happens to them if it is not