Alignment, in plain terms: making capable models do what we actually want
“AI safety” gets talked about as either science fiction or box-ticking. The real, present-day version is narrower and more practical than either.
Say “AI safety” and people picture one of two things: a distant, superhuman intelligence out of a film, or a compliance checkbox on a procurement form. The version that matters to anyone building today sits between those, and it has a plain name: alignment.
Alignment is the problem of getting a system to do what you actually want — not merely what you literally asked for. If that sounds trivial, it is worth remembering how often it is not true of software, of contracts, or of instructions given to a very literal new hire.
The gap between the instruction and the intent
A capable model optimises for the objective it is given, and objectives are slippery. Ask for engagement and you can get outrage. Ask for a summary that “sounds confident” and you can get confident fabrication. The model is not malfunctioning; it is doing exactly what was specified, which turns out to be not quite what was meant. That gap — between the letter of the instruction and the spirit of the intent — is where alignment lives.
Most alignment failures in production are not the model rebelling. They are the model obeying an instruction you did not realise you had given.
Three things it actually means day to day
Stripped of the philosophy, working alignment is mostly three concerns:
- Honesty. Does the system tell you when it does not know, or does it produce a fluent guess indistinguishable from a fact? A confident wrong answer is worse than a refusal.
- Controllability. Can you steer and constrain behaviour reliably — and does it stay within bounds when it meets inputs you did not anticipate, including adversarial ones?
- Corrigibility. When it is going wrong, can you tell, and can you stop it? Observability and an off-switch are alignment features, not afterthoughts.
None of that requires believing in science fiction. All of it shows up the first week a real system meets real users.
Why it gets harder as models get better
The uncomfortable dynamic is that more capable systems make misalignment more consequential, not less. A weak model that misunderstands you produces a bad paragraph. A capable one wired to tools can act on the misunderstanding — send the email, run the query, take the step — before anyone reviews it. Capability and the cost of misalignment climb together.
That is the sober case for taking this seriously now, at ordinary product scale, rather than treating it as someone else's long-term research problem. The practical disciplines — clear objectives, adversarial testing, human review before irreversible actions, and the ability to observe and halt — are the same ones that make any AI product trustworthy. Alignment is not a separate, exotic field bolted on at the end. On a good team it is just what building responsibly looks like.
Sources & further reading
Writes and edits Troiana Signal’s coverage of AI, product building and modern discovery.
Join the discussion
Useful counterpoints, first-hand experience and corrections are welcome. Every response is reviewed before it appears.
No published responses yet. Start with something that adds to the article.

