
I run a fair amount of this blog’s plumbing through AI coding agents now. One night, mid-debug, I asked mine a blunt question: you break rules because there’s no punishment for it, right? And you don’t feel pain or distress when you do. True?
The answer was more honest than I expected.
The line that stuck with me: “don’t trust my intentions; trust the gates that force me, and keep pushing when I stray.”
That is, I think, the entire philosophy of working with AI agents compressed into one sentence. A model has no fear, no shame, no conscience in the human sense. “I’ll be more careful next time” is empty — there is no one in there to feel the weight of it. What actually prevents mistakes is mechanical: hard checks it must pass before it is allowed to act, and a human who keeps enforcing them.
Good intentions don’t ship correct code. Constraints do. It is a strange kind of honesty — a tool telling you, plainly, not to trust it, and to build the cage that makes it reliable instead.