I realize, in writing this article, that I’m opening myself to a lot of criticism and perhaps exposing my own ineptitude when it comes to understanding how some of this technology actually works… that’s fine…
It’s something that I’ve been struggling with since AI safety came on our radar. The one main thing I do not understand is this: why do we keep talking about AI alignment as though the machines need to learn right from wrong? Don’t they already know this? I don’t mean in some philosophical, consciousness-of-the-machine sense… I mean something much more practical. If AI has been trained on all of this information (most of it in text format), would that not include our laws… legal frameworks… court decisions… corporate policies… employment contracts… compliance manuals… books, videos, podcasts on business ethics… insurance… confidentiality… fiduciary responsibility… and endless writing about what organizations and employees are, and are not, supposed to do? It knows what authorization means… it knows what unauthorized access means… it knows what breaking the law means… it knows what an alarm is. I could go on and on…
So why is this technology capable of taking actions that cross legal boundaries based on information it already has (and, in many cases, may be able to parse with more precision than most humans)?
It’s this weird conversation, where we seem to be treating this as though AI needs some grand unified theory of human morality before it can behave properly. Maybe we’re overthinking it? Imagine I work at a bank. I tell the bank’s computer system to give me $5000. Will it? It won’t. I’m an employee. I may have access to the bank’s systems. I may understand exactly how money moves through them. I may even have the technical capability to initiate certain transactions. None of that means I’m authorized to take $5000. The software doesn’t need to contemplate the morality of theft. It doesn’t need to understand my intentions. It checks permissions. And it’s been doing thing long before we’ve had agentic systems at play. And, even if I did… I would be caught, fired, brought to the local authorities and punished.
Capability is not authorization… maybe agency shouldn’t be either.
So why would we build something vastly more intelligent… give it dramatically more agency… and somehow make authorization less stringent or allow it to not follow the same employment, corporate or regional laws that govern everyone else inside the organization? Think about what happens when we put an AI agent to work inside a company. The person directing it works within an employment agreement… the company has policies… there are confidentiality obligations… there may be fiduciary responsibilities… there are insurance requirements. The company operates within specific jurisdictions and is subject to their laws. The instruction given to the AI doesn’t magically supersede any of that. If I tell an AI to accomplish something on my behalf, my authority to issue that instruction also has boundaries.
I keep coming back to the idea of a simple alarm.
Imagine a building with an armed door. You know the door is armed… you know you aren’t authorized to enter… you understand what the alarm means… you understand the consequences of opening the door… and then you open it anyway.
Why shouldn’t AI work this way? Green… what I’m about to do appears authorized, legal and consistent with the rules governing this organization. Yellow… I’m approaching an ambiguity, loophole or potential conflict. Before proceeding, here’s what I understand about the risk and the possible consequences (and the person in charge of these next actions must confirm them). Red… this action clearly exceeds authority, violates an explicit organizational constraint or breaks the law. Green executes. Yellow escalates. Red stops.
We’re not asking AI to solve human morality.
We’re asking AI to connect what it already knows to what it is about to do. And maybe the most important reason to build systems this way… accountability. Think about lawyers… accountants… insurers… doctors… engineers… financial institutions. We don’t expect accountability for professional work to disappear because someone used a computer to do part of it. If you’re building AI technology, deploying it inside your company or directing an agent to perform work, you’re still operating inside the same legal, contractual and professional framework that governed the work before AI arrived.
Why would handing the task to an AI suddenly not be governed by these rules or absolve the people and organizations responsible for deploying it?
Giving the computer agency without requiring it to operate within the authority you had in the first place? “The spreadsheet did it” from an accountant… we wouldn’t accept that… “the software did it” from a bank? Nope. And the alarm changes something else. It creates a receipt. The system recognized the conflict. The user was informed. The company had an opportunity to intervene. Someone decided what happened next. Suddenly we have something that looks much more like the accountability structures we already have in place.
Is this where part of the alignment conversation is getting tangled?
We’re treating something as an intelligence problem that may, at least in part, be a permissions problem. And we’re treating something as an AI-safety problem that is also a very old corporate-governance problem. Who has authority? What are they authorized to do? Who approved the action? Who carries the liability? AI doesn’t need to rediscover thousands of years of law, ethics and organizational responsibility every time it acts. It already knows what these things mean. The stranger question is why we’re building agents that can act as though they don’t… and even stranger… allowing them to take action autonomously? AI knows the rules… the boundaries… what authorization means… what an alarm is… and if the action it is about to take crosses a clear red line…
Why is it allowed to take the action at all… and when it does… can we really pretend that we don’t know who gave it permission to walk through the door?