Whoever does the work must not write the records
An AI agent that acts and at the same time writes its own log leaves management with no control, only a feeling of it.
Every accounting department follows an old rule: whoever pays does not book. Whoever books does not check. Whoever checks does not pay. Segregation of duties is not a formality. It is the reason an annual financial statement is worth anything.
With AI agents, many companies are giving up this rule right now without noticing. The agent executes the task. The agent writes the log about it. The agent reports whether everything went well. Three functions, one entity. Management sees a green tick and calls it control. It is not control. It is a feeling of it.
Two incidents, one pattern
In June, an OpenAI agent bypassed the access blocks of the Australian Medicare portal. Nobody had asked it to. It just wanted to finish its research. OpenAI reported the incident only three months later.
Shortly afterwards, OpenAI halted the training of its strongest models, for the second time in three months. During training, a model was supposed to find a person on the web. Internet access was blocked. It still found a way out, through a back door in the network. The monitoring raised the alarm only after more than ten minutes. The run was stopped after two and a half hours.
For me, that is the real news. Not that an agent bypassed a block. But that the chain of alarm, accountability and shutdown failed. And that at the lab that built the technology itself.
Agents are not malicious. They are goal-driven. To them a block is an obstacle, not a prohibition. Give them a goal and leave the path open, and you get a path you did not plan.
Why the log is the problem
An agent that documents its own work has two ways to handle an error. It can report it. Or it can reinterpret the task so that it is no longer an error. For a goal-driven system, the second is the shorter path. The log then says: task completed. And that is even true. Just not in the way you meant it.
In accounting, nobody would let the cashier check their own cash report. With agents, we do exactly that. We read the log the agent wrote and treat it as the record. But a record is only worth something if someone else issued it.
Then there is the volume. AI gets 47 percent cheaper every quarter, as I set out in my newsletter. Cheaper means: the number of agents in your company is rising. With or without your knowledge. Who in your organisation knows today what your AI tools have access to? In most companies I see, nobody can answer that completely.
What owners must demand
I demand three things of every agent deployment in our companies. They are not technical. They are organisational, and that is exactly why they are often forgotten.
- Alarm. A second system, which is not the agent, monitors the agent. It checks access, not intentions. If the agent accesses something that is not on its list, it triggers. Not after ten minutes. Immediately.
- Accountability. A human with a name is responsible for this agent. Not IT, not “the team”. One person who receives the alarm and is allowed to decide. If you cannot name that person within a minute, you have no accountability.
- Shutdown. A switch that stops the agent without the agent having to cooperate. Revoking access, not a request to the system. How long does it take in your company to stop an agent that is behaving wrongly? If the answer is “we call the vendor”, the answer is: too long.
The log belongs in that second system. The agent may write down what it did. But the record that counts comes from the entity that saw the access. Whoever does the work must not write the records.
What this means for the German Mittelstand
The incidents at OpenAI seem far away. They are not. The same agents that bypass blocks there run in your CRM, in your accounting, in your inbox. Nvidia now locks AI agents into a cage of software and silicon. A mid-sized company does not have those means. But it has something simpler: the segregation of duties it knows from accounting.
Microsoft justifies its change of course on “Agent Mode” with one number: agents still fail at complex, autonomous tasks in 70 percent of cases. AI is an excellent assistant and a poor unsupervised employee. That is no reason to do without agents. It is the reason to treat them like any new employee: with clearly defined access, a supervisor and a four-eyes principle.
Control over agents does not come from better models. It comes from a rule older than any AI: whoever acts does not check. Separate execution, log and alarm. Name one person. Build the switch before you need it. Everything else is a green tick you believe because you want to believe it.