Standalone · AI Governance · Human Oversight

How Will We Ever Work Alongside the Machine?

~15 min read · By Jeroen Janssen · September 2026

At three in the morning, Stephen Normandin showered, picked up his phone and tried to find work.

For almost four years he had delivered packages around Phoenix for Amazon Flex. He was sixty-three and an Army veteran. He knew the city and he knew the job. Before leaving a depot he loaded his 2002 Toyota Corolla so the first deliveries sat beside him and the last were buried at the back. His ratings had been good enough that Amazon asked whether he might train other drivers.

That morning the app would not let him in.

An email said his standing had fallen below an acceptable level and that he was no longer eligible to deliver. A name appeared at the bottom. It was not the name of anyone who had seen the apartment gates that were locked during his predawn routes, or the customers who did not pick up, or the Amazon locker that failed to open. When the locker failed he had spent half an hour with support, and support told him to return the packages. His rating fell anyway.

He appealed. A reply thanked him for providing more context. He sent more. The same reply came back under a different name. He copied Jeff Bezos and asked how the decision had been made. More acknowledgements, then a final refusal.

He was not reinstated. He tried other delivery companies, then used pandemic stimulus money to start repairing small engines.

Normandin’s account comes from Bloomberg (Soper, 2021). Amazon disputes the wider portrayal. The reporting rested on interviews with fifteen drivers and former Amazon managers and engineers, and no court has found that a particular algorithm ended his work. Those limits are worth keeping.

What they leave standing is this. Amazon could watch Normandin work, score him, cut off his income and process his appeal. He could not find out which deliveries had counted against him, which rule turned them into a number below the line, whether any person had read what he wrote, or what evidence would have changed the answer.

Amazon had a complete record of all of it. That was never his problem.

Six things went wrong, and they were different things

He could not find out how he was being described, or correct it.

He objected, and nothing he wrote obliged anyone to answer him.

He had no part in the decision that ended his work, and no way to reopen it.

Something acted on him, and he could not tell what had authority to do so.

Nobody could halt it once it started.

And nobody counted what four years of route knowledge had been worth, or where that value went when he stopped.

Those are six separate failures, and the unit is one episode: a single matter, from the moment it enters the organisation to the moment its consequences land on somebody. An organisation can be excellent at one of the six and hopeless at the next, which is why a single score for “governance” tells you nothing. I have spent two years building an instrument that can test some of them by machine.

It is less than I wanted.

What I wanted was something that protects a man like Normandin before the fact. Something that makes it impossible for a decision to reach him already made, with everyone in the chain able to point somewhere else. The model scored him. The policy set the threshold. The reviewer applied the policy. The manager who signed off the policy left the company two years ago. Every person in that sequence acted correctly, no person decided anything, and the whole of it lands on one man at three in the morning with nobody to ask.

Every person in that sequence acted correctly, no person decided anything, and the whole of it lands on one man at three in the morning with nobody to ask.

That is the thing worth building and I have not built it. I do not know how to, and I am no longer sure it can be built out of records at all. What I have built tests whether a record can answer certain questions afterwards. It is a much smaller thing, and the distance between the two is what the rest of this is about.

What the instrument does

It is unglamorous. It is a vocabulary and a set of constraints, and it works on the record a system leaves behind.

An ordinary audit log records that something happened. A send occurred. An action was permitted. A reviewer confirmed a decision at 09:41:22. What it does not record is what kind of thing happened in law, or how any two entries relate to each other. So when a supervisor asks whether protected data left the company, the log can show that a file was sent and cannot show what class of data it held or whether the destination lay outside the boundary.

The instrument adds two things to a record. It labels events with the legal category they belong to. It preserves the connections between them: what was derived from what, on whose authority, valid over which period. Then a validator can be pointed at the record and asked a question, and it either answers or reports exactly which piece is missing and which obligation that piece served.

That is what it is designed to do. Whether it works is a separate question and I have not answered it. The test is written and deposited, it has not been run, and if a plain log turns out to answer as well as a labelled record then most of what follows is wrong. I would rather learn that from the experiment than from a regulator.[1]

With that on the table, three and a half of Normandin’s six fall inside what the instrument tries to do.

Authority. A delegation can be recorded with a start, an end and a scope, and every withdrawal of it tied to the grant it cancels and to a time. Then “did this system have the authority to do that, at that moment” is a question a machine settles in milliseconds.

Human review. In April 2023 the Amsterdam Court of Appeal looked at Uber, which deactivated drivers by algorithm and had a team in Krakow sign off on the deactivations. The court held the reviewers’ authority and competence had not been established and that the review was “not much more than a purely symbolic act.” A lower court had read the same facts the other way. Because of that judgment there is now a property in my vocabulary recording whether a reviewer’s authority was actually used, and a constraint that flags any decision where it was not. It took a court calling a human review symbolic before anyone thought to make its absence visible to software.

The law underneath is unsettled, and I can say roughly how much. Of 25 EU enforcement cases involving automated decisions, six were decided on entirely different grounds. That is 24 percent, and the honest range around it runs from about 11 to about 43 percent.[2]

I gave you the range for a reason. A governance number you can trust tells you how it was made, how wide it is, and what would have counted as being wrong. Almost nothing you will be shown this year does.

Stopping. A withdrawal of permission can be made a real event that propagates, so that at the moment it fires you can compute which agents, which queued work and which downstream decisions are affected.

Whether that matters is not an abstract question. On 18 March 2018 an automated test vehicle detected Elaine Herzberg 5.6 seconds before it struck and killed her in Tempe. The factory emergency braking had been switched off during automated operation, and stopping was left to an operator who was looking at the console. A human was present. Nothing stopped the car. The investigators did not reduce it to one distracted operator: their findings named inadequate risk assessment, ineffective oversight of the safety driver, and an inadequate safety culture at the operator’s parent company (NTSB, 2019).

Seeing, in one direction. The instrument can establish what class of data crossed which boundary, under what legal basis, derived from what. That answers the regulator’s question.

It does not answer Normandin’s. He wanted to know how he was being described, and to change it. Nothing here faces the person being described. Contesting a decision like his takes technical literacy, which is rare, or an organisation with it acting for you, which in Europe means a handful of bodies with a waiting list. So this one counts as half. A company can prove to a supervisor what it saw. The person it was looking at is exactly as far away as before.

That is three capabilities and half of a fourth. None of them would have helped Normandin.

The two nobody can check

Nobody can check whether he was heard.

Every exchange in his appeal would pass validation. There is a challenge, a response, a closure, timestamps in the right order. The record is complete and correctly formed. Nothing in it can register that the response answered nothing.

I assumed for most of a year that this was a modelling job I had not got round to. An objection has parts, and all of them look recordable: who raised it, what they claimed, what answer came back, whether it arrived before the decision closed. I sat down to write the constraints and got about half of them working.

Then I hit the part that stopped me. The objections that matter most are the ones nobody makes. A reviewer in Krakow who cannot afford to refuse leaves the same trace as a reviewer who read the file and agreed: a confirmation, a timestamp, a clean record. Whatever a person was afraid to say leaves no mark anywhere, because it never happened. There is no constraint you can write over a record that recovers something the record never contained.

There is no constraint you can write over a record that recovers something the record never contained.

Half of it is a job someone should do and I would like it done. The other half is not a job at all. Whether a person can afford to speak depends on their standing in the organisation, which no record of an episode contains. Any system built to check whether people were heard will be blind to the cases where nobody dared.

And nobody counted what it cost him.

Where the value of the work landed. Where the burden landed. Who absorbed the extra hours. Who decided that split, and on what authority.

I went through fourteen governance and agent vocabularies published between 2024 and 2026, scored against eighteen things a European supervisor plausibly needs. Not one of them represents any of that.

There is nothing difficult here. Money saved, hours released, headcount moved, who signed off on the reallocation: these are ordinary facts, no harder to record than provenance. They are absent because no regulator asks for them and no buyer scores them, so nobody has built the field.

So the checking now being built will be able to prove a decision was properly authorised, and will say nothing about who paid for it. Normandin paid for it. Four years of knowing which gates were locked at four in the morning went out of the business with him, and no record exists of what that was worth.

What can be said against this

Plenty of experienced people think the whole enterprise is a mistake, and their argument is not weak.

Oversight, they say, is a management responsibility. What matters is that a competent person is accountable, that actions are recorded well enough to review, and that the record cannot be quietly altered. Tamper-evidence, a named responsible human and a review process is what oversight has meant in regulated industry since long before software. Adding an ontology adds cost and specialists, and creates a new way to fail, in which people trust the schema and stop thinking. The AI Act asks for automatic event recording. Do that, put a qualified person on top, and you have met the duty.

Their strongest point has nothing to do with cost. Every previous attempt to formalise judgement has produced paperwork, and they expect this one to do the same.

I think they are wrong about the first thing and right about the second, which is not a comfortable position.

They are wrong because a sealed log proves nobody altered the record and says nothing about whether the record ever held the fact. Normandin’s file was almost certainly intact, retained and fully reviewable, and every question he asked was unanswerable from it. Under EU law a company can meet the letter of the logging duty and still be unable to show that meaningful human oversight was possible. That gap does not close by recording more.

Their second point is what worries me, and the rest of this is me agreeing with it about my own work.

What happens when four out of six can be checked

My files carry a disclaimer. A passing result means the record is structurally complete against a stated profile. It does not mean the company complied with anything, and it does not establish that the underlying decision was right. I wrote that disclaimer. It will not survive its first procurement department.

The sequence I expect is ordinary. Someone ships a tool. The tool produces a pass or fail. It covers authority, human review, stopping, and half of visibility. It is silent on whether anyone could object and on where the value went, and it does not mention that it is silent, because software reports what it can see and has no way to report its own blind spots. A buyer starts asking suppliers for a pass. An audit firm starts certifying against it. Within two years the pass is what “AI governance” means in a board pack.

The fair objection is that none of the six can be checked today either, so a tool covering four leaves the other two exactly where they were. Half a loaf.

I think that is wrong, for a reason about institutions rather than technology, and I will label it as a judgement rather than a finding. Today, if you stand up in a governance meeting and say we have no idea whether people can object here, that is a live problem and someone has to answer it. Once a tool exists, the same question sounds settled, because the absence of a red flag reads as a green one. Buyers ask for what can be evidenced. Boards review what buyers asked for. A condition with no field does not stay an open problem. It stops being mentioned.

Look at who that suits. A large company already running certified management systems absorbs the new requirements as an extension of what it has. Consultancies that can read the file formats become necessary. Regulators with technical staff gain reach and the rest fall further behind. A mid-sized company under high-risk obligations chooses between a specialist firm and a Big Four practice, which on my own estimates runs €25,000 to €100,000 a year against €500,000 to €2,000,000. Those are estimates rather than observed prices, and the distance between them is the point. Someone in Normandin’s position gains nothing, because nothing in any of this is built to face him.

I run one of those consultancies. I built the instrument. Everything I have just described is good for my business.

So is this essay. Being the person who says out loud what governance cannot deliver is a position, and it sells. I would rather you weighed what I have written knowing that than mistook the candour for disinterest.

Who is writing the rules

The high-risk obligations under the AI Act were pushed back this summer, to 2 December 2027 for standalone systems and 2 August 2028 for systems built into products. The official reason is that the technical standards and conformity procedures are not finished.

Those standards are being drafted right now, and they will decide what a record has to contain for the next decade. Writing them takes people who can read the file formats these things are specified in. There are a few hundred such people in Europe. Nearly all of us are paid, directly or through our firms, by the companies the standards will be used to inspect.

Nothing sinister in that. It is simply who has the skill. It does mean the question of which conditions get a field and which do not is currently being decided by people with an interest in the answer, and I am one of them.

Would any of it have saved him?

Probably not, and the reason matters more than the answer.

It is tempting to read what happened to Normandin as a man against a machine, and the serious literature invites it. Harari (2024) draws the line as sharply as anyone: AI “can process information by itself, and thereby replace humans in decision making. AI isn’t a tool—it’s an agent.” That is accurate about the systems in this essay, and it is why the record problem exists at all. Nobody asks who decided when a hammer is used.

The reading is still wrong, and Dignum (2019) states the correction in a sentence: AI does not happen to us, we make it happen, we are responsible. No algorithm decided Normandin should lose his income. A company decided how drivers would be scored. A company built something to apply that at a scale no person could manage. A company staffed an appeals process with people who had nothing to answer him with. Every one of those was a choice, and the software carried them out.

Crawford’s formulation holds both halves at once. AI, she writes (2021), is among other things “a form of exercising power.” Power belongs to whoever holds it. Nobody at Amazon needed to intend harm for this to happen. They needed only to build a chain in which every link was defensible on its own.

That is why the machine cannot be the villain here. If it is, you go looking for a technical fix, and there isn’t one. A record protects nobody. A record is evidence. It does its work afterwards, in front of a regulator or a court, and only when somebody with authority is already asking the question.

His appeal was recorded perfectly well. What was missing was anyone obliged to answer it. Obliged in the sense that failing to answer would have cost them something.

That obligation is not a data structure, and I cannot build it. Nobody can. It comes from law, from a union, from a works council, from a regulator with the appetite to use its powers, or from a company that decides on its own that a person who has worked for it for four years is owed a reply from a human being. Those are old instruments. They are not fashionable and they are the only things that have ever actually worked.

What the technical work can do is narrower and still worth doing. It can make an absence visible. It can put on the record, in a form a supervisor can query, that no person exercised authority here, that no answer was ever given, that this delegation had expired. It cannot create the obligation to care about the answer. It can remove the excuse that nobody knew.

It cannot create the obligation to care about the answer. It can remove the excuse that nobody knew.

You do not need anything I have built to find out where you stand. Take one consequential decision from the last three months, a real one with a person on the other end, and produce the evidence rather than the policy. Could that person have found out how they were being described, and corrected it? Did anyone object, what answer did they get, and did it arrive before the decision closed? What were the options and the criteria, who closed it, and what would have reopened it? What acted, under whose authority, and was that authority valid at the time? Could anyone have stopped it without damaging their career? Where did the benefit land, where did the cost land, and who chose that?

Four of those you will answer badly, which is a start. Two you will not answer at all, and it will not feel like a failure, because nothing in your governance stack is built to report it. Nobody is going to require you to. You will reach that in your own organisation faster than in my papers, and then it is yours to decide what to do about it.

Normandin went back to repairing small engines. He is not in anyone’s compliance evidence and he never asked to be. Four years in, he asked one reasonable question: who decided this, and did they read what I wrote. Within a few years a well-run European company will be able to answer the first half of that by machine. The second half nobody is building. Until someone is obliged to answer it, no amount of record-keeping is going to help the next man who wakes at three in the morning and finds the app has locked him out.

Notes

1.The design under test: three expert panels, the same real EU enforcement cases, one panel working from an action log, one from a log plus a provenance graph, one from a fully labelled record, and a comparison of whether the answers diverge. The materials are fixed and the pack is deposited. I cannot run it alone. It needs panels with real expertise in data protection, assurance or audit, and it needs somebody other than me doing the coding, because a result I produce about my own instrument is worth very little. That is a real request and it is open.

2.Six of 25 cases, a 24 percent central estimate with a Wilson 95 percent interval of [11.5%, 43.4%], and a sensitivity range across the borderline codings of [16%, 44%]. The sample was coded once, by me. A second coder has been commissioned, with the agreement threshold set in advance at Cohen’s κ ≥ 0.60 and a commitment to publish the result if it falls below. Until that comes back this is one person’s structured reading of the enforcement record rather than a validated finding. The criterion, the case sample and the machine-checkable artefacts are set out in Janssen (2026b) and Janssen (2026c). Read further there.

References

  • Crawford, K. (2021) Atlas of AI: Power, Politics, and the Planetary Costs of Artificial Intelligence. New Haven, CT: Yale University Press.
  • Dignum, V. (2019) ‘AI does not happen to us: we make it happen, we are responsible’, WFE Focus, August. Available at: focus.world-exchanges.org (Accessed: 2 September 2026).
  • Harari, Y.N. (2024) Nexus: A Brief History of Information Networks from the Stone Age to AI. London: Fern Press.
  • Janssen, J. (2026a) Who Gets to Decide? Six Powers for Human Work in the Age of AI. Deventer: Apparens. Available at: books.apple.com (Accessed: 2 September 2026).
  • Janssen, J. (2026b) From Record to Finding: Machine-Checkable Evidentiary Adequacy for Agentic AI Oversight under EU Law. Zenodo. doi: 10.5281/zenodo.21025237.
  • Janssen, J. (2026c) A Supervisory-Evidence Ontology for Agentic AI under EU Law: Candidate Minimum Conceptual Set and Temporal Extension. Zenodo. doi: 10.5281/zenodo.19758440.
  • National Transportation Safety Board (2019) Collision Between Vehicle Controlled by Developmental Automated Driving System and Pedestrian, Tempe, Arizona, March 18, 2018. Highway Accident Report NTSB/HAR-19/03. Washington, DC: NTSB. Available at: ntsb.gov (Accessed: 2 September 2026).
  • Soper, S. (2021) ‘Fired by bot at Amazon: “It’s you against the machine”’, Bloomberg, 28 June. Available at: bloomberg.com (Accessed: 2 September 2026).

Legislation and cases

  • Gerechtshof Amsterdam, 4 April 2023, ECLI:NL:GHAMS:2023:793 (case 200.295.742/01).
  • Regulation (EU) 2016/679 (General Data Protection Regulation), Article 22.
  • Regulation (EU) 2024/1689 (Artificial Intelligence Act), Articles 12 and 14.
  • Regulation (EU) 2026/1744 (Digital Omnibus on AI), amending Article 113 of Regulation (EU) 2024/1689.
All articlesAlle artikelen Related essay →Gerelateerd essay →