Whenever AI in the media comes up, the reassuring answer is always the same: there’ll be a human in the loop checking everything. With automated systems, and more and more agents, that generate, summarize, translate and publish at a pace no newsroom can keep up with, that answer falls short. I think the useful question is a different one: who checks, what they check and when.
The split that seemed reasonable
For a while, the split seemed clear. Routine, low-risk work went to machines: transcripts, summaries, headline variations, format changes. Sensitive work went to people: politics, investigations, anything where a mistake has consequences.
That’s largely the industry consensus. In a report published in August 2026 by the Reuters Institute, based on conversations with twenty newsroom leaders, experts and academics from thirteen countries, almost all agree that human oversight is still the main safeguard against AI-generated content (Sharma, 2026). The report itself warns, however, that this model is starting to show cracks.
Agents change the risk of routine work
The problem is that an agent can turn low-risk content into high-risk content without anyone deciding it. An election result in a small town looks like routine information. Multiplied across thousands of towns and published live, it’s election coverage at scale, and a mistake repeats thousands of times before anyone sees it.
We saw this up close in 2023 with the RTVE Elecciones proof of concept. Narrativa’s system generated automated news from the vote-count data, mainly for the small towns that almost never get coverage, and we provided the synthetic voice and the conversational experience. On the night of the general election, about 1,000 hours of audio were produced in three hours. No newsroom can listen to all that before publishing it. In an automated system like this, and even more so in an agent, the safeguard has to be somewhere else. At RTVE it was in the official source data, in journalistic oversight of the system that generated the news, in a pronunciation dictionary of towns and parties, and in months of prior testing.
Checking everything doesn’t scale
The obvious objection is to just put more people on review. The Reuters Institute report points to why that’s not enough: oversight falls on editors and middle managers who were already stretched, and Scroll, in India, has had to limit how much AI-assisted content each person is asked to review because reviewing was causing fatigue (Sharma, 2026). More is produced than can be reviewed. A tired person reviewing hundreds of similar texts ends up approving them without reading them.
It’s not just fatigue. A 2026 study by Steven Shaw and Gideon Nave, at the Wharton School, describes what they call “cognitive surrender”: the tendency to hand over judgment to AI when its answers sound fluent and confident. In three experiments with 1,372 participants, people accepted AI outputs even when they were deliberately wrong (Shaw and Nave, 2026, cited in Sharma, 2026). The review exists on paper, but it protects no one.
European regulation adds pressure. Since August 2026, the AI Act requires disclosure that a text on matters of public interest was generated by AI, unless it has gone through human review and someone takes editorial responsibility (Regulation (EU) 2024/1689, Article 50(4)). That exception only makes sense if the review is real.
From human in the loop to expert in the loop
The Reuters Institute report itself reaches a similar conclusion: human oversight is still necessary, but we have to accept its limits and design systems that don’t depend on a person catching every mistake (Sharma, 2026).
My proposal is to move from a person who checks every piece to someone who knows the subject and oversees the system. Some concrete ideas:
- Review upfront what’s going to be repeated. Templates, the agent’s instructions, data sources and rules get reviewed more carefully than each individual piece, because a flaw there multiplies.
- Put the right person on it. An election result should be overseen by someone who knows the electoral system; a health story, by someone who knows medicine. A generic reviewer won’t catch the mistakes that matter.
- Review by sampling and with alerts. Instead of reading everything, read a sample and automatically flag anything out of the ordinary: odd figures, unknown names, sudden changes.
- Look for patterns, not just one-off errors. Some biases are invisible in one piece and obvious in a thousand. As James Fletcher, the BBC’s responsible AI lead, explains, in one translated text you can’t tell whether masculine pronouns are used more than feminine ones; in ten thousand, you can (Sharma, 2026).
- Decide what doesn’t get automated. For some content the risk isn’t worth the speed, and it’s best to put that in writing before an agent starts producing it.
- Be able to correct quickly. If something goes wrong, you need to know what was published, where, and how to take it down or correct it.
What’s left for people
Automating the routine part can free up time for what only people do well: deciding what’s news, verifying, telling stories with context. But that only happens if oversight is designed from the start and done by someone who knows the subject. The expert brings the judgment, and that’s the key. If it’s bolted on at the end as a box to tick, the newsroom ends up working for the agent.
If you work in media and you’re thinking about how to oversee a system like this, it’s the kind of challenge we take on in evaluating assistants and agents and in our applied research projects.
References