In January 2026, a student in the AI master’s program at UNIR, where I teach generative AI, asked in the course forum why a detector flagged the master’s thesis they had written themselves as AI-written. The short answer is that these detectors aren’t reliable, and that knowing for sure whether a text was written by AI is still, in general, an unsolved problem. The long answer explains why, and why the situation is different for audio and images.
There are two ways to try: mark the text as it’s generated, or guess afterward.
Marking text as it’s generated
A language model writes by choosing the next word (actually, the next word fragment, or token) according to a set of probabilities. A watermark uses that moment to leave a statistical signal in the text.
The best-known idea was published by John Kirchenbauer and colleagues at the University of Maryland in 2023. Before choosing each word, the system pseudo-randomly splits the vocabulary into a “green” list and a “red” list and slightly favors the green words. A reader notices nothing, but someone who knows the key can count how many green words there are and calculate whether it’s too much of a coincidence (Kirchenbauer et al., 2023).
Google uses a system of this kind, SynthID Text, in Gemini’s responses. In a paper published in Nature in 2024 they explained how it works and tested it on nearly 20 million real responses, with users noticing no loss of quality (Dathathri et al., 2024). The Hugging Face team has a good introduction to these techniques for text, images and audio (Hugging Face, 2024).
Why watermarks don’t solve everything
A watermark only helps if the model that generated the text added it and if you have that provider’s detector. A text written with a model that doesn’t watermark, or one whose detector isn’t public, carries no signal to look for.
What’s more, the mark weakens if the text is rewritten, translated or mixed with your own writing. And there’s a balance to strike: the stronger the mark, the easier it is to detect, but the more it can affect the quality of the text. That’s the hard trade-off WaterBench found when comparing several methods (Tu et al., 2024).
Open models make things even more complicated. If anyone can download the model, anyone can generate without a watermark. And a 2025 study showed that watermarks built into open models themselves disappear with common modifications such as fine-tuning them on other data, compressing them or merging them with other models (Gloaguen et al., 2025).
Guessing afterward: detectors
The detectors most people use don’t look for watermarks. They’re classifiers that estimate whether a text “looks” generated, mostly based on how predictable it is: language models tend to choose the most likely words, so a very predictable text looks suspicious to them.
The problem is that plenty of human writing is predictable too. Academic work is written in formal, orderly, unsurprising language, exactly the style models imitate. That’s why my student got a false positive. A Stanford study found that several detectors systematically classified texts by non-native English speakers, with a more limited vocabulary, as AI-generated (Liang et al., 2023).
OpenAI itself withdrew its detector in July 2023 because of its low accuracy: it identified only 26% of AI-written texts as likely generated, and flagged 9% of human-written texts as AI (OpenAI, 2023). On top of that, every model writes differently, and a detector trained on some doesn’t recognize others well.
Research keeps looking for alternatives. RepreGuard, published in 2025, proposes looking at a model’s internal activations as it processes the text, instead of at the final text, and gets good results in its tests (Chen et al., 2025). It’s promising, but it’s still far from being a tool you can trust to accuse anyone.
In audio and images, the problem is further along
With voice and images there’s much more room to hide a signal people can’t perceive. Meta’s AudioSeal marks audio with an imperceptible signal, detects it even in specific segments of a recording and holds up well against common manipulations (San Roman et al., 2024). Google applies SynthID to images, video and the audio of some of its products, and offers a portal to check whether a piece of content carries its mark (Google DeepMind).
This matters especially for cloned voices, which we’ve discussed in how many minutes of audio it takes to clone a voice and before you clone a voice, ask whose it is. And regulation is pushing in that direction: since August 2026, the EU AI Act requires providers to mark synthetic content so that a machine can detect it (Regulation (EU) 2024/1689, Article 50).
The same limitations remain: the mark is only there if someone put it there, and only someone with the right detector can find it.
What to do with a detector
If you teach or assess writing, a detector can at most be a hint to start a conversation. Never proof. It’s more useful to ask about the process: drafts, version history, sources consulted or an oral defense of the work.
If your organization generates content with AI, the most solid route is the opposite: mark it and say so. It’s easier to show that something was generated by AI when whoever generated it declares it than to try to guess afterward.
References
- Kirchenbauer, J., Geiping, J., Wen, Y. et al. (2023). A Watermark for Large Language Models. ICML 2023.
- Dathathri, S. et al. (2024). Scalable watermarking for identifying large language model outputs. Nature, 634.
- Hugging Face (2024). AI Watermarking 101: Tools and Techniques.
- Tu, S., Sun, Y., Bai, Y. et al. (2024). WaterBench: Towards Holistic Evaluation of Watermarks for Large Language Models. ACL 2024.
- Gloaguen, T., Jovanović, N., Staab, R. and Vechev, M. (2025). Towards Watermarking of Open-Source LLMs. arXiv.
- Liang, W., Yuksekgonul, M., Mao, Y., Wu, E. and Zou, J. (2023). GPT detectors are biased against non-native English writers. Patterns.
- OpenAI (2023). New AI classifier for indicating AI-written text (updated in July 2023 when it was withdrawn).
- Chen, X., Wu, J., Yang, S. et al. (2025). RepreGuard: Detecting LLM-Generated Text by Revealing Hidden Representation Patterns. arXiv.
- San Roman, R., Fernandez, P., Défossez, A. et al. (2024). Proactive Detection of Voice Cloning with Localized Watermarking (AudioSeal). ICML 2024.
- Google DeepMind. SynthID.
- Regulation (EU) 2024/1689 on Artificial Intelligence. Article 50: Transparency obligations.