5 minNews
AI Watermarks and Detectors May Create False Sense of Security, Researcher Warns
A new analysis argues that AI watermarks and detection tools are unlikely to solve the problem of identifying machine-generated content online, and may instead foster misplaced confidence in unmarked material.
As tech companies roll out watermarks and detection tools to help internet users identify AI-generated content, a new analysis warns that these measures may backfire by creating a false sense of confidence in what appears online. The argument, published by a researcher who studies digital literacy, contends that both watermarks and detectors are vulnerable to simple workarounds and that their limitations are often understated by the companies promoting them.
The push for transparency tools has accelerated in recent months. Claude now watermarks AI-generated text to comply with European Union transparency rules, while OpenAI and Google add invisible fingerprints to AI-generated images. Substack has introduced a feature that scans pieces for signs of AI. These efforts respond to a growing crisis of trust, as deepfakes have become increasingly difficult to distinguish from authentic recordings. A 2023 AI-generated video of the actor Will Smith eating spaghetti showed obvious distortions, but by 2025, AI was producing lifelike renditions that experts say are convincing enough to warrant family codewords for verification.
The researcher, who led a 2022 study on how American universities teach online information evaluation, draws a parallel to the early internet. When polished websites were rare, visual cues like design quality and broken links were reliable indicators of credibility. But as inexpensive software made slick graphics ubiquitous, those signals stopped meaning much. The study found that 96% of leading colleges and universities still offered outdated advice on evaluating online information, long after bad actors could easily create convincing fake websites.
The core problem, the analysis argues, is what it calls the inverse illusion: the tendency to believe that if the presence of a signal proves one thing, its absence proves the opposite. A site with misspellings may indeed be unreliable, but a beautiful site with a dot-org domain can also be harmful. The researcher notes that nearly half of hate groups had dot-org domains in a 2019 study. The same logic applies to AI content: the absence of visible flaws does not mean content is genuine.
Watermarks and detectors face practical limitations. The researcher, who is not a software engineer, says they stripped metadata from AI-generated images simply by screenshotting them. Anthropic confirms that file metadata can be removed through format conversion, re-saving, screenshots, or other means. Stronger watermarks like SynthID can persist after screenshots, but the researcher says a free online tool was able to remove one. Google acknowledges that the accuracy of detecting watermarked AI text is greatly reduced when users thoroughly rewrite what they generate, and that it is not designed to stop motivated adversaries.
Open-weight AI models that can run locally, outside platform terms and conditions, further guarantee the spread of unmarked content. Third-party detectors have a spotty track record, with studies finding they do not work consistently across text, images, and audio. Yet their findings are often used as the basis for public accusations. Every detector must confront an arms race with humanizer tools and other workarounds.
The analysis argues that the biggest problem remains the inverse illusion. Just because content lacks a watermark does not mean it was not produced or edited with AI. As Anthropic notes, the lack of a detected mark does not mean content was not AI-generated or processed. Deferring judgment to detectors leaves users vulnerable to bad actors who know how to launder content and make it pass muster.
Rather than hunting for visual clues or outsourcing judgment to detectors, the researcher recommends turning to reputation and context. It is easy to fake content, but much harder to fake a good reputation validated by credible sources. The advice comes amid widespread uncertainty: in a recent pilot, the researcher's group showed 117 students a confident chatbot answer about local history with hallucinated facts, and half said they were not sure if it was true. One student said AI is sometimes right and sometimes wrong, and you never know which is which.
