If you have asked Gemini whether a picture is fake, or read a news story about "invisible watermarks", you have met SynthID. It is Google DeepMind's system for marking content made by Google's SI (AI) models and for checking that mark later. It is often misunderstood, either as a universal SI detector or as a mark that one screenshot wipes off. This guide explains what SynthID is according to Google's own documentation, how it works for text, images, audio and video, who can check for it, where it fails, and how it differs from C2PA Content Credentials.
In one sentence: SynthID hides a signal inside the content itself (the pixels, the sound wave or the choice of words) while it is being generated by a Google model, and a matching detector can later look for that signal. It tells you whether something came from Google's SI tools, not whether it came from SI in general.
What is SynthID?
Google DeepMind describes SynthID as "a tool to watermark and identify content generated through AI". The watermark is designed to be imperceptible to people but detectable by SynthID's own technology. It is applied across Google's generative consumer products rather than being something users switch on themselves.
The project started with images. On 29 August 2023, Google DeepMind and Google Cloud launched SynthID in beta for Vertex AI customers using Imagen, Google's text-to-image model. In May 2024 DeepMind extended it to text and to video generated by Veo. Audio followed for music and speech tools, and Google now lists four media types: images, video, audio and text.
The scale is large. In May 2025 Google said more than 10 billion pieces of content had been watermarked with SynthID. By November 2025 the company gave a figure of more than 20 billion. Those are Google's own numbers, and they only count content produced by Google's systems or by partners that adopted the technology.
How SynthID works, medium by medium
A classic watermark is a logo in the corner of a photo. It is visible, and it can be cropped out in seconds. SynthID takes the opposite approach: the signal is spread through the content and is meant to survive the everyday edits that content goes through when it is shared.
Images and video
For images, DeepMind's 2023 announcement explains that SynthID uses two deep learning models, one to embed the watermark and one to identify it, trained together on a diverse set of images. The mark is woven into the pixels in a way the eye does not notice. DeepMind says it remains detectable after common changes such as adding filters, changing colours and saving with lossy compression like JPEG. It also says, plainly, that SynthID "isn't foolproof against extreme image manipulations".
Video uses the same idea frame by frame: for Veo, the watermark is embedded into the pixels of every frame. Google's SynthID page says image and video watermarks are built to resist cropping, filters, frame rate changes and lossy compression.
The original image detector did not give a yes or no answer. It returned one of three results: watermark detected, watermark not detected, or watermark possibly detected, with the advice to treat the last case with caution.
Audio
For sound, Google says SynthID adds an inaudible watermark to content from its Lyria music models and to NotebookLM audio. According to the SynthID page, the mark is designed to survive added noise, MP3 compression and changes of speed.
Text
Text is the hardest case, because there are no pixels to hide anything in. A large language model writes one token at a time (a token is a word or a piece of a word), and at each step it assigns a probability to every possible next token. SynthID Text works at that exact point. It slightly adjusts which tokens get picked, following a pseudorandom pattern derived from secret keys. A single word choice means nothing, but across a few hundred words the pattern becomes statistically detectable.
The method was described in a peer-reviewed paper, "Scalable watermarking for identifying large language model outputs", by Sumanth Dathathri and colleagues, published in Nature in October 2024. The paper calls the sampling algorithm tournament sampling. In simple terms:
- The model proposes several candidate tokens, drawn from its normal probability distribution.
- The candidates face each other in a knockout tournament over several rounds.
- In each match, a pseudorandom scoring function (the "g-function", seeded by the watermarking keys and the preceding text) decides the winner.
- The final winner becomes the next token.
Because every candidate was already a plausible choice for the model, the text reads normally. The detector, which knows the keys, recomputes the scores and checks whether the winning tokens score higher than chance would allow. Detection does not need to run the language model itself, which makes it cheap.
The quality question was tested at scale. Google DeepMind deployed the watermark in Gemini and assessed nearly 20 million live chat responses, watermarked and not watermarked. According to the paper, the non-distortionary mode of the watermark did not decrease text quality.
On 23 October 2024, Google DeepMind and Hugging Face released SynthID Text as open source in Hugging Face Transformers version 4.46.0. Any developer can now watermark the output of their own model. Two settings matter: a list of secret integer keys, and an n-gram length that trades robustness against detectability (5 is the suggested default). The accompanying detector is a Bayesian classifier that returns one of three states: watermarked, not watermarked, or uncertain.
| Medium | Where the mark lives | Designed to survive | Known weak spots |
|---|---|---|---|
| Images | In the pixels | Filters, colour changes, JPEG compression, cropping | Extreme manipulations |
| Video | In the pixels of every frame | Cropping, filters, frame rate changes, compression | Extreme manipulations |
| Audio | In the waveform | Noise, MP3 compression, speed changes | Heavy transformations |
| Text | In the choice of words | Light edits, partial deletion | Short or factual answers, thorough rewrites, translation |
Who can check for SynthID?
A watermark is only useful if someone can read it. Google has opened detection in stages.
- SynthID Detector. Announced on 20 May 2025, it is a verification portal where you upload an image, audio track, video or piece of text. It scans for the watermark and can highlight the parts of the file most likely to carry it. At launch it was rolling out to early testers, with a waitlist for journalists, media professionals and researchers.
- The Gemini app. Since 20 November 2025, you can upload an image to Gemini and ask "Was this created with Google AI?" or "Is this AI-generated?". Gemini checks for the SynthID watermark and explains what it found. At the time, Google said it would extend verification to video and audio; its SynthID page now invites users to upload an image, video or audio clip to the chat.
- Partners. Google announced that NVIDIA would watermark videos generated with its Cosmos models on build.nvidia.com, and that GetReal Security, a content verification platform, would be able to detect SynthID marks.
- Developers. With the open source release, anyone can build a text watermark and detector for their own model, using their own keys.
Note what this means in practice: for text, a detector can only find a watermark whose keys it knows. The keys used in Gemini are not public, so an open source detector you train yourself will not read Gemini's watermark.
The limits of SynthID
SynthID is a real step forward for provenance, but it answers a narrow question. Google's own documentation lists several limits, and a few more follow from how the system is built.
- It only marks cooperating models. SynthID is embedded at generation time, by the generator. Content from Midjourney, Stable Diffusion, open-weight models or any tool that has not adopted SynthID carries no mark at all.
- No watermark does not mean human. A "not detected" result tells you the content was probably not made with a SynthID-enabled tool. It says nothing about the dozens of other generators in use.
- Factual and short text is weak. Google explains that the watermark is "less effective on responses to factual prompts because there are fewer opportunities to adjust the token distribution without affecting the factual accuracy". If the right answer is one word, there is no room to hide a pattern.
- Rewriting and translation erode it. Detector confidence "can be greatly reduced when an AI-generated text is thoroughly rewritten, or translated to another language", according to the SynthID Text documentation published with the open source release.
- Extreme edits can break image marks. DeepMind does not claim the image watermark survives every manipulation.
- Detection access is controlled. For Google's own content, the verdict comes from Google's tools. You are trusting the company that made the content to tell you it made it.
Do not over-read a negative result. The most common mistake is to treat "no SynthID watermark" as proof that an image or text is authentic. It is not. It only rules out one family of generators, and only if the mark was not destroyed along the way.
SynthID vs C2PA Content Credentials
SynthID is often mentioned in the same breath as C2PA, and Google uses both. They solve different parts of the same problem.
The Coalition for Content Provenance and Authenticity (C2PA) publishes an open technical standard for establishing the origin and edits of digital content. Its Content Credentials work, in the coalition's words, "like a nutrition label for digital content": a manifest attached to the file that records who or what created it and how it was changed, sealed with a cryptographic signature. Google sits on the C2PA steering committee alongside companies such as Adobe, Microsoft, OpenAI, Meta and Sony.
The key difference is where the information lives. C2PA is metadata attached to the file: rich, readable by anyone with a C2PA reader, verifiable, but easy to strip. A screenshot, a re-save in an editor or an upload to many social platforms can remove it. SynthID is a signal inside the content: it carries far less information (essentially "made with a Google SI tool"), but it is built to survive the copies and edits that strip metadata.
That is why Google now combines them. From November 2025, images generated by Nano Banana Pro in the Gemini app, Vertex AI and Google Ads include C2PA metadata as well as SynthID, and Google said it plans to let Gemini verify C2PA credentials for content created outside its own ecosystem.
| SynthID | C2PA Content Credentials | Statistical SI detector | |
|---|---|---|---|
| What it is | Invisible watermark in the content | Signed provenance metadata | Classifier that reads traces of generation |
| Added when | At generation, by the generator | At creation or edit, by the tool | Nothing is added; it analyses after the fact |
| Covers | Google models and partners | Any tool that implements the standard | Any generator, in principle |
| Survives screenshots and re-saves | Designed to, within limits | Usually not | Not applicable, but quality loss weakens the signal |
| Information carried | "Made with Google SI", roughly | Creator, tool, edit history | A probability, with clues |
| Who can check | Google tools and approved partners | Anyone with a C2PA reader | Anyone using the detector |
| Result if absent | Not made with a SynthID tool (or mark lost) | No provenance information | Still gives an estimate |
None of the three is enough on its own. Watermarks and credentials give strong evidence when they are present and weak evidence when they are missing. Statistical detection works on anything but only ever gives a probability. For a deeper look at the third column, see our explainer on how SI detectors work.
Does SI or Not read SynthID?
No, and we prefer to say so clearly. Our SI image detector does not look for SynthID watermarks: that requires Google's detector, which is not publicly available as a library for images. What our tool does read is the provenance that travels as metadata: C2PA Content Credentials manifests, IPTC and XMP flags that mark synthetic media, camera EXIF data and software tags. On top of that it runs image classifiers that estimate, from the pixels alone, how likely a picture is to come from a diffusion generator.
In practice that makes the two checks complementary. If you suspect an image came from Google's tools, ask Gemini or use SynthID Detector if you have access. If the file still has Content Credentials, our detector will show them, and you can check for SI on any image for free. If both come back empty, the classifier score and the visual clues described in our guide on how to spot SI images are what you have left. As always, treat every score as one clue among others, not as proof.
The bottom line
SynthID is Google DeepMind's watermark for SI-generated content, now covering text, images, audio and video, with more than 20 billion items marked according to Google. For text it works by nudging word choices through tournament sampling, a method published in Nature and released as open source. It is robust to everyday edits but not to thorough rewrites, translation or extreme manipulation, and it only marks content from tools that use it. Paired with C2PA Content Credentials and with independent SI detection, it is one useful layer of evidence. On its own, it is not an SI detector for the whole internet, and Google does not claim it is.



