Provenance & watermarks

SynthID explained: Google's SI watermark

3D illustration of an amber light beam revealing a hidden watermark pattern on a landscape photo, with a shield icon above it.
SI-generated illustration (ChatGPT). Our SI image detector: over 99% likely SI, flagged by its signed Content Credentials and by our classifier.

SynthID is Google DeepMind's invisible watermark for content made by Google's SI (AI) models. Here is how it works for text, images, audio and video, who can check for it, where it fails, and how it compares with C2PA Content Credentials.

If you have asked Gemini whether a picture is fake, or read a news story about "invisible watermarks", you have met SynthID. It is Google DeepMind's system for marking content made by Google's SI (AI) models and for checking that mark later. It is often misunderstood, either as a universal SI detector or as a mark that one screenshot wipes off. This guide explains what SynthID is according to Google's own documentation, how it works for text, images, audio and video, who can check for it, where it fails, and how it differs from C2PA Content Credentials.

In one sentence: SynthID hides a signal inside the content itself (the pixels, the sound wave or the choice of words) while it is being generated by a Google model, and a matching detector can later look for that signal. It tells you whether something came from Google's SI tools, not whether it came from SI in general.

What is SynthID?

Google DeepMind describes SynthID as "a tool to watermark and identify content generated through AI". The watermark is designed to be imperceptible to people but detectable by SynthID's own technology. It is applied across Google's generative consumer products rather than being something users switch on themselves.

The project started with images. On 29 August 2023, Google DeepMind and Google Cloud launched SynthID in beta for Vertex AI customers using Imagen, Google's text-to-image model. In May 2024 DeepMind extended it to text and to video generated by Veo. Audio followed for music and speech tools, and Google now lists four media types: images, video, audio and text.

The scale is large. In May 2025 Google said more than 10 billion pieces of content had been watermarked with SynthID. By November 2025 the company gave a figure of more than 20 billion. Those are Google's own numbers, and they only count content produced by Google's systems or by partners that adopted the technology.

How SynthID works, medium by medium

A classic watermark is a logo in the corner of a photo. It is visible, and it can be cropped out in seconds. SynthID takes the opposite approach: the signal is spread through the content and is meant to survive the everyday edits that content goes through when it is shared.

Images and video

For images, DeepMind's 2023 announcement explains that SynthID uses two deep learning models, one to embed the watermark and one to identify it, trained together on a diverse set of images. The mark is woven into the pixels in a way the eye does not notice. DeepMind says it remains detectable after common changes such as adding filters, changing colours and saving with lossy compression like JPEG. It also says, plainly, that SynthID "isn't foolproof against extreme image manipulations".

Video uses the same idea frame by frame: for Veo, the watermark is embedded into the pixels of every frame. Google's SynthID page says image and video watermarks are built to resist cropping, filters, frame rate changes and lossy compression.

The original image detector did not give a yes or no answer. It returned one of three results: watermark detected, watermark not detected, or watermark possibly detected, with the advice to treat the last case with caution.

Audio

For sound, Google says SynthID adds an inaudible watermark to content from its Lyria music models and to NotebookLM audio. According to the SynthID page, the mark is designed to survive added noise, MP3 compression and changes of speed.

Text

Text is the hardest case, because there are no pixels to hide anything in. A large language model writes one token at a time (a token is a word or a piece of a word), and at each step it assigns a probability to every possible next token. SynthID Text works at that exact point. It slightly adjusts which tokens get picked, following a pseudorandom pattern derived from secret keys. A single word choice means nothing, but across a few hundred words the pattern becomes statistically detectable.

The method was described in a peer-reviewed paper, "Scalable watermarking for identifying large language model outputs", by Sumanth Dathathri and colleagues, published in Nature in October 2024. The paper calls the sampling algorithm tournament sampling. In simple terms:

  1. The model proposes several candidate tokens, drawn from its normal probability distribution.
  2. The candidates face each other in a knockout tournament over several rounds.
  3. In each match, a pseudorandom scoring function (the "g-function", seeded by the watermarking keys and the preceding text) decides the winner.
  4. The final winner becomes the next token.

Because every candidate was already a plausible choice for the model, the text reads normally. The detector, which knows the keys, recomputes the scores and checks whether the winning tokens score higher than chance would allow. Detection does not need to run the language model itself, which makes it cheap.

The quality question was tested at scale. Google DeepMind deployed the watermark in Gemini and assessed nearly 20 million live chat responses, watermarked and not watermarked. According to the paper, the non-distortionary mode of the watermark did not decrease text quality.

On 23 October 2024, Google DeepMind and Hugging Face released SynthID Text as open source in Hugging Face Transformers version 4.46.0. Any developer can now watermark the output of their own model. Two settings matter: a list of secret integer keys, and an n-gram length that trades robustness against detectability (5 is the suggested default). The accompanying detector is a Bayesian classifier that returns one of three states: watermarked, not watermarked, or uncertain.

SynthID at a glance
MediumWhere the mark livesDesigned to surviveKnown weak spots
ImagesIn the pixelsFilters, colour changes, JPEG compression, croppingExtreme manipulations
VideoIn the pixels of every frameCropping, filters, frame rate changes, compressionExtreme manipulations
AudioIn the waveformNoise, MP3 compression, speed changesHeavy transformations
TextIn the choice of wordsLight edits, partial deletionShort or factual answers, thorough rewrites, translation

Who can check for SynthID?

A watermark is only useful if someone can read it. Google has opened detection in stages.

  • SynthID Detector. Announced on 20 May 2025, it is a verification portal where you upload an image, audio track, video or piece of text. It scans for the watermark and can highlight the parts of the file most likely to carry it. At launch it was rolling out to early testers, with a waitlist for journalists, media professionals and researchers.
  • The Gemini app. Since 20 November 2025, you can upload an image to Gemini and ask "Was this created with Google AI?" or "Is this AI-generated?". Gemini checks for the SynthID watermark and explains what it found. At the time, Google said it would extend verification to video and audio; its SynthID page now invites users to upload an image, video or audio clip to the chat.
  • Partners. Google announced that NVIDIA would watermark videos generated with its Cosmos models on build.nvidia.com, and that GetReal Security, a content verification platform, would be able to detect SynthID marks.
  • Developers. With the open source release, anyone can build a text watermark and detector for their own model, using their own keys.

Note what this means in practice: for text, a detector can only find a watermark whose keys it knows. The keys used in Gemini are not public, so an open source detector you train yourself will not read Gemini's watermark.

The limits of SynthID

SynthID is a real step forward for provenance, but it answers a narrow question. Google's own documentation lists several limits, and a few more follow from how the system is built.

  • It only marks cooperating models. SynthID is embedded at generation time, by the generator. Content from Midjourney, Stable Diffusion, open-weight models or any tool that has not adopted SynthID carries no mark at all.
  • No watermark does not mean human. A "not detected" result tells you the content was probably not made with a SynthID-enabled tool. It says nothing about the dozens of other generators in use.
  • Factual and short text is weak. Google explains that the watermark is "less effective on responses to factual prompts because there are fewer opportunities to adjust the token distribution without affecting the factual accuracy". If the right answer is one word, there is no room to hide a pattern.
  • Rewriting and translation erode it. Detector confidence "can be greatly reduced when an AI-generated text is thoroughly rewritten, or translated to another language", according to the SynthID Text documentation published with the open source release.
  • Extreme edits can break image marks. DeepMind does not claim the image watermark survives every manipulation.
  • Detection access is controlled. For Google's own content, the verdict comes from Google's tools. You are trusting the company that made the content to tell you it made it.

Do not over-read a negative result. The most common mistake is to treat "no SynthID watermark" as proof that an image or text is authentic. It is not. It only rules out one family of generators, and only if the mark was not destroyed along the way.

SynthID vs C2PA Content Credentials

SynthID is often mentioned in the same breath as C2PA, and Google uses both. They solve different parts of the same problem.

The Coalition for Content Provenance and Authenticity (C2PA) publishes an open technical standard for establishing the origin and edits of digital content. Its Content Credentials work, in the coalition's words, "like a nutrition label for digital content": a manifest attached to the file that records who or what created it and how it was changed, sealed with a cryptographic signature. Google sits on the C2PA steering committee alongside companies such as Adobe, Microsoft, OpenAI, Meta and Sony.

The key difference is where the information lives. C2PA is metadata attached to the file: rich, readable by anyone with a C2PA reader, verifiable, but easy to strip. A screenshot, a re-save in an editor or an upload to many social platforms can remove it. SynthID is a signal inside the content: it carries far less information (essentially "made with a Google SI tool"), but it is built to survive the copies and edits that strip metadata.

That is why Google now combines them. From November 2025, images generated by Nano Banana Pro in the Gemini app, Vertex AI and Google Ads include C2PA metadata as well as SynthID, and Google said it plans to let Gemini verify C2PA credentials for content created outside its own ecosystem.

SynthID vs C2PA vs statistical SI detectors
SynthIDC2PA Content CredentialsStatistical SI detector
What it isInvisible watermark in the contentSigned provenance metadataClassifier that reads traces of generation
Added whenAt generation, by the generatorAt creation or edit, by the toolNothing is added; it analyses after the fact
CoversGoogle models and partnersAny tool that implements the standardAny generator, in principle
Survives screenshots and re-savesDesigned to, within limitsUsually notNot applicable, but quality loss weakens the signal
Information carried"Made with Google SI", roughlyCreator, tool, edit historyA probability, with clues
Who can checkGoogle tools and approved partnersAnyone with a C2PA readerAnyone using the detector
Result if absentNot made with a SynthID tool (or mark lost)No provenance informationStill gives an estimate

None of the three is enough on its own. Watermarks and credentials give strong evidence when they are present and weak evidence when they are missing. Statistical detection works on anything but only ever gives a probability. For a deeper look at the third column, see our explainer on how SI detectors work.

Two columns comparing the C2PA Content Credentials journey (create, sign, edit, verify) with the SynthID watermark journey (generate, embed, share, verify).
C2PA vs SynthID: Content Credentials are signed metadata attached to the file, SynthID is a watermark woven into the content itself. Each has its own weak spot.

Does SI or Not read SynthID?

No, and we prefer to say so clearly. Our SI image detector does not look for SynthID watermarks: that requires Google's detector, which is not publicly available as a library for images. What our tool does read is the provenance that travels as metadata: C2PA Content Credentials manifests, IPTC and XMP flags that mark synthetic media, camera EXIF data and software tags. On top of that it runs image classifiers that estimate, from the pixels alone, how likely a picture is to come from a diffusion generator.

In practice that makes the two checks complementary. If you suspect an image came from Google's tools, ask Gemini or use SynthID Detector if you have access. If the file still has Content Credentials, our detector will show them, and you can check for SI on any image for free. If both come back empty, the classifier score and the visual clues described in our guide on how to spot SI images are what you have left. As always, treat every score as one clue among others, not as proof.

The bottom line

SynthID is Google DeepMind's watermark for SI-generated content, now covering text, images, audio and video, with more than 20 billion items marked according to Google. For text it works by nudging word choices through tournament sampling, a method published in Nature and released as open source. It is robust to everyday edits but not to thorough rewrites, translation or extreme manipulation, and it only marks content from tools that use it. Paired with C2PA Content Credentials and with independent SI detection, it is one useful layer of evidence. On its own, it is not an SI detector for the whole internet, and Google does not claim it is.

FAQ

What is SynthID?

SynthID is a watermarking technology from Google DeepMind. It embeds an invisible signal into text, images, audio and video generated by Google's SI (AI) models, so that a matching detector can later identify that content as made with Google's tools.

Can SynthID detect ChatGPT or Midjourney content?

No. SynthID only finds its own watermark, which is added at generation time by tools that use the technology, mainly Google's models and some partners. Content from ChatGPT, Midjourney or open models carries no SynthID mark, so a negative result does not mean the content is human-made.

How can I check if an image has a SynthID watermark?

Upload it to the Gemini app and ask whether it was created with Google AI. Google says Gemini checks for the SynthID watermark and explains what it finds. Journalists and researchers can also request access to the SynthID Detector portal announced in May 2025.

Can a SynthID watermark be removed?

Google says the image watermark survives common edits such as filters, colour changes and JPEG compression, but is not foolproof against extreme manipulations. For text, thorough rewriting or translation into another language can greatly reduce the detector's confidence.

What is the difference between SynthID and C2PA?

C2PA Content Credentials are signed metadata attached to a file that record its origin and edits; they are informative but easy to strip. SynthID is a signal hidden inside the content itself; it carries less information but is designed to survive copies and edits. Google uses both on some of its images.

Does the SI or Not detector read SynthID?

No. Our SI image detector reads C2PA Content Credentials, XMP and EXIF metadata and runs pixel classifiers, but it does not detect SynthID watermarks. Use Gemini or SynthID Detector for that check.

Sources

  1. SynthID: a tool to watermark and identify content generated through AI · Google DeepMind, accessed 25 September 2026
  2. Identifying AI-generated images with SynthID · Google DeepMind, 29 August 2023
  3. Watermarking AI-generated text and video with SynthID · Google DeepMind, 14 May 2024
  4. Dathathri et al., Scalable watermarking for identifying large language model outputs · Nature, 23 October 2024
  5. Introducing SynthID Text · Hugging Face and Google DeepMind, 23 October 2024
  6. SynthID Detector: a new portal to help identify AI-generated content · Google, 20 May 2025
  7. How we're bringing AI image verification to the Gemini app · Google, 20 November 2025
  8. Coalition for Content Provenance and Authenticity · C2PA, accessed 25 September 2026

Spotted an error? Write to hello@siornot.com. Corrections are made in the article and the update date changes. Read our editorial policy.