Synthetic Content

Synthetic content refers to digital content that has been generated or substantially altered in terms of content, either entirely or to a significant extent, by artificial intelligence. This includes text, images, audio, videos and multimodal content. The decisive factor is whether AI has a substantial influence on the result – not whether it merely assists with minor corrections.

As generative AI becomes increasingly widespread, this distinction is becoming more important. This is because the nature of its creation raises further questions: Is the content authentic or manipulated? Can its origin be traced? And are there any transparency or labelling requirements for its specific use?

Key facts about synthetic content at a glance

  • Synthetic content is digital content that is entirely or predominantly generated or modified by AI.
  • A simple spell-check or standard image editing does not automatically make content synthetic content.
  • The terms ‘synthetic content’, ‘synthetic media’ and ‘AI-generated content’ overlap and are not used consistently across the board.
  • Deepfakes are a subset of synthetic content; synthetic data, on the other hand, refers to artificially generated datasets.
  • Synthetic content is not inherently problematic. Risks arise in particular from deception, disinformation, identity theft and a lack of transparency.
  • Since 2 August 2026, the transparency requirements set out in Article 50 of the EU AI Act have been in force. Whether and how labelling is required depends on the specific case and on the role of the provider or operator.

What is synthetic content?

‘Synthetic content’ is a collective term for content whose key components have been generated or modified using artificial intelligence. For example, a language model can generate a complete text, an image generator can create an image, or an AI system can synthetically replicate a real voice.
Even existing content can be classified as synthetic content if the AI processing substantially alters its presentation or message. Not every use of an AI-powered tool is sufficient for this. An automatic spell-checker is of a different nature to replacing a person in a photograph or generating a scene that never took place.

When is content considered synthetic?

There is no fully standardised definition of synthetic content for every technical and regulatory context. For practical classification purposes, therefore, the degree of AI-based generation or alteration is particularly relevant. If a significant part of the result has been generated, replaced or altered in meaning by AI, there is strong evidence that it constitutes synthetic content. If, on the other hand, the system merely assists with standard editing without substantially influencing the message or presentation, classification is less straightforward.

What counts as synthetic content – and what doesn’t?

ExampleSynthetic Content?Classification

ChatGPT creates a complete product description

Yes

The text is substantially generated by generative AI.

An image generator creates a product image

Yes

The image is created entirely synthetically.

AI replaces the background of a photo with a new scene

Usually yes

A substantial part of the image is altered using generative AI.

A real voice is cloned using AI

Yes

The audio is synthetically generated or reproduced.

AI inserts a person into a scene that never occurred

Yes

The depiction of reality is substantially altered.

Spell-checking software corrects typos

Generally no

The editing merely provides assistance without substantially changing the content.

A photo is cropped using conventional editing tools

No

No AI generation or substantial AI-based alteration takes place.

The crucial distinction therefore does not lie between ‘AI used’ and ‘no AI used’. What matters is whether AI substantially generates or alters the actual content.

Synthetic content, synthetic media, deepfakes or synthetic data?

TermMeaningRelationship to Synthetic ContenExample

Synthetic Content

AI-generated or substantially AI-modified digital content

Umbrella term used in this article

AI-generated product description

AI-generated Content

Content created by generative AI

Strong overlap

Article generated by an LLM

Synthetic Media

Synthetically generated or modified media

Sometimes used synonymously, sometimes more specifically for audiovisual media

AI-generated video

Deepfake

Realistic-looking AI-generated or manipulated representation

Subcategory

Artificial video of a real person

Synthetic Data

Artificially generated datasets

Different term and different use case

Simulated training or test data

The terminology is not fully standardised. Definitions may overlap depending on the scientific, technical or regulatory context.

  • ‘Synthetic media’ is not necessarily defined more narrowly than ‘synthetic content’. In some specialist literature, the term may also encompass text, images, audio, video and multimodal content.
  • ‘Synthetic Data’, on the other hand, must be clearly distinguished. Here, the focus is not on the publication of digital content, but rather, for example, on the generation of artificial data for training, simulation or testing.

What synthetic content is not by definition

Not every form of AI support

Standard corrections or minor technical adjustments are not, in themselves, sufficient to classify something as synthetic content.

Not necessarily a problem

‘Synthetic’ primarily describes the manner in which something is created or altered – not the quality, accuracy or legitimacy of the content.

Not quite a deepfake or synthetic data

Deepfakes are a specific type of synthetic content, whereas synthetic data refers to artificially generated datasets for other purposes.

What types of synthetic content are there?

  • AI-generated texts: Language models can generate product descriptions, summaries, translations or complete editorial drafts. The more the AI produces the actual content, the clearer it is that it should be classified as synthetic content.

  • AI-generated and manipulated images: Image generators can create new images or alter existing ones. If people, objects, backgrounds or scenes are substantially generated or replaced, the result may be classified as synthetic content.

  • Synthetic voices and audio: AI can generate or mimic voices, or automatically alter speech. Applications range from digital narrators and translations to voice cloning.

  • AI videos and deepfakes: Generative systems can produce individual sequences or entire videos. Deepfakes are a particularly relevant sub-category because they can depict real people in a deceptively realistic manner in situations that did not actually take place.

  • Multimodal content: Multimodal systems combine several media forms, such as text, images, audio and video, within a single production process. Synthetic content is therefore not restricted to any particular format.

Why is synthetic content used?

Synthetic content can speed up production processes and make them scalable. Texts, images and audiovisual content can be created, translated and adapted for different target audiences more quickly. Generative systems are therefore used in areas such as marketing, e-commerce, media production, software and internal communications.
Further opportunities arise through personalisation and the creation of creative variations. For example, an initial concept can be automatically adapted for different formats, markets or target audiences.
Speed, however, is not a measure of quality. The more content is produced automatically, the more important editorial review, clear lines of responsibility and transparent processes become.

What risks does synthetic content pose?

Synthetic content can be used for legitimate creative and commercial purposes. However, the same technical capability can also be used to make content appear credible, even though the events, people or statements depicted are not authentic.

Of particular relevance here are:

  • Disinformation and manipulation: Realistically generated images, voices or videos can convey a false impression of real events. The technical quality of a piece of content says nothing about whether its message is true.
  • Identity theft: The voices or faces of real people can be synthetically replicated. This raises questions regarding personal rights, reputation and deception.
  • Copyright and usage rights: There are also potential issues regarding copyright, usage rights and training or source material.

For organisations, therefore, it is not only relevant what can be technically generated, but also the legal, editorial and communicative conditions under which the content is used.

How can synthetic content be identified?

It is difficult to detect this reliably based solely on the visible or audible result. Typical artefacts, unusual phrasing or details that appear unnatural may be clues, but they do not constitute conclusive evidence.
This is why provenance – that is, the traceable origin and editing history of a piece of content – is becoming increasingly important. Technical approaches include, amongst other things, metadata, machine-readable tags, watermarks and cryptographically secured provenance information.

What are Content Credentials?

Content credentials are based on technical provenance standards and can document information about how a digital asset was created and modified. Depending on the specific implementation, details such as editing steps or information on the asset’s origin can be recorded in a traceable manner.
This approach therefore differs from a traditional AI content detector. Rather than merely estimating retrospectively whether content is likely to have been generated by AI, the aim is to make the creation and editing history technically traceable.

What are the limitations of AI content detectors?

In many cases, detectors merely provide probability assessments. Content can be edited, and generative models are constantly evolving.
Detection and provenance therefore address different questions: detection attempts to identify an origin, whilst provenance documents information about creation and modification.

Recognition and identification are not the same thing

A technical label or proof of provenance does not automatically equate to disclosure that is visible to humans. Machine-readable labelling, provenance and perceptible labelling fulfil different functions. This is particularly relevant to the transparency obligations under the EU AI Act.

Does synthetic content have to be labelled in Germany?

Synthetic content does not have to be labelled with the same visible indicator across the board in every use case. However, the transparency obligations set out in Article 50 of the EU AI Act will apply from 2 August 2026. These obligations vary depending on the type of system and content, as well as on whether a party is acting as a provider or operator of an AI system. To determine the classification, it is therefore first necessary to clarify: Who is using which AI system for what content?

Who is required to carry out checks under the EU AI Act?

SituationWhat is relevant?

Provider of a generative AI system

Certain generated or manipulated outputs must be technically marked so that they can be identified as AI-generated or AI-manipulated.

Deployer publishes a deepfake

A clear and perceptible disclosure to natural persons must be considered.

AI-generated or AI-manipulated text provides information on matters of public interest

A disclosure obligation may be relevant for the deployer.

The text has been subject to human review or editorial control and editorial responsibility exists

This is an important statutory exception to the disclosure requirement for certain texts.

AI is used only for standard assistive editing

Depending on the specific case, an exception to certain marking requirements may apply.

When is a technical label sufficient – and when is a visible notice required?

Providers of generative AI systems must, in particular, ensure that outputs covered by the AI Act are machine-readable. Operators may also have their own disclosure obligations:

  • In the case of deepfakes, the information must be clear and perceptible to humans.
  • A machine-readable label from the provider embedded within the content does not automatically replace this disclosure obligation.

In the case of certain AI-generated or manipulated texts on matters of public interest, it also depends on whether human verification or editorial control takes place and who bears editorial responsibility for the publication. The specific legal assessment depends on the individual case. In relevant publication processes, therefore, the focus should not be solely on whether AI was used, but rather on:

  • what role the company plays
  • what type of content is involved
  • what form of transparency is required

Measures for verifying synthetic content

Check 1: Is my content synthetic?

A few key questions are sufficient for content classification:

  1. Was the content generated entirely by AI?
  2. Did AI generate or replace significant parts of the content?
  3. Have the message, presentation, tone, characters or scenes been substantially altered by AI?
  4. Did the editing go beyond standard, supportive corrections?

If at least one of the first three points applies and the editing is not merely standard, supportive editing, there is a strong case for classifying it as synthetic content. This checklist addresses only the nature of the content; it does not indicate whether there is a labelling requirement.

Check 2: Do I need to check labelling or transparency?

The second step concerns the specific application:

  1. Does your company act as the provider or operator of the AI system in use?
  2. Is it a deepfake?
  3. Is AI-generated or manipulated text being published to provide information on matters of public interest?
  4. Is there a qualified human review or editorial oversight in place?
  5. Who bears editorial responsibility for the publication?
  6. Are the necessary machine-readable tags or provenance information available?
  7. Is additional disclosure visible to humans also required?

The classification as ‘synthetic content’ and the regulatory labelling assessment are two separate decisions. A fully labelled AI image is no more or less synthetic than the same image without provenance information.

Synthetic Content in SEO and Content Marketing

When it comes to SEO, it is not solely a question of whether content has been created entirely by humans or with the help of AI. What matters is whether the content fulfils the search intent, is factually sound and offers users genuine informational value.
With automated production, therefore, the demands placed on editorial review, source verification and quality management are increasing in particular. Large volumes of content can be generated more quickly using AI. However, this does not automatically ensure relevance and reliability. For businesses, AI-generated content is therefore primarily a question of process: where does automation make sense – and where is human expertise still required?

Conclusion: What matters is how it is produced, where it comes from and how it is used

Synthetic content primarily describes how digital content has been created or significantly altered. Mere AI support for standard corrections is not automatically sufficient for this. For practical assessment, three separate questions then follow: Is the content trustworthy? Is its origin traceable? And are there any transparency or labelling requirements in its specific use?
It is precisely this distinction that prevents misconceptions. Not all synthetic content is problematic, and not all AI-generated content requires the same visible label. Context, the nature of the content and accountability remain crucial.

Frequently Asked Questions about Synthetic Content

Is a ChatGPT text considered synthetic content?

Yes, if ChatGPT generated the text in its entirety or in substantial parts. If the system is used solely for spell-checking or minor linguistic adjustments, the classification is less clear-cut.

Are AI-generated images synthetic content?

Yes. Fully generated images are typical examples of synthetic content. Existing images may also fall under this category if significant parts of them are generated by AI or their content is altered.

What is the difference between synthetic content and synthetic media?

The terms are not used consistently across the board. In this article, “synthetic content” is used as an umbrella term for digital content that is generated by AI or significantly altered by AI. “Synthetic media” can be used in a similar way, but is sometimes more strongly associated with image, audio, and video content.

Is every deepfake synthetic content?

Yes, deepfakes fall under the category of synthetically generated or manipulated content. Conversely, however, not all synthetic content is a deepfake.

Does synthetic content have to be labeled?

It's not a one-size-fits-all situation. Under Article 50 of the EU AI Act, the requirements vary depending on, among other things, whether an actor is a provider or an operator, and whether, for example, a deepfake or AI-generated or manipulated text relates to matters of public interest.

Can synthetic content be reliably detected?

Not always based solely on appearance. Detectors can provide clues, while provenance mechanisms and content credentials can offer additional information about the content’s origin and editing history.

What are Content Credentials?

Content credentials are a technical approach to documenting the origin and revision history of digital content. They are intended to make it easier to trace how a piece of content was created or modified.