Synthetic Content
Synthetic content refers to digital content that has been generated or substantially altered in terms of its content, either entirely or to a significant extent, by artificial intelligence. This includes text, images, audio, videos and multimodal content. The key factor is whether AI has a substantial influence on the result – not whether it merely assists with minor corrections.
With the increasing prevalence of generative AI, the term is gaining in significance. Synthetic content can encompass both entirely newly generated content and content that has been substantially altered by AI. Depending on the context, questions regarding provenance, authenticity and labelling also play a role.
Key facts about synthetic content at a glance
- Synthetic content is digital content that is entirely or predominantly generated or modified by AI.
- A simple spell-check or standard image editing does not automatically make content synthetic content.
- The terms ‘synthetic content’, ‘synthetic media’ and ‘AI-generated content’ overlap and are not used consistently across the board.
- Deepfakes are a subset of synthetic content; synthetic data, on the other hand, refers to artificially generated datasets.
- Synthetic content is not inherently problematic. Risks arise in particular from deception, disinformation, identity theft and a lack of transparency.
- Since 2 August 2026, the transparency requirements set out in Article 50 of the EU AI Act have been in force. Whether and how labelling is required depends on the specific case and on the role of the provider or operator.
What is synthetic content?
‘Synthetic content’ is a collective term for content whose key components have been generated or modified using artificial intelligence. For example, a language model can generate a complete text, an image generator can create an image, or an AI system can synthetically replicate a real voice.
Even existing content can be classified as synthetic content if the AI processing substantially alters its presentation or message. Not every use of an AI-powered tool is sufficient for this. An automatic spell-checker is of a different nature to replacing a person in a photograph or generating a scene that never took place.
When is content considered synthetic?
There is no fully standardised definition of synthetic content for every technical and regulatory context. For practical classification purposes, therefore, the degree of AI-based generation or alteration is particularly relevant. If a significant part of the result has been generated, replaced or altered in meaning by AI, there is strong evidence that it constitutes synthetic content. If, on the other hand, the system merely assists with standard editing without substantially influencing the message or presentation, classification is less straightforward.
Synthetic content, synthetic media, deepfakes or synthetic data?
| Term | Meaning | Relationship to Synthetic Conten | Example |
|---|---|---|---|
Synthetic Content | AI-generated or substantially AI-modified digital content | Umbrella term used in this article | AI-generated product description |
AI-generated Content | Content created by generative AI | Strong overlap | Article generated by an LLM |
Synthetic Media | Synthetically generated or modified media | Sometimes used synonymously, sometimes more specifically for audiovisual media | AI-generated video |
Deepfake | Realistic-looking AI-generated or manipulated representation | Subcategory | Artificial video of a real person |
Synthetic Data | Artificially generated datasets | Different term and different use case | Simulated training or test data |
The terminology is not fully standardised. Definitions may overlap depending on the scientific, technical or regulatory context.
- ‘Synthetic media’ is not necessarily defined more narrowly than ‘synthetic content’. In some specialist literature, the term may also encompass text, images, audio, video and multimodal content.
- ‘Synthetic Data’, on the other hand, must be clearly distinguished. Here, the focus is not on the publication of digital content, but rather, for example, on the generation of artificial data for training, simulation or testing.
What synthetic content is not by definition
Standard corrections or minor technical adjustments are not, in themselves, sufficient to classify something as synthetic content.
‘Synthetic’ primarily describes the manner in which something is created or altered – not the quality, accuracy or legitimacy of the content.
Deepfakes are a specific type of synthetic content, whereas synthetic data refers to artificially generated datasets for other purposes.
What types of synthetic content are there?
AI-generated texts: Language models can generate product descriptions, summaries, translations or complete editorial drafts. The more the AI produces the actual content, the clearer it is that it should be classified as synthetic content.
AI-generated and manipulated images: Image generators can create new images or alter existing ones. If people, objects, backgrounds or scenes are substantially generated or replaced, the result may be classified as synthetic content.
Synthetic voices and audio: AI can generate or mimic voices, or automatically alter speech. Applications range from digital narrators and translations to voice cloning.
AI videos and deepfakes: Generative systems can produce individual sequences or entire videos. Deepfakes are a particularly relevant sub-category because they can depict real people in a deceptively realistic manner in situations that did not actually take place.
Multimodal content: Multimodal systems combine several media forms, such as text, images, audio and video, within a single production process. Synthetic content is therefore not restricted to any particular format.
Why is synthetic content used?
Synthetic content is used to create digital content more quickly, on a larger scale or with greater variety. Typical areas of application include marketing, e-commerce, media production, translation, software and internal communication. Generative AI, for example, enables the automated creation of various text, image, audio or video variations, as well as their adaptation to different target audiences or markets. However, technical scalability says nothing about the quality or accuracy of the content generated.
What risks does synthetic content pose?
Synthetic content can be used for legitimate creative and commercial purposes. However, the same technical capability can also be used to make content appear credible, even though the events, people or statements depicted are not authentic.
Of particular relevance here are:
- Disinformation and manipulation: Realistically generated images, voices or videos can convey a false impression of real events. The technical quality of a piece of content says nothing about whether its message is true.
- Identity theft: The voices or faces of real people can be synthetically replicated. This raises questions regarding personal rights, reputation and deception.
- Copyright and usage rights: There are also potential issues regarding copyright, usage rights and training or source material.
For organisations, therefore, it is not only relevant what can be technically generated, but also the legal, editorial and communicative conditions under which the content is used.
How can synthetic content be identified?
It is difficult to detect this reliably based solely on the visible or audible result. Typical artefacts, unusual phrasing or details that appear unnatural may be clues, but they do not constitute conclusive evidence.
This is why provenance – that is, the traceable origin and editing history of a piece of content – is becoming increasingly important. Technical approaches include, amongst other things, metadata, machine-readable tags, watermarks and cryptographically secured provenance information.
What are Content Credentials?
Content credentials are based on technical provenance standards and can document information about how a digital asset was created and modified. Depending on the specific implementation, details such as editing steps or information on the asset’s origin can be recorded in a traceable manner.
This approach therefore differs from a traditional AI content detector. Rather than merely estimating retrospectively whether content is likely to have been generated by AI, the aim is to make the creation and editing history technically traceable.
What are the limitations of AI content detectors?
In many cases, detectors merely provide probability assessments. Content can be edited, and generative models are constantly evolving.
Detection and provenance therefore address different questions: detection attempts to identify an origin, whilst provenance documents information about creation and modification.
Recognition and identification are not the same thing
A technical label or proof of provenance does not automatically equate to disclosure that is visible to humans. Machine-readable labelling, provenance and perceptible labelling fulfil different functions. This is particularly relevant to the transparency obligations under the EU AI Act.
Does synthetic content have to be labelled in Germany?
Synthetic content does not have to be labelled with the same visible indicator across the board in every use case. However, the transparency obligations set out in Article 50 of the EU AI Act will apply from 2 August 2026. These obligations vary depending on the type of system and content, as well as on whether a party is acting as a provider or operator of an AI system. To determine the classification, it is therefore first necessary to clarify: Who is using which AI system for what content?
Who is required to carry out checks under the EU AI Act?
| Situation | What is relevant? |
|---|---|
Provider of a generative AI system | Certain generated or manipulated outputs must be technically marked so that they can be identified as AI-generated or AI-manipulated. |
Deployer publishes a deepfake | A clear and perceptible disclosure to natural persons must be considered. |
AI-generated or AI-manipulated text provides information on matters of public interest | A disclosure obligation may be relevant for the deployer. |
The text has been subject to human review or editorial control and editorial responsibility exists | This is an important statutory exception to the disclosure requirement for certain texts. |
AI is used only for standard assistive editing | Depending on the specific case, an exception to certain marking requirements may apply. |
Conclusion: What matters is how it is produced, where it comes from and how it is used
Synthetic content primarily describes how digital content has been created or significantly altered. Mere AI support for standard corrections is not automatically sufficient for this. For practical assessment, three separate questions then follow: Is the content trustworthy? Is its origin traceable? And are there any transparency or labelling requirements in its specific use?
It is precisely this distinction that prevents misconceptions. Not all synthetic content is problematic, and not all AI-generated content requires the same visible label. Context, the nature of the content and accountability remain crucial.
Frequently Asked Questions about Synthetic Content
Yes, if ChatGPT generated the text in its entirety or in substantial parts. If the system is used solely for spell-checking or minor linguistic adjustments, the classification is less clear-cut.
Yes. Fully generated images are typical examples of synthetic content. Existing images may also fall under this category if significant parts of them are generated by AI or their content is altered.
The terms are not used consistently across the board. In this article, “synthetic content” is used as an umbrella term for digital content that is generated by AI or significantly altered by AI. “Synthetic media” can be used in a similar way, but is sometimes more strongly associated with image, audio, and video content.
Yes, deepfakes fall under the category of synthetically generated or manipulated content. Conversely, however, not all synthetic content is a deepfake.
It's not a one-size-fits-all situation. Under Article 50 of the EU AI Act, the requirements vary depending on, among other things, whether an actor is a provider or an operator, and whether, for example, a deepfake or AI-generated or manipulated text relates to matters of public interest.
Not always based solely on appearance. Detectors can provide clues, while provenance mechanisms and content credentials can offer additional information about the content’s origin and editing history.
Content credentials are a technical approach to documenting the origin and revision history of digital content. They are intended to make it easier to trace how a piece of content was created or modified.











