AI Development and the New Regulations Being Introduced to Combat AI-Generated Content

Artificial intelligence has developed rapidly from systems that could perform simple automated tasks into powerful generative tools capable of producing realistic text, images, audio, video and computer code. Services such as ChatGPT, Claude, Gemini and other generative AI systems can now create content within seconds that would previously have required hours of human work. While this development has created major opportunities in education, business, entertainment and research, it has also created a growing problem: it is becoming increasingly difficult to determine whether something was created by a person or by an AI system.

The problem is no longer limited to obviously artificial images or poorly written computer-generated text. Modern AI can produce realistic photographs of people who do not exist, imitate voices, generate convincing videos and create large quantities of written material. This has increased concerns about misinformation, fraud, impersonation, election manipulation, academic cheating and the mass production of misleading online content.

Governments and technology companies are therefore moving toward a new approach: rather than relying entirely on people to detect AI-generated material after it has been published, AI systems themselves may increasingly be required to identify the content they produce.

How AI-generated content is becoming harder to identify

Earlier generations of AI-generated content often contained obvious weaknesses. AI images could have distorted hands, strange facial features or unrealistic objects. AI-written text could also sound repetitive or unnatural.

Those weaknesses are becoming less reliable as AI models improve. Generative systems are now trained on enormous amounts of data and can reproduce human language, artistic styles, voices and visual characteristics with increasing accuracy.

This creates an important problem for online platforms. A photograph, article or video can be distributed without any obvious indication that AI was involved in its creation.

AI also makes it possible to produce content at a scale that would be difficult for humans to achieve. A person can use an AI system to generate thousands of articles, comments, social-media posts or images. This means the concern is not simply whether one piece of content was generated by AI, but whether automated systems could flood the information environment with synthetic material.

The move toward AI watermarking

One of the main technologies being developed to address this problem is digital watermarking.

A watermark does not necessarily have to be visible. Instead, information can be embedded into AI-generated content in a way that allows specialised software to identify it later.

For text, an AI company can alter the statistical pattern of word or token selection while keeping the text readable and maintaining its meaning. A detector can then analyse the text for the statistical pattern associated with the AI system.

For images, audio and video, other forms of watermarking can be used. Another important approach is content provenance, which records information about where a digital file came from and how it was created or modified.

The European Union is increasingly incorporating these technologies into regulation. Under Article 50 of the EU AI Act, providers of systems that generate synthetic audio, images, video or text must use machine-readable marking so that AI-generated or manipulated content can be detected, subject to the scope and exceptions defined by the law. These transparency obligations began applying on 2 August 2026.

The European Union’s approach

The European Union has taken one of the most comprehensive regulatory approaches to AI.

The EU AI Act does not simply attempt to ban AI-generated content. Instead, it establishes transparency requirements designed to help people understand when they are dealing with AI or viewing AI-generated or manipulated material.

Under the rules, AI providers must design systems so that people are informed when they are directly interacting with an AI system, unless this is obvious. Providers of generative AI systems also have obligations concerning machine-readable marking of synthetic content.

There are additional requirements for organisations deploying AI. Deepfake content must be clearly disclosed to people, while AI-generated or manipulated text concerning matters of public interest can also require disclosure when it has been published without human review or editorial control.

The EU has also developed a Code of Practice on Transparency of AI-Generated Content to provide technical and practical guidance for complying with these requirements. The code addresses marking and detection of AI-generated or manipulated text, images, audio and video.

This represents a significant change in how AI-generated content is regulated. Instead of expecting internet users to determine whether something looks artificial, the technology producing the material can be required to leave information that allows its origin to be investigated.

AI companies are already implementing watermarking

Regulation is already influencing the design of commercial AI systems.

Anthropic, for example, announced that Claude would use invisible machine-readable watermarking for AI-generated text and provenance information for images. The company has said that the measures are being implemented in response to the EU’s transparency requirements.

This is particularly significant because watermarking can potentially be implemented at the point where content is generated. Instead of adding a label after an article, image or video has been uploaded to a website, the AI model can create the identification signal at the source.

If this approach becomes widespread, AI-generated content could increasingly carry information about its origin throughout its digital life.

Why watermarking is not a perfect solution

Watermarking has important limitations.

AI-generated text can be copied, translated, paraphrased or rewritten. Images can be cropped, compressed or modified. Videos can be edited and audio can be manipulated.

Recent experiments have already demonstrated that some AI text watermarks can be weakened or removed through rewriting and other techniques. This means a watermark should not necessarily be treated as absolute proof that a particular piece of text was created by AI.

There is also another problem: absence of a watermark does not necessarily prove that something was written by a human.

Someone could generate text using an AI system that does not watermark its output. Alternatively, a watermark could be removed during editing or transformation.

This means future AI detection is likely to rely on several technologies rather than one universal detector.

AI detection versus content provenance

A major distinction is emerging between AI detection and content provenance.

AI detection attempts to examine an existing piece of content and determine whether AI was probably involved in producing it.

Content provenance takes a different approach. It attempts to maintain a record of the content’s origin and subsequent modifications.

For example, an image could contain information showing that it was generated by an AI system, which software created it and whether it was subsequently edited.

This approach can be more useful than attempting to determine AI authorship purely from the final appearance of a file. Standards such as C2PA are being developed around this broader concept of digital content provenance.

Regulation will probably expand beyond deepfakes

The first major concern was realistic fake videos and images of public figures. However, regulators are increasingly concerned about a much wider category of synthetic content.

AI can be used to create fake news articles, fabricated evidence, impersonated voices, fake customer reviews and large volumes of political material.

The European Commission specifically identifies misinformation, manipulation, fraud, impersonation and consumer deception as risks associated with increasingly difficult-to-distinguish AI content.

This means future regulation is likely to focus not simply on whether content was generated by AI, but also how the content is used and whether it could deceive people.

A computer-generated image used in a fictional film is very different from an AI-generated image falsely presented as evidence of a real event.

The future of AI-generated content

The development of AI is unlikely to stop because of these regulations. Instead, the relationship between AI development and regulation is likely to evolve together.

AI companies will continue improving the quality of generated content, while regulators will attempt to develop methods for maintaining transparency and accountability.

This could eventually result in a digital environment where AI-generated material is accompanied by several layers of information: a visible disclosure for users, an invisible machine-readable marker for automated systems and provenance information describing where the content originated.

The biggest challenge will be keeping these systems effective as AI becomes more sophisticated.

Watermarks can potentially be removed. Detection systems can produce false positives. Provenance information can be lost when files are copied between platforms. At the same time, overly aggressive detection could wrongly classify legitimate human writing as AI-generated.

For this reason, the future of AI regulation is unlikely to depend on a single “AI detector.” Instead, it is moving toward a combination of watermarking, machine-readable labels, provenance records, platform policies, human review and legal disclosure requirements.

The EU’s implementation of Article 50 is an important early example of this approach. Its transparency requirements are now being applied, while AI companies are developing technical systems to comply with them.

Ultimately, the objective is not necessarily to prevent people from using AI to create content. The more practical goal is to make AI involvement more transparent, particularly when synthetic content could otherwise be mistaken for authentic human-created material. As AI becomes increasingly capable of producing convincing content, knowing where information came from may become almost as important as the information itself.