GAN vs Diffusion Models: How AI Detectors Tell Them Apart

GAN-generated images (StyleGAN, BigGAN) leave a specific, detectable fingerprint in the frequency domain. Diffusion model images (Midjourney, Stable Diffusion, DALL-E, Flux) leave different and subtler artifacts. Detection works differently for each — here's the technical breakdown, without the jargon where possible.

Quick Answer

GAN images leave periodic grid artifacts in the frequency domain — detectable by Discrete Fourier Transform analysis — from the convolution-upsampling process. Diffusion images don't have a single strong fingerprint; detectors instead analyze the characteristic noise distribution from the denoising process, unnatural high-frequency smoothness, and VAE compression signatures. GANs are easier to detect; diffusion models require more complex multi-signal classifiers.

How GANs generate images — and why that leaves a detectable mark

A GAN (Generative Adversarial Network) generates images through a series of upsampling convolutions. The generator starts with a small noise vector and progressively scales it up to full resolution through convolutional layers. At each scale, it applies learned filters — convolutions — that multiply pixel values according to a fixed kernel pattern.

This convolution process creates a characteristic side effect: because the same convolutional kernel is applied at fixed intervals across the image, the resulting image has periodic patterns in the frequency domain. The patterns are invisible to the eye — they represent tiny, repeating variations in pixel values at fixed spatial frequencies. But when you apply a Discrete Fourier Transform (DFT) to convert the image from spatial representation to frequency representation, the GAN fingerprint shows up as peaks at specific intervals.

Real photographs don't have these periodic frequency peaks. A real photo has complex, non-periodic noise from the camera sensor and optical system. The presence of periodic peaks at intervals corresponding to the GAN's architecture is a reliable indicator of GAN origin.

Why GAN detection generalizes across architectures

StyleGAN, BigGAN, ProGAN, and most other GAN architectures all use some form of convolution-upsampling. The frequency peaks shift position depending on the specific architecture and the number of upsampling stages, but the fact that periodic peaks exist is consistent. A classifier trained on frequency maps from multiple GAN architectures generalizes reasonably well to new GAN architectures — they all leave a recognizable family of marks.

How diffusion models generate images — and why detection is harder

Diffusion models (Stable Diffusion, Midjourney, DALL-E, Flux, Imagen) work entirely differently from GANs. Instead of progressive convolution, they start with random noise and iteratively remove noise step by step, guided by a text prompt or conditioning signal. Each step is computed by a neural network (the denoising model, often a U-Net) that predicts what noise to remove.

After denoising, the image is decoded from a compressed latent space by a VAE (variational autoencoder) decoder. The entire process produces no periodic frequency grid — the denoising steps don't apply the same operation repeatedly at fixed spatial intervals the way GAN convolutions do.

Without a GAN-style frequency fingerprint, detectors rely on several weaker, overlapping signals:

Noise distribution signature

Real camera sensor noise is heterogeneous — noisier in dark regions (high ISO) and smoother in bright regions. Diffusion model output has a different noise distribution: more spatially uniform because the denoising process applies consistent smoothing across regions without the physical constraints of a sensor. Statistical analysis of residual noise distribution across image regions produces a signal that classifiers can learn.

High-frequency smoothness anomaly

Diffusion images often have unnaturally smooth high-frequency content in specific frequency bands. Real photographs have fine-grained texture noise at every scale from the optical system. Diffusion images lack this at specific sub-pixel frequencies — not because they look smooth, but because the statistical distribution of very-high-frequency content differs from camera noise characteristics.

VAE decoder compression artifacts

Most diffusion models encode and decode images through a VAE. The VAE decoder introduces characteristic compression artifacts in the pixel domain — different from JPEG artifacts — that are specific to the decoder architecture. Models trained on images from specific diffusion pipelines (Stable Diffusion's KL-F8 VAE, SDXL's VAE) can learn these decoder-specific fingerprints.

Head-to-head: GAN vs diffusion detection properties

GAN images Diffusion model images
Primary fingerprint Periodic frequency peaks (DFT) Noise distribution + VAE artifacts
Fingerprint strength Strong, single signal Weak, multi-signal
Cross-architecture generalization Good — all GANs use convolution Moderate — VAE signatures vary by model
Survives JPEG compression? Partially — frequency peaks degrade Poorly — noise signatures are fragile
Relative detection difficulty Easier Harder
Dominant generator type (2026) Declining — diffusion now dominant Dominant — Midjourney, SD, DALL-E, Flux

The shift from GAN-dominated to diffusion-dominated AI image generation (roughly 2022-2024) was a genuine setback for detection: the stronger, more reliable GAN fingerprint is less common now, and the harder-to-detect diffusion artifacts are more common. Modern detection models like Sightengine's genai API are trained on large, diverse datasets of diffusion outputs specifically to address this.

Why GAN detection generalizes better — and why diffusion detection is catching up

The key insight is that GAN detection became a well-solved problem because of the strong, consistent convolution artifact. A detector trained on StyleGAN generalizes to ProGAN, VGAN, and other architectures because they all share the underlying convolution mechanism. The training data doesn't need to cover every possible GAN — the signal is architectural, not model-specific.

Diffusion detection doesn't have this advantage. The VAE decoder signature for Stable Diffusion is different from the one for SDXL, which is different from the one for Flux, which is different from the one for DALL-E. Each pipeline produces somewhat different artifacts. Diffusion detectors need diverse training data covering many generator pipelines and must be updated as new models emerge.

The trajectory is improving: as the major diffusion pipelines stabilize (SD, SDXL, Flux, DALL-E 3, Midjourney v6/v7), detection models built on comprehensive training sets of their outputs are achieving accuracy comparable to GAN detectors on the specific generators they've been trained on. The challenge remains with genuinely novel models or significantly fine-tuned versions of existing models.

Common questions

How do AI detectors tell apart GAN images and diffusion model images?

GAN detection uses Discrete Fourier Transform frequency analysis to find periodic grid artifacts from the convolution-upsampling process — peaks at fixed intervals that real photos don't have. Diffusion detection is harder and uses overlapping signals: the noise distribution from denoising, high-frequency smoothness anomalies, and VAE decoder compression artifacts. Modern detection models handle both generator types by training on large diverse datasets and combining multiple signals.

Which is easier to detect — GAN images or diffusion model images?

GAN images. They leave a strong, consistent frequency fingerprint from the convolution-upsampling process that generalizes across architectures. Diffusion images rely on weaker, overlapping signals that don't generalize as cleanly and are more sensitive to post-processing. The shift from GAN-dominated to diffusion-dominated AI image generation (2022-2024) made detection harder, though modern models have substantially closed the gap on major diffusion pipelines.

Can AI detectors identify which specific model generated an image?

Partially. Illuminarty attempts generator identification (Midjourney, Stable Diffusion, DALL-E, etc.) based on model-specific frequency and noise signatures. This works reasonably for images from well-known generators without post-processing. It's less reliable for fine-tuned models or new generators the detector hasn't been trained on. FauxSpy's genai model focuses on the binary AI-or-real classification rather than generator identification.

Related reading

Try the detector — covers both GANs and diffusion models

🕵️ Add to Chrome — Free 🦊 Add to Firefox — Free