How AI Image Detection Actually Works

Most people assume AI image detectors look for obvious visual glitches — extra fingers, blurry ears, artifacts around hair. That's not what they do. The actual detection happens at a level below what your eyes can see: invisible statistical patterns in pixel data that every AI generator leaves behind. This page explains the real mechanism, plainly.

Quick Answer

AI image detection works by finding invisible statistical artifacts in pixel data — frequency-domain fingerprints left behind by the AI generator's mathematical process — that real photographs don't have. Models trained on thousands of known AI-generated images learn to recognize these signatures and return a probability score. No detector looks for visual glitches; they look for patterns human eyes can't see.

The core idea: every AI generator leaves a fingerprint

When an AI model generates an image, it doesn't paint pixels the way a photographer captures them. It runs a mathematical process — convolution in GANs, iterative denoising in diffusion models — that produces characteristic statistical patterns in the pixel array. These patterns are invisible to the eye but show up clearly when you apply the right analysis.

Think of it like a manufacturing process. A photo printed on an inkjet printer and a photo printed on a laser printer look the same to a casual viewer. But under a microscope, the inkjet print has a characteristic dot pattern and the laser print has a different characteristic toner pattern. An expert with a microscope can tell them apart not by how the photo looks, but by the physical signature of the process that made it.

AI detectors are the "microscope." They don't look at whether the photo seems too perfect or the lighting seems off — those visual tells are unreliable. They look for the process signature that the AI generator left in the data.

How GAN detection works: frequency domain analysis

The first generation of AI image detectors was built specifically for GANs (Generative Adversarial Networks) — the technology behind tools like StyleGAN that generate photorealistic faces. GANs work through a series of convolutional operations that upsample a low-resolution noise vector into a full-resolution image. That convolution process creates repeating patterns at fixed pixel intervals — a grid-like artifact in the frequency domain.

What is a GAN fingerprint?

A GAN fingerprint is the characteristic frequency artifact produced by the upsampling convolutions in the generator network. To find it, detectors apply a technique called the Discrete Fourier Transform (DFT) to the image — this converts the pixel data from a spatial representation (where pixels are) into a frequency representation (how often patterns repeat across the image). In the frequency domain, GAN-generated images show distinctive peaks at intervals that correspond to the generator's architecture. Real photographs don't have these periodic peaks.

The peaks are small — they represent a tiny fraction of the total pixel energy. But they're consistent. Because all GAN generators using similar upsampling architectures leave similar-shaped peaks, a classifier trained on frequency maps can identify GAN-generated images even when the visual output looks completely realistic. A 2019 paper by Zhang et al. (CVPR 2019) first demonstrated this approach and showed it could achieve over 99% accuracy on a controlled dataset of GAN images.

Why GAN detection is easier than diffusion detection

GAN images have a predictable, architecture-specific fingerprint. Once you've trained a detector on images from a specific GAN architecture (StyleGAN, BigGAN, etc.), it generalizes reasonably well to other GAN architectures because the convolution-upsampling process is structurally similar across them. The frequency peaks shift position depending on the architecture, but the fact that peaks exist is consistent.

How diffusion model detection works: noise pattern analysis

Modern AI image generators — Midjourney, DALL-E, Stable Diffusion, Imagen, Flux — use diffusion models rather than GANs. Diffusion is a different process: start with random noise, then iteratively remove noise step-by-step guided by a text prompt or conditioning signal until the image emerges. This process doesn't produce the same fixed-frequency artifacts as GAN convolution, which means frequency-domain GAN detectors don't transfer reliably to diffusion images.

What diffusion detectors look for instead

Diffusion model detection relies on several overlapping signals rather than a single frequency fingerprint:

Denoising noise distribution

The iterative denoising process in diffusion models produces a characteristic statistical distribution of residual noise across the image — more uniform than real camera sensor noise, which is heterogeneous (noisier in dark areas, less noisy in bright areas). Detectors can quantify this by analyzing the local noise variance pattern across image regions.

High-frequency smoothness

Diffusion models tend to produce images that are smoother at fine scales than photographs of equivalent content. Real photos have fine-grained texture noise from the camera sensor and optical system at every scale. Diffusion images often lack this fine-scale variation in specific frequency bands — not because they look smooth, but because the statistical distribution of very-high-frequency content is different from real images.

Cross-region consistency artifacts

Real photos have natural optical relationships between regions — the depth of field falloff, the light source consistency, the lens distortion at edges — that constrain how different parts of the image relate to each other. Diffusion images generate each region more independently, sometimes producing subtle statistical inconsistencies in how light behaves across the image that a classifier can learn to detect.

Generator-specific compression artifacts

Many diffusion models save their output in a specific format or run through a specific VAE (variational autoencoder) decoder that introduces characteristic compression patterns. Detectors trained on large datasets of specific generator outputs can learn these format-specific signatures even when they're subtle.

Because diffusion detection depends on multiple overlapping signals rather than a single strong fingerprint, it's generally harder than GAN detection — and accuracy numbers reflect this. The advance of the field has been in building classifiers that combine all of these signals together, using neural networks that learn which combination of features best predicts AI origin across diverse generator architectures.

How deepfake detection works: it's different from AI image detection

Deepfake detection is related to but distinct from AI image detection. A deepfake typically involves a real photo or video with a face surgically replaced or manipulated — not a fully AI-generated image. This means the detection problem is different: instead of finding a generator fingerprint across the whole image, the detector looks for localized inconsistencies at the boundary between the real and manipulated regions.

Face-swap boundary analysis

The edges of a replaced face often have blending artifacts — statistical noise patterns that differ between the pasted region and the surrounding original image. Detectors segment the face region, analyze the edge quality, and compare noise characteristics between the face and background.

Biometric inconsistency

Face-swap algorithms sometimes introduce subtle biometric inconsistencies: eye gaze direction doesn't match the geometry of the 3D scene, pupil reflections are inconsistent with the light source, facial asymmetry is slightly different on one frame than another in a video. Video deepfake detectors can track these across frames.

Temporal coherence (video only)

In video deepfakes, the swapped face often flickers slightly between frames — the face-swap algorithm processes each frame somewhat independently, so subtle variations accumulate in ways that look unnatural in motion. Video deepfake detectors specifically look for these frame-to-frame inconsistencies.

FauxSpy uses different detection models for these two problems: Sightengine's genai and type models for AI image detection on the free tier, and the deepfake model (which powers Pro) for face manipulation and video deepfake analysis. The distinction matters: a fully AI-generated face scores high on AI image detection; a deepfake of a real person's face in a real photo scores high on deepfake detection; a photo of a real person scores low on both.

Why detectors can be fooled — and what the numbers actually mean

The honest answer: AI image detection accuracy is conditional. On standard test sets — images from known AI generators with no post-processing — the leading tools in 2026 benchmark at 88-98% accuracy. Hive Moderation reports 89-98% depending on the test; Illuminarty approximately 91%; FauxSpy (Sightengine genai) approximately 88%. False positive rates — real photos incorrectly flagged as AI — run approximately 3% for leading tools under controlled test conditions.

Those numbers drop significantly in the real world for several reasons:

What degrades accuracy Why it works Effect on detection
JPEG compression (below quality 90) Destroys high-frequency artifacts that detectors key on Moderate degradation
Screenshot-then-crop Applies re-compression and resampling in one step Moderate degradation
Adding Gaussian noise Obscures generator noise pattern with random noise High degradation
Non-integer resize (e.g. 87%) Interpolation shifts frequency artifacts to unpredicted positions Moderate degradation
Adversarial post-processing (targeted) Specifically optimized to fool a known detector Severe degradation
Heavy photo editing (crop, filters, retouch) Overwrites generator signature with editing software artifacts Moderate degradation

Research from Imagera in 2026 found that against images specifically optimized to evade detection, no commercial detector tested exceeded 45-60% accuracy — compared to 88-98% on unmodified AI images. This gap is the central challenge for the field: the same techniques used to improve image quality (compression, denoising, upscaling) also happen to strip out the signals detectors rely on.

This is why FauxSpy returns an Inconclusive result when confidence falls below a threshold rather than forcing a binary verdict. A low-confidence "probably real" verdict is less useful — and potentially more harmful — than an honest "we can't tell." Accuracy on unmodified images is high; the Inconclusive category handles the edges where the evidence is genuinely ambiguous.

What C2PA (Content Credentials) adds — and what it doesn't

C2PA (Coalition for Content Provenance and Authenticity) is a complementary approach to AI image detection that works fundamentally differently: instead of analyzing the image after the fact, it embeds a cryptographically signed provenance record at the moment of generation or capture. OpenAI embeds C2PA metadata in DALL-E and ChatGPT images; Adobe embeds it in Firefly-generated content; major camera manufacturers are adding it to hardware.

C2PA is promising but has a significant real-world limitation: metadata gets stripped when images are uploaded to most social platforms. Instagram, X/Twitter, Facebook, Reddit, and WhatsApp all strip or ignore EXIF and IPTC metadata during upload processing — which is also where C2PA data lives. An image with intact C2PA credentials is provably from a specific source. The same image shared through a typical social media workflow has no credentials at all, even if it was originally credentialed at the source.

This means C2PA and statistical AI image detection are complementary, not competing, approaches:

  • C2PA excels at — verifying images shared through credentialing-aware workflows (journalism, legal proceedings, authenticated media) where the chain of custody is intact
  • Statistical detection excels at — analyzing images in the wild, on social media, in profile photos, in marketplace listings, where metadata has been stripped or was never present

For most consumer use cases — checking a profile photo, a news image, a marketplace listing — the image has passed through at least one platform that strips metadata, so C2PA provides no useful signal. Statistical detection is the relevant tool.

How FauxSpy's detection specifically works

FauxSpy's detection is powered by Sightengine's API — a computer vision API that provides specialized models for AI-generated content detection. This is the same API used directly by developers building moderation systems; FauxSpy wraps it in a right-click browser interface that requires no technical setup.

Free tier: Uses Sightengine's genai model (detects AI-generated images across GAN and diffusion generators) and the type model (classifies content type). Returns a probability score from 0.0 to 1.0 for AI origin. FauxSpy translates scores above a threshold into "Likely AI," scores below into "Likely Real," and borderline scores into "Inconclusive" rather than forcing a verdict.

Pro tier: Adds Sightengine's deepfake model — a separate model specifically trained for face-swap detection and video manipulation. The deepfake model catches cases the genai model doesn't: a real photo with an AI-generated face transplanted into it, for example, might pass the genai model (the base image is real) but fail the deepfake model (the face region has manipulation artifacts).

Processing time using Sightengine is under 1 second for standard-resolution images — the result appears in the overlay before you'd have time to open a separate tab.

Common questions

How does AI image detection work?

AI image detection analyzes statistical patterns in pixel data that human eyes can't see. Every AI generator leaves behind invisible fingerprints in the frequency domain and noise structure of the image — GAN generators leave periodic frequency peaks from their convolution process; diffusion models leave characteristic noise distributions from their denoising process. Detectors are neural networks trained on large datasets of known AI and real images to recognize these patterns and return a probability score.

What is a GAN fingerprint and how do detectors find it?

A GAN fingerprint is a repeating, grid-like artifact left by the generator network's convolution process. Detectors apply a Discrete Fourier Transform (DFT) to the image to convert pixel data into a frequency representation, then look for the characteristic peaks at fixed intervals that GANs produce. Real photographs don't have these periodic peaks. The technique was first described in Zhang et al. (CVPR 2019) and still underlies many modern GAN detectors.

Can AI image detectors be fooled?

Yes. Against images specifically optimized to evade detection — processed with noise addition, JPEG recompression, or non-integer resizing — Imagera's 2026 research found that no commercial detector exceeded 45-60% accuracy. The same tools achieve 88-98% on unmodified AI-generated images. The gap exists because the post-processing techniques that strip out AI artifacts are the same operations used for normal image editing and sharing.

What's the difference between GAN detection and diffusion model detection?

GAN detection uses Discrete Fourier Transform frequency analysis to find periodic grid artifacts. Diffusion model detection is harder — diffusion images don't have a fixed grid pattern — so detectors instead look for characteristic noise distributions from the denoising process, unnatural high-frequency smoothness in specific frequency bands, and cross-region statistical inconsistencies. Both types are handled by modern detectors, but diffusion detection requires larger and more diverse training datasets.

How accurate are AI image detectors in 2026?

On standard unmodified AI-generated images: Hive Moderation 89-98%, Illuminarty approximately 91%, FauxSpy (Sightengine genai API) approximately 88%. False positive rates on real photos run approximately 3% for leading tools under controlled conditions. These numbers drop significantly against post-processed or evasion-optimized images — accuracy is conditional on whether the image has been processed after generation.

Does JPEG compression affect AI image detection accuracy?

Yes, meaningfully. JPEG compression overwrites the fine-grained frequency artifacts that detectors rely on — particularly at quality settings below 90. Images saved from social media (which re-compress uploaded images) or taken as screenshots are harder to detect than the original AI-generated file. This is why FauxSpy's Inconclusive result appears more often on compressed or re-uploaded images than on fresh AI generator output.

Related reading

Try the detector — free, no account

Right-click any image in your browser and run the check instantly. 3 free checks per day, no email required.

🕵️ Add to Chrome — Free 🦊 Add to Firefox — Free