Most people assume AI image detectors look for obvious visual glitches — extra fingers, blurry ears, artifacts around hair. That's not what they do. The actual detection happens at a level below what your eyes can see: invisible statistical patterns in pixel data that every AI generator leaves behind. This page explains the real mechanism, plainly.
When an AI model generates an image, it doesn't paint pixels the way a photographer captures them. It runs a mathematical process — convolution in GANs, iterative denoising in diffusion models — that produces characteristic statistical patterns in the pixel array. These patterns are invisible to the eye but show up clearly when you apply the right analysis.
Think of it like a manufacturing process. A photo printed on an inkjet printer and a photo printed on a laser printer look the same to a casual viewer. But under a microscope, the inkjet print has a characteristic dot pattern and the laser print has a different characteristic toner pattern. An expert with a microscope can tell them apart not by how the photo looks, but by the physical signature of the process that made it.
AI detectors are the "microscope." They don't look at whether the photo seems too perfect or the lighting seems off — those visual tells are unreliable. They look for the process signature that the AI generator left in the data.
The first generation of AI image detectors was built specifically for GANs (Generative Adversarial Networks) — the technology behind tools like StyleGAN that generate photorealistic faces. GANs work through a series of convolutional operations that upsample a low-resolution noise vector into a full-resolution image. That convolution process creates repeating patterns at fixed pixel intervals — a grid-like artifact in the frequency domain.
A GAN fingerprint is the characteristic frequency artifact produced by the upsampling convolutions in the generator network. To find it, detectors apply a technique called the Discrete Fourier Transform (DFT) to the image — this converts the pixel data from a spatial representation (where pixels are) into a frequency representation (how often patterns repeat across the image). In the frequency domain, GAN-generated images show distinctive peaks at intervals that correspond to the generator's architecture. Real photographs don't have these periodic peaks.
The peaks are small — they represent a tiny fraction of the total pixel energy. But they're consistent. Because all GAN generators using similar upsampling architectures leave similar-shaped peaks, a classifier trained on frequency maps can identify GAN-generated images even when the visual output looks completely realistic. A 2019 paper by Zhang et al. (CVPR 2019) first demonstrated this approach and showed it could achieve over 99% accuracy on a controlled dataset of GAN images.
GAN images have a predictable, architecture-specific fingerprint. Once you've trained a detector on images from a specific GAN architecture (StyleGAN, BigGAN, etc.), it generalizes reasonably well to other GAN architectures because the convolution-upsampling process is structurally similar across them. The frequency peaks shift position depending on the architecture, but the fact that peaks exist is consistent.
Modern AI image generators — Midjourney, DALL-E, Stable Diffusion, Imagen, Flux — use diffusion models rather than GANs. Diffusion is a different process: start with random noise, then iteratively remove noise step-by-step guided by a text prompt or conditioning signal until the image emerges. This process doesn't produce the same fixed-frequency artifacts as GAN convolution, which means frequency-domain GAN detectors don't transfer reliably to diffusion images.
Diffusion model detection relies on several overlapping signals rather than a single frequency fingerprint:
Denoising noise distribution
The iterative denoising process in diffusion models produces a characteristic statistical distribution of residual noise across the image — more uniform than real camera sensor noise, which is heterogeneous (noisier in dark areas, less noisy in bright areas). Detectors can quantify this by analyzing the local noise variance pattern across image regions.
High-frequency smoothness
Diffusion models tend to produce images that are smoother at fine scales than photographs of equivalent content. Real photos have fine-grained texture noise from the camera sensor and optical system at every scale. Diffusion images often lack this fine-scale variation in specific frequency bands — not because they look smooth, but because the statistical distribution of very-high-frequency content is different from real images.
Cross-region consistency artifacts
Real photos have natural optical relationships between regions — the depth of field falloff, the light source consistency, the lens distortion at edges — that constrain how different parts of the image relate to each other. Diffusion images generate each region more independently, sometimes producing subtle statistical inconsistencies in how light behaves across the image that a classifier can learn to detect.
Generator-specific compression artifacts
Many diffusion models save their output in a specific format or run through a specific VAE (variational autoencoder) decoder that introduces characteristic compression patterns. Detectors trained on large datasets of specific generator outputs can learn these format-specific signatures even when they're subtle.
Because diffusion detection depends on multiple overlapping signals rather than a single strong fingerprint, it's generally harder than GAN detection — and accuracy numbers reflect this. The advance of the field has been in building classifiers that combine all of these signals together, using neural networks that learn which combination of features best predicts AI origin across diverse generator architectures.
Deepfake detection is related to but distinct from AI image detection. A deepfake typically involves a real photo or video with a face surgically replaced or manipulated — not a fully AI-generated image. This means the detection problem is different: instead of finding a generator fingerprint across the whole image, the detector looks for localized inconsistencies at the boundary between the real and manipulated regions.
Face-swap boundary analysis
The edges of a replaced face often have blending artifacts — statistical noise patterns that differ between the pasted region and the surrounding original image. Detectors segment the face region, analyze the edge quality, and compare noise characteristics between the face and background.
Biometric inconsistency
Face-swap algorithms sometimes introduce subtle biometric inconsistencies: eye gaze direction doesn't match the geometry of the 3D scene, pupil reflections are inconsistent with the light source, facial asymmetry is slightly different on one frame than another in a video. Video deepfake detectors can track these across frames.
Temporal coherence (video only)
In video deepfakes, the swapped face often flickers slightly between frames — the face-swap algorithm processes each frame somewhat independently, so subtle variations accumulate in ways that look unnatural in motion. Video deepfake detectors specifically look for these frame-to-frame inconsistencies.
FauxSpy uses different detection models for these two problems: Sightengine's genai and type models for AI image detection on the free tier, and the deepfake model (which powers Pro) for face manipulation and video deepfake analysis. The distinction matters: a fully AI-generated face scores high on AI image detection; a deepfake of a real person's face in a real photo scores high on deepfake detection; a photo of a real person scores low on both.
The honest answer: AI image detection accuracy is conditional. On standard test sets — images from known AI generators with no post-processing — the leading tools in 2026 benchmark at 88-98% accuracy. Hive Moderation reports 89-98% depending on the test; Illuminarty approximately 91%; FauxSpy (Sightengine genai) approximately 88%. False positive rates — real photos incorrectly flagged as AI — run approximately 3% for leading tools under controlled test conditions.
Those numbers drop significantly in the real world for several reasons:
| What degrades accuracy | Why it works | Effect on detection |
|---|---|---|
| JPEG compression (below quality 90) | Destroys high-frequency artifacts that detectors key on | Moderate degradation |
| Screenshot-then-crop | Applies re-compression and resampling in one step | Moderate degradation |
| Adding Gaussian noise | Obscures generator noise pattern with random noise | High degradation |
| Non-integer resize (e.g. 87%) | Interpolation shifts frequency artifacts to unpredicted positions | Moderate degradation |
| Adversarial post-processing (targeted) | Specifically optimized to fool a known detector | Severe degradation |
| Heavy photo editing (crop, filters, retouch) | Overwrites generator signature with editing software artifacts | Moderate degradation |
Research from Imagera in 2026 found that against images specifically optimized to evade detection, no commercial detector tested exceeded 45-60% accuracy — compared to 88-98% on unmodified AI images. This gap is the central challenge for the field: the same techniques used to improve image quality (compression, denoising, upscaling) also happen to strip out the signals detectors rely on.
This is why FauxSpy returns an Inconclusive result when confidence falls below a threshold rather than forcing a binary verdict. A low-confidence "probably real" verdict is less useful — and potentially more harmful — than an honest "we can't tell." Accuracy on unmodified images is high; the Inconclusive category handles the edges where the evidence is genuinely ambiguous.
C2PA (Coalition for Content Provenance and Authenticity) is a complementary approach to AI image detection that works fundamentally differently: instead of analyzing the image after the fact, it embeds a cryptographically signed provenance record at the moment of generation or capture. OpenAI embeds C2PA metadata in DALL-E and ChatGPT images; Adobe embeds it in Firefly-generated content; major camera manufacturers are adding it to hardware.
C2PA is promising but has a significant real-world limitation: metadata gets stripped when images are uploaded to most social platforms. Instagram, X/Twitter, Facebook, Reddit, and WhatsApp all strip or ignore EXIF and IPTC metadata during upload processing — which is also where C2PA data lives. An image with intact C2PA credentials is provably from a specific source. The same image shared through a typical social media workflow has no credentials at all, even if it was originally credentialed at the source.
This means C2PA and statistical AI image detection are complementary, not competing, approaches:
For most consumer use cases — checking a profile photo, a news image, a marketplace listing — the image has passed through at least one platform that strips metadata, so C2PA provides no useful signal. Statistical detection is the relevant tool.
FauxSpy's detection is powered by Sightengine's API — a computer vision API that provides specialized models for AI-generated content detection. This is the same API used directly by developers building moderation systems; FauxSpy wraps it in a right-click browser interface that requires no technical setup.
Free tier: Uses Sightengine's genai model (detects AI-generated images across GAN and diffusion generators) and the type model (classifies content type). Returns a probability score from 0.0 to 1.0 for AI origin. FauxSpy translates scores above a threshold into "Likely AI," scores below into "Likely Real," and borderline scores into "Inconclusive" rather than forcing a verdict.
Pro tier: Adds Sightengine's deepfake model — a separate model specifically trained for face-swap detection and video manipulation. The deepfake model catches cases the genai model doesn't: a real photo with an AI-generated face transplanted into it, for example, might pass the genai model (the base image is real) but fail the deepfake model (the face region has manipulation artifacts).
Processing time using Sightengine is under 1 second for standard-resolution images — the result appears in the overlay before you'd have time to open a separate tab.
AI image detection analyzes statistical patterns in pixel data that human eyes can't see. Every AI generator leaves behind invisible fingerprints in the frequency domain and noise structure of the image — GAN generators leave periodic frequency peaks from their convolution process; diffusion models leave characteristic noise distributions from their denoising process. Detectors are neural networks trained on large datasets of known AI and real images to recognize these patterns and return a probability score.
A GAN fingerprint is a repeating, grid-like artifact left by the generator network's convolution process. Detectors apply a Discrete Fourier Transform (DFT) to the image to convert pixel data into a frequency representation, then look for the characteristic peaks at fixed intervals that GANs produce. Real photographs don't have these periodic peaks. The technique was first described in Zhang et al. (CVPR 2019) and still underlies many modern GAN detectors.
Yes. Against images specifically optimized to evade detection — processed with noise addition, JPEG recompression, or non-integer resizing — Imagera's 2026 research found that no commercial detector exceeded 45-60% accuracy. The same tools achieve 88-98% on unmodified AI-generated images. The gap exists because the post-processing techniques that strip out AI artifacts are the same operations used for normal image editing and sharing.
GAN detection uses Discrete Fourier Transform frequency analysis to find periodic grid artifacts. Diffusion model detection is harder — diffusion images don't have a fixed grid pattern — so detectors instead look for characteristic noise distributions from the denoising process, unnatural high-frequency smoothness in specific frequency bands, and cross-region statistical inconsistencies. Both types are handled by modern detectors, but diffusion detection requires larger and more diverse training datasets.
On standard unmodified AI-generated images: Hive Moderation 89-98%, Illuminarty approximately 91%, FauxSpy (Sightengine genai API) approximately 88%. False positive rates on real photos run approximately 3% for leading tools under controlled conditions. These numbers drop significantly against post-processed or evasion-optimized images — accuracy is conditional on whether the image has been processed after generation.
Yes, meaningfully. JPEG compression overwrites the fine-grained frequency artifacts that detectors rely on — particularly at quality settings below 90. Images saved from social media (which re-compress uploaded images) or taken as screenshots are harder to detect than the original AI-generated file. This is why FauxSpy's Inconclusive result appears more often on compressed or re-uploaded images than on fresh AI generator output.
Right-click any image in your browser and run the check instantly. 3 free checks per day, no email required.
🕵️ Add to Chrome — Free 🦊 Add to Firefox — Free