GLARE: Towards Generalizable Detection of Latent Diffusion Images with Global-Local Reconstruction Error
Abstract
Latent diffusion models (LDMs) have emerged as a lead-ing paradigm for image generation by performing denoising in a com-pressed latent space, which provides high fidelity and efficiency but alsointensifies forensic and security concerns. While most existing detectionapproaches rely on supervised learning, training-free detectors are ap-pealing for their simplicity and ease of deployment. However, their per-formance remains limited and degrades markedly on novel generators.Our analysis reveals a more pronounced context dependence in LDM-generated images compared to real ones. Motivated by this, we intro-duce GLARE, a novel training-free method designed to exploit this de-pendence as a detection signal. Specifically, GLARE measures the rela-tive difference between full-image and patch-wise reconstruction errorswithin a shared autoencoder. Dividing the image into patches implicitlyremoves global context, thereby inducing a characteristic shift in recon-struction error that serves as a discriminative feature. Furthermore, alightweight semantic complexity calibration is incorporated to compen-sate for content-induced variation. Extensive experiments across a widerange of generators demonstrate the effectiveness and remarkable gener-alization capability of GLARE. Our method outperforms state-of-the-artsupervised and training-free baselines significantly and shows strong ro-bustness against common post-processing operations.