From smooth to sharp: Frequency-Decoupled Latent Optimization for Realistic Image Generation
Abstract
Latent generative models compress images into learned em-beddings prior to synthesis, and the generation quality critically dependson how faithfully these embeddings preserve visual detail. We observethat while such embeddings are effective at reconstructing low frequencystructure, they struggle to recover sharp high frequency details that areessential for perceptual realism. Conventional reconstruction objectivesimplicitly prioritize coarse structural information over high frequencycontent, which can lead to overly smoothed outputs and degraded vi-sual quality in textured regions. Motivated by this observation, we pro-pose DeBaT, a Decoupled frequency Band Tokenizer that explicitly sep-arates the learning of low and high frequency band embeddings. Thisdecoupling enables accurate reconstruction of fine details while preserv-ing global coherence. Integrated into a latent diffusion based generativemodel, DeBaT allows for sharper and more realistic samples than previ-ous latent tokenizers, confirming that the explicit decoupling of high andlow frequency bands eases the preservation of visual details in learnedembedding spaces.