- Research Article
- 10.1109/mmul.2026.3657223
Text-vision guided Latent Prediction for High-Fidelity Blind Face Restoration
- Jan 01, 2026
- IEEE MultiMedia
- Juan Cai + 3 more +3
Blind face restoration (BFR) aims to recover high-quality (HQ) facial images from degraded inputs with unknown distortions, while preserving identity consistency. Existing methods either rely solely on visual priors or incorporate generic textual cues, struggle with noise sensitivity or insufficient semantic guidance. In this paper, we present a vision-language guided BFR framework that consists of three key components: (1) a Facial Key Attribute Estimation module that leverages structured text to compactly describe facial attributes; (2) a Dual-modal Based Latent Prediction module built on a shared-attention Transformer, which fuses visual and textual embeddings to predict dictionary elements representing the target face; (3) a Text-based Adaptive Feature Filter that dynamically suppresses noise in skip-connected features by aligning them with textual semantics, enhancing the generator’s ability to leverage multi-scale spatial information. Extensive experiments demonstrate that our method outperforms state-of-the-art approaches on multiple benchmarks, particularly under severe degradation.
Read more