- Research Article
8
- 10.19139/soic-2310-5070-2546
CAT-VAE: A Cross-Attention Transformer-Enhanced Variational Autoencoder for Improved Image Synthesis
- Jul 13, 2025
- Statistics, Optimization & Information Computing
- Khadija Rais + 3 more +3
Deep generative models are increasingly useful in medical image analysis to solve various issues, including class imbalance in classification tasks, motivating the development of multiple methods, where the Variational Autoencoder (VAE) is recognized as one of the most popular image generators. However, the utilization of convolutional layers in VAEs weakens their ability to model global context and long-range dependencies. This paper presents CAT-VAE, a hybrid approach based on VAE and Cross-Attention Transformers (CAT), in which a cross-attention mechanism is employed to promote long-range dependencies and improve the quality of the generated images. On the Ultrasound breast cancer dataset, the CAT-VAE achieved better image quality (FID 8.7659 for Malignant and 7.8761 for Normal) compared to the standard VAE. An experiment was conducted where a CNN classifier model was trained without data augmentation, with augmentation based on VAE, and using synthetic data generated by CAT-VAE. The CNN achieved the highest accuracy (97.00%) when trained with CAT-VAE synthetic images. A classification accuracy of 86.67% was achieved with mixed datasets of real and synthetic images, demonstrating that CAT-VAE improves generalization and resilience. These results highlight CAT-VAE's ability to produce diverse and realistic synthetic datasets.
Read more