- Research Article
- 10.1142/s021946782750077x
Visual Question Answering Model Using Transformer-Based Multi-Head Cross Attention Network with Optimized Gold Rush Octave Convolution Network
- Oct 28, 2025
- International Journal of Image and Graphics
- J Jinu Sophia + 2 more +2
The ability of Visual Question Answering (VQA) models to provide insightful responses to images based on their content holds the potential to completely transform machine interpretation and interaction. VQA systems face challenges due to noisy image data, making accurate question answering difficult. Several deep learning algorithms are used for predicting answers from images, but they don’t yield enough outcomes. To get over these problems, this work is suggested. In this manuscript, a visual question–answer model-based Transformer-Based Multi-Head Cross Attention Network with optimized Gold Rush Octave Convolution Network (T-MHCAN-OGRCNN) is proposed. To enhance VQA accuracy, three datasets are taken: VQA v2.0, OK-VQA, and FVQA, which include images with embedded questions and answers. These images often contain noise, necessitating pre-processing for improved results. Then, Task-Oriented Homogenized Automatic (TOHA) method for noise reduction is used. For word embedding, Multiscale Parallel MobileNetV3 (WE-MP-MobileNetV3) is employed. Following that, visual and textual features are extracted using Transformer-Based Multi-Head Cross Attention Network (T-MHCAN). The core VQA task is handled by a Fully Octave Convolution Network (OCNN), whose weight parameters are optimized using Gold Rush Optimizer (GRO). Therefore, this method incorporates cutting-edge strategies to improve precision and resilience of VQA systems. The suggested T-MHCAN-OGRCNN technique is implemented on Python platform. Three datasets were used to evaluate the effectiveness of the suggested T-MHCAN-OGRCNN, which achieves superior results than current techniques with 99.9% accuracy and 99.8% recall. This demonstrates approach’s exceptional effectiveness and room for growth in industry.
Read more