- Research Article
1
- 10.3390/info16020084
Fine-Tuning QurSim on Monolingual and Multilingual Models for Semantic Search
- Jan 23, 2025
- Information
- Tania Afzal + 3 more +3
Transformers have made a significant breakthrough in natural language processing. These models are trained on large datasets and can handle multiple tasks. We compare monolingual and multilingual transformer models for semantic relatedness and verse retrieval. We leveraged data from the original QurSim dataset (Arabic) and used authentic multi-author translations in 22 languages to create a multilingual QurSim dataset, which we released for the research community. We evaluated the performance of monolingual and multilingual LLMs for Arabic and our results show that monolingual LLMs give better results for verse classification and matching verse retrieval. We incrementally built monolingual models with Arabic, English, and Urdu and multilingual models with all 22 languages supported by the multilingual paraphrase-MiniLM-L12-v2 model. Our results show improvement in classification accuracy with the incorporation of multilingual QurSim.
Read more