- Conference Article
- 10.1109/icca66035.2025.11430855
Automated Feature Extraction and UML Modeling from Real-World Unstructured Egyptian Arabic Text
- Dec 22, 2025
- Farah El Degheidy + 3 more +3
Conventional requirements engineering (RE) is often a manual process that is tedious, time-consuming, susceptible to errors, and hard to scale in case of continuous updates. Recent advances in natural language processing (NLP) transformed the RE process via automation. However, automation works better with structured data and formal requirement formats, limiting the ability to elicit and extract requirements from unstructured data; the rapid growth of digital platforms accentuates the need for treating user posts, comments, opinions and feedback as primary data sources for requirements mining and modeling. However, research conducted on unstructured Egyptian Arabic text is largely underexplored, due to its inherent linguistic challenges. This paper fills a critical gap in Arabic NLP, by providing an automated feature extraction and UML modeling pipeline utilizing state of the art embedding models, density-based clustering algorithm, and Large Language Models (LLMs), specifically engineered for handling Egyptian-Arabic text. The proposed pipeline demonstrates strong positive results, producing quality features and meaningful use case diagrams, and further confirms the feasibility of applying this advanced NLP-based pipeline to real world Egyptian dialectal data.
Read more