- Conference Article
- 10.1109/icaic67076.2026.11395781
Software Engineering Challenges in the Deployment of Generative AI Models at Scale
- Feb 18, 2026
- Venkata Nagendra Satyam + 2 more +2
The rapid development of one of the biggest advances in the sphere of Generative AI has led to natural language generation still being limited to large-scale implementation due to latency, reproducibility, integration complexity, and scalability challenges. This project offers a simple and repeatable deployment pipeline which is based on the Kaggle Fitness Exercises dataset. The final corpus was then fine-tuned with GPT-2 on Hugging Face and PyTorch using extensive preprocessing, that is, data cleaning, instruction merging, normalization, tokenization, and outlier filtering. The model was evaluated on five configurations with a loss of evaluation of approximately 0.4083, 0.4094, and perplexity of 1.5043, 1.5058, and sensitivity and ablation studies indicate the most sensitive hyperparameter is the learning rate. The fine-tuning of the model was performed through a FastAPI service, and it was then containerized using Docker, which made it suitable for deployment on any cloud provider. The outcome was an even training process, low perplexity, and excellent performance reproducibility, while the tests in the production environment showed moderate in inference and throughput, which in turn led to the decision to carry out further optimization on the runtime. The proposed method offers excellent reproducibility, minimum operational cost, and easy deployment compared to the more complex orchestration and full MLOps pipelines.
Read more