- Conference Article
- 10.1109/aece67531.2025.11386596
Security Risks in AI-Generated Code: A Comparative Analysis of Industry and Academic Benchmarks
- Nov 21, 2025
- Neha Kumari + 3 more +3
In today’s era of Generative AI, the Large Language Model (LLMs)applications are largely used to generate functioning source code. These applications simply does not require any programming skill except a prompt. However, the LLM generated codes are prone to safety and efficiency concerns. A common safety concerns in popular programming languages source code like Java, C++ and C are SQL injection, pointer dereferencing, improper input validation, etc. There are several academic studies and industry approaches to enable a strong and safe AI code generation such as prompt inversion, formal verification, and many more. In this paper, we analyze both the industry and academic approach to reduce and fix the threats over AI generated code and discussed key findings. The platform like Veracode that uses static analysis, dynamic testing and the platform CrowdStrike working to develop self-learning, multi-agent AI systems that employ Red Teaming capabilities. The academic benchmark like CodeLMSec that proposed black-box vulnerability probing framework, FormAI-v2 applies a large-scale formal verification for C language vulnerabilities, and the CodeSecEval benchmark that is based on prompt inversion to evaluate secure code generation and repair. Our studies on the mentioned benchmark highlights that the currently LLM generated source code persists vulnerabilities due to lack of benchmark dataset. The unsafe code fixes are approached in one direction only, either based on static analysis or dynamic analysis.
Read more