• Home
  • Search
  • Evaluating API-Level Deep Learning Fuzzers: A Comprehensive Benchmarking Study
  • Cite Icon1
  • https://doi.org/10.1145/3729533Copy DOI Icon

Evaluating API-Level Deep Learning Fuzzers: A Comprehensive Benchmarking Study

Show More
  • Abstract
  • Literature Map
  • References
  • Citations
  • Similar Papers
Abstract

In recent years, the practice of fuzzing Deep Learning (DL) APIs has received significant attention in the software engineering community. Many API-level DL fuzzers have been proposed to test individual DL APIs by generating malformed input. Although these fuzzers have been effective in detecting bugs and outperforming prior work, there remains a gap in benchmarking them against ground-truth, real-world bugs in DL libraries. Existing comparisons among these API-level DL fuzzers primarily focus on the bugs detected but do not offer a comprehensive, in-depth evaluation of the fuzzers’ effectiveness. In this work, we perform the first in-depth evaluation of state-of-the-art API-level DL fuzzers that generate tests for single DL APIs, focusing on their effectiveness against real-world bugs. We manually created an extensive benchmark dataset, including 517 real-world DL bugs collected from PyTorch and TensorFlow libraries that can be triggered by malformed inputs. We then apply seven state-of-the-art DL fuzzers— FreeFuzz , DeepRel , NablaFuzz , DocTer , ACETest , TitanFuzz , and FuzzGPT —to our benchmark dataset, following their respective instructions. Our results show that these fuzzers detect only 6.5% (34 out of 517) of the unique real-world bugs in the dataset. Our analysis identifies two dominant factors that impact the effectiveness of these fuzzers in detecting real-world bugs. These findings suggest opportunities for improving the performance of fuzzers in future work. Overall, this study extends previous work on DL fuzzers by providing an extensive evaluation and benchmarking platform for fuzzing DL libraries.

Similar Papers
  • Research Article
  • Citations3

Enhancing Differential Testing With LLMs For Testing Deep Learning Libraries

  • May 14, 2025
  • ACM Transactions on Software Engineering and Methodology
  • Meiziniu Li +5
  • Conference Article
  • Citations1

BugsInDLLs : A Database of Reproducible Bugs in Deep Learning Libraries to Enable Systematic Evaluation of Testing Techniques

  • Jun 11, 2025
  • M M Abid Naziri +5
  • Conference Article
  • Citations322

A comprehensive study on deep learning bug characteristics

  • Aug 12, 2019
  • Md Johirul Islam +3
  • Conference Article
  • Citations25

Towards offensive language detection and reduction in four Software Engineering communities

  • Jun 21, 2021
  • Jithin Cheriyan +2
  • Book Chapter
  • Citations12

Deep Learning With PyTorch

  • Jan 01, 2020
  • Anmol Chaudhary +3
  • Research Article
  • Citations1

"Identification of promising genotypes in varietal trials of sugarcane using deep learning "

  • Dec 31, 2021
  • Journal of Sugarcane Research
  • Syed Sarfaraz Hasan +1
  • Conference Article
  • Citations279

Toward Deep Learning Software Repositories

  • May 01, 2015
  • Martin White +3
  • Conference Article
  • Citations124

Deep learning on mobile devices: a review

  • May 13, 2019
  • Yunbin Deng
  • Book Chapter
  • Citations52

Demand Forecasting in Supply Chain Management Using Different Deep Learning Methods

  • Aug 12, 2020
  • Asma Husna +2
  • PDF
  • Research Article
  • Citations1011

Fastai: A Layered API for Deep Learning

  • Feb 16, 2020
  • Information
  • Jeremy Howard +1
  • PDF
  • Research Article
  • Citations47

Deep Ensemble Learning Approaches in Healthcare to Enhance the Prediction and Diagnosing Performance: The Workflows, Deployments, and Surveys on the Statistical, Image-Based, and Sequential Datasets.

  • Oct 14, 2021
  • International Journal of Environmental Research and Public Health
  • Duc-Khanh Nguyen +2
  • Research Article
  • Citations8

CG-ERNet: a lightweight Curvature Gabor filtering based ear recognition network for data scarce scenario

  • May 06, 2021
  • Multimedia Tools and Applications
  • Aman Kamboj +2
  • PDF
  • Research Article
  • Citations28

Deep Transfer Learning based Fusion Model for Environmental Remote Sensing Image Classification Model

  • Jan 07, 2022
  • European Journal of Remote Sensing
  • Anwer Mustafa Hilal +6
  • Conference Article
  • Citations4

DEFIne

  • Feb 04, 2017
  • Nina Dethlefs +1
  • Book Chapter
  • Citations2

Predictive Modeling of Surface Roughness Using Machine and Deep Learning Frameworks from Experimental Data of Chemically Etched Polished Silicon Wafer with DDMAF

  • Sep 09, 2022
  • Kheelraj Pandey +1
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.