• Home
  • Search
  • Enhancing Differential Testing With LLMs For Testing Deep Learning Libraries
  • Cite Icon3
  • https://doi.org/10.1145/3735637Copy DOI Icon

Enhancing Differential Testing With LLMs For Testing Deep Learning Libraries

Show More
  • Abstract
  • Literature Map
  • References
  • Citations
  • Similar Papers
Abstract

Differential testing offers a promising strategy to alleviate the test oracle problem by comparing the test results between alternative implementations. However, existing differential testing techniques for deep learning (DL) libraries are limited by the key challenges of finding alternative implementations (called \(counterparts\) ) for a given API and subsequently generating diverse test inputs. To address the two challenges, this paper introduces DLL ens , an LLM-enhanced differential testing technique for DL libraries. The first challenge is addressed by an observation that DL libraries are commonly designed to support the computation of a similar set of DL algorithms. Therefore, the counterpart of a given API’s computation could be successfully synthesized through certain composition and adaptation of the APIs from another DL library. DLL ens incorporates a novel counterpart synthesis workflow, leveraging a large language model (LLM) to search for valid counterparts for differential testing. To address the second challenge, DLL ens incorporates a static analysis technique that extracts the path constraints from the implementations of a given API and its counterpart to guide diverse test input generation. The extraction is facilitated by LLM’s knowledge of the concerned DL library and its upstream libraries. DLL ens incorporates validation mechanisms to manage the LLM’s hallucinations in counterpart synthesis and path constraint extraction. We evaluate DLL ens on two popular DL libraries, TensorFlow and PyTorch. Our evaluation shows that DLL ens synthesizes counterparts for 1.84 times as many APIs as those found by state-of-the-art techniques on these libraries. Moreover, under the same time budget, DLL ens covers 7.23% more branches and detects 1.88 times as many bugs as state-of-the-art techniques on 200 randomly sampled APIs. DLL ens has successfully detected 71 bugs in recent TensorFlow and PyTorch libraries. Among them, 59 are confirmed by developers, including 46 confirmed as previously unknown bugs, and 10 of these previously unknown bugs have been fixed in the latest version of TensorFlow and PyTorch.

Similar Papers
  • Conference Article
  • Citations1

BugsInDLLs : A Database of Reproducible Bugs in Deep Learning Libraries to Enable Systematic Evaluation of Testing Techniques

  • Jun 11, 2025
  • M M Abid Naziri +5
  • Conference Article
  • Citations322

A comprehensive study on deep learning bug characteristics

  • Aug 12, 2019
  • Md Johirul Islam +3
  • Research Article
  • Citations4

DeepProtein: deep learning library and benchmark for protein sequence learning

  • May 19, 2025
  • Bioinformatics
  • Jiaqing Xie +2
  • Book Chapter
  • Citations52

Demand Forecasting in Supply Chain Management Using Different Deep Learning Methods

  • Aug 12, 2020
  • Asma Husna +2
  • Research Article
  • Citations56

Automated Building Information Modeling Compliance Check through a Large Language Model Combined with Deep Learning and Ontology

  • Jul 01, 2024
  • Buildings
  • Nanjiang Chen +3
  • PDF
  • Research Article
  • Citations1011

Fastai: A Layered API for Deep Learning

  • Feb 16, 2020
  • Information
  • Jeremy Howard +1
  • Research Article
  • Citations6

Assessing the risk of takeover catastrophe from large language models.

  • Jun 30, 2024
  • Risk analysis : an official publication of the Society for Risk Analysis
  • Seth D Baum
  • Book Chapter
  • Citations12

Deep Learning With PyTorch

  • Jan 01, 2020
  • Anmol Chaudhary +3
  • Conference Article
  • Citations4

DEFIne

  • Feb 04, 2017
  • Nina Dethlefs +1
  • Research Article
  • Citations1

Evaluating API-Level Deep Learning Fuzzers: A Comprehensive Benchmarking Study

  • Jan 20, 2026
  • ACM Transactions on Software Engineering and Methodology
  • Nima Shiri Harzevili +4
  • Conference Article
  • Citations124

Deep learning on mobile devices: a review

  • May 13, 2019
  • Yunbin Deng
  • Book Chapter
  • Citations2

Predictive Modeling of Surface Roughness Using Machine and Deep Learning Frameworks from Experimental Data of Chemically Etched Polished Silicon Wafer with DDMAF

  • Sep 09, 2022
  • Kheelraj Pandey +1
  • Research Article

Hundred-layer photonic deep learning

  • Nov 24, 2025
  • Nature Communications
  • Tiankuang Zhou +4
  • Research Article
  • Citations1

Overview of deep learning and large language models in machine translation: a special perspective on the Arabic language

  • Jun 23, 2025
  • Journal of Electrical Systems and Information Technology
  • Sanaa Abou Elhamayed +1
  • Research Article

A systematic literature review of large language models in phishing attack generation and detection

  • Jul 01, 2026
  • Array
  • Dinushan Sivaneswaran +5
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.