• Home
  • Search
  • An Empirical Study on Language Models for Generating Log Statements in Test Code
  • Cite Icon1
  • https://doi.org/10.1145/3759915Copy DOI Icon

An Empirical Study on Language Models for Generating Log Statements in Test Code

Show More
  • Abstract
  • Literature Map
  • References
  • Citations
  • Similar Papers
Abstract

Log statements play a critical role in modern software development, capturing essential runtime information necessary for software maintenance. Recently, new techniques have been developed to automate logging activities, allowing log statements to be injected into code by identifying specific code locations, selecting the appropriate log level, and generating meaningful log messages that describe the behavior being logged. Although automated logging in production code has attracted significant attention, little focus has been given to the injection of logs in test code. To fill this gap, we conduct an empirical study on 5,206,759 Java test methods collected from 6,405 GitHub projects to explore and disclose the effectiveness and limitations of Pre-trained Language Models (PLMs) and Large Language Models (LLMs) for generating and injecting test log statements. Our findings demonstrate that general-purpose LLMs like GPT-3.5-Turbo, when properly instructed to inject logging statements in test methods, performs comparably to the best-performing PLMs on predicting log level. Additionally, GPT-3.5-Turbo substantially outperforms the best in PLMs on predicting log position with a 33.97% improvement while also achieving superior performance in predicting log messages in terms of BLEU and ROUGE . This work takes the first step toward evaluating the capability of PLMs and LLMs to generate test log statement. This work takes the first step toward evaluating the capability of PLMs and LLMs to generate test log statements. To facilitate future research, we have open-sourced all data and source code used in this work.

Similar Papers
  • Research Article
  • Citations1

Using a large language model to provide individualized feedback for pre-service physics teachers’ written reflections

  • Nov 21, 2025
  • Disciplinary and Interdisciplinary Science Education Research
  • Stefan Sorge +2
  • Video Transcripts

Robust Transfer Learning with Pretrained Language Models through Adapters

  • Aug 01, 2021
  • Underline Science Inc.
  • Wenjuan Han +2
  • Conference Article

Exploring Layer-wise Representations of English and Chinese Homonymy in Pre-trained Language Models

  • Jan 01, 2025
  • Matthew King-Hang Ma +3
  • Research Article
  • Citations2

Adaptive Boosting LLMs for Text Classification.

  • Jan 01, 2026
  • IEEE transactions on neural networks and learning systems
  • Mengyao Wang +6
  • PDF
  • Research Article
  • Citations10

AGI-P: A Gender Identification Framework for Authorship Analysis Using Customized Fine-Tuning of Multilingual Language Model

  • Jan 01, 2024
  • IEEE Access
  • Raheem Sarwar +6
  • Book Chapter

Leverage LLMs on Knowledge Tagging for Math Questions in Education

  • Oct 10, 2025
  • Hang Li +2
  • Research Article
  • Citations16

A Review of the Opportunities and Challenges with Large Language Models in Radiology: The Road Ahead.

  • Nov 21, 2024
  • AJNR. American journal of neuroradiology
  • Neetu Soni +4
  • Research Article
  • Citations14

Fine-Tuned Understanding: Enhancing Social Bot Detection With Transformer-Based Classification

  • Jan 01, 2024
  • IEEE Access
  • Amine Sallah +6
  • Conference Article
  • Citations3

Chinese-Korean Weibo Sentiment Classification Based on Pre-trained Language Model and Transfer Learning

  • May 06, 2022
  • Hengxuan Wang +3
  • Research Article
  • Citations1

Detecting Hard-Coded Credentials in Software Repositories via LLMs

  • Jul 07, 2025
  • Digital Threats: Research and Practice
  • Chidera Biringa +1
  • Research Article
  • Citations10

UniproLcad: Accurate Identification of Antimicrobial Peptide by Fusing Multiple Pre-Trained Protein Language Models

  • Apr 11, 2024
  • Symmetry
  • Xiao Wang +3
  • Research Article
  • Citations15

Automatic quantitative stroke severity assessment based on Chinese clinical named entity recognition with domain-adaptive pre-trained large language model

  • Feb 27, 2024
  • Artificial Intelligence In Medicine
  • Zhanzhong Gu +9
  • Conference Article
  • Citations13

Integrating Pre-Trained Speech and Language Models for End-to-End Speech Recognition

  • Jan 01, 2024
  • Yukiya Hono +5
  • Conference Article
  • Citations5

SuicidEmoji: Derived Emoji Dataset and Tasks for Suicide-Related Social Content

  • Jul 10, 2024
  • Tianlin Zhang +5
  • PDF
  • Conference Article
  • Citations31

PANLP at MEDIQA 2019: Pre-trained Language Models, Transfer Learning and Knowledge Distillation

  • Jan 01, 2019
  • Wei Zhu +6
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.