• Home
  • Search
  • Hybrid Attention Approach for Source Code Comment Generation
  • https://doi.org/10.5755/j01.itc.54.2.36699Copy DOI Icon

Hybrid Attention Approach for Source Code Comment Generation

Show More
  • Abstract
  • Literature Map
  • Similar Papers
Abstract

Currently, developers are often obligated to enhance code quality. High-quality code is often accompanied with comprehensive summaries, including code documentation and function explanations, which are invaluable for maintenance and further development. Regrettably, few software projects provide sufficient code comments owing to the high costs associated with human labeling. Contemporary researchers in software engineering concentrate on the methods for automated comment generating. Initial algorithms depended on handwritten templates or information retrieval methods. With the advancement of machine learning, researchers construct automated models for machine translation instead. Nonetheless, the produced code comments remain inadequate owing to the significant disparity between code structure and normal language. This study introduces a unique deep learning model, At-ComGen, which utilizes hybrid attention for the automated creation of source code comments. Utilizing two separate LSTM encoders, our approach integrates essential tokens from source code functions with the code structure, represented by a corresponding Abstract Syntax Tree. In contrast to earlier data-driven models, our methodology utilizes code syntax and semantics in the generation of comments. The hybrid attention method, used for comment creation for the first time to our knowledge, enhances the quality of code comments. The tests demonstrate that At-ComGen is efficacious and surpasses other prevalent methodologies. Machine comments from Seq2Seq and CODE-NN disregard code structure underlying DeepCom and At-ComGen. At-ComGen has 59.3%, 36.4%, 43.3%, and 13.1% higher comment BLEU values than baseline models for a 5-line function. Even though model performance reduces with comment length, At-ComGen's comments often outperform others. 5–10-word machine comments work best. For reference length 10, At-ComGen has 38.2%, 23.7%, 9.3%, and 4.4% greater BLEU values than the other baseline models.

Similar Papers
  • Conference Article
  • Citations22

A Survey on Research of Code Comment

  • Jan 12, 2019
  • Bai Yang +2
  • Research Article
  • Citations1

Improving Just-In-Time Comment Updating via AST Edit Sequence

  • Oct 01, 2022
  • International Journal of Software Engineering and Knowledge Engineering
  • Jiawen Huang +4
  • Conference Article
  • Citations29

Learning Semantic Features for Software Defect Prediction by Code Comments Embedding

  • Nov 01, 2018
  • Xuan Huo +3
  • Research Article
  • Citations148

Neural Network-based Detection of Self-Admitted Technical Debt

  • Jul 29, 2019
  • ACM Transactions on Software Engineering and Methodology
  • Xiaoxue Ren +5
  • PDF
  • Research Article
  • Citations21

Deep Code-Comment Understanding and Assessment

  • Jan 01, 2019
  • IEEE Access
  • Deze Wang +5
  • Conference Article
  • Citations166

Deep learning similarities from different representations of source code

  • May 28, 2018
  • Michele Tufano +5
  • Conference Article
  • Citations30

API2Com: On the Improvement of Automatically Generated Code Comments Using API Documentations

  • May 01, 2021
  • Ramin Shahbazi +2
  • Conference Article
  • Citations20

Source code retrieval on StackOverflow using LDA

  • May 01, 2015
  • Achmad Arwan +2
  • Conference Article
  • Citations1

ASKDetector: An AST-Semantic and Key Features Fusion based Code Comment Mismatch Detector

  • Apr 15, 2024
  • Haiyang Yang +4
  • Conference Article
  • Citations1345

Clone detection using abstract syntax trees

  • Mar 16, 1998
  • I.D Baxter +4
  • Dissertation

Disfluency detection and reconstruction of spontaneous speech transcripts

  • Jan 01, 2016
  • Linlin Wang
  • Conference Article
  • Citations3

Deep Just-In-Time Consistent Comment Update via Source Code Changes

  • Nov 25, 2022
  • Shikai Guo +3
  • Book Chapter
  • Citations2

Evaluating a LSTM Neural Network and a Word2vec Model in the Classification of Self-admitted Technical Debts and Their Types in Code Comments

  • Jan 01, 2021
  • Rafael Meneses Santos +3
  • Conference Article
  • Citations3

Summarizing Source Code from Structure and Context

  • Jul 18, 2022
  • Shifu Hou +2
  • Research Article
  • Citations1

Penerapan Abstract Syntax Tree dan Algoritma Damerau-Levenshtein Distance untuk Mendeteksi Plagiarisme pada Berkas Source Code

  • Feb 05, 2019
  • Jurnal Telematika
  • Stephanie Rusdianto +1
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.