• Home
  • Search
  • Labeling source code with information retrieval methods: an empirical study
  • Cite Icon47
  • https://doi.org/10.1007/s10664-013-9285-5Copy DOI Icon

Labeling source code with information retrieval methods: an empirical study

Show More
  • Abstract
  • Literature Map
  • References
  • Citations
  • Similar Papers
Abstract

To support program comprehension, software artifacts can be labeled--for example within software visualization tools--with a set of representative words, hereby referred to as labels. Such labels can be obtained using various approaches, including Information Retrieval (IR) methods or other simple heuristics. They provide a bird-eye's view of the source code, allowing developers to look over software components fast and make more informed decisions on which parts of the source code they need to analyze in detail. However, few empirical studies have been conducted to verify whether the extracted labels make sense to software developers. This paper investigates (i) to what extent various IR techniques and other simple heuristics overlap with (and differ from) labeling performed by humans; (ii) what kinds of source code terms do humans use when labeling software artifacts; and (iii) what factors--in particular what characteristics of the artifacts to be labeled--influence the performance of automatic labeling techniques. We conducted two experiments in which we asked a group of students (38 in total) to label 20 classes from two Java software systems, JHotDraw and eXVantage. Then, we analyzed to what extent the words identified with an automated technique--including Vector Space Models, Latent Semantic Indexing (LSI), latent Dirichlet allocation (LDA), as well as customized heuristics extracting words from specific source code elements--overlap with those identified by humans. Results indicate that, in most cases, simpler automatic labeling techniques--based on the use of words extracted from class and method names as well as from class comments--better reflect human-based labeling. Indeed, clustering-based approaches (LSI and LDA) are more worthwhile to be used for source code artifacts having a high verbosity, as well as for artifacts requiring more effort to be manually labeled. The obtained results help to define guidelines on how to build effective automatic labeling techniques, and provide some insights on the actual usefulness of automatic labeling techniques during program comprehension tasks.

Similar Papers
  • Research Article
  • Citations1

Special issue on program comprehension

  • Jul 26, 2014
  • Empirical Software Engineering
  • Michael W Godfrey +1
  • Book Chapter
  • Citations58

A Comprehensive Survey on Topic Modeling in Text Summarization

  • Jan 01, 2022
  • G Bharathi Mohan +1
  • Conference Article
  • Citations2

Predicting Budget from Transportation Research Grant Description: An Exploratory Analysis of Text Mining and Machine Learning Techniques

  • Oct 01, 2017
  • SHILAP Revista de lepidopterología
  • Ayush Singhal +2
  • Conference Article

The text mining model building of open questionnaire based on LSA

  • Oct 01, 2016
  • Jiuru Zhao +2
  • Conference Article
  • Citations37

Experimenting with Latent Semantic Analysis and Latent Dirichlet Allocation on Automated Essay Grading

  • Dec 14, 2020
  • Jalaa Hoblos
  • Conference Article
  • Citations14

Redacting sensitive information in software artifacts

  • Jun 02, 2014
  • Mark Grechanik +4
  • Research Article
  • Citations8

A Comparative Analysis of TF-IDF, LSI and LDA in Semantic Information Retrieval Approach for Paper-Reviewer Assignment

  • Nov 30, 2019
  • Journal of Engineering and Applied Sciences
  • A Ayodele Adebiyi +3
  • Conference Article
  • Citations39

TopicView: Visually Comparing Topic Models of Text Collections

  • Nov 01, 2011
  • Patricia J Crossno +3
  • Conference Article

Tag Based Answer Recommendation System

  • Mar 01, 2019
  • Anjali Chauhan +3
  • Conference Article
  • Citations71

TFIDF, LSI and multi-word in information retrieval and text categorization

  • Oct 01, 2008
  • Wen Zhang +2
  • Conference Article
  • Citations12

Normalizing source code vocabulary to support program comprehension and software quality

  • May 18, 2013
  • Latifa Guerrouj
  • Research Article

Topmodpy: A Simple Python Script for Topic Modeling

  • Dec 10, 2020
  • SSRN Electronic Journal
  • Kamakshaiah Musunuru
  • Book Chapter
  • Citations6

Correlation Between K-means Clustering and Topic Modeling Methods on Twitter Datasets

  • Oct 02, 2021
  • Poonam Vijay Tijare +1
  • Book Chapter
  • Citations5

Optimizing People Sourcing Through Semantic Matching of Job Description Documents and Candidate Profile Using Improved Topic Modelling Techniques

  • Aug 14, 2020
  • Lorick Jain +3
  • Conference Article
  • Citations20

LDA-Based Retrieval Framework for Semantic News Video Retrieval

  • Sep 01, 2007
  • Juan Cao +3
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.