- Research Article
64
- 10.1016/j.scico.2014.11.010
Theory-oriented software engineering
- Nov 27, 2014
- Science of Computer Programming
- Klaas-Jan Stol + 1 more +1
Theory-oriented software engineering
The availability of open source software projects has created an enormous opportunity for empirical evaluations in software engineering research. However, this availability requires that researchers judiciously select an appropriate set of evaluation targets and properly document this rationale. This selection process is often critical as it can be used to argue for the generalizability of the evaluated tool or method. To understand the selection criteria that researchers use in their work we systematically read 55 research papers appearing in six major software engineering conferences. Using a grounded theory approach we iteratively developed a codebook and coded these papers along five different dimensions, all of which relate to how the authors select evaluation targets in their work. Our results indicate that most authors relied on qualitative and subjective features to select their evaluation targets. Building on these results we developed a tool called RepoGrams, which supports researchers in comparing and contrasting source code repositories of multiple software projects and helps them in selecting appropriate evaluation targets for their studies. We describe RepoGrams’s design and implementation, and evaluate it in two user studies with 74 undergraduate students and 14 software engineering researchers who used RepoGrams to understand, compare, and contrast various metrics on source code repositories. For example, a researcher interested in evaluating a tool might want to show that it is useful for both software projects that are written using a single programming language, as well as ones that are written using dozens of programming languages. RepoGrams allows the researcher to find a set of software projects that are diverse with respect to this metric.
Theory-oriented software engineering
Theory-oriented software engineering
Influences on the design of exception handling ACM SIGSOFT project on the impact of software engineering research on programming language design
There has long been a close association between research in software engineering and the design of programming languages. Part of the IMPACT project involves an exploration of the interrelations of these two fields and documentation in a report of how fundamental research in software engineering has been a valuable resource for programming language features commonly used today. The resulting report investigates the relationship by considering features in currently used languages, including exceptions, control and data abstractions, types, inheritance, concurrency and visualization mechanisms.This paper, exerpted from the report, focuses on the influence of software engineering research on the development of exceptions. The paper demonstrates that there is a symbiotic relationship between software engineering research and the design of exception handing in programming languages. Publication of these partial results is aimed at soliciting feedback and comments from both the programming languages and software engineering communities.
Read moreTowards Knowledge Evolution in Software Engineering
This article presents an epistemological reading of knowledge evolution in software engineering (SE) both within a software project and into SE theoretical frameworks principally modeling languages and software development life cycles (SDLC). The article envisages SE as an artificial science and notably points to the use of iterative development as a more adequate framework for the enterprise applications. Iterative development has become popular in SE since it allows a more efficient knowledge acquisition process especially in user intensive applications by continuous organizational modeling and requirements acquisition, early implementation and testing, modularity,… SE is by nature a human activity: analysts, designers, developers and other project managers confront their visions of the software system they are building with users’ requirements. The study of software projects’ actors and stakeholders using Simon’s bounded rationality points to the use of an iterative development life cycle. The later, indeed, allows to better apprehend their rationality. Popper’s knowledge growth principle could at first seem suited for the analysis of the knowledge evolution in the SE field. However, this epistemology is better adapted to purely hard sciences as physics than to SE which also takes roots in human activities and by the way in social sciences. Consequently, we will nuance the vision using Lakatosian epistemology notably using his falsification principle criticism on SE as an evolving science. Finally the authors will point to adaptive rationality for a lecture of SE theorists and researchers’ rationality.
Read moreAssessing the Credibility of Grey Literature: A Study with Brazilian Software Engineering Researchers
In recent years, the use and investigations about Grey Literature (GL) increased, in particular, in Software Engineering (SE) research. However, its understanding is still scarce and sometimes controversial, such as interpreting GL types and assessing their credibility. This study aimed to understand the credibility aspects that SE researchers consider in assessing GL and its types. To achieve this goal, we surveyed 53 SE researchers (who answered that they have used GL in our previous investigation), receiving a total of 34 valid responses. Our main findings show that: 1) GL source produced or cited by a renowned source is the main credibility criteria used to assess GL, 2) most of the GL types tend to have a Low to Moderate level of Control and Expertise, 3) there is a positive statistical correlation between the level of Control and Expertise for most GL types, and 4) the different respondent profiles shared similar opinions about the credibility criteria. Our investigation contributes to helping future SE researchers that intend to use GL with more credibility. Additionally, shows the need for future studies to better understand the GL types in SE research.
Read moreThere is no random sampling in software engineering research
Representative sampling is considered crucial for predominately quantitative, positivist research. Researchers typically argue that a sample is representative when items are selected randomly from a population. However, random sampling is rare in empirical software engineering research because there are no credible sampling frames (population lists) for the units of analysis software engineering researchers study (e.g. software projects, code libraries, developers, projects). This means that most software engineering research does not support statistical generalization, but rejecting any particular study for lack of random sampling is capricious.
Read moreA Metrics Suite for Static Structure of Large-Scale Software Based on Complex Networks
In order to overcome the limitations of traditional software metrics and meet the urgent demand in large- scale software development, in this study we constructed the static structure model for large-scale software by structural mapping and visualization. In particular, through application of the complex networks theory, we proposed a two-dimension metrics suite based on complex networks parameters from the perspective of software engineering. The feasibility and validity of the metrics were testified through statistical analyzing the data from a set of software projects; and these experiments show that the metrics suite can not only uncover the features of software structure such as scale free, small world, but also can get some latent rules about design and structure. These features can be conveniently used by designers and developers to understand the system, to control complexity and to improve process.
Read moreFrom Aristotle to Ringelmann: a large-scale analysis of team productivity and coordination in Open Source Software projects
Complex software development projects rely on the contribution of teams of developers, who are required to collaborate and coordinate their efforts. The productivity of such development teams, i.e., how their size is related to the produced output, is an important consideration for project and schedule management as well as for cost estimation. The majority of studies in empirical software engineering suggest that - due to coordination overhead - teams of collaborating developers become less productive as they grow in size. This phenomenon is commonly paraphrased as Brooks' law of software project management, which states that "adding manpower to a software project makes it later". Outside software engineering, the non-additive scaling of productivity in teams is often referred to as the Ringelmann effect, which is studied extensively in social psychology and organizational theory. Conversely, a recent study suggested that in Open Source Software (OSS) projects, the productivity of developers increases as the team grows in size. Attributing it to collective synergetic effects, this surprising finding was linked to the Aristotelian quote that "the whole is more than the sum of its parts". Using a data set of 58 OSS projects with more than 580,000 commits contributed by more than 30,000 developers, in this article we provide a large-scale analysis of the relation between size and productivity of software development teams. Our findings confirm the negative relation between team size and productivity previously suggested by empirical software engineering research, thus providing quantitative evidence for the presence of a strong Ringelmann effect. Using fine-grained data on the association between developers and source code files, we investigate possible explanations for the observed relations between team size and productivity. In particular, we take a network perspective on developer-code associations in software development teams and show that the magnitude of the decrease in productivity is likely to be related to the growth dynamics of co-editing networks which can be interpreted as a first-order approximation of coordination requirements.
Read morePassages
Federico Biancuzzi and Shane Warden's Masterminds of Programming: Conversations with the Creators of Major Programming Languages is a treasure. The book consists of interviews with the creators of, in order, C++, Python, APL, Forth, BASIC, AWK, Lua, Haskell, ML, SQL, Objective C, Java, C#, UML, Perl, PostScript, and Eiffel. Each chapter asks similar, but not identical questions, and the above-mentioned masterminds, including Larry Wall, James Gosling, Brian Kernighan, Bertrand Meyer, Robin Milner, Simon Peyton-Jones, Guido van Rossum, and Bjarne Stroustrop give a wide variety of answers. Some of the masterminds are charming, and many are contentious, even cranky; they are also, almost all, full of deep insights into the deepest problems of software engineering. This insight comes in two forms; first, programming languages are the mechanisms by which software engineering solutions are almost always produced. Second, perhaps even more importantly, creating and evolving a widely-used programming language is a heroic, Herculean, critical software engineering task. All of these masterminds have succeeded in a massive software engineering task; they are not mere ivory tower thinkers about software engineering, but have, in some cases, entire lives shaped by a single, extremely complex, software project. More on that key point below.
Read moreA risk based economical approach for evaluating software project portfolios
Software engineers have been applying economical concepts to shed light upon the value-related aspects of software development processes. Based on credit risk analysis concepts, we present an approach to estimate the probability distribution of losses and earnings that can be incurred by a software development organization according to its software project portfolio. Such approach is built upon an analogy that compares software projects to unhedged loans issued to unreliable borrowers. As loans may not be paid back, software projects may fail, leading their development organizations to losses. By applying this approach, an organization may estimate the variability of its expected profits related to a set of software projects. Initial calibrating data were acquired by accomplishing an experimental study.
Read moreOn the Use of Grey Literature
Background: The use of Grey Literature (GL) has been investigated in diverse research areas. In Software Engineering (SE), this topic has an increasing interest over the last years. Problem: Even with the increase of GL published in diverse sources, the understanding of their use on the SE research community is still controversial. Objective: To understand how Brazilian SE researchers use GL, we aimed to become aware of the criteria to assess the credibility of their use, as well as the benefits and challenges. Method: We surveyed 76 active SE researchers participants of a flagship SE conference in Brazil, using a questionnaire with 11 questions to share their views on the use of GL in the context of SE research. We followed a qualitative approach to analyze open questions. Results: We found that most surveyed researchers use GL mainly to understand new topics. Our work identified new findings, including: 1) GL sources used by SE researchers (e.g., blogs, community website); 2) motivations to use (e.g., to understand problems and to complement research findings) or reasons to avoid GL (e.g., lack of reliability, lack of scientific value); 3) the benefit that is easy to access and read GL and the challenge of GL to have its scientific value recognized; and 4) criteria to assess GL credibility, showing the importance of the content owner to be renowned (e.g., renowned author and institutions). Conclusions: Our findings contribute to form a body of knowledge on the use of GL by SE researchers, by discussing novel (some contradictory) results and providing a set of lessons learned to both SE researchers and practitioners.
Read moreEmpirical Research in Software Engineering — A Literature Survey
Empirical research is playing a significant role in software engineering (SE), and it has been applied to evaluate software artifacts and technologies. There have been a great number of empirical research articles published recently. There is also a large research community in empirical software engineering (ESE). In this paper, we identify both the overall landscape and detailed implementations of ESE, and investigate frequently applied empirical methods, targeted research purposes, used data sources, and applied data processing approaches and tools in ESE. The aim is to identify new trends and obtain interesting observations of empirical software engineering across different sub-fields of software engineering. We conduct a mapping study on 538 selected articles from January 2013 to November 2017, with four research questions. We observe that the trend of applying empirical methods in software engineering is continuously increasing and the most commonly applied methods are experiment, case study and survey. Moreover, open source projects are the most frequently used data sources. We also observe that most of researchers have paid attention to the validity and the possibility to replicate their studies. These observations are carefully analyzed and presented as carefully designed diagrams. We also reveal shortcomings and demanded knowledge/strategies in ESE and propose recommendations for researchers.
Read moreOn IT and SwE Research Methodologies and Paradigms
In this chapter, the authors review the landscape of research methodologies and paradigms available for Information Technology (IT) and Software Engineering (SwE). The aims of the chapter are two-fold: (i) create awareness in current research communities in IT and SwE on the variety of research paradigms and methodologies, and (ii) provide an useful map for guiding new researchers on the selection of an IT or SwE research paradigm and methodology. To achieve this, the chapter reviews the core IT and SwE research methodological literature, and based on the findings, the authors illustrate an updated IT and SwE research framework that comprehensively integrates findings and best practices and provides a coherent systemic (holistic) view of this research landscape.
Read moreGuest editorial: special section on mining software repositories
Mining software repositories is an increasingly popular and important area of software engineering research aimed at retrieving, integrating, and analyzing data available in various kinds of software repositories—such as version control systems, issue trackers, mailing lists, and user reviews, —to distill useful information about software projects and, whenever possible, to provide recommendations to software engineers. This area of research has been especially active in the last 10-15 years, and researchers have contributed to this area (i) by developing techniques and tools to mine software repositories, (ii) conducting various kinds of evolutionary studies on open source and industrial projects, and (iii) proposing approaches and tools that exploit data from software repositories to provide actionable support to software engineers. Noticeably, the large amount of data available in software repositories—as well as of techniques developed to mine such data—has allowed researchers to perform a lot of empirical research. Three noticeable examples of empirical research in the area of mining software repositories are reported in this special section. In the article “Towards Improving Statistical Modelling of Software Engineering Data: Think Locally, Act Globally!”, the authors study and discuss pros and cons of local defect prediction models—i.e., of models built on subsets of homogeneous data—and global defect prediction models, built on the entire available dataset. Also, they show how it is possible to build a hybrid approach combining local and global models together. The article “Understanding the Impact of Rapid Releases on Software Quality The Case of Firefox” investigates how quality assurance in software projects is impacted when applying rapid release cycles. Specifically, the authors study projects with regular and rapid
Read moreWhat Is the Process? A Metamodel of the Requirements Elicitation Process Derived from a Systematic Literature Review
Requirements elicitation is a fundamental process in software engineering, essential for aligning software products with user needs and project objectives. As software projects become more complex, effective elicitation methods are vital for capturing accurate and comprehensive requirements. Despite the variety of available elicitation methods, practitioners face persistent challenges such as capturing tacit knowledge, managing diverse stakeholder needs, and addressing ambiguities in requirements. Moreover, although elicitation is recognized as a core process for gathering and analyzing system objectives, there is a lack of a unified and systematic framework to guide practitioners—especially newcomers—through the activity. To address these challenges, we provide a comprehensive analysis of existing elicitation methods, aiming to contribute to better alignment between software products and project objectives, ultimately improving software engineering practices. We do so by performing a systematic literature review identifying crosscutting steps, common techniques, tools, and approaches that define the core activities of the elicitation process. We synthesize our findings into a metamodel that structures software elicitation processes. This review uncovers various elicitation methods—such as collaborative workshops, interviews, and prototyping—each demonstrating unique strengths in different project contexts. It also highlights significant limitations, including stakeholder misalignment and incomplete requirements capture, which continue to reduce the effectiveness of elicitation processes. Finally, our study seeks to contribute to understanding requirements elicitation methods by providing a comprehensive view of their current strengths and limitations through a metamodel enabling the structuring and optimization of elicitation processes.
Read moreAutomated Static Analysis Tools: A Multidimensional view on Software Quality Evolution
Software use is ubiquitous. The quality and the evolution of quality over long periods of time is therefore of notable importance. Software engineering research investigates software quality in multiple areas. One of these areas are predictive models, in which measurements of past changes to the source code or file contents are used to assess the quality of changes, files or even product releases. However, these predictive models have yet to transition from research to practice on a larger scale. In contrast, Automated Static Analysis Tools (ASATs) are used in practice and are also part of several software quality models. ASATs are able to warn developers about parts of the source code that violate best practices or match common defect patterns. One downside of ASATs are false positives, i.e., warnings about parts of the code which are not problematic. Developers have to manually assess the warnings and annotate the code or the ASAT configuration to mitigate this. Within this thesis, we investigate the evolution of software quality with a focus on a general purpose ASAT for Java. Our main objective is to determine if the use of an ASAT can improve software quality, as measured by defects, significantly enough to mitigate additional effort by the developers to use the ASAT. We combine multiple software engineering research techniques and data validation studies to improve the signal-to-noise ratio to increase the validity and stability of our results. We focus on a general purpose ASAT for the Java programming language due to the maturity of the language and the large number of projects available for this language. Both the language and the general purpose ASAT have been available for a long time, which allows us to include longer periods of time for our analyses. We study how the ASAT is applied, how the generated warnings evolve over long time periods, and how it affects the quality of the source code in terms of defects. In addition, we include the perspective of the developers regarding software quality improvement by measuring changes when developers intend to improve the quality of the source code. Our studies yield surprising insights. While our results show that ASATs have a positive impact on software quality, the magnitude of the impact is much smaller than expected. Moreover, we can show that corrective changes are the main driver of complexity in software projects. They introduce more complexity than feature additions or any other type of maintenance. In addition, we find that software quality estimation models benefit more from size and complexity metrics than static analysis warnings of an ASAT. Our study of developer intents to increase software quality mirrors this result.
Read more