Automatically identifying commits that induce fixes is an important task, as it enables researchers to quickly and efficiently validate many types of software engineering analyses, such as software metrics or models for predicting faulty components. Previous work on SZZ, an algorithm designed by Sliwerski et al and improved upon by Kim et al, provides a process for automatically identifying the fix-inducing predecessor lines to lines that are changed in a bug-fixing commit. However, as of yet no one has verified that the fix-inducing lines identified by SZZ are in fact responsible for introducing the fixed bug. Also, the SZZ algorithm relies on annotation graphs, which are imprecise in the face of large blocks of modified code, for back-tracking through previous revisions to the fix-inducing change.
In this work we outline several improvements to the SZZ algorithm: First, we replace annotation graphs with line-number maps that track unique source lines as they change over the lifetime of the software; and second, we use DiffJ, a Java syntax-aware diff tool, to ignore comments and formatting changes in the source. Finally, we begin verifying how often a fix-inducing change identified by SZZ is the true source of a bug.
[1]
Andreas Zeller,et al.
Mining version archives for co-changed lines
,
2006,
MSR '06.
[2]
Andreas Zeller,et al.
When do changes induce fixes?
,
2005,
ACM SIGSOFT Softw. Eng. Notes.
[3]
Thomas Zimmermann,et al.
Automatic Identification of Bug-Introducing Changes
,
2006,
21st IEEE/ACM International Conference on Automated Software Engineering (ASE'06).
[4]
Jaime Spacco,et al.
Branching and merging in the repository
,
2008,
MSR '08.
[5]
Gerardo Canfora,et al.
Identifying Changed Source Code Lines from Version Repositories
,
2007,
Fourth International Workshop on Mining Software Repositories (MSR'07:ICSE Workshops 2007).