Repository navigation
Quickstart
This is a small screenshot-rich guide that walks you through a simple plagiarism check with AC2.
- From source (requires Git, a JDK, and Maven installed):
git clone https://github.com/manuel-freire/ac2
cd ac2
mvn clean package
java -jar ac-ui/target/*SNAPSHOT*.jar
note that the last line is bash-specific, and may have to be modified for other shells. If mvn package succeeds, it will build a file with a name similar to ac-ui-2.1.3-SNAPSHOT-efb28.jar in the ac-ui/target folder. The exact name depends on the configured version and the git hash of the latest revision.
- From a .jar file (download the latest release, or build it as in the above step): double-click (if java is correctly installed), or execute
java -jar <name-of-jar>on the command-line.
Once you have launched AC2, you will see a screen similar to that of Fig. 1.

Fig 1.: choosing submissions
The left-hand side (LHS) panel is the input, where you begin, by clicking the +... icon to add folders or zip-files with sources to compare (AC2 supports multiple archive formats).
The right-hand side (RHS) panel contains the output which will later be analysed by AC2, a set of source-code submissions that will be compared. Each submission will have an ID (typically the name of its folder in the left-hand side), and a set of 1 or more files.
The >> button in the LHS allows you to "send" all selected LHS items to the RHS, bypassing any filters. The - button removes selected items, and Clean empties the output. You can even double-click on items to create temporary edits (which are not persisted to the filesystem: use an external editor if you need those; operations in this screen are memory-only: filesystem files or folders will not be modified, except for Extract...).
The center panel allows you specify two filters:
- the top filter controls what gets chosen from the LHS as a submission. In the screenshot, only folders that begin by
p5are considered actual submission. The search is recursive from each of the roots in the LHS. Once a match is found, it is sent onward, and its contents is no longer searched. - the bottom filter controls what parts of each submission will be included in the RHS, and used for similarity comparisons with other submissions. In the screenshot, only
.cfiles are being included, and all others are ignored. Submissions must contain at least 1 file, and can contain any number of other files. Folders (and sub-folders) within submissions will be ignored, but any files present in the RHS will be used in comparisons.
In each of the filter panels, you can specify a whole boolean tree of filters by using the + (one condition) and +... (a set of conditions that are AND-ed or OR-ed together), remove conditions using -, and test them using ? (which selects matches in the corresponding panel). Press Confirm to apply a filter. The top filter is considered to be chained with the bottom filter, so pressing its Confirm button will apply both filters at once.
Conceptually, the 4 panels of this screen form a pipeline:
LHS Top filter Bottom filter RHS
(filesystem ---> (choose ---> (choose ---> (what to
roots) submissions) source files) analyze)
Once the RHS contains the submissions and sources that you want to compare, click on Analyze distances... to compare them. Extract ... simply dumps the contents of the LHS to the folder that you specify; this can save you steps if you need to repeat an analysis or share it with a colleague.
** Beware **: if you try to compare identical (exact same contents, bit-by-bit) submissions, an error will be shown in the terminal screen, and only one of them will be used in actual comparisons.
Once you press on Analyze distances ..., you will be asked to run an analysis on the submissions and sources that were on the RHS of the submission selection screen (see Fig. 2). The default analysis, Zip NCD, uses normalized compression distance (NCD) to detect redundancy (= similarity) between submissions. If AC2 knows the programming language of its source files, they will be tokenized prior to comparison. This will allow AC2 to ignore comments, whitespace and exact identifiers, so that simply renaming variables or altering formatting will not affect similarity calculations. AC2 [supports many programming languages, and more can be added.

Fig 2.: choosing analysis. The default is a good start
After you run an analysis, you can explore its results, looking for overly-similar submissions which point to plagiarism. Or even overly-dissimilar submissions which point to out-of-the-box thinking!

Fig 3.: graph + distance histogram view. Move the slider at the bottom to choose a similarity threshold for the graph
The output of an analysis is a big table of distances: how similar or different each submission is to each other submission. Two submissions with 0 distance are identical (=very similar), while two submissions with a distance of 1 are considered very different (=very dissimilar). AC2 comes with multiple visualizations that allow you to get insights from that big table of distances. You can switch between them using the tabs on the top of the results screen (Fig. 3):
- Individual histograms: for each submission, shows a histogram displaying its distance to all other submissions.
- Graph + distance histogram: choose a similarity threshold in the histogram at the bottom by dragging the slider. This histogram shows the distances between all pairs of submissions. All distances below the chosen threshold will be displayed as a network. In Fig. 3, the threshold of 0.41 has been chosen, so the network displays the submissions for those 6 selected similarities (= network edges).
- Graph + dendrogram: similar to the above, but uses a dendrogram with a slider instead of a histogram.
- Visual table: uses colors instead of numbers to display a two-dimensonal table. Red is very similar, green is very dissimilar. Note that all submissions are identical (=0 distance, red) to themselves. Reordering the submissions (use the combo-box) may reveal clusters.
- Table: lists all pairs of submissions and their similarities as a table. This is the same data that feeds all other tabs.
The goal of AC2's visualizations is to help you decide what is "too similar". To actually identify plagiarism, you will have to look at the code. You can double-click on any edge in a graph, or any submission in the individual histograms tab, or any pair in the table, and you will see a side-by-side comparison of the sources that make each submission (see Fig. 4).
In the submission comparison screen of Fig. 4, you can toggle word-wrap, there is syntax highlighting support for several languages, and you can use additional highlight to mark the largest identical runs of text to make it easy to jump from one to another. Right-clicking on a highlighted run of text brings up a context menu that can scroll the other view to its peer on the other side.

Fig 4.: comparing submissions side-by-side. Highlighting large similar runs is a good idea to quickly locate similarities.
Thanks for reading this quickstart guide.
Please file an issue if you think of ways to improve this documentation.