Project details for Harry

Screenshot Harry 0.3.2

by konrad - November 19, 2014, 20:24:21 CET [ Project Homepage BibTeX Download ]

view ( today), download ( today ), 0 subscriptions


Harry is a small tool for comparing strings and measuring their similarity. The tool supports several common distance and kernel functions for strings as well as some excotic similarity measures. The focus of Harry lies on implicit similarity measures, that is, comparison functions that do not give rise to an explicit vector space. Examples of such similarity measures are the Levenshtein distance and the Jaro-Winkler distance.

Harry is implemented using OpenMP, such that the computation time for a set of strings scales linear with the number of available CPU cores. Moreover, efficient implementations of several similarity measures, effective caching of similarity values and low-overhead locking further speedup the computation.

Harry complements the tool Sally that embeds strings in a vector space and allows computing vectorial similarity measures, such as the cosine distance and the bag-of-words kernel.

A tutorial is available here:

Changes to previous version:

Several minor bugfixes.

BibTeX Entry: Download
Supported Operating Systems: Linux, Unix, Posix, Mac Os X
Data Formats: Svmlight, Binary, Txt
Tags: Sequence Analysis, String Kernels, Similarity Measures, String Distances
Archive: download here


No one has posted any comments yet. Perhaps you'd like to be the first?

Leave a comment

You must be logged in to post comments.