Skip to content

Latest commit

 

History

7 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 

Repository files navigation

PhonoGrams research

Papers, notes, and algorithm specs behind the PhonoGrams Go libraries for personal-name matching, aimed at Watchman (OFAC / AML watchlist search):

The first four are character n-gram / edit-distance scorers (Similarity in [0, 1]). Double Metaphone and Beider-Morse are phonetic encodings used as Watchman boosts. They are not token-set metrics (Jaccard on words).

Watchman originally had no first-party Go impl of Editex, positional Kondrak N-SIM (including n=3), Double Metaphone, or Beider-Morse. Those four now live in this org at v0.1.0 and are wired as ?algorithm=editex, nsim, nsim-3, double-metaphone, and beider-morse.

Start here

Doc What it is
docs/algorithms.md Recurrences and cost scales
docs/survey.md Digest of the paper collection and what still matters for Watchman
docs/watchman.md How to plug these into Watchman without regressing Jaro-Winkler
docs/bibliography.md Full list of PDFs in docs/papers/

The two algorithms in one paragraph

Kondrak (SPIRE 2005) defined N-DIST and N-SIM: edit distance / LCS run over n-grams instead of characters, with the first letter repeated so prefixes count. BI-DIST / BI-SIM are the n=2 cases. The FDA's POCA software used BI-SIM on look-alike drug names.

Soft-Bidist (Hadwan et al. 2021) keeps the BI-DIST DP and replaces the positional 0/0.5/1 substitution with a nine-case scale (exact, transpose, first-only, second-only, …). Soft-Bisim (Millán-Hernández et al. 2019) does the same to BI-SIM. Both papers tune the nine weights statistically; the Go packages ship the published best weights, with Soft-Bisim's exact-match weight raised to 1 so identical strings score 1.0.

About

Papers, algorithm specs, and Watchman notes for PhonoGrams name-matching libraries

Topics

Resources

Stars

1 star

Watchers

1 watching

Forks

Contributors