Millán-Hernández, García-Hernández, Ledeneva, Hernández-Castañeda, Soft Bigram Similarity to Identify Confusable Drug Names, MCPR 2019 (LNCS 11524).
Kondrak N-SIM / BI-SIM with first-letter affixing (the first character is repeated so prefixes form a bigram):
S[i, 0] = S[0, j] = 0
S[i, j] = max(
S[i-1, j], // skip X
S[i, j-1], // skip Y
S[i-1, j-1] + s(bigram_i, bigram_j)
)
Soft-Bisim(X, Y) = S[n, m] / max(n, m)
This is a similarity (maximize), not a distance (minimize). Sister package soft_bigram is the distance dual.
| Case | Pattern | Paper | DefaultWeights |
|---|---|---|---|
| w8 all equal | aa vs aa | 0.6 | 1.0 |
| w9 exact | ab vs ab | 0.8 | 1.0 |
| w4 transpose | ab vs ba | 0 | 0 |
| w3 / w6 repeats | doubled-letter overlaps | 0.4 / 0.2 | 0.4 / 0.2 |
| w7 first match | a* vs a* | 0.4 | 0.4 |
| w1 second match | *a vs *a | 0 | 0 |
| w5 single cross | one swapped letter | 0 | 0 |
| w2 all different | no shared letters | 0.1 | 0.1 |
DefaultWeights raises w8 and w9 to 1 so identical strings score 1.0. The published genetic-algorithm scale (PaperWeights) does not, because it was trained as a ranker for LASA retrieval, not as a normalized similarity. Watchman needs identity = 1.
Kondrak positional BI-SIM is Bisim() / KondrakWeights: s = (id(a1,b1) + id(a2,b2)) / 2.
Same integration point as Soft-Bidist: inner token pair scorer under BestPairsJaroWinkler. Soft-Bisim is usually the better first try for look-alike aliases (it rewards common subsequences). Soft-Bidist is the better first try for typographical noise (insert/delete/transpose). A blend is what Kondrak used for the FDA POCA system (Avg(Prefix, NED, Bisim, Aline)); replacing Bisim with Soft-Bisim was the MCPR 2019 result.