Skip to content

Latest commit

 

History

History
40 lines (29 loc) · 1.88 KB

File metadata and controls

40 lines (29 loc) · 1.88 KB

Soft-Bisim

Millán-Hernández, García-Hernández, Ledeneva, Hernández-Castañeda, Soft Bigram Similarity to Identify Confusable Drug Names, MCPR 2019 (LNCS 11524).

Recurrence

Kondrak N-SIM / BI-SIM with first-letter affixing (the first character is repeated so prefixes form a bigram):

S[i, 0] = S[0, j] = 0
S[i, j] = max(
    S[i-1, j],                              // skip X
    S[i,   j-1],                            // skip Y
    S[i-1, j-1] + s(bigram_i, bigram_j)
)
Soft-Bisim(X, Y) = S[n, m] / max(n, m)

This is a similarity (maximize), not a distance (minimize). Sister package soft_bigram is the distance dual.

Scale s (w1–w9)

Case Pattern Paper DefaultWeights
w8 all equal aa vs aa 0.6 1.0
w9 exact ab vs ab 0.8 1.0
w4 transpose ab vs ba 0 0
w3 / w6 repeats doubled-letter overlaps 0.4 / 0.2 0.4 / 0.2
w7 first match a* vs a* 0.4 0.4
w1 second match *a vs *a 0 0
w5 single cross one swapped letter 0 0
w2 all different no shared letters 0.1 0.1

DefaultWeights raises w8 and w9 to 1 so identical strings score 1.0. The published genetic-algorithm scale (PaperWeights) does not, because it was trained as a ranker for LASA retrieval, not as a normalized similarity. Watchman needs identity = 1.

Kondrak positional BI-SIM is Bisim() / KondrakWeights: s = (id(a1,b1) + id(a2,b2)) / 2.

Watchman

Same integration point as Soft-Bidist: inner token pair scorer under BestPairsJaroWinkler. Soft-Bisim is usually the better first try for look-alike aliases (it rewards common subsequences). Soft-Bidist is the better first try for typographical noise (insert/delete/transpose). A blend is what Kondrak used for the FDA POCA system (Avg(Prefix, NED, Bisim, Aline)); replacing Bisim with Soft-Bisim was the MCPR 2019 result.