Application: Spelling Correction

By now we've seen a couple of different applications of dynamic programming with Changemaking and Longest Common Subsequence (LCS), but we should look at one last, very applicable one that we're used to dealing with on a daily basis: Spelling Correction

For example, consider the options that are presented by Word when I misspell "intentionally" during the formative days of my thesis creation:


[Brainstorm] How do you think these suggested words are selected, given that the spelling error can occur anywhere in the word?

We would need some metric of "difference" between two strings such that the ones that are of minimal difference from the misspelled are suggested!

It turns out that defining this metric of string differences is a close lemma from our LCS problem, and whose extension is used in a variety of applications.

The Edit Distance between two strings is the minimal number of primitive string-manipulations that are required to turn one string into another.

What do we mean by "primitive string-manipulations" you might ask? Well, to answer that, we can think about what the common mistakes are when mistyping a word:

What are the common mistakes that people make when they misspell words (and no, I definitely didn't have to open the dictionary to make sure I didn't misspell "misspell")?

Several common blunders: (1) adding an extra letter, (2) missing a needed letter, (3) putting the wrong letter where another belonged, and (4) swapping the positions of two adjacent letters.

You might note that (3) and (4) could be composed from (1) and (2), and you'd be write err... right!

It turns out that (3) and (4) would later be added because they best predict over 80% of all spelling mistakes by humans (i.e., from empirical validation) and are also useful in some branches of computational biology (with certain common gene mutations).

As such, the full set of edit distance string operations are: (1) deletion, (2) insertion, (3) replacement, and (4) transposition. Considering each to have a uniform cost, the edit-distance problem is to find the minimal number required to transform one string into another.

  // Deletion
  DRINKE -> DRINK
  
  // Insertion
  MRE -> MORE
  
  // Replacement
  OVEL -> OVAL
  
  // Transposition
  TENE -> TEEN

Of course, all of the above are just single manipulations needed to transform each typo into a word.

In practice, we'll need to string those together for potentially multiple typos.

Consider the edit distance to turn the string "fkc" into (you guessed it) "hack":

$$FKC \rightarrow HKC \rightarrow HAKC \rightarrow HACK$$

The total edit distance dist("FKC", "HACK") here is 3 for the actions: replacement (\(F \rightarrow H\)), insertion \(+A\), then transposition \(KC \rightarrow CK\).

Note also that these could have been accomplished in a different order in general, but we'll see them take a particular form when we translate this to a dynamic programming approach.

Note: Strings that have an edit distance of 0 are a special case of 0-cost replacements where each letter is kept the same!

For example, dist("ABC", "ABC") = 0 consists of 3 "replacements" (in quotes) where each letter is replaced with itself for a cost of 0.

That said, we now have the duty of translating the above intuitions into a workable algorithm.


Edit Distance Memoization Structure


Let me just rip this bandaid off: we can use dynamic programming!

We'll deduce how we envision the memoization structure and define the recurrence, but then the rest will be left to you as an exercise.

Memoization Structure + Ordering: looks almost exactly the same as with LCS: with one "start" String along one axis and the other "destination" String along the other.

The idea is that we'll count the *minimal* number of our 4 primitive operations required to turn the "start" String into the "destination" one.

Let's determine the edit distance between the start string "fkc" and determine its distance from "hack"

The table structure will look the same as with LCS, though with a slightly different base case configuration around the gutters, in which (to make):

Why do the gutters contain the numbers they do?

Because each cell contains the number of letters that would be needed to be removed / added from to get from each string to the empty String.

These gutters now serve as our base cases for the recurrence that follows.


Edit Distance Recurrence


Base Case - Gutters: to convert any String into the empty String, or vice versa, requires that many characters' worth of insertions or deletions: \begin{eqnarray} T[r][c] = \begin{cases} r, & \text{if}~c == 0 \\ c, & \text{if}~r == 0 \\ \end{cases} \end{eqnarray}

With the base case out of the way, we're going to borrow the notation from LCS to think about the recursive cases.

  • \(r, c\) (lowercase) as the numerical index in each col / row.

  • \(R, C\) (uppercase) as the strings along the Rows and Columns, respectively.

  • \(R[r], C[c]\) the "newly added" characters at row r and column c in strings R and C, respectively (just like in LCS).

  • \(R[0, r], C[0, c]\) indicating the prefixes / substrings of length \(r, c\) in each of the Row and Column strings, respectively.

  • \(T[r][c]\) the value in the Table at row r and column c, which should contain an int indicating the minimum number of operations required to turn the Column substring \(C[0, c]\) into the Row substring \(R[0, r]\) (or vice versa, it doesn't really matter since edit distance is symmetrical).


Intuitions for Recurrence:

  • Remember that since we've ordered the table with smaller subproblems above and to the left of any cell, our recurrence must only reference those.

  • Just like in LCS, we're going to examine only the most recently "appended" characters to each of the Row and Col Strings as a way of working backwards to the base case.

These intuitions combine to some recurrence cases that are more intuitive than others; visualized:

Recursive Case: since all of these are options from a single cell, we want to take the minimum edit distance from the edit-path corresponding to each of the 4 string operations, such that: \begin{eqnarray} T[r][c] = min \begin{cases} \text{Deletion Case}, & \text{if}~r \ge 1 \\ \text{Insertion Case}, & \text{if}~c \ge 1 \\ \text{Replacement Case}, & \text{if}~r,c \ge 1 \\ \text{Transposition Case}, & \text{if}~r,c \ge 2 \\ \end{cases} \end{eqnarray}

Let's fill in these values and then depict them in the table that follows!

Case

Description

Recurrence

Insertion

Adding 1 character to C, so add 1 to col from left.

\(T[r][c-1] + 1\)

Deletion

Removing 1 character from R, so add 1 to row above.

\(T[r-1][c] + 1\)

Replacement

Making \(C[c] = R[r]\), which only has some cost if they are not already equal. So below, we use the ternary expression \((C[c] \ne R[r]) ? 1 : 0\), which returns 1 if the condition is true, false otherwise.

\(T[r-1][c-1] + (C[c] \ne R[r]) ? 1 : 0\)

Transposition

Only possible if the two adjacent characters are equal but swapped, in which case it's a single action that then recurses on the rest of the substrings.

\(T[r-2][c-2] + 1 ~ \text{if}~ R[r] == C[c-1] ~\text{and}~ C[c] == R[r-1] \)

Using this recurrence, let's complete a Tabulation / Bottom-up example to the problem \(dist("FKC", "HACK")\)

So there we have it! An edit distance of 3, just as we considered from before.


Extra Practice


Find the edit distance: Dist("WXYYXW", "WYXXYX")

Click for solution!


In practice, many spell correctors will pick an arbitrary edit-distance cutoff of ~2-3 and then only examine those words in the dictionary that can be reached up until that edit distance.

This is why, when you *really* mess up a word, your computer looks at you with a confused red-underline -- knowing that it isn't misspelled but having no real clue for how to correct you because there are so many possibilities, and the computation of which would be pretty brutal!

There are other techniques for spelling correction not covered herein, but edit distance is a good tool to have in your back pocket for many string-distance tasks!

So, onwards to one last consideration with spelling correction and string comparison...



  PDF / Print