Application: Longest Common Subsequence
Hey look! Letters!
GTACAC GACATG
What's interesting about those letters?
a) They're parts of a genetic sequence! b) They have a decent amount in common with one another -- not just in the letters themselves, but also the order in which they occur.
In genetic analysis, it's often important to find commonalities in gene sequences for closer examination, but genes are not perfect and the coded sequences are subject to noise:
Genes themselves can have tiny amounts of corruption / noise that prevent finding perfect, sequential matches.
Genes are spliced imperfectly, and we may only find partial matches at different parts of each string.
It turns out that dynamic programming is very commonly applied in this very scenario!
The Longest Common Subsequence (LCS) problem is specified such that, given two Strings \(S_1, S_2\), find the longest sequence of characters that both strings possess in the same order (left to right), though not necessarily in contiguous blocks.
For example, what is the lcs("AXBYCZ", "SATBCU")?
lcs("AXBYCZ", "SATBCU") = "ABC" since that is the longest sequence (i.e. order matters!) of letters that can be found in
both strings, even though there may be other letters in between.
On the surface, this seems like a pretty easy problem -- why is it somewhat tricky?
Because we don't always initially know *where* in each string the longest subsequence is going to begin!
If that problem isn't consumately computer science, I don't know what is!
Consider the following tricky case:
ABCDABADE ^^ ^^^ ^^ ACBDACBDE ^ ^^^ ^^^
The LCS of these two strings is "ABDABDE", but note that there are many different potential-longest substrings, and we need to find the best.
Luckily, we can start to think about applying our dynamic programming tools to sift through these possibilities, beginning, of course, with one key property:
How does the LCS problem exhibit the optimal substructure property?
Consider the simple case of lcs("ABC", "ABC"), in order to know that "ABC" is the LCS of these two Strings, we can observe
that the LCS of any prefix can be added to the LCS of any postfix such that: e.g., lcs("ABC", "ABC") = lcs("AB", "AB") + "C" = smallerSubproblem + oneMoreStep
In other words, we want to compare all pairs of subsequences between the two strings, finding the maximal-length subsequences between each.
So, now that we've identified this as a problem with optimal substructure, let's think about how we might solve it with both bottom-up and top-down DP... starting with what I'd consider the most intuitive.
Formalizing LCS as DP
Before we start filling out tables, let's think about how to structure the LCS problem as one that Dynamic Programming (either bottom-up or top-down) can solve.
As a summary
We want to be able to look at all combinations of left-to-right subsequences between the 2 input Strings.
Of those left-to-right subsequences, we want to find the longest, not knowing where in each String it begins.
LCS Search Intuition
As with Changemaker when we first considered the operation as a sort of hybrid minimax-search tree, we can envision an approach to solving LCS through the lens of search before thinking about how to formalize as a dynamic programming problem.
Note: remember that this is just a conceptual approach to inspire the dynamic programming formalization to come!
Let's think of the pieces that would look search-like here, which can be a bit trickier to analogize than with Changemaker:
State: we have two Strings composing the state, \(R,C\) as in the arguments to the LCS method: \(lcs(R,C)\)
Initial State: \(R,C\) with all original letters in each.
Terminal State: At least one of \(R,C\) reduced to the empty string
""(because the LCS of anything and an empty string is empty too)Actions/Transitions: Work toward the terminals one letter at a time: Examine the last letter in each of \(R,C\) at one state:
If they match, they might be part of the LCS, so pair them up, add them to that path's LCS, remove them from both \(R,C\) and continue on the remainder of each.
If they mismatch, the LCS must ignore one or the other, so branch and take the max length subsequence of the children.
Consider what this would (conceptually) look like in a search-like tree on the subproblem \(lcs(AXB, ABX)\).
What is/are the solution(s) to the subproblem above?
The LCS will be of length 2, either "AX" or "AB"
So, let's now consider how this tree-based representation maps to our dynamic programming by specifying it using our 3 components!
LCS Memoization Table & Ordering
To help us think about the table, let's first have a motivating example and use that to answer some targetted questions:
Consider the memoization table that would be associated with the problem lcs("AXBCZ", "XABZC") in the questions that follow.
Intuition 1: Compared to lcs("AXBCZ", "XABZC"), what would a smaller subproblem look like, and how does that map to the ordering of
problem specific rows / cols?
We can think about each row / col adding a new letter to the substring of previous rows / cols in left-to-right order.
This is a lot like Changemaker where each row added a new coin denomination compared to the row before it.
Intuition 2: What are the SMALLEST two substrings we could solve (think: top-left of the table)?
Kind of a trick question: the empty Strings!.
These will constitute so called "gutters" of the memoization table that will be convenient for the recurrence's base cases!
Intuition 3: What should be recorded in the table's cells (data type and purpose) to find the LCS of two Strings?
Cells contain the number of strings (int) of the longest common subsequence, which can then be traced back to recover the actual String just like in Changemaker!
Draw and then interpret the memoization table that would be employed in solving:lcs("AXBCZ", "XABZC")
Things to note:
Any given cell will find the LCS of the substring of all previous rows / cols that came before it (see optimal substructure highlighted in red box above).
As such, the solution to the original, full problem will (per usual) be located in the bottom-right of the table.
There may be multiple solutions to each LCS problem, each equally valid!
What possible solutions are there to the LCS problem posed above (hint: there are 4)?
lcs("AXBCZ", "XABZC") = "ABC" = "XBC" = "XBZ" = "ABZ"
As such, we should be sensitive that, if we are simply looking for a single solution, then any one of maximal length will suffice, but we could also use this approach to collect *all* solutions (left as an exercise)!
So, with the table format specified, let's consider how to fill it out!
Completing the Table
Recall: we want to record the length of the LCS in the substrings associated with each cell's row and column, after which (similar to with Changemaker) we can "walk back" a solution from the bottom-right of the table.
First, some labels:
Let index \(r\) correspond to letters in the string \(R\) along the rows and index \(c\) of letters in the string \(C\) along the columns.
As such, \(R[r]\) would represent the *new* letter added to substring before \(r\), with the special case of \(R[0] = \emptyset\). E.g., if \(R = AXBCZ"\), \(R[0] = \emptyset\), \(R[1] = A\), \(R[2] = X\), etc.
The answer to any cell of the Table \(T\) can be expressed as \(T[r][c]\)
So, let's start with the easy cells to complete: the base cases.
Which cells will we know the answers to at the start without having to do any sort of computation?
The gutters! The LCS of any String with the empty string MUST be of length 0!
Case 0 - Base Case - Rule: The LCS of any String with the empty String \(\emptyset\) must be 0, formally: $$T[r][c] = 0 ~ \text{if} ~ r == 0~||~c == 0$$
For the rest of the table (i.e., between two non-empty Strings), we can focus on the newly added letter at each row and column substring.
Specifically, to figure out the value of \(T[r][c]\), we can look at the newly added letter at \(R[r]\) and \(C[c]\) and see if they are the same or different!
If they match, pair them up and move on to the rest of the substring.
If they don't, the LCS must ignore one (or both) of them, so consider each branching path.
This is pretty much exactly the distinction between our branching vs. non-branching children in the search-tree intuition above!
Case 1 - Mismatched Letters - Intuition: when the letters at a given row \(r\) and col \(c\) disagree, these are the max nodes in our tree-based conceptual understanding requiring that we:
Take the max of ignoring each of the letters individually.
Phrase this "ignoring" operation as a function of smaller subproblems (again, must be above or to the left of \(r, c\)).
Given that we're after the *longest* common subsequence, how will this intuition translate into a rule to decide the value of the cell when the two letters disagree?
Take the max of the cells above and to the left of it!
Case 1: Mismatched Letters - Rule: \(R[r] \ne C[c] \Rightarrow T[r][c] = max(T[r-1][c], T[r][c-1])\).
Alright, one down -- now the other case!
Case 2 - Matched Letters - Intuition: when the letters at a given row \(r\) and col \(c\) agree, then we can add 1 letter to the LCS of whatever the LCS was to the prefixes before the match.
To fill a cell \(T[r][c]\), where \(R[r] == C[c]\) what cell has the LCS of "whatever the LCS was to the prefixes before matching at row \(r\) and col \(c\)?"
\(T[r-1][c-1]\)! In words: the cell diagonally up and to the left of the matching cell.
Pictorially, this gives us:
Case 2: Matched Letters - Rule: \(R[r] = C[c] \Rightarrow T[r][c] = T[r-1][c-1] + 1\)
Bottom-Up LCS
Nice! So with these rules in place, let's fill out our table!
Awesome! Not too bad now that we know these simple rules... but now, reasonably, we must ask: how do we use this result?
Examining the result above, we can note a couple of facts:
The bottom-right cell will have the maximal number of letters in the LCS to the original problem.
There may be clues to either side of each cell that tell which letters were added to the LCS along a given path, and which were not... let's think about that next!
Gathering the Solution
Since we now have the complete memoization table in-hand, collecting the actual LCS String is not difficult, and once again decomposes into answering 1 question: "Which matched letters are in the LCS?"
Answering this question is fun, and uses the same intuition as for Changemaker. We simply start at the bottom-right of the table, and then walk backwards to our solution.
That said, there's something that we didn't (but could have!) considered with Changemaker: we should be aware of multiple solutions that exist to the LCS problem.
The steps are as follows, and essentially "undo" each of the 2 cases that we used in generating the table's entries:
Obtaining LCS from Memoization Table:
Start at the bottom-right cell of the table.
Undoing Case 1 - Mismatched Letters: If \(R[r] \ne C[c]\) then \(T[r][c]==T[r-1][c]\) OR \(T[r][c]==T[r][c-1]\), so recurse to the cell that has the same value.
Note: if BOTH adjacent cells have the same value as the current, either one are acceptable to recurse to.
Undoing Case 2 - Matched Letters: If \(R[r] == C[c]\), then collect that matched letter as part of the LCS, and then recurse on the top-left cell, \(T[r-1][c-1]\).
Let's see that depicted on our example!
Some things to note here:
We know we're done collecting our solution as soon as we hit a base case: the gutters.
Since we're collecting the LCS letters from the bottom-right, the subsequence we assemble will be in reverse order, and must simply be flipped after collection. Above, we assembled the solution "ABC" though collected those letters in order "CBA"
There were several choice points in the path we took wherein we could have chosen a different path to get a different (but equally valid) solution. To collect ALL possible solutions to an LCS problem, we would recurse at ALL choice points.
And there you have it! Bottom-up LCS laid out before us in all of its splendor!
Top-Down LCS
With the bottom-up approach in pocket, let's now think about the flip with top-down.
Recall: we use the same memoization structure and recurrence in the TD approach, but might be able to save some work by specifically targetting ONLY the subproblems we need.
Those steps are (pictorially):
Start at the largest subproblem (bottom-right of table) and identify which recurrence case is needed to solve it.
For each cell needed by that recursion case, draw an arrow from the cell that needs the subproblem to the one that has the answer.
In a depth-first fashion, you'll discover the value in any cell when ALL outbound arrows from it have those cells / subproblems solved.
Depicted on our running example:
Reminder: Where, in the table above, do solutions to the overlapping subproblems save computation?
Whenever two arrows point INTO the same cell! One of these arrows / recursive calls will have had to compute the answer, and the other will simply find it waiting there.
Above, we trace the needed subproblems to solve each cell, which would eventually give us the following table:
Implementation
When it comes to the implementation details of top-down, remember that this can be implemented recursively!
When choosing the parameters of our recursive helper, we can be clever and not have to use a bunch of substring operations:
Intuition: rather than thinking about the top-down approach of "whittling into smaller prefix substrings" as actual substring operations, we can simply use sliding index parameters that refer to:
The position of the letter in the String.
The corresponding row / col of the String in the bottom-up memoization table!
Because of point 2 above, we need complete only *select* parts of the bottom-up table using top-down for this problem, thus providing us with a space-efficient memoization data structure.
We can therefore parameterize our recursive call as the following:
/** * Completes the memoization table using top-down dynamic programming * @param rStr The String along the memoization table's rows * @param rInd The current letter's index in rStr * @param cStr The String along the memoization table's cols * @param cInd The current letter's index in cStr * @param memo The memoization table */ int lcsRecursiveHelper (String rStr, int rInd, String cStr, int cInd, int[][] memo);
Left as an exercise: how do you specify the recurrence for the top-down helper method described above?
And there you have it! An intuitive binding of search and top-down DP to see two different ways of approaching the same problem.