Beyond Spellchecking

Note that Spelling Correction as intimated by the section above is a procedure only required after we've identified that a word has been misspelled to begin with!

A separate problem is in determining which words have been misspelled, and in a way that is extremely quick and efficient.

Really, this is a special case of a class of problems that share a similar need:

  • How to determine if a giant genetic sequence is a known strand or not?

  • Is someone sending you email in your known contacts list, without having to request that information from the server?

  • Has a URL shortener already allocated a given short-URL to someone else before giving one to you?

  • Have you already swiped left / right on an arbitrary individual on Tinder yet?

All of the above cases seem to suggest what kind of data structure / implementation? Why might they have slightly lighter requirements?

These seem like prime applications of a HashSet testing for membership of some candidate, but unlike a HashSet (which stores the objects to test for membership), we don't necessarily need to store the individual items if we can develop some filter to simply determine whether or not values belong to hypothetical set (i.e., to test for set membership without needing to store the set).

Having such a filter is useful in the same way that a Lawyer having a Receptionist is useful: consulting the Lawyer is expensive, so knowing whether or not they can help you before paying their fees would be big savings (for you, not the Lawyer)!

By the same intuition, sometimes we want to avoid consulting some data source's resources when we can just check a compact version of the database local to our client.

This is usually the case when:

  • The dataset is located in some remote location like a server and roundtrip communication is expensive (either in time or server resources).

  • The dataset is large, and can be memory prohibitive to load all at once for purposes of querying.

As such, we'll explore a neat trick with a special type of data-structure-meets-algorithmic-paradigm today, which are used all over the place!


Intuitions


Before we dive into the specifics of this new approach, we'll generate some intuitions that we can use to springboard into the formalisms.

Let's start with what we're trying to do:

The chief objective at-hand: find a space and time efficient data structure + algorithm for testing set membership without needing to consult the actual set.

Since we already indicated that our objective looks a lot like a HashSet, let's refresh our memory with what we're dealing with...

What was a hash function's purpose in the context of a hash table?

Given some key to store / look up, a hash function provided a means of converting the fields of a given object into an integer index corresponding to a bucket in which to store / find that key.

Recall a HashTable storing Strings with a (pretty lousy) hash function f(s) = s.length() % b for \(b=8\) buckets:


So how can we use this motivation moving forward?

Intuition 1: a hash function \(f\), as we saw in the context of hash tables, provides a semi-unique numerical fingerprint for a particular key based on its data. $$f(key) = index$$

Intuition 2: if we were to hash a key using *multiple* hash functions \(f_1, f_2\), then the resulting tuple of hash indexes will look more unique / result in fewer collisions than any one alone. $$(f_1(key), f_2(key)) = (index_1, index_2)$$

Intuition 3: rather than have buckets that store the keys themselves (which can take up a lot of space, as in a set), what if we merely stored whether or not a contained item had been hashed to that bucket (requiring only 1 bit)?

Putting these intuitions together produces our new approach...



Bloom Filters

In order to accommodate these settings in which the sets and their candidate values may be massive, and to still have savings in time and space over traditional hash tables, we'll explore a new programming paradigm:

Probabilistic programming is a programming paradigm that takes a departure from the absolutes we've been guaranteeing in the past, and instead attempts to create solutions that may not be optimal 100% of the time, but to a high enough degree that some tangible tradeoffs / incentives are obtained.

Bloom Filters are space-efficient, probabilistic data structures used to test set membership, and are used in a swath of different applications with large sets that are read-heavy.

Let's check out the components of Bloom Filters before we examine their theoretical guarantees.


Components


First things first, borrowing from the intuition of the hash table, we'll still maintain state in our Bloom Filters using an array of "buckets", but as opposed to storing our keys directly in those buckets, we'll only need to store an array of bits.

Component 1: an array of some \(m\) bits each initialized to 0, indicating that the Filter is empty.

Component 2: a set of some \(k\) hash functions, each of which map a stored key to one of the \(m\) bits in the array.

Some notes on the above:

  • Choices of both \(m\) and \(k\) will depend on some other properties we'll define later; for now, let's just consider that we have an array of some length \(m\) and some number of hash functions \(k\).

  • Note that these are *bits* stored in each index of the array; by comparison, a single character of a String stored in a HashSet or Trie will be at least 8 bits, which can grow arbitrarily large for arbitrary Strings.

  • Because each "bucket" is so small, typically \(m\) is much much larger than \(k\).

Assuming we have these two pieces, the operations are straightforward.


Operations


Inserting a key into a Bloom Filter is a simple, 2-step process:

  1. For each of the \(k\) hash functions, obtain an index for the key that will be between \([0, m-1]\): $$(f_1(key), f_2(key), ..., f_k(key))$$

  2. Set the bit at each of those found indexes to 1.

Intuitively, setting these bits to 1 is like leaving a "breadcrumb" that we've hashed a key into this position in the past.

Consider hashing two Strings, \(A, B\) into a Bloom Filter with \(m = 8\) bits (very small, but good for illustration), and with \(k=2\) hash functions \(f_1, f_2\) that produce the following indexes:

What do we note that's troubling with the above? Where do we foresee problems down the line?

The two Strings had a collision with one of their hash functions at index 4. This will make it difficult to disentangle which keys were responsible for setting which bits to 1.

To illustrate this problem, let's consider the other primary operation:

Querying a Bloom Filter to determine whether or not a key is contained within is a similar 2-step process to insertion:

  1. For each of the \(k\) hash functions, obtain an index for the key that will be between \([0, m-1]\): $$(f_1(key), f_2(key), ..., f_k(key))$$

  2. Examine the bits at each hashed index in the bit array:

    • If *any* bit is 0: the key is *certainly* not contained within (or else it would've been set to 1 during insertion).

    • If *all* bits are 1: the key is *likely* contained within (though not positive due to possible collisions).

Herein is the cost that Bloom Filter's pay for their spatial parsimony: they can exhibit false positives for certain keys that happen to hash to the set bits of other keys, whether or not they were inserted themselves.

Consider, using the same 8-bit array as above having inserted Strings \(A, B\) that we are now querying for Strings \(C, D\):

Note that if we were to Query any one of Strings \(A, B, C, D\) on the Bloom Filter, we would end up with the following answers and their truth values:

True

False

Positive

Querying \(A, B\) on the filter will give True Positives, because they were indeed stored within, and the bits at their hashed indexes are 1.

Querying \(D\) on the filter will give a False Positive, since it was never stored within, but the bits at its hashed indexes are 1.

Negative

Querying \(C\), on the other hand, yields a True Negative, since one of its hashed index bits is 0.

There will *never* be a false negative using a Bloom Filter.


False Positives


False positives are the primary risk run by using a Bloom Filter, so let's take a deeper look at these.

To build some intuition to start:

What will increase the likelihood of a false positive query? What will decrease it?

The more keys we store, the more bits will be flipped to 1, thus increasing our likelihood of false positives. However, assuming we have good hash functions (evenly distributing keys), a larger number of bits (i.e., \(m\)) will decrease that likelihood.

This means that we can express the likelihood of a false positive in terms of \(m, n, k\):

In a Bloom Filter, the likelihood \(p\) of a false positive is given by the following equation, for \(n\) inserted elements, \(m\) bits in the array, and \(k\) hash functions:

What is the false positive likelihood of a Bloom filter with 8 bits, 2 hash functions, and 2 stored keys?

$$p = (1 - [1 - \frac{1}{8}]^{2*2} )^2 \approx 0.17$$

17%... not too good, which is why we see that increasing \(m\) substantially reduces that likelihood.

For example, doubling our number of available bits yields a better result:

$$p = (1 - [1 - \frac{1}{16}]^{2*2} )^2 \approx 0.05$$

Using the above equation, you can solve for the optimal value of \(m, k\) for a desired false-positive likelihood \(p\), but that is out of scope for this class.


Theoretical Guarantees and Miscellany


Some assorted, concluding properties of Bloom Filters:

  • Time Complexity: never greater than \(O(k)\) for \(k\) presumably fast hash functions that are generally assumed to be \(O(1)\).

  • Space Complexity: exactly \(O(m)\) for the \(m\) bits required to form the bit array.

  • In addition to this very sparse space, Bloom Filters can accommodate a potentially infinite number of stored keys with a fixed size, though the chance of false positives grows with each insertion.

  • There are tons of variants of Bloom Filters used in a variety of different contexts, but most use the above definition as a starting point.

  • One of the earliest Bloom Filters applied for Phone Spell Checking used only 32Kb to store the entire dictionary!


Whew! So that's Bloom Filters in a nutshell... do you feel as though you've bloomed into a new state of understanding?



  PDF / Print