PHP/ir

Information Retrieval and other interesting topics

Tokenisation

In: special, interest

09 Sep 2009

Taking a string and separating it into tokens is one of those smaller problems in search that seems initially simple - split on spaces - but can quickly become overwhelmed with edge cases. Ignoring the problem of other languages, some of which don't even necessarily use a space, the exceptions tend to fall into two categories, punctuation related and normalisation.