This is a crude Python port of the Twokenize class from ark-tweet-nlp.
It produces nearly identical output to the original Java tokenizer, except in a few infrequent situations. In particular, Python does not support partial case-insensitivity in regular expressions and this causes some tokenization differences for ``Eastern" style emoticons, particularly when the left and right halves are of different cases. For example: