Building a biomedical tokenizer using the token lattice design pattern and the adapted Viterbi algorithm Public Deposited

Downloadable Content

Download PDF
  • Weber-Jahnke, Jens
    • Other Affiliation: Department of Computer Science, University of Victoria, Victoria, Canada
  • Barrett, Neil
    • Other Affiliation: Department of Computer Science, University of Victoria, Victoria, Canada
  • Abstract: Background: Tokenization is an important component of language processing yet there is no widely accepted tokenization method for English texts, including biomedical texts. Other than rule based techniques, tokenization in the biomedical domain has been regarded as a classification task. Biomedical classifier-based tokenizers either split or join textual objects through classification to form tokens. The idiosyncratic nature of each biomedical tokenizer’s output complicates adoption and reuse. Furthermore, biomedical tokenizers generally lack guidance on how to apply an existing tokenizer to a new domain (subdomain). We identify and complete a novel tokenizer design pattern and suggest a systematic approach to tokenizer creation. We implement a tokenizer based on our design pattern that combines regular expressions and machine learning. Our machine learning approach differs from the previous split-join classification approaches. We evaluate our approach against three other tokenizers on the task of tokenizing biomedical text. Results: Medpost and our adapted Viterbi tokenizer performed best with a 92.9% and 92.4% accuracy respectively. Conclusions: Our evaluation of our design pattern and guidelines supports our claim that the design pattern and guidelines are a viable approach to tokenizer construction (producing tokenizers matching leading custom-built tokenizers in a particular domain). Our evaluation also demonstrates that ambiguous tokenizations can be disambiguated through POS tagging. In doing so, POS tag sequences and training data have a significant impact on proper text tokenization.
Date of publication
  • doi:10.1186/1471-2105-12-S3-S1
Resource type
  • Article
Rights statement
  • In Copyright
Rights holder
  • Neil Barrett et al.; licensee BioMed Central Ltd.
  • English
Is the article or chapter peer-reviewed?
  • Yes
Bibliographic citation
  • BMC Bioinformatics. 2011 Jun 09;12(Suppl 3):S1
  • Open Access
  • BioMed Central Ltd

This work has no parents.