yang-tsay-chan-2002-longest-match-segmentation
IN premise — summaries/2026-08-24/wiki-Tokenization_lexical_analysis-chunk-2.md
Created 2026-08-24T17:11:23+00:00
Yang, Tsay, and Chan (2002) formally analyzed the applicability and limitations of the longest-match rule for dictionary-based word segmentation, published in Computer Languages, Systems & Structures, Vol. 28(3), pp. 273–288.
Summary
This entry records that a peer-reviewed 2002 study mapped out exactly where the "greedily pick the longest matching word" strategy works well and where it produces wrong splits, giving any segmentation system a published baseline for knowing its own failure modes. It matters because it pins down the known limits of the simplest common approach before more complex methods are justified.