The Functionalizer: Lossless Functional Decomposition for Subword Tokenization

Standard subword tokenizers either treat every orthographic variation of a word (such as hello, Hello, HELLO, and H\'ello) as unrelated vocabulary…

Standard subword tokenizers either treat every orthographic variation of a word (such as hello, Hello, HELLO, and H\'ello) as unrelated vocabulary entries, which fragments the embedding space, or discard this variation through lossy normalization. We present the Functionalizer,…

Read the original source — arxiv.org

paper · Shared by tscosj

0 comments

No comments yet.