Standard subword tokenizers either treat every orthographic variation of a word (such as hello, Hello, HELLO, and H\'ello) as unrelated vocabulary entries, which fragments the embedding space, or discard this variation through lossy normalization. We present the Functionalizer,…
Read the original source — arxiv.org
paper · Shared by tscosj
0 comments
No comments yet.