Span-level masked language modeling (MLM) has shown to be advantageous to pre-trained language models over the original single-token MLM, as entities/phrases and their dependencies are critical to language understanding. Previous works only consider span length with some…
Read the original source — arxiv.org
paper · Shared by tscosj
0 comments
No comments yet.