LLMs handle ordinary typing variation fluently: a typo or missing punctuation leaves both user intent and the model's response substantively unchanged. Yet probes that detect malicious prompts by reading the model's hidden states tell a different story: the same edit rotates the…
Read the original source — arxiv.org
paper · Shared by tscosj
0 comments
No comments yet.