arXiv:2609.13149v1 Announce Type: new Abstract: For local large language model agents, active context is a scarce resource: memory capacity, prefill latency, cache growth, and service objectives all constrain how many input tokens each call can afford. We present BudgetBench, an…
Read the original source — arxiv.org
paper · Shared by tscosj
0 comments
No comments yet.