CapMem: A Benchmark for Caption-Based Episodic Memory in Egocentric Video

Wearable assistants require episodic memory over egocentric video, yet current vision-language models face bounded frame budgets, growing visual-token…

Wearable assistants require episodic memory over egocentric video, yet current vision-language models face bounded frame budgets, growing visual-token costs, and long-context retrieval failures. Under these practical constraints, we study whether textual captions can serve as…

Read the original source — arxiv.org

paper · Shared by tscosj

0 comments

No comments yet.