Wearable assistants require episodic memory over egocentric video, yet current vision-language models face bounded frame budgets, growing visual-token costs, and long-context retrieval failures. Under these practical constraints, we study whether textual captions can serve as…
Read the original source — arxiv.org
paper · Shared by tscosj
0 comments
No comments yet.