Search papers, labs, and topics across Lattice.
5
0
7
15
Achieving 58.2% accuracy on conversational-memory tasks, ReFind shows that unstructured chat logs can rival complex memory systems when paired with intelligent search controls.
Personalized evaluation rubrics reveal that user satisfaction can diverge significantly from generic quality assessments in role-playing agents.
Current text-to-image models struggle with realistic multi-person interactions, scoring poorly on anatomical accuracy despite high VLM checklist ratings.
BM25 outperforms all competitors at scale, revealing that lexical retrieval is the most efficient choice for large corpus sizes.
Memory systems can recall facts perfectly but fail spectacularly at indirect queries, revealing a critical blind spot in AI memory interfaces.