Search papers, labs, and topics across Lattice.
17
15
13
22
Static profiles lead to identity essentialism in LLMs, but a new longitudinal memory framework reveals a path to richer, more diverse social simulations.
Ranking systems can achieve greater accuracy and stability by leveraging structural relationships in permutations, as demonstrated by the SRPO framework.
Evaluator score gaps are the secret sauce for optimizing LLM policies, and DynamicRubric turns this insight into a powerful co-evolution framework that outshines existing methods.
Generative retrieval can revolutionize statute retrieval by effectively bridging the gap between everyday legal language and formal statutes.
Despite high retrieval rates, LLMs miss over 47% of relevant studies in meta-analysis, exposing a critical gap in their systematic reasoning capabilities.
ReGrad enables scalable and reversible knowledge injection without the risk of catastrophic forgetting, outperforming traditional methods in both general and domain-specific tasks.
SKIM compresses procedural skills in LLMs by 30-60% without sacrificing performance, revolutionizing how we manage reusable natural language skills.
LexRubric reveals that even state-of-the-art LLMs struggle with open-ended legal tasks, exposing critical gaps in their contextual understanding and reasoning abilities.
Reliable civil court judgments can now be simulated with a framework that adapts to the complexities of legal claims and remedies.
Untangling task-solving skills from factual knowledge in PRAG adapters makes them play better together, boosting performance when you combine multiple documents.
Explicitly enumerating skills in-context doesn't scale for agentic LLMs, but retrieving skills on demand can substantially improve performance – if the LLM can figure out when and which skill to load.
Humans are still way better than LLMs at trial-and-error problem solving, and this new dataset of human problem-solving trajectories shows us why.
Injecting demographic attributes directly into LLM hidden states can drastically improve the diversity and realism of public opinion simulations.
Current search paradigms fall short for analytical tasks, motivating a new "analytical search" framework that treats search as an evidence-driven, multi-step reasoning process.
LLMs still can't convincingly mimic human personas, especially when it comes to syntactic style and memory, despite advancements in other areas.
LLMs still struggle to learn effectively from user feedback during service, as revealed by a new benchmark spanning multiple domains and languages.
LLMs still struggle to synthesize coherent scientific surveys, as evidenced by a new benchmark revealing significant performance gaps even with advanced agentic frameworks.