Search papers, labs, and topics across Lattice.
9
0
10
0
Evolving rubrics from a single query can dramatically enhance LLM evaluation by eliminating reliance on external annotations and improving answer quality discrimination.
Explicit evidence graphs in VeriGraph enable LLMs to achieve 87.61% claim grounding, transforming how we verify AI-generated conclusions.
Arbor's innovative approach to autonomous research enables a cumulative learning process that outperforms existing models by over 2.5 times in real-world tasks.
LLM agents are shockingly vulnerable to multi-stage "trojan" attacks that inject malicious instructions into their workspace, achieving near-perfect success rates where standard prompt injection defenses fail.
Scaling out peer agents with a shared reasoning hub, AgentFugue, unlocks a new dimension of capability gains in long-horizon tasks, proving that collective reasoning is more than just parallel compute.
Agent-World reveals that self-evolving environments can dramatically boost agent performance, outperforming established models by leveraging dynamic task synthesis.
Forget trajectory-level rollouts: MuSEAgent learns faster and reasons better by distilling past interactions into reusable, state-aware decision experiences.
Observational user feedback, often dismissed as too noisy and biased, can actually power effective RLHF with the right causal modeling, achieving a 49.2% gain on WildGuardMix.
Current multimodal models are stuck in bi-modal interactions, but OmniGAIA and OmniAtlas offer a path towards truly omni-modal AI assistants capable of reasoning and tool use across video, audio, and images.