Search papers, labs, and topics across Lattice.
Affiliation:
3
0
6
Instead of relying on brittle prompt histories or rigid hand-coded workflows, LLM agents can now autonomously construct, debug, and evolve their own graph-structured execution policies purely from contrastive trial and error.
PoTRE's innovative use of diverse reasoning agents leads to a remarkable 49.92% accuracy on Humanity's Last Exam, setting a new benchmark for LLM performance in complex reasoning tasks.
MLLMs are riddled with shared vulnerabilities across modalities, meaning a single weakness can be exploited to jailbreak safety filters, hijack instructions, or even poison training data.