Search papers, labs, and topics across Lattice.
Affiliation:
4
0
7
Instead of relying on brittle prompt histories or rigid hand-coded workflows, LLM agents can now autonomously construct, debug, and evolve their own graph-structured execution policies purely from contrastive trial and error.
Machine translation benchmarks have functionally saturated, but pairing human-authored failure cases with deterministic verification rules reveals critical multimodal blind spots that automated metrics consistently miss.
ClinEnv reveals that LLMs struggle significantly with management decisions in clinical scenarios, achieving only 0.17 F1 for these critical actions despite better performance in diagnosis.
Achieve SOTA medical image segmentation with barely any labels by combining fine-grained visual features with text-based semantic guidance.