Search papers, labs, and topics across Lattice.
Affiliation:
9
0
11
0
A single adversarial argument can reduce LLM accuracy to near zero, exposing a critical vulnerability in their belief systems.
When faced with noisy tools, state-of-the-art LLMs fail to prevent educators from over-relying on incorrect suggestions, exposing a significant flaw in AI decision support.
Human calibration, rather than just data scale or prompts, is the key to achieving scenario-specific turn-taking in AI dialogues.
Fine-grained credit assignment in multi-agent systems can dramatically boost performance, revealing error sources with unprecedented precision.
Agents struggle to maintain planning accuracy in complex tool ecosystems, with GPT-5.4's performance plummeting from 51.90% to 11.36% under severe blocking conditions.
Seemingly innocuous prompts can covertly hijack robotic actions, steering them toward adversarial outcomes while maintaining the facade of intended commands.
LLMs can guide phoneme editing to create synthetic accented speech from just a handful of examples, substantially improving ASR accuracy where training data is scarce.
Current depression patient simulators are more like Pollyannas than patients, resolving negative emotions too quickly and following a predictable trajectory from negative to positive.
Forget prompt engineering and fine-tuning: this "Reasoning Inception" method injects targeted reasoning into LLM agents at test time to fix conversational errors on the fly.