Search papers, labs, and topics across Lattice.
4
0
7
5
OpenEnded, a corpus of English practice speech from Mandarin speakers in open-response tasks, is introduced and it is shown that ALM-generated pseudo-labels improve training over original ALM scoring, while VoxPA achieves the best performance among all baselines.
An on-policy expert-correction pipeline is developed, automated by a meta-level MLE agent, that localizes the failing turn in the weaker model's own rollout and asks the expert to rewrite only that turn, which preserves the model's planning style and combines the gains of harness evolution and model adaptation.
Tool-using LLMs face a near-universal robustness gap, but combining Bayesian Tool Memory with reinforcement learning can boost recovery performance by over 40% in failure scenarios.
Optimizing privacy in LLM interactions reveals that larger models can achieve a superior balance of privacy and utility, while smaller models falter.