Search papers, labs, and topics across Lattice.
12
2
12
10
The top-performing systems in multilingual financial question answering are separated by less than one percentage point, showcasing the intense competition and subtlety in model performance.
Systems achieved up to 97.5% accuracy in multilingual financial question answering, revealing the potential for high-performance AI across diverse languages.
A unified framework reveals that the choice of tokenization and vocabulary topology can significantly influence the performance of discrete diffusion models, unlocking new avenues for optimization.
BPO achieves up to 6.1% higher success rates in sandbox-native RL tasks while cutting down on the number of required policy updates by 38%.
Mandate Salience Decay can lead to a 4.4x behavioral gap in financial agents over time, revealing critical vulnerabilities in their long-term deployment.
Grounding item representations in user behavior can dramatically elevate the accuracy of sequential recommendations, bridging the gap between semantic understanding and real-world interactions.
Set representation models can be made robust to inference-time corruptions like outliers and missing data by training against a learned barycentric adversary.
LLMs can achieve more consistent and reliable cross-jurisdictional financial reporting by acting as constrained verifiers within a structured, agentic workflow, rather than as free-form generators.
LLM agents can now autonomously generate complex skills with multi-file dependencies, rivaling human-authored skills, thanks to a co-evolutionary verification process that doesn't need ground truth labels.
Even state-of-the-art LLMs struggle to adapt to mid-task changes in long-horizon web navigation, highlighting a critical gap in their ability to handle realistic user interactions.
Diffusion language models can achieve better reasoning performance by explicitly balancing generation quality and exploration, outperforming methods that prioritize only one.
Forget brute-force distillation: this method uses pedagogical principles to distill LLMs, boosting student model performance on complex reasoning tasks by up to 22.3%.