Search papers, labs, and topics across Lattice.
Mohamed bin Zayed University of Artificial Intelligence, Abu Dhabi, United Arab Emirates
24
19
17
49
Models that seem equally accurate can drastically differ in their ability to clarify ambiguous requests, impacting user experience and efficiency.
Systems achieved up to 97.5% accuracy in multilingual financial question answering, revealing the potential for high-performance AI across diverse languages.
The top-performing systems in multilingual financial question answering are separated by less than one percentage point, showcasing the intense competition and subtlety in model performance.
Multimodal unlearning could revolutionize how we handle sensitive data in AI, enabling targeted removal without sacrificing model performance.
MIRAGE restores factuality in long-form RAG systems, even when faced with heavily polluted retrieval data, outperforming prior methods.
Manipulating just a few sparse neurons can steer Arabic LLMs to generate dialect-specific outputs without the need for extensive fine-tuning.
LLMs can exhibit a staggering 68.6% mismatch between their reasoning and final diagnostic outputs, raising serious concerns about their reliability in clinical settings.
Mandate Salience Decay can lead to a 4.4x behavioral gap in financial agents over time, revealing critical vulnerabilities in their long-term deployment.
Inductive biases can significantly enhance performance in AI tasks with long feedback loops, challenging the dominance of purely data-driven approaches.
Bayesian control outperforms traditional orchestration methods, especially when verification costs are high, by providing a more nuanced understanding of candidate correctness.
Cross-language safety evaluations reveal that LLMs exhibit starkly different risk profiles in Bulgarian compared to German, challenging the notion of universal model safety.
Representation choice can drastically alter model performance in table understanding tasks, revealing that structured text often outperforms rendered images.
EvoNote outperforms human-generated health notes 89.6% of the time while slashing correction production time from hours to minutes.
Forget the heavy transformers: surprisingly effective LLM-generated code detection can be achieved with lightweight stylometric features and decision trees, offering near-instant inference.
Forget Shakespeare, LLMs can now sling verses in Arabic dialects, thanks to a new dataset for instruction-guided poetry generation.
Arabic LLMs can speak the language of finance, but they often fail to reason about it, especially when it comes to causality and generation.
Forget tedious multi-turn dialogues: Co-FactChecker's "trace-editing" lets human experts directly shape an LLM's reasoning process, leading to higher quality claim verification.
LLMs may nail the Text-to-SQL execution accuracy, but SQLStructEval reveals they're often generating wildly different query structures for the same question, raising serious reliability concerns.
Deferring to a larger LLM only when a smaller LLM is uncertain can match the performance of the larger model alone, while slashing inference costs.
LLMs can achieve more consistent and reliable cross-jurisdictional financial reporting by acting as constrained verifiers within a structured, agentic workflow, rather than as free-form generators.
Detecting AI-generated code is harder than you think: even state-of-the-art detectors fail to reliably identify machine-written code, especially when faced with distribution shifts or adversarial attacks.
Training VLMs on a unified, multilingual, multitask meme dataset reveals that robust meme understanding requires multimodal training and is highly sensitive to dataset-specific overfitting.
A new open-source Hindi LLM, Nanda, outperforms existing models of similar scale by strategically balancing Hindi and English training data.
LLM360 K2 unveils the black box of large language model training, offering a 65B parameter model that beats LLaMA-65B while using fewer resources, all under a fully transparent, open-source framework.