Search papers, labs, and topics across Lattice.
100 papers published across 9 labs.
Systems achieved up to 97.5% accuracy in multilingual financial question answering, revealing the potential for high-performance AI across diverse languages.
The top-performing systems in multilingual financial question answering are separated by less than one percentage point, showcasing the intense competition and subtlety in model performance.
Public benchmarks from TalentCLEF could revolutionize how NLP systems are developed for fair and effective talent management across industries.
AI-generated legal answers can sometimes surpass traditional reference responses, revealing untapped potential in legal tech.
Early detection of reasoning non-convergence in models could revolutionize adaptive compute allocation strategies, improving efficiency without sacrificing accuracy.
Early detection of reasoning non-convergence in models could revolutionize adaptive compute allocation strategies, improving efficiency without sacrificing accuracy.
LLMs are overzealous tutors, intervening too soon and too often, which may undermine true learning and cognitive engagement.
A unified framework that enables reproducible blood glucose forecasting while maintaining data privacy could revolutionize diabetes research.
Unlearning in offline RL is more complex than previously thought, with common deletion methods showing environment-dependent privacy-utility behavior that can mislead evaluations.
GRPO's hierarchical penalty reduces vulnerability prediction errors by nearly 28%, outperforming larger models under challenging conditions.
Position bias in LLM evaluations is only detectable within a narrow accuracy range, challenging the interpretation of benchmark results for high-performing models.
BoE transforms candidate selection by leveraging partial verification, significantly enhancing outcomes in vision-language tasks where complete evaluations are unattainable.
RUMBA reveals critical insights into how memory mechanisms in LLMs perform across long-term contexts, exposing significant gaps in current models' capabilities.
Multilingual embeddings outperform Greek-specific models in book retrieval, revealing surprising strengths in hybrid methods for complex queries.
Coordinated feedback across RAG components leads to a remarkable 12-15 percentage point improvement in performance, challenging the notion that isolated optimization suffices.
Models that seem equally accurate can drastically differ in their ability to clarify ambiguous requests, impacting user experience and efficiency.
Closed-source models may seem superior, but they only edge out open-weight models by a narrow margin, with major discrepancies in retrieval and synthesis capabilities.
LLMs may ace academic trivia but falter on everyday cultural knowledge, revealing a critical gap in their understanding.
LLMs falter in culturally nuanced dialogue, revealing critical gaps in their understanding of Indonesian cultural commonsense.
Citation-aware evidence governance is essential for generating trustworthy legal research, as retrieval tools alone fail to ensure citation reliability.
RFFI performance hinges more on where you tap in the receiver chain than on the sophistication of your model, with timing recovery proving crucial for accuracy.
IRIS can detect model substitutions and routing dilutions in LLM gateways with unprecedented accuracy using only the output text, challenging the reliability of commercial AI services.
A new benchmark reveals that existing defocus deblurring methods struggle with generalization across datasets, highlighting a critical gap in current evaluation practices.
Emotion-oriented MLLMs can now achieve unprecedented accuracy in visual emotional intelligence, revealing critical insights into their limitations and capabilities.
Webly supervised multi-label recognition can achieve state-of-the-art performance with a novel dual-branch learning approach that tackles label noise head-on.
Despite advances in MLLMs, they still struggle with dynamic reasoning, falling far short of human capabilities in interpreting continuous visual cues.
Transforming the asymmetric inner product problem into a Euclidean search space could redefine efficiency benchmarks in high-dimensional embedding searches.
Coding agents can now be evaluated on their ability to navigate fuzzy requirements and interactive workflows, reflecting real-world software development challenges.
Fine-tuned robotic policies can be biased towards certain instruction factors, but a new bias-aware data strategy can significantly enhance their performance with fewer demonstrations.
FutureSurf reveals that existing models can miss up to 4.1 times the expected accuracy in predicting future surfaces, challenging the assumptions of current dynamic scene reconstruction methods.
Knowledge retrieval and reasoning are the main bottlenecks in KI-VQA, but a new diagnostic benchmark reveals deeper issues in visual grounding and object identification.
Current pathology VLMs can achieve high accuracy without visual input, revealing critical flaws in how we assess their multimodal capabilities.
This new evaluation framework can reliably distinguish between meaningful audio caption variations and genuine corruptions, transforming how we assess automated audio captioning systems.
Internal evidence can mislead auditors, making reports sound plausible while lacking actual relevance, exposing a critical governance risk in AI oversight.
Real-world coding tasks, reverse-engineered from actual commits and scenarios, make Tencent WorkBuddy Bench a game-changer in contamination-resistant evaluation for coding agents.
Red-team evaluations can certify safety for common risks but fail to provide evidence for rare catastrophic failures, highlighting a critical gap in current AI safety assessments.
Majority voting in LLM ensembles might be more about capability than diversity, with only 9.98% of subsets showing performance gains over the best model.
Evaluator score gaps are the secret sauce for optimizing LLM policies, and DynamicRubric turns this insight into a powerful co-evolution framework that outshines existing methods.
Attack ensembles optimized for minimum-norm strategies can provide a more accurate and flexible evaluation of adversarial robustness than traditional fixed-budget methods.
Expert-guided forecast editing can dramatically enhance time-series predictions by intelligently balancing exploitation and exploration, leading to superior outcomes.
LLMs often produce fluent but flawed reasoning, and our new framework reveals the hidden weaknesses in their outputs that traditional evaluation methods miss.
SelectBench reveals that while selective evidence adoption in LLMs can be improved, the gains are modest and highlight significant challenges that remain in ensuring robust performance.
Intensity normalization can enhance MRI segmentation performance, but its benefits are minimal compared to the challenges posed by domain shifts.
Systems achieved up to 97.5% accuracy in multilingual financial question answering, revealing the potential for high-performance AI across diverse languages.
The top-performing systems in multilingual financial question answering are separated by less than one percentage point, showcasing the intense competition and subtlety in model performance.
Surface accuracy metrics can mislead researchers about the true reliability of multimodal search systems, with silent failures lurking beneath the surface.
The Sagittal T2-weighted MRI sequence outperforms multisequence approaches, achieving a Macro F1-score of 50.31% in spine pathology diagnosis.
New benchmarks reveal that even advanced LLMs struggle with cultural value alignment, but targeted fine-tuning using Sri Lankan values can significantly enhance their performance.
No single Arabic LLM can master hallucination detection, localization, and explanation, revealing critical gaps in current models' capabilities.
LLMs misidentify adjacent values over 50% of the time, revealing critical biases that could skew their understanding of human motivations.
LLMs are assessed on their alignment with human values through a groundbreaking benchmark that captures real-world dilemmas, revealing critical insights into their decision-making processes.
A groundbreaking benchmark that captures the complexities of real-world dynamic tracking with 795K RGBT frame pairs and extensive annotations for robust evaluation.
Current MLLMs excel at visual reproduction but falter in generating the necessary data semantics and interaction logic for coordinated multi-view interfaces.
SafeGen boosts VLMAD performance by over 24% in safety-critical scenario generation, bridging the sim-to-real gap that has long plagued autonomous driving systems.
Current models falter in reasoning-guided remote sensing image editing, with the best achieving only 24.28% accuracy, exposing a critical gap in AI capabilities.
Strong performance in static evaluations masks a critical flaw: LLMs struggle to adapt to evolving user intent during multi-turn interactions.
A lightweight vision-language model can effectively identify UI principle violations, achieving an impressive F1 score of 84% across multiple quality dimensions.
The Maskability Index reveals that the right prompting strategy can significantly boost the performance of pretrained language models in knowledge extraction tasks.
State-of-the-art LLMs fail to capture nuanced user preferences, lagging behind simple baselines in predicting choices in interactive narratives.
Agents struggle to match expert financial judgments, with the best only achieving 52.4% accuracy, exposing critical limitations in AI's ability to process real-time market information.
Sound probabilistic safety bounds reveal that even state-of-the-art LLMs can be rigorously evaluated for harmful output risks, transforming our approach to model safety.
Language models exhibit a significant shift in self-reporting behavior post-training, revealing that their "inner life" is not a fixed trait but a product of their training regime.
Credentialing in education is at risk as universities struggle to define what AI delegation means for learning evidence and assessment integrity.
No current LLM-based agent can reliably avoid executing unsafe actions when using third-party skills, with a staggering 17% failure rate even in optimal conditions.
Public benchmarks from TalentCLEF could revolutionize how NLP systems are developed for fair and effective talent management across industries.
LLMs may not rival formal verification tools in security analysis, but they could serve as useful pre-screening filters for identifying vulnerabilities.
Fact verification methods can degrade significantly when faced with controlled evidence poisoning, revealing vulnerabilities that traditional benchmarks overlook.
NCIP reveals that leveraging prediction variability across multiple checkpoints can dramatically enhance early fault discovery in DNNs, outperforming traditional confidence-based methods.
The new homogeneity-parsimony scores not only unify existing evaluation criteria but also reveal Pareto-optimal clustering solutions that traditional metrics miss.
Definition blindness in OWVAD evaluation allows models to score highly while failing to respond to user-defined anomalies, but new metrics and scoring rules can fix this oversight.
Current navigation agents falter in cross-context scenarios, with a significant performance drop when moving from outdoor to indoor-to-outdoor environments.
AI-generated covers often hide severe harmonic errors behind acceptable key consistency, challenging the reliability of global quality scores.
UniRank reveals the potential for reproducible benchmarking in ranking models, leveling the playing field between academic research and industrial applications.
LLMs struggle with complex temporal reasoning, showing significant accuracy drops on intricate waveform queries, underscoring a critical gap in their capabilities.
GPT-4.1 can predict election outcomes and healthcare opinions with striking accuracy, revealing the untapped potential of persona simulation in AI.
Current reranking methods fall short, but Rubric4Setwise transforms evaluation into actionable selection signals, achieving unprecedented performance across diverse document sets.
Even the most advanced autonomous agents struggle with fundamental document manipulation tasks, revealing critical vulnerabilities in their operational capabilities.
LLMs struggle with financial reasoning under real-world conditions, revealing critical flaws in their ability to handle complex, long-horizon tasks.
Existing staypoint detection algorithms falter in noisy environments, but new unsupervised methods show promise for substantial improvements.
Textual data from financial reports can significantly enhance fraud detection, as evidenced by our framework's superior performance on a novel benchmark task.
Circuit extraction claims can be misleading, as the perceived mechanisms behind model behavior are heavily influenced by how circuits are reported and compared.
NSMA achieves unprecedented performance in adaptive bitrate streaming by seamlessly integrating neural learning with symbolic reasoning, revealing hidden dynamics that traditional metrics overlook.
Models that excel in prediction often falter in mechanistic reasoning, revealing hidden flaws in their logic and understanding of cellular context.
AI agents struggle to surpass 50% accuracy in genomic surveillance tasks, revealing significant gaps in their analytical capabilities.
Despite apparent progress in SVG repairs, top models struggle to meet rigorous editing specifications, achieving only 15% success.
OpenRTAG reveals that traditional GNNs struggle significantly more than LLM-GNNs under realistic data quality degradation, highlighting critical vulnerabilities in graph learning models.
Current LLMs achieve only a 12.30% success rate in generating executable scientific code, highlighting a significant gap in their capabilities.
MIRA-Ev reveals that traditional MCQA methods fail to capture the complexities of clinical reasoning, highlighting the need for a more nuanced evaluation framework.
Changing the diagnostic reader can dramatically shift performance metrics, revealing that history quality is often obscured in traditional evaluations.
Evaluator bias can dramatically alter the perceived safety of medical AI, with LLM judges showing a leniency that could misrepresent model performance. WHY_IT MATTERS: This insight challenges the reliability of current evaluation methods for medical AI, emphasizing the need for standardized assessment frameworks to ensure safety in clinical applications.
Vision-language models consistently fail to recognize missing object parts, even when external evidence contradicts their visual perceptions, highlighting a fundamental flaw in their design.
Audio-language models struggle significantly with hidden evaluation tasks, with accuracy plummeting by nearly 12 percentage points on average, highlighting the challenge of true audio comprehension.
DeepDebug achieves a 32% improvement in task recovery accuracy, showcasing a powerful new approach to debugging LLM agent failures.
Open-ended generation models are failing to capture factual completeness, with top performers only achieving 58.7% accuracy on a new benchmark designed to assess this critical aspect.
Small language models may excel at tasks but often disregard conflicting instructions, revealing a critical gap in their reliability.
Instruction adherence collapses at 80 simultaneous commands, revealing a critical threshold for prompt design that every practitioner should heed.
Structural generalization is mathematically unattainable for pure Transformers, revealing a critical limitation in their learning capabilities compared to neuro-symbolic systems.
Models exhibit a date-trigger reflex that is more indicative of training generation than size, challenging assumptions about LLM behavior consistency across versions.
MLLMs struggle to discern genuine consensus from pseudo-consensus in multi-party meetings, revealing critical gaps in their Theory of Mind reasoning abilities.
Domain-specific fine-tuning boosts RF reasoning in LLMs, especially for smaller models, while semantic retrieval proves superior for context alignment.
Google Telephony not only rivals but sometimes outperforms human listeners in recognizing diverse speech, challenging traditional views on ASR limitations.
Predicting user reactions rather than evaluating responses directly allows for verifiable self-evolution in dialogue agents, achieving over 75% accuracy in a challenging sales context.
AI-generated legal answers can sometimes surpass traditional reference responses, revealing untapped potential in legal tech.
GPT-generated surveys can effectively capture social attitudes, matching human designs in key areas while revealing unique insights into belief structures.
Hidden sabotage in training data eludes detection more than 50% of the time, revealing critical vulnerabilities in automated AI R&D.