Search papers, labs, and topics across Lattice.
100 papers published across 6 labs.
Engaging with an LLM can significantly diminish belief in conspiracy theories during crises, with effects lasting well beyond the initial interaction.
LLMs may ace proverb completion but falter dramatically when faced with multiple-choice questions, exposing a troubling reliance on memorized patterns over true comprehension.
Intonation and tone in Mandarin are not just separate features; they dynamically interact to shape meaning and speaker intent in profound ways.
NeuPAT recovers nearly all lost language capabilities in multimodal LLMs while ensuring robust performance across diverse tasks.
Even the most advanced LLMs struggle to maintain narrative consistency, with a staggering 68% of generated content conflicting under user interventions.
NeuPAT recovers nearly all lost language capabilities in multimodal LLMs while ensuring robust performance across diverse tasks.
Even the most advanced LLMs struggle to maintain narrative consistency, with a staggering 68% of generated content conflicting under user interventions.
Multi-turn interactions can be effectively optimized in LLMs using tailored RL strategies, overcoming significant challenges in credit assignment and reward design.
Learned routers can outperform fixed-model baselines by 14.6%, revealing a new frontier in efficient LLM deployment.
U-OPSD enables LLMs to self-improve without any external supervision, achieving up to 10.7% performance gains on challenging reasoning benchmarks.
SiPE not only boosts syntactic accuracy but also enhances general language understanding, setting a new standard for integrating syntax into Transformer models.
Dynamic layer routing can boost LLM accuracy by 5% without the need for weight updates or expensive search loops.
Simple transformations don't universally translate across text embedding models, revealing critical compatibility issues that challenge existing assumptions in the field.
Decoupling expert personas in LLMs can drastically improve their accuracy and appropriateness in high-stakes domains like healthcare and finance.
TYTAN achieves 100% coverage and retrieval correctness in constructing analytic schemas, eliminating the manual bottleneck in data analysis.
LLMs exhibit systematic biases in political conflict scenarios, influenced by country identities and user affiliations, challenging assumptions of neutrality in AI-generated content.
ECHO achieves a remarkable 94.9% tool-execution pass rate while maintaining patient data privacy, setting a new standard for locally-deployable health assistants.
Engaging with an LLM can significantly diminish belief in conspiracy theories during crises, with effects lasting well beyond the initial interaction.
Anacreon achieves a groundbreaking ordinal alignment score of 0.775, showcasing its ability to simulate individual responses with unprecedented accuracy.
Verified survey-country metadata boosts LLM predictive accuracy, but random labels can mislead forecasts without any benefit.
Hybridizing autoregressive and speculative decoding, BALANCE boosts task throughput in edge LLM inference while managing latency and memory constraints.
A single lightweight adapter can enhance language model personalization without the need for per-user fine-tuning or costly forward passes, achieving robust performance across diverse user contexts.
Achieving interoperability among defense ontologies is now grounded in a robust framework that includes consensus alignments and validated mappings.
Hierarchical Latent Prediction reduces error accumulation in language models, enabling coherent long-horizon reasoning and more efficient decoding.
SteerWrite achieves state-of-the-art performance in personalized co-writing without the overhead of training, revolutionizing how LLMs can be adapted to specialized domains.
G-STEER refines user research queries with unprecedented efficiency, asking one-third as many questions while maximizing personalization and target coverage.
A new benchmark reveals that even the best LLMs lag significantly behind human experts in reviewing national standards, but structured coordination can bridge this gap.
Concentrating distillation on reasoning pivots boosts multilingual reasoning performance, outperforming traditional methods across 17 languages.
LLMs can achieve 7.32% better goal completion in social negotiations by strategically optimizing reward signals based on dialogue context.
State-of-the-art multimodal models falter in interpreting implicit social cues, revealing a critical gap in AI's understanding of human communication.
Context biasing methods can slash biased word error rates by up to 88%, outperforming speech LLMs in challenging ASR scenarios.
LLMs may generate plausible narratives, but they fall short in capturing the rich diversity and stylistic irregularities found in human writing, revealing a critical gap in their cultural reach.
Character-noised continued pre-training boosts zero-shot dialect robustness while preserving standard performance, revealing distinct adaptation mechanisms in language models.
ProVerif outpaces Tamarin in verification speed by over 92%, while ensuring soundness in the majority of comparative tasks.
AssertMate outperforms existing LLM-based assertion generation methods, achieving higher accuracy and bug detection through a novel multi-agent approach.
A handful of critical factors can dramatically elevate the success rate of Agile projects, challenging the notion that more complexity leads to better outcomes.
Current LLMs can generate precise LTL specifications from unstructured requirements, achieving significant performance without fine-tuning.
ALTER transforms how we generate CT reports by accurately modeling longitudinal changes across multiple anatomical regions, achieving state-of-the-art results in the process.
VLMs can transform under-resourced historical languages by automating data extraction at unprecedented scales, as demonstrated by the mapping of Armenian commercial advertisements in Paris.
Ontology-based frameworks can revolutionize how we personalize learning in higher education, making student profiles central to educational success.
EXCISE corrects exclusion inversion in retrieval systems, boosting exclusion success rates from 5.8% to an impressive 69.1%.
Novice users benefit significantly from explanations in product recommendations, while expert users remain unaffected by additional information complexity.
Faked character detection in handwritten Chinese text recognition can be achieved without sacrificing recognition performance, thanks to DTRNet's innovative dual decoding approach.
Software engineers are increasingly dependent on LLMs, risking overreliance that could undermine traditional practices like peer consultation and documentation.
Task-Conditional Flow Matching redefines multilingual embedding adaptation by tailoring optimization strategies to task-specific needs, achieving unprecedented improvements in embedding quality.
FormBharo shows that rule-based controls can enable smaller, cost-effective models to excel in form completion, even when faced with challenging real-world speech inputs.
Parser-derived supervision can replace costly human annotations, achieving up to 93.8% parse success in Danish and 80% preference in native speaker comparisons.
MetaboLLM-GIN not only outperforms conventional models in predicting stress hyperglycemia but also transforms biochemical knowledge into interpretable metabolite graphs.
Some LLMs can outperform humans in legal argumentation, but all struggle with the complex planning required for notary exams.
PPMI graph averaging can boost Random Indexing accuracy by over 50%, but struggles against neural embeddings in broader contexts.
A new neural network structure can effectively model hierarchical data while addressing the biases introduced by missing data, but faces challenges with stability.
GROM achieves rapid, effective unlearning in seconds while maintaining model performance, outpacing traditional methods that struggle with computational efficiency and content recovery.
Achieving over 90% accuracy in extracting structured information from unstructured documents, this framework outpaces human experts by a factor of 30 in speed.
PSRS affects up to 56% of responses in LLMs, revealing a critical vulnerability in AI alignment that can lead to harmful outcomes.
Some large language models exhibit surprising human-like sensitivity to discourse factors in anaphor resolution, but falter on semantic interference.
ECG-LENS outperforms existing ECG report generation systems by integrating multi-lead signal modeling with advanced clinical context, achieving unprecedented report quality.
The rise of unspaced em-dashes in congressional press releases signals a significant stylistic shift potentially driven by the adoption of large language models in governmental communication.
Synthetic clinical communication can effectively bootstrap NLP systems, outperforming traditional zero-shot approaches and paving the way for more robust healthcare applications.
EchoPrompt reveals that by restoring latent prompts, we can significantly enhance the detection of LLM-generated text, achieving state-of-the-art results without any training.
ASR systems are not just failing technically; they perpetuate colonial hierarchies that silence marginalized voices, necessitating a radical rethinking of how we design these technologies.
HallDetect flags hallucinations by identifying just one confidently contradicted claim, revolutionizing how we ensure factual accuracy in LLM outputs.
Autogrammar can automatically learn context-free grammars that boost language model performance, achieving near-perfect precision and tripling execution speed on DSL tasks.
Current MLLMs struggle with creative decoding, achieving only 50.7% accuracy in understanding cross-concept relations, revealing a critical gap in their cognitive capabilities.
OPD$^2$ not only boosts multilingual math reasoning but also narrows the performance gap between English and Korean models, revealing the hidden potential of language-specific training signals.
MameLoshnLM not only sets a new standard for Yiddish NLP but also reveals the inadequacies of multilingual models in handling linguistically rich yet underrepresented languages.
Factorized Hypothesis Search reveals that maintaining multiple interpretations of evidence can dramatically improve taxonomy retrieval accuracy, outperforming conventional methods.
OneEmo outperforms larger models in emotional intelligence tasks while using significantly fewer parameters, showcasing the power of unified multimodal reasoning.
PC-Agents mimic human personality dynamics but fall short in capturing the full complexity of personality evolution after life events.
Language models can exhibit human-like preferences for modifier order even when trained on data that lacks direct evidence for such structures.
Certified deferral reveals that even well-calibrated small language models struggle to meet safety thresholds in risk-sensitive applications.
Every LLM evaluated fabricates user attributes, with a staggering 41.6% of claims showing over-inference, challenging the reliability of self-reported model confidence.
Korean models struggle significantly with writing-system-intensive tasks, revealing a 68.7 pp accuracy gap in the Korean Cipher compared to English.
Achieving 90% sparsity, BnBERT-iPET rivals larger models while drastically reducing computational costs for Bengali NLP tasks.
Attention-free models can outperform transformers in language generation, especially at smaller dataset scales, challenging the dominance of attention mechanisms in NLP.
ODRA's innovative approach to modeling patient resistance leads to synthetic therapy sessions that are not only more realistic but also preferred by licensed psychologists.
MLLMs can now refuse to localize non-existent objects without sacrificing their accuracy on valid requests, thanks to a novel reinforcement learning strategy.
Real-time collaboration between AI and human annotators can transform the efficiency of artwork annotation, reducing cognitive load and enhancing knowledge accumulation.
Guideline-driven training can elevate a medical triage agent's performance without expert annotations, achieving a remarkable 74.1% agreement with operational benchmarks.
EndoVLM achieves superior performance in endoscopic image analysis by effectively aligning clinical reports with visual data, outperforming traditional models and showcasing impressive zero-shot capabilities.
Verification-first coordination in language model ensembles can boost accuracy by over 6% while ensuring diverse responses are retained only when warranted.
LLMs may ace proverb completion but falter dramatically when faced with multiple-choice questions, exposing a troubling reliance on memorized patterns over true comprehension.
LoRA+ outshines other PEFT methods, achieving the best energy efficiency in 19 out of 24 configurations, paving the way for practical on-device personalization of language models.
MIDAS achieves up to an 18.2% improvement in summarization quality without the need for manual prompt engineering, revolutionizing how we handle domain-specific summarization tasks.
The rise of the far-right in Germany has led to a dramatic shift towards intuition-based rhetoric among political elites, threatening the integrity of democratic discourse.
Injecting targeted language features at inference time can boost multilingual model performance by over 10 percentage points without retraining.
The study reveals that historical transcription practices have led to misinterpretations of grammatical structures in Mapudungun, reshaping our understanding of its syntax.
Simple, off-the-shelf NLP pipelines can achieve competitive accuracy in tagging morphologically complex, low-resource languages like Scottish Gaelic.
EdgeLM reveals that selecting edge evidence can significantly improve LLM performance in table understanding tasks, outperforming traditional similarity-based retrieval methods.
Expert evaluations reveal that while fluency is achieved, institutional completeness and report identity remain significant hurdles for financial report generation in LLMs.
Cash assistance interventions show a remarkable Level-of-Evidence score of 0.865, indicating strong convergence in humanitarian outcomes.
Multi-agent frameworks can significantly enhance the translation of puns, outperforming traditional discriminator-guided methods in preserving humor and creativity.
LLMs struggle with the complexities of Islamic scholarship, and the newly introduced ISTB reveals significant gaps in their performance across different levels of scholarly demand.
STRIVE automates the generation of event plausibility sets, achieving a remarkable 75% quality rate, but still struggles with boundary cases that require human judgment.
System prompts can be optimized to ensure equitable response quality, reducing worst-case information loss by over 13% while maintaining average performance.
Recovering nearly half of the original text records as clean units reveals that context dependence can effectively signal boundaries in language models, outperforming conventional methods.
Successful pun translation hinges on discovering new sound-meaning collisions rather than merely translating words, revealing a critical bottleneck in the retrieval process.
CS educators are not just talking tech; they're deeply engaged in discussions about mathematics and humanities, revealing a rich tapestry of interests that could reshape educational support strategies.
A new structured representation for manipulation tasks reveals critical labeling anomalies that traditional methods miss, enhancing both readability and verification.
The rise of Generative AI has fundamentally altered how users seek information, revealing surprising shifts in interface preferences and cognitive engagement.
LLMs can both combat misinformation and generate it, revealing a paradox that demands urgent research attention.
DBLAST significantly enhances the accepted draft length in stochastic decoding scenarios, particularly when the target distribution's entropy is high.
Example-guided prompting boosts document-level simplification quality, outperforming traditional methods and revealing the nuanced interplay between retrieval and model capabilities.
Recoverable eviction can drastically reduce missed attention and improve information retention in long-context decoding, outperforming traditional methods.
SFC redefines semantic understanding in spoken language tasks, achieving superior accuracy and adaptability in open-domain contexts.
Intonation and tone in Mandarin are not just separate features; they dynamically interact to shape meaning and speaker intent in profound ways.