Search papers, labs, and topics across Lattice.
Jiuge-Tuiqiao transforms AI from a passive generator into an active collaborator, enhancing user creativity in classical Chinese poetry.
FormuEvo can accelerate solver performance by up to 5.5× by intelligently evolving MIP formulations through LLM-guided optimization.
Skill selection in LLMs can be optimized to achieve a 0.73 task success rate while using 28% fewer tokens than existing methods.
DeepWeaver transforms the way LLMs synthesize evidence, leading to answers that are not only more comprehensive but also better grounded in citations.
The empirical analysis reveals that larger models may excel in generating narratives but often fail to maintain coherence and depth, exposing a critical trade-off in LLM storytelling capabilities.
Groundedness Drift reveals that explanations can mislead in the presence of backdoor attacks, highlighting vulnerabilities in language model classifiers.
Token-level credit assignment can drastically improve the effectiveness of generative document retrieval, leading to better alignment between generation and relevance.
Cross-path reasoning reveals untapped research ideas, outperforming traditional methods and rivaling human insights in scholarly evolution.
Auditing Chinese web content reveals pervasive pollution that shifts over time, challenging the integrity of LLM training data.
Even the top-performing LLM struggles with complex legal temporal reasoning, revealing significant gaps in AI's understanding of time-sensitive legal contexts.
SteerWrite achieves state-of-the-art performance in personalized co-writing without the overhead of training, revolutionizing how LLMs can be adapted to specialized domains.
AFD-Ledger reveals that optimizing deployment for AFD can drastically cut evaluation costs while exposing the nuanced performance dynamics between homogeneous and heterogeneous setups.
Hard prompt compressors can leave critical context gaps, leading to a staggering 60% of examples suffering from referential dangling, which severely impacts accuracy in multi-hop question answering.
Accepting selective mismatches in autoregressive decoding can boost throughput by over 15% without any additional training or model adjustments.
Language models exhibit a surprising bias towards cities with expansive infrastructure and rapid growth, revealing their implicit urban assumptions.
DataSpace reveals that even the best multimodal models struggle with data agent accuracy, achieving only 66.34% in complex heterogeneous environments.
SkillTrace redefines skill composition for LLM agents, achieving unprecedented success rates by leveraging a structured query-skill graph.
Latent Softmax achieves up to 17.5% lower phoneme error rates in multilingual ASR by intelligently modeling tonal distinctions without sacrificing cross-lingual sharing.
Vibe-FDTR achieves near-perfect accuracy in thermal property analysis while slashing computational costs and execution time, revolutionizing how researchers approach FDTR data.
Agents can now autonomously teach themselves creative skills using high-quality human texts, bypassing the need for expensive human feedback.
CoMem achieves a 7.83x prefill speedup and drastically reduces memory usage while maintaining high performance on long-context tasks, challenging conventional memory management in LLMs.
CURL effectively harnesses LLMs to stabilize CATE estimation, leading to improved performance in personalized interventions across multiple benchmarks.
UrbanDS outperforms traditional data science agents by effectively navigating complex urban datasets through a novel graph-guided multi-agent architecture.
Evolving solver and rubric skills in tandem reveals hidden weaknesses and boosts performance by up to 5% without relying on fixed evaluation criteria.
PILA boosts ad effectiveness in LLM-native advertising without sacrificing response quality, offering a game-changing solution for monetization.
A unified taxonomy reveals how diverse memory mechanisms in LLMs can be systematically understood and leveraged for future innovations.
MemSFT enables LLMs to gain specialized domain knowledge without sacrificing their general performance, effectively sidestepping the alignment tax.
UNIFUSION achieves unprecedented performance in generative tasks by seamlessly adapting autoregressive models to uniform-noise diffusion, outperforming all evaluated models on key metrics.
Semref leverages LLMs to automatically refine architecture recovery results, achieving up to 118.57% improvement in accuracy over traditional methods.
REFACT reduces token consumption while enhancing the density and faithfulness of reasoning traces in large language models, ensuring that every cited fact meaningfully supports the answer.
Adding depthwise convolutions to Transformers can boost accuracy on downstream tasks while barely increasing model size.
Current LLMs struggle with multi-granularity event analysis, revealing critical performance gaps that could hinder their application in complex narrative tasks.
Long-form article generation can achieve a significant quality boost through a modular improvement loop that adapts based on structured evaluations.
Temporal updates in LLMs can be made without sacrificing historical accuracy, achieving over 23% improvement in consistency with a single optimized representation.
Generative retrieval can revolutionize statute retrieval by effectively bridging the gap between everyday legal language and formal statutes.
UGP enables multilingual ASR models to achieve near-zero forgetting while adapting to low-resource languages, revolutionizing how we approach continual learning in diverse linguistic contexts.
Training language models with just 100 million tokens could redefine efficiency benchmarks in Chinese NLP.
HCC-STAR not only surpasses leading models in treatment accuracy but also offers a significant survival advantage, highlighting the potential of AI in precision oncology.
Over 40% improvement in analytical efficiency could revolutionize how researchers conduct trajectory inference in single-cell transcriptomics.
NPUs can waste up to 40% of energy due to suboptimal configurations, but a new profiling tool reveals how to cut this waste significantly.
ACE achieves a remarkable 70% success rate in constraint retrieval tasks without any task-specific retraining, showcasing the power of zero-shot workflow reasoning in robotic manipulation.
Online imitation learning can outperform offline methods, but only when the student can effectively represent the expert—realizability is key.
Token-level learning dynamics, not just model size, dictate scaling laws in language models, revealing actionable insights for training optimization.
ICMPG achieves a groundbreaking balance between semantic fidelity and physical realism in motion synthesis, outperforming traditional methods in both standard and zero-shot scenarios.
Sentence-level contextual entrainment can skew inference probabilities, but selectively disabling just a few attention heads can mitigate this effect without sacrificing performance.
TokenMinds reveals that combining discrete SID-based user tokens with dense embeddings can significantly enhance user modeling in recommender systems at scale.
Query language can dramatically shift historical credit in LLMs, revealing a hidden bias that favors dominant narratives over marginalized voices.
MOCAP slashes LLM inference latency by over 76% while boosting throughput more than threefold, redefining efficiency for long-context processing on wafer-scale chips.
Advanced RS MLLMs struggle with negation, but a novel learning method can dramatically enhance their understanding using minimal unlabeled data.
Adaptive weighting in model merging can drastically improve multilingual reasoning performance, outperforming traditional methods across 21 languages.