Search papers, labs, and topics across Lattice.
100 papers published across 5 labs.
Models can autonomously bootstrap their own dense token-level supervision simply by extrapolating the trajectory of their own RL updates away from a trailing checkpoint.
LLMs can match or outperform traditional CKD screening methods with minimal examples, but their stability diminishes with increased input complexity.
Personalizing item selection during denoising boosts recommendation accuracy by leveraging user interactions more effectively than traditional methods.
ALRA achieves a notable accuracy improvement over existing distillation methods by intelligently combining student and teacher token selections, revealing the importance of local relational alignment in model training.
Increased generative-process diversity in language models significantly reduces correlated failures, a factor overlooked by traditional semantic similarity assessments.
Models can autonomously bootstrap their own dense token-level supervision simply by extrapolating the trajectory of their own RL updates away from a trailing checkpoint.
LLMs can match or outperform traditional CKD screening methods with minimal examples, but their stability diminishes with increased input complexity.
Personalizing item selection during denoising boosts recommendation accuracy by leveraging user interactions more effectively than traditional methods.
ALRA achieves a notable accuracy improvement over existing distillation methods by intelligently combining student and teacher token selections, revealing the importance of local relational alignment in model training.
Increased generative-process diversity in language models significantly reduces correlated failures, a factor overlooked by traditional semantic similarity assessments.
ESPO not only boosts accuracy by nearly 4 percentage points but also slashes prompt length by nearly half, redefining efficiency in prompt optimization.
Auxiliary views can enhance LLM learning efficiency, revealing that reallocation of training tokens can lead to better factual recall even when using weaker teacher models.
Translation divergence among agents reveals significant decision flexibility in multilingual models, challenging the notion of a single optimal output.
CAPA enables LLMs to speak up in meetings, slashing silence rates from over half to just 2.5% while maintaining high accuracy in contributions.
Preference optimization boosts safe responses in banking agents from 52% to 80%, while reinforcement learning enhances edge-case performance significantly with fewer tokens generated.
Dialogue state tracking accuracy in industrial robots skyrockets with the new IRWOZ 2.0 dataset, achieving a BLEU-4 score increase of over 200%.
Instruction duplication boosts diagnostic accuracy in language models by nearly 3 percentage points, significantly reducing failure rates without the need for retraining.
LLMs can achieve over 80% accuracy in extracting architectural insights from code commits, but often miss the critical rationale behind design choices.
Existing hallucination detection methods falter in the nuanced landscape of scientific peer reviews, revealing a critical gap in ensuring review reliability.
Visual summaries from KnowVis not only clarify complex concepts but also boost student retention and understanding, outperforming traditional video summarization techniques.
Minor changes in prompt phrasing can drastically alter LLM outputs, overshadowing the benefits of fine-tuning in drug toxicity prediction.
Single-pass annotation can miss critical factual errors in chatbot responses, but a multi-perspective approach reveals a more comprehensive picture of accuracy.
LLMs not only understand plural references but also mirror human preferences in pronoun usage based on entity similarity and conjunctions.
Multilingual LLMs suffer substantial reasoning failures when faced with structurally altered inputs, revealing a critical vulnerability in their design.
Selective retrieval can enhance mental health QA by preserving response quality and safety while still improving specificity.
No single large language model dominates idiom comprehension in Bangla, revealing nuanced strengths and weaknesses across different tasks.
CHARM not only outperforms existing moral detection systems but also reveals that moral framing significantly influences online endorsement behavior during the COVID-19 pandemic.
LLMs not only introduce new vulnerabilities but also amplify existing web threats, necessitating a paradigm shift in how we secure online interactions.
Existing reflection models fail to authenticate student engagement in the GenAI era, but the 5P model offers a structured solution to enhance learning and integrity.
OCR-EDR transforms OCR error analysis into actionable repairs, achieving a remarkable 94.78% diagnostic accuracy and a 30.99-point boost in formula performance.
Hyperbolic geometry can revolutionize recommendation systems by dramatically improving the accuracy of long-tail item suggestions.
Machine translation benchmarks have functionally saturated, but pairing human-authored failure cases with deterministic verification rules reveals critical multimodal blind spots that automated metrics consistently miss.
Reranked retrieval methods significantly outperform traditional approaches in matching graduate applicants with relevant academic advisors, revealing critical insights into effective faculty profiling.
Traditional single-session assessments can mislead clinicians by failing to accurately track worsening depression trends, underscoring the need for longitudinal approaches.
Feminine loanwords in Latvian are significantly more tied to fixed suffixes, revealing a striking asymmetry in gender assignment that challenges traditional views on loanword integration.
Current PII detection systems fail dramatically under real-world conditions, with each architecture revealing unique vulnerabilities that standard benchmarks overlook.
TraveL captures the nuances of traveler behavior and regional correlations, leading to a 14.7% improvement in travel time distribution estimation over existing methods.
The first open Armenian LLM, arm-gemma-e4b, not only outperforms all predecessors but also highlights the critical balance between fluency and knowledge retention in low-resource language models.
A standardized protocol for AI agents could revolutionize interoperability, enabling seamless communication across diverse systems and frameworks.
LLMs often prioritize safety, leading to conservative predictions about recipe suitability for diabetes, but some models can significantly outperform others when reasoning with dietary guidelines.
Achieving up to 83% success in adapting retail supply chain operations shows that intelligent decision-making can significantly enhance operational efficiency.
Harnessing the power of document structure, STAIR achieves a remarkable 82.6% Recall@1, setting a new standard for low-hallucination information retrieval in LLMs.
A novel benchmark dataset reveals that a multi-stage RAG pipeline can dramatically boost financial question answering accuracy by over 34 percentage points.
LLMs can accurately predict typological features while providing interpretable rationales, even for low-resource languages, when given the right contextual evidence.
Placement of synthetic data in representation space can be more crucial than sheer volume, leading to significant performance gains in low-resource NLP tasks.
Language models trained on synthetic data may not reflect the true nature of the languages they represent, challenging our assumptions about their utility in linguistic studies.
Pattern over-generalization in knowledge graph embeddings can be effectively mitigated with PogRE, leading to superior link prediction performance.
Localization-sensitive questions can see up to 40% performance improvement by dynamically selecting reasoning strategies based on question type.
Once a critical threshold of LLM adoption is crossed, even minor increases can trigger rapid cognitive decline across populations, underscoring the need for strategies to maintain cognitive autonomy.
Over 80% of emergency messages are in English, but BEACON ensures that non-English speakers receive critical evacuation guidance in their native language, potentially saving lives.
The strongest LLM only achieves a macro-F1 score of 67.3 on a benchmark designed to capture the complexity of opposing emotions, revealing significant gaps in current emotion recognition systems.
Real Human-Human dialogue data can transform turn-taking in dialogue systems, achieving better proficiency without sacrificing semantic quality.
By reallocating KV cache resources based on attention head specialization, SGD-KV cuts memory usage by 75% while handling contexts up to 1M tokens with state-of-the-art accuracy.
Large language models can surpass human performance in understanding frame semantics, revealing their advanced implicit comprehension abilities.
Robustness claims in language models can be misleading when based solely on output behavior, as perturbations reveal complex, multi-level effects that traditional metrics overlook.
Contextual grammar errors in Tamil can be corrected with 52.5% accuracy after targeted training, a significant leap from 1.0%.
A shared emotional landscape on Arabic YouTube reveals that sentiment is predominantly negative, shaped by historical trauma and political contexts rather than national identities.
Off-the-shelf LLMs miss nearly 53% of critical requirement issues, raising serious questions about their reliability in engineering contexts.
Grading multi-topic essays without explicit labels can achieve unprecedented accuracy through graph-augmented retrieval techniques.
LLM4AIGQ transforms user preference extraction in e-commerce by generating guidance queries that accurately reflect multi-interests, overcoming the pitfalls of traditional methods.
Shifting the focus from topical relevance to answerability, CLEAR reveals that many conversational retrievers miss the mark by overlooking critical answer-supporting passages.
Achieving 89.84% labeling accuracy without any pre-labeled training data could revolutionize how software issues are categorized and managed.
Auxiliary draft models are no longer necessary for speculative decoding: distilling lightweight diffusion heads directly into standard LLMs yields lossless 3× generation speedups even at peak batch sizes.
Recurring frontier API calls can be replaced by one-minute compile-time distillation, turning plain-text prompts into local, versionable neural functions that achieve 83.6% accuracy on benchmarks where instantaneous compilers fail entirely.
Propagating local edits across multi-turn artifacts does not require expensive sequential reflection loops—reranking just three parallel samples reliably yields up to a 9.7% consistency boost across open and frontier models.
Incorporating spectral structural priors into uncertain knowledge graph completion can dramatically enhance prediction accuracy and training stability without adding extra parameters.
A foundational ontology reveals the hidden contradictions in human-robot dialogue, setting the stage for smarter and more reliable interactions.
LLMs systematically prioritize atmospheric elements over character-driven action in storytelling, revealing a fundamental divergence from human authorship.
Language models can now self-select relevant context, slashing attention costs by over 50% while maintaining performance.
Pseudo-triplet construction enables TTS systems to generate nuanced voice modifications that respond to performance directions while maintaining speaker identity.
Security mechanisms based on LLM linguistic outputs are fundamentally unreliable due to inherent "linguistic illegibility," necessitating new isolation techniques for model safety.
Loom achieves a remarkable 26x speedup in Root Cause Analysis while maintaining competitive accuracy, redefining efficiency in NLP applications.
Different link prediction models capture unique knowledge, but combining them reveals a saturation point, leaving many queries unanswered.
Small Language Models can outperform larger, general-purpose LLMs in predicting evolving user preferences, revolutionizing eCommerce personalization.
DKL boosts RAG accuracy by over 25 points in retrieval failure cases without the need for costly instruction fine-tuning.
LLMs systematically underrepresent the diversity of their training data, with a notable conditional diversity gap that can be mitigated through a novel entropy-constrained projection method.
CoSPOT achieves superior online time series forecasting by leveraging frequency-domain insights, allowing it to adapt to unseen patterns with minimal parameter updates.
Tailoring steering directions to specific inputs boosts LLM truthfulness by nearly 10% compared to static methods.
Structured reasoning in LLMs can dramatically boost diagnostic accuracy in telecom root cause analysis, overcoming the limitations of vanilla models.
Compliance with smaller requests can dramatically increase after a refusal in some models, but backfires in others, revealing crucial differences in how language models process human-like persuasion techniques.
DIFFIE reveals that leveraging diffusion stochasticity can significantly enhance the efficiency and effectiveness of OpenIE systems, outperforming traditional methods.
SCX Router achieves a 1.5% performance boost over the best fixed model by intelligently selecting the most suitable LLM for each task in real-time.
Fine-tuned encoders reveal unexpected asymmetries in cross-lingual transfer, exposing the nuanced challenges of political evasion in Romanian discourse.
NE-R1 achieves a remarkable 2.52% F1 score improvement in in-domain NER tasks by intelligently balancing parametric and external knowledge retrieval.
PEARL achieves state-of-the-art performance in inductive knowledge graph completion by intelligently aligning relational paths with their contextual subgraphs.
EmoStance achieves a remarkable 62.2% win rate in generating more contextually aware and responsive empathetic dialogues by harnessing emoji distributions as weak supervision.
Sentiment in social-media conversations is significantly influenced by discourse moves, with specific interventions capable of mitigating negativity or amplifying hostility.
Task-level natural-language priors can transform low-resource LLM training, boosting performance even with minimal data.
Optimizing inspection strategies in reverse logistics can yield significant financial gains, with one method adding nearly $54k per batch in aircraft maintenance alone.
PILL achieves up to 6.0 BLEU-2 improvement on text infilling while running 1.82x faster than the best existing method, revolutionizing efficiency in diffusion language models.
Achieving a staggering 99.50% acceptance rate in synthetic dialogue generation reveals the transformative power of feedback-guided refinement in meeting complex communicative constraints.
Ragebait is not just prevalent; it thrives in politically charged discussions, spreading faster and provoking stronger negative reactions than typical posts.
More extensive Cantonese-specific training significantly enhances model predictions of reading behavior, revealing nuanced insights into language processing.
The effectiveness of persona prompting in aligning LLMs with human survey responses hinges on strategic attribute selection, not just quantity.
HyperStyler achieves superior style transfer fidelity and semantic preservation with just 2.4% more parameters than T5-large, while being over 1.8x faster than LLMs.
LLMs can enhance ontology rankers for rare-disease diagnosis, boosting recall significantly while retaining valuable evidence trails.
NER-driven extraction transforms complex radiology reports into clear, accurate summaries, while RAG can inadvertently introduce errors.
Dutch language models exhibit alarming biases, with some favoring stereotypical representations of transgender identities up to 97% of the time.
Loneliness in older adults can be detected through a powerful combination of speech patterns and vocal characteristics, revealing critical insights into their emotional states.
A novel layered taxonomy reveals critical insights into the complexities of grammatical error annotation in Chinese learner writing, bridging computational and pedagogical perspectives.
LLMs can significantly enhance user interactions by mastering question clarification, as revealed by a novel tri-agent evaluation framework that benchmarks their performance.
Despite widespread availability, Python type hints are adopted inconsistently, with maintainers favoring explicit annotations for public interfaces over inferred types.
Transforming verbose queries into concise keywords can elevate segmentation accuracy from 20% to nearly 77% in complex dynamic scenes.
A novel approach reveals that existing MRI report generators can be enhanced to accurately diagnose brain tumors by leveraging latent features, achieving up to 92% accuracy in identifying meningiomas.
InsightSeg transforms past correction episodes into actionable insights, enabling segmentation agents to prevent errors before they occur.