Search papers, labs, and topics across Lattice.
100 papers published across 2 labs.
Agents trained with the Preference Tree Optimization framework achieve unprecedented improvements in goal-oriented dialogue, outperforming traditional methods in both satisfaction and strategic planning.
Bengali speakers face a staggering 67:1 training-token deficit compared to English, highlighting a critical inequity in AI language support.
Attention mechanisms can amplify the visibility of latent variables by up to 17x, depending on task demand, challenging our understanding of how language models manage internal representations.
Achieving the accuracy of a 12-layer GPT-2 Small with just a 3-layer Gated Recurrent Transformer showcases a groundbreaking efficiency in model design.
Personalization in AI co-scientists could be the key to unlocking novel research insights that generic systems overlook.
Attention mechanisms can amplify the visibility of latent variables by up to 17x, depending on task demand, challenging our understanding of how language models manage internal representations.
Achieving the accuracy of a 12-layer GPT-2 Small with just a 3-layer Gated Recurrent Transformer showcases a groundbreaking efficiency in model design.
Personalization in AI co-scientists could be the key to unlocking novel research insights that generic systems overlook.
Controlled genre expansion, rather than just story-centric data, is crucial for unlocking robust creative writing skills in LLMs.
Popular facts are harder to forget, but AdaPop can effectively unlearn them while retaining essential knowledge, outperforming existing methods by a significant margin.
RMM reveals that optimizing attention-side computations can lead to substantial runtime gains in Transformer models without sacrificing accuracy.
DMDIntel reveals that LLMs can be interpreted with greater accuracy by leveraging dynamic mode decomposition, outperforming traditional methods in token attribution.
Classical smishing detection models can fail catastrophically under adversarial attacks, while transformer-based models exhibit surprising resilience, challenging assumptions about clean-text performance.
LLMs may outperform embedding models in reasoning, but the cost difference is staggering—up to 1,431 times more expensive for marginal gains.
REAG transforms acceptance testing for LLM-based software by achieving a 3.91 to 4.30 improvement in oracle quality while ensuring 98.8% accuracy in verdict reliability.
Achieving phoneme conversion in under 0.15 ms could revolutionize real-time Thai text-to-speech applications.
Segment-level consolidation in LycheeMemory V2 slashes construction costs by up to 86% while maintaining high accuracy in long-term memory tasks.
Every exactly calibrated convex surrogate for the Jaccard measure demands exponentially more prediction dimensions than previously understood.
DARTree achieves a staggering 9.73× speedup in autoregressive decoding while accepting nearly 99% more tokens per round than existing methods.
Users often ignore provider-set defaults in LLM services, opting instead for their own optimal token allocations unless convenience is prioritized.
M&M's analysis reveals that M&G's approach may mislead researchers about the true capabilities of MAML in capturing Bayesian learning principles.
ReconSpan achieves superior text preservation with adaptive tokenization, outperforming random chunking by retaining more contextual information.
PRISM reveals that large language models exhibit cognitive specialization akin to that of aphasic patients, challenging our understanding of LLM interpretability.
LLMs can recognize when they lack knowledge about a referent but still choose to fabricate specific details instead of opting for safer, generic responses.
AaLLM can generate innovative circuit topologies that outperform conventional designs while slashing design time by up to 40x.
Transformers can length-generalize on specific regular languages, revealing a hidden algebraic property that classical finite decomposition theory fails to capture.
Predicting both courses and grades together can cut prediction errors by nearly half, revolutionizing how we assess student performance.
Authority-aware retrieval in parliamentary transcripts leads to a 0.97 coverage rate across political groups and perfect quotation accuracy, setting a new standard for multi-view RAG systems.
Transforming vague security intents into fully compliant network topologies, TopoIntent achieves perfect CIS Controls satisfaction in under 1.5 iterations.
Gendered language in prompts can lead to significantly poorer responses from LLMs, revealing a hidden bias that affects workplace communication.
StateBridge reveals that training-free hidden-state alignment can significantly enhance communication efficiency in LLM multi-agent systems, outperforming traditional methods.
Self-referential prompts lead to a striking 64% increase in response instability compared to verifiable questions, revealing the unpredictable nature of LLMs' subjective reports.
Deferring translation in multilingual question answering can significantly reduce errors and improve efficiency, leading to more accurate results across diverse languages.
SPADE slashes cloud model calls by 76% while preserving accuracy, revolutionizing the deployment of large language models in edge environments.
Training data influence shifts dramatically over the course of language model pretraining, with literature data dominating early and STEM data taking over later.
CROP reveals that prioritizing task relevance in token supervision can boost performance by nearly 3 points, challenging conventional selection methods in OPD.
Emotional companionship capabilities of LLMs can be significantly influenced by their attachment styles, which can be shaped through targeted prompting.
Achieving first place in KBQA competitions, HybridRAG-BN demonstrates that effective retrieval and verification can significantly enhance answer accuracy for low-resource languages like Bangla.
EviReform reveals that leveraging evidence from retrieved passages can significantly boost multi-hop retrieval performance, achieving up to 5.59 points in Recall@5.
Analyzing 57.5K transactional prompts reveals a surprising Zipf-like distribution in usage patterns, challenging assumptions about prompt effectiveness across different contexts.
Multilingual models may excel in general tasks, but they falter significantly when faced with region-specific cultural knowledge, as shown by BavGround's rigorous evaluation.
Gender imputation can yield valid measurements for studying disparities, but its legitimacy hinges on the context and the populations affected.
Achieving 58.2% accuracy on conversational-memory tasks, ReFind shows that unstructured chat logs can rival complex memory systems when paired with intelligent search controls.
The model Gemma 3 4B IT reveals a striking distinction in its representation of falsehoods and impossibilities, challenging our understanding of how AI interprets language.
CRAFT achieves unprecedented accuracy in reconstructing symptom timelines from sparse clinical narratives, transforming how we interpret temporal data in healthcare.
Vietnamese social media sentiment analysis faces unique challenges, with implicit sources and vocabulary ambiguities complicating emotion detection.
Client simulations can now exhibit realistic resistance and emotional depth, transforming how novice counselors are trained and evaluated.
Users frequently trust AI-generated legal advice based on its appearance and emotional appeal, rather than its accuracy, revealing a troubling gap in verification practices.
Misrecognition, rather than mere misrepresentation, is the critical issue in ensuring fairness in generative AI, reshaping how we approach social justice in technology.
Achieving over 98% accuracy in classifying infected mosquitoes from video frames showcases the power of combining vision and language models for biological analysis.
AnnoIndex achieves a remarkable F1 score of 0.87 by transforming unstructured text into a structured format, enabling precise analytical queries that traditional methods struggle with.
Multilingual embeddings can boost cross-lingual retrieval accuracy to over 96%, far surpassing traditional translation methods.
No LLM can dominate all tasks in evidence synthesis, revealing the critical need for a human-in-the-loop approach to ensure comprehensive analysis.
Achieving a 5.28% improvement in segmentation accuracy while running 4.7% faster than existing methods, DiCoR redefines efficiency in referring remote sensing image segmentation.
MLLMs can leak sensitive personal information when visual evidence is lacking, but the Dynamic Relational Unlearning Framework (DRUF) significantly mitigates this risk without sacrificing performance.
A groundbreaking watermarking method that not only traces LLM-generated content but also detects tampering with unprecedented accuracy.
Instruction tuning boosts model confidence but often at the cost of rationale diversity and calibration accuracy.
Achieving state-of-the-art results for Danish with a model that uses only permissible data challenges the notion that larger datasets are necessary for competitive performance.
LITTLELEARNER reveals that even a well-defined knowledge scope can yield a competent language model, but it won't expand its capabilities beyond its educational boundaries.
Evaluations reveal that current models struggle with long-range narrative integration and cultural reasoning, highlighting a critical gap in video understanding capabilities.
Class imbalance in consumer reviews can significantly hinder positive sentiment detection, but SVM and Bidirectional LSTM offer robust solutions for accurate sentiment analysis.
Transforming EEG decoding into a continuous semantic embedding task unlocks new levels of generalization in EEG-language models.
Automated construction of Dynamic Master Logic models can transform technical documentation into actionable insights for complex system diagnostics.
Larger, R&D-focused firms are not just adopting ChatGPT Enterprise faster; they’re also leveraging it more intensely across diverse job functions, particularly among early-career employees.
TELLME boosts language model performance by over 23% while slashing the need for extensive domain-specific datasets through innovative quiz-based training.
Gist-based context compression can boost reasoning tasks but dramatically falters on temporal questions, revealing a critical blind spot in current language model designs.
Supervised fine-tuning outperforms complex reinforcement learning techniques in ensuring multilingual API reliability, challenging the notion that more sophisticated methods are always necessary.
Rephrasing benchmark problems can flip model answers, revealing that stronger LLMs are paradoxically more fragile to wording changes than weaker ones.
Autonomous Semantic Solitons emerge from a new dynamical framework, enabling LLMs to generate diverse outputs while avoiding stagnation.
Prioritizing lexically rich narratives can accelerate ASR transcription quality improvements by overcoming the cold start problem in language documentation.
LLMs exhibit striking inconsistencies in political evaluations, influenced by prompt design and model persona, raising critical questions about their role in shaping public opinion during elections.
Proactively committing mid-entropy pivot positions can accelerate dLLM decoding by up to 18 times while improving accuracy.
Financial time series data can boost classification accuracy in monetary-policy stance detection, achieving over 70% F1 score with minimal human annotations.
Language-Conditional Dequantization recovers up to 83% of the perplexity gap for non-Latin languages, challenging the notion that quantization is uniformly detrimental across languages.
The reliance on knowledge bases as gold standards in machine translation may inflate performance metrics, masking the true quality of translations in low-resource settings.
Achieving 99.5% citation validity and 96.4% figure editability, Spark-to-Paper transforms how research papers can be generated with unprecedented reliability and efficiency.
Groundedness Drift reveals that explanations can mislead in the presence of backdoor attacks, highlighting vulnerabilities in language model classifiers.
Lightweight NLP models can uncover hidden patterns in public procurement, revealing potential irregularities with impressive accuracy.
Separating firm-specific signals from macroeconomic indicators can lead to significantly higher portfolio returns in small-cap trading strategies.
Bengali speakers face a staggering 67:1 training-token deficit compared to English, highlighting a critical inequity in AI language support.
Training LLMs with long contexts can paradoxically weaken their ability to retain knowledge, leading to poorer performance when context is not available.
A structured hierarchy of perspective-related concepts reveals how to effectively navigate and operationalize perspectives in NLP research.
A specialized clinical RAG system outperformed cutting-edge LLMs on a comprehensive medical benchmark, proving that context-specific design can yield superior results.
Achieving a BLEU score of 29.26, this system outperforms leading models while translating across 12 distinct Bangla dialects without relying on a standard pivot.
Agents trained with the Preference Tree Optimization framework achieve unprecedented improvements in goal-oriented dialogue, outperforming traditional methods in both satisfaction and strategic planning.
A-CRC-QA achieves superior reliability in selective question answering by effectively controlling error rates without the need for retraining.
Compressing a reliable large model via quantization yields Small Language Models that are not only more trustworthy but also more adaptable than those trained from scratch.
DexterSQL boosts Text-to-SQL accuracy by over 2.7% through innovative schema exploration and rule-based corrections that tackle common generation pitfalls.
Engaging with AI can trigger a profound "philosophical vertigo," reshaping our understanding of reality and authority in unsettling ways.
Hate speech targeting migrants surges in neighborhoods experiencing recent demographic shifts, revealing a troubling pattern of exclusionary discourse.
Indian foundation models excel in traditional benchmarks but struggle with newer evaluations, revealing critical gaps in the national AI ecosystem.
Modern LLMs rarely refuse to discuss restricted content, instead opting for nuanced warnings that reveal a significant shift in content moderation strategies.
Transparency about an AI's persuasive intent can cut its influence by nearly 50%, challenging the efficacy of current disclosure regulations.
Source-linked safety knowledge support could revolutionize how medical device developers manage compliance and risk in an increasingly complex regulatory landscape.
Despite the promise of multi-agent systems in software engineering, key frameworks still lack advanced features and show no significant performance difference in summarization tasks.
Continuous non-autoregressive models outperform discrete methods in speech enhancement, revealing a critical shift in paradigm effectiveness.
Token-level credit assignment can drastically improve the effectiveness of generative document retrieval, leading to better alignment between generation and relevance.
Personalized feedback-driven recommendations can boost literature discovery effectiveness by over 10% in aligning with user preferences.
Evolving skills in frozen LLMs can lead to substantial performance gains, allowing smaller models to rival their larger counterparts without parameter updates.
Language models may ace wrong-case citation detection but fail to verify pinpoint accuracy, missing 40% of critical errors even in advanced configurations.
A unified framework reveals 15 critical trends and 21 unexplored opportunities in the lifecycle of privacy documents, spotlighting urgent challenges in AI-driven environments.
AI-generated novels show a striking lack of formal diversity, with repeated generations compressing sentence structure far more than human authors do.
HPSE transforms how LLMs integrate new knowledge, enabling them to answer atomic questions and perform multi-hop reasoning with edited facts.
Emotion-driven feedback can transform multi-turn dialogue systems, boosting empathetic responses and model performance significantly.
ASR-roundtrip evaluation can miss nearly half of the critical reading errors in Chinese news TTS, revealing a significant gap in current assessment methods.