Search papers, labs, and topics across Lattice.
100 papers published across 3 labs.
OCR technologies are evolving rapidly, yet significant challenges remain in recognizing diverse scripts and handwritten text, demanding innovative solutions for real-time applications.
The first-ever line-level dataset of historical Arabic manuscripts with detailed margin annotations unlocks new avenues for understanding non-linear reading orders in ancient texts.
LLMs reveal a complex geometric structure of moral knowledge that captures the tension in moral dilemmas, rather than offering simplistic moral resolutions.
LAMA not only boosts platform revenue but also maintains user response quality, challenging the status quo of traditional advertising models.
FiUni reveals that task-free continual learning can be achieved by intelligently leveraging Fisher information, enabling LLMs to adapt without explicit task boundaries.
LLMs reveal a complex geometric structure of moral knowledge that captures the tension in moral dilemmas, rather than offering simplistic moral resolutions.
LAMA not only boosts platform revenue but also maintains user response quality, challenging the status quo of traditional advertising models.
FiUni reveals that task-free continual learning can be achieved by intelligently leveraging Fisher information, enabling LLMs to adapt without explicit task boundaries.
Internal activation dynamics can reveal factual errors in LLM outputs without the costly overhead of multiple decoding passes.
Tailoring academic advising responses to individual student profiles boosts relevance and user satisfaction, outperforming traditional methods.
Quantizing small language models can be predictable and efficient, with a novel metric that identifies optimal layers for speed and quality trade-offs.
Repeated paraphrasing can significantly enhance the indistinguishability of AI-generated text from human writing, challenging existing detection methods.
Pruned LLMs can still generate coherent text by strategically guiding token selection, reducing repetition loops and enhancing overall quality.
Clinical language models can be made auditable and robust against deployment shifts by effectively suppressing misleading artifacts, revealing a clear path to trustworthy AI in healthcare.
TransMeme achieves a 33.1% improvement in meme transcreation quality by expertly balancing cultural adaptation and multimodal coherence.
Misleadingness in public discourse is more complex than mere fabrication, with emotional arousal and communicative intent playing critical roles in shaping reader interpretations.
Bilingual pretraining can create hidden-state mismatches that mislead researchers about the comparability of multilingual models.
Senior employees leverage generative AI more effectively, but surprisingly, training doesn't boost sophistication across the board.
LLMs exhibit a troubling performance gap in Braille comprehension, with Grade 2 Braille proving especially challenging, highlighting urgent needs for inclusive AI design.
By embedding online learning directly into LLM call latencies, this framework can cut query costs by nearly 8x, revolutionizing how we optimize semantic data processing.
BPMN4CAI transforms how we model conversational AI, enabling adaptive decision-making and enhanced context management in business processes.
DSA orchestrates LLM agents to enhance stock research by managing evidence and model capabilities, ensuring rigorous report generation without sacrificing operational integrity.
Achieving a Term-Typing F1 of 0.9200 reveals the potential of retrieval-augmented generation in overcoming traditional ontology learning hurdles.
LLMs can outperform humans in recall for screening tasks, but their effectiveness hinges on the workflow design rather than the model itself.
ColRel reveals that even weak metadata can yield meaningful insights when enhanced with contextual business knowledge, transforming how we navigate complex data lakes.
Distinguishing between AI-generated content and human expression can improve detection accuracy by over 6% in mixed-origin text scenarios.
LiveSim transforms user behavior simulation by dynamically adapting to the evolving interactions in live-streaming environments, leading to unprecedented accuracy in risk analysis.
Professional editing can dramatically skew AI text detector outputs, with false positive rates for human-written texts reaching as high as 100% depending on the detector used.
Compact models can now outperform massive systems like GPT-4o in document translation by leveraging a new metric that prioritizes structural fidelity.
Forgetting in language models just got smarter—GRAPHSU slashes knowledge leakage by nearly 50% by targeting support routes, not just forget seeds.
Fragment-based reasoning in machine translation could redefine how LLMs leverage in-context samples for improved accuracy and reliability.
Unbiased sampling of source prefixes in TLMs can reduce runtime by several orders of magnitude while maintaining accuracy in estimating target prefix probabilities.
LLMs exhibit a distinct inter-sentence transition variance that can be exploited for highly accurate AI-generated text detection, achieving over 97% accuracy.
Auditable pair-level evidence consolidation reveals that traditional methods outperform LLMs in precision for identifying historical text reuse.
Monolingual models can achieve cross-lingual alignment without joint training, revealing the power of linguistic structure over shared parameters.
KinyaEmbed outperforms state-of-the-art multilingual models by over 40% in Kinyarwanda sentence embeddings, setting a new standard for low-resource languages.
LLMs can not only track belief states but also geometrically organize them in a way that mirrors the underlying statistical dynamics of their latent variables.
Zero-shot LLM agents struggle to predict wellbeing scores from longitudinal data, often performing no better than a basic mean baseline.
Parsing Korean requires a nuanced approach, as fine-grained morphological details significantly enhance accuracy over simpler representations.
Authentic telemedicine conversations in Bengali reveal that existing models can significantly improve their performance on clinical reasoning tasks with the right data.
TabuLM outperforms multilingual baselines by over 11 points, showcasing the power of morphology-aware embeddings in low-resource language models.
Affirmation, not Reflection, emerges as the key behavioral marker for high-quality text-based counseling sessions, challenging established norms in the field.
Cascaded batch prompting resolves the unpredictability of conventional batch prompting, achieving both superior performance and efficiency in large language models.
The first-ever line-level dataset of historical Arabic manuscripts with detailed margin annotations unlocks new avenues for understanding non-linear reading orders in ancient texts.
PragAlign significantly outperforms traditional methods in multilingual reply assistance, revealing critical insights into cultural and linguistic appropriateness.
ITL achieves precise document alignment with structured frameworks, ensuring that every result is traceable to the underlying terminology.
Cross-lingual word-to-speech mappings can be effectively learned from visual grounding without the need for transcriptions or extensive model training.
Homophones in English are not just phonetically identical; they reveal significant differences in pronunciation that reflect their meanings.
DARD achieves a remarkable 2.71× speedup in dLLM inference while enhancing output quality by 4.35 points on CIDEr, redefining the efficiency landscape for large language models.
Prioritizing semantic anchors over easy tokens can drastically reduce error accumulation in multimodal language models.
Sparse updates from surgical alignment can boost reasoning quality in LLMs, even when accuracy takes a hit.
Achieving 88.5% accuracy in identifying buggy classes from bug reports, this multi-objective approach could revolutionize the efficiency of software debugging.
LLMs can optimize model-based test generation, outperforming traditional tools in both path efficiency and scalability.
OCR technologies are evolving rapidly, yet significant challenges remain in recognizing diverse scripts and handwritten text, demanding innovative solutions for real-time applications.
UniGeo achieves a remarkable 13.59-point improvement in retrieval accuracy for text-guided drone geo-localization by leveraging a unified multimodal framework.
Aggregating insights from neighboring users' reviews can transform how we tackle data sparsity in recommender systems, leading to better predictions and richer explanations.
SFT can dramatically reduce instruction sensitivity in smaller models, but its effectiveness diminishes in larger architectures, revealing a nuanced relationship between model size and fine-tuning outcomes.
Real-time updates in conversational recommendations can transform the e-commerce landscape by ensuring users always see the latest products without lag.
Navigating the fragmented landscape of explainability tools just got easier with Virgil, a system that empowers both experts and non-experts alike.
Collapse in On-Policy Self-Distillation narrows reasoning paths, revealing critical biases that could undermine model performance.
Traces of a forgotten language can enhance re-learning speed by 14%, challenging the notion of critical periods as purely maturational.
Bridging the gap between attribution and counterfactual explanations, this framework achieves unprecedented stability and reliability in time series model interpretability.
Existing KPA benchmarks fail to deliver reliable evaluations, but a new structure-aware benchmark reveals significant improvements in coherence and quality of key points.
Transforming scientific papers into multi-turn generation trajectories not only doubles the training data but also boosts academic writing benchmarks while maintaining reasoning skills.
Low-probability tokens disproportionately influence model updates, and a simple reweighting strategy can significantly enhance performance without sacrificing generalization.
Language-model agents in SwarmWorld can self-organize to build resilient technological societies, surpassing isolated search methods in innovation and adaptability.
Data citation in large language models is not just a verification issue—it's a complex challenge that demands new frameworks for credit and provenance.
A runtime monitoring system that achieves 85% accuracy in detecting air traffic control procedural violations could significantly enhance aviation safety protocols.
MoganBert-TR outperforms traditional masked language models by up to 3.7x in Turkish retrieval tasks, redefining benchmarks for Turkish NLP.
Forgetting previous episodes in mixed-topic conversations can lead to significant accuracy drops, but TSIM shows how to maintain context integrity and improve performance in chat assistants.
Whisper's adaptation for Baniwa ASR reveals that multilingual models can effectively bridge the gap for low-resource languages, achieving competitive error rates.
EmoVec enables real-time emotional control in LLMs, enhancing their affective responses without retraining.
Local mismatches in dialogue can paradoxically lead to stronger preservation of prior beliefs rather than immediate revisions, challenging conventional assumptions about listener behavior.
Conceptual disruptions in reading trigger immediate, localized processing costs, while referential disruptions unfold gradually, revealing fundamental differences in how meaning is processed by humans and LLMs.
Encoder models hold their ground against generative LLMs in ASR evaluation, but the latter enhance interpretability and hypothesis selection.
LLMs' personalities are dynamic and layer-dependent, revealing that quantization can significantly disrupt their behavioral consistency.
Romanization during pretraining can dramatically enhance multilingual model performance, outpacing traditional text-based approaches.
Active learning can dramatically reduce the annotation burden in summarization tasks, with LOBSTER achieving up to 665x faster query selection without sacrificing performance.
Spectral analysis reveals that the ability to detect machine-generated text hinges on text length and generation style, challenging conventional detection methods.
LLMs can now proactively correct user input errors, significantly boosting their accuracy and reliability in real-world applications.
Instruction-tuned LLMs can outperform specialized models in hate speech detection, achieving state-of-the-art results across multiple domains and languages.
Claim-locked reporting boosts the accuracy of LLM-generated statistical reports by over 37%, ensuring that evidence integrity is maintained throughout the writing process.
Expert-informed Skill prompting achieves remarkable cross-dataset consistency, outperforming traditional fine-tuning in stability while still allowing for high accuracy on native data.
Memory management strategies that adapt to individual users can significantly enhance agent performance, outperforming traditional static approaches.
ClueWeaver transforms how compact language models tackle long narratives, achieving superior evidence retrieval and reasoning transparency.
VietAIDetector achieves superior detection of AI-generated Vietnamese text without requiring any domain-specific training data, setting a new standard for language-specific AI content verification.
Leveraging speech act information can dramatically enhance derailment forecasting accuracy, especially in low-data environments.
TOPAS slashes job completion times by up to 49.4% in multi-agent LLM serving by intelligently balancing prefix caching and request scheduling.
Unsupervised query correction can achieve superior performance by cleverly encoding phonetic and visual similarities, avoiding the pitfalls of intent drift.
Mixed-policy reinforcement learning can enable language models to absorb knowledge more effectively than traditional supervised fine-tuning, especially in complex reasoning scenarios.
Broad financial competence scores can mislead practitioners, as they may overlook critical operational reliability in professional workflows.
Switching between natural language and structured graphs can boost multi-agent LLM performance by over 12 percentage points while slashing token usage by more than threefold.
Language communities shape war narratives on Wikipedia, revealing that even supposedly neutral accounts are influenced by cultural perspectives.
Multi-turn context significantly enhances harm detection in AI conversations, yet LLMs still falter in understanding relational nuances and severity.
LLMs can reveal critical information from OOXML documents that is invisible in Microsoft Office, with up to 76% of trials exposing hidden facts.
Spear-phishing exploiting greed outperforms all other social engineering tactics, revealing critical insights into psychological manipulation.
EAVA not only predicts software vulnerabilities but also provides actionable evidence, bridging the gap between automated assessments and human validation.
Event Tokens can dramatically enhance LLM recommendation systems by capturing rich interaction context, leading to significant performance gains across various benchmarks.
Users may prefer a less effective interface, believing it enhances their knowledge acquisition, revealing critical insights into user perception versus actual performance.
CloSeR achieves state-of-the-art GCD performance by elegantly decoupling closed-set recognition from open-set discovery, minimizing objective conflicts and enhancing semantic coherence.
LLMs can show drastically different skill levels depending on the language used, revealing a hidden barrier to true multilingual capabilities.
VoiceMem achieves a 30-point accuracy boost over traditional memory systems while delivering real-time, emotionally aware interactions without added latency.
Admissibility in compositional generalization reveals distinct structural profiles that can diagnose training corpus limitations without the need for predictive modeling.
Nominal facility availability can mislead urban planners, as residents with mobility limitations face greater travel burdens than expected.
Existing ASR systems falter in recognizing code-switched speech, with the best performing system still achieving a staggering 35.93% word error rate.