Search papers, labs, and topics across Lattice.
100 papers published across 5 labs.
Automating persona generation for user simulators can drastically reduce the labor involved in testing interview dialogue systems while enhancing the diversity of user interactions.
Subtle prompt changes can destabilize LLMs significantly, but four key factors can mitigate this sensitivity by targeting low-order interactions.
Hallucinations in LLMs may blur the line between advanced simulation and true cognitive processes, challenging our understanding of AI consciousness.
Turn-taking prediction in Turkish conversations can be significantly improved using a novel multimodal dataset and a hybrid rule-based approach that captures complex interaction cues.
Lexical convergence in LLMs can be significantly influenced by peer-ranked feeds, but distributed sources fail to provide a reliable advantage in shaping agent opinions.
Turn-taking prediction in Turkish conversations can be significantly improved using a novel multimodal dataset and a hybrid rule-based approach that captures complex interaction cues.
Lexical convergence in LLMs can be significantly influenced by peer-ranked feeds, but distributed sources fail to provide a reliable advantage in shaping agent opinions.
Early-stage solutions can reliably inform final outcomes, leading to a 56.9% reduction in the primal gap for MILP problems.
Models trained on incomplete multimodal data can significantly improve generalization to unseen combinations, achieving over 5% higher accuracy than existing approaches.
Skill selection in LLMs can be optimized to achieve a 0.73 task success rate while using 28% fewer tokens than existing methods.
Injecting new subjects into a language model's output can boost perceived novelty by over 1.4 points, transforming how we understand model engagement.
Iterative proxy correction boosts sentiment analysis accuracy by refining initial proxies and adapting to incomplete multimodal inputs.
Automating persona generation for user simulators can drastically reduce the labor involved in testing interview dialogue systems while enhancing the diversity of user interactions.
Self-consistency outperforms traditional confidence measures, reducing financial NER errors from 34.3% to below 2% in high-confidence scenarios.
Users show markedly different levels of acceptance for their own AI agents compared to those of others, revealing critical insights into the dynamics of AI integration in online dating.
PersonalBench reveals that current LLM personalization methods can differentiate author styles but fall short of achieving human-level writing quality, with a stark similarity gap.
Falcon-7B with Summarize Chains not only automates financial news summarization but does so with remarkable accuracy, outperforming traditional methods and exposing critical weaknesses in RAG techniques.
Cross-lingual fairness gaps in language model watermarking are not just language-specific but are fundamentally tied to the structural properties of language families.
Open-vocabulary word-level information can be reliably decoded from EEG during silent reading, revealing a scalable approach to understanding inner speech.
Restricting evidence visibility in language model societies can boost compositional generalization, leading to a 20-point performance increase over fully visible models.
Nested sequential Monte Carlo methods can dramatically improve inference-time control in text generation, outperforming traditional techniques in steering towards desired outcomes.
LLMs can detect 80% of borderline authentication anomalies, far surpassing traditional methods that struggle with such cases.
Uncertainty in LLM reasoning is largely a sampling artifact, allowing for more efficient analyses that cut costs without sacrificing accuracy.
G-CARL not only improves the accuracy of medical report interpretations but also ensures they are tailored to patient queries, outperforming traditional methods in both factuality and user satisfaction.
No current LLM can accurately identify missing legal information in user queries, with all evaluated models struggling to balance responses to both deficient and complete questions.
Subtask-level skills can boost LLM performance beyond baseline levels, while task-level skills often hinder it—highlighting a critical distinction in skill transferability.
Frontier LLMs struggle with contract scrubbing, achieving only modest recall rates despite their prowess in general benchmarks.
SABET-QA outperforms existing methods by effectively refining reasoning states across multiple hops, making it a game-changer for complex temporal queries.
Fine-tuning a language model for crowd simulation can significantly improve accuracy using only aggregate mobility data, achieving a 25% reduction in destination-share error.
By clarifying ambiguous patient queries, this framework boosts diagnostic accuracy by over 57 percentage points, transforming how healthcare chatbots interact with patients.
Achieving a Micro-F1 score of 0.8085, this framework effectively balances adaptation and retention in LLMs for smart contract vulnerability detection.
Event memory enables real-time soccer commentary that adapts seamlessly to the evolving context of live matches, outperforming traditional methods.
OenoBench reveals that even leading LLMs struggle with knowledge retention, achieving only 53%-84% accuracy on wine-related questions despite extensive training.
Sarcastic-aware contrastive regularization enables the model to discern nuanced sarcasm, outperforming traditional methods that struggle with modality inconsistencies.
Fine-tuning with sparse attention can outperform traditional exact attention models while running efficiently on modest hardware.
Performance gaps in multilingual medical evaluations reveal that proprietary models outperform open-source ones, but translation quality can swing results dramatically.
Static profiles lead to identity essentialism in LLMs, but a new longitudinal memory framework reveals a path to richer, more diverse social simulations.
UniLang enables pretrained LLMs to seamlessly integrate machine-native symbols, outperforming traditional models in diverse structured prediction tasks.
Prosody can boost decision-making accuracy in dialogue systems by nearly 25% when it conveys user concerns that words alone cannot express.
TextRefine achieves superior text editing in product posters by ensuring high fidelity and optimal placement, overcoming common pitfalls of existing models.
Madhhab-aware filtering can more than double retrieval accuracy for school-specific fiqh questions, revealing a crucial dimension in Islamic jurisprudence retrieval.
IAR achieves a remarkable boost in retrieval-free question answering, outperforming traditional methods by effectively internalizing document knowledge into LLMs.
Achieving a 25.3% relative error reduction in character error rates, Phoenix redefines manuscript transcription as an auditable evidence management task rather than mere text replacement.
Expert deliberations can be mined for predictive financial signals, with sentiment analysis proving more effective than keyword frequency in forecasting asset performance.
Sequence-pooled normalization can deliver nearly the same performance as larger receptive fields, redefining our understanding of context in convolutional sequence labeling.
A novel computational framework reveals how team communication dynamics in VR can be dissected into coherent phases that align with specific actions, enhancing our understanding of collaboration.
Subtle prompt changes can destabilize LLMs significantly, but four key factors can mitigate this sensitivity by targeting low-order interactions.
Achieving a 95.3% functional accuracy and a 93.0% hallucination-free rate, this multi-agent platform redefines the standards for conversational business intelligence.
LLMs can autonomously generate prompts that rival expert-written ones, but still fall short in accurately interpreting scientific context and discovering literature.
ORBITER significantly boosts decision-making reliability in last-mile delivery, outperforming existing models by up to 9.2% through enhanced reasoning about spatiotemporal cues.
PWAL outperforms traditional methods by boosting accuracy by up to 30.86 percentage points while providing a transparent trace of logical decision-making in enthymeme completion.
A new characterization of string tuple factorizations reveals deeper structural insights into the multiple context-free grammar properties of $O_2$.
Iterative fine-tuning of OCR can drastically cut down the time and expertise needed for transcribing complex historical manuscripts.
Achieving a mean F1 score of 0.911 in Tangut word segmentation reveals the power of combining traditional lexicons with modern machine learning techniques in resource-scarce scenarios.
Complete evidence retrieval can be dramatically improved by strategically allocating context, achieving a 23.8% boost without increasing the retrieval budget.
Online spaces with politically diverse audiences are hotspots for cross-partisan dialogue, revealing unexpected avenues for political engagement.
Interactive correction of fish tracking predictions via natural language guidance shows promise, yet highlights the need for better integration of user input.
A single agentic framework can seamlessly tackle multiple intelligent document processing tasks, outperforming specialized models in the process.
FRAGMENT's factorized graph representation enables precise document generation and editing, outperforming conventional models in capturing complex dependencies.
Social.Wiki redefines web ownership by enabling communities to collaboratively create and govern their online spaces without centralized control.
The retention of meaning in language models can vary from half to zero based solely on the positions of edits, challenging existing assumptions about watermark robustness.
Creators prefer steering adaptive lyrics through explicit controls, reshaping their creative process and audience engagement.
Real-time risk triage in mental health supervision is now possible, reducing response times from days to seconds.
Clients transformed their view of AI from a mere tool for answers to a collaborative partner in problem-solving.
A single model can now achieve state-of-the-art performance in both word and sentence alignment across multiple languages, simplifying the alignment process significantly.
Fine-tuned small language models can outperform massive counterparts by leveraging high-quality filtered data, challenging the notion that bigger is always better in machine translation.
Languages with overt inflection share more agreement circuitry, revealing that multilingual LLMs reuse computational structures instead of relying on distinct solutions for each language.
Failure-mode contextual bandits can boost model accuracy by over 4% on standard benchmarks while eliminating the need for additional human annotation.
LLM-Detector achieves superior anomaly detection in tabular data without the need for fine-tuning, making it a game-changer for real-world applications.
SynFlow reveals the intricate interplay of syntax, morphology, and semantics in lexical change, offering a holistic view that traditional methods miss.
LLMs are becoming less diverse in their creative outputs, potentially stifling human agency in co-creative processes.
Politically charged topics on Reddit exhibit significant directional drift over time, while music and sports discussions remain surprisingly stable.
DeepWeaver transforms the way LLMs synthesize evidence, leading to answers that are not only more comprehensive but also better grounded in citations.
Fidelity-oriented translations significantly enhance perceived quality and trust compared to readability-focused outputs, especially for simpler narratives.
Medical QA systems can achieve unprecedented accuracy by leveraging a multi-agent framework that integrates adaptive memory and structured reasoning.
Exploitation, not exploration, is the critical bottleneck in test-time scaling for language models, with selection processes yielding near-random results despite rich candidate pools.
Tailor your text analysis with a revolutionary pipeline that preserves metadata while enhancing OCR text processing for 983,004 volumes.
Despite the promise of adaptive inference, internal representation statistics fail to provide reliable difficulty signals for multilingual NLI across African languages.
Balancing hate speech detection and user privacy is not just a challenge—it's a necessary trade-off that could redefine online safety standards.
Unlocking 16.3 billion tokens from historical newspapers could revolutionize access to archival data for AI research and applications.
Hybrid models that blend linguistic features with transformer architectures can significantly enhance the evaluation of German NLG, outperforming traditional baselines.
Fine-tuning Whisper models for multilingual medical ASR reveals that the best performance hinges on the adaptation strategy, with surprising shifts in internal representations based on language context.
A neuro-symbolic approach reveals how LLMs can effectively generate implicit premises, transforming the landscape of argument reconstruction.
Fine-tuning outperforms zero-shot inference, but the real game-changer is the use of synthetic data to elevate performance in culturally specific tasks.
Hallucinations in LLMs may blur the line between advanced simulation and true cognitive processes, challenging our understanding of AI consciousness.
A single attention head in the Mistral-7B model captures demographic identity with surprising fidelity, yet its causal use reveals a disconnect that complicates LLMs' ability to simulate real-world populations.
LLM-based query reformulation not only bridges the gap in retrieval performance but also sets a new standard for legal question answering in Greek statutory contexts.
Generative AI doesn't just reflect biases; it embeds a dominant cultural epistemology that marginalizes minority voices at the very foundation of knowledge creation.
SiNMULI achieves 99.89% accuracy in malicious URL detection, outperforming traditional models while being lightweight and interpretable.
AUTOSIGMA achieves superior rule generation by dynamically converting unstructured CTI into context-aware Sigma rules, outperforming traditional methods and LLMs.
Local LLM-based SSH honeypots can achieve superior shell emulation accuracy with the right prompting and fine-tuning strategies, but their effects can conflict in unexpected ways.
Swapping to CTIFoundry allows smaller models to outperform flagship models, achieving higher accuracy with fewer tool calls in cyber threat intelligence investigations.
Gaussian noise isn't enough; a new method reveals that even noise-protected embeddings can leak sensitive information.
Flama unifies API development and LLM services in a single, type-driven framework that streamlines production workflows and enhances developer efficiency.
Achieving up to 2.3x throughput gains, HYDRA reveals that co-designing architecture and runtime policies is essential for optimizing hybrid LLM workloads on chiplet systems.
GateDiffInt reveals how structured intent extraction can significantly boost conversion rates by effectively managing noise in user behavior data.
Whisper's fine-tuning can reduce Mizo ASR error rates to as low as 7.22%, showcasing its potential for low-resource languages.
Reducing feature duplication from 37.2% to 6.8% while maintaining state-of-the-art performance with just 5.4 features showcases a breakthrough in metadata-free AutoFE.
Selective re-scanning in recurrent networks can drastically reduce memory usage while improving task performance, challenging the conventional wisdom of fixed-size state fidelity.
Harnessing the unique signals from Mixture-of-Experts architectures, InnerExpert achieves unprecedented accuracy in detecting hallucinations at the token level.
SRT recovers up to 37% of lost knowledge in language models while enhancing new information retention, challenging the limitations of traditional replay methods.
Randomized splits can inflate model performance by up to 6.5 times, revealing the critical need for temporal leakage audits in financial NLP.
FlightLLM reveals how combining LLMs with structured prompts and statistical classifiers can transform complex flight safety data into interpretable insights.
EvoTS-Agent not only outperforms existing models in financial change-point detection but does so with a flawless execution success rate across multiple datasets.
A small subset of question types drives the majority of student inquiries in AI interactions, revealing a dynamic evolution in questioning as tasks progress.