Search papers, labs, and topics across Lattice.
100 papers published across 4 labs.
An unsupervised framework for identifying and characterizing competing narratives in political discourse on social media, focusing on German politicians'tweets, employs a multi-stage pipeline that integrates natural language processing techniques such as topic modeling, event detection, and event linking.
Results show that structured schema context and lightweight prompting can substantially reduce reliance on fine-tuning for scalable conversational access to KGs, though closing the remaining gap to fully fine-tuned approaches will likely require reducing dependence on curated question-query exemplars.
Non-language data can serve as a partial substitute for language data for the training objective of next-token prediction but does not reliably support broader linguistic generalization, and transfer from non-language data is less efficient than additional language data.
This work develops Generative Marketing Mix Modeling (GMMM) to estimate the causal effects of Generative Engine Optimization and Generative Engine Marketing and investigates the empirical performance of the proposed method using simulated answers to product recommendation in English and Japanese.
The Explainability Assistant is introduced, an open-source conversational XAI system that leverages the function-calling capabilities of modern Large Language Models (LLMs) to overcome limitations and achieves 94% intent-parsing accuracy, supports flexible natural language interaction, and adapts to different ML problem types without task-specific fine-tuning.
Results show that structured schema context and lightweight prompting can substantially reduce reliance on fine-tuning for scalable conversational access to KGs, though closing the remaining gap to fully fine-tuned approaches will likely require reducing dependence on curated question-query exemplars.
Non-language data can serve as a partial substitute for language data for the training objective of next-token prediction but does not reliably support broader linguistic generalization, and transfer from non-language data is less efficient than additional language data.
This work develops Generative Marketing Mix Modeling (GMMM) to estimate the causal effects of Generative Engine Optimization and Generative Engine Marketing and investigates the empirical performance of the proposed method using simulated answers to product recommendation in English and Japanese.
The Explainability Assistant is introduced, an open-source conversational XAI system that leverages the function-calling capabilities of modern Large Language Models (LLMs) to overcome limitations and achieves 94% intent-parsing accuracy, supports flexible natural language interaction, and adapts to different ML problem types without task-specific fine-tuning.
It is found that cache performance depends on transfer granularity, intermediate memory use, and when transfers enter the request schedule, not only on device bandwidth, not only on device bandwidth.
E-CONAN benchmarks that are composed of sentences pairs from various sources, derived from a combination of sources, offers a broader and more robust assessment compared to XNLI and ArNLI, and are used to evaluate state-of-the-art multilingual pretrained models using zero-shot classification.
Where recent studies report that probe-detected errors are resistant to interventions, it is found that in-context binding is a setting in which probes are actionable.
This work analyzes over 790,000 translation outputs from 12 LLMs across 22 language pairs (LPs) and identifies 12 recurring noise patterns, which are group into formatting and content noise, and constructs TransClean, a controlled benchmark of 9,900 pairs of noisy and clean translation outputs.
A unified evaluation framework testing gradient and loss-based balancing strategies under controlled settings, and a theoretical diagnosis explaining why these methods fail, as they conflate fitting speed with discriminative contribution are provided.
A controlled benchmark of membership inference vulnerability for text classification on the GLUE SST-2 sentiment dataset is presented and lightweight training adjustments can improve the privacy-utility trade-off without complex defenses.
Financial NLP has a standard workflow: validate a sentiment tool against human labels, then trust it to extract market signal, but it is found that the relationship between construct and predictive validity depends on the sampling convention and score representation.
A multi-signal detection pipeline combining fine-tuned DeBERTa-v3 classification, Monte Carlo Dropout uncertainty quantification, and temperature-scaled calibration for response-level hallucination detection is presented.
LOCUS, a method that selects a task-aware low-rank adaptation subspace to minimize output-token cost subject to a utility constraint, is presented, a method that selects a task-aware low-rank adaptation subspace to minimize output-token cost subject to a utility constraint.
The proposed Koan-Safe, which combines a prompt-only intent parser, a replaceable generator, and structural repair with default safety parameters, is proposed, which supports explicit trade protections and execution-based evaluation, while distinguishing declared safety from a general guarantee.
A directory-aware semantic data management system that tightly integrates semantic and structural access to support structural-context-efficient, evidence-gap-driven multi-round retrieval, and introduces an adaptive escalation strategy that answers from one-round experience-augmented retrieval when the evidence is sufficient, and invokes agentic multi-round retrieval only otherwise.
The model incorporates three core components: spatiotemporal embedding module, spatiotemporal fusion module, and LLM backbone, which adopts a differentiated parameter adaptation strategy to balance training efficiency and traffic data adaptability.
Tightening top-p, which cuts the low-probability tail at generation time, nearly stops collapse within three generations and stabilizes six checkpoints spanning the whole spectrum together, while data-side filtering slows collapse without stopping it.
A lightweight Text Risk Score (TRS) is proposed, which estimates synthesis risk from interpretable text features without manual annotation or model training and shows a positive correlation with content errors and duration abnormalities, demonstrating its usefulness as a low-cost pre-synthesis indicator for complex-text risk diagnosis in low-resource multilingual TTS.
Direct LLM-based projection enables high-quality multilingual clinical annotation transfer and provides a practical approach for extending clinical NLP resources to languages with fewer annotated datasets and language-specific tools.
Results show that CoSTAR enables small, locally deployable LLMs to achieve performance competitive with substantially larger LLMs for privacy-sensitive COBOL legacy systems.
An unsupervised framework for identifying and characterizing competing narratives in political discourse on social media, focusing on German politicians'tweets, employs a multi-stage pipeline that integrates natural language processing techniques such as topic modeling, event detection, and event linking.
A new completeness estimation approach is introduced that quantifies the degree of sentiment-relevant information preserved in incomplete data to guide the reconstruction of missing semantics and enables more accurate semantic reconstruction, leading to more precise sentiment prediction.
It is shown the path to generalizing L2 reasoning to held-out languages goes through broader language coverage, readily available multilingual non-reasoning data, and a sufficient English reasoning backbone, indicating that reasoning is a language-agnostic behavior that can be transferred across typologically diverse languages through careful data mixing.
Intermediate activations in split LLM fine-tuning trivially leak raw training prompts past standard privacy defenses, but a learned obfuscation pipeline closes this leakage vector without destroying model utility.
Infinite language generation from positive examples succeeds if and only if finite witnesses preserve an infinite common core, unlocking an infinite complexity hierarchy verified entirely in Lean.
Supervised compression can shrink 768-dimensional language embeddings down to just 5 dimensions without accuracy loss, allowing 5-qubit variational circuits to match full-scale 384-dimensional classical baselines.
Automated micro-step prompt refinement can more than double an agent's domain-specific reasoning and scientific rigor scores without fine-tuning underlying model weights.
Compact transformers trained purely on synthetic clinical text can match or exceed real-data de-identification performance while proving more robust to format perturbations across out-of-distribution primary care domains.
Balanced $k$-shot sampling inadvertently breaks classical small-sample discriminant estimators through an exact algebraic degeneracy, yet even after deriving closed-form repairs, a standard cross-validated logistic probe remains superior on LLM embeddings.
Eliminating the language model output projection matrix in favor of geodesic decoding on a Poincaré ball cuts perplexity by over 50% compared to tied Transformers and SSMs at sub-million parameter scale.
Scaling LLM document auditing to large batches causes detection recall to crash from 60% to 2.8%, triggering confident hallucinations of nonexistent errors rather than graceful abstention.
Recycling just four transformer layers through twelve recurrent iterations matches full-depth baselines on core linguistic benchmarks, exposing precisely where compute can—and cannot—substitute for raw parameter capacity.
Span-level supervision fixes the token-representation mismatch in cross-lingual sentence encoders without sacrificing—and in fact improving—sentence retrieval performance.
Full rationale supervision burns up to 254× more tokens without guaranteeing reliability, whereas isolating representations with high rationale-boundary sensitivity selectively immunizes medical LLMs against choice-order brittleness.
Grounding LLM candidate reranking in explicit ontology hierarchies closes a stubborn 5% accuracy gap in fine-grained biomedical concept normalization.
Open-weight embedding models fine-tuned via contrastive ranking beat frontier LLMs like GPT-5.2 and Claude-Sonnet-4.6 by over 10 percentage points on real-world entity disambiguation.
Pre-filtering text streams with emotion-aware semantic screening slashes the inference overhead of transformer-based toxicity classifiers without degrading cyberbullying detection sensitivity.
Even state-of-the-art long-context models and memory architectures collapse when forced to synthesize evolving user preferences into actionable advice rather than simply regurgitating past conversational facts.
Gating multi-teacher distillation on candidate correctness induces catastrophic label collapse and zero minority-class recall while failing to outperform simple hard filtering or improve factual grounding.
Romanized and regional dialects represent a massive evaluation blind spot for frontier LLMs—even across globally dominant low-resource languages like Bangla.
Foundation models already encode whether they are right or wrong: a lightweight attention layer over frozen internal token states reliably predicts classification errors across both text and multimodal domains.
Outrage might dominate political media, but it hinders lawmaking: high speech emotional intensity and anger correlate with lower legislative effectiveness, whereas enthusiasm, pride, and emotional diversity predict institutional success.
Fluent surface text masks deep grammatical fragility: across 600K test cases, leading LLMs consistently break down when forced to execute fine-grained Arabic morphosyntactic control, particularly under cliticization and rare inflections.
When privacy rules forbid raw audio access, general LLMs falter on sparse ASR transcription mistakes—making two-stage, detector-gated span correction the far more effective paradigm.
Jointly generating inline argument and entity tags beats sequential extraction pipelines by over 40% relative F1, revealing that argumentative structure cannot be accurately decoupled from the specific real-world entities being debated.
Weakly-supervised LLM annotations can bridge the phonemic data gap in low-resource languages, slashing Filipino sentence-level G2P error rates from nearly 20% down to 0.54% while successfully disambiguating prosodic homographs.
Discarding "easy" examples where the base model already follows glossary constraints yields an 11-point accuracy leap at fixed data volume, outperforming RL-based alignment methods with standard supervised fine-tuning alone.
Instead of stuffing raw context with dialogue history or brittle retrieval vectors, treating user personalization as particle-filtered hypothesis tracing over natural language resolves the tension between fleeting intent and long-term preferences.
Omitted dates in news articles quietly undermine temporal RAG retrieval, yet a deterministic, rule-based normalizer matches LLM-grade disambiguation at zero inference overhead.
Retrieval-augmented prompting fails to beat hand-tuned static exemplars for legal moderation, and even top LLMs miss up to 57% of criminal hate speech despite aggressive over-policing.
Hyper-parallel decoding on a distilled 4B LLM matches frontier model information extraction accuracy while operating at an 8% compute footprint.
LLMs can parse complex clinical and security logs using just 1% to 2% of the original context length without sacrificing predictive performance or explanation faithfulness.
The sparse autoencoder features optimal for document-topic mixture estimation actually undermine topic readability, revealing that monosemantic topic modeling requires decoupling inference from semantic labeling.
Standard taxonomy induction systematically corrupts AI research lineages by forcing transitional breakthroughs into leaf nodes—a failure mode resolved by enforcing monotonic temporal constraints over hierarchical citation graphs.
Standard multimodal fusion treats modality weights uniformly across emotional states, but coupling facial geometry with class-adaptive gating and valence-arousal transition priors drives up to a 4.36 F1 gain on conversational emotion tracking.
NCP-ArchPreview is introduced, a latent-space language model that pushes autoregressive pretraining beyond standard next-token prediction (NTP) and learns through Next Concept Prediction (NCP) to predict discrete concepts that span multiple tokens, introducing an explicit and more challenging concept-level objective while preserving standard token-level autoregressive generation.
This work introduces Brain2Semantics2Text, a method that reconstructs text through an intermediate semantic embedding space and describes the core principles of the approach, its implementation, and the strategies used to mitigate the challenges of learning a reliable neural-to-semantic mapping.
In-context learning is not limited to discriminative prediction: frozen transformers run full iterative diffusion and energy-based sampling within a single forward pass.
By mapping context and generation into knowledge graphs, S3KG untangles whether an LLM is genuinely reasoning over provided context or merely parroting memorized pretraining associations.
Weight-space model editing no longer requires fine-tuning, as forward-pass activation vectors can be analytically converted into composable rank-one weight edits for precise behavioral control.
Human reviewers completely stopped rewarding flowery academic prose once LLMs made it cheap to fake, but LLM evaluators trained on historical data still pay a premium for it—revealing a hidden vulnerability in static LLM-as-a-judge pipelines.
Scalar rewards in on-policy distillation tell tokens to gain or lose probability without specifying where that mass should actually go—framing distillation as explicit pairwise probability transport eliminates this background leakage and beats reverse-KL across reasoning tasks.
Fact-Ablated Evaluation (FAE) is introduced, a new evaluation framework that iteratively ablates the cited evidence to assess whether LLMs revise their predictions accordingly and highlights that strong fact-checking performance can still coexist with weak evidence dependency, while REAL encourages veracity predictions to remain more closely tied to the availability of supporting evidence.
Open-source LLMs swing wildly from perfect fidelity to total failure (0.00 to 1.00 provenance coverage) when transcribing structured financial evidence, exposing the necessity of deterministic, shape-constrained knowledge graphs for auditable domain reasoning.
Cheap segment-then-caption video pipelines can overcome boundary drift and context fragmentation once temporal shifts and historical context are governed by interventional dependency modeling rather than raw sequential attention.
Prompting LLMs with historical period tags steers grammar toward century-scale approximations in GPT and Qwen while failing entirely in Llama, revealing that models internalize historical style as coarse temporal heuristics rather than distinct period representations.
Semantic drift over eight centuries is not a uniform, corpus-wide evolution, but is driven by a sparse cluster of high-impact contextual outliers that contextualized LLMs can isolate.
A systematic ablation across five runs spanning two encoder variants, four stopping strategies, and three gate values is reported, along with negative results from GRPO policy training, BDI-II post filtering, MentalLongformer encoding, and DeBERTa ensembling.
Standard translation metrics and 235B-parameter LLM judges actively reward the worst social media translations due to cultural blindness, but progressively decaying masked cultural scaffolding enables an 8B model to rival Gemini-3.1-Pro.
Standard semantic triples collapse all negative knowledge into a single binary absence, blinding structured representations to the critical distinctions between contradictory, opposite, and intermediary states required for counterfactual reasoning.
Arbitrarily splitting identical retrieved records into separate chunks swings an LLM's evidence weighting by up to 32 percentage points without altering a single word of context.
While long-tail knowledge is notoriously hard for LLMs to retain, structurally popular facts suffer the worst collateral damage during knowledge updates and act as super-spreaders of downstream hallucinations.
Strict glossary constraints artificially inflate translation consistency scores while actively destroying the natural lexical variation that human translators preserve.
Structural alignment—not affective warmth or semantic relevance—is the primary conversational feedback mechanism driving grammatical acquisition in sample-constrained language models.
Collapsing heterogeneous modalities like tables and images into unified text tokens outperforms specialized modular extractors across multimodal QA benchmarks, despite inevitable information loss during serialization.
Monolithic RAG pipelines drastically underperform on multilingual financial QA compared to a bifurcated architecture that routes numeric queries to filing extraction and synthesis queries to rule-based news filtering.
Steering vectors for distinct model behaviors remain naturally orthogonal in the residual stream, enabling training-free, multi-attribute control over language, safety, and style simply by stratifying injection across model depth.
Institutional editing sets a hard ceiling on text forensics: simple character-level stylometry reliably unmasks ghostwriters in tweets and legal filings across languages, but breaks down completely when speechwriters homogenize their styles around a shared corporate persona.