Search papers, labs, and topics across Lattice.
100 papers published across 2 labs.
Social scientists face a critical disconnect between their data needs and the metadata capabilities of current search systems, limiting research potential.
On-policy distillation is massively data-overfed: a single training prompt recovers most full-dataset performance gains, while just 16 prompts saturate 98.9% of reachable state space to match full-data distillation.
Bypassing retraining, OSR achieves competitive classification performance while significantly improving computational efficiency and privacy preservation.
Naive interpretations of educational outcomes can lead to misleading conclusions, as demonstrated by OBER+'s ability to reveal a 25-point collapse in attainment that actually reflects misaligned outcome definitions.
Natural language descriptions of dataset differences can reveal critical insights into the safety and robustness of autonomous driving systems.
Bypassing retraining, OSR achieves competitive classification performance while significantly improving computational efficiency and privacy preservation.
Naive interpretations of educational outcomes can lead to misleading conclusions, as demonstrated by OBER+'s ability to reveal a 25-point collapse in attainment that actually reflects misaligned outcome definitions.
Natural language descriptions of dataset differences can reveal critical insights into the safety and robustness of autonomous driving systems.
Auxiliary views can enhance LLM learning efficiency, revealing that reallocation of training tokens can lead to better factual recall even when using weaker teacher models.
Dialogue state tracking accuracy in industrial robots skyrockets with the new IRWOZ 2.0 dataset, achieving a BLEU-4 score increase of over 200%.
Catalogue photography can significantly aid in the cold-start problem for carbide burr recognition, but its effectiveness in real-world applications is limited without targeted adjustments.
Synthetic semantic supervision allows small transformers to outperform larger models in code representation tasks, offering a scalable alternative to traditional methods.
Achieving precise concept removal in text-to-video models without sacrificing generation quality could redefine safety standards in generative AI.
The resulting 82M-parameter model, Wayu-Paxa-TTS-Edge, enables on-device Thai TTS without reference audio and achieves the lowest pause-placement error and intra-word pause rates among the three systems.
DropClick can maintain high segmentation performance with minimal user input, saving over 46% of annotation effort while still achieving competitive results.
Achieving mask-free video virtual try-on is now possible with BooM-VVT, which significantly reduces reliance on costly video-level data while enhancing garment realism.
DeepSSIM++ reveals that you can achieve a 46-point boost in memorization detection accuracy without sacrificing computational efficiency in medical generative models.
Text2Thermal synthesizes thermal images directly from text, achieving state-of-the-art results while eliminating the need for RGB image registration.
Localized chord corruption can produce over 2.88 times the output change compared to traditional ACR replay, challenging assumptions about chord recognition in music generation systems.
A revolutionary two-dimensional curation model enables precise web preservation by disentangling topical and geographic dimensions, setting a new standard for national libraries.
On-policy distillation is massively data-overfed: a single training prompt recovers most full-dataset performance gains, while just 16 prompts saturate 98.9% of reachable state space to match full-data distillation.
An innovative embedding technique reveals hidden energy inefficiencies in mobile networks, allowing operators to target their resources effectively.
Causal discovery in federated settings can achieve high accuracy without compromising privacy, even in the presence of noise, by leveraging higher-order cumulants.
WeatherNext 3 achieves state-of-the-art probabilistic medium-range forecasting by integrating real-time satellite data, outperforming traditional models in accuracy and resolution.
Scaling task diversity, not just dataset size, is the key to achieving robust zero-shot generalization in offline multi-agent reinforcement learning.
QAT-FM reduces coupling construction costs while achieving competitive generative performance, transforming how we approach high-dimensional generative tasks.
The first open Armenian LLM, arm-gemma-e4b, not only outperforms all predecessors but also highlights the critical balance between fluency and knowledge retention in low-resource language models.
A staggering 11.8% of existing compaction data is physically impossible, reshaping our understanding of soil compaction testing.
Achieving high-quality model performance with just 10% of the required labels could revolutionize the scalability of RLVR in large language models.
REMIND can identify and correct noisy annotations in infant pose estimation, achieving up to 93% accuracy in real clinical settings.
Achieving top-tier regression performance while drastically reducing computational costs, Xiaomi-TabLDM sets a new benchmark for efficiency in tabular data modeling.
Placement of synthetic data in representation space can be more crucial than sheer volume, leading to significant performance gains in low-resource NLP tasks.
Language models trained on synthetic data may not reflect the true nature of the languages they represent, challenging our assumptions about their utility in linguistic studies.
Real Human-Human dialogue data can transform turn-taking in dialogue systems, achieving better proficiency without sacrificing semantic quality.
Self-regulated learning strategies are key to minimizing digital distractions in online education, revealing a surprising disconnect with peer engagement methods.
ATIBA could revolutionize manuscript submission by automating integrity checks that are often inconsistently applied or overlooked.
Gated communities in urban China are not just spatial constructs; they create profound social inequities that can be mapped and quantified at scale.
Achieving up to 96% accuracy in parking space classification with minimal annotated data could revolutionize intelligent transportation systems.
A groundbreaking pipeline has generated over 8,000 hours of action-conditioned video, overcoming the limitations of real-world data collection.
MudraGen generates culturally rich and anatomically precise two-hand dance gestures, setting a new standard in the preservation of Indian classical dance.
XR-2 shows that scaling demonstration data and incorporating real-time corrections can dramatically enhance bimanual manipulation success rates in household tasks.
Decoupling task synthesis from on-policy rollouts solves the learning-signal saturation bottleneck in terminal agents, boosting long-horizon RL performance by up to 18 percentage points on Terminal-Bench 2.1.
Visual place recognition models often crumble under weather and lighting shifts, but injecting 160K geometrically verified, route-aware synthetic hard positives boosts R@1 retrieval by up to 9.2% across foundation backbones.
Frontier search agent performance does not require complex multi-agent swarms or test-time search verifiers: a single ReAct policy trained via iterative SFT-RL climbing hits 56.4% on Humanity's Last Exam and 92.9% on DeepSearchQA.
Recurring frontier API calls can be replaced by one-minute compile-time distillation, turning plain-text prompts into local, versionable neural functions that achieve 83.6% accuracy on benchmarks where instantaneous compilers fail entirely.
It is found that synthetic-only training can be competitive, and the 0.9B-parameter PaddleOCR-VL-1.6 is adapted into Wayu-Paxa-OCR-Zero, a Thai OCR model adapted without OCR labels from real Thai document pages, showing that synthetic-only training can be competitive.
Incorporating spectral structural priors into uncertain knowledge graph completion can dramatically enhance prediction accuracy and training stability without adding extra parameters.
AI-derived measures can recover critical individual and group-level effects, challenging the traditional reliance on survey data for social and occupational insights.
RINSE achieves state-of-the-art performance in zero-shot graph anomaly detection by effectively estimating target normality without requiring any target labels or gradients.
IFW-BLS achieves superior robustness against noise and outliers, outperforming traditional models by intelligently down-weighting unreliable samples.
RepEmp reveals that the future capacity to model and plan is more critical than immediate fidelity in representation selection, reshaping our understanding of model construction.
Mislabeling in encrypted traffic benchmarks can lead to a staggering 40% drop in classifier accuracy, challenging the validity of existing models.
Bridging the gap between graph topology inference and generative modeling could unlock new avenues for innovation in graph learning.
LLMs systematically underrepresent the diversity of their training data, with a notable conditional diversity gap that can be mitigated through a novel entropy-constrained projection method.
SMart achieves a remarkable 19.5% reduction in regression error by leveraging multi-source datasets and innovative recovery tasks in time series representation learning.
Disease distribution shifts, not skin tone, are the primary culprits behind the poor performance of dermatology AI models in unfamiliar clinical settings.
OBJECTION slashes the False Guilty Rate in legal AI predictions from over 82% to under 17% by challenging prosecutorial bias in real-time.
Task-level natural-language priors can transform low-resource LLM training, boosting performance even with minimal data.
Git4Data achieves up to 10x faster version control for AI agents in relational databases, transforming how we manage data states.
Achieving a staggering 99.50% acceptance rate in synthetic dialogue generation reveals the transformative power of feedback-guided refinement in meeting complex communicative constraints.
Output format can distort perceived model capabilities, with a 40-point accuracy gain in one format vanishing in another.
The shift to differential privacy could redefine how National Statistical Organisations balance data utility and individual privacy in the age of big data.
Traditional palm-vein recognition systems can see their accuracy plummet by 75% under dirty conditions, but a novel matching technique recovers much of that lost robustness.
Generative RGB-to-IR translation can boost infrared vehicle detection performance by over 10 mAP points in unseen UAV domains, challenging the notion that real data is irreplaceable.
Information density imbalance, not just instance count, is a key driver of category bias in object detection models, revealing new avenues for enhancing fairness and accuracy.
FuDU transforms uncertainty into a strategic asset, enabling real-time defect detection with unprecedented reliability in industrial applications.
Generative image editing can boost object detection performance by over 20% in challenging camouflage scenarios, transforming how we approach domain shift problems in AI.
Achieving a 91.1% OCR exact-match rate, this method outperforms larger models while dramatically enhancing the detection of rare traffic signs.
Point-supervised change detection can achieve performance on par with fully supervised methods by effectively refining noisy pseudo-labels through a novel two-stage optimization process.
Realistic anomalies can be generated for unseen products without needing any target-product samples, revolutionizing how we approach anomaly detection in industrial settings.
Forget classes can be effectively recovered in a source-free setting, with some methods even surpassing traditional retraining performance.
Real-world speech recordings with overlapping noise events are now timestamped and categorized, filling a crucial gap in audio datasets.
Social scientists face a critical disconnect between their data needs and the metadata capabilities of current search systems, limiting research potential.
Achieving a 12.31% accuracy boost in open-set WiFi RF fingerprinting, C$^2$T-OpenMax redefines class representation geometry for better device authentication.
Achieving optimal accuracy in answering linear queries with minimal randomness could redefine standards for differential privacy in practical applications.
SHELF reveals that sparse methods can outperform traditional approaches in bibliographic tasks, challenging assumptions about model performance consistency.
LLMs can autonomously inject vulnerabilities into smart contracts, yielding a surprising 16.58% survival rate of confirmed vulnerabilities across diverse types.
Governance metrics in agricultural open-source software reveal surprising independence from actual security risks, challenging assumptions about the sector's vulnerability.
Achieving 0% detectable speech while preserving 85% activity recognition accuracy could revolutionize privacy in acoustic monitoring for elderly care.
Tracing the origins of synthetic media is now possible without altering generator architectures, thanks to a novel self-referential framework that ensures round-trip consistency.
PrivateHub can reduce the accuracy of private application detection by up to 50% without sacrificing the performance of non-private applications.
Combining anatomical priors with active learning not only boosts segmentation accuracy but also reveals the potential of model uncertainty as a guide for dataset expansion.
Internet video can finally solve the physical robot data bottleneck once manipulation behaviors are indexed by actor-centric 3D hand trajectories rather than fragile visual pixels.
Boundary-mutation testing uncovers critical vulnerabilities in secret scanners, revealing that some rules can fail completely under realistic context variations.
Influence-guided response rewriting is introduced, which uses IF to identify intervention targets and replaces their responses with behavior-aligned or behavior-opposed supervision while keeping instructions fixed, motivating intervention-aware evaluation of TDA methods.
Current face recognition models miss key features in twin identification, achieving only 76% accuracy despite the potential for significant improvement through skin marks and asymmetry.
Augmentation reliability hinges more on distributional approximation error than on predictive performance, challenging conventional beliefs about generative methods in imbalanced classification.
Algorithm-dependent learnability reveals that focusing on the optimizer's trajectory can significantly enhance offline optimization performance, leading to state-of-the-art results on complex tasks.
Even with extensive data and advanced algorithms, predicting childbirth remains frustratingly imprecise, with chance playing a larger role than expected.
Privacy in synthetic data is often treated as an implicit assumption rather than an explicit, testable claim, leading to uneven protections for sensitive information.
CATeye reveals that effectively isolating invariant attributes and edges can lead to substantial improvements in detecting evolving fraud patterns in e-commerce.
SAGE boosts worst-group accuracy by up to 7.7 percentage points, tackling the pervasive issue of spurious correlations in machine learning without relying on prior group labels.
Attention-based models can significantly enhance in-table predictions, outperforming traditional architectures when trained on synthetic tabular data.
Hidden preferences can be stealthily transferred during model distillation, but targeted regularization can significantly curb this effect without sacrificing performance.
Pseudo-label quality, not quantity, is the key to successful self-training in domain adaptation, revealing a critical insight for real-world applications like waste sorting.
A searchable catalog of over a thousand AI model findings could revolutionize how researchers access and build upon existing knowledge in the field.
Informative missing labels can significantly enhance classification accuracy by providing insights into uncertainty, leading to reduced expected error rates in semi-supervised settings.
Multilingual models struggle to transfer factual knowledge across languages, with most English-acquired facts failing to make the leap to Persian after targeted data interventions.
Finetuned LLMs can transform unstructured text into structured event logs, vastly improving the efficiency of process mining.
A novel Mixture-of-Experts VLM slashes document processing costs by over 80%, outperforming human annotation and existing models.
Deep graph generative models can replicate real-world network structures and uncover effective strategies for epidemic control, outperforming traditional models.
Athena outperforms existing methods by leveraging graph structures, achieving superior identification of vulnerability-affected libraries with significantly fewer parameters.
Trade restrictions emerge as the dominant risk in semiconductor supply chains, identified through a novel pipeline that combines LLMs and expert validation to score over 76,000 risk items.
Language models may score high on benchmarks, but LLMPEDIA reveals their true factual accuracy is only 68.4%, exposing critical gaps in their encyclopedic knowledge.
Experts prefer citation-grounded disaster narratives that enhance situational awareness, revealing a significant shift in how humanitarian information can be synthesized and utilized.