Search papers, labs, and topics across Lattice.
100 papers published across 5 labs.
A single architecture can achieve state-of-the-art results in pelvic imaging across multiple modalities, addressing critical data integrity issues in women's health research.
Existing methods falter on a new benchmark that captures the full complexity of basketball game events, revealing critical gaps in current visual understanding approaches.
SATS achieves a remarkable 65.6% increase in model efficiency while improving prediction accuracy in time series analysis.
Optimizing feature engineering can double model accuracy and dramatically reduce error rates in ocean colour machine learning, but a one-size-fits-all approach won't work.
Training on geographically isolated samples can cut costs by over 140x while boosting model performance across multiple Earth observation tasks.
A single architecture can achieve state-of-the-art results in pelvic imaging across multiple modalities, addressing critical data integrity issues in women's health research.
Existing methods falter on a new benchmark that captures the full complexity of basketball game events, revealing critical gaps in current visual understanding approaches.
SATS achieves a remarkable 65.6% increase in model efficiency while improving prediction accuracy in time series analysis.
Optimizing feature engineering can double model accuracy and dramatically reduce error rates in ocean colour machine learning, but a one-size-fits-all approach won't work.
Training on geographically isolated samples can cut costs by over 140x while boosting model performance across multiple Earth observation tasks.
SAGE-XGBoost achieves unprecedented accuracy in hazard mapping by leveraging spatially augmented graph embeddings, outperforming traditional models by over 33 percentage points.
Automated ENC change classification achieves up to 94% accuracy, transforming maritime safety assessments and operational efficiency.
Reliable detection of malicious agent skills hinges on a comprehensive benchmark that reveals stark differences in threat composition across sources.
Semantic structured partitioning in SCPaT leads to enhanced forecasting accuracy by intelligently modeling interactions among heterogeneous temporal patterns.
LFU emerges as the clear champion among eviction policies, with alternatives failing to deliver meaningful improvements in cache performance.
By pooling data across groups, this framework achieves faster convergence rates in nonparametric regression, challenging traditional assumptions about group-specific learning.
FedCurv-DR not only reduces knowledge forgetting but also ensures fairness and energy efficiency in AI models applied to evolving cultural heritage data.
MS-WDRO outperforms traditional methods by leveraging the Wasserstein metric to fuse heterogeneous data sources, achieving superior graph recovery even in sample-scarce environments.
LLMs can detect 80% of borderline authentication anomalies, far surpassing traditional methods that struggle with such cases.
Anomaly detection can now adapt in real-time to evolving patterns without retraining, achieving state-of-the-art results on diverse datasets.
Anomaly scoring can make or break the effectiveness of flow-matching methods, with trajectory-based scores outperforming traditional approaches in contaminated datasets.
MidTool reveals that dedicated mid-training can significantly enhance LLMs' ability to utilize tools effectively, outperforming traditional post-training approaches.
A closed-form conditional-Gaussian estimator outperforms deep learning approaches in reconstructing complex cardiac shapes, achieving unprecedented accuracy in shape completion.
Treating specification changes as the primary unit of change can drastically reduce deployment times and defects in data platforms.
SciDSK transforms how AI agents interact with scientific datasets, enabling more effective discovery and interpretation through a structured, reusable skill representation.
OenoBench reveals that even leading LLMs struggle with knowledge retention, achieving only 53%-84% accuracy on wine-related questions despite extensive training.
Jointly leveraging selection-based and reinforcement-learning-based alignment can transform medical image captioning into a more trustworthy diagnostic tool.
Static profiles lead to identity essentialism in LLMs, but a new longitudinal memory framework reveals a path to richer, more diverse social simulations.
TGL-APT achieves over 95% F1-score while cutting training time and memory usage by nearly 40%, revolutionizing APT investigation efficiency.
Transferable presentation attack representations can be learned without relying on facial content, achieving high performance across standard benchmarks.
Usability testing reveals that DesCartes Builder empowers domain experts to create real-time digital twins with unprecedented ease and reliability.
A single training example can boost model performance, but its impact fades significantly within just a few training steps.
MDTIM not only separates missing from observed values but also directly predicts original time series values, leading to superior imputation performance.
Learning RGGs in probabilistic metric spaces reveals that edges can exist with a defined probability, transforming how we understand graph connectivity in complex datasets.
Transporting causal effects across different network populations can significantly enhance intervention strategies in social networks and public health.
Colorist outperforms complex generative models by safely generating clinically relevant domain variations while preserving anatomical structures.
Co-observation in training data can dramatically enhance generalization in continual learning, revealing a new dimension beyond forgetting and plasticity.
CutMix boosts the reliability of semantic segmentation models under distribution shifts, even if it doesn't significantly improve accuracy.
SVM and Simple Cart emerged as top performers in heart disease prediction, revealing the critical importance of model choice in clinical settings.
Directly improving reliability in randomized data releases, ProxyGuard boosts statistical power dramatically while ensuring valid inference.
In low-budget Federated Active Learning, homogeneous data demands more coordination than heterogeneous data, flipping the conventional wisdom on its head.
Treating missing data as a contextual signal, MARCUS dramatically improves rent prediction accuracy, slashing MAE by over 51% in some cases.
PU-HNO outperforms traditional methods by learning stable propagation structures from noisy labels, revolutionizing indoor radio map generation.
Tailoring data augmentation to individual learning states boosts performance by an average of 4.5% on natural images, revealing a new frontier in generative data strategies.
Achieving competitive medical image segmentation with just 10 annotated cases could revolutionize the efficiency of clinical workflows.
Achieving a mean F1 score of 0.911 in Tangut word segmentation reveals the power of combining traditional lexicons with modern machine learning techniques in resource-scarce scenarios.
Unlocking 22.6 million visual elements from historical books could revolutionize how we engage with digitized library collections and their applications in AI and research.
CDGP achieves state-of-the-art anomaly localization without any human pixel annotations, revolutionizing the efficiency of industrial visual inspections.
A novel dataset and model that significantly improve JSL sign recognition by addressing confusable signs, crucial for effective communication between deaf children and their hearing parents.
VLMs reveal hidden temporal insights in art history, but their biases could mislead interpretations of historical timelines. WHY_IT MATTERS: This research challenges the reliability of pretrained models in historical contexts, highlighting the need for critical evaluation of data representation in AI systems.
Fine-tuned small language models can outperform massive counterparts by leveraging high-quality filtered data, challenging the notion that bigger is always better in machine translation.
Failure-mode contextual bandits can boost model accuracy by over 4% on standard benchmarks while eliminating the need for additional human annotation.
Multi-view reasoning can be pretrained and reused, enabling competitive performance without the need for constant retraining on downstream tasks.
Incremental learning can render retraining policies nearly irrelevant, with no significant accuracy gains over a no-retrain baseline in many scenarios.
LLM-Detector achieves superior anomaly detection in tabular data without the need for fine-tuning, making it a game-changer for real-world applications.
EventTime outperforms traditional forecasting models by effectively quantifying the financial impact of cybersecurity breaches, revealing a new frontier in event-driven market analysis.
Tailor your text analysis with a revolutionary pipeline that preserves metadata while enhancing OCR text processing for 983,004 volumes.
Unlocking 16.3 billion tokens from historical newspapers could revolutionize access to archival data for AI research and applications.
Fine-tuning outperforms zero-shot inference, but the real game-changer is the use of synthetic data to elevate performance in culturally specific tasks.
Active learning for deterministic register automata is now polynomial-time solvable across both dense and non-dense ordered domains, unlocking new possibilities in automata theory.
Six incorrect references and three legal misqualifications highlight the critical need for precision in regulatory cross-references within the EU AI Act.
Generative AI doesn't just reflect biases; it embeds a dominant cultural epistemology that marginalizes minority voices at the very foundation of knowledge creation.
SiNMULI achieves 99.89% accuracy in malicious URL detection, outperforming traditional models while being lightweight and interpretable.
AUTOSIGMA achieves superior rule generation by dynamically converting unstructured CTI into context-aware Sigma rules, outperforming traditional methods and LLMs.
Adding noise to shared anchor representations instead of private data significantly enhances learning accuracy while maintaining privacy in collaborative settings.
IriSig-Spoof reveals that achieving high accuracy in satellite RFF can mask significant vulnerabilities in spoofing detection, challenging assumptions about model reliability in real-world scenarios.
Swapping to CTIFoundry allows smaller models to outperform flagship models, achieving higher accuracy with fewer tool calls in cyber threat intelligence investigations.
Reusing existing data assets in federated data-sharing pipelines can drastically cut down on complexity and redundancy, making scalability feasible.
NeuroAssertion doubles the number of assertions and mutation coverage, transforming hardware verification by ensuring critical design behaviors are not overlooked.
Aray synthesizes benign artifacts for YARA validation with a 97.6% success rate, eliminating the need for actual malware samples.
Self-supervised pre-training on just one real table can yield surprisingly strong generalization in Tabular Foundation Models, challenging conventional wisdom about dataset size.
Irregular time series forecasting can be revolutionized with DNBNet, which eliminates bias and adapts to diverse temporal patterns for superior predictive performance.
Delta2Gamma achieves a striking 92.4% accuracy in Alzheimer's detection using low-cost EEG, surpassing traditional supervised methods and dedicated EEG techniques.
Task-specific pseudo-label refinement and synthetic augmentation can dramatically enhance segmentation performance even with minimal labeled data.
Fine-tuning civilian models on military-relevant datasets can significantly boost target detection performance, but small object detection remains a critical hurdle.
A novel capability-driven data infrastructure enables multimodal models to achieve unprecedented versatility and transferability in image generation tasks.
Automated skill assessment in cataract surgery can achieve 87% accuracy while providing explainable metrics that align closely with expert evaluations.
Political ideologies on social media can be tracked and predicted with unprecedented accuracy, revealing the complexities of online polarization.
An adaptive ensemble-size rule can cut memory usage by 37 MB and fit time by 0.4 seconds while maintaining classification accuracy in time series analysis.
QPID reveals that targeting association instability in track states can significantly enhance the efficiency of active learning in multi-object tracking.
Generative supervision can transform the landscape of moiré removal, yielding substantial performance gains even in the most challenging real-world scenarios.
Forgetting a class can be done with up to 99% accuracy while preserving the performance of other classes, revealing a new frontier in machine unlearning.
CoinVE-200K enables video editing models to understand and execute complex, multi-faceted editing instructions with unprecedented accuracy and quality.
Valid inference from AI-generated data is possible without gold-standard labels, challenging the conventional reliance on costly benchmarks.
Realistic attack synthesis using tabular diffusion models can boost DDoS detection accuracy in 5G systems from catastrophic failures to perfect scores.
MemCatalyst reveals that targeted data poisoning can drastically boost membership inference accuracy in Vision-Language Models with minimal resource expenditure.
Automated security patch backporting tools show stark performance drops in real-world scenarios, revealing hidden challenges that could reshape tool development.
The integration of data engineering and software engineering practices could fundamentally redefine how we approach the software lifecycle in AI systems.
DAS achieves a remarkable average score of 4.34 in academic survey automation, surpassing its closest competitor by a significant margin.
Automated data curation and imbalance-aware training strategies significantly enhance LALMs' performance on culturally diverse folk music, yet deep musical understanding remains elusive.
Synthetic data can dramatically enhance drone detection in thermal imagery, but real data is still essential for bridging performance gaps.
Fine-tuning on poetry boosts idiom comprehension in LLMs, but cultural fine-tuning paradoxically undermines proverb interpretation accuracy.
A mere 3% of documents contain half of the text in web PDF corpora, revealing a stark imbalance that challenges how we assess corpus size and utility.
Regular medical check-ups and mental health stress are critical predictors of chronic kidney disease, revealing new avenues for early intervention.
Synthetic weather variations can drastically enhance drone detection accuracy in real-world conditions, reducing false alarms and missed detections.
ITNTs unlock the potential of tensor networks for general-purpose data science and large-scale optimization, transforming how we approach nonlinear operations on massive datasets.
TabPFN-Rel not only tops the leaderboard on RelArena-$\alpha$ but also challenges the notion that specialized architectures are always superior to flattened relational databases.
Analytical prior information can cut prediction errors by over 89% compared to direct learning methods when simulation data is limited.
Time-aware validation reveals that traditional model evaluation methods can significantly misrepresent the performance of fuel consumption predictions, leading to misguided operational decisions.
Learning-to-UnLearn achieves retraining-level accuracy with a fraction of the computational cost by automating the unlearning process.
Achieving accurate state of health estimation with just 1% of labeled data reveals a breakthrough in leveraging unlabeled data for battery health monitoring.
Density-reweighted EOT can recover geometrically faithful correspondences even when datasets have drastically different sampling densities, challenging traditional assumptions in dataset alignment.
Adaptive splitting can boost ensemble diversity and performance, addressing a critical gap in existing decision tree models for data streams.
Outlier-robust GPR can achieve superior prediction accuracy while handling contamination, challenging the limitations of traditional Gaussian likelihood models.
FETERS achieves state-of-the-art early time-series classification with just five labeled examples, outperforming existing methods on 44 datasets.