Search papers, labs, and topics across Lattice.
100 papers published across 6 labs.
Saliency-guided augmentation can drastically improve the robustness of behavior cloning models to visual domain shifts without sacrificing in-domain performance.
A multi-layered defense combining technology, training, and policy can drastically cut unauthorized data access risks.
FedTVD redefines client weighting in federated learning, achieving significant performance gains by balancing data quality and quantity.
Bengali speakers face a staggering 67:1 training-token deficit compared to English, highlighting a critical inequity in AI language support.
LOB-ID reveals that traditional metrics fail to capture the intricate structures of market data, providing a more sensitive evaluation of generative models in finance.
LOB-ID reveals that traditional metrics fail to capture the intricate structures of market data, providing a more sensitive evaluation of generative models in finance.
Uniform Herding not only boosts accuracy but also reduces forgetting, outperforming traditional methods by effectively managing exemplar representation across tasks.
Generated retinal images can inherit critical clinical information, but they may not perform well against real-world classifiers, exposing a crucial representation gap.
A new generative framework can synthesize over 332 million individuals with realistic geographic and demographic attributes, outperforming traditional methods in joint distribution reconstruction.
Trusting the wrong pseudo-labels can lead to confirmation bias, but CW-BASS v2 intelligently adapts to the confidence landscape of strong foundation models, yielding significant performance improvements.
Certified-optimal samplers can achieve prescribed KL error with high probability while dramatically improving sampling efficiency through adaptive scheduling based on data geometry.
OCSVM-based drift detection achieves comparable accuracy to constant retraining while drastically improving training efficiency in malware classification.
Wasserstein Filtering can significantly enhance the robustness of generative models by effectively isolating and removing outliers from contaminated datasets.
ORBIT enables unprecedented control over training distributions, leading to superior zero-shot forecasting performance in time series models.
MK-TGAN outperforms traditional generative models by integrating biological knowledge, producing synthetic transcriptomic data that is both realistic and useful for research.
Complex models may shine in simulations, but they struggle in real-world fall detection, revealing the critical role of representation choice.
Minimum description length outperforms traditional methods in high-dimensional neighbourhood selection, achieving lower false positive rates even when the true model is misspecified.
FlowLOB achieves high-fidelity limit order book simulations with remarkable efficiency, outperforming traditional methods in both realism and controllability.
Synthetic medical data can maintain up to 97.3% of the utility of real data, challenging assumptions about the limitations of synthetic datasets in healthcare applications.
EEG-PRIME achieves zero-shot transfer capabilities in EEG decoding, outperforming traditional models without the need for domain-specific calibration.
Training data influence shifts dramatically over the course of language model pretraining, with literature data dominating early and STEM data taking over later.
A trust-tiered librarian can eliminate 6,845 contradictions in research reports while generating evidence-grounded narratives at unprecedented speed.
Client simulations can now exhibit realistic resistance and emotional depth, transforming how novice counselors are trained and evaluated.
GDI transforms defect classification by generating single-defect samples, leading to a remarkable 63.6% boost in F1-Score for rare defects.
Achieving state-of-the-art results for Danish with a model that uses only permissible data challenges the notion that larger datasets are necessary for competitive performance.
LITTLELEARNER reveals that even a well-defined knowledge scope can yield a competent language model, but it won't expand its capabilities beyond its educational boundaries.
Noise-aware calibration methods can transform unreliable predictions into trustworthy confidence estimates, even in privacy-preserving contexts.
Current unsupervised feature selection evaluations are often misleading, masking supervised influences that compromise their validity.
Real-valued pairwise supervision can dramatically improve clustering performance, as shown by ECI-PP's ability to outperform state-of-the-art methods in uncertain environments.
Higher epiplexity in training data can significantly boost model performance in unseen tasks, revealing a new avenue for data-driven generalization strategies.
Organizing time-series data into granular balls can drastically reduce inference time while boosting resilience against mislabeled samples.
Noise and redundancy in training data can cripple SVM performance, but a novel loss function transforms the geometric twin SVM into a robust feature selector that excels in challenging environments.
Achieving 72.3% accuracy in freshness prediction with just three labeled days per fillet reveals a breakthrough in few-shot learning for food quality assessment.
Current multimodal models excel at basic character recognition but struggle significantly with deeper linguistic and historical analyses, revealing critical gaps in their capabilities.
Text descriptions can replace visual data in video anomaly detection, leading to a dramatic boost in performance without the need for annotated video datasets.
Refining low-confidence pseudo-labels with diffusion techniques boosts histopathology segmentation accuracy by over 6% compared to traditional methods.
Extracting defect-spacing information from borehole archives using weak labels and advanced segmentation techniques yields unprecedented accuracy in crack localization.
The reliance on knowledge bases as gold standards in machine translation may inflate performance metrics, masking the true quality of translations in low-resource settings.
Bengali speakers face a staggering 67:1 training-token deficit compared to English, highlighting a critical inequity in AI language support.
As GenAI automates routine tasks in geovisualization, the real challenge lies in navigating the new landscape of judgment and accountability.
Time-window-based evidence aggregation in Slips leads to a significant boost in detection performance, outperforming traditional methods without increasing false positives.
Side-channel attacks on FGAC can now reconstruct high-entropy data, revealing a critical vulnerability in widely used database security mechanisms.
Confidence-aware pseudo-labeling boosts performance in HD map construction, achieving a +6.1 mAP gain even with scarce labeled data.
Leveraging dialogue as a conditioning signal significantly enhances video-to-music generation, outperforming existing models with a novel dataset that ensures reproducibility.
Saliency-guided augmentation can drastically improve the robustness of behavior cloning models to visual domain shifts without sacrificing in-domain performance.
A unified benchmark that combines synthetic data with real-world testing could revolutionize how we train robots for complex manipulation tasks.
Last-layer embeddings in protein language models often underperform, with shallow layers proving more effective for specific datasets like deep mutational scans.
By integrating structured website exploration with task-trajectory synthesis, SynWeaver enables web agents to achieve unprecedented levels of generalization across diverse websites.
TailBooster not only generates synthetic data for rare extreme events but also ensures that the data adheres to operational constraints, leading to substantial improvements in predictive accuracy.
CalibDCD reveals that post-training shifts can significantly degrade data contamination detection, but targeted calibration can restore accuracy.
Pretraining incoherence, not self-attention deficits, explains why ViTs can outperform CNNs in low-data settings.
Achieving fully-labeled model performance with 70% less annotation effort could revolutionize automated cancer detection in PET/CT imaging.
Existing unlearning methods for MLLMs can lead to significant knowledge locality degradation, while the new SGPE method strikes a competitive balance between forgetting and preserving relevant knowledge.
ConVAWG reveals that synthetic dialogue generation can effectively model the nuanced dynamics of Violence Against Women and Girls, filling a critical gap in the study of relational abuse.
LLMs can be harnessed to transform weakly supervised hierarchical text classification, boosting performance on imbalanced datasets through innovative data augmentation techniques.
Achieving a WER of 23.44% for Burmese medical ASR, this work sets a new benchmark that challenges the capabilities of larger models.
A multi-layered defense combining technology, training, and policy can drastically cut unauthorized data access risks.
Operationally useful explainable IoT intrusion detection hinges on a delicate balance of predictive quality, explanation cost, and stability, not just accuracy.
Introducing skip semantics for TGG rules could revolutionize how overlapping queries are handled in view-based development, enhancing both usability and precision.
A unified framework that not only generates but also interprets piano performances reveals critical insights across all skill levels, challenging traditional assessment methods.
No single perspective is sufficient for art interpretation; MMArt reveals that each perspective uniquely enhances understanding and retrieval of visual art.
Multi-part optical music recognition is revolutionized with the OSSQ-OMR dataset, revealing that LSTM models can outperform Transformers by 2.6 times on scanned scores.
Trust in scientific data can be systematically built through contextual histories rather than just popularity metrics.
Existing provenance systems often compromise on security, with many tools failing to ensure the integrity of captured events.
Models trained on IoTVulBench outperformed existing benchmarks by up to 0.42 MCC, showcasing the critical role of domain-specific training and curriculum design in vulnerability detection.
Weak supervision from loosely aligned broadcast transcripts can yield sign representations that significantly outperform traditional methods in both localization and translation tasks.
Released tokenizer vocabularies can yield precise estimates of hidden corpus compositions, revealing insights that were previously obscured.
Recognition accuracy for isolated handshapes drops significantly when models encounter unseen signers, highlighting a critical gap in current methodologies.
$π$-SUB establishes a new standard for underwater image enhancement, achieving unprecedented realism and generalizability that could redefine benchmark datasets in the field.
Automating data selection with DataMaster not only reduces manual effort but also enhances performance across diverse applications, challenging traditional heuristic methods.
Steering persona features can amplify emergent misalignment rates in language models beyond what traditional fine-tuning achieves.
BPG reduces forgetting to an impressive 0.22% while achieving state-of-the-art accuracy in domain incremental learning.
Machine learning models can reduce path loss prediction errors in LPWAN by over 30% compared to traditional methods, revolutionizing network planning for smart cities.
Misalignment in data augmentation can lead to significant training errors, but AlbumentationsX ensures that images and their annotations are always transformed together, preserving label integrity.
PhysDGM not only generates high-fidelity synthetic time-series data but also boosts predictive performance by up to 48% while slashing data collection costs by an order of magnitude.
No synthetic generation method can fully replace original training data, but Grasynda-P strikes a compelling balance between forecasting accuracy and privacy risk.
RTSKG reveals that integrating urban entities into a cohesive knowledge graph can dramatically improve the accuracy of ridership predictions and related urban analyses.
Workflow Cards nearly double the quality of insights from workflow executions, bridging a critical documentation gap in machine learning practices.
Auditing Chinese web content reveals pervasive pollution that shifts over time, challenging the integrity of LLM training data.
Merging records in a knowledge graph can irreversibly corrupt data, highlighting the urgent need for precise identity management in automated systems.
ReLTEx reduces hallucinations in LLM-generated taxonomies, leading to expansions that are not only more reliable but also semantically coherent.
A novel benchmark synthesis reveals a persistent compositional gap in human action recognition, challenging existing models to rethink their approach to hierarchical reasoning.
A lightweight model that predicts children's age alongside phonemes can outperform larger models, revolutionizing phoneme recognition in children's speech.
A segmentation model trained on synthetic data outperforms one trained on real images, all while safeguarding proprietary designs from potential IP breaches.
Generating realistic fault data through HIL simulation could revolutionize how we validate automotive software systems in real time.
FedTVD redefines client weighting in federated learning, achieving significant performance gains by balancing data quality and quantity.
A novel reliability metric and optimization framework could save cloud services millions while enhancing user experience.
Denoising bioacoustic signals using ridge-guided training synthesis can dramatically enhance the clarity of vocalizations, leading to better classification outcomes in noisy environments.
Achieving perfect precision in retrieval tasks, Guardian Crawler sets a new standard for evidence-grounded summarization in noisy web contexts.
The standard evaluation protocol for generative time-series models can drastically misrepresent model performance, turning top performers into cautionary tales.
Temporal deep learning methods can effectively reconstruct missing tropical cyclone parameters, revealing a surprising resilience in Rmax variability preservation with fewer data samples.
CAN-FLOW generates cardiac anatomies that not only look realistic but also align closely with clinical metadata, outperforming traditional methods.
ICED achieves competitive performance across multiple density estimation tasks without the need for retraining or hyperparameter tuning, revolutionizing the efficiency of tabular data analysis.
Label-flipping attacks on federated GANs can skew generation distributions significantly while remaining undetectable by traditional label-agnostic metrics.
Limiting input variables to just ten homeowner-accessible features can drastically reduce predictive accuracy, revealing the critical role of comprehensive data in energy estimation models.
CPDA aligns class-conditional latent paths, outperforming 30 existing methods in unsupervised time-series domain adaptation.
Explicitly optimizing for target function smoothness can dramatically enhance the performance of models on tabular data.
Original depth images typically outperform synthetic ones in sign language recognition, but surprising instances show the opposite, challenging assumptions about data quality.
TaxoScale reveals novel security and privacy concerns from over 600,000 mobile app reviews, outperforming traditional methods in taxonomy generation.
Deleting structural announcements from text makes subsequent prose significantly harder for models to predict, highlighting the critical role of document arrangement in training efficacy.
Single-corpus evaluations can obscure up to 80 points of macro-F1 variability, revealing the hidden pitfalls of current infant cry analysis methods.
Synthetic datasets generated by Generative AI can match real-world data performance, achieving up to 93% accuracy in encrypted traffic classification.
Achieving a 96.30% attack success rate, DFCS reveals that strategic sample selection based on feature diversity can dramatically enhance the effectiveness of backdoor attacks.
Existing document parsers may score high on benchmarks, but they still falter on real-world tables, with a top parser achieving only 85.03 TEDS.
Synthesizing labeled training images from a few unlabeled photographs can outperform traditional methods, achieving superior detection accuracy in challenging environments.