Search papers, labs, and topics across Lattice.
100 papers published across 3 labs.
MnemoDyn outperforms state-of-the-art transformer models in reconstructing resting-state fMRI data, revealing new insights into brain dynamics.
PPE achieves a staggering 83.3% recall in predicting Ebola outbreak zones, outperforming traditional models by over 10 percentage points.
Efficiently estimating high information projections could revolutionize how we approach dimensionality reduction in complex datasets.
High-performing cough-based TB models fail to generalize across datasets, revealing that data collection artifacts overshadow disease-related signals.
Transforming scientific papers into multi-turn generation trajectories not only doubles the training data but also boosts academic writing benchmarks while maintaining reasoning skills.
PPE achieves a staggering 83.3% recall in predicting Ebola outbreak zones, outperforming traditional models by over 10 percentage points.
Efficiently estimating high information projections could revolutionize how we approach dimensionality reduction in complex datasets.
High-performing cough-based TB models fail to generalize across datasets, revealing that data collection artifacts overshadow disease-related signals.
Transforming scientific papers into multi-turn generation trajectories not only doubles the training data but also boosts academic writing benchmarks while maintaining reasoning skills.
EXAONE Tabular outperforms tuned ensembles and larger models, achieving top rankings in multiple benchmarks while drastically reducing inference costs.
Unsupervised diffusion pretraining can elevate medical image segmentation performance, achieving up to a 74% improvement in boundary precision without relying on extensive labeled datasets.
Achieving over 98% accuracy with a compact model, CropCop sets a new standard for plant-health recognition while ensuring data integrity through rigorous auditing.
Data citation for large language models is not just a verification issue; it’s a multifaceted challenge that could redefine how we acknowledge and trace the origins of AI-generated information.
Multi-factor difficulty estimation boosts segmentation performance, achieving state-of-the-art results across multiple architectures.
MoganBert-TR outperforms traditional masked language models by up to 3.7x in Turkish retrieval tasks, redefining benchmarks for Turkish NLP.
Unmatched beliefs in Theory-of-Mind tracking are often valid, and mislabeling them can lead to significant errors in model selection and calibration.
PolyMemDB resolves long-term factual conflicts in AI memory, drastically reducing hallucinations and enhancing user personalization.
Fine-grained dataset distillation can achieve superior performance by focusing on localized evidence rather than just global statistics.
VISA's self-evolving framework not only enhances multimodal instruction synthesis but also adapts in real-time to improve training data quality and model performance.
Active learning can dramatically reduce the annotation burden in summarization tasks, with LOBSTER achieving up to 665x faster query selection without sacrificing performance.
LLMs can extract critical information from OOXML files that is invisible in Microsoft Office, revealing a hidden layer of semantic divergence that could compromise decision-making in financial and compliance contexts.
A staggering 304-day gap in critical trading data reveals the hidden pitfalls of relying on public cryptocurrency archives for timely decision-making.
A groundbreaking dataset and framework that significantly enhance ulcer tissue segmentation accuracy, even with limited labeled data.
RefVideo-6M revolutionizes video editing datasets by providing 5 million high-quality editing samples that prioritize visual references over flawed automatic edits.
Achieving zero-shot action localization with minimal annotation, this method effectively estimates unseen actions by innovatively leveraging weakly-supervised pretraining.
SE-CLIP outperforms traditional semi-supervised methods by effectively mining high-confidence samples, transforming how we adapt vision-language models to satellite imagery.
Automating LiDAR annotation can cut down the labor-intensive process of data labeling while achieving near-human accuracy in segmentation tasks.
Incorporating noisy, in-the-wild imagery into CVL models significantly boosts their performance on clean datasets, challenging traditional reliance on high-end sensors.
A unified model for traffic behavior reduces prediction errors by over 36% across multiple intersections, challenging the need for separate local models.
Unlabelled GNSS data can dramatically enhance PVT accuracy in urban environments, achieving significant improvements even under harsh conditions.
Human-written album reviews can dramatically enhance music retrieval models, yielding substantial gains in performance on complex queries.
CloSeR achieves state-of-the-art GCD performance by elegantly decoupling closed-set recognition from open-set discovery, minimizing objective conflicts and enhancing semantic coherence.
Benchmark rankings for multilingual embedding models are severely compromised by dataset scarcity, with many relying on a single source, which could mislead evaluations.
Models trained on LAION-BVD achieve state-of-the-art performance in multimodal tasks, showcasing the dataset's potential to redefine video understanding.
Current power outage prediction models may be overhyped, as they often fail to generalize beyond inflated performance metrics derived from flawed evaluation methods.
INCEPT outperforms existing EEG models by prioritizing representation stability over signal reconstruction, achieving top rankings in 26 of 30 evaluation metrics across diverse tasks.
SSL for tabular data can outperform traditional methods, but its effectiveness is highly variable and context-sensitive, particularly in the presence of missing values.
Mechanistic control over data generation reveals hidden model dynamics, leading to more diverse and effective datasets that enhance downstream performance.
MnemoDyn outperforms state-of-the-art transformer models in reconstructing resting-state fMRI data, revealing new insights into brain dynamics.
COCI transforms unstructured Calls for Papers into structured metadata, bridging the gap between informal scholarly communication and formal knowledge systems.
Shifting to structural causal modeling reveals how targeted interventions can be explicitly evaluated, unlocking new insights into student learning dynamics.
Achieving a 57% improvement in image quality for synthetic CMR generation reveals the untapped potential of metadata-aware conditioning in medical imaging.
BALIGN filters out high-risk preference samples, preserving foundational model capabilities while optimizing alignment, achieving the best of both worlds.
Traditional LLMs may excel at semantics, but they miss the mark on capturing the unique workflows of individual culinary creators, revealing a critical gap in procedural generation.
Knowledge-guided change synthesis can dramatically enhance the realism and utility of synthetic data for remote sensing applications.
Hard admissibility constraints can shrink candidate spaces by 480x without sacrificing accuracy, revolutionizing how we approach enterprise data mapping.
Digital job language is not just a trend; it reveals deep disparities in how different sectors are adapting to digitalization, with managerial roles leading the charge.
Persian isn't just low-resource; it's annotation-scarce, revealing a complex landscape of NLP resource availability that challenges conventional assumptions.
Models can achieve near-perfect accuracy in resolving conflicting cues, yet their internal mechanisms can differ dramatically, challenging our understanding of model interpretability.
ROBE achieves a remarkable 0.10 increase in F1 score for long-tail event classes, proving that tailored expert classifiers can outperform conventional models in niche historical contexts.
Refined annotations can boost PE segmentation performance more than changes in model training, challenging assumptions about model superiority.
Achieving near state-of-the-art CXR detection with under 10% of the usual annotations could revolutionize medical imaging workflows.
A novel training pipeline that decouples object geometry from semantics and enhances robustness, achieving state-of-the-art performance in cross-city object detection.
Transforming raw gameplay footage into high-quality training data by effectively removing user interfaces could revolutionize how world models are trained.
Synthetic question generation from knowledge graphs boosts retrieval precision and reasoning performance, even in the absence of labeled data.
Achieving a 75× increase in annotation throughput, RefLAM transforms the scalability of historical Arabic manuscript digitization.
Generating over 203,000 unique web interaction trajectories, BrowserForge significantly boosts model performance on real-world tasks by leveraging the vastness of the open web.
A federated model could revolutionize how medical centers share and enhance device knowledge, ensuring local control while fostering collaborative improvement.
Inconsistencies in cybersecurity research are often driven by flawed evaluation designs rather than the technologies being tested.
The overwhelming majority of ICS cybersecurity datasets ignore critical early-stage threats, limiting the effectiveness of intrusion detection research.
HRV Studio achieves near-perfect agreement with leading HRV analysis tools, revolutionizing reproducibility in cardiovascular research.
UPT can either refine model capabilities or amplify errors, depending on the internal signals used during adaptation.
SPECMINE reveals the intricate relationship between AI-generated specifications and code implementation, providing a treasure trove of data for understanding Spec-Driven Development.
Automating the transformation of cultural heritage records into a fully resolvable knowledge graph boosts metadata connectivity by 86%, unlocking previously inaccessible data.
Mapping unlabeled data into a predictive probability space with entropy weighting can dramatically improve active learning efficiency, surpassing traditional methods.
Adapting ECG classification models at inference time can yield a significant performance boost, even in the presence of noisy data and domain shifts.
Generating high-quality synthetic samples for minority classes can dramatically enhance classifier performance in imbalanced time-series tasks.
Light data augmentation and progressive backbone unfreezing can significantly boost the accuracy of bee detection systems, achieving over 97% precision.
Counterfactual annotations reveal how specific interaction failures can lead to drastically different surgical team outcomes, opening new avenues for performance improvement.
Compact and reliable prediction sets can be achieved in healthcare AI even with limited labeled data, thanks to a novel integration of conformal risk minimization and optimal transport.
Non-English Wikipedia entries can dramatically enrich English biographies, revealing overlooked narratives, especially for women from diverse backgrounds.
TCN-AE outperforms traditional TSFMs in anomaly detection while being more resource-efficient, challenging the assumption that larger models are always better.
Even with an impressive $R^2$ of 0.75, EO-ML methods can mislead policymakers due to inherent uncertainties, underscoring the need for robust uncertainty quantification in poverty mapping.
FedCC enables clients to handle ambiguity in data, leading to a remarkable 67.3% accuracy even when facing severe label distribution skews.
Pruning data with MCL not only boosts efficiency but also reveals hidden semantic structures that traditional methods overlook.
Historical query profiles can boost Text-to-SQL performance by up to 25%, far surpassing traditional context retrieval methods.
AI-driven measurement could redefine empirical research by shifting the focus from finding measures to selecting among diverse, potentially conflicting options.
PatchWrite achieves a flawless preservation of manuscript integrity, maintaining 100% accuracy in edits while traditional methods fail completely.
Conceptual framing of bias definitions can lead to significant discrepancies in annotation, impacting both human and LLM assessments.
Detoxifying Arabic text is not just about removing harmful content—it's about preserving meaning and dialectal nuance, and AraDetox shows how to do it effectively.
PQC evolution reshapes encrypted traffic in ways that undermine the reliability of existing classifiers, revealing a critical vulnerability in current evaluation practices.
Memory contamination can mislead performance metrics in anomaly detection, revealing that less contaminated memories don't always yield better results.
Systematic reuse of models in Digital Twin ecosystems can now be achieved with a framework that clarifies compatibility and integration challenges, improving interoperability across diverse systems.
CLEANCON achieves near-zero memory contamination while paradoxically showing that less contamination doesn't guarantee better anomaly detection performance.
Bridging fisheye and standard perspectives, MIVIFI enables robust multi-view image generation that overcomes data scarcity challenges in autonomous vehicle training.
Tokenization can unlock sustained performance gains in user representation learning, even as raw data scaling hits diminishing returns.
Fine-tuning small LLMs on industrial datasets can boost accuracy by over 13 percentage points, revealing a stark contrast in performance between open and closed model-generated data.
Evolving data snapshots can boost LLM performance by over 20% in both intrinsic and extrinsic evaluations, reshaping how we approach human-centered AI alignment.
Weird generalization is not just a quirk of model training; it’s a fragile phenomenon that could be weaponized if not carefully managed.
Identical item wording can yield vastly different psychometric outcomes, revealing hidden instabilities in AI-assisted item development.
Targeted geometric regularization can significantly enhance the performance of LLMs on low-resource languages, addressing a critical gap in language representation.
LLMs misrepresent intersectional identities by reducing complex opinions to single features, undermining their utility as synthetic survey respondents.
DDKD outperforms larger fine-tuned models in cross-domain data-to-text generation, revealing that size isn't everything in model performance.
RAD achieves state-of-the-art anomaly detection by preserving relational context and incorporating symbolic rules, outperforming traditional methods that flatten data.
Temporal portability reveals that user feature stability on Twitter diminishes over time, challenging assumptions about longitudinal data reuse.
Tighter bounds on instance encoding invertibility reveal that deterministic encoders can be just as secure as their randomized counterparts, transforming our understanding of data privacy techniques.
Phishing detection just got a major upgrade with a dataset that captures the raw evidence behind 67,502 web scans, revealing critical differences between phishing and benign sites.
Adapting to new malware threats with just a few examples is now feasible without sacrificing previously learned knowledge.
Model robustness against training-time data contamination varies dramatically, challenging the assumption that clean data performance predicts real-world reliability.
LLMCrater transforms the way we generate FAIR metadata by continuously enriching it throughout the research lifecycle, not just at publication time.
Hybrid panels could revolutionize survey research by leveraging AI to improve participant engagement and data quality in real-time.
Cross-source generalization in mass-shooting risk classification collapses due to feature deficiencies, revealing that model choice is less critical than data quality.
VarIS revolutionizes facial mesh capture, enabling photorealistic results with minimal human oversight and streamlined production workflows.
Agreement among multiple OCSR models delivers a staggering AUROC of 0.916, far surpassing traditional pixel-space verification methods.
A single architecture can achieve state-of-the-art results in pelvic imaging across multiple modalities, addressing critical data integrity issues in women's health research.