Search papers, labs, and topics across Lattice.
26 papers published across 1 lab.
SATS achieves a remarkable 65.6% increase in model efficiency while improving prediction accuracy in time series analysis.
Full Relation not only outperforms MHA in validation NLL across all tested scales but also accelerates processing speed by up to 4.41 times with FlashRelation.
NAPE achieves state-of-the-art performance in audio representation learning by simplifying the pre-training process to a single autoregressive prediction task.
Achieving up to 47.26x speedup in long-context LLM serving could redefine efficiency benchmarks in AI inference.
Optimal learning rates for training massive Mixture-of-Experts models can be accurately predicted from small proxy models, achieving high fidelity even at trillion-token scales.
SATS achieves a remarkable 65.6% increase in model efficiency while improving prediction accuracy in time series analysis.
Full Relation not only outperforms MHA in validation NLL across all tested scales but also accelerates processing speed by up to 4.41 times with FlashRelation.
NAPE achieves state-of-the-art performance in audio representation learning by simplifying the pre-training process to a single autoregressive prediction task.
Achieving up to 47.26x speedup in long-context LLM serving could redefine efficiency benchmarks in AI inference.
Optimal learning rates for training massive Mixture-of-Experts models can be accurately predicted from small proxy models, achieving high fidelity even at trillion-token scales.
Distinct neural architectures exhibit surprisingly similar collective dynamics, revealing a universal infrared organization that transcends their microscopic differences.
Hybrid quantum models can achieve competitive forecasting performance with fewer parameters and distinct learning trajectories compared to classical counterparts.
Exploitation, not exploration, is the critical bottleneck in test-time scaling for language models, with selection processes yielding near-random results despite rich candidate pools.
Reducing feature duplication from 37.2% to 6.8% while maintaining state-of-the-art performance with just 5.4 features showcases a breakthrough in metadata-free AutoFE.
Nuisance interpolation in feature priming can lead to significant underperformance in online regression, revealing a fundamental limitation in existing regret guarantees.
A centralized model repository can enhance CSI feedback efficiency, achieving significant performance gains while slashing local training requirements.
Surrogates can now handle over 3 million controllable variables, achieving up to 26.5x speedups over traditional simulation methods.
The minimax risk in deep Gaussian regression exhibits a surprising quadratic dependence on depth, challenging conventional assumptions about model capacity.
Diffusion models require ten times more data than language models to achieve optimal performance, challenging existing paradigms in visual generation.
FLEXRec shows that compact LLMs can achieve state-of-the-art recommendation accuracy without the computational burden of larger models.
Fine-grained MoE designs can outperform dense vision encoders while dramatically reducing latency, challenging the status quo in image and video understanding.
Traditional accuracy metrics can obscure critical improvements in reasoning fluency, as shown by fine-tuned models that reason correctly in low-resource languages despite initial benchmark null results.
The curse of multilinguality may not be a fundamental barrier; instead, it’s a byproduct of data and training practices.
Localized TabICLv2 retains nearly all the accuracy of its predecessor while slashing inference time, making it a game-changer for large-scale tabular data tasks.
Preventing unnecessary use of LLMs could dramatically reduce their climate impact, revealing a crucial path for sustainable AI development.
GigaBrain-0.7 achieves unprecedented task adaptability and completion rates, outperforming previous models in both home and industrial settings.
TinyCast shatters the size-accuracy barrier in zero-shot forecasting, achieving superior performance with a fraction of the parameters of its competitors.
Gated Recurrent Transformers achieve 63% fewer parameters and 59% less peak memory while matching the accuracy of much larger models, redefining efficiency in language modeling.
Surprisingly, larger LLMs benefit from increased repetition of high-quality domain data, challenging conventional wisdom about data diversity in training.
Achieving high-fidelity 3D object generation with up to 300 parts, MegaParts redefines the limits of part-aware modeling through token-efficient autoregressive techniques.
Forecast collapse in time-series models reveals a critical calibration-ranking tradeoff that could mislead financial decision-making.