Search papers, labs, and topics across Lattice.
38 papers published across 2 labs.
Raven achieves superior long-context recall by intelligently routing memory updates, outperforming traditional models that struggle with interference and eviction.
Converged diffusion loss in visual generation improves linearly with structured language, leading to a new training paradigm that outperforms both open-weight and closed-weight models.
Repeated sampling outperforms self-reflection methods in language models, revealing a critical flaw in approaches that rely on self-critique and rewriting.
Diminishing returns in inference-time scaling reveal that more computation doesn't always equate to better performance in local computer-use agents.
ROCS can triple the throughput of recommendation systems without sacrificing prediction accuracy, revolutionizing how we handle user requests in large-scale settings.
Converged diffusion loss in visual generation improves linearly with structured language, leading to a new training paradigm that outperforms both open-weight and closed-weight models.
Repeated sampling outperforms self-reflection methods in language models, revealing a critical flaw in approaches that rely on self-critique and rewriting.
Diminishing returns in inference-time scaling reveal that more computation doesn't always equate to better performance in local computer-use agents.
ROCS can triple the throughput of recommendation systems without sacrificing prediction accuracy, revolutionizing how we handle user requests in large-scale settings.
CCFormer delivers a 3.57% increase in click-through rates and a 1.71% boost in advertising revenue, all while cutting model training time by over 2x.
Chimera achieves 7.3x compute efficiency over traditional models while enabling zero-shot extrapolation from short video clips to significantly longer sequences.
Sequence models outperform LLMs in predictive process monitoring, challenging the assumption that larger models always yield better results.
Agents can now autonomously teach themselves creative skills using high-quality human texts, bypassing the need for expensive human feedback.
CoMem achieves a 7.83x prefill speedup and drastically reduces memory usage while maintaining high performance on long-context tasks, challenging conventional memory management in LLMs.
SCSE transforms recurrent computation in Looped Transformers by ensuring that deviations from a learned anchor enhance performance without compromising stability.
Achieving high-fidelity 3D generation with just 1.5% of the training data could revolutionize resource allocation in 3D modeling.
Expert subspaces in MoE models overlap significantly, yet selected routes yield better token representations than their strongest unselected counterparts, challenging traditional views on redundancy.
Independently scaling long-term memory in language models can yield better performance with fewer parameters than simply increasing model size.
Shifting the burden of intelligence from runtime search to a belief-based framework allows for professional-level Go play on consumer hardware without the need for extensive MCTS.
Predicting the next token's KV entries can boost long-context LLM throughput by over 2.5 times without sacrificing latency or quality.
Knowledge transfer in MLLM fusion is not a blanket inheritance but a selective process favoring high-level reasoning over perception.
A photonic memory architecture can slash LLM inference latency by over 50%, paving the way for scalable, high-performance KV cache management.
BM25 outperforms all competitors at scale, revealing that lexical retrieval is the most efficient choice for large corpus sizes.
Exploration as a pretraining axis can boost generative model performance by up to 36%, revolutionizing how we approach end-to-end generation.
Raven achieves superior long-context recall by intelligently routing memory updates, outperforming traditional models that struggle with interference and eviction.
Logarithmic-depth networks can efficiently learn complex Boolean functions that constant-depth networks cannot, revealing a critical algorithmic separation in neural network capabilities.
Localized latent reasoning can slash inference latency while maintaining competitive accuracy, challenging the need for larger models or lengthy reasoning chains.
Contextual influence in human language diminishes predictably with distance, following a scaling law that could redefine our understanding of linguistic structure.
LLMs and human cognition share strikingly similar principles of cognitive organization, challenging the view of AI as an alien intelligence.
AngelSpec achieves nearly double the inference speed of traditional methods while improving output quality by intelligently adapting to the specific demands of different tasks.
Token effectiveness varies dramatically with model size and data strategies, revealing that classic compute-optimal methods are often misguided in real-world applications.
A unified taxonomy reveals how diverse memory mechanisms in LLMs can be systematically understood and leveraged for future innovations.
Kimi K3's innovative architecture achieves a 2.5x scaling efficiency improvement, enabling robust performance across diverse long-horizon tasks.
A single dial in the S-LCU framework allows researchers to balance the trade-off between computational complexity and trainability in quantum circuits.
MMOE achieves faster convergence and superior generation quality in diffusion transformers by effectively integrating expert routing strategies, challenging the notion that more parameters always lead to better performance.
LOCKS achieves near-full KV performance with only 2% of the tokens attended, revolutionizing long-context decoding efficiency.
MXAttention achieves near-FP16 generation quality in video models while cutting computational costs, revolutionizing efficient attention mechanisms.
Dynamic head grouping and adaptive rank allocation can drastically cut Key cache parameters without sacrificing performance in LLMs.
Degenerate covariance structures can trigger multiple descent in regression, revealing unexpected peaks in prediction risk that challenge conventional wisdom.
Agents that learn to optimize their own weights can achieve unprecedented stability and efficiency over long time horizons.
Static leaderboards can mislead deployment decisions, as a system with higher factuality may actually be less efficient when compute costs are factored in.
Larger vision-language models dramatically improve internal uncertainty signals, but their verbalized confidence often fails to keep pace, revealing a critical disconnect.
A frozen 12B model achieves 100% accuracy on new problem instances at zero generation tokens, challenging the need for constant retraining in language models.