Search papers, labs, and topics across Lattice.
86 papers published across 6 labs.
Performance gains of up to 59% are achieved by aligning block size with cache shape, challenging conventional wisdom about architectural optimization.
Minor architectural tweaks can lead to a staggering 47% drop in long context performance, challenging assumptions about model design.
Achieving similar performance to larger models with significantly less data and faster inference speeds could redefine efficiency benchmarks in foundation models.
Fixing deployment bugs in Nanbeige4.2-3B transforms it from a non-functional model to one that can tackle real agentic tasks with a significant performance boost.
Traditional load balancing in EP MoE serving can lead to inefficiencies, but a new makespan-aware dispatcher achieves up to 15.5% throughput gains by adapting to varying compute and memory constraints.
Achieving similar performance to larger models with significantly less data and faster inference speeds could redefine efficiency benchmarks in foundation models.
Fixing deployment bugs in Nanbeige4.2-3B transforms it from a non-functional model to one that can tackle real agentic tasks with a significant performance boost.
Traditional load balancing in EP MoE serving can lead to inefficiencies, but a new makespan-aware dispatcher achieves up to 15.5% throughput gains by adapting to varying compute and memory constraints.
Transforming feature relationships through geometric guidance leads to consistent performance gains in both image classification and reasoning tasks across multiple architectures.
HybridSB-MoE achieves superior speech enhancement by intelligently fusing spectral and waveform models, outperforming traditional methods in both efficiency and quality.
Despite promising theoretical foundations, RoPE-aligned rotations fail to improve quantization accuracy, highlighting a critical misalignment in current methods.
MergeOver reduces peak activation memory by over 37% while maintaining competitive accuracy, making it a game-changer for deploying Vision Transformers on edge devices.
Learnable wavelet activations can significantly enhance plasticity in continual learning, preventing catastrophic forgetting while maintaining performance across tasks.
AaLLM can generate innovative circuit topologies that outperform conventional designs while slashing design time by up to 40x.
INT4 quantization can distort the architecture landscape, but a zero-shot FP32 surrogate offers a surprising advantage in maintaining Pareto optimality.
Independently trained depth slices can be recombined to match the performance of monolithic models, revealing a new avenue for efficient language model training.
Post-norm normalization significantly enhances performance in LLMs when depth is introduced through a curriculum, outperforming pre-norm by an order of magnitude.
SCOPE achieves a remarkable 1.99× speedup in video attention while enhancing fidelity, challenging the effectiveness of traditional sparse attention methods.
An agent-driven approach to hardware design has achieved a 61.1% IPC speedup, outperforming expert-designed prefetchers by up to 23.6%.
YAVIN achieves over 20x speedup in secure edge processing while maintaining robust cryptographic protections against untrusted memory.
Dryas can dynamically adjust its filtering capabilities in under a second, revolutionizing how engineers debug and analyze high-speed interconnects without interrupting their operation.
Executable contracts can transform how we evolve legacy hardware, ensuring validated designs adapt without starting from scratch.
Regime information can destabilize neural network training, but routing it smartly can significantly boost volatility forecasting accuracy.
Reducing parameter redundancy in KANs, HYDRA achieves superior predictive performance while enhancing interpretability in hyperbolic spaces.
TESLA not only solves the parity problem with minimal data but also excels in robustness, outperforming traditional methods under label noise.
SATADL can accurately predict air quality for up to 48 hours during sensor failures, outperforming traditional models in both accuracy and reliability.
TradingMoE boosts trading returns by over 30% by dynamically selecting the most relevant experts based on evolving market conditions.
VGG16 outshines other CNN architectures in Alzheimer's detection, but classifying early-stage dementia remains a significant hurdle.
Gradient manipulation in multi-task learning can be significantly enhanced by recognizing the matrix structure of model parameters, leading to superior optimization outcomes.
Normalizing dual-encoder networks not only clarifies their interpretability but also reveals hidden structure in learned representations, challenging existing assumptions in the field.
Achieving a 3.2x speedup in video diffusion without sacrificing fidelity could redefine efficiency benchmarks in the field.
XBRIDGE reduces communication latency by 11x while ensuring that heterogeneous LLMs maintain entity identity and contextual relevance.
STAR achieves state-of-the-art performance in 3D scene understanding by effectively balancing semantic consistency with geometric heterogeneity.
HyGA achieves unprecedented improvements in attention mechanisms by effectively integrating multiple gating strategies, leading to enhanced representational capacity and training efficiency.
A single-layer speech enhancement model outperforms naive architectures and achieves competitive quality with a significant speedup through progressive knowledge distillation.
Expert rerouting and attention-aware data packing can boost MoE model throughput by nearly 15%, reshaping efficiency in reinforcement learning.
APEX achieves over 99% overlap accuracy in expert prefetching, slashing per-token latency by up to 26% while enhancing energy efficiency for edge MoE inference.
Performance gains of up to 59% are achieved by aligning block size with cache shape, challenging conventional wisdom about architectural optimization.
By buffering intermediate values in fast DRAM, NITRO slashes inference latency by up to 85%, revolutionizing the efficiency of NAND flash-based computing.
Achieving a staggering reduction in hardware footprint while maintaining near-perfect accuracy, Uni-SFU redefines the efficiency of activation function implementations in neural networks.
Pre-attention spikes and inter-spike plateaus reveal a surprising organization in hybrid linear attention models that could redefine our understanding of activation dynamics in LLMs.
An AI coding agent can autonomously refactor a complex codebase, achieving zero bugs and correcting over 200 defects without human oversight.
Achieving a remarkable -16.85% BD-Rate improvement over VVC, this MoE-based approach revolutionizes learned image compression efficiency.
Faults in Barrett Modular Multiplication can be detected efficiently with minimal overhead, safeguarding critical cryptographic operations against adversarial attacks.
Malicious modifications to router weights can turn Mixture-of-Experts models into trigger-controlled bottlenecks, revealing a critical vulnerability in AI serving architectures.
Ark, an open-source coding agent, solves 80% of software maintenance tasks while offering a clear architectural framework that could redefine how we study coding agents.
Real-time monitoring of patient vitals is now feasible with a cloud-native framework that seamlessly integrates diverse wearables without vendor lock-in.
Energy and latency can diverge by 3x under high computational demand, necessitating platform-specific models for accurate CNN inference cost predictions.
Charge-CIM slashes ADC energy consumption by 91.7% while doubling throughput, reshaping the landscape of energy-efficient deep learning accelerators.
Achieving a 1.85x speedup in matrix multiplication on Ascend NPUs could redefine performance benchmarks for dynamic tensor operations.
Tool architecture can enhance coding agent performance, with structured interfaces yielding up to 4.7 times more consistency and 41.6% fewer steps in task execution.
ODE-inspired update dynamics can significantly boost sign language translation performance without the need for larger models.
Quantum mechanics can provide an exact realization of softmax attention, transforming how we understand attention mechanisms in AI.
Compressing tabular models by 85% without sacrificing performance could revolutionize how we deploy foundation models in resource-constrained environments.
Optimal sparsity in sparse MoE models emerges only when considering the cluster's systems constraints, challenging traditional compute-centric design paradigms.
Proxy models can slash RL post-training costs by up to 87.5% while maintaining critical fault reproduction capabilities.
Inverse-distance attention can achieve exact retrieval with constant resources, outperforming softmax's logarithmic scaling in complexity.
Sharing expert knowledge before routing leads to a substantial reduction in computational demand while boosting model performance.
Minor architectural tweaks can lead to a staggering 47% drop in long context performance, challenging assumptions about model design.
Replacing fixed attention mechanisms with a learned, input-generated operator reveals surprising invariance properties that could redefine our understanding of attention in LLMs.
Causal attention in Post-Norm Transformers amplifies token similarity, leading to a collapse that training dynamics fail to repair, revealing critical insights into model behavior.
Recursive self-improvement in Macaron-V1 leads to continual learning that adapts to real-world experiences, setting a new standard for open agent models.
Motif 3 achieves unprecedented efficiency and performance in language modeling by leveraging a novel Mixture-of-Experts architecture that activates only a fraction of its parameters per token.
Expert allocation in retinal pathology detection is not only disease-dependent but also enhances interpretability, revealing how models can effectively disentangle complex co-occurring conditions.
MoNo's innovative approach to optimal transport ensures stable latent spaces, preventing token collapse and enabling efficient learning of long-range physical interactions in PDEs.
ICNNs can outperform traditional neural network surrogates by providing tighter relaxations and faster optimization in real-world applications.
Multiple ordered projections can reveal hidden structural dependencies that single sequential models miss, enhancing RNN performance on complex tasks.
A random transformer can achieve universal approximation without any pretraining, challenging conventional beliefs about model training requirements.
MixFormer achieves substantial performance improvements in long-context tasks by dynamically managing memory with a novel Mixture-of-Memory-Experts approach.
LEED uncovers hidden over-smoothing patterns in GNNs, enabling a new level of precision in node representation analysis and virtual node selection.
LITEWAY slashes model size and energy consumption for wearable human activity recognition without sacrificing performance, achieving up to 9.52x size reduction.
Task-oriented multi-agent systems can achieve unprecedented performance by leveraging specialized roles and subtasks, revealing the power of structured collaboration in AI.
ArchAgent v2 outperforms hand-designed solutions by achieving a 3.8% IPC speedup through innovative multi-level data prefetching strategies.
ZetaGPT achieves position-aware language modeling without explicit positional encodings, paving the way for more efficient and compact architectures.
ANTMAN achieves zero false positives and rapid detection of stealthy branch predictor attacks, redefining runtime security for RISC-V architectures.
DistMoE enables MLLMs to adapt to diverse visual-language domains without centralized data, achieving competitive performance while preserving client-specific knowledge.
A SysML-driven MBSE framework transforms UAV design by ensuring seamless integration and traceability across complex system components.
FSGen achieves a staggering 1.4x power efficiency improvement and 10x speedup for LLM accelerators, redefining the landscape of AI chip design.
Unified audio generation just got a major upgrade—SonicWeave's innovative routing mechanism boosts compositional quality and expert specialization, outperforming traditional models.
Achieving over 20% performance improvement in data-free knowledge distillation, UniDFKD eliminates reliance on architecture-specific priors, paving the way for more robust model training across diverse architectures.
Zero-initialization emerges as the key driver behind the superior performance of adaLN-Zero in diffusion transformers, reshaping our understanding of conditioning mechanisms in image generation.
Bridging the gap between high-level reasoning and real-time interaction, this architecture enables embodied agents to think and act in complex virtual environments seamlessly.
Latent feedback in full-bandwidth transformers enables deeper contextual understanding without sacrificing efficiency, leading to improved performance across multiple tasks.
Tying scales in PTQTP leads to a uniform nine-level quantizer that matches official serving performance while reducing file size and improving decoding speed.
The evolution of Mixture-of-Experts architectures reveals a critical shift towards decoupling semantic routing from computational budgets, reshaping our understanding of model efficiency.
Infrastructural complexity, not execution performance, is the main barrier to adopting distributed computing, revealing a critical need for new architectural paradigms.
C2C-Explorer boosts LLM inference efficiency, achieving a 44.1% increase in goodput and a staggering 98.4% reduction in memory usage.
FlashBoot slashes weight loading times from 20.1 seconds to just 0.4 seconds, revolutionizing how large models are deployed at scale.
Eco-SoC achieves a remarkable 42% reduction in switching activity while offsetting its carbon footprint in just over a year of deployment.
Voltage droop violations are tackled with a 76x reduction in energy-delay product, all while keeping ML inference accuracy intact.
Achieving over 99% performance retention while significantly accelerating large recommendation models by intelligently merging experts based on functional similarity.