Search papers, labs, and topics across Lattice.
95 papers published across 6 labs.
SiPE not only boosts syntactic accuracy but also enhances general language understanding, setting a new standard for integrating syntax into Transformer models.
Dynamic layer routing can boost LLM accuracy by 5% without the need for weight updates or expensive search loops.
Geometry-adaptive prompting can transform how we approach few-shot learning in dynamic graphs, leading to substantial performance gains.
MicroEvo achieves a staggering 10.6x increase in search efficiency while improving Pareto-front quality by 36.2%, revolutionizing microarchitecture design exploration.
Dense-Cast achieves a remarkable MAE of 0.235 mm for half-hourly precipitation nowcasting, setting a new benchmark for accuracy in this challenging domain.
SiPE not only boosts syntactic accuracy but also enhances general language understanding, setting a new standard for integrating syntax into Transformer models.
Dynamic layer routing can boost LLM accuracy by 5% without the need for weight updates or expensive search loops.
Geometry-adaptive prompting can transform how we approach few-shot learning in dynamic graphs, leading to substantial performance gains.
MicroEvo achieves a staggering 10.6x increase in search efficiency while improving Pareto-front quality by 36.2%, revolutionizing microarchitecture design exploration.
Dense-Cast achieves a remarkable MAE of 0.235 mm for half-hourly precipitation nowcasting, setting a new benchmark for accuracy in this challenging domain.
A runtime observability framework reveals how modern memory architectures can be rigorously monitored and quantified, exposing hidden risks in AI model performance.
A single model can now adaptively balance short-term precision and long-term accuracy in weather forecasts, revolutionizing how we approach atmospheric predictions.
JTA reveals that validation gaps in safety-critical software can be systematically addressed by treating scenarios, test systems, and systems under test as interconnected entities.
ShimGen not only matches but surpasses manually-designed protocols in consistency, revealing critical performance gains in heterogeneous memory systems.
Achieving a staggering 7,300x reduction in instruction overhead for sensor reads could redefine efficiency benchmarks in embedded systems.
U-Nets unexpectedly show greater robustness to resolution changes than anticipated, challenging assumptions about neural operator architectures in inverse imaging.
ASTELD uncovers a critical gap in the autonomous AI landscape: no evaluated systems achieve both local-first deployment and enterprise-grade security.
Regime-aware modeling with MGSB boosts leak detection performance by over 20% in out-of-distribution scenarios compared to traditional methods.
Recurrent Vision Transformers can outperform standard models in accuracy-to-parameter trade-offs when memory constraints are prioritized, challenging conventional wisdom about architectural efficiency.
Novel aggregation techniques for Memory Augmented Autoencoders in Federated Learning can enhance anomaly detection performance, even in shallow models.
Pruning Echo State Networks dynamically can enhance forecasting accuracy while significantly reducing model complexity.
Elbow-based routing can cut inference latency by over 5% in MoE models while preserving accuracy, revolutionizing expert selection efficiency.
Merging experts based on their phase roles can enhance MoE-VLM performance by up to 9.6%, challenging the effectiveness of traditional global aggregation methods.
MESH achieves a 62.5% reduction in optimizer-state memory for Mixture-of-Experts training while preserving performance, challenging assumptions about memory efficiency in deep learning.
Achieving a 3.56× increase in decoding throughput without sacrificing accuracy, BinaryPC revolutionizes efficiency in long-context LLMs.
ORACLE achieves up to 104.4x faster circuit design while meeting nearly all target specifications, revolutionizing multi-objective optimization in analog circuit design.
Eigenius not only validates scientific conclusions but also uncovers discrepancies in published research, revolutionizing how we ensure data integrity in AI-driven science.
Current LLMs fail to meet the rigorous demands of PCB routing, showing major weaknesses in path planning and constraint adherence.
MCHA achieves unprecedented performance speedups for parallel-sequential computing tasks, outperforming NVIDIA A100 GPUs by up to 2456.96×.
EdgeXpert slashes LLM inference latency by over 56% while cutting energy use by nearly 45%, all without sacrificing accuracy.
Achieving up to 91.4x speedup in O-RAN fronthaul decompression could revolutionize the efficiency of 5G networks.
AFD-Ledger reveals that optimizing deployment for AFD can drastically cut evaluation costs while exposing the nuanced performance dynamics between homogeneous and heterogeneous setups.
PowerScope achieves intra-cycle power estimation with 80x speedup and competitive accuracy, revolutionizing power analysis workflows.
Achieving up to 71% lower error in activation functions while using less area and power could revolutionize the efficiency of neural network accelerators.
Fragmented agentic AI workflows expose significant inefficiencies in conventional server architectures, necessitating a radical rethink of resource allocation strategies.
Formal verification of a compiler for asynchronous dataflow could redefine the reliability and efficiency of parallel computing architectures.
K-EXAONE 2.0 achieves over three times the capacity of its predecessor while enhancing multilingual capabilities and long-context reasoning.
Muon optimizer's unique advantage lies in its ability to enhance token efficiency specifically when applied to the output projection, challenging assumptions about conditioning in state-space models.
LAEF achieves superior point-of-care ECG diagnostics by leveraging lead-agnostic architecture, outperforming traditional models even with minimal lead data.
Energy-aware DNN design can now be optimized offline, paving the way for truly autonomous AI on intermittent power sources.
Dynamic processing in Transformers can exceed static processing, revealing a deeper layer of interaction that mirrors human language processing.
Tiny input changes can destabilize UAV tracking models, revealing a new attack surface that undermines their efficiency and accuracy.
MuEvo not only evolves heuristic ensembles but also dynamically adapts component priorities, leading to superior performance in complex optimization tasks.
AgenticECO clears 7 out of 9 defect cases with minimal disturbance, revolutionizing ECO processes in 3D-ICs by ensuring repair attribution without the chaos of full rerouting.
LoopMTP boosts reasoning accuracy by up to 8.1% by effectively guiding looped transformer iterations with multi-token prediction.
MoEGen achieves instance-specific adaptations without the storage burden of full LoRA experts, revolutionizing how we think about parameter-efficient fine-tuning.
AcceptMoE slashes host-to-device traffic by over 73% while boosting throughput by more than double, all without a major accuracy trade-off.
Achieving competitive accuracy while slashing computational costs, this method redefines the feasibility of 3D object detection on lightweight devices.
SLAMFormer-$\infty$ can handle trajectory sequences over 17 km without losing performance, revolutionizing long-range SLAM capabilities.
Real-time AI monitoring can boost tomato crop disease detection rates to 95%, transforming agriculture scalability.
Fovea achieves a remarkable 7.80x speedup in wafer architecture selection while ensuring optimal performance across diverse workloads.
CAMTA achieves nearly an order of magnitude improvement in Softmax function approximation while offering unprecedented runtime configurability in hardware.
iFAN boosts segmentation accuracy by aligning query competition with mask quality, achieving significant performance gains without extra computational overhead.
ALiBi positional encoding can blind attention heads, drastically impairing token retrieval without impacting standard performance metrics.
MoE dLLMs can outperform leading models with significantly fewer training tokens, challenging assumptions about data efficiency in large-scale language models.
Recovering multi-head attention parameters without orthogonality assumptions could revolutionize how we learn complex attention mechanisms in neural networks.
Energy efficiency in microservices is often treated as an operational issue, neglecting its critical role in architectural design.
Relocating the KV cache to processing-near-memory nodes can boost LLM throughput by over 6x while supporting evolving sparse attention methods.
Ternary LLMs can now achieve efficient attention computation without the overhead of high-precision K/V processing, revolutionizing their performance.
Achieving a 1.43x speedup in MoE training while slashing communication costs by up to 74% could redefine efficiency benchmarks in large-scale LLM training.
Static power consumption can skew efficiency estimates by up to 3.85X, revealing critical oversights in current PIM-GPU design practices for LLM inference.
Non-uniform segmentation in FPGA-based non-linear function interpolation can significantly boost accuracy while slashing hardware costs by up to 50%.
A centralized monitoring architecture can drastically simplify performance data collection across diverse hardware components, enhancing real-time system optimization.
Achieving real-time decoding of quantum error correction codes in under 1 microsecond could revolutionize fault-tolerant quantum computing.
SpecDrop shows that fixed, category-conditioned routing can outperform learned routers, achieving up to 6.53% accuracy gains without additional parameters.
Maglev achieves superior validation loss and performance benchmarks by cleverly combining full attention with sliding-window mechanisms, all while conserving memory.
Geometry-guided width allocation in Transformers can lead to substantial improvements in model performance, reducing validation loss more effectively than traditional methods.
KANs outperform MLPs in FTN BPSK detection, achieving 18.6 times lower bit error rates with drastically fewer parameters.
Hierarchical local attention in TextNCA reveals that the arrangement of attention windows can dramatically influence language modeling performance, even more than the model's iterative nature.
Uncertainty alone can mislead expert activation; VI-MoLE redefines routing by focusing on certified value-of-information, leading to more informed and efficient model decisions.
Fragmented experimental data can be transformed into actionable insights for plastic upcycling, achieving unprecedented accuracy without biased imputation.
Native computer use can be achieved at scale, enabling agents to outperform leading systems while significantly enhancing security against adversarial attacks.
CARNet achieves superior signal detection performance across diverse channel conditions by dynamically selecting specialized expert networks based on real-time channel estimates.
A single shared intervention in the query-key channel can stabilize low-precision transformer training, eliminating the need for individual fault repairs.
Real-time EEG control of exoskeletons is now feasible, achieving over 55% success in gait initiation with a novel BCI architecture that tackles motion artifacts head-on.
Retaining only the tangential component of the feed-forward network can preserve model quality while enhancing prediction accuracy in Transformer dynamics.
STEAM achieves superior EEG decoding performance with a novel hierarchical pre-training approach that enhances model specialization without starting from scratch.
Convex neural energy elements transform neural operators into reusable components that guarantee stability and accuracy, even in complex geometries.
DART achieves a remarkable 75% reduction in inference cache size while enhancing associative recall in long-context sequence modeling.
HMM achieves a remarkable 34.3% improvement in retrieval success while only adding 2% to the model's parameters.
Contextual cues can dramatically enhance emotion recognition, achieving better results with still images than traditional video-based methods.
A language model can autonomously navigate complex neural architecture design, achieving significant performance improvements while revealing the critical role of workflow design in research productivity.
Divisive normalization is not just a biological curiosity; it’s essential for stabilizing continuous working memory in neural networks, preventing the fragmentation that plagues traditional architectures.
REFLEX redefines MoE inference by aligning expert computation with the distinct refinement needs of tokens, achieving efficiency gains without compromising quality.
Bole accelerates hybrid-attention LLMs by up to 4.72 times while slashing memory usage by up to 99 times, transforming the landscape of autoregressive decoding.
RING achieves superior knowledge integration without the latency of external retrieval, redefining efficiency in large-scale language models.
By leveraging unmixing-informed spectral prompts, USP-Mamba achieves superior hyperspectral image reconstruction while maintaining spatial continuity and contextual integrity.
Achieving up to 97% reduction in energy consumption for attention mechanisms, the SMM Transformer redefines efficiency in multimodal SNN applications.
Loop-Mamba not only restores old photos more effectively but also introduces a novel degradation-aware framework that outperforms state-of-the-art methods by leveraging persistent memory and semantic guidance.
FedJigsaw redefines model personalization in federated learning by enabling clients to collaboratively assemble tailored models, leading to significant performance gains and reduced resource overhead.
LSA slashes indexing overhead while maintaining full attention performance, enabling efficient long-context processing for models with up to one million tokens.
DeGS redefines the performance landscape for 3D Gaussian Splatting, achieving up to 7.25x throughput improvements while maintaining high processing element utilization.
Achieving 2.19x throughput and 2.03x energy efficiency improvements on FPGA for SSMs by optimizing projection rank could redefine performance benchmarks in latency-sensitive applications.
LEAP achieves a 7.6x speedup in toggle propagation prediction while maintaining near-perfect accuracy, revolutionizing power analysis in VLSI design.
PRECOG achieves a staggering 4500× reduction in prefill latency, transforming retrieval-augmented generation from sluggish to interactive on edge devices.
Celty achieves up to 5.3x speedup in LLM inference by harnessing dual-sparsity, revolutionizing how we approach GPU kernel design for sparse workloads.
Machine-learning predictors can mislead design decisions, with over 22% of configurations showing unexpected performance reversals that traditional simulation still captures best.
A dynamic selection approach among microarchitectural policies can recover over 70% of performance potential without executing inactive policies.
Achieving high-quality reranking with a 30B MoE model is now feasible on an academic budget, outperforming traditional dense models in efficiency.
ReBA achieves over fivefold improvement in load balancing for Vision-Language MoE without sacrificing accuracy, revealing a critical interplay between token mix and load profiles.