Search papers, labs, and topics across Lattice.
98 papers published across 6 labs.
Targeted middle-layer recurrence can dramatically enhance Transformer reasoning without the need for full-layer looping, leading to superior performance in complex tasks.
Loopie outperforms traditional Transformers by leveraging a novel architecture that maximizes efficiency without sacrificing reasoning power, achieving gold-medal performance in competitive settings.
Runtime load balancing in FVAttn slashes attention processing time by over four times, transforming video generation efficiency.
Evolving neural architectures can achieve high accuracy without backpropagation, leveraging shared neurons and asynchronous processing to unlock new computational capabilities.
ExaGEMM achieves a staggering 13.29x latency reduction for low-bit GEMM on CPUs, revolutionizing how we approach efficient ML inference.
Loopie outperforms traditional Transformers by leveraging a novel architecture that maximizes efficiency without sacrificing reasoning power, achieving gold-medal performance in competitive settings.
Runtime load balancing in FVAttn slashes attention processing time by over four times, transforming video generation efficiency.
Evolving neural architectures can achieve high accuracy without backpropagation, leveraging shared neurons and asynchronous processing to unlock new computational capabilities.
ExaGEMM achieves a staggering 13.29x latency reduction for low-bit GEMM on CPUs, revolutionizing how we approach efficient ML inference.
Training deep residual networks is safe only if the input-magnitude exponent of their velocity fields is constrained to one or less, revealing a critical architectural design principle.
Near-zero forgetting in continual learning is achievable with a geometric approach that preserves old knowledge while integrating new information seamlessly.
VideoSEMA outperforms heavier models while maintaining efficiency, achieving top accuracy even as image resolution scales up.
A modern multimodal assistant can achieve impressive performance on legacy hardware, revealing surprising insights about weight precision and long-context processing.
Despite its simplicity, DynaBase rivals state-of-the-art models in zero-shot reconstruction, revealing that less can indeed be more in forecasting dynamical systems.
Achieving up to 40x energy efficiency gains in deep learning inference by integrating ADC-free nonlinear operations directly into FPGA architectures.
FlashDecoder decodes video frames 3.6x-4.7x faster than traditional methods while maintaining high-quality reconstruction at 1080p.
Targeted middle-layer recurrence can dramatically enhance Transformer reasoning without the need for full-layer looping, leading to superior performance in complex tasks.
HGA enables fine-tuning on sequences of 16,384 tokens with just 15.28 GB of VRAM, outperforming traditional methods constrained by memory limits.
A novel pipeline converts human-written satellite mission plans into executable workflows while enforcing safety policies, reducing the risk of mission failures.
Pattern-guided exploration can slash FPGA design search spaces by over 80% while preserving optimal performance.
Achieving 100% routability while reducing wirelength by up to 23% could redefine standards in advanced package design.
Achieving 17x faster memory allocation without sacrificing programmability could revolutionize how we handle memory in serverless architectures.
Capturing design pattern variability could revolutionize mobile app generation, ensuring both customization and architectural integrity.
Scaling Hyper-Connections beyond four streams is now feasible, yielding substantial performance gains without prohibitive costs.
Skip connections and normalization layers are not just about controlling magnitude; they are crucial for preserving gradient rank and influencing model performance.
OT-ICA achieves superior performance in independent component recovery by maximizing the Wasserstein distance, challenging the dominance of traditional proxy-based methods.
Moderate-depth entangled Quantum Kitchen Sinks outperformed classical methods in RF anomaly detection, achieving an impressive AUROC of 0.8778 on real-world signals.
Achieving 90% of asymptotic performance with just 64 parameters challenges conventional wisdom about model complexity in molecular property prediction.
A task's representability determines not just generalization but also memorization, leading to a binary success-failure outcome in training.
The sharpness of neural network solutions is fundamentally tied to class distribution, revealing critical insights for optimization and generalization in deep learning.
Achieving reliable approximations of solutions to ill-posed inverse problems using a straightforward neural network training scheme could revolutionize how we tackle complex parameter-dependent challenges.
KANs deliver superior classification performance but come with a hefty computational price tag that may not justify their use in all scenarios.
EXPLORE boosts analog circuit topology generation success rates from 12% to 65% by intelligently combining Monte Carlo Tree Search with transformer decoding.
MyAG reveals that a graph-based approach can revolutionize the design of LLM agent systems, enabling unprecedented flexibility and efficiency in execution.
Achieving up to 129x speedups for hash-based PQC algorithms without compromising resource efficiency, HORCRUX redefines the capabilities of RISC-V in the post-quantum era.
Achieving a 4.2x throughput boost for MXFP formats could redefine efficiency benchmarks in low-precision deep learning inference on FPGAs.
MixCompress redefines the landscape of learned image compression by dynamically scaling model capacity to achieve superior performance across variable bit rates.
Achieving up to 30.45% better compression than traditional codecs, DCVC-MB redefines efficiency in neural video coding.
Achieving a staggering 24.0x speedup in quantum error correction simulations while maintaining high fidelity opens new avenues for resource-efficient quantum computing.
Kaleido achieves a remarkable 5.9x speedup in video diffusion transformers by intelligently reusing computations based on latent space correlations.
Looping Transformers can achieve better performance with shared parameter updates, revealing that scaling rules must adapt to parameter visits for stable recurrent depth.
Inhibited Self-Attention sharpens focus on relevant objects in Vision Transformers, significantly improving performance by reducing reliance on spurious correlations.
Autonomous AI agents can now seamlessly orchestrate IoT environments, transforming how we manage smart buildings and beyond.
Global spatial information in traffic forecasting can be captured just as effectively with simpler methods, raising questions about the necessity of complex attention mechanisms.
CoCo loss achieves faster convergence and tighter class clustering, outperforming traditional methods in learning optimal embeddings.
AVQ Attention refines key representations dynamically, achieving better accuracy and efficiency without increasing computational complexity.
SinAE achieves near-lossless reconstruction across diverse atomic systems, drastically improving generative performance while simplifying the architecture.
The implicit bias of diagonal linear networks under infinitesimal initialization reveals a surprising alignment with a modified \( \mathcal{l}_1 \) norm, reshaping our understanding of their training dynamics.
Routing attribution reveals that the choice of data views can dramatically influence model interpretability, even when performance gains are minimal.
ICW achieves up to 28 times improvement in accuracy for power flow surrogates while adapting 21 to 34 times faster than conventional methods.
ORRAM achieves the same bit density as traditional DFFRAMs but with a vastly superior feature set and no reliance on custom components.
Early-stage 3D IC design can now achieve accurate performance predictions by integrating thermal and cache effects into the layout optimization process.
Achieving over 2.2x inference speedup in Vision Transformers while maintaining accuracy reveals the untapped potential of hardware-software co-design in optimizing model performance.
Real-time adaptive encoding in a low-power AFE can revolutionize wireless neural signal processing for brain-computer interfaces.
MambaPSA achieves a remarkable 17.6% boost in CPU inference speed while maintaining competitive accuracy in YOLO26, showcasing the potential of state space models in object detection.
Native integration of messaging within CRM systems transforms how enterprises manage communications, enhancing efficiency and scalability.
EMAGN reduces the complexity of traffic forecasting models from quadratic to linear, enabling larger configurations without sacrificing performance.
Suppressing structural KEY tokens in KV cache eviction can dramatically enhance LLM accuracy, recovering up to 98% of lost performance with minimal computational cost.
EcoSpec reveals that optimizing for expert activation costs can lead to faster decoding in large-scale MoE models, achieving significant speedups without sacrificing model performance.
GDPR compliance and operational efficiency can be achieved simultaneously in elderly care IoMT systems, challenging the notion that they are mutually exclusive.
Microflow reveals hidden performance bottlenecks that traditional tools miss, transforming how we analyze hardware-software interactions.
WaterMoE achieves watermarking with only 1% additional latency while maintaining fidelity nearly indistinguishable from unwatermarked outputs.
StratMamba achieves a 5% faster navigation speed and 0.915 path efficiency, setting a new benchmark for obstacle avoidance in robotic navigation.
ArchSim enables researchers to scale simulation studies effortlessly while maintaining high accuracy, revolutionizing how computer architecture research is conducted.
Achieving up to 34.05 FPS with only a 5% accuracy drop, this DPU-aware YOLO architecture redefines efficiency for edge-based object detection.
Achieving $O(W)$ storage efficiency and high cache hit rates in a large-scale LLM serving system could redefine performance benchmarks for hybrid architectures in production.
SlimPer achieves deeper relevance understanding in recommendation systems without the computational bloat, leading to measurable boosts in user engagement on platforms like Instagram.
Transforming our understanding of Transformers, this work reveals that learnability may be as crucial as expressivity in optimizing large language models.
Sparse neuron activations in Transformers reveal that only a few upstream signals are crucial for maintaining activation fidelity, challenging assumptions about dense parameter reliance.
ARMT-augmented models can handle inputs far beyond their original context limits while using 30% less compute, revolutionizing efficiency in long-context processing.
RepTran achieves a remarkable 74.7% repair rate for Transformer models, significantly outperforming existing repair methods.
UMoE transforms underperforming expert pools into high-performing domain-specific models, achieving up to 6.0 points improvement on key benchmarks without increasing computational costs.
Achieving 84.85% accuracy with only 174,000 parameters on CIFAR-10, this framework makes advanced NAS accessible on consumer hardware.
Trained models adaptively reallocate their state space based on input, enabling pruning strategies that can halve output error at reduced computational budgets.
FactorDiff reveals that dynamically routing sample factors to specialized experts can dramatically enhance reasoning performance in discrete diffusion models.
Enforcing convexity in neural networks can lead to exponential increases in size without improving expressiveness.
DAG-FM achieves state-of-the-art causal discovery performance by dynamically adapting to diverse causal mechanisms, outperforming both classical algorithms and recent models.
TECO achieves a groundbreaking balance between accuracy and efficiency by pruning CNNs across multiple dimensions, setting a new standard for embedded hardware performance.
Backpropagation is not just an algorithm; it’s a nilpotent linear system that reveals profound insights into network architectures and gradient flow.
Hierarchical orchestration can drastically reduce decision-space explosion in LLMs, enhancing both routing accuracy and memory efficiency.
INCs can outperform traditional dynamic controllers, achieving lower costs in challenging control scenarios while ensuring mathematical rigor in their design.
HCRMap slashes end-to-end latency by up to 46.7% in 3.5D MoE inference by intelligently managing expert residency based on real-time pressure metrics.
Architectural decisions are often made by non-architects, challenging the assumption that formal titles dictate influence in software design.
Curriculum fine-tuning can significantly enhance the success rate of LLMs in neural architecture synthesis, but distinct failure modes require different repair strategies.
CUST slashes memory usage and boosts inference speed while enhancing image super-resolution performance by intelligently aggregating patch similarities.
MMA-Former's innovative attention mechanism enables spatially adaptive feature extraction, significantly improving PNI prediction accuracy from 3D MRI scans.
ReviewDSE achieves a 1.78% reduction in wirelength while exposing and repairing design flaws that traditional methods miss.
Achieving near-perfect anomaly detection while automating network access control could redefine security protocols for IoT devices.
Certified token eviction in attention mechanisms can now be achieved without sacrificing output integrity, enabling precise memory management in AI systems.
AtomFlow slashes quantum computing latency to 25.3 ms by integrating control tasks on a single FPGA, paving the way for faster neutral atom operations.
Modern LLM performance hinges on dependency structures rather than individual instruction latencies, revealing a critical insight for GPU optimization.
Achieving up to 9 percentage points improvement in latency prediction accuracy, HiFi-LLP revolutionizes hardware-aware neural architecture search efficiency.
Rethinking feature fusion, this study reveals that focusing on the differences between feature streams can dramatically enhance U-Net performance across diverse applications.
Gauntlet outperformed human reviewers in technical critique of computer architecture papers, revealing that LLMs can achieve significant analytical depth through structured multi-agent collaboration.
Achieving nearly 800x speedup in 3D scene modeling without sacrificing quality, AsySplat redefines efficiency in long-sequence novel view synthesis.
Single-head attention is a precise connection propagation step, while multi-head attention reveals complex edge-dependent dynamics that reshape our understanding of transformer geometry.
Multi-agent collaboration not only outperformed other architectures in accuracy but also underscores the critical role of design in shaping journalistic AI's effectiveness and transparency.
Manufacturing variations can halve accuracy, but ARMOR-IMC restores performance while simultaneously defending against power side-channel attacks.
DP-Splat adapts the number of Gaussian components to scene complexity, achieving up to 7.6x fewer components while maintaining or exceeding predictive accuracy.
Visual programming for AMD Ryzen AI NPUs eliminates the need for specialized coding skills, making advanced machine learning accessible to a broader audience.
TOLiD achieves superior transfer learning in LiDAR representation by effectively bridging the modality and architecture gaps between vision and 3D data.
Identifying vulnerability nodes in PCHB circuits reveals that targeted hardening strategies can significantly enhance resilience without excessive overhead.
Radial Rotary Complex Attention transforms interatomic potential modeling, achieving state-of-the-art performance by enhancing extrapolation capabilities.