Search papers, labs, and topics across Lattice.
82 papers published across 6 labs.
NanoSleep achieves superior sleep stage classification accuracy while being small enough for deployment on wearable devices, striking a crucial balance between performance and efficiency.
ODEONN achieves a 45× reduction in energy-delay product while maintaining over 98% accuracy compared to traditional software simulations.
Achieving best-in-class energy efficiency of 39 pJ/b, this innovative detector redefines the capabilities of MU-MIMO-OFDM systems.
Full Relation not only outperforms MHA in validation NLL across all tested scales but also accelerates processing speed by up to 4.41 times with FlashRelation.
A new Bayesian estimator achieves competitive performance with established methods while being remarkably simple and tuneless, revealing critical scaling laws in data discovery.
ODEONN achieves a 45× reduction in energy-delay product while maintaining over 98% accuracy compared to traditional software simulations.
Achieving best-in-class energy efficiency of 39 pJ/b, this innovative detector redefines the capabilities of MU-MIMO-OFDM systems.
Full Relation not only outperforms MHA in validation NLL across all tested scales but also accelerates processing speed by up to 4.41 times with FlashRelation.
A new Bayesian estimator achieves competitive performance with established methods while being remarkably simple and tuneless, revealing critical scaling laws in data discovery.
A novel decision-making framework for expert model management achieves zero false spawns and reuses, revolutionizing how streaming systems adapt to new data.
QUASAR achieves unprecedented data efficiency in satellite authentication, requiring only 10% of the training data while outperforming classical methods in accuracy.
Software 3.0 could redefine the entire software engineering landscape by merging reasoning and context into a unified architecture.
Axon achieves up to 107% speedup on JAX, revolutionizing how LLMs can be efficiently deployed across different frameworks without sacrificing optimization.
Recalibrating the CFG scale in response to CIM noise can restore over 87% of generation quality lost due to nonidealities in Diffusion Transformers.
Achieving up to 42.1x faster clock rates for quantum error correction could redefine the efficiency of fault-tolerant quantum computing.
Network bandwidth emerges as the key bottleneck for on-premises Earth observation data access, with implications for infrastructure investment strategies.
Voltage droops can now be corrected in real-time, potentially unlocking higher operational frequencies in energy-efficient VLSI circuits.
Achieving over 90% accuracy in predicting MBIST costs without the need for full RTL synthesis could revolutionize design efficiency in memory IP development.
A 14M parameter model that outperforms larger transformers while being 12 times smaller, reshaping the efficiency landscape of audio-visual processing.
Daedalus-150M outperforms larger competitors while being optimized for CPU inference, achieving faster decoding and lower memory usage.
Achieving up to 47.26x speedup in long-context LLM serving could redefine efficiency benchmarks in AI inference.
Optimal learning rates for training massive Mixture-of-Experts models can be accurately predicted from small proxy models, achieving high fidelity even at trillion-token scales.
Trust your predictions: Lévy Attention quantifies uncertainty in real-time without sacrificing accuracy, outperforming traditional methods in sparse datasets.
Sparse subnetworks in relational GNNs can maintain expressivity, challenging the notion that larger models are always necessary for performance.
A new graphical notation reveals the inner workings of interpretable AI architectures, translating complex designs into clear, reproducible PyTorch code.
A modified DMRG method outperforms gradient descent in optimizing tensor networks for quantum state representation, revealing new potential in machine learning applications.
GraphK can generate graphs with variable sizes while maintaining structural integrity, outperforming traditional methods in both accuracy and efficiency.
Distinct neural architectures exhibit surprisingly similar collective dynamics, revealing a universal infrared organization that transcends their microscopic differences.
FlashAttention-V achieves up to 42x speedup in transformer inference on CPUs, transforming how we leverage vector architectures for small language models.
NanoSleep achieves superior sleep stage classification accuracy while being small enough for deployment on wearable devices, striking a crucial balance between performance and efficiency.
WhiteMatter achieves superior performance with fewer resources by allowing each attention layer to dynamically access all previous layer representations, challenging traditional fixed connection patterns.
SparsePR cuts attention-reconstruction error while speeding up video generation by over 2.5x without sacrificing quality.
Achieving over 2.5% better segmentation performance than existing methods, OptiModNet does so with a fraction of the computational cost.
Achieving state-of-the-art performance in lightweight semantic segmentation, SiConMo reveals that simplicity in design can outperform complex architectures.
APEX achieves ANN-equivalent accuracy with 40% energy savings, revolutionizing the efficiency of Spiking Neural Networks in practical applications.
Achieving up to 2.3x throughput gains, HYDRA reveals that co-designing architecture and runtime policies is essential for optimizing hybrid LLM workloads on chiplet systems.
A disciplined, iterative verification process can prevent costly post-silicon bugs and ensure complex processors hit performance targets.
DPA4C achieves quantum-trained accuracy at speeds previously reserved for empirical potentials, revolutionizing molecular dynamics simulations.
The mapping of non-maximal probabilities to GMM components significantly influences the performance of S-JEPA encoders, revealing that structure matters as much as values in representation learning.
Achieving an 80% reduction in delay estimation error while being over 300 times smaller and twice as fast than existing methods could revolutionize pre-route analysis in IC design.
Selective re-scanning in recurrent networks can drastically reduce memory usage while improving task performance, challenging the conventional wisdom of fixed-size state fidelity.
Harnessing the unique signals from Mixture-of-Experts architectures, InnerExpert achieves unprecedented accuracy in detecting hallucinations at the token level.
Current AI agents can match human performance in some tasks, but they largely recycle existing human-designed algorithms rather than creating novel solutions.
CORAM achieves up to 1.35 points of performance improvement in model merging, redefining the boundaries of orthogonal transformations in AI.
MoFE not only mitigates phase-lag effects in cryptocurrency forecasting but also translates its predictive edge into substantial trading profits.
J64 reveals hidden reasoning states that can significantly boost model accuracy and decision-making, while R64 provides a lightweight, effective proxy for deployment.
Q-Interference reveals that phase-aware attention can significantly enhance token interactions without the memory overhead of traditional methods.
Recirculation enables foundation models to achieve a 23% reduction in perplexity and a 21% increase in accuracy without any added generation latency.
The minimax risk in deep Gaussian regression exhibits a surprising quadratic dependence on depth, challenging conventional assumptions about model capacity.
FESC achieves private long-document inference on a single GPU for sequences up to 2,048 tokens, setting a new standard for efficiency and accuracy in encrypted machine learning.
Coiflet wavelet convolutions can outperform Haar in efficiency, slashing parameters and FLOPs while maintaining competitive accuracy.
Achieving 87.35% accuracy on CIFAR10-DVS in only 10 inference steps reveals a breakthrough in training efficiency for spiking neural networks.
Achieving a remarkable 98.25% accuracy in retinal disease detection, RetiWave-Mamba redefines the standards for automated OCT analysis.
The interaction between selection and aggregation in sparse attention networks can lead to super-resolution gains that are greater than the sum of their parts, challenging conventional wisdom in model design.
MoNe slashes compute and memory costs by 80% for long-context inference while enabling Transformers to handle context lengths far beyond their original limits.
GADR reveals that even in chaotic meetings, a structured approach can yield clear and actionable architectural decisions, outperforming conventional methods.
NeuroAbs accelerates hardware verification by intelligently abstracting RTL designs, achieving significant efficiency gains over traditional approaches.
Achieving 11.8× faster inference and 81× faster adaptation on edge FPGAs, MAGMA revolutionizes GMM performance in dynamic environments.
Reducing NoC congestion by over 95% in FPGA designs could revolutionize how we approach chip integration and performance optimization.
Despite executing target instructions, LLMs often fail to deliver competitive performance in GPU kernel optimization, especially on complex tasks.
Fine-grained MoE designs can outperform dense vision encoders while dramatically reducing latency, challenging the status quo in image and video understanding.
Traditional accuracy metrics can obscure critical improvements in reasoning fluency, as shown by fine-tuned models that reason correctly in low-resource languages despite initial benchmark null results.
TileMix achieves a breakthrough in LLM inference by enabling mixed-precision attention that boosts throughput while maintaining long-context quality.
Palmyra x6 outperforms previous models in enterprise agentic tasks while maintaining a strong safety profile and low bias.
Nexus achieves high-resolution text-to-image generation with a fraction of the computational cost, rivaling leading models in quality.
PCT-Prompt transforms dense prediction in point clouds by integrating prompt-guided features, leading to substantial performance gains over standard Transformers.
A unified layer equation for GNNs reveals how over 200 architectures can be systematically compared and optimized for performance.
Achieving competitive accuracy with fewer parameters, Self-Routed Tensor Adapters redefine the landscape of visual model adaptation across heterogeneous domains.
Lifelong learning just got a major upgrade: SoftModel's dynamic topology adapts in real-time, breaking free from the constraints of fixed neural architectures.
AsyTO achieves state-of-the-art forecasting accuracy while keeping model complexity linear, challenging the notion that more parameters always lead to better performance.
Occlusion-aware gating in SIGMA-Lane significantly improves temporal consistency in video lane detection, outperforming traditional methods under heavy occlusion.
Achieving over 99% accuracy in DDoS detection while maintaining sub-millisecond inference times could revolutionize security in operational technology networks.
Achieving high-quality video generation with a staggering \(\sim67\times\) reduction in computational costs could revolutionize the efficiency of video diffusion models.
Static memory models are holding back performance—Proteus reveals that incrementally activating memory can drastically enhance context retention and reduce interference.
Activation-state transfer between LLMs is effective but only works under specific architectural conditions, revealing a surprising limitation in cross-model communication.
GoalEvolve achieves a 30.67% improvement in post-route TNS while simultaneously reducing leakage and dynamic power, showcasing a transformative approach to physical design algorithm evolution.
Energy optimization in circuit synthesis can be dramatically improved with Renesis, achieving up to 91% of the default energy across various benchmarks.
Quantum-safe web services can be achieved using TOTPs, ensuring robust security against future quantum threats.
FreeToken transforms personal machines into powerful platforms for running massive AI models, enabling users to deploy frontier-scale intelligence without specialized infrastructure.
Microservices can outperform monolithic architectures under heavy load, but they come with new challenges in failure management.
Achieving a wavelength-hierarchy correlation of 0.852, QuantumPhaseNet outperforms classical models, but reveals no quantum advantage in overall efficiency.
Routing changes in MoE models may not influence behavior as expected, with the routing term contributing less than half of the natural context effect.
Reordering attention and feed-forward operations in LLMs can significantly enhance computational efficiency without sacrificing performance.
DeltaLog slashes recurrent-state write traffic by up to 7.83x while boosting decoding speed, revolutionizing how linear attention models manage memory.
Gated Recurrent Transformers achieve 63% fewer parameters and 59% less peak memory while matching the accuracy of much larger models, redefining efficiency in language modeling.
Achieving similar performance to larger models with significantly less data and faster inference speeds could redefine efficiency benchmarks in foundation models.
Fixing deployment bugs in Nanbeige4.2-3B transforms it from a non-functional model to one that can tackle real agentic tasks with a significant performance boost.