Search papers, labs, and topics across Lattice.
89 papers published across 5 labs.
Inverted self-attention reveals hidden causal relationships in multivariate time series, outperforming traditional methods by reducing spurious correlations.
Raven achieves superior long-context recall by intelligently routing memory updates, outperforming traditional models that struggle with interference and eviction.
ReBA achieves over fivefold improvement in load balancing for Vision-Language MoE without sacrificing accuracy, revealing a critical interplay between token mix and load profiles.
Achieving 1,500 tokens per second, DiffusionGemma redefines the speed-capability trade-off in language models, outpacing conventional autoregressive approaches.
SliGFM achieves a breakthrough in graph learning by ensuring that heterogeneous features are not only unified but also semantically rich and transferable across domains.
ReBA achieves over fivefold improvement in load balancing for Vision-Language MoE without sacrificing accuracy, revealing a critical interplay between token mix and load profiles.
Achieving 1,500 tokens per second, DiffusionGemma redefines the speed-capability trade-off in language models, outpacing conventional autoregressive approaches.
SliGFM achieves a breakthrough in graph learning by ensuring that heterogeneous features are not only unified but also semantically rich and transferable across domains.
S-CEReBrO processes EEG signals 100 times longer than traditional methods while slashing memory usage by 55%.
Sparsifying FFNs can lead to nearly double the decoding speed without sacrificing model quality, thanks to a novel channel-selection strategy.
MSCM-net achieves superior hyperspectral image classification by leveraging multi-scale convolution and Mamba, striking a balance between accuracy and efficiency.
CoRE-UIR achieves a remarkable 1.05 dB PSNR improvement while being 11.83 times faster than the leading image restoration method, BaryIR.
Simplifying deep networks during training can lead to substantial parameter reductions while preserving accuracy, challenging the notion that bigger always means better.
SPFM-Net achieves a breakthrough in invisible watermark attacks, balancing effective signal removal with high visual fidelity through innovative semantic and frequency-guided techniques.
The fastest operational rates in analog computing are fundamentally constrained by the largest combined unity-gain bandwidth, revealing a critical limit on performance that mirrors numerical method restrictions.
A novel architecture enables seamless integration of Edge, Fog, and Cloud resources, transforming how researchers can experiment with distributed Cyber-Physical Systems.
CCFormer delivers a 3.57% increase in click-through rates and a 1.71% boost in advertising revenue, all while cutting model training time by over 2x.
TopoFormer achieves state-of-the-art performance in graph learning by seamlessly integrating topological structures into attention-based models, outperforming traditional methods.
Inverted self-attention reveals hidden causal relationships in multivariate time series, outperforming traditional methods by reducing spurious correlations.
Identical tokens can mask significant divergence in autoregressive states, revealing that numerical compatibility in sparse MoE models is more complex than previously understood.
Transforming passive nanoparticle networks into tunable nonlinear systems reveals new pathways to enhance computational efficiency in neuromorphic computing.
Few-shot prompting boosts LLM performance in generating microservice architectures, achieving an impressive F1 score of 0.97 for service identification.
CoMem achieves a 7.83x prefill speedup and drastically reduces memory usage while maintaining high performance on long-context tasks, challenging conventional memory management in LLMs.
SCSE transforms recurrent computation in Looped Transformers by ensuring that deviations from a learned anchor enhance performance without compromising stability.
Exact deletion from language-model memory can be achieved through clever memory representation, allowing models to effectively manage and amend their learned records.
ReTopK achieves a 3.07x speedup in attention computation with only a 0.50% increase in perplexity, revolutionizing long-context processing efficiency.
Securing critical logic paths without exposing sensitive parameters could redefine the landscape of logic-locking in SoC designs.
A novel gating mechanism that integrates contextual information from both encoder and decoder dramatically boosts performance on minority class detection in nucleus segmentation tasks.
Expert subspaces in MoE models overlap significantly, yet selected routes yield better token representations than their strongest unselected counterparts, challenging traditional views on redundancy.
Finite precision in transformers can drastically alter memory operations, revealing a surprising hierarchy of expressivity based on attention mechanisms.
Surrogate-guided ensemble search can identify architectures that are not only strong individually but also diverse, outperforming traditional methods on benchmark datasets.
BATS achieves a remarkable 53% reduction in GPU memory usage while maintaining competitive segmentation accuracy, revolutionizing resource-efficient volumetric segmentation.
An attention-based model for Alzheimer's classification achieves 88.95% accuracy by directly analyzing resting-state fMRI data, bypassing traditional feature engineering pitfalls.
Sparsity in spiking networks isn't a given; it heavily depends on the task, with recurrent models struggling to drop below 50% activity without sacrificing quality.
CMP's innovative architecture cuts catastrophic forgetting in continual learning by leveraging sparse representations and local updates, outperforming traditional Transformer models.
Experts in federated learning can now achieve specialization without sacrificing performance across heterogeneous tasks, thanks to FedWeave's innovative asymmetric aggregation strategy.
GRU decoders outperform Mamba hybrids in brain-to-text systems, revealing critical insights into the interaction between output targets and model architecture.
Emergent attribution in B1ade-1B shows that grounding behavior can arise from RL training without explicit supervision, achieving a citation rate that surpasses its training distribution.
KinRT achieves over 23% improvement in expert routing accuracy by leveraging kinematic archetypes, transforming how MoE systems can operate without kinematic signals at inference.
Reservoir computing can revolutionize branch prediction in CPUs, but it currently lags 15x behind established methods in adaptability.
A memory-leak detector that blocks harmful resource recommendations could revolutionize how Kubernetes manages container workloads, preventing costly over-provisioning.
Achieving a foundational link between ISA leakage contracts and cycle-level execution, Granite eliminates the need for intermediate specifications in verifying cryptographic confidentiality.
GARI enables a flexible, learnable interface that maintains transformation consistency across diverse data types, challenging the rigidity of traditional equivariant architectures.
WALoMA outperforms traditional models by achieving an 87.80% composite score while using just 14.68% of the parameters, redefining efficiency in wireless task performance.
Spectral descriptors can effectively guide CNN architecture design, leading to significant performance improvements in NIR chemometrics.
Raven achieves superior long-context recall by intelligently routing memory updates, outperforming traditional models that struggle with interference and eviction.
CondPSE achieves a remarkable leap in synthetic graph structural discrimination, raising accuracy from 42.9% to 97.3%, but struggles to translate this advantage into real-world applications.
Logarithmic-depth networks can efficiently learn complex Boolean functions that constant-depth networks cannot, revealing a critical algorithmic separation in neural network capabilities.
A novel spectral framework reveals that while internal coherence maintains task relevance, it does not guarantee successful execution transitions in transformer models.
Factor gradient flow in positive quadratic networks mirrors Riemannian dynamics, revealing surprising insights into curvature and recovery that challenge conventional training assumptions.
CoSA achieves nearly 5× faster attention computation without sacrificing accuracy, revolutionizing long-context processing in LLMs.
A novel balanced soft mixture-of-expert model achieves unprecedented accuracy in glaucoma detection, outpacing existing multi-modal and uni-modal approaches.
Activating experts based on token uncertainty allows CARE to outperform traditional fixed routing methods while using fewer resources.
Operating four automated trucks with a single driver can slash freight costs by over 56%, revolutionizing the economics of long-haul trucking.
Modern software architecture is evolving to prioritize continuous governance and AI-assisted decision-making, yet critical research gaps in security and empirical validation remain unaddressed.
Traditional round-robin scheduling in MoE architectures can lead to an exponential incast problem, but a new proactive scheduling framework effectively eliminates this bottleneck.
Route-block interventions can drastically alter model outputs, revealing hidden dependencies in expert alignment that challenge conventional assumptions about MoE preprocessing.
Achieving 63.6% power savings and 40.4% area reduction, MDTransformer redefines efficiency in photonic transformer accelerators without sacrificing performance.
Ventaglio accelerates sparse tensor contractions by up to 7.4 times, pushing performance limits closer to theoretical roofline bounds.
A unified taxonomy reveals how diverse memory mechanisms in LLMs can be systematically understood and leveraged for future innovations.
Achieving over 2x speedups in video generation without sacrificing visual quality, Sol-Attn redefines the efficiency of sparse attention mechanisms.
Kimi K3's innovative architecture achieves a 2.5x scaling efficiency improvement, enabling robust performance across diverse long-horizon tasks.
MMOE achieves faster convergence and superior generation quality in diffusion transformers by effectively integrating expert routing strategies, challenging the notion that more parameters always lead to better performance.
CAEs may compress data better, but they compromise long-term forecast stability, revealing a crucial trade-off for real-time flow control applications.
Gated feature reweighting in FPGA implementations increases hardware complexity without enhancing classification accuracy, challenging the assumption that more sophisticated models always yield better performance.
Capturing fine-scale physical structures in PDE predictions is now achievable with a novel function-projection approach that outperforms traditional methods.
The study reveals that energy distribution across frequency planes can create unexpected consensus dynamics, defying traditional assumptions in self-attention models.
DraftExpert achieves a 1.45x boost in decoding throughput while maintaining high draft acceptance rates, revolutionizing MoE inference for end-device applications.
Quantum models fail to show any advantage over classical counterparts in time-series forecasting, challenging assumptions about their superiority.
SpecFormer transforms recommendation systems by mitigating embedding and attention collapse, leading to superior performance and scalability.
LLM-generated shuttling compilers can cut development time from months to days while achieving superior performance compared to traditional hand-crafted solutions.
PIVOT achieves up to 4x faster indexing for token-level sparse attention without sacrificing accuracy, transforming how we handle query processing in large models.
Clustering languages into groups allows MoLGE to outperform traditional multilingual ASR models while keeping the parameter count low.
Prompt-based adaptation can significantly enhance the merging of specialized models, yielding better performance without the pitfalls of traditional weight merging.
Achieving over 95% mIoU with just 1.175M parameters, LCMamNet sets a new standard for lightweight infrared small target detection.
ERF organization, not just scale, is the key to unlocking superior performance in infrared small target detection.
Bi-encoder performance hinges on model variant, while cross-encoders maintain a robust edge by jointly encoding record pairs, revealing critical insights for entity matching strategies.
Achieving over 99% accuracy in RF signal recognition tasks while maintaining a latency of just 98 µs per frame could revolutionize spectrum intelligence applications on edge devices.
A simple yet powerful adaptation of attention for triangle meshes outperforms existing methods, achieving state-of-the-art results in geometry-processing tasks.
Achieving up to 4.5 TOPS while consuming less than 250 mW, the SpiNNaker2 chip redefines energy efficiency in brain-inspired computing.
Ignoring dynamic signal integrity effects can skew inter-chiplet performance predictions by a significant margin, revealing critical flaws in current simulation practices.
Retention loss doesn't have to mean accuracy loss; innovative compensation techniques can recover nearly all inference performance even after prolonged storage.
Evolving the source code of FPGA placement and routing tools can yield up to 2.7% better performance than conventional hyperparameter tuning.
Memory technology can alter execution time by over an order of magnitude, revealing that the best host memory isn't always the best PIM substrate.
Learning kernelized attention through a Coulomb particle model boosts performance without sacrificing the efficiency of linear attention.
Achieving machine-checked equivalence across multiple representations of floating-point arithmetic could redefine standards for hardware verification in AI systems.
Achieving 16x faster-than-real-time audio synthesis with a memory footprint of just 21 MB, this architecture redefines on-device speech synthesis capabilities.
Sketching at just 25% completion can yield better spatial rationality than fully specified baselines in 3D scene generation.
ATLAS slashes inference latency for transformer models under FHE by tailoring approximation settings layer-by-layer, rather than relying on rigid uniform configurations.
Detecting dormant hardware Trojans with 90.94% accuracy could redefine security protocols in integrated circuits by addressing vulnerabilities that remain hidden until triggered.
Semref leverages LLMs to automatically refine architecture recovery results, achieving up to 118.57% improvement in accuracy over traditional methods.
Tailored attention mechanisms can outperform traditional softmax approaches by leveraging structured interactions, as demonstrated by VIA's success in predicting retrosynthesis reaction centers.
Warp divergence incurs a predictable performance cost across NVIDIA GPU architectures, even as reconvergence mechanisms evolve dramatically.
UltraViT achieves a groundbreaking 1.7x speed increase for on-device vision-language model encoding without sacrificing performance.