Search papers, labs, and topics across Lattice.
89 papers published across 4 labs.
UltraViT achieves a groundbreaking 1.7x speed increase for on-device vision-language model encoding without sacrificing performance.
Achieving approximation errors of \(10^{-14}\) in quantum state preparation could revolutionize how we approach circuit optimization in quantum computing.
Polynomial-time training is now achievable for a broader class of neural networks, including those with ReLU activation and specific structural constraints.
Learning-guided strategies can slash SCM SAT encoding time by orders of magnitude while preserving near-optimal quality.
Naju achieves unprecedented long-sequence memory retention and overwriting capabilities, outperforming traditional models while ensuring linear efficiency.
UltraViT achieves a groundbreaking 1.7x speed increase for on-device vision-language model encoding without sacrificing performance.
Achieving approximation errors of \(10^{-14}\) in quantum state preparation could revolutionize how we approach circuit optimization in quantum computing.
Polynomial-time training is now achievable for a broader class of neural networks, including those with ReLU activation and specific structural constraints.
Learning-guided strategies can slash SCM SAT encoding time by orders of magnitude while preserving near-optimal quality.
Naju achieves unprecedented long-sequence memory retention and overwriting capabilities, outperforming traditional models while ensuring linear efficiency.
MemTools transforms agent memory research by enabling seamless integration and evaluation of diverse memory types across architectures.
Federated Cognitive Digital Twins can revolutionize decision-making in smart cities by distributing intelligence and enhancing real-time responsiveness.
Equivariant networks can now achieve superior performance across accuracy, efficiency, and speed, challenging the dominance of non-equivariant architectures.
Achieving a 5% reduction in wirelength compared to leading floorplanners could redefine efficiency in VLSI design.
DMG achieves a remarkable 4.9X performance boost while slashing compute-side cache requirements by nearly 19X, redefining efficiency in graph processing systems.
Achieving softmax-level expressiveness at a fraction of the cost, SANA-Video 2.0 is 120x faster than its closest competitor while generating high-quality video.
KroQuant achieves superior output quality in post-training quantization of diffusion transformers while being up to 14% faster than conventional methods.
Agentic Designer achieves unprecedented structural adherence in interior layouts by employing a multi-agent system that iteratively verifies geometric constraints before placement.
TF-MossFormer outperforms existing speech separation models by effectively balancing local and global context through innovative attention mechanisms.
Sinusoidal recurrence can achieve higher fidelity in image representation with fewer resources, outperforming traditional feed-forward models.
ELSAA achieves efficient attention approximation, allowing Transformers to handle longer inputs without sacrificing interaction quality.
The Receptron model achieves competitive accuracy on edge devices while sidestepping the complexities of multi-layer networks, making it a game-changer for IoT applications.
Hybridizing contrastive learning with selective state space models accelerates convergence and enhances explainability in user-centric transactional modeling.
Treating AI models as a single, scalable architecture may be a fundamental mistake, as distinct cognitive tasks require qualitatively different structures for optimal performance.
Embedding scattering mechanisms into complex-valued networks can significantly enhance the accuracy and consistency of PolSAR image classification.
Solar Open 2 outperforms its predecessors and competitors with a groundbreaking 1M-token context window, redefining the capabilities of large language models in agentic tasks.
MoAKE achieves superior action quality assessment by unifying diverse action evaluations into a single model, overcoming the limitations of traditional one-by-one approaches.
ESSDs can significantly outperform local SSDs, but only if software adapts to their unique performance characteristics.
Unveiling the hidden NUMA architecture of GPUs could revolutionize how we optimize memory efficiency in high-performance computing.
SHFormer achieves a remarkable 1 dB improvement in MRI reconstruction quality by effectively capturing high-frequency details across diverse modalities.
Achieving a 71.81% zero-shot Pass@1 score, the DeepSeek-V4-Flash model outperforms leading competitors by leveraging a novel optimization framework on Ascend SuperPOD.
Emergent task decomposition in robot policies reveals that learned experts can specialize and be reused, challenging traditional notions of task hierarchy in AI.
Causal emergence in active inference agents hinges on architectural design, with the global latent $g$ revealing surprising dynamics that challenge conventional interpretations of integrated information.
Modern FPGA implementations of hardware priority queues can drastically outperform traditional software methods, revealing critical insights into architectural relevance.
Cumsum-composable phase transport enables high-accuracy keyword spotting with a fraction of the model size and training time, revolutionizing streaming speech applications.
Multi-Head Attention Residuals achieve superior validation loss by allowing Transformers to leverage multiple attention heads, revealing that subspace disagreement is a key factor in model performance.
Generalizing batch normalization to Lie groups could revolutionize how we train neural networks on manifold-valued data.
Staleness-Adaptive Trust Regions reshape update geometry in asynchronous reinforcement learning, achieving record performance while controlling for high-staleness updates.
NMLNs can now outperform diffusion-based models on large graphs, thanks to a novel parallel noising algorithm and enhanced expressive capacity.
Achieving robust neural network performance on non-Euclidean spaces could redefine stability benchmarks in machine learning applications.
Sparse observations can enhance ocean modeling performance, challenging the reliance on complete datasets that limit model capabilities.
Spectral Higher-Order Neural Networks can drastically reduce computational costs while improving performance on notoriously difficult tasks like N-bit parity.
Selective adaptation in clinical AI can enhance prediction accuracy for critical interventions while preserving the integrity of patient physiology.
A novel gating architecture that balances cost and quality in SFT data procurement, achieving near-optimal accuracy while minimizing expenses.
A vast number of functionally equivalent neural networks can exhibit striking geometric diversity, challenging assumptions about model uniqueness in approximation tasks.
A renal-inspired architecture achieves a four-fold concentration increase, challenging conventional iterative refinement methods in neural networks.
Relative positional encodings not only enable extrapolation in transformers but also reveal a profound connection between implicit bias and sequence length generalization.
SCM achieves an impressive 84.87% accuracy on conversational memory tasks, showcasing a novel approach to optimizing agent memory retrieval and synthesis.
Encoding spatial relationships in Transformers can significantly boost performance in complex routing problems, achieving better solutions than conventional methods.
Reciprocal cross-stream addressing in Dual Attention Residuals leads to significant improvements in Transformer performance without sacrificing depth-wise diversity.
Compiled pipelines using a novel visibility control abstraction outperform traditional HLS tools while matching the efficiency of hand-written RTL designs.
The verification framework reveals that excess out-of-order executions can be systematically managed, paving the way for robust multiprocessor designs.
IMMoE boosts anomaly detection performance by over 11% even when critical view information is missing, redefining expectations for real-world applications.
Achieving shape modeling with under 100 parameters while allowing for zero-shot user editing could revolutionize how we deploy AI in mobile and AR applications.
VMamba's unique encoding strategy enables it to outperform MambaOut in high-resolution tasks by leveraging the organization of semantic evidence across token magnitude and direction.
LANav outperforms traditional Transformer-based navigation policies by 6.3 percentage points, revealing the power of structured state updates in complex environments.
Unstructured grid codes can outperform GPUs on Spatial Dataflow Architectures, challenging conventional assumptions about hardware limitations.
Flipping low-order bits in DNNs may not matter, but a single flip at the exponent-mantissa boundary can lead to catastrophic failures—this paper reveals how to protect against that efficiently.
Achieving up to 2.64x speedup in Mixture-of-Experts execution by cleverly overlapping computation and communication could redefine efficiency benchmarks in large-scale AI models.
Selective state-space adaptation can boost reasoning accuracy in language models by over 18% on challenging tasks, revealing the power of dynamic context management.
Exact Coulomb friction can be solved with unprecedented accuracy and robustness by separating linear responses from non-associated couplings, revolutionizing frictional contact dynamics.
Optimizing for realizability in MoE systems leads to tangible improvements in deployment efficiency, with moefs achieving a 0.9% edge over hand-tuned plans in training.
SparHiXcel-v2 achieves over 2.5 TOPS and 210 GOP/s/W on VGG16, striking a remarkable balance between sparsity flexibility and hardware efficiency.
BaseRT achieves up to 6.4x faster prompt processing for LLMs on Apple Silicon, redefining performance expectations for on-device AI inference.
Memory-efficient optimizer state allocation can lead to dramatic improvements in training performance without sacrificing accuracy, as shown by SkewAdam's superior perplexity results.
Online adaptation of neural state-space models can now be achieved with high accuracy and computational efficiency, filling a critical gap in real-time system identification.
Fixing the residual mixing matrix to identity during finetuning can lead to unexpected performance improvements in Transformer models.
Substantial information processing capacity in reservoir computing may be hidden in low-energy modes, making them vulnerable to noise and potentially undermining performance.
Cyclic depth folding enables Transformer blocks to excel in both shallow and deep roles, achieving better performance than traditional fixed-order training.
AMO outperforms traditional PDE solvers by leveraging a novel integration of reproducing kernels and adaptive Fourier decomposition, revolutionizing the approach to solving complex equations across varied geometries.
Local, sparse learning can outperform backpropagation in resisting catastrophic forgetting, achieving up to 19 times better backward transfer in text domain tasks.
Halving parameters while ensuring gradient stability, MRSNorm redefines normalization by preserving phase properties and preventing numerical explosion in deep networks.
Task-specific transformer architectures can dramatically boost learning efficiency but often at the cost of versatility, challenging the notion that one design fits all.
Jointly pruning dependent units in LLMs can lead to more effective model compression without the need for fine-tuning, challenging conventional independent pruning methods.
Identical embeddings for non-isomorphic graphs reveal the limitations of global-attention graph transformers in capturing the nuanced structure of mixed-integer linear programs.
VNVSpec transforms high-level user requirements into actionable, machine-readable specifications that can be directly linked to test results, addressing a critical gap in software verification.
Hyperbolic geometry at the loss layer can stabilize training in expert AI models, overcoming the pitfalls of traditional methods that dilute structural integrity.
Attention-only transformers can match standard architectures in performance when resource allocation favors attention depth, challenging assumptions about the necessity of feed-forward layers.
Adding depthwise convolutions to Transformers can boost accuracy on downstream tasks while barely increasing model size.
L1 augmented attention achieves a remarkable 14.5% reduction in perplexity by integrating L1 geometry into Transformer models, challenging the dominance of traditional dot product methods.
A robust architecture for agentic commerce that ensures transaction integrity and security, even in the face of state changes and unauthorized access.
Architectural choices, not model identity, drive performance differences in Geospatial Foundation Models, challenging conventional ranking methods.
DuSPiT generates images with unprecedented detail and structure by decoupling global reasoning from local appearance in a dual-branch architecture.
SEAM-V's innovative architecture can boost vector processing speeds by up to 3x, revolutionizing performance in data-parallel workloads.
ExpertPlex slashes over 95% of duplicate model weights and boosts goodput by up to 2.01 times, transforming how we serve large language models.
Custom reduced-precision floating-point formats can outperform fixed-point designs in FFT implementations, cutting power by nearly 20% without sacrificing performance.
Long-time behaviors in deep linear transformers can reveal surprising dynamics like clustering and oscillations, all rooted in a hidden Hamiltonian structure.
Logic-based neural architectures can outperform traditional models in EEG classification while dramatically reducing latency and memory usage.
OrderMoE slashes average latency and cross-server traffic by intelligently grouping experts based on functional similarity, achieving efficient edge inference without sacrificing quality.
ThAME achieves a staggering 15.7x speedup in MoE inference, revolutionizing the efficiency of Large Language Models.
ThRIve achieves thermally robust CNN inference with a remarkable 5.4x improvement in energy efficiency while maintaining accuracy within 2% of ideal performance.
Reducing shift overhead by nearly 50% in RTM-based caches could revolutionize energy efficiency in data-centric applications.
Fine-tuning CNNs on PIM architectures can be both energy-efficient and accurate, with ADEPT cutting off-chip memory access while maintaining performance.
Analyzing Transformers through continuous stochastic differential geometry reveals surprising predictive insights into their stability limits and optimization dynamics.