Search papers, labs, and topics across Lattice.
100 papers published across 1 lab.
Achieving an 11% boost in energy efficiency for edge AI applications by smartly balancing throughput and latency through hierarchical operator parallelism.
LevelSyn achieves a remarkable 6.89% power reduction and 27.48% timing improvement by integrating physical awareness into logic synthesis, revolutionizing design closure efficiency.
Predicting CPU workload from customer service requests can improve forecasting accuracy and operational efficiency in cloud environments.
Achieving up to 9.9 times the throughput of traditional methods, this GPU-accelerated solver transforms the efficiency of constant optimization in symbolic regression.
Causal metric learning in CauseCollab dramatically improves semantic consistency across heterogeneous modalities, leading to state-of-the-art performance in collaborative perception tasks.
Achieving an 11% boost in energy efficiency for edge AI applications by smartly balancing throughput and latency through hierarchical operator parallelism.
LevelSyn achieves a remarkable 6.89% power reduction and 27.48% timing improvement by integrating physical awareness into logic synthesis, revolutionizing design closure efficiency.
Predicting CPU workload from customer service requests can improve forecasting accuracy and operational efficiency in cloud environments.
Achieving up to 9.9 times the throughput of traditional methods, this GPU-accelerated solver transforms the efficiency of constant optimization in symbolic regression.
Causal metric learning in CauseCollab dramatically improves semantic consistency across heterogeneous modalities, leading to state-of-the-art performance in collaborative perception tasks.
A novel deadline-aware metric and a hybrid reinforcement learning architecture dramatically improve network control efficiency, enabling timely delivery of critical information in dynamic environments.
Trust in AI-generated cryptographic artifacts can be achieved without relying on the authorship of the AI, as demonstrated by a zero-failure rate in defect detection.
NACRE achieves secure containerization without sacrificing performance, maintaining syscall metrics within 3.5% of the baseline while enhancing confidentiality.
Iapetus achieves a remarkable 91.6% task completion rate while slashing mean latency and battery draw by over 50% in satellite-based ViT inference.
Optimizing GPU resource allocation can cut multi-agent workflow completion times by nearly 37% while saving substantial GPU usage.
No existing distributed authorization design can simultaneously ensure native-signature compatibility, unilateral-signing resistance, and threshold-layer agility, revealing a critical architectural tension in post-quantum systems.
Achieving over 2x throughput improvements in attention mechanisms could redefine efficiency benchmarks for large-scale model training.
Current AI methods for data center optimization overlap significantly, making it impossible to rank their effectiveness—CLEAR-DC aims to change that.
Privacy noise in federated learning can severely compromise the detection of rare attacks, challenging the assumption that privacy and robustness can be optimized independently.
Achieving retrieval speeds of up to 31.5 times faster than prior methods, Spruce redefines the landscape of private outsourced retrieval for large document collections.
Causal discovery in federated settings can achieve high accuracy without compromising privacy, even in the presence of noise, by leveraging higher-order cumulants.
WebXR could be the key to a more accessible and sustainable Metaverse, challenging the dominance of commercial game engines.
Barnacle reduces transaction latency in DAG-based consensus by dynamically adjusting leader counts, outperforming static configurations even in degraded network conditions.
sp-DBA accelerates transform-domain computations by up to 28.1 times, revolutionizing how we approach scientific simulations on parallel hardware.
Einsummable achieves a remarkable 35% speedup over hand-tuned PyTorch implementations by automatically optimizing multi-GPU computation distribution without manual intervention.
The admission gate in cache prefetching outperforms the predictor, revealing that better proxies do not guarantee improved system performance.
RASER slashes workflow makespan by 39% while ensuring resilience and high CPU utilization in HPC environments, all without needing kernel access.
JuPyLive transforms Jupyter notebooks into a bridge for effortless migration between local and HPC resources, enabling researchers to scale their workflows with just one click.
Lantern's innovative approach allows for a dramatic throughput increase in transaction processing without the need for prior knowledge, challenging conventional concurrency control limitations.
Achieving an 80% reduction in low-latency VM offloading while managing only an 11.97% worst-case latency increase could revolutionize how we integrate low-latency applications into renewable energy systems.
Cloud computing's carbon footprint could be significantly reduced without sacrificing latency, but existing strategies fall short of this goal.
BU-MBAR achieves faster and more stable convergence for MBAR equations by dynamically adjusting bin widths, revolutionizing statistical analysis in thermodynamics.
Low-level accelerator optimization is yielding to agentic search: closed-loop LLMs driven by hardware profiling can now match human experts at generating high-performance TPU kernels.
Decoy activations fail to protect split-LLM training because zero-valued backward gradients unmask private frames with 100% accuracy, rendering forward-only privacy audits dangerously deceptive.
GRADSOLVE achieves up to 14.1x faster gradient computations for ODE ensembles, revolutionizing the speed-accuracy trade-off in GPU-based solvers.
Achieving 2.27x faster decoding for 2-bit LLM weights could redefine efficiency benchmarks in low-bit quantization methods.
SAPE-FL achieves superior model performance in heterogeneous environments by intelligently balancing global and local learning through similarity-aware personalization.
Reinforcement learning can cut weather forecast errors by up to 45% while maintaining numerical stability in operational models.
Federated LoRA adaptation can boost chest X-ray classification performance across diverse international cohorts, achieving near-centralized model performance without compromising data privacy.
Optimizing inspection strategies in reverse logistics can yield significant financial gains, with one method adding nearly $54k per batch in aircraft maintenance alone.
Federated learning can now thrive in privacy-sensitive scientific collaborations, unlocking new avenues for public-private partnerships in AI model development.
Exponential triggering delays can significantly alter packet stream dynamics, revealing performance metrics that classical models overlook.
Achieving speedups of up to 100x in nonbonded force evaluation could revolutionize large-scale molecular dynamics simulations.
Privacy and robustness in federated learning are not just additive; they interact in ways that can significantly impact model performance in adversarial settings.
Automating enclave partitioning can cut host-enclave transitions in OpenSSL by over 50%, streamlining secure application development.
Achieving 78,732 feasible configurations with a scheduling accuracy within 0.10% of exhaustive search, this method revolutionizes GPU resource allocation for concurrent AI workloads.
Achieving 18.5x faster city-scale 3D rendering on mobile VR devices while slashing energy consumption by 92.4% could redefine immersive experiences in urban environments.
Federated meta-analysis can yield accurate GWAS results without ever centralizing sensitive genotype data.
Achieving over 2x energy efficiency in analog vision processing could redefine the performance benchmarks for future AI hardware.
BASP slashes communication overhead in LLM training, boosting efficiency by over 30% while maintaining model performance.
Skywing enables resilient decentralized mathematical computations even in unreliable environments, challenging the limitations of traditional high-performance computing frameworks.
Achieving fully fluctuating participation in sleepy consensus without relying on heavy cryptographic primitives like VDFs could revolutionize blockchain protocols.
AceSpec transforms edge-cloud LLM inference by turning catastrophic network stalls into efficient local memory lookups, achieving a remarkable 3.52× speedup even in low-bandwidth environments.
Containers can now pinpoint their CPU contention culprits in real-time without the need for kernel patches or costly tracing.
RACE-AIMC cuts energy use by 69% while maintaining accuracy comparable to digital systems by intelligently selecting the best AIMC accelerator and certifying its reliability.
Achieving 2.3x the power-normalized throughput of a high-performance processor, this FPGA design revolutionizes real-time adaptive beamforming for ultrasound applications.
RT-HiSS achieves unprecedented speedups for high-dimensional similarity searches, revolutionizing the efficiency of GPU-based algorithms.
BBYT slashes candidate-selection time by over 18% while maintaining the accuracy of timing-aware logic rewrites, revolutionizing how we approach timing evaluations in circuit design.
As conversation lengths increase, LLM systems can achieve a hit ratio that converges to a predictable limit, revolutionizing memory management strategies.
Achieving up to 65.5% parameter reduction and nearly 2x inference speedup for 3D point cloud models without needing source code is a game-changer for edge deployment.
The ARFT dataset reveals that synchronized acoustic and RF data can significantly enhance positioning accuracy in complex environments.
An honest-but-curious server can identify clients with near-perfect accuracy from federated learning updates, exposing critical privacy vulnerabilities in vehicular networks.
Federated learning can outperform centralized models in multi-agent safety without compromising data privacy, achieving a remarkable 43% reduction in attack success rates.
A lightweight security layer for TRDP communications enables robust cryptographic protection without compromising real-time performance.
Label-free transfer in blockchain compliance screening achieves up to 99.67% recall on held-out positives, revolutionizing how unlabelled addresses are assessed across multiple chains.
Strategic placement of Byzantine nodes can exponentially increase their impact on honest participants, revealing a critical vulnerability in decentralized federated learning systems.
FPGAs could revolutionize Transformer model deployment by offering superior energy efficiency and latency compared to traditional hardware.
ModalShare reallocates bandwidth based on each modality's contribution, boosting accuracy by over 15 percentage points compared to traditional equal keep-ratios.
Outperforming prior open models, Instella-MoE sets a new standard for efficiency and performance in language modeling with its innovative architecture and training methods.
Byzantine manipulation can no longer compromise honest sensors in CRSF, ensuring robust privacy and correctness in sensor fusion.
NymHS revolutionizes mixnet privacy by enabling secure hidden services that protect both sender and receiver identities, while also achieving remarkable performance improvements.
Griotte reveals that formal verification can ensure compartmentalization security in systems where cross-communication is necessary, challenging traditional assumptions about isolation.
Achieving a 95% reduction in communication overhead, CERF enables seamless collaboration among heterogeneous agents without the need for retraining.
Early denoising stages in diffusion models are up to 2.58× more vulnerable to hardware noise, but a novel approach can mitigate this without retraining.
The Leiden method can automatically generate optimal block structures for parallel preconditioning, eliminating the need for predefined parameters while matching the performance of established methods.
Achieving over 24 times better energy efficiency in LLM inference could redefine the scalability of AI applications.
DART slashes P99 sojourn times by up to 23% by intelligently navigating configuration reconfigurations and job servicing decisions.
Understanding how concurrent AI training jobs disrupt each other could redefine resource allocation strategies in supercomputing environments.
Localizing cyclic ambiguity in transaction ordering can lead to a staggering 10.5× increase in throughput while ensuring fairness in blockchain systems.
Substantial behavioral diversity in SPEC CPU 2026 reveals critical scale-dependent effects that could redefine performance optimization strategies for next-gen datacenter processors.
CoSMO achieves an 18.6% to 21.2% improvement in task completion rates by rethinking how edge nodes manage and offload tasks based on semantic state rather than mere freshness.
On-the-Fly3R achieves robust online 3D reconstruction from unordered UAV images, significantly improving accuracy in large-scale mapping scenarios.
Achieving a 98% reduction in AI traffic without sacrificing performance challenges conventional wisdom about the necessity of precise data representation in VLMs.
A single self-hosted LLM can absorb 50% of enterprise traffic while outperforming larger models in critical quality metrics.
A comprehensive dataset of Ethereum smart contract vulnerabilities reveals critical insights into the security landscape of blockchain applications.
Quantum-safe IPsec tunnels can now remain operational even when QKD infrastructure fails, ensuring uninterrupted secure communication.
JENGA reveals that RowHammer countermeasures can inadvertently increase task execution times, undermining the safety assumptions of real-time systems.
Farmers can now leverage AI for precision agriculture without compromising their data privacy, thanks to a novel federated learning framework tailored for rural infrastructure.
Despite meeting baseline requirements, none of the Bitcoin validators achieved full proof completeness, exposing critical gaps in agent-assisted development practices.
A columnar relational engine can outperform traditional graph databases by two-to-four orders of magnitude, challenging the need for specialized graph engines in enterprise workloads.
By eliminating bidirectional communication, L-shaped SFT allows edge devices to participate in fine-tuning without the need for constant connectivity, drastically cutting down on communication overhead.
CREDIT achieves 91.7% accuracy in predicting DSMEM profitability, delivering consistent speedups that outperform leading compilation frameworks.
A new protocol achieves a contraction rate of \( 1/\sqrt{2} \), pushing the boundaries of convergence in multidimensional approximate agreement closer to optimality.
FALCON achieves robust fault tolerance in edge AI by seamlessly integrating in-memory computing with stochastic techniques, ensuring reliable performance even in the face of significant noise and variations.
AInfer-PD slashes rollout completion times by up to 35.3% in distributed MoE environments, revolutionizing the efficiency of reinforcement learning inference.
CAPSUM slashes service deployment costs by up to 45.5% by intelligently balancing prediction accuracy and capacity constraints.
Energy costs for computation are not just about operations; they include significant 'rent' and 'fare' that can dramatically increase with context length, reshaping our understanding of computational efficiency.
Optimizing satellite learning architectures reveals that tailored federated learning can drastically improve wildfire detection across varying orbital altitudes.
RT-CDF achieves up to 4.39x faster sorting on GPUs for small to medium integer ranges, but struggles with larger ranges due to histogram overhead.
Achieving nearly 4x energy savings and over 4x latency improvements for LLMs on edge devices without sacrificing performance opens new avenues for deploying advanced AI in real-world applications.
Reducing ciphertext expansion and online costs in Federated Learning could revolutionize how sensitive data is aggregated without compromising privacy.
Rule-based pricing mechanisms consistently outperform reinforcement learning approaches in peer-to-peer electricity trading, even as energy storage boosts RL performance.
Local PIR can significantly boost retrieval capacity, achieving gains that scale with the number of identical sub-graphs in the system.
Cross-model KV sharing can boost accuracy and cut prefill costs dramatically, challenging the notion that KV states are strictly model-local.
ERIQ reveals that 54 previously unknown logic bugs lurk in popular DBMSs, exposing a critical oversight in existing testing frameworks.