Search papers, labs, and topics across Lattice.
100 papers published across 6 labs.
Retry policies can inadvertently amplify failures, reducing success rates by over 25% when not designed with system-level considerations in mind.
Achieving a balance of low complexities in data dispersal, storage, retrieval, and recovery could redefine efficiency standards in distributed storage systems.
Aggregate latency masks crucial insights into LLM inference efficiency, revealing that backend changes and quantization effects significantly influence performance metrics.
RNN-guided load balancing slashes global workload imbalance to 3.5%, boosting simulation efficiency in multicellular growth models.
SCAFFOLD's struggles in federated optimization stem from its inability to reliably estimate the global gradient at the Edge of Stability, revealing critical limitations in its design.
RNN-guided load balancing slashes global workload imbalance to 3.5%, boosting simulation efficiency in multicellular growth models.
SCAFFOLD's struggles in federated optimization stem from its inability to reliably estimate the global gradient at the Edge of Stability, revealing critical limitations in its design.
MeMark embeds watermarks in the internal states of SNNs, ensuring ownership evidence survives even after significant model alterations.
KOPE's ability to retain and leverage past optimization experiences leads to a 1.54x speedup in kernel optimization tasks compared to existing methods.
QEF-GT-AdamW achieves superior robustness and convergence in decentralized learning over wireless networks, even when faced with severe communication constraints.
Federated learning can now predict QoS failures in wireless networks, achieving near-centralized performance while preserving user data privacy.
Selective transmission of task-relevant image latents can significantly reduce communication costs while maintaining high accuracy in edge AI applications.
TOPAS slashes job completion times by up to 49.4% in multi-agent LLM serving by intelligently balancing prefix caching and request scheduling.
Conditional total correlation serves as the precise information cost of parallelism in adaptive sampling, revealing critical insights into the efficiency of decoding strategies.
A single transition in Sigries can compromise security, reducing its effectiveness to just 1 second, while the new FiRM approach eliminates this vulnerability and complexity.
GIFT enforces user data isolation in LLM serving with less than 11% throughput overhead, revolutionizing privacy protection in shared infrastructures.
Real-time anomaly detection can achieve over 96% accuracy in identifying threats to DeFi stablecoins, potentially safeguarding billions in assets.
Retry policies can inadvertently amplify failures, reducing success rates by over 25% when not designed with system-level considerations in mind.
Slasher can dynamically adjust Azure datacenter power consumption with minimal impact on workloads, addressing critical energy management challenges in cloud computing.
APC-RLNC achieves up to 9.8 percentage-point improvements in packet delivery while reducing latency by up to 23% in dynamic wireless environments.
Spatial filtering in IoT messaging can be achieved without altering existing protocols, unlocking new possibilities for location-aware applications.
Achieving up to 32 times greater throughput in energy-harvesting wireless computing by optimizing resource allocation in real-time.
HSMA-TRSM achieves over 2x speedup on key GPU platforms by optimizing shared memory usage for complex triangular solves, revealing untapped performance potential in BLAS operations.
A machine learning predictor in BOOSTEDSOSA slashes runtime estimation errors by up to 63.85%, revolutionizing scheduling efficiency in high-performance computing.
Maia 200 achieves unprecedented AI acceleration while slashing energy costs, redefining the landscape for high-performance computing.
Achieving rapid earthquake magnitude estimation with a single station, SeisMamba outperforms traditional models in both speed and accuracy, even in unseen regions.
FedV-KGQA enables multi-hop reasoning across disjoint knowledge graph silos without compromising data sovereignty, achieving near-centralized performance.
A new feature-major codebook layout accelerates self-organizing map training by up to 621x, enabling the largest reported atlas of 1.05 million neurons on a single GPU.
Achieving over five times improvement in Latency-Error-Energy metrics could redefine efficiency standards for edge-deployable virtual sensing systems.
Achieving 90% accuracy in optical card reading without manual intervention could redefine the capabilities of physical Turing Machines.
Simthesizer achieves up to 284.96x faster simulation speeds while maintaining a mere 2.51% throughput error, revolutionizing how we model LLM serving systems.
Formally-verified NoCs can now be generated effortlessly, eliminating the tedious verification process while ensuring strong liveness guarantees.
Achieving almost stable matching in constant rounds on general bipartite graphs with minimal shared randomness could revolutionize distributed algorithms in large-scale networks.
Achieving up to 100x performance improvements in datacenter replication, scarHW challenges the limits of traditional consensus protocols.
Achieving over 100 times faster tensor decomposition on a single GPU could revolutionize simulations of complex quantum systems.
LLMs can optimize database queries on GPUs, achieving over 2.5x speedup by leveraging advanced execution strategies and kernel fusion.
A federated model could revolutionize how medical centers share and enhance device knowledge, ensuring local control while fostering collaborative improvement.
Current hardware fuzzing techniques are falling short, with significant gaps in input generation and feedback mechanisms that could undermine verification reliability.
Achieving a balance of low complexities in data dispersal, storage, retrieval, and recovery could redefine efficiency standards in distributed storage systems.
Targeted mixed-precision strategies can cut CFD simulation time and energy costs by over a third without sacrificing accuracy.
ROS2 Connect achieves lower latency and higher stability for remote robotic operations, revolutionizing how distributed systems communicate over WANs.
FLINT transforms LLM inference by integrating high-bandwidth flash, overcoming memory constraints that limit model deployment and performance.
Reducing satellite training time by nearly 19% while slashing energy consumption by up to 88% could revolutionize satellite-based machine learning.
Achieving a 98.6% energy reduction in cloud removal for LEO satellites could revolutionize Earth observation capabilities.
Aggregate latency masks crucial insights into LLM inference efficiency, revealing that backend changes and quantization effects significantly influence performance metrics.
pigzpp achieves up to 16 times faster compression than Python's gzip while maintaining full compatibility with existing gzip standards.
A probabilistic framework reveals that synchronization overhead can negate the expected advantages of full-tree parallelism in distributed search systems.
Energy-aware strategies in O-RANs can cut energy consumption by over 5% while maintaining performance, challenging traditional deployment models.
Best-effort computing can achieve a 92% scaling efficiency in evolutionary models while gracefully handling hardware failures—transforming our approach to high-performance digital evolution.
Eliminating thermal-induced tuning stalls in optical interconnects can boost MoE model performance by up to 3.8x, unlocking new potential for large-scale AI systems.
Trustless document co-signing is achievable over a lossy optical channel without intermediaries, redefining secure mobile communications.
HMT slashes hashing costs by 2.4x compared to Ethereum's Merkle Patricia Trie while adapting dynamically to changing access patterns.
Turning BGP hijack filtering from altruism into a market-driven transaction could fundamentally change how networks protect themselves against malicious announcements.
Trust evaluations can now adaptively integrate multi-source evidence while quantifying uncertainty, leading to more reliable collaborator selection in distributed systems.
Achieving up to 5.35x faster data processing for multimodal AI datasets could revolutionize how researchers handle and access diverse data formats.
TailSieve achieves up to 2.59x speedup in LLM rollouts by intelligently routing long-tail requests, transforming how we handle high-concurrency decoding.
Regularization in federated learning can be significantly influenced by the choice of update masks, revealing a critical trade-off between generalization and training efficacy.
Achieving a faster convergence rate in federated multiobjective optimization could redefine how we approach complex, conflicting objectives in distributed learning environments.
FedCC enables clients to handle ambiguity in data, leading to a remarkable 67.3% accuracy even when facing severe label distribution skews.
SplitLite slashes communication costs in federated learning by up to 93.5% without sacrificing performance, revealing a hidden structure in model training data.
Quantum computers could compromise major cryptocurrencies, but practical migration strategies to post-quantum security are within reach.
NICWhisper reveals that electromagnetic emissions from network interface cards can effectively identify network threats, achieving over 80% accuracy without analyzing packet data.
Major gaps in firmware security practices could leave the TianoCore community vulnerable, but targeted improvements could significantly enhance UEFI firmware integrity.
VIPER slashes design-space exploration time from hours to under a minute while achieving less than 10% error in performance predictions for Processing-in-Memory systems.
A decentralized bidding system for LLM agents not only enhances efficiency but also reduces manipulation risks, outperforming traditional orchestration methods.
The shift from ad-hoc to standardized carbon accounting in climate modeling reveals significant discrepancies in energy and emissions reporting that could reshape climate research practices.
SxSSD allows trusted applications to dynamically define FTL policies while preserving the security isolation of traditional SSDs, striking a crucial balance between flexibility and safety.
Achieving lower latency in secure IoT data aggregation, Phi-PHE-BC redefines the performance landscape for homomorphic blockchain architectures.
Critical compliance gaps in IoT device security were uncovered, revealing vulnerabilities in traffic encryption and availability that could jeopardize future 5G and 6G networks.
Decoupling Bluetooth pairing from authorization could revolutionize how we manage device access and security in IoT environments.
Local timing disturbances in microservice architectures can lead to unpredictable system-level effects, challenging the reliability of software-defined vehicles.
ShardMeter reveals that larger training islands can lead to diminishing returns, fundamentally changing how we approach resource allocation in distributed AI training.
CED-EF achieves faster convergence in decentralized optimization while using significantly less communication bandwidth, reshaping the landscape of multi-agent learning.
Achieving a semantic demand lower bound of 43.59375 GiB reveals that traditional memory limits can be surpassed without sacrificing execution accuracy in AI inference.
Disabling thread affinity can significantly improve the scaling of parallel applications with work imbalance on hybrid CPU architectures.
Achieving 242.97 GFLOPS on a RISC-V accelerator not only outpaces traditional CPUs but also does so with a fraction of the power consumption.
Achieving a 25% footprint reduction while enhancing performance metrics positions this SRAM design as a game-changer for future memory technologies at the 2nm node.
Energy-efficient distributed algorithms can now operate with nodes that selectively sleep, drastically cutting down energy consumption while maintaining performance.
Polymer-linked nanoparticle networks can harness heat for computing, potentially revolutionizing neuromorphic hardware design.
Achieving 2–5 orders of magnitude speedup in path selection for Virtual Payment Channels could revolutionize off-chain transaction efficiency in Payment Channel Networks.
Achieving a 99% correlation with real silicon, this simulation framework reveals crucial insights into the architectural evolution of GPUs for AI workloads.
DFL-C not only ensures global model consistency in decentralized federated learning but also outperforms existing solutions in the face of Byzantine attacks.
Hardware hacking competitions reveal that combining diverse verification techniques can uncover a broader spectrum of vulnerabilities in SoC designs than previously thought.
GCA reduces communication overhead in federated learning by up to 99.15% while simultaneously enhancing data protection and improving model accuracy.
Fully-non-leaking wait-free implementations are possible for some concurrent objects, but not all—revealing critical limitations in information security for concurrent systems.
Achieving a staggering 99.7% reduction in computational overhead, this framework revolutionizes real-time anomaly detection in geospatial data streams.
Quantum threats could render existing blockchain privacy protocols obsolete, but Obscura-PQ offers a robust, efficient solution that maintains privacy even against future adversaries.
Bidders can now prove eligibility without revealing their bids, striking a crucial balance between privacy and transparency in decentralized auctions.
Reinforcement learning strategies can enable legitimate receivers to achieve near-optimal secrecy rates in competitive RIS auctions, significantly outpacing conventional bidding approaches.
Deeper partitions in federated fine-tuning may boost throughput and privacy, but they can also cause LLM performance to collapse catastrophically.
The synchronization tax can consume over 50% of communication time in GPU scale-up domains, fundamentally challenging our understanding of bandwidth scaling.
Prefix delegation can reduce pod deployment time by over 90% compared to traditional individual-IP allocation methods, enabling efficient fleet-scale Kubernetes operations.
Targeted approximation in floating point multipliers can yield up to 92% hardware footprint savings without sacrificing CNN accuracy.
Lexical convergence in LLMs can be significantly influenced by peer-ranked feeds, but distributed sources fail to provide a reliable advantage in shaping agent opinions.
Achieving high-accuracy disease detection without ever exposing sensitive patient data could revolutionize collaborative medical research.
AEGIS slashes token recovery rates to nearly zero while preserving model performance, tackling all three critical channels of information leakage in federated learning.
Introducing required affinity in Kubernetes scheduling transforms the pod-deployability problem into a PSPACE-complete challenge, exposing hidden complexities in resource management.
ODEONN achieves a 45× reduction in energy-delay product while maintaining over 98% accuracy compared to traditional software simulations.
Achieving best-in-class energy efficiency of 39 pJ/b, this innovative detector redefines the capabilities of MU-MIMO-OFDM systems.
Workload intensity, not garbage collector choice, is the key driver of energy consumption in Java applications, challenging conventional wisdom about collector rankings.
FedCurv-DR not only reduces knowledge forgetting but also ensures fairness and energy efficiency in AI models applied to evolving cultural heritage data.
MS-WDRO outperforms traditional methods by leveraging the Wasserstein metric to fuse heterogeneous data sources, achieving superior graph recovery even in sample-scarce environments.
FleetSieve cuts GPU resource usage by over 5% while ensuring LLM latency meets stringent service level objectives.
Achieving a 2.3x performance improvement in LLM serving while increasing cache hit rates to over 93% could redefine efficiency benchmarks in large-scale AI deployments.
Efficiently routing queries in AI systems can be achieved without the costly overhead of exhaustive value estimation, thanks to novel policies that balance accuracy and cost.