Search papers, labs, and topics across Lattice.
100 papers published across 2 labs.
ROLoad-PMP achieves robust security for low-level software with less than 1.40% hardware overhead while enabling lightweight defenses that outperform existing solutions.
FedTVD redefines client weighting in federated learning, achieving significant performance gains by balancing data quality and quantity.
Performance gains of up to 59% are achieved by aligning block size with cache shape, challenging conventional wisdom about architectural optimization.
TANGCO achieves up to 246% improvement in resilience against cascading failures compared to traditional heuristics, showcasing the power of topology-aware optimization in networked systems.
Efficient belief synchronization in 6G networks can be achieved without requiring homogeneous AI models, preserving privacy and reducing costs.
TANGCO achieves up to 246% improvement in resilience against cascading failures compared to traditional heuristics, showcasing the power of topology-aware optimization in networked systems.
Efficient belief synchronization in 6G networks can be achieved without requiring homogeneous AI models, preserving privacy and reducing costs.
Traditional load balancing in EP MoE serving can lead to inefficiencies, but a new makespan-aware dispatcher achieves up to 15.5% throughput gains by adapting to varying compute and memory constraints.
Sustainability-optimal data center choices can significantly differ from latency-optimal ones, revealing the hidden costs of ignoring local grid carbon intensity.
AutoQuREO uncovers hidden resource trade-offs in quantum computing that existing tools struggle to reveal, transforming the landscape of quantum resource estimation.
Aggregating quantum information across nodes without losing fidelity is now possible, even for non-Gaussian codes, thanks to a novel measurement-based approach.
Reducing signing-phase time by over 700 times for multi-file uploads could revolutionize secure cloud storage in a post-quantum world.
Integration-first coverage strategies can reveal untested code paths in embedded systems, transforming how we assess testing completeness.
HBF can significantly boost LLM serving efficiency by enabling more expert replicas and reducing loading times, all while preserving critical execution paths.
ROLoad-PMP achieves robust security for low-level software with less than 1.40% hardware overhead while enabling lightweight defenses that outperform existing solutions.
Existing defenses against backdoor attacks in Vertical Federated Learning are fundamentally flawed, often relying on unrealistic assumptions that mask their true vulnerabilities.
Network features dominate in predicting 6G beamforming performance, outpacing other feature groups in critical metrics.
A hybrid cloud-edge system achieves near-perfect diagnostic recall while slashing operational costs and latency in rural healthcare settings.
INT4 quantization can distort the architecture landscape, but a zero-shot FP32 surrogate offers a surprising advantage in maintaining Pareto optimality.
SPADE slashes cloud model calls by 76% while preserving accuracy, revolutionizing the deployment of large language models in edge environments.
Enforcing verifiable delays could thwart most existing MEV threats without sacrificing system liveness.
OmniSphinx allows for seamless emulation of multiple mix network formats, drastically simplifying the infrastructure needed for secure communication.
RealmEye reveals that even in highly isolated environments, effective introspection can be achieved without compromising the security of confidential VMs.
The report reveals critical vulnerabilities in Energy Internet systems and offers innovative solutions to bolster their resilience against cyber-physical threats.
Operator-level scaling can reduce GPU usage by over a third while maintaining strict service level objectives for LLMs.
Achieving consensus in just two communication steps could revolutionize the efficiency of blockchain protocols under certain Byzantine conditions.
Achieving triangle-free graph coloring in $\log^{O(1)}\log n$ rounds could redefine the efficiency benchmarks for distributed algorithms in sparse graph structures.
An agent-driven approach to hardware design has achieved a 61.1% IPC speedup, outperforming expert-designed prefetchers by up to 23.6%.
Achieving a 5.1x speedup in a legacy weather simulation while ensuring scientific validity reveals the critical role of validation in AI-assisted GPU porting.
Meshlib slashes service mesh latency by eliminating sidecar proxies without sacrificing security, achieving the lowest end-to-end latency in its class.
Achieving a space complexity reduction for LL/SC implementations opens the door to efficient dynamic hashing on standard hardware, without sacrificing performance.
LAAB transforms performance reporting for mathematical libraries, ensuring that every benchmark is traceable and relevant to real-world scientific applications.
YAVIN achieves over 20x speedup in secure edge processing while maintaining robust cryptographic protections against untrusted memory.
Dryas can dynamically adjust its filtering capabilities in under a second, revolutionizing how engineers debug and analyze high-speed interconnects without interrupting their operation.
Function-space analysis reveals that traditional parameter comparisons can lead to suboptimal updates in federated learning, with LIGHTYEAR showing significant performance gains.
Achieving a 12-CNOT decomposition for double qubit excitation could redefine efficiency benchmarks in quantum gate implementations.
By decoupling generation and robustness in avatar training, Avatar-Forever achieves high-quality, real-time video generation without the pitfalls of traditional distillation methods.
LazyTrain boosts training efficiency by 1.24× and enables larger batch sizes, revolutionizing resource allocation for large language models on limited hardware.
Ambient IoT devices can transform ISAC systems into secure environments, achieving a 14-dB SNR advantage for legitimate users against eavesdroppers.
Xamt uncovers critical API discrepancies across deep learning libraries, revealing that traditional testing methods may miss up to 72 significant issues.
Agents can traverse complex maze environments more efficiently by leveraging local communication and leader-follower dynamics, achieving optimal performance with fewer resources.
PLB achieves up to 28% better goodput retention for Premium services during failures, redefining how we approach QoS in replicated databases.
HyperFlux achieves unprecedented core allocation speed, moving cores between VMs in just 13 microseconds, revolutionizing tail latency management in serverless environments.
Expert rerouting and attention-aware data packing can boost MoE model throughput by nearly 15%, reshaping efficiency in reinforcement learning.
FQTree slashes hardware costs for boosted decision trees by up to 57% without sacrificing accuracy, revolutionizing their deployment in latency-sensitive applications.
OpenMP Offloading can deliver up to 4x performance gains in multi-GPU setups, outperforming traditional single-GPU implementations across major architectures.
User-assisted collaborative inference can slash dedicated resource usage while boosting performance as demand scales.
INT8 support on NVIDIA's Blackwell Ultra GPU is effectively non-existent despite being listed in specifications, revealing a critical gap between hardware promises and practical usability.
Descriptive metadata can boost job execution success rates in multi-cluster AI workflows from 48% to 87%, revolutionizing operational efficiency.
Performance gains of up to 59% are achieved by aligning block size with cache shape, challenging conventional wisdom about architectural optimization.
By buffering intermediate values in fast DRAM, NITRO slashes inference latency by up to 85%, revolutionizing the efficiency of NAND flash-based computing.
Lonic achieves up to 66.28x energy efficiency improvements over leading GPUs, revolutionizing the training landscape for spiking neural networks.
Aggressive TCNOT scheduling can lead to significant quantum speedups, but it also risks overwhelming classical decoders, revealing a critical trade-off in FTQC performance.
Replacing SSDs with High-Bandwidth Flash in LLM serving can paradoxically slow down performance by over 5 times due to mismatched workload characteristics.
Achieving a staggering reduction in hardware footprint while maintaining near-perfect accuracy, Uni-SFU redefines the efficiency of activation function implementations in neural networks.
Keeping GPU-computed decisions on-device can boost performance by up to 2.39x, revolutionizing LLM-agent control efficiency.
Dion3 slashes optimizer step time by up to 6x while maintaining or improving loss performance compared to its predecessor, Muon.
Achieving a 1.84 dB PSNR improvement at 20% packet loss, this method redefines resilience in learned image compression.
Statistically secure bit commitment is now feasible with hybrid hardware, bridging classical and quantum cryptography.
Homomorphic multiplication is revealed to be a critical vulnerability, amplifying transient bit-flip errors and risking silent data corruption in privacy-preserving computations.
Faults in Barrett Modular Multiplication can be detected efficiently with minimal overhead, safeguarding critical cryptographic operations against adversarial attacks.
Malicious modifications to router weights can turn Mixture-of-Experts models into trigger-controlled bottlenecks, revealing a critical vulnerability in AI serving architectures.
MoMQ slashes frame completion times by over 70% and meets stringent interactive latency targets by leveraging video metadata for smarter packet scheduling.
Real-time monitoring of patient vitals is now feasible with a cloud-native framework that seamlessly integrates diverse wearables without vendor lock-in.
Prioritizing KV cache entries by importance can achieve over 93.7% accuracy during edge LLM handovers, significantly optimizing bandwidth usage.
Identical nodes can exhibit up to 5% performance variation, challenging assumptions about uniformity in data center hardware.
Achieving high throughput in quantum decoding at both room and cryogenic temperatures could redefine the efficiency of fault-tolerant quantum computing systems.
Independent analyses of branch and cache performance often underestimate gains, with 40% of workloads exceeding expected performance by over 6%.
Energy and latency can diverge by 3x under high computational demand, necessitating platform-specific models for accurate CNN inference cost predictions.
Charge-CIM slashes ADC energy consumption by 91.7% while doubling throughput, reshaping the landscape of energy-efficient deep learning accelerators.
Achieving a 1.85x speedup in matrix multiplication on Ascend NPUs could redefine performance benchmarks for dynamic tensor operations.
Battlefield 5G blocks SIM-transplant and rogue-certificate attacks while only adding 373.4 ms to onboarding latency, proving that enhanced security can be achieved without sacrificing efficiency.
DACER uniquely combines local self-attestation with global repair coordination, enabling real-time firmware restoration in automotive systems under threat.
POLO outperforms traditional dispatch optimization methods by effectively navigating the complexities of partial observability in multi-platform environments.
SCOUT reveals that real-time consensus among model replicas can pinpoint failure origins in LLM training, addressing critical issues that traditional methods overlook.
Optimal sparsity in sparse MoE models emerges only when considering the cluster's systems constraints, challenging traditional compute-centric design paradigms.
Players can achieve coordination in unknown Lipschitz environments without communication by leveraging a clever dithered quantization strategy.
Mixed-state prototypes allow quantum models to learn new classes without expanding circuit complexity, achieving robust performance with fewer qubits.
LLMs can dramatically reduce backend error rates in load balancing, but only if they exceed a critical parameter threshold—below that, they may perform worse than static policies.
Regret guarantees for decentralized algorithms in multi-agent bandits nearly match centralized rates, even under heavy-tailed rewards and information asymmetry.
Curiosity-driven exploration can dramatically enhance policy personalization in federated learning, even in sparse-reward scenarios.
A novel I/O-aware reformulation of wavelet convolution slashes memory usage and accelerates training speed, making it a game-changer for deep learning efficiency.
A decentralized orchestration framework enables seamless service provisioning in Organic 6G networks, minimizing coordination overhead while maintaining high decision quality.
A single 500 ms delay in AMI communication can escalate daily economic losses to over $5,000 during peak pricing periods.
A tested model failed to meet interoperability standards, revealing the critical gaps in EHR data exchange validation.
Parsing, not copying, dominates GPU LZ77 decode time, challenging long-held assumptions and revealing new avenues for optimization.
SeFoRA enables efficient federated fine-tuning of large models by overcoming the challenges of heterogeneous client ranks, achieving superior performance on benchmark tasks.
Native encoding outperforms external-function serialization in thread scaling for TDVRPTW, yielding better solutions and stability across multiple threads.
The GH200 Superchip outperforms the H100 PCIe by making hybrid CPU-GPU work divisions competitive and simplifying memory management for complex workloads.
External monitoring can drastically improve the accuracy of energy estimates for Spark applications, reducing significant underestimations.
FedTVD redefines client weighting in federated learning, achieving significant performance gains by balancing data quality and quantity.
A novel reliability metric and optimization framework could save cloud services millions while enhancing user experience.
The new hierarchical blocking operator reveals that cache occupancy fractions may be more architecture-specific constants than previously thought, challenging existing assumptions in performance prediction.
A new ontology for decentralization reveals that existing definitions often lead to contradictory conclusions in decentralized AI systems.
Uniform weighting in decentralized optimization can yield a vanishing tracking error, but discounted strategies may trap you in a persistent bias floor.
Sharing the right LoRA factor can significantly enhance fine-tuning performance in federated learning, with a novel adaptive strategy that outperforms traditional methods.
FedOrbit achieves up to 16.1 percentage points improvement in accuracy for federated learning in LEO satellites, tackling the unique challenges of non-IID data distributions.
FEAST boosts federated learning accuracy by over 2.4 points compared to the best existing model-heterogeneous weight-sharing baseline, all while slashing parameter traffic by 6.8 times.
Label granularity skew can drastically undermine federated learning performance, but FedBDFT offers a robust solution that significantly improves classification accuracy under this challenge.
FedA2L accelerates convergence in decentralized federated learning by up to 4.94 times while slashing communication rounds by 59%, even under severe data heterogeneity.
SwiftQK slashes QK-Norm latency by up to 93.9%, revolutionizing multi-GPU training efficiency for large language models.
MARA achieves 63.46% task completion in resource-constrained environments, outperforming existing methods by over 8 percentage points.
Achieving 99.96% detection accuracy, SSHafe not only thwarts SSH brute-force attacks but also revolutionizes password rotation with a secure, streamlined process.
DistMoE enables MLLMs to adapt to diverse visual-language domains without centralized data, achieving competitive performance while preserving client-specific knowledge.
Achieving an 87.2% success rate, SiriusDeliver slashes data warehouse delivery times from hours to mere minutes, transforming enterprise analytics workflows.