Search papers, labs, and topics across Lattice.
75 papers published across 6 labs.
The case is made that barriers for new accelerators can be lowered by starting with a baseline AI accelerator, and adding minimal logic to support new operators demanded by new specialized domains, which leads to a versatile chip that can be manufactured at high volume and deployed for a range of popular applications.
The transition from file storage to database management systems transformed stored data into managed resources into AI model management systems, and AI now faces an analogous transition from AI model storage to AI model management.
It is found that cache performance depends on transfer granularity, intermediate memory use, and when transfers enter the request schedule, not only on device bandwidth, not only on device bandwidth.
The results show that inconsistency under uncertainty can be effectively learned, analyzed, and repaired through response-surface modeling, providing a scalable foundation for uncertainty-aware consistency management in CPS development.
A phase-decoupled, model-calibrated controller that hypothesizes that the optimal power setting is a property of the deployed (model, quantization, engine, hardware) combination rather than of the GPU class, that each lane warrants its own profile, and that converting SLO headroom into energy safely requires latency-gated calibration under a runtime SLO guard rather than a fixed recipe.
The case is made that barriers for new accelerators can be lowered by starting with a baseline AI accelerator, and adding minimal logic to support new operators demanded by new specialized domains, which leads to a versatile chip that can be manufactured at high volume and deployed for a range of popular applications.
The transition from file storage to database management systems transformed stored data into managed resources into AI model management systems, and AI now faces an analogous transition from AI model storage to AI model management.
It is found that cache performance depends on transfer granularity, intermediate memory use, and when transfers enter the request schedule, not only on device bandwidth, not only on device bandwidth.
The results show that inconsistency under uncertainty can be effectively learned, analyzed, and repaired through response-surface modeling, providing a scalable foundation for uncertainty-aware consistency management in CPS development.
A phase-decoupled, model-calibrated controller that hypothesizes that the optimal power setting is a property of the deployed (model, quantization, engine, hardware) combination rather than of the GPU class, that each lane warrants its own profile, and that converting SLO headroom into energy safely requires latency-gated calibration under a runtime SLO guard rather than a fixed recipe.
GPU-CFR is proposed, a compiler and runtime built on observation that for a fixed game, everything about a CFR iteration except the numerical values is known before the first iteration runs, and beats every CPU and GPU baseline on the mid-to-large games of the suite without changing the update rule.
A novel error analysis is provided that provides substantially sharper bounds for products of operators, thereby significantly relaxing existing restrictions on the maximum number of local machines while retaining optimal learning rates for the distributed kernel-based robust gradient descent algorithm.
A novel framework for model-based reinforcement learning which introduces approximate inverse process models within the training of reinforcement policies, and proposes a lightweight feedforward architecture for approximate inverse models and integrate them within the policy network of standard RL algorithms.
This paper introduces AccelForge, which improves upon existing accelerator modeling frameworks in capabilities, speed, and ease-of-use and includes fast mappers that enable accurate evaluation in orders of magnitude less (computer and human) time.
Intermediate activations in split LLM fine-tuning trivially leak raw training prompts past standard privacy defenses, but a learned obfuscation pipeline closes this leakage vector without destroying model utility.
Federated learning can cut edge BCI bandwidth by 42% without harming cohort-level accuracy, but aggregate benchmarks mask severe individual personalization failures of up to 12 percentage points under realistic network lag.
Cross-silo federated learning achieves up to 40% faster time-to-accuracy when SDN-level network telemetry—rather than client-side heuristics—determines which institutions train synchronously versus asynchronously.
Extreme label skew degrades federated multimodal performance nearly three times more than a sevenfold increase in participating clients, exposing data heterogeneity rather than network scale as the primary failure mode for decentralized clinical AI.
Decoupling multi-center biomedical AI from centralized aggregators, directed trust propagation paired with generative replay reveals coordinated higher-order molecular subnetworks of aging directly across private, fragmented clinical datasets.
Standard quadratic regularization fails to contain runaway client drift under extreme heterogeneity, but tuning the proximal exponent to $p \in [5, 7]$ slashes severe-stress federated loss by over 23%.
Predicting device-specific circuit fidelity before compilation with a fast GNN enables multi-QPU clusters to match brute-force hardware assignment quality without the crippling overhead of compiling across every target machine.
Monolithic federated updates are a fundamental bottleneck for embodied AI: decoupling vision, language, and action streams cuts uplink payloads by 96% and beats standard FedAvg by 22 percentage points under real-world wireless interference.
Event-driven neuromorphic execution slashes language model decode energy to 0.044 Joules per token by directly converting activation sparsity into skipped memory traffic rather than just idle compute cycles.
Frontier-scale agentic RL on 700B+ MoEs is no longer locked inside hyperscaler proprietary stacks, achieving stable 263-second step times across 64 GB300 GPUs via a verified, open-source training infrastructure.
GraphFAS (Graph Feature Automated Selection), a distributed feature selection procedure based on Boruta that decouples feature aggregation from model training, enabling direct integration with tabular models and direct compatibility with TreeSHAPbased explanations.
Analog over-the-air model aggregation can match the theoretical $\mathcal{O}(1/\sqrt{T})$ convergence of ideal FedAvg without requiring instantaneous channel state information or strict phase alignment.
Interactive communication yields zero minimax advantage in 1-bit distributed mean estimation: purely non-adaptive queries match the adaptive rate for any heavy-tailed moment condition $k > 1$.
Federated GANs for semantic communication suffer catastrophic drift under non-IID data, but keeping discriminators strictly local while conditioning generation on physical SNR yields up to a 58.2% recovery in reconstructed text fidelity.
Exploiting spatial proximity between hospitals enables direct AUC maximization under strict privacy constraints, preventing the discriminative performance collapses typical of standard federated learning on skewed clinical data.
Standard random fuzzing misses critical multi-core hardware bottlenecks, but treating micro-architectural stress-testing as an intrinsic curiosity problem reliably surfaces elusive, worst-case contention edge cases under fixed experimental budgets.
Battery-free edge nodes can now achieve cryptographic-grade uplink security without active RF generation by modulating orthogonal polarizations directly over the power beam that energizes them.
SemBridge establishes consumer observation as a semantic layer between placement and collective execution at a boundary between vendor runtimes that cannot share a native communicator.
Quantum error correction telemetry leaks logical inputs under realistic asymmetric noise, exposing secret states on physical hardware with over 92% total variation distance.
Static key-switching schemes leave massive hardware efficiency on the table—dynamically arbitrating between KLSS and HKS based on on-chip memory constraints cuts FHE bootstrapping latency by up to 3.31×.
Industry-standard GPU memory sanitizers harbor major blind spots across core memory spaces, finally quantifiable through 149 targeted CUDA stress tests.
Network topology alone can completely compromise secure aggregation in decentralized learning, enabling colluding nodes to crack private weights and extract training data using lattice reduction.
Mid-circuit measurements break standard flat quantum compilers; restructuring qubit mapping around hierarchical control flow slashes routing SWAP counts by up to 52% and error rates by up to 40% across monolithic and chiplet QPUs.
Sub-5-microsecond hybrid execution across CPUs, GPUs, and FPGAs is now achievable directly from pure Python, removing the rigid manual hardware-programming bottleneck currently stalling real-time quantum error correction.
FP8 Ozaki II emulates FP64 matrix multiplication by tensor-core products over a CRT residue system; converting the operands into residue planes costs integer-pipe and memory resources before tensor instructions issue.
Coupling wet-lab robotics directly to an exascale supercomputer via a unified cloud substrate collapses half a day of manual scientific analysis down to sub-minute interactive queries.
A controlled measurement study of self-hosted LLM inference across edge and near-edge deployment nodes: an NVIDIA Jetson AGX Orin and a near-edge server with CPU-only and GPU-enabled inference modes, highlighting that compute-side inference metrics alone can lead to suboptimal placement for latency-sensitive interactive web services.
The approach provides a systematic method for implementing large logical fanout operations across distributed error-corrected quantum processors, and a distributed implementation of the global gate GCZ involving logical qubits (encoded using BB-code blocks), exploiting the concurrency in transversal distributed fanouts.
Uncovering the trade-offs between accuracy and energy consumption could redefine how we optimize resource allocation in high-performance computing.
Guarantees data integrity in CDC systems even when merging live logs with copied rows, ensuring no updates are lost or incorrectly applied.
SHARDLP shatters previous performance records, solving a 1.185-billion-variable LP in under 10 minutes—21 times faster than CPU counterparts.
ContinuumBench is presented, a benchmark that controls workload, connectivity, and calibration assumptions and metrics over completed tasks hide unfinished work in cloud-edge controllers and compares placement-only and scale-capable controllers under declared regimes and stressors.
This work builds on Wang et al.'s taxonomy of Continuum Orchestration Systems employing DRL techniques and extends it with two further dimensions, measuring how LLMs are exploited and bridging the incommensurable per-tier signals and the LLM Orchestrator.
PENDA (processing element via norm-of-difference architecture), which leverages the law of cosines to recast multiplications as squared-difference operations to preserve exactness while optimizing hardware.
This work proves a dimension-independent weighted stability theorem for a directed noncommutative analogue of Mantel's theorem, which connects distributed quantum computing with noncommutative extremal combinatorics by identifying local collision probabilities with the weighted multiplicative energy of matrix-space decompositions.
Communication compression can significantly enhance performance in HPC and LLM workloads, but existing benchmarks fail to capture its true potential—CC-Bench changes that narrative.
Co-training speculative draft models directly inside 122B, 256K-token distributed RL runs removes the massive rollout bottleneck without causing pipeline stalls or context-parallel memory blowups.
Validators no longer need to maintain massive state trees on the critical consensus path to support trustless inclusion proofs: off-chain recursive SNARKs can maintain a billion-leaf state with only 2–4 seconds of latency.
Component-level metrics routinely report all-green health even while multi-step asynchronous user tasks quietly strand and fail.
Decentralized federated networks can survive Byzantine adversaries without enforcing consensus: predicting and clipping peer deviations against a client's own historical trajectory guarantees robust personalized convergence.
Data providers no longer have to feed continuous ML training streams on blind trust: dormant consumers are automatically cut off, and any data leak can be cryptographically traced back to the exact recipient using as few as 40 tabular rows.
Post-quantum downgrades break zero existing tests or automated audit scalars, leaving autonomous coding agents vulnerable to silently stripping hybrid key exchanges in 100% of tested engineering scenarios.
Frontier AI hardware controls are functionally unpoliceable by government regulators alone, but a deployable framework of private-sector telemetry and third-party auditing can close critical verification gaps within a single year.