Search papers, labs, and topics across Lattice.
Distributed training, model parallelism, AI accelerator design, and large-scale compute infrastructure.
#17 of 24
4
The case is made that barriers for new accelerators can be lowered by starting with a baseline AI accelerator, and adding minimal logic to support new operators demanded by new specialized domains, which leads to a versatile chip that can be manufactured at high volume and deployed for a range of popular applications.
The transition from file storage to database management systems transformed stored data into managed resources into AI model management systems, and AI now faces an analogous transition from AI model storage to AI model management.
It is found that cache performance depends on transfer granularity, intermediate memory use, and when transfers enter the request schedule, not only on device bandwidth, not only on device bandwidth.
The results show that inconsistency under uncertainty can be effectively learned, analyzed, and repaired through response-surface modeling, providing a scalable foundation for uncertainty-aware consistency management in CPS development.
A phase-decoupled, model-calibrated controller that hypothesizes that the optimal power setting is a property of the deployed (model, quantization, engine, hardware) combination rather than of the GPU class, that each lane warrants its own profile, and that converting SLO headroom into energy safely requires latency-gated calibration under a runtime SLO guard rather than a fixed recipe.
GPU-CFR is proposed, a compiler and runtime built on observation that for a fixed game, everything about a CFR iteration except the numerical values is known before the first iteration runs, and beats every CPU and GPU baseline on the mid-to-large games of the suite without changing the update rule.
A novel error analysis is provided that provides substantially sharper bounds for products of operators, thereby significantly relaxing existing restrictions on the maximum number of local machines while retaining optimal learning rates for the distributed kernel-based robust gradient descent algorithm.
A novel framework for model-based reinforcement learning which introduces approximate inverse process models within the training of reinforcement policies, and proposes a lightweight feedforward architecture for approximate inverse models and integrate them within the policy network of standard RL algorithms.
This paper introduces AccelForge, which improves upon existing accelerator modeling frameworks in capabilities, speed, and ease-of-use and includes fast mappers that enable accurate evaluation in orders of magnitude less (computer and human) time.
Intermediate activations in split LLM fine-tuning trivially leak raw training prompts past standard privacy defenses, but a learned obfuscation pipeline closes this leakage vector without destroying model utility.
Federated learning can cut edge BCI bandwidth by 42% without harming cohort-level accuracy, but aggregate benchmarks mask severe individual personalization failures of up to 12 percentage points under realistic network lag.
Cross-silo federated learning achieves up to 40% faster time-to-accuracy when SDN-level network telemetry—rather than client-side heuristics—determines which institutions train synchronously versus asynchronously.
Extreme label skew degrades federated multimodal performance nearly three times more than a sevenfold increase in participating clients, exposing data heterogeneity rather than network scale as the primary failure mode for decentralized clinical AI.
Decoupling multi-center biomedical AI from centralized aggregators, directed trust propagation paired with generative replay reveals coordinated higher-order molecular subnetworks of aging directly across private, fragmented clinical datasets.
Standard quadratic regularization fails to contain runaway client drift under extreme heterogeneity, but tuning the proximal exponent to $p \in [5, 7]$ slashes severe-stress federated loss by over 23%.
Predicting device-specific circuit fidelity before compilation with a fast GNN enables multi-QPU clusters to match brute-force hardware assignment quality without the crippling overhead of compiling across every target machine.
Monolithic federated updates are a fundamental bottleneck for embodied AI: decoupling vision, language, and action streams cuts uplink payloads by 96% and beats standard FedAvg by 22 percentage points under real-world wireless interference.
Event-driven neuromorphic execution slashes language model decode energy to 0.044 Joules per token by directly converting activation sparsity into skipped memory traffic rather than just idle compute cycles.
Frontier-scale agentic RL on 700B+ MoEs is no longer locked inside hyperscaler proprietary stacks, achieving stable 263-second step times across 64 GB300 GPUs via a verified, open-source training infrastructure.
GraphFAS (Graph Feature Automated Selection), a distributed feature selection procedure based on Boruta that decouples feature aggregation from model training, enabling direct integration with tabular models and direct compatibility with TreeSHAPbased explanations.
Analog over-the-air model aggregation can match the theoretical $\mathcal{O}(1/\sqrt{T})$ convergence of ideal FedAvg without requiring instantaneous channel state information or strict phase alignment.
Interactive communication yields zero minimax advantage in 1-bit distributed mean estimation: purely non-adaptive queries match the adaptive rate for any heavy-tailed moment condition $k > 1$.
Federated GANs for semantic communication suffer catastrophic drift under non-IID data, but keeping discriminators strictly local while conditioning generation on physical SNR yields up to a 58.2% recovery in reconstructed text fidelity.
Exploiting spatial proximity between hospitals enables direct AUC maximization under strict privacy constraints, preventing the discriminative performance collapses typical of standard federated learning on skewed clinical data.
Standard random fuzzing misses critical multi-core hardware bottlenecks, but treating micro-architectural stress-testing as an intrinsic curiosity problem reliably surfaces elusive, worst-case contention edge cases under fixed experimental budgets.
Battery-free edge nodes can now achieve cryptographic-grade uplink security without active RF generation by modulating orthogonal polarizations directly over the power beam that energizes them.
SemBridge establishes consumer observation as a semantic layer between placement and collective execution at a boundary between vendor runtimes that cannot share a native communicator.