Search papers, labs, and topics across Lattice.
This paper investigates the implications of best-effort computing in high-performance computing (HPC) for digital evolution, particularly focusing on the challenges posed by emerging hardware like the Cerebras Wafer-Scale Engine. By conducting two case studies鈥攐ne on multicellularity evolution using CPU-cluster multiprocessing and another on WSE-based simulations鈥攖he authors demonstrate that best-effort strategies can achieve significant performance improvements while maintaining robust quality of service despite hardware anomalies. The findings suggest that embracing non-deterministic computing paradigms can enhance the scalability and reliability of evolutionary optimization tasks in complex biological systems.
Best-effort computing can achieve a 92% scaling efficiency in evolutionary models while gracefully handling hardware failures鈥攖ransforming our approach to high-performance digital evolution.
Developments in high-performance computing (HPC) technology continue to drastically increase quantities of available processing power. In the context of digital evolution, this explosive growth offers opportunities to advance both hypothesis-driven explorations of multi-scale biological phenomena and application-driven evolutionary optimization targeting hard problem domains. A particular opportunity arises from emerging next-generation AI/ML hardware accelerator platforms, such as the 880,000-processor Cerebras Wafer-Scale Engine (WSE). Such hardware, however, constrains on-device data storage and movement --- a challenge compounded by vulnerability to failures arising over numerous device components. Best-effort relaxations that depart from a traditional deterministic computing paradigm can help accommodate such constraints, but complicate reproducibility and risk introducing artifactual biases. We explore these concerns, developing a framework to measure runtime behavior of best-effort code and examining case studies of best-effort computing in digital evolution projects. The first case study applies best-effort CPU-cluster multiprocessing to a multicellularity evolution model, which provides 92% scaling efficiency at 64 processes ($2.1\times$ speedup) and exhibits robust median quality of service, even under hardware anomalies. The second case study examines WSE-based simulations, demonstrating best-effort strategies to track spatiotemporal population history --- through sparse, asynchronous device-to-host sampling that tolerates hardware faults. In sum, across potential forms and scopes of best-effort relaxation, we argue that digital evolution is uniquely positioned to contribute in developing post-deterministic HPC paradigms.