Search papers, labs, and topics across Lattice.
This paper outlines twelve practical strategies for small research software teams to effectively manage IT disasters, particularly in light of recent disruptions such as government attacks and natural disasters. It emphasizes the importance of proactive disaster planning and recovery, especially for teams lacking dedicated IT resources. Key findings suggest that leveraging institutional support from research computing groups and data librarians can significantly enhance resilience against such crises.
Small research software teams can dramatically improve their disaster resilience by implementing practical strategies and leveraging institutional support systems.
In 2025, the US government launched an unprecedented series of attacks on its own scientific research groups. A year later GitHub dropped below 90% availability for the first time, while wildfires in Canada, France, Spain, and elsewhere forced researchers from the homes and labs. These events and others have reminded us just how fragile research computing systems can be, and that planning for disasters is one of the most effective ways to prevent them. This paper is a short guide to disaster planning and recovery for a small research software team. The tips assume you are doing everything yourself on top of your regular job, and that you aren't an experienced system administrator. Some of the tips do require that kind of expertise, but most research institutions have research computing groups, data librarians, and environmental health-and-safety offices whose entire job is to help with exactly these problems. This paper tells you what"done"looks like; they can often provide it.