Search papers, labs, and topics across Lattice.
This paper conducts a comprehensive narrative review of 151 works in enterprise cybersecurity, categorizing methodological practices into eleven families and detailing their strengths, weaknesses, and failure modes. It emphasizes the importance of methodological pluralism and provides executable protocols for each family, complete with evaluation criteria and reporting checklists. The findings reveal that inconsistencies in intrusion-detection algorithm rankings are largely due to variations in evaluation design rather than the algorithms themselves, underscoring the need for careful alignment of research methods with specific decision contexts.
Inconsistencies in cybersecurity research are often driven by flawed evaluation designs rather than the technologies being tested.
Enterprise cybersecurity research draws on a wider range of methods than any single community routinely teaches. Researchers face a selection problem before they face a technical one: a study may simultaneously need a systematic review, a design-science artifact, a controlled detection experiment, an interview study, or an attack-graph model. This paper addresses that problem in two ways. First, it provides a narrative review and synthesis of methodological practices across a verified corpus of 151 works. We organise these practices into eleven methodology families, detailing for each what questions it answers, the strength of its supporting evidence, and its common failure modes. Second, we convert each family into an executable protocol comprising ordered steps, required instruments, evaluation criteria, common validity threats, and a reporting checklist. Every protocol is also visually mapped to make the sequence, decisions, and threats legible at a glance. We also treat contradictions in the literature as evidence. For example, reported rankings of intrusion-detection algorithms are wildly inconsistent across individually careful studies. We argue this pattern is most parsimoniously explained by variations in evaluation design rather than the algorithms themselves, as these studies differ in design dimensions known to shift results by more than the margins separating the algorithms. Ultimately, the evidence supports methodological pluralism disciplined by explicit validity reasoning. We conclude that researchers must match their evaluation design to the decision under study, triangulate technical against organisational evidence, explicitly state the population a result generalises to, and report the conditions under which the result would not hold.