Search papers, labs, and topics across Lattice.
This paper introduces SIREN-Bench, a co-simulation platform that integrates SUMO and CARLA to generate and evaluate emergency vehicle (EMV) interactions with civilian traffic. By providing behavior-level control over EMV privileges and civilian responses, SIREN-Bench enables the assessment of critical safety interactions through realistic simulations. The results reveal distinct behavior-dependent failure modes across various tasks, highlighting the challenges in trajectory prediction and detection during emergency scenarios, and establishing SIREN as a valuable tool for autonomous driving research.
Behavior-driven evaluation reveals that no learned predictor outperforms a constant-velocity reference in emergency vehicle scenarios, exposing critical gaps in current models.
Emergency vehicles (EMVs) can reorganize surrounding traffic as civilian vehicles brake, change lanes, or form rescue corridors in response to their passage. Evaluating these safety-critical interactions requires behavior-level control over both EMV privileges and civilian responses, together with consistent sensing and ground truth. Existing datasets and simulation benchmarks do not directly provide this combination. We present \textbf{SIREN}, a behavior-driven SUMO--CARLA co-simulation platform for generating EMV--civilian interactions. SIREN couples SUMO's network-level traffic evolution and behavior logic with CARLA's continuous vehicle control and synchronized onboard sensing; depending on the active behavior, the interaction is controlled by SUMO, CARLA, or jointly. We instantiate the platform as \textbf{SIREN-Bench-v1}, comprising seven parameterized interaction templates across emergency levels L1--L3 and three behavior families, with synchronized sensor observations and simulator-native annotations. We demonstrate the benchmark through three representative tasks: 3D object detection, trajectory prediction, and vision-language risk understanding. Evaluations of nine trajectory predictors, four LiDAR-based detectors, and five vision-language models reveal behavior-dependent failure modes. Traffic-clearance interactions are hardest for detection, privileged intersection traversal is hardest for prediction, and no learned predictor outperforms the constant-velocity reference on average. Vision-language models perform substantially better on normal traffic than on near-miss and collision events. These results demonstrate the value of behavior-centered benchmarking and establish SIREN as an extensible data-generation and evaluation platform for autonomous-driving and transportation safety research.