Search papers, labs, and topics across Lattice.
This paper introduces SG-NCA, a novel scene graph generation framework utilizing Neural Cellular Automata (NCA) to facilitate efficient object detection and relationship modeling in surgical videos. By significantly reducing the parameter count to 55 times fewer than existing methods, SG-NCA addresses critical constraints of hardware footprint, hygiene, and latency in operating room settings. Evaluated on cataract and cholecystectomy surgeries, SG-NCA achieves performance on par with established models while ensuring compatibility with fanless edge devices for enhanced privacy and usability in surgical environments.
Achieving state-of-the-art scene graph generation with 55x fewer parameters, SG-NCA is tailored for the sterile demands of the operating room.
Scene graph generation from surgical video enables a holistic and structured understanding of surgical scenes by modeling objects and their semantic relationships. Despite recent advances, state-of-the-art approaches rely on large, parameter-heavy deep learning models that are impractical for deployment in the operating room (OR) due to hardware footprint, hygiene constraints, latency, and data privacy concerns. To the best of our knowledge, this is the first scene graph generation method built on NCAs and the first NCA framework capable of learning structured representations. We introduce SG-NCA, a lightweight scene graph generation framework based on Neural Cellular Automata (NCA), designed for inference in fanless devices critical for OR hygiene protocols. SG-NCA is the first scene graph generation combining NCA-based multiclass segmentation for efficient object detection and feature extraction with a lightweight relation predictor. We evaluate SG-NCA on videos of cataract surgery and cholecystectomy, demonstrating performance comparable to established baselines while requiring 55x fewer parameters. We showcase deployment on fanless edge devices better suited for the OR and demonstrate downstream applications such as surgical video captioning, highlighting SG-NCA's potential for affordable, privacy-preserving, and OR-ready intraoperative scene understanding.