Search papers, labs, and topics across Lattice.
The paper introduces BUSSARD, a normalizing flow-based model, for detecting anomalous relationships in scene graphs generated from images. BUSSARD leverages a language model to embed object and relationship tokens from scene graphs, mapping object-relation-object triplets to a Gaussian base distribution via bijective transformations learned by the normalizing flow. Experiments on the SARD dataset demonstrate a 10% AUROC improvement and 5x speedup compared to the state-of-the-art, along with improved robustness to synonym usage.
Normalizing flows can flag anomalous relationships in scene graphs with 10% better accuracy and 5x faster speed than existing methods, while also exhibiting superior robustness to semantic variations.
We propose Bijective Universal Scene-Specific Anomalous Relationship Detection (BUSSARD), a normalizing flow-based model for detecting anomalous relations in scene graphs, generated from images. Our work follows a multimodal approach, embedding object and relationship tokens from scene graphs with a language model to leverage semantic knowledge from the real world. A normalizing flow model is used to learn bijective transformations that map object-relation-object triplets from scene graphs to a simple base distribution (typically Gaussian), allowing anomaly detection through likelihood estimation. We evaluate our approach on the SARD dataset containing office and dining room scenes. Our method achieves around 10% better AUROC results compared to the current state-of-the-art model, while simultaneously being five times faster. Through ablation studies, we demonstrate superior robustness and universality, particularly regarding the use of synonyms, with our model maintaining stable performance while the baseline shows 17.5% deviation. This work demonstrates the strong potential of learning-based methods for relationship anomaly detection in scene graphs. Our code is available at https://github.com/mschween/BUSSARD .