Search papers, labs, and topics across Lattice.
This paper introduces OpenRTAG, a comprehensive benchmark designed to evaluate the robustness of text-attributed graph learning under various data quality degradations. By organizing TAG quality issues into a unified 3x3 taxonomy, the benchmark facilitates standardized assessments across nine datasets and three downstream tasks, addressing the fragmented understanding of TAG robustness in the literature. Key findings reveal significant differences in model sensitivity and performance across traditional GNNs, LLM-GNNs, and a representative GFM when faced with composite degradation scenarios.
OpenRTAG reveals that traditional GNNs struggle significantly more than LLM-GNNs under realistic data quality degradation, highlighting critical vulnerabilities in graph learning models.
Text-attributed graphs (TAGs) are an important graph data form that combine relational structure with rich node text. However, real-world TAGs are often imperfect, with quality issues arising from text, structure, and labels, and typically manifesting as sparsity, noise, and imbalance. These dimensions define nine representative degradation scenarios that can substantially affect TAG learning. Although prior studies have explored specific mitigation strategies, existing evidence remains fragmented across degradation types, datasets, tasks, and model families, leaving TAG robustness insufficiently understood. To address this gap, we present OpenRTAG, a robustness benchmark for text-attributed graph learning. OpenRTAG organizes TAG quality issues into a unified 3 * 3 taxonomy and supports standardized evaluation across nine TAG datasets and three downstream tasks. It systematically evaluates scenario validity and model sensitivity, compares traditional GNNs, LLM-GNNs, and a representative GFM, investigates the effectiveness, efficiency, and robustness of scenario-matched baselines, and further examines model behavior under composite degradation scenarios. OpenRTAG provides a standardized testbed for understanding robustness in TAG learning under realistic low-quality settings.