Search papers, labs, and topics across Lattice.
This paper addresses the complex challenge of data citation for large language models, emphasizing its distinction from traditional document-level citation. It argues that effective data citation must ensure verifiability, traceability of provenance, and proper credit allocation to data creators. The authors propose three critical research directions: training data attribution, data citation during inference, and knowledge graph fact citation, highlighting the need for interdisciplinary collaboration to advance these areas.
Data citation for large language models is not just a verification issue; it鈥檚 a multifaceted challenge that could redefine how we acknowledge and trace the origins of AI-generated information.
Large language models increasingly mediate access to information, and a growing body of work asks whether they cite the sources behind their outputs. That work treats citation as a verification device and applies it to textual documents. Scholarly citation serves two further functions, credit and provenance, and it applies to data as much as to text. This paper argues that data citation for large language models is an open challenge, distinct from document-level citation grounding and harder to solve. We ask how such models should cite data so that outputs stay verifiable, provenance stays traceable, and credit reaches data creators and curators. We set out three research directions. Training data attribution has to turn influence estimates into references for corpora absorbed into model parameters. Data citation at inference time has to identify datasets, subsets, and query results at the right granularity and fixity. Citing knowledge graph facts has to define what a reference to a single triple denotes and how credit propagates along provenance. Progress on all three depends on joint work across the database, information retrieval, knowledge representation, and artificial intelligence communities.