Search papers, labs, and topics across Lattice.
This paper introduces SciSchema.org, a comprehensive set of 16 expert-annotated schemas designed to standardize the description of scientific processes across multiple disciplines, including Biology, Chemistry, and Psychology. By employing a human-in-the-loop schema-mining approach, the authors utilized large language models to generate candidate structures from existing literature, which were then refined through expert feedback to create robust final schemas. The resulting schemas facilitate improved reproducibility, automation, and comparison of scientific processes, ultimately enhancing the interoperability of scientific data.
A groundbreaking collection of 16 schemas that standardizes scientific process descriptions, paving the way for enhanced reproducibility and automation in research.
Scientific processes are often described in heterogeneous article discourse, with details needed for comparison, reproducibility, reuse, and automation dispersed across prose, tables, figures, protocols, and supplementary files. We present the first release of SciSchema.org, a multidisciplinary collection of 16 expert-annotated schemas spanning Biology&Biotechnology, Materials&Chemistry, Imaging&Measurement, Physics, and Psychology. Each schema defines reusable fields for describing process instances, including inputs, outputs, materials, instruments or software, parameters, conditions, procedural steps, measurements, and provenance-related information. The schemas were created through a human-in-the-loop schema-mining workflow in which large language models generated candidate structures from process specifications, scientific articles, and expert feedback, followed by domain-expert construction of final master schemas. The dataset contains final schemas in JSON Schema and SHACL formats, intermediate model-generated schemas, expert-feedback records, source-paper metadata, community-development materials, and analysis scripts. Technical validation assessed schema structure, development provenance, expert review, and syntactic conformance. The collection supports structured annotation, metadata enrichment, scientific knowledge graphs, information extraction, semantic publishing, and cross-study comparison.