Search papers, labs, and topics across Lattice.
This study addresses the challenge of data interoperability and findability in livestock population data by extracting and analyzing the semantics of existing reporting categories rather than relying on resource-intensive metadata creation. By focusing on cattle data from multiple sources, the authors demonstrate that the intrinsic age, sex, and production modifiers in the datasets provide a level of granularity that surpasses that of established vocabularies like AGROVOC. The findings suggest that leveraging existing data semantics can significantly enhance data discoverability and interoperability without necessitating the creation of new metadata or standards.
Existing cattle data contains rich semantic information that can vastly improve data interoperability without the burden of creating new metadata.
Livestock population data disaggregated by age, sex, and production are important inputs to calculations and models that inform our understanding of global health, yet these data are fragmented across disparate sources. Bridging data siloes to improve the findability of data requires interoperability. Conventional approaches to improving the findability and interoperability of data include indexing standardized metadata. However, creating metadata is time and resource-intensive and is often difficult in domains such as livestock, which lack standards that address the needs of broad user groups. When metadata exist, they typically need to be standardized against a pre-existing vocabulary, ontology, or thesaurus, requiring a technique known as `crosswalking'. To overcome issues in the absence of metadata, the lack of standards, and the resource-intensive solutions that currently exist, this study uses a bottom-up approach. By leveraging real-world reporting categories in datasets, the composition and semantics of terms already present in the data were extracted and analyzed. Using cattle data as a pilot, we find the age, sex, and production modifiers present across cattle terms from five datasets from four data sources capture granularity not present in AGROVOC, the largest agricultural vocabulary in the world. We discuss how the composition and semantics of these terms can be used to improve the interoperability and findability of data without first requiring metadata to be generated or standards to be created. Rather than forcing datasets to conform to an existing vocabulary, this approach uses the semantics embedded in terms already present in datasets, allowing systems to make data more discoverable and interoperable while maintaining culturally and dataset-specific terminology.