It seems that large livestock genomic data sets may need stricter oversight and standardisation, after analysis of a global database of cattle diversity found omissions and errors in metadata – the information which makes such data actionable.
While genomic data consists of DNA or RNA sequences, metadata sets the all-important context for each sequence, including biological, technical and analytical elements. Without this, the raw data is uninterpretable strings of letters.
Researchers from the University of Göttingen took a closer look at the 1,000 Bull Genomes Project, a collaborative database which includes whole genome sequences for thousands of bulls, covering a large proportion of global cattle diversity. They analysed the metadata from all the records to date, which includes each animal’s subspecies, breed and individual identifier.
The analysis revealed that 603 individuals in the database had been assigned an incorrect subspecies or breed, which the scientists corrected themselves using relevant databases. However, a number of these individual records had errors at multiple levels of the metadata, meaning an overall higher level of inaccuracy.
Although the findings were not outside what might have been expected from previous studies on RNA sequencing data, they did underline a persistent problem that could have knock-on effects for those trying to use such datasets to improve breeding outcomes in cattle.
“Community collection efforts, as the 1000 Bull Genomes Project, can on the one hand reduce the workload and financial burden of the individual contributors, on the other hand, they can be prone to irregularities and misassignments in the metadata,” they wrote in the journal Genetics Selection Evolution.
Interested in livestock genetics? You might like these articles
Only findable breeds can be meaningfully used for further work, and they pointed to the case of the Qinchuan breed, which had one letter wrong in the metadata of samples, and when corrected, resulted in the sample size increasing from two to 39 individuals. Similarly, they noted that the N’Dama breed was listed as both NDama and Ndama, which could result in the erroneous detection of two breeds when using case-sensitive software.
Such datasets, where records are added and edited by different people around the world, call for an increased need for curation efforts, increasing quality control, they urged. They made several recommendations for how the metadata could be improved, including better indications of crossbreeds, using a custom vocabulary, or using the established vocabulary of the Livestock Breed Ontology employed by the Functional Annotation of Animal Genomes (FAANG) initiative, which standardise the language used around breeds of farmed animals.
“Clean, consistent metadata is essential for meaningful reuse of genomic datasets, even when obtained from established consortia. Therefore, its careful examination should not be foregone, even when working with well-established, prestigious datasets,” they added.
“Even in widely used datasets, like the 1000 Bull Genomes Project, metadata errors occur on different levels…and may affect downstream research applications.”
Key takeaways
- Researchers found metadata errors in 603 cattle genomic records.
- Breed and subspecies errors could affect downstream cattle research.
- The quoted cost excludes the camera, onboard computer and laser-weeding equipment.
- A single letter error increased one breed’s sample size from two to 39.
- Researchers call for stronger curation and metadata standardisation.
- Clean metadata is vital for making livestock genomic data reusable.
Want to read more stories like this? Sign up to our newsletter for bi-weekly updates on sustainable farming and agtech innovation.







