by Alice Aakerberg and Clare Manning
Over the last half-year, SEBI-Livestock has been working with partners across the University of Edinburgh to explore how data science and AI can help address persistent livestock data challenges
Collaborators include the Data Science Unit (DSU) at the School of Informatics, who utilise cutting-edge data science and AI tools to drive humanitarian change; Global Agriculture and Food Systems (GAAFS), bringing expertise across food systems, science-policy-industry interfaces, governance, ethics, and data-driven innovation; CAMARADES, from the Centre for Clinical Brain Sciences, contributing knowledge in systematic review and meta-analysis.
In March, we identified five use cases for data innovation: national livestock metrics, price data for modelling, insights from evaluation reports, livestock market potential and the animal health evidence base. Drawing on the diverse expertise from our partners, we gathered requirements for the use cases and explored various technical solutions to tackle our data obstacles.
In August, we reconvened to take a closer look at the national livestock metrics use case, this time joined by colleagues from the Jameel Observatory, who use data to support low- and middle-income communities facing climate and environmental risks. Our university partners were invited to share examples of AI or machine-learning techniques applied in their work which may be relevant to our national metrics use case.
Instead of jumping to a technical solution, we agreed that we first needed to gain a better understanding of the current data landscape: what data already exists, where the gaps are, how reliable and accessible different sources are, and what could realistically be achieved with the data.
From a broad challenge to a specific question
SEBI-L supports the Gates Foundation to understand how livestock sectors are changing in countries and commodities where investments are being made. This includes indicators such as livestock populations, production and productivity, food availability and greenhouse-gas emissions intensity.
National datasets can tell us how many cattle a country has or how much milk is produced. But when we ask more detailed questions – how are cattle distributed between pastoral, mixed and commercial production systems, how productivity differs between them, and how do these patterns vary across places and over time? – significant gaps and inconsistencies in the available data become apparent.
These more detailed questions are important for understanding livestock-sector transformation, so during the August workshop we narrowed the challenge to a potential proof of concept: can we better estimate the distribution of cattle population and production across different production systems in Kenya?
Kenya provides a useful starting point because there is already relevant research, data and expertise to build on, while important gaps remain.
Understanding the data landscape
Before modelling, we need a clearer picture of the data we already have. Building on the broader national-metrics requirements identified in March, we agreed to develop a shared data evaluation table for the Kenya cattle proof of concept. This will record geographic and temporal coverage, accessibility, update frequency, quality, and missingness for each source, and potential links to other datasets.
The search will extend beyond conventional datasets to government publications, institutional reports, project evaluations and other grey literature. The goal is not simply to collect more data, but to assess whether the available sources are suitable for stitching together into a more complete picture.
Stitching together an incomplete picture
A major part of the workshop explored data stitching: combining different, imperfect sources of information to produce estimates that cannot be obtained from one dataset alone.
The DSU shared examples from their work with UNICEF and WHO, where similar data challenges arise. Surveys can provide detailed information but may only be available for certain places or times, while national statistics provide wider coverage at a coarser level. Other sources, including satellite imagery, land-cover information, building data and OpenStreetMap, can provide additional information at much finer spatial and temporal scales.
The challenge is not simply to combine as much data as possible. Sources differ in their quality, coverage and granularity, and relationships that appear useful in one place may not hold elsewhere.
One approach discussed was the use of graph neural networks (GNNs), which can model relationships between neighbouring areas, different administrative levels and locations over time, while incorporating environmental and other contextual information.
For livestock data, these approaches could eventually help combine national statistics, subnational surveys, production-system observations and other contextual information to generate estimates at a level of detail that is currently missing.
What happens next?
Next, we will populate the evaluation table with the existing Kenya cattle and national metrics sources, systematically identify additional evidence, and then assess whether the available information is sufficient for an initial data stitching exercise.
Success would mean producing a defensible estimate for a national metric we currently cannot measure, but identifying where the evidence is insufficient would also help pinpoint priorities for future data collection.
For us as members of the data team, one important takeaway was that data innovation does not always begin with a complicated algorithm. Sometimes the critical first step is bringing fragmented information together, understanding its limitations, and asking a sufficiently precise question. March helped us identify where innovation could add value; August moved us from a broad opportunity to a concrete challenge. Now we get to see how far the available data can take us.
Alice Aakerberg and Clare Manning are Data Analyst Programmers with SEBI-Livestock.