Yield prediction for hybrid cereals at new locations
How does a hybrid cereal perform at a location where it has never been grown? We built a model that answers that question from genetic markers, soil and weather data – and applied it to 20,000 new locations to determine the best-performing variety for 2017.
The situation
The decision about which hybrid is grown where is made before the season – that is, before weather and soil conditions are known. Inferring from trial results at known locations to unknown ones is therefore not a marginal question but the core task of the breeding and commercial business.
The data
Several datasets spanning a total of 15 years covering more than 2,000 hybrid types and locations, together with genetic markers and soil and weather parameters.
Challenges and solutions
Identify outliers across several sources. A thorough analysis was carried out for each dataset. Outliers within locations could be identified through overlapping geographic data, climate information and a range of events, and then removed. The point being: an outlier at one location is only recognisable when set against other sources.
Reduce the dimensionality of the genetic data. The genetic dataset showed a large degree of overlap. Dimensionality could therefore be reduced substantially – without loss of information. A high-dimensional marker dataset with only a few thousand observations would otherwise not have been modellable.
Predict the weather to predict the yield. For unknown locations no weather data exists. By combining various space-time models, weather could be forecast at 95% accuracy – a model of its own, placed upstream of the yield model.
Approach
The reduced genetic dataset was merged with weather, soil and yield data into a single training set. On that basis, using various algorithms, we predicted hybrid performance at 75% accuracy.
Outcome
The model was successfully applied to the new hybrids at 20,000 new locations. That made it possible to determine the best-performing variety for 2017.
Transferability
The structure – reduce high-dimensional feature data, forecast environmental conditions for unknown locations, and combine both into a performance prediction – corresponds structurally to predicting compound behaviour under unknown conditions. The task of inferring from observed points to unobserved ones is the same in both fields.
Today
This project shows most clearly how the field has shifted. In 2017 every building block was bespoke work; today established standards exist for most of them.
Genomic prediction has become a field in its own right with recognised methods – genomic selection using relationship matrices rather than general dimensionality reduction, and models that represent genotype-by-environment interaction explicitly rather than learning it implicitly. Environmental description is now handled by reanalysis data such as ERA5 in combination with crop growth models; for weather forecasting itself, machine learning models have replaced classical space-time statistics and clearly outperform them. Phenotype can be captured through satellite and drone data at a resolution unavailable in 2017. And for modelling itself, tabular foundation models are candidates for datasets of this size – models that did not exist then.
What has stayed the same: the bottleneck is not the algorithm but outlier cleaning, and the question of how to describe environmental conditions for a location where nothing has yet grown.
Last updated: 30 July 2026