Most agronomically important traits do not arise from a single gene. Instead, yield, flowering time, root architecture, stress tolerance, and seed composition emerge from layered interactions among thousands of genetic variants, dynamic regulatory networks, developmental programs, and environmental cues. Capturing this complexity is one of the central problems in modern biology and a major bottleneck in plant breeding.
Genomic prediction has become the workhorse of quantitative genetics. By associating genome-wide markers with phenotypes, breeders can estimate the genetic value of individuals without waiting for field trials. Yet DNA sequence alone is a static blueprint. It records what alleles are present, but not when, where, or how strongly those alleles are expressed under a given set of conditions. For traits that depend heavily on regulation and environment, genomic models may miss a large fraction of the relevant biological signal.
A promising way to fill this gap is to integrate additional molecular layers, particularly transcriptomic and phenomic data. Gene expression profiles reflect the current physiological state of a plant, while image-derived vegetative indices can capture canopy development, biomass accumulation, and stress responses in real time. The question is whether combining these layers meaningfully improves prediction and whether the gains are robust across diverse environments and germplasm.
To test the value of integrated data, a team led by researchers at Michigan State University and the University of Nebraska assembled a multi-omics resource in Zea mays, or maize. They used approximately 750 accessions from the Wisconsin Diversity Panel and built predictive models for 129 agronomic traits measured across nine field environments.
The study compared three input layers. The genomic layer consisted of 34,153 high-quality SNP markers derived from whole-genome resequencing. The transcriptomic layer comprised 28,223 gene-expression values from leaf tissue sampled around flowering at two locations, Michigan and Nebraska. The phenomic layer included 950 vegetative indices extracted from drone imagery captured repeatedly during the growing season.
Two modeling strategies were compared. The first was rrBLUP, a linear mixed model widely used in genome-wide association studies and genomic selection. The second was support vector regression, a nonlinear machine-learning algorithm that can capture interactions among features. All models were evaluated with five-fold cross-validation so that every genotype was predicted from a model that had never seen it before.
Principal-component analysis confirmed that the three data types captured distinct information. Genotype and transcript abundance were moderately correlated, consistent with genetic control of some expression variation. Phenomic data, however, were almost orthogonal to both genotypes and transcripts, suggesting that drone-derived vegetative indices track environment-responsive processes that are not directly encoded in sequence or expression data.

Figure 1. Multi-omics models consistently outperformed single-omics approaches across 129 maize phenotypes. (Creach, et al. 2026)
Across the full set of 129 traits, models that combined genomic and transcriptomic data usually achieved the highest predictive accuracy. The average Pearson correlation between predicted and observed values was markedly higher for multi-omics models than for models trained on any single data layer. Adding phenomic data did not always improve performance further, and in some cases it slightly reduced accuracy, but it provided useful information for specific trait classes.
As expected, prediction quality was strongly related to trait heritability. Flowering-time traits, which have a robust genetic basis, were predicted most accurately. In contrast, root traits, which are heavily influenced by soil conditions and show lower heritability, remained more difficult to predict regardless of input type. Nevertheless, root traits benefited the most from including all three data layers, with multi-omics models reaching the highest mean accuracy for that category.
Performance was also trait specific. Seed composition traits were best predicted with genomic markers alone, probably because these traits are strongly genetically determined and leaf transcriptomes collected at flowering carry little information about seed biochemistry. Disease-related traits, on the other hand, gained more from transcriptomic data, likely because expression profiles capture biotic-stress responses that static markers cannot see.
Surprisingly, the nonlinear SVR approach did not consistently beat the linear rrBLUP model. A representative deep-learning framework for multi-omics prediction even performed substantially worse, most likely because the number of accessions was modest relative to the very high dimensionality of the combined feature space. These results suggest that, for many datasets, gains may come more from richer biological data than from more complex algorithms.
When the authors examined which features received the highest weights in combined genomic-plus-transcriptomic models, transcriptomic features were consistently over-represented among the top predictors. For every phenotype, expression values were more likely than SNP markers to rank in the top 5% of weighted features. This pattern was especially pronounced for flowering time and vegetative traits.
Importantly, the information contributed by genetic markers and gene expression was largely nonredundant. Markers linked to the strongest transcript predictors rarely overlapped, even when allowing for linkage disequilibrium within a 50 kb window. The finding supports the idea that transcriptome analysis adds a genuinely different dimension of biological information rather than simply acting as a proxy for genotype.
To understand the biology behind the predictions, the authors focused on yield measured in bushels per acre in the Nebraska environment. The cumulative distribution of gene weights followed a polygenic pattern: the top 1%, 10%, and 50% of transcripts accounted for only about 6%, 36%, and 86% of total model weight, respectively. No single gene dominated the prediction. Instead, accuracy arose from the coordinated contribution of hundreds to thousands of genes.
Among the most predictive transcripts were genes involved in photosynthesis, redox metabolism, protein homeostasis, and nutrient transport. Higher expression of a ferredoxin gene, HSP90.7, FLAP1, and the nitrate transporter NPF14 was associated with higher yield across both Michigan and Nebraska. In contrast, stress-responsive genes such as dehydrin15 and a DnaJ/Hsp40 family protein were elevated in lower-yielding lines. These observations are biologically interpretable, but they represent only a small fraction of the signal the model uses. This reinforces the view that yield is an emergent systems-level trait.
One of the most intriguing findings concerned prediction across environments. Because transcriptomic data were collected from the same panel grown in Michigan and Nebraska, the authors could ask whether gene expression measured at one site could predict phenotypes measured at another. The answer was yes, and sometimes surprisingly well.
For days to anthesis, a model trained on Michigan expression data actually predicted Nebraska phenotypes more accurately than Michigan phenotypes, at least when no Michigan phenotypic records were included. This suggests that the transcriptomic dataset captured stable, genotype-driven regulatory programs that generalize beyond the local environment. Combining expression data from both states further improved cross-environment prediction, pointing to the value of multi-site sampling for capturing genotype-by-environment interactions that DNA markers alone cannot reveal.
At the same time, the expression-phenotype relationship for individual flowering-time genes was environment dependent. The MADS-box gene ZMM15 correlated strongly with anthesis timing in Nebraska but not in Michigan, whereas an Aux/IAA transcription factor showed the opposite pattern. Across a broader set of known flowering genes, many showed strong correlations in one environment but weak or no correlation in the other. These gene-level results mirror the broader predictive signal: while a core developmental program is shared, environment-specific regulation shapes how individual genes relate to phenotype.
The study makes a strong case for incorporating transcriptomic and phenomic data into modern breeding pipelines. Marker-assisted breeding and genomic selection will remain essential, but adding a single transcriptomic snapshot around flowering can materially improve predictions for many traits. Because the models also identify functionally relevant genes and pathways, they can guide downstream gene function analysis and plant genetic engineering efforts.
The work also highlights the value of high-throughput phenotyping. Although image-derived vegetative indices were the weakest predictors overall, they added information for root architecture and other hard-to-measure traits. Combining drone surveys with molecular profiling could reduce the cost and labor of measuring below-ground and environmentally sensitive traits.
For service providers, these findings translate into a clear opportunity. Multi-omics analysis is no longer just a research luxury; it is becoming a practical tool for predicting complex traits, understanding genotype-by-environment interactions, and accelerating genetic gain. The key is matching the right data layer to the right trait, rather than assuming that more data types always help.
Static genomic markers have transformed breeding, but they are only one window into biological complexity. By integrating genomic, transcriptomic, and phenomic data, the new maize study shows that predictive models can capture regulatory and environmental signals that DNA alone misses. The largest gains came from adding gene expression, which provided nonredundant, biologically interpretable information across 129 traits and multiple environments. Linear models performed as well as or better than more elaborate machine-learning architectures, underscoring the importance of data quality and biological relevance over algorithmic complexity. As multi-omics datasets become larger and more temporally resolved, they are likely to become a standard component of predictive breeding and functional genomics.