Chapter 2

Multiomics Techniques and Their Uses

Background

Multiomics approaches including but not limited to some combination of genome-wide association studies, transcriptomic profiling, proteomics interaction mapping, metagenomic sequencing, and metabolomic biochemical profiling, provide insights into the molecular dynamics of disease pathogenesis that is unattainable with any individual omics technology. However, harmonizing these heterogeneous datasets into comprehensive biological models remains challenging due to omics-specific analytical platforms, missing values, and differences in dynamic ranges and data dimensionality. In response, a new generation of computational methods have emerged, which combines statistical learning, artificial intelligence (AI), network biology, and causal inference to integrate diverse molecular datasets. These technologies are transforming multiomics from a descriptive research tool to an engine for mechanism-driven drug discovery, biomarker development, clinical trial optimization, and precision medicine.

Machine Learning

Among the most influential advances are latent representation learning methods, which transform multiple omics datasets into a shared low-dimensional latent representation while preserving biologically meaningful sources of variation. Algorithms including Multi-Omics Factor Analysis Plus (MOFA+), totalVI, canonical correlation analysis (CCA), and partial least squares (PLS)-based approaches identify shared biological variation despite substantial differences between sequencing, mass spectrometry, and ELISA2-4. As a result, interactions between genes, proteins, metabolites, and other omics are modeled simultaneously, enabling characterization of changes in signaling cascades, transcriptomic landscapes, and metabolic fluxes that orchestrate disease pathogenesis.

The rapid adoption of multimodal AI has helped accelerate multiomics integration. Deep neural networks, multimodal transformers, and biological foundation models are capable of learning complex, non-linear relationships across several omics datasets. Since these models are usually trained by self-supervised learning, they can discover latent biological patterns without the need for extensive manual annotation. Thanks to transfer learning, models developed in one disease area can be adapted for another, reducing the amount of experimental data needed for new applications. At the same time, AI approaches, including Shapley Additive Explanations (SHAP), integrated gradients, attention mechanisms, and feature attribution algorithms, help identify which molecules are responsible for a model’s predictions, improving biological interpretation5-7.

Computational Frameworks

Network biology has become central to many fields in the life sciences. Frameworks including Similarity Network Fusion (SNF), graph neural networks, Bayesian networks, and biological knowledge graphs enable prior biological knowledge to be combined with experimental multiomics data8-11. These network-based methods identify disease modules, characterize signaling pathways, predict previously unrecognized drug-target interactions, and reveal compensatory mechanisms that frequently underlie therapeutic resistance. Since they preserve biological context, graph-based approaches have become particularly valuable for identifying novel therapeutic targets and understanding how molecular perturbations impact biological systems.

Batch Correction and Data Harmonization

As datasets are often generated across multiple clinical sites, sequencing centers, contract research organizations, and international consortia, appropriate data harmonization is essential to avoid the masking of genuine biological signals. Algorithms such as Harmony, ComBat, mutual nearest neighbors, and deep learning-based domain adaptation methods substantially reduce platform- and batch-specific technical variation while preserving biologically meaningful signals across integrated multiomics datasets12-15. These methods enable investigators to combine legacy datasets with newly generated data, dramatically increasing statistical power while facilitating large-scale biomarker discovery and validation.

Single-Cell and Spatial Multiomics

Traditionally, molecular data is averaged across millions of cells. Single-cell and spatial multiomics measure several molecular layers within individual cells while preserving their spatial organization inside tissues. Analytical platforms such as Seurat v5 and MultiVI integrate diverse single-cell and spatial omics modalities, including transcriptomic, proteomic, metabolomic, and spatial imaging data16, 17, while downstream frameworks such as CellChat and NicheNet leverage these integrated datasets to infer cell-cell communication and reconstruct tissue cellular ecosystems at higher resolution18, 19. Collectively, these analytical frameworks have enabled investigators to reconstruct tumor microenvironments at single-cell resolution, identify rare drug-resistant cellular populations, map complex immune-cell communication networks, and visualize patients-specific differences in therapeutic response and disease progression20-24.

Causal Inference

Traditional association studies identify molecular features that correlate with disease, but they have limited ability to distinguish causal drivers from secondary consequences. Modern computational approaches, including Mendelian randomization, Bayesian causal networks, structural equation modeling, dynamic causal modeling, perturbation-based network inference, and CRISPR-informed computational models, can be combined with harmonized multiomics datasets to reconstruct causal biological pathways25-29. This shift from correlation to causation substantially improves confidence that newly identified therapeutic targets are biologically meaningful and more likely to translate into clinical benefit.

Future Directions

Beyond individual applications and sectors, harmonized multiomics is reshaping disease biology itself. Instead of classifying diseases solely by histology or clinical presentation, investigators are defining molecular endotypes based on integrated genomic, transcriptomic, proteomic, metabolomic, and immune signatures. This systems-level approach has already transformed precision oncology, where molecular classification guides therapeutic selection. Similar strategies are rapidly emerging for neurodegenerative disease, autoimmune disorders, and inflammatory bowel disease. Ultimately, novel harmonization and analytical strategies as described above are shifting disease research and drug development from descriptive molecular profiling towards predictive and mechanistic biology. These advances promise to not only accelerate drug discovery but also to improve target selection, reduce clinical trial attrition, identify robust biomarkers, and deliver truly individualized therapies.

With some of the most innovative and cutting edge multiomics techniques in mind, we will now discuss key studies that show how multiomics was used to spur scientific discovery or provide a means to improve outcomes for patients with cancer (Chapter 3), neurodegenerative diseases (Chapter 4), or inflammatory bowel disease (Chapter 5). We hope that seeing multiomics in action helps researchers gain a new appreciation for its scientific value and impact.