Menachemi, N. & Collum, T. H. Benefits and drawbacks of electronic health record systems. Risk Manag. Healthc. Policy 4, 47–55 (2011).
Google Scholar
Jha, A. K. et al. Use of electronic health records in U.S. hospitals. N. Engl. J. Med. 360, 1628–1638 (2009).
Google Scholar
Pendergrass, S. A. & Crawford, D. C. Using electronic health records to generate phenotypes for research. Curr. Protoc. Hum. Genet. 100, e80 (2019).
Google Scholar
Chishtie, J. et al. Use of epic electronic health record system for health care research: scoping review. J. Med. Internet Res. 25, e51003 (2023).
Google Scholar
Tam, V. et al. Benefits and limitations of genome-wide association studies. Nat. Rev. Genet. 20, 467–484 (2019).
Google Scholar
Uffelmann, E. et al. Genome-wide association studies. Nat. Rev. Methods Primers 1, 59 (2021).
Google Scholar
Verma, A. et al. PheWAS and Beyond: the landscape of associations with medical diagnoses and clinical measures across 38,662 individuals from Geisinger. Am. J. Hum. Genet. 102, 592–608 (2018).
Google Scholar
Sun, J. et al. Translating polygenic risk scores for clinical use by estimating the confidence bounds of risk prediction. Nat. Commun. 12, 5276 (2021).
Google Scholar
Dudbridge, F. Power and predictive accuracy of polygenic risk scores. PLoS Genet. 9, e1003348 (2013).
Google Scholar
Jayasinghe, D., Eshetie, S., Beckmann, K., Benyamin, B. & Lee, S. H. Advancements and limitations in polygenic risk score methods for genomic prediction: a scoping review. Hum. Genet. 143, 1401–1431 (2024).
Google Scholar
Kirchler, M. et al. Large language models improve transferability of electronic health record-based predictions across countries and coding systems. npj Digit. Med. 9, 177 (2026).
Google Scholar
Meng, X. et al. The application of large language models in medicine: a scoping review. iScience 27, 109713 (2024).
Google Scholar
Liao, K. P. et al. Development of phenotype algorithms using electronic medical records and incorporating natural language processing. BMJ 350, h1885 (2015).
Google Scholar
Guevara, M. et al. Large language models to identify social determinants of health in electronic health records. npj Digit. Med. 7, 6 (2024).
Google Scholar
Kline, A. et al. Multimodal machine learning in precision health: A scoping review. npj Digit. Med. 5, 171 (2022).
Google Scholar
Subramanian, I., Verma, S., Kumar, S., Jere, A. & Anamika, K. Multi-omics data integration, interpretation, and its application. Bioinform. Biol. Insights 14, 1177932219899051 (2020).
Google Scholar
Amirahmadi, A., Ohlsson, M. & Etminani, K. Deep learning prediction models based on EHR trajectories: a systematic review. J. Biomed. Inform. 144, 104430 (2023).
Google Scholar
Tong, L. et al. Integrating multi-omics data with EHR for precision medicine using advanced artificial intelligence. IEEE Rev. Biomed. Eng. 17, 80–97 (2024).
Google Scholar
Gallagher, C. S., Ginsburg, G. S. & Musick, A. Biobanking with genetics shapes precision medicine and global health. Nat. Rev. Genet. 26, 191–202 (2025).
Google Scholar
The All of Us Research Program Investigators The ‘All of Us’ Research Program. N. Engl. J. Med. 381, 668–676 (2019).
Google Scholar
Bick, A. G. et al. Genomic data in the All of Us Research Program. Nature 627, 340–346 (2024).
Google Scholar
Bycroft, C. et al. The UK Biobank resource with deep phenotyping and genomic data. Nature 562, 203–209 (2018).
Google Scholar
Allen, N. E. et al. Prospective study design and data analysis in UK Biobank. Sci. Transl. Med. 16, eadf4428 (2024).
Google Scholar
Kurki, M. I. et al. FinnGen provides genetic insights from a well-phenotyped isolated population. Nature 613, 508–518 (2023).
Google Scholar
Leitsalu, L. et al. Cohort profile: Estonian Biobank of the Estonian Genome Center, University of Tartu. Int. J. Epidemiol. 44, 1137–1147 (2015).
Google Scholar
Nagai, A. et al. Overview of the BioBank Japan Project: study design and profile. J. Epidemiol. 27, S2–S8 (2017).
Google Scholar
Kim, Y., Han, B.-G. & the KoGES group Cohort profile: the Korean genome and epidemiology study (KoGES) consortium. Int. J. Epidemiol. 46, e20 (2017).
Google Scholar
Gaziano, J. M. et al. Million Veteran Program: a mega-biobank to study genetic influences on health and disease. J. Clin. Epidemiol. 70, 214–223 (2016).
Google Scholar
McGregor, T. L. et al. Inclusion of pediatric samples in an opt-out biorepository linking DNA to de-identified medical records: pediatric BioVU. Clin. Pharmacol. Ther. 93, 204–211 (2013).
Google Scholar
Carey, D. J. et al. The Geisinger MyCode community health initiative: an electronic health record–linked biobank for precision medicine research. Genet. Med. 18, 906–913 (2016).
Google Scholar
Verma, A. et al. The Penn Medicine BioBank: towards a genomics-enabled learning healthcare system to accelerate precision medicine in a diverse population. J. Pers. Med. 12, 1974 (2022).
Google Scholar
Chen, Z. et al. China Kadoorie Biobank of 0.5 million people: survey methods, baseline characteristics and long-term follow-up. Int. J. Epidemiol. 40, 1652–1666 (2011).
Google Scholar
Cook, M. B. et al. Our Future Health: a unique global resource for discovery and translational research. Nat. Med. 31, 728–730 (2025).
Google Scholar
Cancer Genome Atlas Research Network et al. The Cancer Genome Atlas Pan-Cancer analysis project. Nat. Genet. 45, 1113–1120 (2013).
Google Scholar
Perez-Riverol, Y. et al. Discovering and linking public omics data sets using the Omics Discovery Index. Nat. Biotechnol. 35, 406–409 (2017).
Google Scholar
Zhou, W. et al. Global Biobank Meta-analysis Initiative: powering genetic discovery across human disease. Cell Genom. 2, 100192 (2022).
Google Scholar
Beesley, L. J. et al. The emerging landscape of health research based on biobanks linked to electronic health records: existing resources, statistical challenges and potential opportunities. Stat. Med. 39, 773–800 (2020).
Google Scholar
Robinson, J. R., Wei, W.-Q., Roden, D. M. & Denny, J. C. Defining phenotypes from clinical data to drive genomic research. Annu. Rev. Biomed. Data Sci. 1, 69–92 (2018).
Google Scholar
Ueda, D. et al. Fairness of artificial intelligence in healthcare: review and recommendations. Jpn. J. Radiol. 42, 3–15 (2024).
Google Scholar
Al-Sahab, B., Leviton, A., Loddenkemper, T., Paneth, N. & Zhang, B. Biases in electronic health records data for generating real-world evidence: an overview. J. Healthc. Inform. Res. 8, 121–139 (2024).
Google Scholar
Boyd, A. D. et al. Equity and bias in electronic health records data. Contemp. Clin. Trials 130, 107238 (2023).
Google Scholar
Kachuri, L. et al. Principles and methods for transferring polygenic risk scores across global populations. Nat. Rev. Genet. 25, 8–25 (2024).
Google Scholar
Ko, S. et al. Unsupervised discovery of ancestry-informative markers and genetic admixture proportions in biobank-scale datasets. Am. J. Hum. Genet. 110, 314–325 (2023).
Google Scholar
Venkatesh, R. et al. Importance of genetic ancestry in pharmacogenomics for precision medicine. Pharmacogenomics 26, 747–762 (2025).
Google Scholar
Popejoy, A. B. et al. The clinical imperative for inclusivity: race, ethnicity, and ancestry (REA) in genomics. Hum. Mutat. 39, 1713–1720 (2018).
Google Scholar
Marees, A. T. et al. A tutorial on conducting genome-wide association studies: quality control and statistical analysis. Int. J. Methods Psychiatr. Res. 27, e1608 (2018).
Google Scholar
Carss, K. et al. Whole-genome sequencing of 490,640 UK Biobank participants. Nature 645, 692–701 (2025).
Google Scholar
Turner, S. et al. Quality control procedures for genome-wide association studies. Curr. Protoc. Hum. Genet. 1, 19 (2011).
McLaren, W. et al. The Ensembl Variant Effect Predictor. Genome Biol. 17, 122 (2016).
Google Scholar
Wang, K., Li, M. & Hakonarson, H. ANNOVAR: functional annotation of genetic variants from high-throughput sequencing data. Nucleic Acids Res. 38, e164 (2010).
Google Scholar
Landrum, M. J. et al. ClinVar: public archive of interpretations of clinically relevant variants. Nucleic Acids Res. 44, D862–D868 (2016).
Google Scholar
Hamosh, A., Scott, A. F., Amberger, J. S., Bocchini, C. A. & McKusick, V. A. Online Mendelian Inheritance in Man (OMIM), a knowledgebase of human genes and genetic disorders. Nucleic Acids Res. 33, D514–D517 (2005).
Google Scholar
Chen, S. et al. A genomic mutational constraint map using variation in 76,156 human genomes. Nature 625, 92–100 (2024).
Google Scholar
Price, A. L. et al. Principal components analysis corrects for stratification in genome-wide association studies. Nat. Genet. 38, 904–909 (2006).
Google Scholar
Taliun, D. et al. Sequencing of 53,831 diverse genomes from the NHLBI TOPMed Program. Nature 590, 290–299 (2021).
Google Scholar
Auton, A. et al. A global reference for human genetic variation. Nature 526, 68–74 (2015).
Google Scholar
Reich, D., Price, A. L. & Patterson, N. Principal component analysis of genetic data. Nat. Genet. 40, 491–492 (2008).
Google Scholar
Zhang, D., Yin, C., Zeng, J., Yuan, X. & Zhang, P. Combining structured and unstructured data for predictive models: a deep learning approach. BMC Med. Inform. Decis. Mak. 20, 280 (2020).
Google Scholar
O’Malley, K. J. et al. Measuring diagnoses: ICD code accuracy. Health Serv. Res. 40, 1620–1639 (2005).
Google Scholar
Chang, E. & Mostafa, J. The use of SNOMED CT, 2013-2020: a literature review. J. Am. Med. Inform. Assoc. 28, 2017–2026 (2021).
Google Scholar
McDonald, C. J. et al. LOINC, a universal standard for identifying laboratory observations: a 5-year update. Clin. Chem. 49, 624–633 (2003).
Google Scholar
Liu, S., Ma, W., Moore, R., Ganesan, V. & Nelson, S. RxNorm: prescription for electronic drug information exchange. IT Prof. 7, 17–23 (2005).
Google Scholar
Nelson, S. J., Zeng, K., Kilbourne, J., Powell, T. & Moore, R. Normalized names for clinical drugs: RxNorm at 6 years. J. Am. Med. Inform. Assoc. 18, 441–448 (2011).
Google Scholar
CPT® code set overview. American Medical Association (2026).
Reinecke, I. et al. The usage of OHDSI OMOP – a scoping review. Stud. Health Technol. Inform. 283, 95–103 (2021).
Google Scholar
Schuemie, M. et al. Health-Analytics Data to Evidence Suite (HADES): open-source software for observational research. Stud. Health Technol. Inform. 310, 966–970 (2024).
Google Scholar
Klann, J. G., Joss, M. A. H., Embree, K. & Murphy, S. N. Data model harmonization for the All Of Us Research Program: Transforming i2b2 data into the OMOP common data model. PLoS One 14, e0212463 (2019).
Google Scholar
Vasilevsky, N. A. et al. Mondo: integrating disease terminology across communities. Genetics 232, iyaf215 (2026).
Google Scholar
Robinson, P. N. et al. The Human Phenotype Ontology: a tool for annotating and analyzing human hereditary disease. Am. J. Hum. Genet. 83, 610–615 (2008).
Google Scholar
Köhler, S. et al. The Human Phenotype Ontology in 2021. Nucleic Acids Res. 49, D1207–D1217 (2021).
Google Scholar
Qualls, L. G. Evaluating foundational data quality in the national patient-centered clinical research network (PCORnet®). EGEMS 6, 3 (2018).
Google Scholar
Antunes, R. S., André da Costa, C., Küderle, A., Yari, I. A. & Eskofier, B. Federated learning for healthcare: systematic review and architecture proposal. ACM Trans. Intell. Syst. Technol. 13, 54 (2022).
Google Scholar
Hegselmann, S. et al. Large language models are powerful electronic health record encoders. npj Digit. Med. 9, 530 (2026).
Google Scholar
Chang, C. C. et al. Second-generation PLINK: rising to the challenge of larger and richer datasets. GigaScience 4, (2015).
Zhou, W. et al. Efficiently controlling for case-control imbalance and sample relatedness in large-scale genetic association studies. Nat. Genet. 50, 1335–1341 (2018).
Google Scholar
Danecek, P. et al. Twelve years of SAMtools and BCFtools. Gigascience 10, giab008 (2021).
Google Scholar
Ramirez, A. H. et al. The All of Us Research Program: data quality, utility, and diversity. Patterns 3, 100570 (2022).
Google Scholar
Shih, C. C. et al. A five-safes approach to a secure and scalable genomics data repository. iScience 26, 106546 (2023).
Google Scholar
Silva, S. et al. Fed-BioMed: A general open-source frontend framework for federated learning in healthcare. In Domain Adaptation and Representation Transfer, and Distributed and Collaborative Learning (eds Albarqouni, S. et al.) 201–210 (Springer, 2020).
Beutel, D. J. et al. Flower: a friendly federated learning research framework. Preprint at (2022).
Xu, J. et al. Federated learning for healthcare informatics. J. Healthc. Inform. Res. 5, 1–19 (2021).
Google Scholar
Silva, S. et al. Federated learning in distributed medical databases: meta-analysis of large-scale subcortical brain data. In Proc. 2019 IEEE 16th International Symposium on Biomedical Imaging (ISBI 2019) 270–274 (IEEE, 2019).
Cooray, L., Sendanayake, J., Vithanaarachchi, P. & Priyadarshana, Y. H. P. P. Deep federated learning: a systematic review of methods, applications, and challenges. Front. Comput. Sci. 7, 1617597 (2025).
Google Scholar
Kirby, J. C. et al. PheKB: a catalog and workflow for creating electronic phenotype algorithms for transportability. J. Am. Med. Inform. Assoc. 23, 1046–1052 (2016).
Google Scholar
Newton, K. M. et al. Validation of electronic medical record-based phenotyping algorithms: results and lessons learned from the eMERGE network. J. Am. Med. Inform. Assoc. 20, e147–e154 (2013).
Google Scholar
Hripcsak, G. et al. Observational health data sciences and informatics (OHDSI): opportunities for observational researchers. Stud. Health Technol. Inform. 216, 574–578 (2015).
Google Scholar
Kho, A. N. et al. Electronic medical records for genetic research: results of the eMERGE consortium. Sci. Transl. Med. 3, 79re1 (2011).
Google Scholar
Banda, J. M., Seneviratne, M., Hernandez-Boussard, T. & Shah, N. H. Advances in electronic phenotyping: from rule-based definitions to machine learning models. Annu. Rev. Biomed. Data Sci. 1, 53–68 (2018).
Google Scholar
Zitnik, M. et al. Machine learning for integrating data in biology and medicine: principles, practice, and opportunities. Inf. Fusion. 50, 71–91 (2019).
Google Scholar
Ahuja, Y., Zou, Y., Verma, A., Buckeridge, D. & Li, Y. MixEHR-Guided: a guided multi-modal topic modeling approach for large-scale automatic phenotyping using the electronic health record. J. Biomed. Inform. 134, 104190 (2022).
Google Scholar
Yang, S., Varghese, P., Stephenson, E., Tu, K. & Gronsbell, J. Machine learning approaches for electronic health records phenotyping: a methodical review. J. Am. Med. Inform. Assoc. 30, 367–381 (2023).
Google Scholar
Luo, L. et al. PhenoTagger: a hybrid method for phenotype concept recognition using human phenotype ontology. Bioinformatics 37, 1884–1890 (2021).
Google Scholar
Liao, K. P. et al. High-throughput multimodal automated phenotyping (MAP) with application to PheWAS. J. Am. Med. Inform. Assoc. 26, 1255–1262 (2019).
Google Scholar
Adamson, B. et al. Approach to machine learning for extraction of real-world data variables from electronic health records. Front. Pharmacol. 14, 1180962 (2023).
Google Scholar
Alzoubi, H. et al. A review of automatic phenotyping approaches using electronic health records. Electronics 8, 1235 (2019).
Google Scholar
Gao, Y. & Cui, Y. Clinical time-to-event prediction enhanced by incorporating compatible related outcomes. PLoS Digit. Health 1, e0000038 (2022).
Google Scholar
Li, Y. et al. Validation of risk prediction models applied to longitudinal electronic health record data for the prediction of major cardiovascular events in the presence of data shifts. Eur. Heart J. Digit. Health 3, 535–547 (2022).
Google Scholar
Ramachandram, D. & Taylor, G. W. Deep multimodal learning: a survey on recent advances and trends. IEEE Signal Process. Mag. 34, 96–108 (2017).
Google Scholar
Guo, A., Beheshti, R., Khan, Y. M., Langabeer, J. R. & Foraker, R. E. Predicting cardiovascular health trajectories in time-series electronic health records with LSTM models. BMC Med. Inform. Decis. Mak. 21, 5 (2021).
Google Scholar
Pham, T., Tran, T., Phung, D. & Venkatesh, S. DeepCare: A deep dynamic memory model for predictive medicine. In Advances in Knowledge Discovery and Data Mining (eds Bailey, J. et al.) 30–41 (Springer, 2016).
Choi, E. et al. RETAIN: an interpretable predictive model for healthcare using reverse time attention mechanism. In Proc. 30th International Conference on Neural Information Processing Systems (eds Lee, D. D. et al.) 3512–3520 (Curran Associates Inc., 2016).
Li, Y. et al. BEHRT: transformer for electronic health records. Sci. Rep. 10, 7155 (2020).
Google Scholar
Niu, H. et al. EHR-BERT: a BERT-based model for effective anomaly detection in electronic health records. J. Biomed. Inform. 150, 104605 (2024).
Google Scholar
Lin, K.-W., Kuo, Y.-C., Wang, H.-Y. & Tseng, Y.-J. KAT-GNN: a knowledge-augmented temporal graph neural network for risk prediction in electronic health records. Preprint at (2025).
Gao, Y. et al. Precision adverse drug reactions prediction with heterogeneous graph neural network. Adv. Sci. 12, 2404671 (2025).
Google Scholar
Lahoti, A. et al. Mamba-3: improved sequence modeling using state space principles. Preprint at (2026).
Reátegui, R. & Ratté, S. Comparison of MetaMap and cTAKES for entity extraction in clinical notes. BMC Med. Inform. Decis. Mak. 18, 74 (2018).
Google Scholar
Savova, G. K. et al. Mayo clinical Text Analysis and Knowledge Extraction System (cTAKES): architecture, component evaluation and applications. J. Am. Med. Inform. Assoc. 17, 507–513 (2010).
Google Scholar
Aronson, A. R. Effective mapping of biomedical text to the UMLS Metathesaurus: the MetaMap program. In Proc. AMIA Symposium 2001 17–21 (American Medical Informatics Association, 2001).
Kang, T. et al. EliIE: an open-source information extraction system for clinical trial eligibility criteria. J. Am. Med. Inform. Assoc. 24, 1062–1071 (2017).
Google Scholar
Lee, J. et al. BioBERT: a pre-trained biomedical language representation model for biomedical text mining. Bioinformatics 36, 1234–1240 (2020).
Google Scholar
Yu, X., Hu, W., Lu, S., Sun, X. & Yuan, Z. BioBERT based named entity recognition in electronic medical record. In Proc. 2019 10th International Conference on Information Technology in Medicine and Education (ITME) 49–52 (IEEE, 2019).
Wu, J. et al. A hybrid framework with large language models for rare disease phenotyping. BMC Med. Inform. Decis. Mak. 24, 289 (2024).
Google Scholar
Alsentzer, E. et al. Publicly available clinical BERT embeddings. In Proc. 2nd Clinical Natural Language Processing Workshop (eds Rumshisky, A. et al.) 72–78 (Association for Computational Linguistics, 2019).
Yang, X. et al. A large language model for electronic health records. npj Digit. Med. 5, 194 (2022).
Google Scholar
Khan, S. N., Danishuddin, Khan, M. W. A., Guarnera, L. & Akhtar, S. M. F. Multi-modal AI in precision medicine: integrating genomics, imaging, and EHR data for clinical insights. Front. Artif. Intell. 8, 1743921 (2026).
Google Scholar
Tang, J., Yin, X., Lai, J., Luo, K. & Wu, D. Fusion of X-ray images and clinical data for a multimodal deep learning prediction model of osteoporosis: algorithm development and validation study. JMIR Med. Inform. 13, e70738 (2025).
Google Scholar
Wei, H., Liu, B., Zhang, M., Shi, P. & Yuan, W. VisionCLIP: an Med-AIGC based ethical language-image foundation model for generalizable retina image analysis. Preprint at (2024).
Sellergren, A. et al. MedGemma technical report. Preprint at (2025).
Prottasha, M. S. I. & Rafi, N. W. MedGemma vs GPT-4: open-source and proprietary zero-shot medical disease classification from images. Preprint at (2025).
Wang, L. Cross-lingual NLP: bridging language barriers with multilingual model. In Proc. 2024 International Conference on Electronics and Devices, Computational Science (ICEDCS) 1005–1012 (IEEE, 2024).
Sindane, T., Marivate, V. & Modupe, A. Cross-lingual embedding methods and applications: a systematic review for low-resourced scenarios. Nat. Lang. Proc. J. 12, 100157 (2025).
Qin, L. et al. A survey of multilingual large language models. Patterns 6, 101118 (2025).
Google Scholar
Guo, P. et al. Steering large language models for cross-lingual information retrieval. In Proc. 47th International ACM SIGIR Conference on Research and Development in Information Retrieval 585–596 (Association for Computing Machinery, 2024).
Shah, A. Efficient cross-lingual transfer for language models. In Proc. 2025 9th International Symposium on Innovative Approaches in Smart Technologies (ISAS) 1–9 (IEEE, 2025).
Brandes, N., Linial, N. & Linial, M. PWAS: proteome-wide association study—linking genes and phenotypes by functional variation in proteins. Genome Biol. 21, 173 (2020).
Google Scholar
Cheng, B. et al. Integrated analysis of proteome-wide and transcriptome-wide association studies identified novel genes and chemicals for vertigo. Brain Commun. 4, fcac313 (2022).
Google Scholar
Evans, P. et al. Transcriptome-wide association studies (TWAS): methodologies, applications, and challenges. Curr. Protoc. 4, e981 (2024).
Google Scholar
Gamazon, E. R. et al. A gene-based association method for mapping traits using reference transcriptome data. Nat. Genet. 47, 1091–1098 (2015).
Google Scholar
Barbeira, A. N. et al. Exploiting the GTEx resources to decipher the mechanisms at GWAS loci. Genome Biol. 22, 49 (2021).
Google Scholar
Gusev, A. et al. Integrative approaches for large-scale transcriptome-wide association studies. Nat. Genet. 48, 245–252 (2016).
Google Scholar
Georgakis, M. K. et al. Genetically downregulated interleukin-6 signaling is associated with a favorable cardiometabolic profile. Circulation 143, 1177–1180 (2021).
Google Scholar
Topaloudi, A. et al. PheWAS and cross-disorder analysis reveal genetic architecture, pleiotropic loci and phenotypic correlations across 11 autoimmune disorders. Front. Immunol. 14, 1147573 (2023).
Google Scholar
Verma, A. et al. Human-disease phenotype map derived from PheWAS across 38,682 individuals. Am. J. Hum. Genet. 104, 55–64 (2019).
Google Scholar
Ritchie, M. D., Holzinger, E. R., Li, R., Pendergrass, S. A. & Kim, D. Methods of integrating data to uncover genotype–phenotype interactions. Nat. Rev. Genet. 16, 85–97 (2015).
Google Scholar
Žitnik, M. & Zupan, B. Data fusion by matrix factorization. IEEE Trans. Pattern Anal. Mach. Intell. 37, 41–53 (2015).
Google Scholar
Iribarren, C. et al. Polygenic risk and incident coronary heart disease in a large multiethnic cohort. Am. J. Prev. Cardiol. 18, 100661 (2024).
Google Scholar
Ritchie, S. C. et al. Combined clinical, metabolomic, and polygenic scores for cardiovascular risk prediction. Eur. Heart J. 47, 1861–1873 (2026).
Google Scholar
Kim, H., Lee, G. & Chung, W. Multimodal deep learning approaches for improving polygenic risk scores with imaging data. Sci. Rep. 16, 4012 (2026).
Google Scholar
Tong, L., Mitchel, J., Chatlin, K. & Wang, M. D. Deep learning based feature-level integration of multi-omics data for breast cancer patients survival analysis. BMC Med. Inform. Decis. Mak. 20, 225 (2020).
Google Scholar
Farhadizadeh, M. et al. Challenges and proposed solutions in modeling multimodal data: a systematic review. Preprint at (2025).
Wu, K.-H. H. et al. Integrating large scale genetic and clinical information to predict cases of heart failure. Commun. Med. 5, 493 (2025).
Google Scholar
Mataraso, S. J. et al. A machine learning approach to leveraging electronic health records for enhanced omics analysis. Nat. Mach. Intell. 7, 293–306 (2025).
Google Scholar
Eijpe, A. et al. Disentangled and interpretable multimodal attention fusion for cancer survival prediction. In Medical Image Computing and Computer Assisted Intervention – MICCAI 2025 (eds Gee, J. C. et al.) 117–127 (Springer, 2025).
Lin, J. et al. Integration of biomarker polygenic risk score improves prediction of coronary heart disease. JACC Basic Transl. Sci. 8, 1489–1499 (2023).
Google Scholar
Türkmen, D. et al. Polygenic scores for cardiovascular risk factors improve estimation of clinical outcomes in CCB treatment compared to pharmacogenetic variants alone. Pharmacogenomics J. 24, 12 (2024).
Google Scholar
Liu, Y., Huse, J. & Kannan, K. Expression graph network framework for biomarker discovery. Brief. Bioinform. 26, bbaf559 (2025).
Google Scholar
Wang, T. et al. MOGONET integrates multi-omics data using graph convolutional networks allowing patient classification and biomarker identification. Nat. Commun. 12, 3445 (2021).
Google Scholar
Amar, J. et al. Integrating genomics into multimodal EHR foundation models. Preprint at (2025).
Nguyen, E. et al. Sequence modeling and design from molecular to genome scale with Evo. Science 386, eado9336 (2024).
Google Scholar
Brixi, G. et al. Genome modelling and design across all domains of life with Evo 2. Nature 652, 1349–1361 (2026).
Google Scholar
Dalla-Torre, H. et al. Nucleotide Transformer: building and evaluating robust foundation models for human genomics. Nat. Methods 22, 287–297 (2025).
Google Scholar
Nguyen, E. et al. HyenaDNA: long-range genomic sequence modeling at single nucleotide resolution. In Proc. 37th International Conference on Neural Information Processing Systems (eds Oh, A. et al.) 43177–43201 (Curran Associates Inc., 2023).
Cui, H. et al. scGPT: toward building a foundation model for single-cell multi-omics using generative AI. Nat. Methods 21, 1470–1480 (2024).
Google Scholar
Hao, M. et al. Large-scale foundation model on single-cell transcriptomics. Nat. Methods 21, 1481–1491 (2024).
Google Scholar
Steinberg, E. et al. Language models are an effective representation learning technique for electronic health record data. J. Biomed. Inform. 113, 103637 (2021).
Google Scholar
Wornow, M., Thapa, R., Steinberg, E., Fries, J. A. & Shah, N. H. EHRSHOT: an EHR benchmark for few-shot evaluation of foundation models. In Advances in Neural Information Processing Systems (eds Oh, A. et al.) 67125–67137 (Curran Associates, Inc., 2023).
Renc, P. et al. Zero shot health trajectory prediction using transformer. npj Digit. Med. 7, 256 (2024).
Google Scholar
Steinberg, E., Fries, J., Xu, Y. & Shah, N. MOTOR: a time-to-event foundation model for structured medical records. In International Conference on Learning Representations (eds Kim, B. et al.) 41038–41077 (ICLR, 2024).
Zhao, T. et al. BiomedParse: a biomedical foundation model for image parsing of everything everywhere all at once. Nat. Methods 22, 166–176 (2025).
Google Scholar
Wu, C., Zhang, X., Zhang, Y., Wang, Y. & Xie, W. Towards generalist foundation model for radiology by leveraging web-scale 2D&3D medical data. Nat Commun. 16, 7866 (2025).
Google Scholar
Cardoso, M. J. et al. MONAI: an open-source framework for deep learning in healthcare. Preprint at (2022).
Yang, L. et al. Advancing multimodal medical capabilities of Gemini. Preprint at (2024).
Ahlqvist, E. et al. Novel subgroups of adult-onset diabetes and their association with outcomes: a data-driven cluster analysis of six variables. Lancet Diabetes Endocrinol. 6, 361–369 (2018).
Google Scholar
Qiu, J. et al. Deep representation learning for clustering longitudinal survival data from electronic health records. Nat. Commun. 16, 2534 (2025).
Google Scholar
Soman, K. et al. Early detection of Parkinson’s disease through enriching the electronic health record using a biomedical knowledge graph. Front. Med. 10, 1081087 (2023).
Google Scholar
Li, C. et al. Improving cardiovascular risk prediction through machine learning modelling of irregularly repeated electronic health records. Eur. Heart J. Digit. Health 5, 30–40 (2024).
Google Scholar
Shen, L. et al. Genetic analysis of quantitative phenotypes in AD and MCI: imaging, cognition and biomarkers. Brain Imaging Behav. 8, 183–207 (2014).
Google Scholar
Crane, P. K. et al. Cognitively defined Alzheimer’s dementia subgroups have distinct atrophy patterns. Alzheimers Dement. 20, 1739–1752 (2024).
Google Scholar
Ferraro, P. M. et al. Clinical and biological underpinnings of longitudinal atrophy pattern progression in Alzheimer’s disease. J. Alzheimer’s Dis. 103, 243–255 (2025).
Google Scholar
Higginbotham, L. et al. Unbiased classification of the elderly human brain proteome resolves distinct clinical and pathophysiological subtypes of cognitive impairment. Neurobiology of Disease 186, 106286 (2023).
Hähnel, T. et al. Progression subtypes in Parkinson’s disease identified by a data-driven multi cohort analysis. npj Parkinsons Dis. 10, 95 (2024).
Google Scholar
Krix, S. et al. MultiGML: multimodal graph machine learning for prediction of adverse drug events. Heliyon 9, e19441 (2023).
Google Scholar
Adam, G. et al. Machine learning approaches to drug response prediction: challenges and recent progress. npj Precis. Oncol. 4, 19 (2020).
Google Scholar
Ozery-Flato, M., Goldschmidt, Y., Shaham, O., Ravid, S. & Yanover, C. Framework for identifying drug repurposing candidates from observational healthcare data. JAMIA Open 3, 536–544 (2020).
Google Scholar
Hernán, M. A., Wang, W. & Leaf, D. E. Target trial emulation: a framework for causal inference from observational data. JAMA 328, 2446–2447 (2022).
Google Scholar
Kaufman, H., Rappoport, N., Gilad, A. & Linial, M. Advancing causal inference in medicine using biobank data. J. Biomed. Inform. 171, 104903 (2025).
Google Scholar
Yao, M., Wang, A., Li, X. & Liu, Z. Mendelian randomization methods for causal inference: estimands, identification and inference. Stat. Med. 45, e70394 (2026).
Google Scholar
Al Khzem, A. H. & Wali, S. M. Drug repurposing as an effective drug discovery strategy: a critical review. Drug Des. Dev. Ther. 19, 12019–12034 (2025).
Google Scholar
Lau-Min, K. S. et al. Impact of integrating genomic data into the electronic health record on genetics care delivery. Genet. Med. 24, 2338–2350 (2022).
Google Scholar
Rasmussen-Torvik, L. J. et al. Design and anticipated outcomes of the eMERGE-PGx project: a multi-center pilot for pre-emptive pharmacogenomics in electronic health record systems. Clin. Pharmacol. Ther. 96, 482–489 (2014).
Google Scholar
Murugan, M. et al. Genomic considerations for FHIR®; eMERGE implementation lessons. J. Biomed. Inform. 118, 103795 (2021).
Google Scholar
Cheng, D. T. et al. Memorial Sloan Kettering-integrated mutation profiling of actionable cancer targets (MSK-IMPACT): a hybridization capture-based next-generation sequencing clinical assay for solid tumor molecular oncology. J. Mol. Diagn. 17, 251–264 (2015).
Google Scholar
Klein, H. et al. MatchMiner: an open-source platform for cancer precision medicine. npj Precis. Oncol. 6, 69 (2022).
Google Scholar
Vassy, J. L. et al. Genomic risk model to implement precision prostate cancer screening in clinical care: the ProGRESS study. Nat. Cancer 7, 352–367 (2026).
Google Scholar
Geisinger. Geisinger launches ‘MyCode-Connect’: a new era of precision health integrating real-world data and multi-omics. Geisinger News Releases (2026).
Iribarren, C. et al. Abstract 4355586: Enhancing the prevent equation with a polygenic risk score: clinical utility evaluation. Circulation 152, A4355586 (2025).
Google Scholar
Rockowitz, S. et al. Children’s rare disease cohorts: an integrative research and clinical genomics initiative. npj Genom. Med. 5, 29 (2020).
Google Scholar
Martin, A. R. et al. Human demographic history impacts genetic risk prediction across diverse populations. Am. J. Hum. Genet. 100, 635–649 (2017).
Google Scholar
Linder, J. E. et al. Returning integrated genomic risk and clinical recommendations: the eMERGE study. Genet. Med. 25, 100006 (2023).
Google Scholar
Dolin, R. H., Boxwala, A. & Shalaby, J. A pharmacogenomics clinical decision support service based on FHIR and CDS Hooks. Methods Inf. Med. 57, e115–e123 (2018).
Google Scholar
Mandel, J. C., Kreda, D. A., Mandl, K. D., Kohane, I. S. & Ramoni, R. B. SMART on FHIR: a standards-based, interoperable apps platform for electronic health records. J. Am. Med. Inform. Assoc. 23, 899–908 (2016).
Google Scholar
Adnan, M., Kalra, S., Cresswell, J. C., Taylor, G. W. & Tizhoosh, H. R. Federated learning and differential privacy for medical image analysis. Sci. Rep. 12, 1953 (2022).
Google Scholar
Ramos, E. et al. Pharmacogenomics, ancestry and clinical decision making for global populations. Pharmacogenomics J. 14, 217–222 (2014).
Google Scholar
Huang, R. et al. Evaluation and bias analysis of large language models in generating synthetic electronic health records: comparative study. J. Med. Internet Res. 27, e65317 (2025).
Google Scholar
Kullo, I. J. et al. Polygenic scores in biomedical research. Nat. Rev. Genet. 23, 524–532 (2022).
Google Scholar
Abràmoff, M. D. et al. Considerations for addressing bias in artificial intelligence for health equity. npj Digit. Med. 6, 170 (2023).
Google Scholar
Azad, T. D., Krumholz, H. M. & Saria, S. Principles to guide clinical AI readiness and move from benchmarks to real-world evaluation. Nat. Med. 32, 802–804 (2026).
Google Scholar
Hassija, V. et al. Interpreting black-box models: a review on explainable artificial intelligence. Cogn. Comput. 16, 45–74 (2024).
Google Scholar
Loh, H. W. et al. Application of explainable artificial intelligence for healthcare: a systematic review of the last decade (2011–2022). Comput. Methods Programs Biomed. 226, 107161 (2022).
Google Scholar
Linardatos, P., Papastefanopoulos, V. & Kotsiantis, S. Explainable AI: a review of machine learning interpretability methods. Entropy 23, 18 (2021).
Google Scholar
Watson, J. et al. Overcoming barriers to the adoption and implementation of predictive modeling and machine learning in clinical care: what can we learn from US academic medical centers? JAMIA Open 3, 167–172 (2020).
Google Scholar
Singh, V., Cheng, S., Kwan, A. C. & Ebinger, J. United States food and drug administration regulation of clinical software in the era of artificial intelligence and machine learning. Mayo Clin. Proc. Digit. Health 3, 100231 (2025).
Google Scholar
Atmaca, U. I. et al. Data-driven medical devices and the EU MDR: mapping gaps in standards for regulatory compliance. npj Health Syst. 3, 21 (2026).
Google Scholar
Meszaros, J., Minari, J. & Huys, I. The future regulation of artificial intelligence systems in healthcare services and medical research in the European Union. Front. Genet. 13, 927721 (2022).
Google Scholar
Karunanayake, N. Next-generation agentic AI for transforming healthcare. Inform. Health 2, 73–83 (2025).
Google Scholar
Liu, F. et al. A foundational architecture for AI agents in healthcare. Cell Rep. Med. 6, 102374 (2025).
Google Scholar
Waight, M. C. et al. Personalized heart digital twins detect substrate abnormalities in scar-dependent ventricular tachycardia. Circulation 151, 521–533 (2025).
Google Scholar
Pan, R., Sun, H., Chen, X., Pedrielli, G. & Huang, J. Human digital twin: data, models, applications, and challenges. Preprint at (2025).
DeMeo, B. et al. Active learning framework leveraging transcriptomics identifies modulators of disease phenotypes. Science 390, eadi8577 (2025).
Google Scholar
Boshar, S. et al. A foundational model for joint sequence-function multi-species modeling at scale for long-range genomic prediction. Preprint at bioRxiv (2025).
Shen, T. et al. Accurate RNA 3D structure prediction using a language model-based deep learning approach. Nat. Methods 21, 2287–2298 (2024).
Google Scholar
Chen, J. et al. Interpretable RNA foundation model from unannotated data for highly accurate RNA structure and function predictions. Preprint at (2022).
Peng, C. et al. A study of generative large language model for medical research and healthcare. npj Digit. Med. 6, 210 (2023).
Google Scholar
Inouye, M. et al. Genomic risk prediction of coronary artery disease in 480,000 adults. J. Am. Coll. Cardiol. 72, 1883–1893 (2018).
Google Scholar
Verma, S. S. et al. Evaluating the frequency and the impact of pharmacogenetic alleles in an ancestrally diverse Biobank population. J. Transl. Med. 20, 550 (2022).
Google Scholar
Zhu, J., Liu, H., Liu, X., Chen, C. & Shu, M. Cardiovascular disease detection based on deep learning and multi-modal data fusion. Biomed. Signal Process. Control 99, 106882 (2025).
Google Scholar
Zhang, X. et al. Data-driven subtyping of Parkinson’s disease using longitudinal clinical records: a cohort study. Sci. Rep. 9, 797 (2019).
Google Scholar
Nelson, C. A., Bove, R., Butte, A. J. & Baranzini, S. E. Embedding electronic health records onto a knowledge network recognizes prodromal features of multiple sclerosis and predicts diagnosis. J. Am. Med. Inform. Assoc. 29, 424–434 (2022).
Google Scholar
Jin, W., Xu, Y. & Wang, Z. Modeling Alzheimer’s disease biomarkers’ trajectory in the absence of a gold standard using a Bayesian approach. Stat. Med. 44, e70283 (2025).
Google Scholar
Lachmann, M. et al. Subphenotyping of patients with aortic stenosis by unsupervised agglomerative clustering of echocardiographic and hemodynamic data. JACC Cardiovasc. Interv. 14, 2127–2140 (2021).
Google Scholar
Chaudhary, K., Poirion, O. B., Lu, L. & Garmire, L. X. Deep learning-based multi-omics integration robustly predicts survival in liver cancer. Clin. Cancer Res. 24, 1248–1259 (2018).
Google Scholar
Schlosser, P. et al. Transcriptome- and proteome-wide association studies nominate determinants of kidney function and damage. Genome Biol. 24, 150 (2023).
Google Scholar
Sun, B. B. et al. Plasma proteomic associations with genetics and health in the UK Biobank. Nature 622, 329–338 (2023).
Google Scholar
Himmelstein, D. S. et al. Systematic integration of biomedical knowledge prioritizes drugs for repurposing. eLife 6, e26726 (2017).
Google Scholar
Minikel, E. V., Painter, J. L., Dong, C. C. & Nelson, M. R. Refining the impact of genetic evidence on clinical success. Nature 629, 624–629 (2024).
Google Scholar
Gao, C. et al. Proteome-wide association study for finding druggable targets in progression and onset of Parkinson’s disease. CNS Neurosci. Ther. 31, e70294 (2025).
Google Scholar
Zitnik, M., Agrawal, M. & Leskovec, J. Modeling polypharmacy side effects with graph convolutional networks. Bioinformatics 34, i457–i466 (2018).
Google Scholar
Miotto, R., Li, L., Kidd, B. A. & Dudley, J. T. Deep Patient: an unsupervised representation to predict the future of patients from the electronic health records. Sci. Rep. 6, 26094 (2016).
Google Scholar
Estiri, H. et al. Transitive sequencing medical records for mining predictive and interpretable temporal representations. Patterns 1, 100051 (2020).
Google Scholar
Li, D., Xing, W., Zhao, J., Shi, C. & Wang, F. Multimodal deep learning for predicting in-hospital mortality in heart failure patients using longitudinal chest X-rays and electronic health records. Int. J. Cardiovasc. Imaging 41, 427–440 (2025).
Google Scholar
Barr, P. B. et al. Polygenic risk factors for comorbid diagnoses in individuals with substance use disorders: a phenome-wide survival analysis. Psychol. Med. 56, e174 (2026).
Google Scholar
Forero, D. A. et al. Current needs for human and medical genomics research infrastructure in low and middle income countries. J. Med. Genet. 53, 438–440 (2016).
Google Scholar
Woldemariam, M. T. & Jimma, W. Adoption of electronic health record systems to enhance the quality of healthcare in low-income countries: a systematic review. BMJ Health Care Inform. 30, e100704 (2023).
Google Scholar
The H3Africa Consortium et al. Enabling the genomic revolution in Africa. Science 344, 1346–1348 (2014).
Google Scholar
Were, M. C. et al. mUzima mobile electronic health record (EHR) system: development and implementation at scale. J. Med. Internet Res. 23, e26381 (2021).
Google Scholar
Robbiati, C. et al. Improving TB surveillance and patients’ quality of care through improved data collection in Angola: development of an electronic medical record system in two health facilities of Luanda. Front. Public Health 10, 745928 (2022).
Google Scholar



