Data availability
Raw mass spectrometry proteomics data have been deposited in the ProteomeXchange Consortium available via iProX (IPX0007409000). The PTDS protein matrix and associated resources are available for academic and non-commercial use through the open-access platform db.prottalks.com. Other data used are available from the Kyoto Encyclopedia of Genes and Genomes pathway at https://www.genome.jp/kegg/pathway.html (ref. 57), from Metascape at https://metascape.org/gp/ (ref. 58), from STRING at https://cn.string-db.org/ (ref. 59) and from Ingenuity Pathway Analysis at http://www.ingenuity.com (ref. 60).
Code availability
The project data analysis codes are available via GitHub at https://github.com/guomics-lab/PTV-1.
References
Bunne, C. et al. How to build the virtual cell with artificial intelligence: priorities and opportunities. Cell 187, 7045–7063 (2024).
Article CAS PubMed PubMed Central Google Scholar
Qian, L., Dong, Z. & Guo, T. Grow AI virtual cells: three data pillars and closed-loop learning. Cell Res. 35, 319–321 (2025).
Article PubMed PubMed Central Google Scholar
Cui, H. et al. Towards multimodal foundation models in molecular cell biology. Nature 640, 623–633 (2025).
Article ADS CAS PubMed Google Scholar
Theodoris, C. V. et al. Transfer learning enables predictions in network biology. Nature 618, 616–624 (2023).
Article ADS CAS PubMed PubMed Central Google Scholar
Hao, M. et al. Large-scale foundation model on single-cell transcriptomics. Nat. Methods 21, 1481–1491 (2024).
Article CAS PubMed Google Scholar
Rosen, Y. R. et al. Universal cell embeddings: a foundation model for cell biology. Nature 656, 183–191 (2026).
Abhinav K. et al. Predicting cellular responses to perturbation across diverse contexts with state. Preprint at bioRxiv https://doi.org/10.1101/2025.06.26.661135 (2025).
Dong, M. et al. Stack: in-context learning of single-cell biology. Preprint at bioRxiv https://doi.org/10.64898/2026.01.09.698608 (2026).
Wang, C. et al. X-Cell: scaling causal perturbation prediction across diverse cellular contexts via diffusion language models. Preprint at bioRxiv https://doi.org/10.64898/2026.03.18.712807 (2026).
Bunne, C. et al. Learning single-cell perturbation responses using neural optimal transport. Nat. Methods 20, 1759–1768 (2023).
Article CAS PubMed PubMed Central Google Scholar
Yeo, G. H. T., Saksena, S. D. & Gifford, D. K. Generative modeling of single-cell time series with PRESCIENT enables prediction of cell trajectories with interventions. Nat. Commun. 12, 3222 (2021).
Article ADS CAS PubMed PubMed Central Google Scholar
Tong, A., Huang, J., Wolf, G., van Dijk, D. & Krishnaswamy, S. TrajectoryNet: a dynamic optimal transport network for modeling cellular dynamics. Proc. Mach. Learn. Res. 119, 9526–9536 (2020).
PubMed PubMed Central Google Scholar
Zhang, Z., Li, T. & Zhou, P. Learning stochastic dynamics from snapshots through regularized unbalanced optimal transport. In International Conference on Learning Representations (ICLR, 2025).
Zhang, Z. et al. Deciphering cell-fate trajectories using spatiotemporal single-cell transcriptomic data. npj Syst. Biol. Appl. 12, 2 (2026).
Article CAS Google Scholar
Boiarsky, R. et al. Deeper evaluation of a single-cell foundation model. Nat. Mach. Intell. 6, 1443–1446 (2024).
Article Google Scholar
Kedzierska, K. Z., Crawford, L., Amini, A. P. & Lu, A. X. Zero-shot evaluation reveals limitations of single-cell foundation models. Genome Biol. 26, 101 (2025).
Article PubMed PubMed Central Google Scholar
Ahlmann-Eltze, C., Huber, W. & Anders, S. Deep-learning-based gene perturbation effect prediction does not yet outperform simple linear baselines. Nat. Methods 22, 1657–1661 (2025).
Article CAS PubMed PubMed Central Google Scholar
Gillet, L. et al. Targeted data extraction of the MS/MS spectra generated by data-independent acquisition: a new concept for consistent and accurate proteome analysis. Mol. Cell. Proteomics 11, O111.016717 https://doi.org/10.1074/mcp.O111.016717 (2012).
Qian, L. et al. AI-empowered perturbation proteomics for complex biological systems. Cell Genom. 4, 100691 (2024).
Article CAS PubMed PubMed Central Google Scholar
Xiao, Q. et al. High-throughput proteomics and AI for cancer biomarker discovery. Adv. Drug Deliv. Rev. 176, 113844 (2021).
Article CAS PubMed Google Scholar
Guo, T., Steen, J. A. & Mann, M. Mass-spectrometry-based proteomics: from single cells to clinical applications. Nature 638, 901–911 (2025).
Article ADS CAS PubMed Google Scholar
Chen, R. T., Rubanova, Y., Bettencourt, J. & Duvenaud, D. K. Neural ordinary differential equations. In Proc. 32nd International Conference on Neural Information Processing Systems (eds Bengio, S. et al.) 6571–6583 (Curran Associates, 2018).
Weinan, E. A proposal on machine learning via dynamical systems. Comm. Math. Stat. 5, 1–11 (2017).
MathSciNet Google Scholar
Jaaks, P. et al. Effective drug combinations in breast, colon and pancreatic cancer cells. Nature 603, 166–173 (2022).
Article ADS CAS PubMed PubMed Central Google Scholar
Lamb, J. et al. The Connectivity Map: using gene-expression signatures to connect small molecules, genes, and disease. Science 313, 1929–1935 (2006).
Article ADS CAS PubMed Google Scholar
Subramanian, A. et al. A next generation connectivity map: L1000 platform and the first 1,000,000 profiles Cell 171, 1437–1452 (2017).
Article ADS CAS PubMed PubMed Central Google Scholar
Li, X. et al. LncRNA NEAT1 promotes autophagy via regulating miR-204/ATG3 and enhanced cell resistance to sorafenib in hepatocellular carcinoma. J. Cell. Physiol. 235, 3402–3413 (2020).
Article CAS PubMed Google Scholar
Hu, J. et al. BTF3 sustains cancer stem-like phenotype of prostate cancer via stabilization of BMI1. J. Exp. Clin. Cancer Res. 38, 227 (2019).
Article PubMed PubMed Central Google Scholar
Phi, L. T. H. et al. Cancer stem cells (CSCs) in drug resistance and their therapeutic implications in cancer treatment. Stem Cells Int. 2018, 5416923 (2018).
Article PubMed PubMed Central Google Scholar
De Greve, J. & Giron, P. Targeting the tyrosine kinase inhibitor-resistant mutant EGFR pathway in lung cancer without targeting EGFR? Transl. Lung Cancer Res. 9, 1–3 (2020).
Article PubMed PubMed Central Google Scholar
Liu, L. et al. The LIS1/NDE1 complex is essential for FGF signaling by regulating FGF receptor intracellular trafficking. Cell Rep. 22, 3277–3291 (2018).
Article CAS PubMed Google Scholar
Park, G. B., Jeong, J. Y., Choi, S., Yoon, Y. S. & Kim, D. Glucose deprivation enhances resistance to paclitaxel via ELAVL2/4-mediated modification of glycolysis in ovarian cancer cells. Anticancer Drugs 33, e370–e380 (2022).
Article CAS PubMed Google Scholar
Yang, T. et al. CKS2 promotes the malignant phenotypes of bladder cancer cells via PI3K/AKT signaling pathway activation. Cell Cycle 24, 687–701 (2025).
Article CAS PubMed PubMed Central Google Scholar
Yang, X. et al. GeneCompass: deciphering universal gene regulatory mechanisms with a knowledge-informed cross-species foundation model. Cell Res. 34, 830–845 (2024).
Article PubMed PubMed Central Google Scholar
Preuer, K. et al. DeepSynergy: predicting anti-cancer drug synergy with deep learning. Bioinformatics 34, 1538–1546 (2018).
Article CAS PubMed PubMed Central Google Scholar
Konstantinopoulos, P. A. et al. A phase II, two-stage study of letrozole and abemaciclib in estrogen receptor-positive recurrent endometrial cancer. J. Clin. Oncol. 41, 599–608 (2023).
Article CAS PubMed Google Scholar
Ingham, M. et al. Phase II study of olaparib and temozolomide for advanced uterine leiomyosarcoma (NCI Protocol 10250). J. Clin. Oncol. 41, 4154–4163 (2023).
Article CAS PubMed PubMed Central Google Scholar
Farago, A. F. et al. Combination olaparib and temozolomide in relapsed small-cell lung cancer. Cancer Discov. 9, 1372–1387 (2019).
Article CAS PubMed PubMed Central Google Scholar
Bruna, A. et al. A biobank of breast cancer explants with preserved intra-tumor heterogeneity to screen anticancer compounds. Cell 167, 260–274 (2016).
Article CAS PubMed PubMed Central Google Scholar
Gao, H. et al. High-throughput screening using patient-derived tumor xenografts to predict clinical trial drug response. Nat. Med. 21, 1318–1325 (2015).
Article CAS PubMed Google Scholar
Wang, J. et al. CDK7 inhibitor THZ1 enhances antiPD-1 therapy efficacy via the p38alpha/MYC/PD-L1 signaling in non-small cell lung cancer. J. Hematol. Oncol. 13, 99 (2020).
Article PubMed PubMed Central Google Scholar
Wang, Z. et al. HDAC6 promotes cell proliferation and confers resistance to gefitinib in lung adenocarcinoma. Oncol. Rep. 36, 589–597 (2016).
Article CAS PubMed Google Scholar
Zecha, J. et al. Decrypting drug actions and protein modifications by dose- and time-resolved proteomics. Science 380, 93–101 (2023).
Article ADS CAS PubMed PubMed Central Google Scholar
Eckert, S. et al. Decrypting the molecular basis of cellular drug phenotypes by dose-resolved expression proteomics. Nat. Biotechnol. 43, 406–415 (2025).
Ruprecht, B. et al. A mass spectrometry-based proteome map of drug action in lung cancer cell lines. Nat. Chem. Biol. 16, 1111–1119 (2020).
Article CAS PubMed Google Scholar
Dibaeinia, P. et al. Virtual cells need context, not just scale. Preprint at bioRxiv https://doi.org/10.64898/2026.02.04.703804 (2026).
Liu, Z. et al. DIA-BERT: pre-trained end-to-end transformer models for enhanced DIA proteomics data analysis. Nat. Commun. 16, 3530 (2025).
Article ADS CAS PubMed PubMed Central Google Scholar
Wallmann, G. et al. AlphaDIA enables DIA transfer learning for feature-free proteomics. Nat. Biotechnol. 44, 1168–1177 (2026).
Tang, X. et al. CellForge: agentic design of virtual cell models. Preprint at https://doi.org/10.48550/arXiv.2508.02276 (2025).
Mitchell, D. C. et al. A proteome-wide atlas of drug mechanism of action. Nat. Biotechnol. 41, 845–857 (2023).
Article CAS PubMed PubMed Central Google Scholar
Cai, X. et al. High-throughput proteomic sample preparation using pressure cycling technology. Nat. Protoc. 17, 2307–2325 (2022).
Article CAS PubMed PubMed Central Google Scholar
Sun, R. et al. Accelerated protein biomarker discovery from FFPE tissue samples using single-shot, short gradient microflow SWATH MS. J. Proteome Res. 19, 2732–2741 (2020).
Article CAS PubMed Google Scholar
Sun, R. et al. A prostate cancer tissue specific spectral library for targeted proteomic analysis. Proteomics 22, e2100147 (2022).
Article PubMed Google Scholar
Demichev, V., Messner, C. B., Vernardis, S. I., Lilley, K. S. & Ralser, M. DIA-NN: neural networks and interference correction enable deep proteome coverage in high throughput. Nat. Methods 17, 41–44 (2020).
Article CAS PubMed Google Scholar
Zhong, Q. et al. Proteomic-based stratification of intermediate-risk prostate cancer patients. Life Sci. Alliance 7, e202302146 (2024).
Sun, R. et al. Proteomic dynamics of breast cancer cell lines identifies potential therapeutic protein targets. Mol. Cell Proteomics 22, 100602 (2023).
Article CAS PubMed PubMed Central Google Scholar
Kanehisa, M. & Goto, S. KEGG: Kyoto Encyclopedia of Genes and Genomes. Nucleic Acids Res. 28, 27–30 (2000).
Article CAS PubMed PubMed Central Google Scholar
Zhou, Y. et al. Metascape provides a biologist-oriented resource for the analysis of systems-level datasets. Nat. Commun. 10, 1523 (2019).
Article ADS PubMed PubMed Central Google Scholar
Szklarczyk, D. et al. The STRING database in 2023: protein–protein association networks and functional enrichment analyses for any sequenced genome of interest. Nucleic Acids Res. 51, D638–D646 (2023).
Article CAS PubMed PubMed Central Google Scholar
Kramer, A., Green, J., Pollard, J. Jr. & Tugendreich, S. Causal analysis approaches in ingenuity pathway analysis. Bioinformatics 30, 523–530 (2014).
Article PubMed Google Scholar
Lundberg, S. M. & Lee, S.-I. A unified approach to interpreting model predictions. In Proc. 31st International Conference on Neural Information Processing Systems (eds von Luxburg, U. et al.) 4768–4777 (Curran Associates Inc., 2017).
Download references
Acknowledgements
We thank X. Zhang, M. Liang, X. Liu, Y. Lai, W. Lin, Y. Hu, M. Chen, C. Li, J. Wang, Y. Wu and D. Yang for assistance with the sample preparation and mass spectrometry maintenance. We thank M. Yan and L. Wan for guiding immunohistochemical staining and PDO culture. We thank Westlake University Supercomputer Center for assistance in data generation and storage, and the Mass Spectrometry & Metabolomics Core Facility at the Center for Biomedical Research Core Facilities of Westlake University for sample analysis. This study was partly supported by the π-HuB project.
Funding
This work is supported by grants from Joint Funds of the National Natural Science Foundation of China (grant no. U24A20476), National Natural Science Foundation of China (Young Scientist Fund, grant no. 32401239, Basic Science Center Program, grant no. 12288101), the Zhejiang Provincial Natural Science Foundation of China (grant nos. LQ24C050002 and LMS26C050002), National Key R&D Program of China (grant nos. 2022YFF0608403 and 2021YFA1301600), Shanghai Municipal Special Program for Basic Research on General AI Foundation Models (grant no. 2025SHZDZX026D07), the State Key Laboratory of Medical Proteomics (grant nos. SKLP-K202501, SKLP-Y202403 and SKLP-K202406), National Natural Science Foundation of China (grant no. 81972492) and the Science and Technology Commission of Shanghai Municipality (STCSM) (grant no. 25JS2850100).
Ethics declarations
Competing interests
T.G. and Y. Zhu are shareholders of Westlake Omics Inc. L.T., Y. Zhan and W.H. are employees of Westlake Omics Inc. Y.L. is an employee of DP Technology Co., Ltd. H.W. is the founder of AiR (AI-in-RNA)-Bio Technology. The other authors declare no competing interests.
Peer review
Peer review information
Nature thanks the anonymous reviewers for their contribution to the peer review of this work. Peer reviewer reports are available.
Additional information
Publisher’s note Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
Extended data figures and tables
Extended Data Fig. 1 Quality control analysis of perturbation proteomics dataset (PTDS).
(A) Classification of 63 FDA-approved drugs by mechanism of action (MOA). Sankey plot shows the relationship among drugs, target proteins and associated pathways. (B)-(C) The reproducibility of pool (n = 618), technical replicate (n = 1238), and biological replicate samples (n = 4492, 2-4 biological replicate samples for each condition (cell-line-drug-time-point pair)) evaluated by Pearson correlation (B) and coefficient of variation (C). (D)-(E) UMAP visualization of proteome-wide patterns. UMAP shows the unsupervised clustering of the whole proteomes by different cell lines (D), perturbation durations (E).
Extended Data Fig. 2 Quality control analysis of multi-time-point perturbation proteomics dataset (mtPTDS).
(A)-(B) Reproducibility metrics. (A) Pearson correlations between pooled samples (n = 41), technical replicates (n = 84), and biological replicates (n = 49). (B) Coefficient of variation distributions for each sample type. (C–F) UMAP embeddings of whole-proteome profiles colored by: (C) mass spectrometry instrument batch; (D) cell line identity; (E) treatment duration; and (F) drug class. These plots reveal minimal batch effects and clear biological separation by cell type, time, and drug MOA.
Extended Data Fig. 3 Dysregulated perturbation score.
(A) Drug targets identified in this perturbation proteomics dataset. (B) Time-resolved expression of TYMS protein following treatment with its inhibitor (capecitabine). Line plot shows mean log2 (protein abundance) ± s.e. across time points. (C) siRNA knockdown validation. Box plots depicting the log2-scaled protein abundance detected by MS to evaluate RNA interference efficacy (n = 3 biological replicates). The knockdown efficiency of TYMS was evaluated in HCC1143 cells. Statistical significance between groups was assessed using the two-sided Student’s t-test. Box plots display median (center line), interquartile range (box), and 1.5× IQR (whiskers). (D) Functional validation of TYMS knockdown. Dose-response curves show enhanced capecitabine sensitivity in TYMS-depleted HCC1143 cells compared to control. Each curve shows the fitted sigmoid function, with the central point corresponding to the IC50 concentration. For each concentration, n = 3 biological replicates from independent cell cultures were used for model fitting; replicate-level measurements are shown in Supplementary Information. (E) Number of overlapping PRPs identified across PertScore thresholds (7, 8, 9, 10, 11, 12) following FDR correction (FDR < 0.05). (F) Perturbation score for different drug classes.
Extended Data Fig. 4 Pre-experiments of ProteinTalks using mtPTDS.
(A) UMAP shows whole proteomics embedding by ProteinTalks grouping by different perturbation durations. (B)-(D) Impact of temporal sampling on model performance. Box plots show the performance of ProteinTalks evaluated by AUPRC (B) and AUROC (C) decreasing with the number of time points, accuracy (D) increasing with the number of time points. Box plots display median (center line), interquartile range (box), and 1.5× IQR (whiskers). (E)-(G) Learning curves showing the performance of ProteinTalks evaluated by AUPRC (E), AUROC (F), and accuracy (G) increasing with the training data size. Points show mean performance; error bars represent s.e. across independent training runs (n = 17 for 10% and 30%, n = 16 for 50%, and n = 15 for 70% training data). Central lines represent the means, and shaded areas indicate mean ± s.d..
Extended Data Fig. 5 Transfer learning across five non-breast cancer cell lines using ProteinTalks.
(A)-(B) The reproducibility of technical replicate samples (n = 86) evaluated by Pearson correlation (A) and coefficient of variation (B). (C)-(D) The reproducibility of pool samples (n = 90) evaluated by Pearson correlation (C) and coefficient of variation (D).
Extended Data Fig. 6 Dynamic SHAP value.
Heatmap represents the protein with the greatest impact on drug efficacy, and among the proteins it influences, it highlights the protein exhibiting the largest variation across six drug classes: alkylating agents (A), HDAC inhibitors (B), topoisomerase inhibitors (C), CDK inhibitors (D), hormonal agents (E), and kinase inhibitors (F). The y-axis denotes proteins affecting drug efficacy prediction, while the x-axis represents the proteins influenced by these proteins. Color intensity indicates SHAP value magnitude. Histograms show column-wise SHAP value sums. Drug-class-specific mechanistic signatures are highlighted.
Extended Data Fig. 7 Interpretation of ProteinTalks model using SHAP values for different drugs.
(A) Beeswarm plots illustrate the protein-level contributions to the model’s predictions for antimitotic agents. Each point represents the SHAP value of a given protein in an individual sample. The y-axis represents the SHAP value, while the x-axis displays the corresponding proteins. Positive SHAP values signify a positive contribution to the predicted efficacy, whereas negative values indicate a negative contribution. Larger absolute SHAP values indicate stronger contribution to the predicted efficacy. (B) Box plot depicting the log2-scaled protein abundance detected by MS to evaluate RNA interference efficacy. The knockdown efficiency of AKR1C3 was evaluated in BT-20 cells (n = 3 biological replicates). Statistical significance between groups was assessed using the two-sided Student’s t-test. Box plots display median (center line), interquartile range (box), and 1.5× IQR (whiskers).
Extended Data Fig. 8 Assessing the clinical relevance of the fine-tuned ProteinTalks (model-1) with PDX transcriptomics data.
(A) Diagram illustrating the construction of the ProteinTalks model. The model underwent training with perturbation proteomic data and was subsequently refined using a subset of transcriptomic data from PDX models. The remaining transcriptomic data were used to evaluate the model’s capability to predict drug efficacy. Model-1-wo, Model-1-without, was trained from scratch solely on the same subset of transcriptomic PDTC data without transfer learning from the perturbation proteomic data. When 90% of the PTDS perturbation proteomic data were used in the first stage, the corresponding models obtained were Model-1-wo and Model-1-90%, respectively. (B)-(D) Comparative display for pan-cancer PDXs generated by the ProteinTalks model, including AUROC (B), AUPRC (C), and Accuracy (D). Statistical significance assessed by two-sided Student t-test (n.s., not significant).
Extended Data Fig. 9 Clinical relevance validation using the TNBC-HMU cohort.
(A)-(B) Reproducibility of technical replicates (n = 41), biological replicates (n = 46), pool (n = 35), and mouse liver samples (n = 43) evaluated by Pearson correlation (A) and coefficient of variation (B). (C)-(D) Kaplan-Meier (KM) curves of RFS (C) and OS (D), stratified by ProteinTalks-predicted risk scores for patients treated with three or four drug combinations. Central lines represent Kaplan-Meier estimates of survival probability, with shaded bands indicating 95% confidence intervals calculated using Greenwood’s formula. P values were calculated using the log-rank test.
Supplementary information
Rights and permissions
Springer Nature or its licensor (e.g. a society or other partner) holds exclusive rights to this article under a publishing agreement with the author(s) or other rightsholder(s); author self-archiving of the accepted manuscript version of this article is solely governed by the terms of such publishing agreement and applicable law.
Reprints and permissions
About this article
Cite this article
Sun, R., Qian, L., Li, Y. et al. An operational perturbation proteomics-based virtual cell model. Nature (2026). https://doi.org/10.1038/s41586-026-11001-9
Download citation
Received:
Accepted:
Published:
Version of record:
DOI: https://doi.org/10.1038/s41586-026-11001-9