문서

연구 주제

가상세포

오믹스 데이터를 학습한 모델로 세포의 반응을 실험 전에 가늠하는 연구 — Virtual Cell.

무엇을 하려는 연구인가

가상세포(virtual cell)는 오믹스 데이터를 학습한 모델로 세포를 계산 가능한 형태로 옮겨 놓고, 실험을 하기 전에 결과를 가늠해 보려는 시도입니다. 어떤 유전자를 눌렀을 때 발현이 어떻게 움직이는지, 어떤 조건이 세포 상태를 어느 방향으로 미는지를 모델 위에서 먼저 돌려 보고, 실제 실험은 확인할 값어치가 가장 큰 조건에 씁니다.

실험을 대체하자는 것이 아닙니다 — 실험의 순서를 바꾸자는 것입니다. 후보를 넓게 훑는 일은 모델이 맡고, 손과 시약은 모델이 추린 소수의 조건에 집중합니다.

왜 지금인가

단일세포·공간 오믹스가 세포 상태를 전례 없는 해상도로 기록하기 시작했고, 대규모 섭동(perturbation) 데이터셋과 그것을 학습한 기반 모델(foundation model)들이 등장하고 있습니다. "세포의 디지털트윈"이 구호가 아니라 검증 가능한 연구 문제가 된 시점입니다. 관건은 모델이 있느냐가 아니라 — 어느 예측을 믿어도 되는지를 데이터로 가려내는 일입니다.

contextBio 는 무엇을 하나

모델의 품질은 학습 데이터의 품질에서 시작합니다. AURORA가 오믹스 원자료를 재현 가능하게 처리해 두는 것이 이 연구의 데이터 쪽 바탕이고, 조직 맥락을 붙이는 일은 암 공간생물학과 이어집니다. 세포를 넘어 환자 수준으로 올라서는 방향은 가상병원에 적었습니다.

지금 어디까지 왔나

두 갈래로 진행 중입니다. 데이터 쪽에서는 AURORA가 처리한 오믹스가 모델이 읽을 수 있는 정돈된 형태로 쌓이도록 처리와 보관의 배선을 잇고 있습니다. 추론 쪽에서는 큐레이션된 데이터베이스·자체 분석 코드·문헌 지식을 엮어 분석을 설계하고 해석하는 코사이언티스트(co-scientist) 에이전트를 연구실 안에서 만들어 검증하고 있습니다 — 아직 서비스로 열기 전의 연구 단계입니다.

참고문헌

89편

이 쪽의 근거는 리뷰 원고 Perspectives on Virtual Cell to Digital Twin Technology 이고, 그 원고가 인용한 문헌 목록입니다.

  1. Bunne C, Roohani Y, Rosen Y, et al. How to build the virtual cell with artificial intelligence: priorities and opportunities. Cell. 2024;187[25]:7045-7063. doi:10.1016/j.cell.2024.11.015
  2. Rood JE, Regev A, et al. From modality-specific to compositional foundation models for cell biology. Cell Syst. 2026. doi:10.1016/j.cels.2026.01.016
  3. Cui H, Wang C, Maan H, et al. scGPT: toward building a foundation model for single-cell multi-omics using generative AI. Nat Methods. 2024;21[8]:1470-1480. doi:10.1038/s41592-024-02201-0
  4. Hao M, Gong J, Zeng X, et al. Large-scale foundation model on single-cell transcriptomics. Nat Methods. 2024;21[8]:1481-1491. doi:10.1038/s41592-024-02305-7
  5. Heimberg G, Kuo T, DePianto DJ, et al. A cell atlas foundation model for scalable search of similar human cells. Nature. 2025;638(8050):1085-1094. doi:10.1038/s41586-024-08411-y
  6. Arc Institute Virtual Cell Program. Predicting cellular responses to perturbation across diverse contexts with State. bioRxiv. 2025. doi:10.1101/2025.06.26.661135
  7. Tahoe Therapeutics, Arc Institute, Chan Zuckerberg Biohub. Tahoe-100M: a giga-scale single-cell perturbation atlas for context-dependent gene function and cellular modeling. bioRxiv. 2025. doi:10.1101/2025.02.20.639398
  8. Pearce JD, Simon LM, Mibar N, et al. A cross-species generative cell atlas across 1.5 billion years of evolution: the TranscriptFormer single-cell model. bioRxiv. 2025. doi:10.1101/2025.04.25.650731
  9. Roohani YH, Hua TJ, Tung PY, et al. Virtual Cell Challenge: toward a Turing test for the virtual cell. Cell. 2025;188[13]:3370-3374. doi:10.1016/j.cell.2025.06.008
  10. Ma J, Hao M, Song L. Grow AI virtual cells: three data pillars and closed-loop learning. Cell Res. 2025. doi:10.1038/s41422-025-01101-y
  11. Wu Z, et al. Virtual cells: from conceptual frameworks to biomedical applications. arXiv. 2025. doi:10.48550/arXiv.2509.18220
  12. Bunne C, Doncevic D, Wu D, et al. AI-powered virtual tissues from spatial proteomics for clinical diagnostics and biomedical discovery (VirTues). arXiv. 2025. doi:10.48550/arXiv.2501.06039
  13. Xu Y, et al. Whole-cell particle-based digital twin simulations from 4D lattice light-sheet microscopy data. Cell. 2026. doi:10.1016/j.cell.2026.06.011
  14. National Academies of Sciences, Engineering, and Medicine. Foundational Research Gaps and Future Directions for Digital Twins. Washington, DC: National Academies Press; 2024. doi:10.17226/26894
  15. Katsoulakis E, Wang Q, Wu H, et al. Digital twins for health: a scoping review and the promise of medical digital twins for precision medicine and medical artificial intelligence. Lancet Digit Health. 2025;7[3]:e229-e240. doi:10.1016/S2589-7500[25]00028-7
  16. Vallée A, et al. A scoping review of human digital twins in healthcare applications and usage patterns. NPJ Digit Med. 2025;8:610. doi:10.1038/s41746-025-01910-w
  17. Viceconti M, Emili L, Afshari P, et al. The future of in silico trials and digital twins in medicine. PNAS Nexus. 2025;4[5]:pgaf123. doi:10.1093/pnasnexus/pgaf123
  18. Laubenbacher R, et al. Multi-scale digital twins for personalized medicine. Front Digit Health. 2026. doi:10.3389/fdgth.2026.1753906
  19. Trayanova NA, Prakosa A. Building digital twins for cardiovascular health: from principles to clinical impact. J Am Heart Assoc. 2024;13[15]:e031981. doi:10.1161/JAHA.123.031981
  20. Polak S, et al. Digital twins, synthetic patient data, and in-silico trials: can they empower paediatric clinical trials? Lancet Digit Health. 2025;7[2]:e155-e162. doi:10.1016/S2589-7500[25]00007-X
  21. EDITH Consortium. A roadmap for the European Virtual Human Twin. Zenodo. 2025. doi:10.5281/zenodo.14769224
  22. U.S. Food and Drug Administration. Assessing the credibility of computational modeling and simulation in medical device submissions; ASME V&V 40. Silver Spring, MD: FDA; 2023.
  23. International Council for Harmonisation. M15 general principles for model-informed drug development (draft guidance). 2024.
  24. Ma C, Zhang H, Rao Y, et al. AI-driven virtual cell models in preclinical research: technical pathways, validation mechanisms, and clinical translation potential. NPJ Digit Med. 2025;8:614. doi:10.1038/s41746-025-02198-6
  25. Yang T, Wang YY, Ma F, Xu BH, Qian HL. Build the virtual cell with artificial intelligence: a perspective for cancer research. Mil Med Res. 2025;12:8. doi:10.1186/s40779-025-00591-6
  26. Callaway E. Can AI build a virtual cell? Scientists race to model life's smallest unit. Nature. 2025;643:19-21. doi:10.1038/d41586-025-02011-0
  27. An G, Cockrell C, et al. Forum on immune digital twins: a meeting report. NPJ Syst Biol Appl. 2024. doi:10.48550/arXiv.2310.18374
  28. Wenteler A, Cabrera CP, Wei W, et al. Benchmarking transcriptomics foundation models for perturbation analysis: one PCA still rules them all. arXiv. 2024. doi:10.48550/arXiv.2410.13956
  29. Park S, et al. Artificial intelligence virtual organoids (AIVOs): organoid-scale digital twins with virtual cells as minimal executable units. Mater Today Bio. 2025. doi:10.1016/j.mtbio.2025.102155
  30. Zhang L, et al. Artificial intelligence-enabled organoid platforms for precision medicine: integrating multi-omics, digital twins, and microphysiological systems. Organoids. 2026;5[3]:20. doi:10.3390/organoids5030020
  31. Chen Y, et al. PharmaFormer predicts clinical drug responses through transfer learning guided by patient-derived organoids. NPJ Precis Oncol. 2025;9:82. doi:10.1038/s41698-025-01082-6
  32. Wang H, et al. Building consensus on the application of organoid-based drug sensitivity testing in cancer precision medicine and drug development. Theranostics. 2024;14[8]:3300-3316. doi:10.7150/thno.96027
  33. Boiko DA, MacKnight R, Kline B, Gomes G. Autonomous chemical research with large language models. Nature. 2023;624(7992):570-578. doi:10.1038/s41586-023-06792-0
  34. United States Congress. Food and Drug Administration Modernization Act 2.0, Pub. L. No. 117-328. 2023. (recognition of organoids, organ-on-chip, microphysiological systems, and AI/ML models as alternatives to animal testing)
  35. Ahlmann-Eltze C, Huber W, Anders S. Deep-learning-based gene perturbation effect prediction does not yet outperform simple linear baselines. Nat Methods. 2025;22[8]:1657-1661. doi:10.1038/s41592-025-02772-6
  36. Csendes G, Sanz G, Szalay KZ, Szalai B. Benchmarking foundation cell models for post-perturbation RNA-seq prediction. BMC Genomics. 2025;26[1]:393. doi:10.1186/s12864-025-11600-2
  37. Kedzierska KZ, Crawford L, Amini AP, Lu AX. Zero-shot evaluation reveals limitations of single-cell foundation models. Genome Biol. 2025;26:101. doi:10.1186/s13059-025-03574-x
  38. Arc Institute. Virtual Cell Challenge 2025 wrap-up: winners and reflections. 2025. https://arcinstitute.org/news/virtual-cell-challenge-2025-wrap-up Accessed 29 July 2026.
  39. Li K, Xiao X, Deng S, et al. Large language models meet virtual cell: a survey. arXiv. 2025. doi:10.48550/arXiv.2510.07706
  40. Liu K, Li M, Bao Y, et al. Medical world models: representing medical states, modelling clinical dynamics and guiding intervention policies. arXiv. 2026. doi:10.48550/arXiv.2606.16721
  41. Chuai G, Chen X, Yang X, et al. Towards building a world model to simulate perturbation-induced cellular dynamics by AlphaCell. bioRxiv. 2026. doi:10.64898/2026.03.02.709176
  42. Wei Z, Ma R, Wang Z, Li Z, Song S, Zheng S. VCWorld: a biological world model for virtual cell simulation. arXiv. 2026 (ICLR 2026). doi:10.48550/arXiv.2512.00306
  43. Minimal life by computer. Nat Biotechnol. 2026 (editorial). doi:10.1038/s41587-026-03110-7
  44. Sun Y, A J, Liu Z, et al. Strategic priorities for transformative progress in advancing biology with proteomics and artificial intelligence. arXiv. 2025. doi:10.48550/arXiv.2502.15867
  45. Dibaeinia P, Babu S, Knudson M, et al. Virtual cells need context, not just scale. bioRxiv. 2026. doi:10.64898/2026.02.04.703804
  46. Eadie AL, Fernandez Lynch H, Scheinerman N, Parikh RB. Advancing FDA New Approach Methodologies from animal models through digital twins. NPJ Digit Med. 2026;9:449. doi:10.1038/s41746-026-02476-x
  47. U.S. Food and Drug Administration; European Medicines Agency. Guiding principles of good AI practice in drug development. January 2026. https://www.fda.gov/about-fda/artificial-intelligence-drug-development/guiding-principles-good-ai-practice-drug-development Accessed 29 July 2026.
  48. Mao X, Zhang S, Wen Q, et al. Benchmarking virtual cell models for in-the-wild perturbation response. arXiv. 2026. doi:10.48550/arXiv.2604.27646
  49. Bereket MD, Leskovec J. Are current AI virtual cell models useful for scientific discovery? bioRxiv. 2026. doi:10.64898/2026.04.23.719015
  50. Cheng X, Li P, Guo H, et al. Harnessing AI to build virtual cells. bioRxiv. 2026. doi:10.64898/2026.04.11.717183
  51. Xie Z, Li W, Chen Y, Peng Z, Xiang L, Wang D. AetherCell: a generative engine for virtual cell perturbation and in vivo drug discovery. bioRxiv. 2026. doi:10.64898/2026.03.13.710968
  52. Kang B, Fan R, Yi M, Cui C, Cui Q. BulkFormer: a large-scale foundation model for bulk transcriptomes. Cell Syst. 2026;17[7]:101657. doi:10.1016/j.cels.2026.101657
  53. Agarwal D, Bisht R. The metric picks the winner: evaluation choice flips model rankings for drug-response prediction in unseen chemistry. arXiv. 2026. doi:10.48550/arXiv.2606.12639
  54. Wu AR. Revisiting the blueprint for an interpretable virtual cell. Nat Rev Genet. 2026;27[5]:347. doi:10.1038/s41576-026-00940-8
  55. Eadie AL, Fernandez Lynch H, Scheinerman N, Parikh RB. The arrival of digital twins and in silico trials in drug development. Nat Med. 2026;32[6]:1967-1971. doi:10.1038/s41591-026-04344-3
  56. Jiang H, Huang X, Bi X, et al. Artificial intelligence-enabled multi-scale virtual cell: perspective, challenges, and opportunities. Brief Bioinform. 2026;27[2]:bbag104. doi:10.1093/bib/bbag104
  57. Xiao Z, Wang W, Long X, et al. scDifformer: diffusion-based post-training for virtual cell modeling across large-scale single-cell data. Nucleic Acids Res. 2026;54[13]:gkag706. doi:10.1093/nar/gkag706
  58. Jia Y, Xing X, Han H, Zhao Y, Zhou S, Chen J. Virtual cell construction for artificial intelligence-driven drug discovery. Br J Pharmacol. 2026. doi:10.1111/bph.70585
  59. Li J, Zhang H, Zhang W, Li N, Zhao F, Shen G. Virtual cell: current perspectives and future prospects. J Int Med Res. 2026;54[5]:3000605261425080. doi:10.1177/03000605261425080
  60. Du Q, Wang M, Zhuang Y, Gu Q, Qi W. Virtual cell/virtual cell like: the key to unlocking a new era of HIV/AIDS treatment? Microb Pathog. 2026;215:108459. doi:10.1016/j.micpath.2026.108459
  61. Martinov MV, Ataullakhanov FI, Protasov ES, Vitvitsky VM. Virtual cell and metabolic control analysis: control coefficients for glycolytic flux are highly dependent on the subsystem selected for analysis. Life (Basel). 2026;16[3]:414. doi:10.3390/life16030414
  62. Wu T, Tang HF, Wang WL. DRIVE: a comprehensive resource deciphering drug-induced transcriptomic and splicing response in cancer cell. Neoplasia. 2026;79:101336. doi:10.1016/j.neo.2026.101336
  63. Wang S, Li Y, Wang X, et al. STGBench: sequencing-level spatial DNA-RNA simulation for multimodal and virtual cell-oriented benchmarking of genomic alterations. Brief Bioinform. 2026;27[4]:bbag354. doi:10.1093/bib/bbag354
  64. Zheng Y, Jin R, Zhang Y, Cheng Y, Wang Y. scAClc: a multi-objective adaptive clustering framework for single-cell transcriptomics via contrastive and resolution-aware representation learning. Anal Chem. 2026;98[15]:11010-11019. doi:10.1021/acs.analchem.5c06601
  65. Yin M, Luo R, Hu S, Lin H, Sun H, Sun Y. Applications of digital twin technology in cardiometabolic disease management: a scoping review. BMC Med Inform Decis Mak. 2026;26[1]:236. doi:10.1186/s12911-026-03673-0
  66. Zarei A, Gholamzadeh M, Asadi F. Application of digital twin technology to enhance chronic diseases management: a systematic review. Int J Telemed Appl. 2026;2026:2299762. doi:10.1155/ijta/2299762
  67. Theivendren P, Panneerselvam S, Kunjiappan S, Mani D, Parthasarathy AK. Medical digital twin insights: enhancing cancer treatment through generative modeling. Comput Biol Chem. 2026;124(Pt 1):109074. doi:10.1016/j.compbiolchem.2026.109074
  68. Niarakis A, Laubenbacher R, An G, et al. Immune digital twins for complex human pathologies: applications, limitations, and challenges. NPJ Syst Biol Appl. 2024;10[1]:141. doi:10.1038/s41540-024-00450-5
  69. Moore R, Mayo L, Startsev D, et al. A comprehensive mechanistic multicellular model of the human immune system spanning 11 diseases. Front Immunol. 2026;17:1732556. doi:10.3389/fimmu.2026.1732556
  70. Abuhassan Q, Al-Ameer HJ, Balogh Z, et al. Digital immune twins and AI-integrated multi-omic biomarkers: redefining personalized immunotherapy in non-small cell lung cancer. Iran J Basic Med Sci. 2026;29[5]:688-716. doi:10.22038/ijbms.2026.92560.19984
  71. Roohani Y, Huang K, Leskovec J. Predicting transcriptional outcomes of novel multigene perturbations with GEARS. Nat Biotechnol. 2024;42[6]:927-935. doi:10.1038/s41587-023-01905-6
  72. Theodoris CV, Xiao L, Chopra A, et al. Transfer learning enables predictions in network biology (Geneformer). Nature. 2023;618(7965):616-624. doi:10.1038/s41586-023-06139-9
  73. Chandak P, Huang K, Zitnik M. Building a knowledge graph to enable precision medicine (PrimeKG). Sci Data. 2023;10[1]:67. doi:10.1038/s41597-023-01960-3
  74. Rosen Y, Roohani Y, Agrawal A, Samotorčan L, Quake SR, Leskovec J. Universal cell embedding provides a foundation model for cell biology (UCE). Nature. 2026. doi:10.1038/s41586-026-10689-z
  75. Rizvi SA, Levine D, Patel A, et al. Scaling large language models for next-generation single-cell analysis (Cell2Sentence-Scale). bioRxiv. 2025. doi:10.1101/2025.04.14.648850
  76. Wei Z, Wang Y, Gao Y, et al. Benchmarking algorithms for generalizable single-cell perturbation response prediction. Nat Methods. 2026;23[2]:451-464. doi:10.1038/s41592-025-02980-0
  77. Gayoso A, Steier Z, Lopez R, Regier J, Nazor KL, Streets A, Yosef N. Joint probabilistic modeling of single-cell multi-omic data with totalVI. Nat Methods. 2021;18[3]:272-282. doi:10.1038/s41592-020-01050-x
  78. Ashuach T, Gabitto MI, Koodli RV, Saldi GA, Jordan MI, Yosef N. MultiVI: deep generative model for the integration of multimodal data. Nat Methods. 2023;20[8]:1222-1231. doi:10.1038/s41592-023-01909-9
  79. Ji B, Hu T, Wang J, et al. CAPTAIN: a multimodal foundation model pretrained on co-assayed single-cell RNA and protein. Nat Commun. 2026;17[1]. doi:10.1038/s41467-026-72882-y
  80. Li Y, et al. scKEPLM: knowledge-enhanced large-scale pre-trained language model for single-cell transcriptomics. bioRxiv. 2024. doi:10.1101/2024.07.09.602633
  81. Wen H, Tang W, Dai X, Ding J, Jin W, Xie Y, Tang J. CellPLM: pre-training of cell language model beyond single cells. bioRxiv. 2023. doi:10.1101/2023.10.03.560734
  82. Tejada-Lapuerta A, Schaar AC, Gutgesell R, et al. Nicheformer: a foundation model for single-cell and spatial omics. Nat Methods. 2025;22[12]:2525-2538. doi:10.1038/s41592-025-02814-z
  83. Liu G, Zhao Y, Zhao Y, et al. SCARF: single cell ATAC-seq and RNA-seq foundation model. bioRxiv. 2025. doi:10.1101/2025.04.07.647689
  84. Cao Y, Zhao X, Tang S, et al. scButterfly: a versatile single-cell cross-modality translation method via dual-aligned variational autoencoders. Nat Commun. 2024;15[1]:2973. doi:10.1038/s41467-024-47418-x
  85. Yuan H, Kelley DR. scBasset: sequence-based modeling of single-cell ATAC-seq using convolutional neural networks. Nat Methods. 2022;19[9]:1088-1096. doi:10.1038/s41592-022-01562-8
  86. Wu J, Wan C, Ji Z, Zhou Y, Hou W. EpiFoundation: a foundation model for single-cell ATAC-seq via peak-to-gene alignment. bioRxiv. 2025. doi:10.1101/2025.02.05.636688
  87. Wang CX, Cui H, Zhang AH, Xie R, Goodarzi H, Wang B. scGPT-spatial: continual pretraining of single-cell foundation model for spatial transcriptomics. bioRxiv. 2025. doi:10.1101/2025.02.05.636714
  88. Blampey Q, Benkirane H, Bercovici N, et al. Novae: a graph-based foundation model for spatial transcriptomics data. Nat Methods. 2025;22[12]:2539-2550. doi:10.1038/s41592-025-02899-6
  89. López-De-Castro M, García-Galindo A, González-Gomariz J, Armañanzas R. Conformal inference for reliable single cell RNA-seq annotation. Bioinformatics. 2025;41[10]:btaf521. https://doi.org/10.1093/bioinformatics/btaf521 Figure 1. The cell-to-organism digital twin continuum. Schematic overview of the central thesis: the AI virtual cell is the foundational building block of the biomedical digital twin, and the two integrate into a multiscale continuum spanning the cell, tissue, organ and whole-body scales. Three enabling technologies (top) make the continuum tractable. The central ladder shows the scales at which published systems now operate; at the cellular rung these are grouped by model class rather than by individual system, comprising foundation models, knowledge-grounded models that inject ontologies and graphs, multimodal integration across transcriptome, chromatin and protein, perturbation world models, and mechanistic whole-cell simulation. Upward arrows denote bottom-up compositional aggregation of learned cellular representations into higher-scale twins; downward arrows denote top-down propagation of physiological and clinical constraints. Applications (right) are aligned with the scale at which they principally operate. At the base, the lab-in-the-loop closes the defining feedback loop of a twin and is drawn as three nested loops: a computational inner loop that triages candidates before any assay is run and reports calibrated uncertainty; an experimental middle loop in which patient-derived organoids act as a physical twin whose measurements are assimilated; and a translational outer loop in which monitored clinical outcomes refine cellular and organ-scale priors. Results from every loop return to the model through active learning. A cross-cutting foundation of credibility and governance underpins clinical translation at and across every scale. Figure 2. The lab-in-the-loop as three nested validation loops. The quality of the lab-in-the-loop is the rate-limiting step of the cell-to-organism continuum; this figure makes its internal structure explicit. A virtual cell or digital twin proposes candidate perturbations, therapies or trajectories. These enter a computational inner loop [1] in which they are triaged in silico before any assay is run, using coverage of the phenotype space to detect mode collapse, rank correlation between predicted and measured responses, adversarial validation in which a classifier trained to separate generated from measured profiles should approach chance performance, and calibrated uncertainty, for which conformal prediction returns label sets whose size flags contexts the model has not seen. Surviving candidates enter an experimental middle loop [2] of CRISPR perturbation, patient-derived organoid and organ-on-chip assays executed under laboratory automation, whose measurements are assimilated by the model. Validated findings enter a translational outer loop [3] of in silico trials, precision oncology and monitored clinical outcomes assessed against risk-based regulatory expectations. Results from each loop return to the model through active learning, in which the acquisition rule ranks candidate experiments by the reduction in predictive error they are expected to yield rather than by the size of the predicted effect. The inner loop cannot be skipped: independent benchmarks report that self-supervised cellular embeddings are matched or exceeded by simple baselines, that accuracy degrades when models are transported across biological contexts, and that metric choice alone can reverse model rankings.

마지막 수정 ·