Data availability
We evaluated xvr and other baseline methods using the following publicly available 2D/3D registration datasets: DeepFluoro57 (https://doi.org/10.7281/T1/IFSXNV), Femur58 (https://zenodo.org/records/15753063) and Ljubljana59 (https://lit.fe.uni-lj.si/en/research/resources/3D-2D-GS-CA). With permission from the original authors, remixed versions of these datasets in a standardized NIfTI and DICOM format are released to the community (https://huggingface.co/datasets/eigenvivek/xvr-data). We used the following 3D imaging datasets to pretrain our patient-agnostic foundation model: CTPelvic1K60 (https://doi.org/10.5281/zenodo.4588402), NITRC MRA Atlas29 (https://www.nitrc.org/projects/icbmmra), TotalSegmentator61 (https://doi.org/10.5281/zenodo.6802613) and the AutoPET partition of the ENHANCE.PET 1.6k62 (https://doi.org/10.57760/sciencedb.34150). The Brigham CTA/DSA and Boston Children’s MRA/DSA datasets are not publicly available due to patient privacy considerations and the terms of the Institutional Review Board approvals governing their use. All datasets were accessed and used in accordance with their respective data-use agreements and licences.
Code availability
The Python package and command line interface for xvr, along with all scripts necessary to replicate the experiments presented in this manuscript, are available at GitHub (https://github.com/eigenvivek/xvr). xvr is implemented in Python (v.3.10+) using DiffDRR (v.0.6.0+) and PyTorch (v.2.2+)63.
References
Unberath, M. et al. The impact of machine learning on 2D/3D registration for image-guided interventions: a systematic review and perspective. Front. Robot. AI 8, 716007 (2021).
Article PubMed PubMed Central Google Scholar
Yip, M. et al. Artificial intelligence meets medical robotics. Science 381, 141–146 (2023).
Article ADS CAS PubMed Google Scholar
Penney, G. P. et al. A comparison of similarity measures for use in 2D/3D medical image registration. IEEE Trans. Med. Imaging 17, 586–595 (1998).
Article ADS CAS PubMed Google Scholar
Knaan, D. & Joskowicz, L. Effective intensity-based 2D/3D rigid registration between fluoroscopic X-ray and CT. In Proc. International Conference on Medical Image Computing and Computer-Assisted Intervention (eds Ellis, R. E. & Peters, T. M.) 351–358 (Springer, 2003).
Grupp, R. B. et al. Automatic annotation of hip anatomy in fluoroscopy for robust and efficient 2D/3D registration. Int. J. Comput. Assist. Radiol. Surg. 15, 759–769 (2020).
Article PubMed PubMed Central Google Scholar
Grimm, M., Esteban, J., Unberath, M. & Navab, N. Pose-dependent weights and domain randomization for fully automatic X-ray to CT registration. IEEE Trans. Med. Imaging 40, 2221–2232 (2021).
Article ADS PubMed Google Scholar
Mahesh, M., Ansari, A. J. & Mettler, F. A. Jr Patient exposure from radiologic and nuclear medicine procedures in the United States and worldwide: 2009–2018. Radiology 307, e221263 (2022).
Article PubMed PubMed Central Google Scholar
Cornelis, F. H., Dzaye, O., Schoellnast, H. & Solomon, S. B. Imaging of interventional therapies in oncology: image guidance, robotics, and fusion systems. In Interventional Oncology (eds Fong, Y. et al.) 1–17 https://doi.org/10.1007/978-3-030-51192-0_19-1 (Springer, 2023).
Jhawar, B. S., Mitsis, D. & Duggal, N. Wrong-sided and wrong-level neurosurgery: a national survey. J. Neurosurg. Spine 7, 467–472 (2007).
Article PubMed Google Scholar
Tonetti, J., Boudissa, M., Kerschbaumer, G. & Seurat, O. Role of 3D intraoperative imaging in orthopedic and trauma surgery. Orthop. Traumatol. Surg. Res. 106, S19–S25 (2020).
Article PubMed Google Scholar
Abumoussa, A. et al. Machine learning for automated and real-time two-dimensional to three-dimensional registration of the spine using a single radiograph. Neurosurg. Focus 54, E16 (2023).
Article PubMed Google Scholar
Naik, R. R., Hoblidar, A., Bhat, S. N., Ampar, N. & Kundangar, R. A hybrid 3D-2D image registration framework for pedicle screw trajectory registration between intraoperative X-ray image and preoperative CT image. J. Imaging 8, 185 (2022).
Article PubMed PubMed Central Google Scholar
Metz, C. T. et al. Patient specific 4D coronary models from ECG-gated CTA data for intra-operative dynamic alignment of CTA with X-ray images. In Proc. International Conference on Medical Image Computing and Computer-Assisted Intervention (eds Guang-Zhong, Y. et al.) 369–376 (Springer, 2009).
Wagner, M., Schafer, S., Strother, C. & Mistretta, C. 4D interventional device reconstruction from biplane fluoroscopy. Med. Phys. 43, 1324–1334 (2016).
Article PubMed PubMed Central Google Scholar
Huynh, E. et al. Artificial intelligence in radiation oncology. Nat. Rev. Clin. Oncol. 17, 771–781 (2020).
Article PubMed Google Scholar
Kim, Y. et al. Telerobotic neurovascular interventions with magnetic manipulation. Sci. Robot. 7, eabg9907 (2022).
Article PubMed PubMed Central Google Scholar
Gu, W., Gao, C., Grupp, R., Fotouhi, J. & Unberath, M. Extended capture range of rigid 2D/3D registration by estimating Riemannian pose gradients. In Proc. International Workshop on Machine Learning in Medical Imaging (eds Liu, M. et al.) 281–291 (Springer, 2020).
Gopalakrishnan, V. & Golland, P. Fast auto-differentiable digitally reconstructed radiographs for solving inverse problems in intraoperative imaging. In Proc. Workshop on Clinical Image-Based Procedures (eds Chen, Y. et al.) 1–11 (Springer, 2022).
Gao, C. et al. A fully differentiable framework for 2D/3D registration and the projective spatial transformers. IEEE Trans. Med. Imaging 43, 275–285 (2023).
Article ADS Google Scholar
Grupp, R. B. et al. Pose estimation of periacetabular osteotomy fragments with intraoperative X-ray navigation. IEEE Trans. Biomed. Eng. 67, 441–452 (2019).
Article PubMed PubMed Central Google Scholar
Bier, B. et al. Learning to detect anatomical landmarks of the pelvis in X-rays from arbitrary views. Int. J. Comput. Assist. Radiol. Surg. 14, 1463–1473 (2019).
Article PubMed PubMed Central Google Scholar
Shrestha, P., Xie, C., Yoshii, Y. & Kitahara, I. Rayemb: arbitrary landmark detection in X-ray images using ray embedding subspace. In Proc. Asian Conference on Computer Vision (eds Cho, M. et al.) 665–681 (Springer, 2024).
Miao, S., Wang, Z. J. & Liao, R. A CNN regression approach for real-time 2D/3D registration. IEEE Trans. Med. Imaging 35, 1352–1363 (2016).
Article ADS Google Scholar
Bui, M., Albarqouni, S., Schrapp, M., Navab, N. & Ilic, S. X-ray posenet: 6 DoF pose estimation for mobile X-ray devices. In Proc. 2017 IEEE Winter Conference on Applications of Computer Vision (WACV) 1036–1044 (IEEE, 2017).
Zhang, B. et al. A patient-specific self-supervised model for automatic X-ray/CT registration. In Proc. International Conference on Medical Image Computing and Computer-Assisted Intervention (eds Greenspan, H. et al.) 515–524 (Springer, 2023).
Gopalakrishnan, V., Dey, N. & Golland, P. Intraoperative 2D/3D image registration via differentiable X-ray rendering. In Proc. IEEE/CVF Conference on Computer Vision and Pattern Recognition 11662–11672 (IEEE, 2024).
Wasserthal, J. et al. TotalSegmentator: robust segmentation of 104 anatomic structures in CT images. Radiol. Artif. Intell. 5, e230024 (2023).
Article PubMed PubMed Central Google Scholar
Jaus, A. et al. Towards unifying anatomy segmentation: automated generation of a full-body CT dataset. In Proc. 2024 IEEE International Conference on Image Processing (ICIP) 41–47 (IEEE, 2024).
Magnetic Resonance Angiography Atlas Dataset. NeuroImaging Tools & Resources Collaboratory https://www.nitrc.org/projects/icbmmra (2017).
Liu, P. et al. Deep learning to segment pelvic bones: large-scale CT datasets and baseline models. Int. J. Comput. Assist. Radiol. Surg. 16, 749–756 (2021).
Article PubMed Google Scholar
Yushkevich, P. A. et al. Fast automatic segmentation of hippocampal subfields and medial temporal lobe subregions in 3 Tesla and 7 Tesla T2-weighted MRI. Alzheimers Dement. 12, P126–P127 (2016).
Article Google Scholar
Zhou, Y., Barnes, C., Lu, J., Yang, J. & Li, H. On the continuity of rotation representations in neural networks. In Proc. IEEE/CVF Conference on Computer Vision and Pattern Recognition 5745–5753 (IEEE, 2019).
Geist, A. R., Frey, J., Zhobro, M., Levina, A. & Martius, G. Learning with 3D rotations, a hitchhiker’s guide to SO(3). In Proc. 41st International Conference on Machine Learning Vol. 235 (eds Salakhutdinov, R. et al.) 15331–15350 (PMLR, 2024).
Lin, C., Hanson, A. J. & Hanson, S. M. Algebraically rigorous quaternion framework for the neural network pose estimation problem. In Proc. IEEE/CVF International Conference on Computer Vision 14097–14106 (IEEE, 2023).
He, K., Zhang, X., Ren, S. & Sun, J. Deep residual learning for image recognition. In Proc. IEEE Conference on Computer Vision and Pattern Recognition 770–778 (IEEE, 2016).
Flepp, R. et al. Automatic multi-view X-ray/CT registration using bone substructure contours. Int. J. Comput. Assist. Radiol. Surg. 20, 1401–1408 (2025).
Article PubMed Google Scholar
Mitrović, U., Špiclin, Ž, Likar, B. & Pernuš, F. 3D-2D registration of cerebral angiograms: a method and evaluation on clinical images. IEEE Trans. Med. Imaging 32, 1550–1563 (2013).
Article ADS PubMed Google Scholar
Xu, M. et al. VesselBoost: a Python toolbox for small blood vessel segmentation in human magnetic resonance angiography data. Aperture Neuro https://doi.org/10.52294/001c.123217 (2024).
Ronneberger, O., Fischer, P. & Brox, T. U-net: convolutional networks for biomedical image segmentation. In Proc. International Conference on Medical Image Computing and Computer-Assisted Intervention (eds Navab, N. et al.) 234–241 (Springer, 2015).
Terzakis, G. & Lourakis, M. A consistently fast and globally optimal solution to the perspective-n-point problem. In Proc. European Conference on Computer Vision (eds Vedaldi, A. et al.) 478–494 (Springer, 2020).
Potente, M. & Mäkinen, T. Vascular heterogeneity and specialization in development and disease. Nat. Rev. Mol. Cell Biol. 18, 477–494 (2017).
Article CAS PubMed Google Scholar
Nuñez, F. B. & Dohna-Schwake, C. Epidemiology, diagnostics, and management of vein of Galen malformation. Pediatr. Neurol. 119, 50–55 (2021).
Article Google Scholar
Varghese, C., Harrison, E. M., O’Grady, G. & Topol, E. J. Artificial intelligence in surgery. Nat. Med. 30, 1257–1268 (2024).
Article CAS PubMed Google Scholar
Jena, R., Sethi, D., Chaudhari, P. & Gee, J. Deep learning in medical image registration: magic or mirage? In Proc. 38th Annual Conference on Neural Information Processing Systems (eds Globerson, A. et al.) 108331–108353 (Curran Associates Inc., 2024).
Dey, N. et al. Learning general-purpose biomedical volume representations using randomized synthesis. In Proc. 13th International Conference on Learning Representations (eds Yue, Y. et al.) (ICLR, 2025).
Gagoski, B. et al. Automated detection and reacquisition of motion-degraded images in fetal HASTE imaging at 3T. Magn. Reson. Med. 87, 1914–1922 (2022).
Article PubMed Google Scholar
Arsigny, V., Commowick, O., Ayache, N. & Pennec, X. A fast and log-euclidean polyaffine framework for locally linear registration. J. Math. Imaging Vis. 33, 222–238 (2009).
Article MathSciNet Google Scholar
Gopalakrishnan, V., Dey, N. & Golland, P. PolyPose: deformable 2D/3D registration via polyrigid transformations. In Advances in Neural Information Processing Systems 60633–60659 (NeurIPS, 2025).
Google Scholar
Hartley, R. & Zisserman, A. Multiple View Geometry in Computer Vision (Cambridge Univ. Press, 2003).
Swinehart, D. F. The Beer-Lambert law. J. Chem. Educ. 39, 333 (1962).
Article CAS Google Scholar
Siddon, R. L. Fast calculation of the exact radiological path for a three-dimensional CT array. Med. Phys. 12, 252–255 (1985).
Article CAS PubMed Google Scholar
Unberath, M. et al. DeepDRR—a catalyst for machine learning in fluoroscopy-guided procedures. In Proc. International Conference on Medical Image Computing and Computer-Assisted Intervention (eds Frangi, A. F. et al.) 98–106 (Springer, 2018).
Milletari, F., Navab, N. & Ahmadi, S.-A. V-net: fully convolutional neural networks for volumetric medical image segmentation. In Proc. 2016 Fourth International Conference on 3D Vision (3DV) 565–571 (IEEE, 2016).
Grupp, R. B., Armand, M. & Taylor, R. H. Patch-based image similarity for intraoperative 2D/3D pelvis registration during periacetabular osteotomy. In Proc. International Workshop on Computer-Assisted and Robotic Endoscopy 153–163 (Springer, 2018).
Ferrara, D. et al. Sharing a whole-/total-body [18F] FDG-PET/CT dataset with CT-derived segmentations: an ENHANCE.PET initiative. Sci. Data 13, 869 (2026).
Sundar, L. K. S. et al. Fully automated, semantic segmentation of whole-body 18F-FDG PET/CT images based on data-centric artificial intelligence. J. Nucl. Med. 63, 1941–1948 (2022).
Article CAS Google Scholar
Grupp, R. et al. Data and code associated with the publication: Automatic annotation of hip anatomy in fluoroscopy for robust and efficient 2D/3D registration. Johns Hopkins Research Data Repository https://doi.org/10.7281/T1/IFSXNV (2020).
Flepp, R. et al. Automatic multi-View X-Ray/CT registration using bone substructure contours. Zenodo https://zenodo.org/records/15753063 (2025).
Mitrović, U. & Špiclin, Ž. Gold standard for 3D-2D registration of cerebral angiograms. Laboratory of Imaging Technologies https://lit.fe.uni-lj.si/en/research/resources/3D-2D-GS-CA (2013).
Pengbo, L. et al. CTPelvic1K Dataset. Zenodo https://doi.org/10.5281/zenodo.4588402 (2020).
Wasserthal, J. Dataset with segmentations of 117 important anatomical structures in 1228 CT images. Zenodo https://doi.org/10.5281/zenodo.6802613 (2023).
Ferrara, D. et al. ENHANCE.PET 1.6k: a whole-/total-body [18F]FDG-PET/CT dataset with CT-derived segmentations. Science Data Bank https://doi.org/10.57760/sciencedb.34150 (2026).
Paszke, A. et al. PyTorch: an imperative style, high-performance deep learning library. In Proc. 33rd International Conference on Neural Information Processing Systems (eds Wallach, H. M. et al.) 8026–8037 (Curran Associates Inc., 2019).
Download references
Acknowledgements
We thank T. v. Walsum for explanation of how C-arm poses are parameterized in DICOM headers. No funders or third parties were involved in the study design, analysis or writing.
Funding
V.G. and P.G. are supported by NIH NIBIB 5T32EB001680−19, the MIT CSAIL-Wistron Program, the MIT-IBM Watson AI Lab, the MIT Jameel Clinic, the MIT Health and Life Sciences Collaborative (HEALS) and the Chou Family Transformative Research Fund. N.D. is supported by the National Institute of Biomedical Imaging and Bioengineering of the National Institutes of Health under award number R01EB033773 and an MGH Neuroscience Transformative Scholar Award. N.H. is supported by NIH K25EB035166. S.F. is supported by NIH R01EB034223.
Ethics declarations
Competing interests
V.G. has received consulting compensation from Noah Medical; the current work is unaffiliated with this consulting role and was performed in his role as a graduate student at the Massachusetts Institute of Technology (MIT). P.G. has received compensation for her services as an advisor from Ver-AI and Xellar Biosystems; she owns equity in VideaHealth, Empallo and Iterative Health as compensation for her services as an advisor; and has received grant support from Wistron and IBM to support her research at MIT; the current work is unaffiliated with her advising roles and was performed in her role as faculty at MIT. The other authors declare no competing interests.
Peer review
Peer review information
Nature thanks Arjun Manrai, Russell Taylor and the other, anonymous, reviewer(s) for their contribution to the peer review of this work. Peer reviewer reports are available.
Additional information
Publisher’s note Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
Extended data figures and tables
Extended Data Fig. 1 Physics-based C-arm simulation in xvr.
a, Our renderer requires two inputs: a 3D volume from which to generate synthetic X-rays and the pose of the C-arm (represented with a camera frustum). Our renderer is differentiable with respect to the C-arm pose, enabling us to use gradient-based optimization to register X-ray images to 3D volumes. b, Optionally, a 3D label map of the preoperative volume can also be used to render X-rays of specific anatomical structures. c, A pictorial overview of trilinear interpolation, one of the ray tracing methods we implemented to render synthetic X-rays, along with Siddon’s method51. d, In addition to developing fully differentiable implementations of ray tracing with trilinear interpolation and Siddon’s method, we also adapt these algorithms to project 3D anatomical labels onto 2D space, enabling structure-specific registration.
Extended Data Fig. 2 Data augmentation pipeline to diversify synthetic X-ray training data.
a, Synthetic X-rays rendered at random C-arm poses from a preoperative volume. b, Projected bone labels from the 3D label map of the preoperative volume overlaid on the synthetic X-rays. c, Examples of data augmentations applied to synthetic X-rays to improve the robustness of pose regression, including Gaussian noise, local contrast enhancement, random masking, and simulated collimation.
Extended Data Fig. 3 Making pose regression neural networks insensitive to intrinsic parameter changes.
a, Using simple resampling operations (i.e., translation, cropping, and bilinear interpolation), an intraoperative X-ray image with some set of intrinsic parameters—height \(H\), width \(W\), pixel spacing \((\Delta X,\Delta Y)\), principal point \(({X}_{0},{Y}_{0})\), and source-to-detector distance—is resampled to the canonical intrinsics used for rendering synthetic X-rays when training the patient-specific neural network. b, As a result, the resampled X-ray has a spatial resolution that matches the network’s training data. c, After this network predicts the pose of the image, xvr can render the predicted X-ray with the intrinsic parameters of the original image. That is, the synthetic X-ray is rendered with the original high-resolution intrinsics for pose refinement via iterative optimization.
Extended Data Fig. 4 Intraoperative pose refinement strategy.
a, Estimation errors of the six pose parameters before (left) and after (right) iterative optimization. All rotational parameters incur roughly ±2.5° of error and in-plane translational parameters (\(x\) and \(z\)) incur roughly ±1.5 mm of error. However, the source-to-object distance (\(y\)) incurs errors of ±15 mm, demonstrating the difficulty in accurately estimating depth. b, Intraoperative pose refinement takes about 1–10 s and successfully overcomes the depth error in the network’s initial pose estimate. c, The loss landscape induced by multiscale normalized cross correlation (mNCC) is smooth in a large neighbourhood around the true pose, broadening the capture radius of our pose refinement strategy. However, mNCC is relatively non-specific about the true pose. d, In contrast, the loss landscape induced by gradient normalized cross correlation (gNCC) results in a more specific optimum at the expense of a rougher landscape further from the true pose. Averaging mNCC and gNCC achieves millimetre-accurate pose refinement.
Full size table
Full size table
Supplementary information
Rights and permissions
Springer Nature or its licensor (e.g. a society or other partner) holds exclusive rights to this article under a publishing agreement with the author(s) or other rightsholder(s); author self-archiving of the accepted manuscript version of this article is solely governed by the terms of such publishing agreement and applicable law.
Reprints and permissions
About this article
Cite this article
Gopalakrishnan, V., Chlorogiannis, DD., Abumoussa, A. et al. Rapid patient-specific neural networks for X-ray to volume registration. Nature (2026). https://doi.org/10.1038/s41586-026-11045-x
Download citation
Received:
Accepted:
Published:
Version of record:
DOI: https://doi.org/10.1038/s41586-026-11045-x