Rapid patient-specific neural networks for X-ray to volume registration

Nature正文已收录本站

Data availability

We evaluated xvr and other baseline methods using the following publicly available 2D/3D registration datasets: DeepFluoro57 (https://doi.org/10.7281/T1/IFSXNV), Femur58 (https://zenodo.org/records/15753063) and Ljubljana59 (https://lit.fe.uni-lj.si/en/research/resources/3D-2D-GS-CA). With permission from the original authors, remixed versions of these datasets in a standardized NIfTI and DICOM format are released to the community (https://huggingface.co/datasets/eigenvivek/xvr-data). We used the following 3D imaging datasets to pretrain our patient-agnostic foundation model: CTPelvic1K60 (https://doi.org/10.5281/zenodo.4588402), NITRC MRA Atlas29 (https://www.nitrc.org/projects/icbmmra), TotalSegmentator61 (https://doi.org/10.5281/zenodo.6802613) and the AutoPET partition of the ENHANCE.PET 1.6k62 (https://doi.org/10.57760/sciencedb.34150). The Brigham CTA/DSA and Boston Children’s MRA/DSA datasets are not publicly available due to patient privacy considerations and the terms of the Institutional Review Board approvals governing their use. All datasets were accessed and used in accordance with their respective data-use agreements and licences.

Code availability

The Python package and command line interface for xvr, along with all scripts necessary to replicate the experiments presented in this manuscript, are available at GitHub (https://github.com/eigenvivek/xvr). xvr is implemented in Python (v.3.10+) using DiffDRR (v.0.6.0+) and PyTorch (v.2.2+)63.

References

  1. Unberath, M. et al. The impact of machine learning on 2D/3D registration for image-guided interventions: a systematic review and perspective. Front. Robot. AI 8, 716007 (2021).

    Article  PubMed  PubMed Central  Google Scholar 

  2. Yip, M. et al. Artificial intelligence meets medical robotics. Science 381, 141–146 (2023).

    Article  ADS  CAS  PubMed  Google Scholar 

  3. Penney, G. P. et al. A comparison of similarity measures for use in 2D/3D medical image registration. IEEE Trans. Med. Imaging 17, 586–595 (1998).

    Article  ADS  CAS  PubMed  Google Scholar 

  4. Knaan, D. & Joskowicz, L. Effective intensity-based 2D/3D rigid registration between fluoroscopic X-ray and CT. In Proc. International Conference on Medical Image Computing and Computer-Assisted Intervention (eds Ellis, R. E. & Peters, T. M.) 351–358 (Springer, 2003).

  5. Grupp, R. B. et al. Automatic annotation of hip anatomy in fluoroscopy for robust and efficient 2D/3D registration. Int. J. Comput. Assist. Radiol. Surg. 15, 759–769 (2020).

    Article  PubMed  PubMed Central  Google Scholar 

  6. Grimm, M., Esteban, J., Unberath, M. & Navab, N. Pose-dependent weights and domain randomization for fully automatic X-ray to CT registration. IEEE Trans. Med. Imaging 40, 2221–2232 (2021).

    Article  ADS  PubMed  Google Scholar 

  7. Mahesh, M., Ansari, A. J. & Mettler, F. A. Jr Patient exposure from radiologic and nuclear medicine procedures in the United States and worldwide: 2009–2018. Radiology 307, e221263 (2022).

    Article  PubMed  PubMed Central  Google Scholar 

  8. Cornelis, F. H., Dzaye, O., Schoellnast, H. & Solomon, S. B. Imaging of interventional therapies in oncology: image guidance, robotics, and fusion systems. In Interventional Oncology (eds Fong, Y. et al.) 1–17 https://doi.org/10.1007/978-3-030-51192-0_19-1 (Springer, 2023).

  9. Jhawar, B. S., Mitsis, D. & Duggal, N. Wrong-sided and wrong-level neurosurgery: a national survey. J. Neurosurg. Spine 7, 467–472 (2007).

    Article  PubMed  Google Scholar 

  10. Tonetti, J., Boudissa, M., Kerschbaumer, G. & Seurat, O. Role of 3D intraoperative imaging in orthopedic and trauma surgery. Orthop. Traumatol. Surg. Res. 106, S19–S25 (2020).

    Article  PubMed  Google Scholar 

  11. Abumoussa, A. et al. Machine learning for automated and real-time two-dimensional to three-dimensional registration of the spine using a single radiograph. Neurosurg. Focus 54, E16 (2023).

    Article  PubMed  Google Scholar 

  12. Naik, R. R., Hoblidar, A., Bhat, S. N., Ampar, N. & Kundangar, R. A hybrid 3D-2D image registration framework for pedicle screw trajectory registration between intraoperative X-ray image and preoperative CT image. J. Imaging 8, 185 (2022).

    Article  PubMed  PubMed Central  Google Scholar 

  13. Metz, C. T. et al. Patient specific 4D coronary models from ECG-gated CTA data for intra-operative dynamic alignment of CTA with X-ray images. In Proc. International Conference on Medical Image Computing and Computer-Assisted Intervention (eds Guang-Zhong, Y. et al.) 369–376 (Springer, 2009).

  14. Wagner, M., Schafer, S., Strother, C. & Mistretta, C. 4D interventional device reconstruction from biplane fluoroscopy. Med. Phys. 43, 1324–1334 (2016).

    Article  PubMed  PubMed Central  Google Scholar 

  15. Huynh, E. et al. Artificial intelligence in radiation oncology. Nat. Rev. Clin. Oncol. 17, 771–781 (2020).

    Article  PubMed  Google Scholar 

  16. Kim, Y. et al. Telerobotic neurovascular interventions with magnetic manipulation. Sci. Robot. 7, eabg9907 (2022).

    Article  PubMed  PubMed Central  Google Scholar 

  17. Gu, W., Gao, C., Grupp, R., Fotouhi, J. & Unberath, M. Extended capture range of rigid 2D/3D registration by estimating Riemannian pose gradients. In Proc. International Workshop on Machine Learning in Medical Imaging (eds Liu, M. et al.) 281–291 (Springer, 2020).

  18. Gopalakrishnan, V. & Golland, P. Fast auto-differentiable digitally reconstructed radiographs for solving inverse problems in intraoperative imaging. In Proc. Workshop on Clinical Image-Based Procedures (eds Chen, Y. et al.) 1–11 (Springer, 2022).

  19. Gao, C. et al. A fully differentiable framework for 2D/3D registration and the projective spatial transformers. IEEE Trans. Med. Imaging 43, 275–285 (2023).

    Article  ADS  Google Scholar 

  20. Grupp, R. B. et al. Pose estimation of periacetabular osteotomy fragments with intraoperative X-ray navigation. IEEE Trans. Biomed. Eng. 67, 441–452 (2019).

    Article  PubMed  PubMed Central  Google Scholar 

  21. Bier, B. et al. Learning to detect anatomical landmarks of the pelvis in X-rays from arbitrary views. Int. J. Comput. Assist. Radiol. Surg. 14, 1463–1473 (2019).

    Article  PubMed  PubMed Central  Google Scholar 

  22. Shrestha, P., Xie, C., Yoshii, Y. & Kitahara, I. Rayemb: arbitrary landmark detection in X-ray images using ray embedding subspace. In Proc. Asian Conference on Computer Vision (eds Cho, M. et al.) 665–681 (Springer, 2024).

  23. Miao, S., Wang, Z. J. & Liao, R. A CNN regression approach for real-time 2D/3D registration. IEEE Trans. Med. Imaging 35, 1352–1363 (2016).

    Article  ADS  Google Scholar 

  24. Bui, M., Albarqouni, S., Schrapp, M., Navab, N. & Ilic, S. X-ray posenet: 6 DoF pose estimation for mobile X-ray devices. In Proc. 2017 IEEE Winter Conference on Applications of Computer Vision (WACV) 1036–1044 (IEEE, 2017).

  25. Zhang, B. et al. A patient-specific self-supervised model for automatic X-ray/CT registration. In Proc. International Conference on Medical Image Computing and Computer-Assisted Intervention (eds Greenspan, H. et al.) 515–524 (Springer, 2023).

  26. Gopalakrishnan, V., Dey, N. & Golland, P. Intraoperative 2D/3D image registration via differentiable X-ray rendering. In Proc. IEEE/CVF Conference on Computer Vision and Pattern Recognition 11662–11672 (IEEE, 2024).

  27. Wasserthal, J. et al. TotalSegmentator: robust segmentation of 104 anatomic structures in CT images. Radiol. Artif. Intell. 5, e230024 (2023).

    Article  PubMed  PubMed Central  Google Scholar 

  28. Jaus, A. et al. Towards unifying anatomy segmentation: automated generation of a full-body CT dataset. In Proc. 2024 IEEE International Conference on Image Processing (ICIP) 41–47 (IEEE, 2024).

  29. Magnetic Resonance Angiography Atlas Dataset. NeuroImaging Tools & Resources Collaboratory https://www.nitrc.org/projects/icbmmra (2017).

  30. Liu, P. et al. Deep learning to segment pelvic bones: large-scale CT datasets and baseline models. Int. J. Comput. Assist. Radiol. Surg. 16, 749–756 (2021).

    Article  PubMed  Google Scholar 

  31. Yushkevich, P. A. et al. Fast automatic segmentation of hippocampal subfields and medial temporal lobe subregions in 3 Tesla and 7 Tesla T2-weighted MRI. Alzheimers Dement. 12, P126–P127 (2016).

    Article  Google Scholar 

  32. Zhou, Y., Barnes, C., Lu, J., Yang, J. & Li, H. On the continuity of rotation representations in neural networks. In Proc. IEEE/CVF Conference on Computer Vision and Pattern Recognition 5745–5753 (IEEE, 2019).

  33. Geist, A. R., Frey, J., Zhobro, M., Levina, A. & Martius, G. Learning with 3D rotations, a hitchhiker’s guide to SO(3). In Proc. 41st International Conference on Machine Learning Vol. 235 (eds Salakhutdinov, R. et al.) 15331–15350 (PMLR, 2024).

  34. Lin, C., Hanson, A. J. & Hanson, S. M. Algebraically rigorous quaternion framework for the neural network pose estimation problem. In Proc. IEEE/CVF International Conference on Computer Vision 14097–14106 (IEEE, 2023).

  35. He, K., Zhang, X., Ren, S. & Sun, J. Deep residual learning for image recognition. In Proc. IEEE Conference on Computer Vision and Pattern Recognition 770–778 (IEEE, 2016).

  36. Flepp, R. et al. Automatic multi-view X-ray/CT registration using bone substructure contours. Int. J. Comput. Assist. Radiol. Surg. 20, 1401–1408 (2025).

    Article  PubMed  Google Scholar 

  37. Mitrović, U., Špiclin, Ž, Likar, B. & Pernuš, F. 3D-2D registration of cerebral angiograms: a method and evaluation on clinical images. IEEE Trans. Med. Imaging 32, 1550–1563 (2013).

    Article  ADS  PubMed  Google Scholar 

  38. Xu, M. et al. VesselBoost: a Python toolbox for small blood vessel segmentation in human magnetic resonance angiography data. Aperture Neuro https://doi.org/10.52294/001c.123217 (2024).

  39. Ronneberger, O., Fischer, P. & Brox, T. U-net: convolutional networks for biomedical image segmentation. In Proc. International Conference on Medical Image Computing and Computer-Assisted Intervention (eds Navab, N. et al.) 234–241 (Springer, 2015).

  40. Terzakis, G. & Lourakis, M. A consistently fast and globally optimal solution to the perspective-n-point problem. In Proc. European Conference on Computer Vision (eds Vedaldi, A. et al.) 478–494 (Springer, 2020).

  41. Potente, M. & Mäkinen, T. Vascular heterogeneity and specialization in development and disease. Nat. Rev. Mol. Cell Biol. 18, 477–494 (2017).

    Article  CAS  PubMed  Google Scholar 

  42. Nuñez, F. B. & Dohna-Schwake, C. Epidemiology, diagnostics, and management of vein of Galen malformation. Pediatr. Neurol. 119, 50–55 (2021).

    Article  Google Scholar 

  43. Varghese, C., Harrison, E. M., O’Grady, G. & Topol, E. J. Artificial intelligence in surgery. Nat. Med. 30, 1257–1268 (2024).

    Article  CAS  PubMed  Google Scholar 

  44. Jena, R., Sethi, D., Chaudhari, P. & Gee, J. Deep learning in medical image registration: magic or mirage? In Proc. 38th Annual Conference on Neural Information Processing Systems (eds Globerson, A. et al.) 108331–108353 (Curran Associates Inc., 2024).

  45. Dey, N. et al. Learning general-purpose biomedical volume representations using randomized synthesis. In Proc. 13th International Conference on Learning Representations (eds Yue, Y. et al.) (ICLR, 2025).

  46. Gagoski, B. et al. Automated detection and reacquisition of motion-degraded images in fetal HASTE imaging at 3T. Magn. Reson. Med. 87, 1914–1922 (2022).

    Article  PubMed  Google Scholar 

  47. Arsigny, V., Commowick, O., Ayache, N. & Pennec, X. A fast and log-euclidean polyaffine framework for locally linear registration. J. Math. Imaging Vis. 33, 222–238 (2009).

    Article  MathSciNet  Google Scholar 

  48. Gopalakrishnan, V., Dey, N. & Golland, P. PolyPose: deformable 2D/3D registration via polyrigid transformations. In Advances in Neural Information Processing Systems 60633–60659 (NeurIPS, 2025).

    Google Scholar 

  49. Hartley, R. & Zisserman, A. Multiple View Geometry in Computer Vision (Cambridge Univ. Press, 2003).

  50. Swinehart, D. F. The Beer-Lambert law. J. Chem. Educ. 39, 333 (1962).

    Article  CAS  Google Scholar 

  51. Siddon, R. L. Fast calculation of the exact radiological path for a three-dimensional CT array. Med. Phys. 12, 252–255 (1985).

    Article  CAS  PubMed  Google Scholar 

  52. Unberath, M. et al. DeepDRR—a catalyst for machine learning in fluoroscopy-guided procedures. In Proc. International Conference on Medical Image Computing and Computer-Assisted Intervention (eds Frangi, A. F. et al.) 98–106 (Springer, 2018).

  53. Milletari, F., Navab, N. & Ahmadi, S.-A. V-net: fully convolutional neural networks for volumetric medical image segmentation. In Proc. 2016 Fourth International Conference on 3D Vision (3DV) 565–571 (IEEE, 2016).

  54. Grupp, R. B., Armand, M. & Taylor, R. H. Patch-based image similarity for intraoperative 2D/3D pelvis registration during periacetabular osteotomy. In Proc. International Workshop on Computer-Assisted and Robotic Endoscopy 153–163 (Springer, 2018).

  55. Ferrara, D. et al. Sharing a whole-/total-body [18F] FDG-PET/CT dataset with CT-derived segmentations: an ENHANCE.PET initiative. Sci. Data 13, 869 (2026).

  56. Sundar, L. K. S. et al. Fully automated, semantic segmentation of whole-body 18F-FDG PET/CT images based on data-centric artificial intelligence. J. Nucl. Med. 63, 1941–1948 (2022).

    Article  CAS  Google Scholar 

  57. Grupp, R. et al. Data and code associated with the publication: Automatic annotation of hip anatomy in fluoroscopy for robust and efficient 2D/3D registration. Johns Hopkins Research Data Repository https://doi.org/10.7281/T1/IFSXNV (2020).

  58. Flepp, R. et al. Automatic multi-View X-Ray/CT registration using bone substructure contours. Zenodo https://zenodo.org/records/15753063 (2025).

  59. Mitrović, U. & Špiclin, Ž. Gold standard for 3D-2D registration of cerebral angiograms. Laboratory of Imaging Technologies https://lit.fe.uni-lj.si/en/research/resources/3D-2D-GS-CA (2013).

  60. Pengbo, L. et al. CTPelvic1K Dataset. Zenodo https://doi.org/10.5281/zenodo.4588402 (2020).

  61. Wasserthal, J. Dataset with segmentations of 117 important anatomical structures in 1228 CT images. Zenodo https://doi.org/10.5281/zenodo.6802613 (2023).

  62. Ferrara, D. et al. ENHANCE.PET 1.6k: a whole-/total-body [18F]FDG-PET/CT dataset with CT-derived segmentations. Science Data Bank https://doi.org/10.57760/sciencedb.34150 (2026).

  63. Paszke, A. et al. PyTorch: an imperative style, high-performance deep learning library. In Proc. 33rd International Conference on Neural Information Processing Systems (eds Wallach, H. M. et al.) 8026–8037 (Curran Associates Inc., 2019).

Download references

Acknowledgements

We thank T. v. Walsum for explanation of how C-arm poses are parameterized in DICOM headers. No funders or third parties were involved in the study design, analysis or writing.

Funding

V.G. and P.G. are supported by NIH NIBIB 5T32EB001680−19, the MIT CSAIL-Wistron Program, the MIT-IBM Watson AI Lab, the MIT Jameel Clinic, the MIT Health and Life Sciences Collaborative (HEALS) and the Chou Family Transformative Research Fund. N.D. is supported by the National Institute of Biomedical Imaging and Bioengineering of the National Institutes of Health under award number R01EB033773 and an MGH Neuroscience Transformative Scholar Award. N.H. is supported by NIH K25EB035166. S.F. is supported by NIH R01EB034223.

Author information

Authors and Affiliations

  1. Harvard-MIT Health Sciences and Technology, Massachusetts Institute of Technology, Cambridge, MA, USA

    Vivek Gopalakrishnan & Polina Golland

  2. Computer Science and Artificial Intelligence Laboratory, Massachusetts Institute of Technology, Cambridge, MA, USA

    Vivek Gopalakrishnan, Neel Dey & Polina Golland

  3. Department of Radiology, Harvard Medical School, Boston, MA, USA

    Vivek Gopalakrishnan, David-Dimitris Chlorogiannis, Nazim Haouchine, Sarah Frisken & Neel Dey

  4. Saint Luke’s Marion Bloch Neuroscience Institute, Kansas City, MO, USA

    Andrew Abumoussa

  5. Pediatric Critical Care Medicine, Massachusetts General Hospital, Boston, MA, USA

    Anna M. Larson

  6. Department of Interventional Neuroradiology, Boston Children’s Hospital, Boston, MA, USA

    Anna M. Larson & Darren B. Orbach

  7. Athinoula A. Martinos Center for Biomedical Imaging, Massachusetts General Hospital, Boston, MA, USA

    Neel Dey

Authors

  1. Vivek Gopalakrishnan
  2. David-Dimitris Chlorogiannis
  3. Andrew Abumoussa
  4. Anna M. Larson
  5. Nazim Haouchine
  6. Darren B. Orbach
  7. Sarah Frisken
  8. Neel Dey
  9. Polina Golland

Contributions

V.G. collected and standardized public data, developed code, trained models, ran experiments, analysed results and created figures. V.G. and N.D. wrote the manuscript. All of the authors reviewed the manuscript and provided revisions and feedback. N.D., S.F. and P.G. provided technical advice. V.G. and N.D. found public benchmarking datasets. D.-D.C., A.A., A.M.L. and D.B.O. provided clinical insights and feedback. D.-D.C., A.M.L. and N.H. manually annotated clinical data. N.D. and P.G. served as co-principal investigators for this study and advised on technical details. P.G. provided research support for the project.

Corresponding authors

Correspondence to Vivek Gopalakrishnan, Neel Dey or Polina Golland.

Ethics declarations

Competing interests

V.G. has received consulting compensation from Noah Medical; the current work is unaffiliated with this consulting role and was performed in his role as a graduate student at the Massachusetts Institute of Technology (MIT). P.G. has received compensation for her services as an advisor from Ver-AI and Xellar Biosystems; she owns equity in VideaHealth, Empallo and Iterative Health as compensation for her services as an advisor; and has received grant support from Wistron and IBM to support her research at MIT; the current work is unaffiliated with her advising roles and was performed in her role as faculty at MIT. The other authors declare no competing interests.

Peer review

Peer review information

Nature thanks Arjun Manrai, Russell Taylor and the other, anonymous, reviewer(s) for their contribution to the peer review of this work. Peer reviewer reports are available.

Additional information

Publisher’s note Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.

Extended data figures and tables

Extended Data Fig. 1 Physics-based C-arm simulation in xvr.

a, Our renderer requires two inputs: a 3D volume from which to generate synthetic X-rays and the pose of the C-arm (represented with a camera frustum). Our renderer is differentiable with respect to the C-arm pose, enabling us to use gradient-based optimization to register X-ray images to 3D volumes. b, Optionally, a 3D label map of the preoperative volume can also be used to render X-rays of specific anatomical structures. c, A pictorial overview of trilinear interpolation, one of the ray tracing methods we implemented to render synthetic X-rays, along with Siddon’s method51. d, In addition to developing fully differentiable implementations of ray tracing with trilinear interpolation and Siddon’s method, we also adapt these algorithms to project 3D anatomical labels onto 2D space, enabling structure-specific registration.

Extended Data Fig. 2 Data augmentation pipeline to diversify synthetic X-ray training data.

a, Synthetic X-rays rendered at random C-arm poses from a preoperative volume. b, Projected bone labels from the 3D label map of the preoperative volume overlaid on the synthetic X-rays. c, Examples of data augmentations applied to synthetic X-rays to improve the robustness of pose regression, including Gaussian noise, local contrast enhancement, random masking, and simulated collimation.

Extended Data Fig. 3 Making pose regression neural networks insensitive to intrinsic parameter changes.

a, Using simple resampling operations (i.e., translation, cropping, and bilinear interpolation), an intraoperative X-ray image with some set of intrinsic parameters—height \(H\), width \(W\), pixel spacing \((\Delta X,\Delta Y)\), principal point \(({X}_{0},{Y}_{0})\), and source-to-detector distance—is resampled to the canonical intrinsics used for rendering synthetic X-rays when training the patient-specific neural network. b, As a result, the resampled X-ray has a spatial resolution that matches the network’s training data. c, After this network predicts the pose of the image, xvr can render the predicted X-ray with the intrinsic parameters of the original image. That is, the synthetic X-ray is rendered with the original high-resolution intrinsics for pose refinement via iterative optimization.

Extended Data Fig. 4 Intraoperative pose refinement strategy.

a, Estimation errors of the six pose parameters before (left) and after (right) iterative optimization. All rotational parameters incur roughly ±2.5° of error and in-plane translational parameters (\(x\) and \(z\)) incur roughly ±1.5 mm of error. However, the source-to-object distance (\(y\)) incurs errors of ±15 mm, demonstrating the difficulty in accurately estimating depth. b, Intraoperative pose refinement takes about 1–10 s and successfully overcomes the depth error in the network’s initial pose estimate. c, The loss landscape induced by multiscale normalized cross correlation (mNCC) is smooth in a large neighbourhood around the true pose, broadening the capture radius of our pose refinement strategy. However, mNCC is relatively non-specific about the true pose. d, In contrast, the loss landscape induced by gradient normalized cross correlation (gNCC) results in a more specific optimum at the expense of a rougher landscape further from the true pose. Averaging mNCC and gNCC achieves millimetre-accurate pose refinement.

Extended Data Table 1 Multiple pose estimation error metrics reported as the median and interquartile range (mm) and submillimeter success rate (%)

Full size table

Extended Data Table 2 Ranking of each method’s final pose estimation error (mTRE) per dataset, reported as the median and interquartile range (mm)

Full size table

Supplementary information

About this article

Check for updates. Verify currency and authenticity via CrossMark

Cite this article

Gopalakrishnan, V., Chlorogiannis, DD., Abumoussa, A. et al. Rapid patient-specific neural networks for X-ray to volume registration. Nature (2026). https://doi.org/10.1038/s41586-026-11045-x

Download citation

  • Received:

  • Accepted:

  • Published:

  • Version of record:

  • DOI: https://doi.org/10.1038/s41586-026-11045-x