Main
The CRISPR–Cas9 toolkit provides unprecedented opportunities to engineer the human genome. However, as most genetic disorders are caused by heterogeneous mutations, generalizable approaches to replace larger DNA sequences could provide universal gene-specific solutions to remedy most pathogenic mutations that cause disease. Furthermore, there are many potential applications for site-specific integration of synthetic transgenes, such as installing antigen receptors to redirect immune cells2,3.
Therapeutic nuclease-based editing of haematopoietic stem and progenitor cells (HSPCs) has enabled the development of exa-cel, the first CRISPR–Cas9 therapy approved by regulatory agencies in North America and Europe1,4,5,6. As the field advances, cell context-specific barriers are emerging. The development of gene correction or knock-in strategies based on homology-directed repair (HDR), which occurs in the S and G2 phases of the cell cycle7, are restricted to actively dividing cells. Nuclease-driven DNA double-strand breaks (DSBs) also generate safety concerns regarding off-target editing8,9 and complex genomic rearrangements10,11 that are exacerbated in proliferating cells12,13. Gene editing technologies that do not require nuclease-driven DSBs and cell cycle progression could overcome these hurdles.
A variety of genome editors have been developed to precisely substitute DNA sequences without requiring DNA DSBs, including base and prime editors14,15,16. Although prime editing offers high product purity with limited off-target activity and genotoxicity16,17,18,19, it currently supports only short-to-medium genomic changes of less than 250 bp20. A two-step genome editing process, in which a prime editor installs a landing pad followed by recombinase-driven DNA integration, can expand this targeting range to large-sized modifications of more than 250 bp, although this is dependent on multicomponent delivery and sequential action21,22,23. Recombinase- and transposase-based technologies21,22,23,24,25,26,27,28,29 currently rely on double-stranded DNA (dsDNA) or adeno-associated virus (AAV) donors that induce toxicity due to cGAS–STING sensing30,31 and p53 activation32,33, respectively. Given their lower immunogenicity, single-stranded DNA (ssDNA) and circular ssDNA (cssDNA) donors are better tolerated by many cells30,34,35 and could potentially facilitate effective site-specific exon recoding or therapeutic transgene integration.
A new paradigm in molecular cloning emerged in 2009 when Gibson et al. reported the enzymatic assembly of kilobase-sized DNA molecules in vitro36. Molecular cloning through in cellulo DNA assembly in commonly used strains of Escherichia coli was reported more than a decade before that, but saw limited adoption37,38,39. DNA assembly in widely used cloning strains (such as E. coli DH5α) depends on an incompletely characterized RecA-independent recombination (RAIR) mechanism39. These DNA assembly modalities require common steps: the generation of complementary ssDNA overlaps, homology-directed annealing, fill-in synthesis and ligation. Whereas RAIR occurs via 5′ ssDNA homology-directed annealing39, in vitro isothermal Gibson assembly occurs through 3′ ssDNA overlaps36. Notably, strategies based on paired prime editing guide RNAs (pegRNAs) require similar DNA repair steps20,21,40,41, suggesting that human cells may possess the ability to permanently incorporate exogenous DNA sequences to their genome via in cellulo DNA assembly.
Inspired by RAIR and isothermal Gibson assembly cloning36,37,38,39, we developed an approach for site-specific DNA assembly and integration in human cells using CRISPR-targeted 3′ flap synthesis. We applied this method, which we term prime assembly (PA), to perform targeted exon recoding, transgene integration and megabase-scale rearrangements at multiple loci using either dsDNA or ssDNA donors. PA was active in human primary CD3+ T cells and CD34+ HSPCs, as well as in non-dividing cells. Our study establishes a new modality to introduce medium to large genetic modifications to the human genome.
Exon recoding using PA
We hypothesized that synthesizing 3′ flaps using prime editing could enable site-specific in cellulo DNA assembly and integration of exogenous ssDNA donors, or a dsDNA donor with 3′ overhangs (Fig. 1a). After homology-directed donor(s) annealing to the 3′ flaps, endogenous cellular DNA repair pathways could ensure the excision of the unedited duplex, free ssDNA fill-in synthesis and ligation (Fig. 1a). PA thus shares characteristics of RAIR and isothermal Gibson assembly cloning36,37,38,39.
a, Schematic representation of PA. F, forward; R, reverse. b, Schematic of recoding of the DC cluster in TINF2 using PA. c, PA and imprecise allele quantification at TINF2 as determined by CRISPResso2 analysis from amplicon sequencing. K562 cells were electroporated with PA vectors and donors with (3′-odsDNA_v1) or without (ssDNA_v1) prior annealing, and genomic DNA was collected 3 days post-nucleofection. Data are mean ± s.e.m. from n = 3 independent biological replicates. d, Same as in c, comparing donors with or without three phosphorothioate bonds at both 5′ and 3′ ends. Data are plotted as mean ± s.e.m. from n = 3 independent biological replicates. e, Same as in c, using ssDNA_v5 donors, with cells cultured in the presence or absence of AZD7648 and PolQi1 for 3 days before genomic DNA was collected. Data are mean ± s.e.m. from n = 3 independent biological replicates. f, Genome-wide Donor-seq integration profile after PA at TINF2 in HEK293T cells from the samples in Extended Data Fig. 1f. Each plot represents the total number of deduplicated UMI reads from n = 3 independent biological replicates.
Source data
As a proof of concept, we first designed a medium-sized therapeutic exon-recoding strategy. Dominant gain-of-function mutations in the DC cluster, a small region of TINF2 exon 6 that encodes 30 amino acids (Fig. 1b and Extended Data Fig. 1), result in very short telomeres and a bone marrow failure syndrome known as dyskeratosis congenita42,43. We reasoned that a single approach could potentially remedy nearly all known pathogenic mutations in the DC cluster. To prevent DNA flaps and donors annealing to the endogenous genomic sequence, we performed codon optimization throughout the targeted 106-bp region to recode the exon without altering the amino acid sequence (Fig. 1b and Extended Data Fig. 1). We electroporated K562 cells with plasmids encoding for PEmax44 and a pair of engineered prime editing guide RNAs (epegRNAs) targeting TINF2, along with various concentrations of synthetic ssDNA donors, either annealed (3′-overhang dsDNA (3′-odsDNA)) or not (ssDNA) prior to electroporation. We observed an average of 30.0% editing with ssDNA donors, as determined by amplicon sequencing (Fig. 1b,c). Electroporating 16 pmol of synthetic ssDNA donors protected with 3 phosphorothioate bonds on both 5′ and 3′ ends, without prior in vitro annealing, yielded the highest editing efficiency of 33.1% (Fig. 1d).
Our initial design (v1) relied on 25-nucleotide (nt) flaps and no fill-in synthesis. We then tested whether shorter flaps could support PA and observed an abrogation of editing with flaps shorter than 14 nt (Extended Data Fig. 1b). We designed ssDNA donors with different complementary overlap lengths (v1–v4) between the two ssDNA donors, and observed similar efficiencies with overlaps of 20–56 nt, suggesting that short fill-in synthesis is not a limiting factor (Extended Data Fig. 1c,d), as previously reported for twin prime editing20,21. We switched to the PEmax-La (PE7) editor45 with standard (La-accessible) pegRNAs, and tested flap lengths ranging from 20–35 nt and ssDNA donors sharing 36 nt of overlaps (v5). We observed similar efficiencies with flaps ranging from 20–35 nt (Extended Data Fig. 1e). Characterization of unintended outcomes using amplicon sequencing revealed the presence of flap integration without insertion or deletion (reminiscent of standard prime editing), imprecise PA alleles composed of flap integration accompanied by deletion, mainly of the sequences located between the flaps, imprecise PA alleles with insertions including but not limited to pegRNA scaffold incorporation, and insertion–deletion mutations (indels) at nick sites (Extended Data Fig. 2). PA requires two single-strand breaks (SSBs) and may generate staggered DSB intermediates, as previously reported for prime editing with two nicks16,44,46. Pharmacological inhibition of non-homologous end joining (NHEJ) using the DNA-PK inhibitor AZD7648 and microhomology-mediated end joining (MMEJ) using the Polθ inhibitor PolQi1 has recently been shown to decrease TwinPE-mediated indels47. Using these inhibitors, the ratio of precise PA to imprecise alleles increased 4.1- and 8.5-fold with AZD7648 alone or combined with PolQi1, respectively (Fig. 1e). NHEJ and MMEJ inhibition also increased this purity ratio by up to 2.6-fold in HEK293T cells, although with little to no effect on flap integration without indels, suggesting that the latter is mechanistically distinct (Extended Data Fig. 1f). We performed in–out PCR to characterize the precision of donor–target junctions and observed 94.1% precise 5′ junctions and 97.6% precise 3′ junctions without AZD7648 and PolQi1, and 97.8% precise 5′ junctions and 98.8% precise 3′ junctions with AZD7648 and PolQi1 (Extended Data Fig. 1g). We recommend out–out genotyping for comprehensive characterization of imprecise on-target repair events, only a minority of which are detectable by in–out junction analysis. Together, these results suggest that imprecise NHEJ and MMEJ repair outcomes, in particular flap integration with deletion, may compete with precise PA.
PA requires complementary base pairing at a targeted locus using two pegRNAs, a process that involves multiple layers of specificity checkpoints20 for productive flap generation, DNA assembly and integration. We reasoned that this platform could provide high genome-wide specificity. We adapted GUIDE-seq (genome-wide, unbiased identification of DSBs enabled by sequencing) and PE-tag8,18 to assess the genome-wide donor integration profile based on unidirectional amplification from donor DNA-specific bait sequences, a method we refer to as Donor-seq. Using this approach, we observed no recurrent guide-dependent off-target ssDNA donor integration (Fig. 1f). Together, these results demonstrate that PA enables efficient medium-sized exon recoding with high genome-wide specificity.
Site-specific transgene integration
Efficient HDR has been reported using 3′-odsDNA donors48. We reasoned that we could repurpose this type of donor to achieve targeted transgene integration by generating 3′ overhangs that are compatible with 3′ PA flaps. We generated 3′-odsDNA donors by PCR amplification and lambda exonuclease digestion under the protection of phosphorothioate bonds to generate complementary 3′ overhangs48 (Supplementary Fig. 1). We electroporated K562 cells with 3′-odsDNA donors and pairs of pegRNAs targeting the TRAC locus. Our leading pair of pegRNAs achieved an average PA allele frequency of 4.9%, as determined by droplet digital PCR (ddPCR) (Extended Data Fig. 3a,b). We amplified the integration junctions and confirmed precise site-specific integration by Sanger sequencing across all biological replicates (Extended Data Fig. 3c). We then targeted eGFP and CD19–CAR transgenes to transcriptionally active loci to confirm functional integration (Fig. 2a and Supplementary Fig. 2). The incorporation of the flip and extension (F + E) scaffold modifications to the IL2RG pegRNAs increased PA from an average of 15.0% to 21.2% eGFP+ cells with NHEJ and MMEJ inhibitors, and from 11.3% to 15.3% when overexpressing the 53BP1 inhibitor i53 (Fig. 2b).
a, Schematic representation of targeted eGFP integration using a dsDNA donor without or with 3′-overhangs (3′-odsDNA). Stabilizing phosphorothioate bonds are illustrated with orange stars. b, Percentage of eGFP+ cells as determined by flow cytometry. K562 cells were electroporated with PA vectors encoding for pegRNAs with standard or F + E scaffold and 3′-odsDNA donor targeting IL2RG, and cultured in the presence or absence of AZD7648 and PolQi1. Where indicated, an i53 overexpression plasmid was co-delivered during electroporation. Data are plotted as mean ± s.e.m. from n = 3 independent biological replicates. c, Same as in b, using 3′-odsDNA or a dsDNA donor. Data are mean ± s.e.m. from n = 3 independent biological replicates. d, Same as in c, using a dsDNA donor for eGFP integration at IL2RG in K562 cells and TRAC in Jurkat cells. Cells were treated with AZD7648 and PolQi1. Data are mean ± s.e.m. from n = 3 independent biological replicates. e, Percentage of eGFP+ cells as determined by flow cytometry. K562 cells were electroporated with PA vectors and dsDNA donor targeting AAVS1, and cultured in the presence of AZD7648 and PolQi1. Where indicated, puromycin selection was performed 3 days post-nucleofection. Data are mean ± s.e.m. from n = 3 independent biological replicates. f, Schematics of split dsDNA donors integration at AAVS1 via PA. g, Same as in e, without puromycin selection. Where indicated, an equimolar ratio of split 3′-odsDNA or dsDNA donors were co-electroporated, and cells were treated with AZD7648 and PolQi1. Data are mean ± s.e.m. from n = 3 independent biological replicates. h, Schematics of integration of four split dsDNA donors at AAVS1 via PA. dsDNA donors are numbered 1 to 4. i, Same as in e,g, using a 3.2 kb dsDNA donor split in 2 (split (2), donor 1 and donor 2-4), 3 (split (3), donor 1, donor 2–3 and donor 4) or 4 (split (4), donor 1, donor 2, donor 3 and donor 4). Data are plotted as mean ± s.e.m. from n = 3 independent biological replicates.
Source data
Leveraging cellular 5′ to 3′ exonuclease activity (Fig. 2a) could facilitate the generation of linear dsDNA donors without requiring additional steps of in vitro exonuclease digestion and purification. We electroporated donors with (3′-odsDNA) or without (dsDNA) preceding in vitro lambda exonuclease digestion and observed a decrease in editing efficiency from an average of 8.9% to 2.7% eGFP+ cells, suggesting that 3′ overhangs can be generated in cellulo, although at the cost of decreased efficiency (Fig. 2c). Notably, inhibition of DNA-PK with AZD7648 markedly increased the average efficiency to 25.9% eGFP+ cells when using a dsDNA donor, suggesting that DNA repair factors might act directly on DNA donors to reduce PA efficiency (Fig. 2c). We then designed reverse pegRNAs at increasing inter-nick distances from 56 bp to 19.2 kb to introduce larger sequence replacements. Although we observed decreased efficiencies at kilobase inter-nick distances, some pegRNA pairs remained considerably efficient up to inter-nick distances of 9.6 kb at IL2RG and 4.9 kb at TRAC, confirming that PA supports kilobase-scale sequence replacement (Fig. 2d).
To test the PA donor size limit, we targeted gene-sized donors of 3.1 kb to 12.1 kb encoding puromycin resistance (PuroR) and PGK1-eGFP transgenes to the AAVS1 locus (Extended Data Fig. 4a). The percentage of eGFP+ cells was inversely proportional to donor size—we achieved an average of 28.1% eGFP+ cells with a 9.1 kb donor and 7.4% eGFP+ cells with a 12.1 kb donor, and readily enriched puromycin-resistant cells (Fig. 2e). We confirmed targeted integration and expression of the AAVS1-Exon1-PuroR chimeric mRNA transcript under the endogenous promoter using quantitative PCR with reverse transcription (Extended Data Fig. 4b, c). We amplified the integration junctions and further confirmed precise site-specific integration by Sanger sequencing (Extended Data Fig. 4d). We then performed Donor-seq and observed an average of 93.4–99.3% precise integration junction reads after targeted transgene integration at IL2RG, TRAC and AAVS1, which increased to more than 99.5% with NHEJ and MMEJ inhibition (Extended Data Fig. 5a). Mapping genome-wide integrations revealed high genome-wide specificity using all three pairs of pegRNAs (Extended Data Fig. 5b), consistent with the high fidelity of prime editing16,17,18,19,20. No guide-dependent off-targets were observed with IL2RG pegRNAs, one guide-dependent off-target accounting for up to 1.145% of unique molecular identifier (UMI) counts was detected with TRAC, and two off-targets accounting for up to 0.734% and 0.002% of UMI counts were detected with AAVS1 (Extended Data Fig. 5c). Together, PA provides a versatile platform for large-scale sequence replacement and kilobase-scale donor integration with high genome-wide specificity.
Multi-fragment assembly
A key benefit of DNA assembly cloning methods36,37,38,39 is their capacity to assemble multiple fragments together. The ability to split large dsDNA donors into shorter fragments could facilitate modular delivery when constrained by DNA synthesis. We first split the promoter and the reporter transgene from the PGK1-eGFP donor into two fragments with or without 3′-overhangs (Fig. 2f). Inhibition of NHEJ and MMEJ markedly increased dsDNA integration from an average of 10.5–68.5% eGFP+ cells, while abrogating random donor integration in donor-only controls from an average of 1.0% eGFP+ cells to background levels (less than 0.1%) (Fig. 2g). Notably, splitting the 3′-odsDNA or dsDNA donor into two fragments had little to no effect on PA efficiency (Fig. 2g). Using Southern blotting, we confirmed targeted genomic integration of the split donors at AAVS1 using probes against both AAVS1 and donor eGFP alleles (Extended Data Fig. 6a,b). We performed in–out and in–in junction amplicon sequencing and observed 90.1% precise 5′ junctions and 90.6% precise 3′ junctions using split 3′-odsDNA donors, which increased to 92.0% precise 5′ junctions and 96.9% precise 3′ junctions with NHEJ and MMEJ inhibition. We observed frequent NHEJ-mediated ligation of the blunt dsDNA donors producing duplicated overlap sequence that could be abrogated using split 3′-odsDNA donors or NHEJ and MMEJ inhibitors (Extended Data Fig. 6c).
We then designed a larger 3.2 kb donor encoding PuroR, eBFP and eGFP reporters and delivered the transgenes together or split into 2, 3 or 4 fragments (Fig. 2h). We measured the percentage of double-positive eBFP+eGFP+ cells 21 days post-nucleofection and observed a decrease from an average of 57.7% to 18.2% when the 3.2 kb donor was split into 4 fragments (Fig. 2i). Using puromycin selection, we enriched an average of 96.0% of cells expressing all 3 reporter transgenes after the assembly and integration of 4 dsDNA fragments (Fig. 2i). We amplified the integration and donor junctions and confirmed precise four-fragment assembly and site-specific integration by in–out PCR-based genotyping and Sanger sequencing (Extended Data Fig. 6d). Finally, we also achieved targeted assembly and integration of four ssDNA donors at ATP1A1 under selective ouabain pressure46 (Supplementary Figs. 3 and 4 and Supplementary Discussion). Thus, reminiscent of Gibson assembly cloning36, PA supports targeted multi-fragment assembly in human cells.
Benchmarking PA
Multiple nuclease-based platforms are available to integrate transgenes, such as HDR, homology-independent targeted integration (HITI) and microhomology-mediated targeted integration (MMTI) such as PITCh (precise integration into target chromosome)49. To benchmark PA, we adapted our strategies to achieve targeted integration with dsDNA donors carrying homology arms for HDR, 3′-overhangs with microhomology for MMTI, and without homology arms for HITI (Fig. 3a). We also adapted our donors for prime editing-assisted site-specific integrase gene editing (PASSIGE) using PE7 (ref. 45) and the recently described recombinase eeBxb1 (ref. 23) (Fig. 3a). In parallel, we tested the effect of using nuclease PE7 (PE7n) to achieve targeted transgene integration, a method we refer to as nuclease PA. At AAVS1, PA outperformed all other gene targeting platforms in HEK293T and K562 cells, as determined by flow cytometry and ddPCR (Fig. 3b and Extended Data Fig. 7). Although slightly higher integration efficiency was observed with PASSIGE at IL2RG, PA outperformed all nuclease-based strategies (Fig. 3b and Extended Data Fig. 7). Although low efficiency (1–2% eGFP+ cells) was observed under basal conditions at the TRAC locus in Jurkat cells, PA outperformed PASSIGE by 10.2-fold, with an average of 17.4% eGFP+ cells in the presence of NHEJ and MMEJ inhibitors (Fig. 3c and Extended Data Fig. 7), offering opportunities for more challenging settings.
a, Schematic representation of site-specific transgene integration using PA, PASSIGE, HITI, HDR or MMTI. Stabilizing phosphorothioate bonds are illustrated with orange stars. b, Percentage of eGFP+ cells as determined by flow cytometry. HEK293T or K562 cells were electroporated with the indicated vector and donor targeting PGK1-eGFP to AAVS1 or eGFP to IL2RG, and cultured for 10 to 21 days to eliminate background signal from non-integrated episomal donors. The percentage of eGFP+ cells was quantified 10 and 21 days post-nucleofection for IL2RG and AAVS1, respectively. Data are mean ± s.e.m. from n = 3 independent biological replicates. c, Percentage of eGFP+ cells as determined by flow cytometry. Jurkat cells were electroporated with the indicated vector and donor targeting eGFP to TRAC, and cultured for 3 days in the presence or absence of AZD7648 and PolQi1. The percentage of eGFP+ cells was quantified 10 days post-nucleofection. Data are mean ± s.e.m. from n = 3 independent biological replicates. d, Percentage of precise junction reads from samples in b as determined by Donor-seq. Data are mean ± s.e.m. from n = 3 independent biological replicates. Nuc-PA, nuclease PA. e, Percentage of deduplicated UMI reads for each guide-dependent off-target identified by Donor-seq from samples in b. n = 3 independent biological replicates. f, Genome-wide Donor-seq integration profile from samples in b. Each plot represents the total number of deduplicated UMI reads from n = 3 independent biological replicates.
Source data
Alongside the on-target ddPCR assays, we designed drop-off assays to broadly quantify the percentage of wild-type alleles disrupted by small indels, flap integration with or without indels, incomplete PASSIGE, and larger deletions and rearrangements (Extended Data Fig. 7). Using complementary ddPCR assays, we observed high levels of wild-type allele disruption with all platforms (Extended Data Fig. 7). The highest levels of wild-type allele drop-off were observed with nuclease PA, which generates two DNA DSBs (Extended Data Fig. 7). We note that standard PA also generated high levels of disruption with an average of 70.4% and 30.0% allele drop-off at AAVS1 and IL2RG, respectively. At the TRAC locus, the percentage of drop-off allele decreased from an average of 50.7% to 16.3% in the presence of NHEJ and MMEJ inhibitors (Extended Data Fig. 7).
We reasoned that PA could decrease off-target donor integration compared to nuclease-based strategies. We performed Donor-seq with PA, nuclease PA, HITI and MMTI. We first analysed the integration junctions and observed an average of 97.7% and 98.9% precise integration reads with PA at AAVS1 and IL2RG, respectively (Fig. 3d). We observed a marked decrease in the percentage of precise junction reads with all three nuclease-based strategies, which generate knock-in alleles with accompanied indels (Fig. 3d). Of note, we identified multiple guide-dependent off-targets occurring at substantially higher frequencies (for instance, up to 149- and 199-fold for IL2RG off-targets 1 and 2, respectively) when using nuclease-based platforms, contrasting with the high specificity observed with PA (Fig. 3e,f). Although for longer donors, donor availability might be limiting PA efficiency and purity at the target site, our results suggest that PA offers a high genome-wide specificity profile.
Transgene integration with ssDNA donors
We then tested whether PA could support site-specific integration with long ssDNA donors and successfully targeted eGFP to the IL2RG and AAVS1 loci (Extended Data Fig. 8 and Supplementary Discussion). Using an asymmetric design with a long ssDNA donor and a short synthetic ssDNA sharing an overlapping region of 50 bp (v3), we achieved an average of 3.5% eGFP+ cells with ssDNA, a 1.6-fold decrease over the corresponding 3′-odsDNA donor (Fig. 4a and Extended Data Fig. 8h). The capacity of PA to accommodate long ssDNA donors with short overlapping regions opens opportunities for targeted therapeutic transgene integration in primary haematopoietic cells that are sensitive to dsDNA30,31. We have previously achieved highly efficient prime and twin prime editing in CD34+ HSPCs by modulating nucleotide metabolism50. Using this protocol50, we electroporated CD34+ HSPCs from three different healthy donors with PE7 mRNA and various concentrations of pegRNAs and ssDNA donors to target an eGFP reporter to the AAVS1 and IL2RG loci. With 100 pmol of each pegRNA and 4 pmol of each ssDNA donor, we observed 60.8–82.9% viability 24 h post-nucleofection, and an average of 1.5% and 1.3% PA allele at the AAVS1 and IL2RG loci, respectively (Fig. 4b). These results demonstrate that in cellulo DNA assembly is applicable in therapeutically relevant haematopoietic cells.
a, Schematic representation of targeted transgene integration at the IL2RG locus using the asymmetric ssDNA donor design and pegRNAs encoding for 32-nt flaps. Stabilizing phosphorothioate bonds are illustrated with orange stars. b, Percentage of PA allele as determined by ddPCR and cell viability as determined by trypan blue staining and manual cell counting. In total, 2.5 × 105 CD34+ HSPCs were electroporated with PE7 mRNA and the indicated concentration of synthetic pegRNAs and ssDNA donors. Genomic DNA was collected 3 days post-nucleofection. Data are mean ± s.e.m. from n = 3 independent biological replicates. Circle, donor 1; diamond, donor 2; square, donor 3. c, Schematic representation of targeted deletions and the integration of a puromycin resistance (PuroR) cassette at chromosome 7 via long-range PA. d, PCR-based genotyping of the expected PuroR transgene integration and megabase deletion allele after PA and puromycin selection with the primers illustrated in orange in c. K562 cells were electroporated with PA vectors and ssDNA donors targeting chromosome 7, and puromycin selection was performed 3 days post-nucleofection. Representative gel from one out of three independent biological replicates. Ctrl, negative control. e, Percentage of PA allele and TMEM248 copy loss as determined by ddPCR, as described in d. Data are mean ± s.e.m. from n = 3 independent biological replicates. f, Same as in e, for megabase deletion at chromosome 7 using PA or PRIME-Del (P-Del) with or without PuroR integration, respectively. Data are mean ± s.e.m. from n = 3 independent biological replicates. g, Genotype of single cell-derived K562 clones after the installation of a 1 Mb deletion at chromosome 7 using PA or HDR as determined by long-read nanopore sequencing. n = 48 single cell-derived clones for each condition.
Source data
Megabase-scale rearrangement
Paired pegRNA approaches, such as twin prime editing and PRIME-Del, enable deletions smaller than 10 kb20,21,41. We reasoned that PA could excise larger genomic sequences while installing a puromycin selection marker. Using this approach, we achieved kilobase- and megabase-scale deletions on chromosomes 7 and 19 using a 3′-odsDNA donor (Supplementary Figs. 5 and 6 and Supplementary Discussion). Consistent with previous studies reporting low frequencies of off-target dsDNA donor integrations48,51, we observed puromycin-resistant cells in our donor-only controls. We reasoned that ssDNA donors could decrease these off-target integration events and facilitate the enrichment of cells with the intended megabase-scale deletions. We also hypothesized that nuclease PA, in a process reminiscent of PE-Cas9-based deletion and repair (PEDAR)40, could facilitate large-scale chromosomal rearrangements. We electroporated cells with PE7 or PE7n vectors and two ssDNA donors, and successfully introduced a megabase-scale deletion on chromosome 7 with few to no puromycin-resistant cells in our donor-only controls (Fig. 4c–e). Of note, we observed an average of 17.5% and 32.2% TMEM248 allele loss, a gene present in the intended 1 Mb deletion, in bulk populations of resistant cells after PA and nuclease PA, respectively (Fig. 4e). We readily detected the 1 Mb deletion allele by long-range PCR and confirmed the expected allele by long-read nanopore sequencing (Fig. 4d). Using nuclease PA, we also achieved targeted megabase-scale inversion as well as chromosome arm deletion and translocation (Extended Data Fig. 9, Supplementary Discussion and Supplementary Fig. 7).
To benchmark against other platforms, we adapted our pegRNA designs and performed megabase deletion with PA and PRIME-Del41, and observed an average basal efficiency of 0.013% and 2.663%, respectively (Fig. 4f). Although lower basal efficiency was observed with PA, PuroR integration and selection enabled an average of 22.1% TMEM248 allele loss (Fig. 4f). We then isolated puromycin-resistant single cell-derived K562 clones using PA with 3′-odsDNA or ssDNA donors, or HDR with a plasmid donor. We observed an increase in the percentage of single cell-derived clones with the intended megabase deletion from 39.6% (19 out of 48 clones) to 100% (48 out of 48 clones) when using ssDNA donors (Fig. 4g). PA compared favourably with HDR, which generated 75% (36 out of 48 clones) of clones with the intended deletion, and 6.3% (3 out of 48 clones) of clones with imprecise megabase deletion (Fig. 4g). PA and nuclease PA thus compare favourably to other platforms for installing large-scale genomic rearrangements, offering a versatile tool for functional genomics applications in human cells.
PA in non-dividing cells
The dependence of HDR on cell cycle progression limits its applications in primary cells, many of which are quiescent or post-mitotic. We designed an assay to assess whether PA could occur in cells arrested in G1 using the CDK4/6 inhibitor palbociclib (PD-0332991, hereafter PD). We treated K562 cells for 24 h with 5 µM PD, the lowest dose that enabled G1 arrest in at least 90% of cells with no effect on viability, and observed a marked decrease in cells progressing through S and G2/M (Fig. 5a–c). We note that a fraction of K562 cells can progress through S and G2/M after more than 24 h of PD treatment (Fig. 5c), possibly owing to fractional resistance52. Treatment with 5 µM PD 24 h before nucleofection and/or 72 h after nucleofection abrogated HDR at AAVS1 from an average of 17% to below the level of detection by Sanger sequencing (Supplementary Fig. 8). We then treated K562 cells for 24 h with 5 µM PD, electroporated cells, and enabled cell cycle progression (without PD) or kept cells arrested in G1 (5 µM PD) for 72 h before genotyping (Fig. 5a). We observed minimal expansion and little to no effect on eGFP integration in PD-treated cells using either PE7 or PE7n (Fig. 5d), suggesting that PA can occur in G1-arrested cells with or without nuclease-driven DSBs.
a, Timeline for the cell cycle experiment. b, Representative flow cytometry plot for cell cycle analysis. K562 cells were cultured for 24 h in the presence or absence of PD. Representative image is from one of four independent biological replicates. c, Cell cycle progression as determined by flow cytometry. K562 cells treated for 24 h with PD were mock electroporated and cultured in the presence or absence of PD for 4 days. Data are mean ± s.e.m. from n = 3 independent biological replicates. d, Same as in c, with K562 cells electroporated with PA vectors and 3′-odsDNA eGFP donors targeting AAVS1 or IL2RG. Cell cycle analyses, measurement of fold expansion, and ddPCR-based genotyping were performed 3 days post-nucleofection. Data are mean ± s.e.m. from n = 3 independent biological replicates. e, Effect of cell cycle arrest on HDR and PA. K562 cells were cultured in the presence or absence of PD before and/or after electroporation. Cell cycle analysis, measurement of fold expansion and genomic DNA purification were performed 3 days post-nucleofection. The percentage of edited alleles was determined by amplicon sequencing. Data are mean ± s.e.m. from n = 3 independent biological replicates. f, Flow cytometry analysis of activated and G0/G1 resting CD3+ T cell size and proliferation using CellTrace Violet. Representative flow plots from one out of four independent biological replicates. g, Percentage of PA allele as determined by ddPCR. A total of 1.5 × 106 activated or resting CD3+ T cells were electroporated with PE7 mRNA, pegRNAs and ssDNA donors targeting TRAC (see Extended Data Fig. 10c). Genomic DNA was collected 3 days post-nucleofection. Data are plotted as mean ± s.e.m. from n = 4 biological replicates. Circle, donor 1; diamond, donor 2; square, donor 3; triangle, donor 4.
Source data
We then installed dyskeratosis congenita-associated mutations via HDR (TINF2T284P or TINF2R282H) or recoded the DC cluster using PA at the TINF2 locus. Treatment with 5 µM PD before and after electroporation decreased HDR efficiency by 3.7- and 3.4-fold for TINF2T284P and TINF2R282H, respectively (Fig. 5e). For PA, we observed a 1.5-fold decrease in editing efficiency in PD-treated cells, and a 1.5-fold increase in imprecise alleles, suggesting that although PA can occur in G1-arrested cells, some DNA repair factors involved in precise ssDNA assembly and integration might be limiting in G1. Notably, in actively dividing cells, we observed 2.1- and 22.5-fold increases in precise:imprecise editing ratio with PA compared with HDR for TINF2T284P and TINF2R282H, respectively (Fig. 5e). In cells treated with PD before and after electroporation, we observed 3.5- and 31.6-fold increases in precise:imprecise editing ratio with PA compared with HDR for TINF2T284P and TINF2R282H, respectively (Fig. 5e). Together, our results suggest that PA outperforms HDR in non-proliferating cells.
To substantiate our findings, we adapted previously reported culture conditions53 to maintain primary CD3+ T cells viable in a resting G0/G1 state using 1 ng ml−1 IL-7 and 1 ng ml−1 IL-15 and confirmed that T cells remained in G0/G1 3 days post-nucleofection in resting culture conditions (Extended Data Fig. 10a,b). In parallel, we stimulated CD3+ T cells with a CD3/CD28/CD2 activator for 24 h in the presence of 300 U ml−1 IL-2 and observed a marked increase in cell size (Fig. 5f) and cell cycle progression to G1, S and G2/M (Extended Data Fig. 10b). Using a CellTrace Violet proliferation assay, activated T cells displayed multiple rounds of cell division 3 days post-nucleofection, whereas resting T cells remained non-dividing (Fig. 5f). We then electroporated T cells with PE7 mRNA, synthetic pegRNAs and ssDNA donors targeting TRAC (Extended Data Fig. 10c) and observed an average of 2.2% and 3.3% PA alleles in activated and resting G0/G1 CD3+ T cells, respectively (Fig. 5g and Extended Data Fig. 10d). Collectively, our results confirm that PA can occur in non-dividing cells.
Discussion
The ability to assemble DNA sequences in vitro36 emerged as a new paradigm for molecular cloning. By leveraging CRISPR-targeted 3′ flap synthesis and endogenous cellular DNA repair pathways, we demonstrate that this modality can extend to human cells. Whereas molecular cloning via in vitro or in cellulo DNA assembly relies on exonuclease degradation to generate complementary 3′ or 5′ ssDNA overhangs36,39, PA depends on DNA synthesis, a programmable framework for precision genome editing. Notably, PA can accommodate ssDNA donors, expanding the versatility of the approach for therapeutic genome editing and functional genomics. We envision that the ability to directly assemble and integrate multiple fragments in human cells could facilitate high-throughput library screening when constrained by DNA synthesis.
PA can be associated with precise targeted integration as well as some imprecise outcomes, mainly flap integration with deletion events. We observed higher editing efficiencies and purity with inhibition of competing end joining repair pathways. Although the optimal nucleofection mix concentrations used for short synthetic ssDNA ranged from 800–1,600 nM, only 100–200 nM of long ssDNA donor or 13–159 nM of 3′-odsDNA or dsDNA (see Supplementary Table 4) could be used without overt toxicity. Our observations suggest that the local concentration of donors available for homology-directed annealing to the 3′ PA flaps may be limiting when delivering larger DNA donors. A promising future direction may be recruitment of DNA donors to the PA site to facilitate DNA assembly without DNA repair inhibitors.
The installation of large-scale structural genomic modifications provides opportunities to haploidize specific genomic loci or model diseases. To our knowledge, this is the first report of megabase-scale deletions achieved in human cells using a prime editor without requiring nuclease-driven DSBs. Coupling PA with a nuclease prime editor also enabled megabase inversion, chromosome arm deletion and translocation, expanding the breadth of applications for functional genomics. Although we observed benefits with nuclease PA for chromosomal-scale rearrangements, we recommend nickase-based PA for targeted genomic integration to maximize product purity and genome-wide specificity.
We successfully programmed targeted integrations of transgenes to three loci in primary CD3+ T cells and CD34+ HSPCs from different healthy donors, demonstrating a novel method to introduce medium to large-sized modifications in therapeutically relevant cells. Owing to their lower immunogenicity30,31, ssDNA donors could be used to recode exonic mutation hotspots or target a functional transgene copy to complement a defective gene at its endogenous locus, providing universal strategies to remedy sundry pathogenic mutations. Given that ex vivo culture with cytokines and cell cycle progression, which are essential for HDR7,54, correlate negatively with haematopoietic stem cell engraftment potential55,56 and increase genotoxicity12,13, PA may provide an alternative method for medium to large-sized therapeutic edits. As PA can accommodate synthetic ssDNA donors, unique molecular indexes could also be introduced for long-term clonal tracking. This approach could overcome AAV-induced toxicity32,33 as well as imprecise concatemeric AAV vector integration at nuclease DSBs57,58. We envision that PA should also enable site-specific antigen receptor transgene integration at endogenous loci for tailored immunity2,3. Multiplexed PA coupled with standard prime editing could further improve the functionality of immune cells by modifying key targets59,60 without requiring multiple nuclease-induced DSBs. Furthermore, PA could permit targeted integration in post-mitotic cells.
An optimal method for therapeutic targeted integration of gene-sized DNA sequences would require only a single protein effector, avoid the cytotoxicity of dsDNA donors, bypass the genotoxicity and imprecision of nuclease-driven DSBs, and maintain activity in non-dividing cells. To our knowledge, PA is the only currently described method that meets these criteria. Together, PA enables in cellulo DNA assembly for precision human genome engineering.
Methods
Cell culture and nucleofection
K562 (CCL-243) and Jurkat (TIB-152) cells were obtained from American Type Culture Collection (ATCC) and cultured at 37 °C under 5% CO2 in RPMI media (ThermoFisher Scientific, 11875093) supplemented with 10% FBS, and 1% Penicillin/Streptomycin (ThermoFisher Scientific, 15140122). HEK293T (CRL-1573) cells were obtained from the ATCC and cultured at 37 °C under 5% CO2 in DMEM media (ThermoFisher Scientific, 11995065) supplemented with 10% FBS, and 1% penicillin/streptomycin. Cell lines were authenticated by the supplier and tested negative for mycoplasma.
For standard K562 nucleofections, 2 × 105 cells were electroporated with 750 ng pCMV-PE7 (ref. 45) (Addgene #214812), 250 ng of each standard pegRNA16 vector (derived from Addgene #132777), and the indicated concentration of PA donor with an Amaxa 4D-nucleofector (Lonza) using the SF cell line nucleofection kit (Lonza, V4XC-2032) (pulse FF-120). For each nucleofection, cells were resuspended in 20 µl of nucleofection buffer, the indicated amount of DNA was added while maintaining a final total volume of less than 24 µl, and nucleofection mixes were transferred to the 16-well strip. For short synthetic ssDNA donors, the optimal concentrations were 800 nM to 1,600 nM of each ssDNA donor. These concentrations could not be achieved with long ssDNA donors due to toxicity, and 100 nM to 200 nM were used. For standard 3′-odsDNA and dsDNA donors, 78 nM to 159 nM were used. Finally, for larger 3′-odsDNA and dsDNA donors ranging from 3.1 to 12.1 kb, 13 nM to 53 nM were used. Where indicated, the F + E scaffold modifications61 were included in the pegRNAs. For initial experiments (Fig. 1 and Extended Data Figs. 1 and 2), K562 cells were electroporated with 750 ng pCMV-PEmax44 (Addgene #174820), 250 ng of each tevopreq1-epegRNA62 vector (derived from Addgene #174038) harbouring the (F + E) scaffold modifications61, and the indicated concentration of PA donor. For nuclease PA experiments, the HNH domain of PE7 (Addgene #214812) was restored to generate pCMV-PE7nuclease via Gibson assembly. For 53BP1 inhibitor overexpression under the EF1α promoter, the hMLH1dn cassette of pEF1a-hMLH1dn (Addgene #174824) was replaced with the i53 (ref. 63) coding sequence. For HDR experiments, K562 cells were electroporated with 750 ng pX330-U6-Chimeric_BB-CBh-hSpCas9 (ref. 64) (Addgene #42230) expressing the sgRNA of interest and 16 pmol ssDNA donor. All high-quality plasmids used for electroporation were purified using the EZNA FastFilter Plasmid DNA Midi Kit (Omega Bio-tek, D6905-04) and DNA concentration and purity was assessed by nanodrop. The composition of all nucleofection mixes is provided in the Supplementary Information.
For Jurkat nucleofections, 1 × 106 cells were electroporated with 500 ng pCMV-PE7, 250 ng of each standard pegRNA, and the indicated concentration of PA donor using the SE cell line nucleofection kit (Lonza, V4XC-1302) (Pulse CL-120). For HEK293T nucleofections, 2 × 105 cells were electroporated with 750 ng pCMV-PE7, 375 ng of each standard pegRNA, and the indicated concentration of PA donor using the SF cell line nucleofection kit (pulse CM-130). For benchmarking experiments, 750 ng of pCMV-PE7, 750 ng eeBxb1 (ref. 23) (Addgene #222339), or 750 ng pX330 vector was used for K562 and HEK293T cells, and 500 ng of each editor expression vector was used for Jurkat cells. The total concentration of DNA was normalized between PA and PASSIGE while maintaining the original vector ratio for the latter23. The composition of all nucleofection mixes is provided in the Supplementary Information.
StemSelect PD-0332991 (Sigma, 5304870001) was dissolved at 10 mM in water and stored at −80 °C. Where indicated, K562 cells were treated with 5 µM PD-0332991. Ouabain octahydrate (Sigma, O3125-250GM) was dissolved at 5 mg ml−1 in water, and working dilutions were prepared in water and stored at −20 °C. Where indicated, ouabain selection was performed with 0.5 µM 3 days post-nucleofection until all non-resistant cells were eliminated. Puromycin (Sigma, P8833-25MG) was dissolved at 1 mg ml−1 in water and stored at −20 °C. Where indicated, puromycin selection was performed with 1 µg ml−1 3 days post-nucleofection until all non-resistant cells were eliminated. AZD7648 (MedChemExpress, HY-111783) and PolQi1 (MedChemExpress, HY-159078) were dissolved at 10 mM in DMSO, and working dilutions were prepared in water and stored at −80 °C. Where indicated, K562 and HEK293T cells were treated during 3 days post-nucleofection with 1 µM AZD7648 and 1.5 µM PolQi1. Jurkat cells were treated during 3 days post-nucleofection with 0.5 µM AZD7648 and 0.5 µM PolQi1.
Primary CD34+ HSPC culture and nucleofection
Cryopreserved human CD34+ HSPCs from mobilized peripheral blood of deidentified healthy donors were obtained from the Fred Hutchinson Cancer Research Center (Seattle, Washington) and their use was determined as exempt from human subjects research requirements by Boston Children’s Hospital Institutional Review Board. CD34+ HSPCs were cultured in X-Vivo-15 media (Lonza, 04-418Q) supplemented with 100 ng ml−1 human Stem Cell Growth Factor (SCF) (R&D Systems, 255-SC-010), 100 ng ml−1 human thrombopoietin (TPO) (Peprotech, 300-18), and 100 ng ml−1 recombinant human FMS-like Tyrosine Kinase 3 Ligand (Flt3-L) (Peprotech, 300-19). CD34+ HSPCs were thawed and cultured for 24 h in the presence of cytokines, and electroporated using the P3 Primary Cell X kit S (Lonza, V4XP-3032) according to the manufacturer’s recommendations. Cells (2.5 × 105) were electroporated with 2,000 ng PE7 mRNA45, an equimolar ratio of simian immunodeficiency virus (SIV) Vpx mRNA50, and the indicated concentration of each pegRNA and ssDNA donors using pulse code DS-130. Following electroporation, 80 µl of media supplemented with cytokines was added to each well and cells were incubated for 10 min prior to transfer to the culture plate. Cells were cultured in a 48-well plate in a final volume of 500 µl of media supplemented with 50 µM of each deoxynucleoside50. Cell viability was assessed 24 h post-nucleofection via Trypan Blue staining and manual counting using a haemocytometer, and genomic DNA was purified 3 days post-nucleofection. Deoxynucleosides (dA, Sigma-Aldrich, D8668; dG, Sigma-Aldrich, D0901; dC, Sigma-Aldrich, D0776; and dT, Sigma-Aldrich, T1895) were resuspended in water at 12.5 mM each, filter-sterilized, and stored at −20 °C.
Primary CD3+ T cell culture and nucleofection
Human CD3+ T cells were isolated from leukocyte reduction system (LRS) cones of deidentified healthy donors from the Blood Donor Center at Boston Children’s Hospital and their use was determined as exempt from human subjects research requirements by Boston Children’s Hospital Institutional Review Board. Peripheral blood mononuclear cells (PBMCs) were collected via density gradient centrifugation by layering the blood diluted with PBS on Ficoll-Paque (Cytiva, 17144002) using SepMate-50 tubes (StemCell Technologies, 85450). PBMCs were aspirated and washed with cold PBS. Bulk T cells were isolated by magnetic labelling with CD3 MicroBeads (Miltenyi Biotec, 130-050-101) and separation via LS columns (Miltenyi Biotec, 130-042-401) using a manual MACS separator according to the manufacturer’s recommendations. CD3+ T cells were either used fresh or cryopreserved.
Primary CD3+ T cells were cultured at a density of 106 cells per ml in ImmunoCult T Cell Expansion Medium (StemCell Technologies, 10981), supplemented with 1% penicillin/streptomycin at 37 °C with 5% CO2. CD3+ T cells were activated with ImmunoCult Human CD3/CD28/CD2 T Cell Activator (25 µl per million cells) (StemCell Technologies, 10990) for 24 h and cultured with 300 U ml−1 IL-2 (StemCell Technologies, 78036.1). Resting T cells were kept in culture with 1 ng ml−1 IL-7 (Miltenyi Biotec, 130-095-367) and 1 ng ml−1 IL-15 (Miltenyi Biotec, 130-095-760). Primary T cells were electroporated using the P3 Primary Cell X kit S (Lonza, V4XP-3032) according to the manufacturer’s recommendations. Cells (1.5 × 106) were electroporated with 2,000 ng PE7 mRNA45, an equimolar ratio of SIV Vpx mRNA50, 200 pmol of each pegRNA, and 4 pmol of each ssDNA donor using pulse code DS-137. Following electroporation, 80 µl of media supplemented with cytokines was added to each well and cells were incubated for 10 min prior to transfer to the culture plate. Cells were cultured in a 48-well plate in a final volume of 500 µl of media supplemented with 50 µM of each deoxynucleoside50.
Preparation of donors for PA and other platforms
Short ssDNA donors were synthesized as ultramers (IDT) at a 4 nmol scale with 5′ phosphorylation. To generate dsDNA donors, ssDNA ultramers were mixed in 50 mM NaCl, 10 mM Tris-HCl (pH 8.0), 1 mM EDTA, and annealed by heating the solution to 95 °C for 10 min, followed by gradual cooling on a thermocycler. The ssDNA and annealed dsDNA donors were then diluted in IDTE buffer (IDT) and stored at −20 °C. dsDNA donors with 3′ overhangs were generated via exonuclease digestion, as previously described48. In brief, donors were amplified from plasmids using Kapa-HiFi polymerase (Roche, 07958897001) with primers harbouring 5′ phosphorylation, and the expected overhang sequence followed by five consecutive phosphorothioate linkages to block lambda exonuclease from digesting the donor further. PCR products were purified using SPRIselect beads (Beckman Coulter, B23318) using a bead:sample ratio of 0.8:1, digested with lambda exonuclease (NEB, M0262S), and purified again using SPRIselect beads. For the 12.1 kb dsDNA donor (Fig. 2e), purification was performed with the Monarch Spin High-Capacity DNA Cleanup Kit (NEB, T1135S). Donor concentration, purity, and integrity was assessed by nanodrop and agarose gel electrophoresis. The eGFP and puromycin resistance transgenes were amplified from AAVS1_Puro_hPGK1_eGFP_Donor (Addgene #178088)46. Alternatively, the eGFP cassette was cloned in a pUC19 backbone with a splicing acceptor (SA) and a self-cleaving 2A peptide (2A) in-frame with TRAC or AAVS1. The CD19-CAR-2A-eGFP donor was amplified from MTOR-F2108L_CD19-CAR-28z-2A-eGFP_AAV6_Donor65 (Addgene #211904). Donor sequences used in this study are provided in the Supplementary Information.
For benchmarking experiments, the PA donors used for targeted eGFP integration at AAVS1, IL2RG and TRAC were cloned alongside a Bxb1 attB site in a pUC19 backbone vector, generating donors of ~5.1 kb and allowing functional eGFP integration and expression at endogenous loci. For nuclease-based approach, PA donors were adapted for nuclease-based integration using the same spacer as the ones used for AAVS1 (F1), IL2RG (F1), and TRAC (F1) pegRNAs. The HDR templates were generated as previously described51 by cloning the PA donors with ~300 bp homology arms. The donors were amplified from plasmids using Kapa-HiFi polymerase and purified with SPRI beads. For microhomology-mediated targeted integration, the donors were amplified from plasmids with primers harbouring 5′ phosphorylation and 24-bp overhang sequences followed by 5 consecutive phosphorothioate linkages. The PCR products were purified, digested with lambda exonuclease to generate 24-bp microhomology overhangs (as described for PA with 3′-odsDNA), and purified again with SPRI beads. Finally, blunt dsDNA PA donors (without exonuclease digestion and 3′ overhangs) were used for homology-independent targeted integration. All donors were purified with SPRIselect beads. Donor concentration, purity, and integrity was assessed by nanodrop and agarose gel electrophoresis.
For experiments requiring long ssDNA, donors were amplified using Kapa-HiFi polymerase with a primer harbouring a 5′ biotin modification for the DNA strand to separate, and a primer harbouring a 5′ phosphorylation for the DNA strand to isolate, as previously described34. PCR amplicons were purified with SPRIselect beads. The single strand of interest was then purified via magnetic separation using Streptavidin C1 Dynabeads (ThermoFisher Scientific, 11205D). In brief, Streptavidin C1 Dynabeads were washed two times, mixed with biotinylated PCR amplicons, and incubated at room temperature for 30 min with agitation. For magnetic separation, Dynabeads coated with biotinylated amplicons were washed twice, and the supernatant was removed and replaced with 0.125 M NaOH melt solution (prepared fresh) to denature the dsDNA. The solution was placed back on the magnet and the supernatant containing the nonbiotinylated strand was removed gently and mixed immediately with Neutralization buffer (freshly prepared by mixing 100 µl 3 M sodium acetate pH 5.2 with 4.8 ml 1× TE buffer). A second round of denaturation and elution was performed with 0.125 M NaOH melt solution using the same neutralization tube. Resulting ssDNA was purified using SPRIselect beads, eluted in IDTE buffer, and ssDNA concentration and purity was assessed by nanodrop. Alternatively, long ssDNA donors were provided by Genscript, resuspended in IDTE buffer, and stored at −20 °C.
In vitro transcription and pegRNA synthesis
The PE7 transcription template vector45 (Addgene #223022) was linearized using BbsI-HF (NEB, R3539L), and mRNA was transcribed using the HiScribe T7 high yield RNA kit (NEB, E2050S) using N1-methylpseudouridine (Trilink, N-1081) instead of uridine, and co-transcriptional capping with CleanCap AG (Trilink, N-7113). The PE7 in vitro transcription plasmid template encodes for a T7 promoter, a minimal 5′-untranslated region (UTR), a PE7 cassette45 harbouring a silent mutation disrupting a restriction site for the linearizing BbsI enzyme, a 2× HBB 3′ UTR, and a 80–90 bp poly(A) sequence. For SIV Vpx mRNA, the template was generated as previously described50. In brief, the SIV Vpx vector template50 (Addgene #216792) was amplified by PCR with a forward primer that correct a T7 promoter inactivating mutation and a reverse primer that appends a 119-nt poly(A) tail to the 3′ UTR. Following IVT, mRNAs were purified using the Monarch RNA Cleanup kit (500 µg) (NEB, T2050L) and eluted in 1× nuclease-free IDTE buffer (10 mM Tris, 0.1 mM EDTA, pH 7.5). The mRNA concentration was quantified using Qubit RNA high sensitivity (HS) kit (ThermoFisher Scientific, Q32852). Synthetic pegRNAs were provided by Integrated DNA Technologies (IDT) and resuspended at 200 pmol µl−1 in nuclease-free IDTE buffer (10 mM Tris, 0.1 mM EDTA, pH 7.5). The pegRNAs contained 2′-O-methyl modifications and phosphorothioate linkages. All pegRNA sequences and chemical modifications are provided in the Supplementary Information.
DNA sequencing
Genomic DNA was collected at the indicated time post-nucleofection using QuickExtract DNA extraction solution (Fisher Scientific, NC9904870) following manufacturer’s recommendations. For Sanger sequencing, primers were designed to amplify a 600−800 bp amplicon66,67,68. PCR amplifications were performed with 30 cycles of amplification with Phusion high-fidelity polymerase (NEB, M0531L). PCR product quality was evaluated by agarose gel electrophoresis, and purification was performed with SPRIselect beads using a bead:sample ratio of 0.8:1 before Sanger sequencing. Sequencing trace quality was manually inspected using Geneious R11 software (v11.1.5), and lower quality reactions with background noise were repeated. For prime editing at B2M and HDR at AAVS1, the percentage of precise alleles and indels were quantified using BEAT67 and TIDE66 webtools from Sanger sequence data files, respectively.
For amplicon sequencing, primers were designed to amplify 200–250 bp amplicons. PCR amplifications were performed with Phusion high-fidelity polymerase, and amplicons were purified with SPRIselect beads using a bead:sample ratio of 0.9:1 or 1:1. PCR product quality was assessed via agarose gel electrophoresis. Indexing (PCR 2) was performed with 1 µl of locus-specific PCR product using TruSeq adapters (Illumina). Following bead purification, PCR product quality was assessed by electrophoresis and TapeStation using a DS1000 High Sensitivity ScreenTape assay (Agilent, 5067-5585), and quantified with a Qubit dsDNA High Sensitivity (HS) assay kit (ThermoFisher Scientific, Q33231). Amplicons were sequenced using paired-end 150-bp reads on an Illumina MiniSeq system in-house, Illumina NovaSeq X system by Novogene, or Illumina NovaSeq X system by the Harvard Biopolymers Core Facility. The percentage of HDR and indel alleles was quantified with CRISPResso2 (ref. 69) using the HDR mode with a quantification window of 5 bp on each side of the cut site, and the percentage of indels was determined as the percentage of NHEJ reads plus imperfect HDR reads. To quantify precise PA or indels at the flap and split donor junctions from in–out and in–in amplicons, CRISPResso2 was run on NHEJ mode with the expected junction amplicon as the reference sequence and the quantification window was set as the full amplicon except the first and last 15 base pairs, and substitutions were not quantified as indels. Precise PA and indels were designated as the percentage of unmodified and modified reads, respectively. For TINF2, genomic DNA (gDNA) from a single cell-derived K562 clone recoded at exon 6 was used as a control template. For PGK1-eGFP integration at AAVS1, a plasmid encoding the expected allele was used as a control template.
For PA out–out amplicons, editing outcomes were first analysed using CRISPResso2 (v2.3.3). Paired-end reads were aligned to five reference amplicons simultaneously: wild type, precise PA, and three flap integration references representing forward flap (5′ junction), reverse flap (3′ junction), and dual flap integrations without donor integration. Each reference amplicon and its corresponding name were supplied via the -a and -an flags, respectively. Four guide RNA sequences, the primary spacers and their flap variants, were provided for cut site annotation. The –min_frequency_alleles_around_cut_to_plot 0 and –write_detailed_allele_table flags were used to retain all alleles and enable downstream analysis. The detailed allele tables generated by CRISPResso2 were subsequently parsed and reclassified into discrete editing outcome categories using a custom Python script (see Code availability). For each allele, deletion and insertion positions were extracted from the allele frequency table, and only deletions and insertions located within the window spanning both nick sites were used for classification, except where noted below. Alleles were then classified based on their alignment reference and the presence of indels. For alleles aligned to the wild-type reference, those with no indels within the window were classified as {Unedited}; those carrying an indel overlapping a ±1 bp window centred on either nick site were classified as {Indels at nick site}; remaining alleles were assigned to {Others}. For alleles aligned to the Precise PA reference, indel-free alleles were classified as {Precise PA}; alleles harbouring deletions only were classified as {Flap integration with deletion}; those with insertions only as {Flap integration with insertion}; and all others as {Others}. Alleles classified as {Others} comprised low-frequency, complex alleles that could not be unambiguously assigned to a single category. For alleles aligned to any flap integration only reference (forward, reverse, or dual), the same sub-classification scheme was applied, with indel-free alleles designated as {Flap integration without indels}. Alleles flagged as ‘AMBIGUOUS’ by CRISPResso2 were reassigned to their most likely reference by removing the AMBIGUOUS prefix prior to classification. Representative allele plots from the different allele categories are provided in Extended Data Fig. 2. We note that large deletions could be missed with amplicon sequencing due to amplification or purification biases and complementary analyses may be required to capture these events.
For long-read nanopore sequencing of kilobase- and megabase-scale deletions, primers were designed to amplify a 2–3 kb amplicon encompassing the PA junctions and the selection marker cassette. PCR amplifications were performed with Phusion high-fidelity polymerase, and amplicons were purified with SPRIselect beads using a bead:sample ratio of 0.8:1. PCR product quality was assessed via agarose gel electrophoresis. Long-read nanopore sequencing was performed by Plasmidsaurus. Uncropped scans of all gels from this study are provided in Supplementary Fig. 9. For clonal analysis, single cell-derived K562 clones were isolated via serial dilution in 96-well plates with 200 µl of media supplemented with 1 µg ml−1 puromycin. For SpCas9 nuclease-induced megabase inversion, single cell-derived clones were isolated without puromycin selection. Single cell-derived K562 clones resistant to puromycin were expanded in a final volume of 1 ml in a 24-well plate, and genomic DNA was collected using the Monarch spin gDNA extraction kit (NEB, T3010L). Nanopore libraries were prepared and sequenced in-house. PCR amplifications were performed with Phusion high-fidelity polymerase, and amplicons were purified with SPRIselect beads. PCR product quality was assessed via agarose gel electrophoresis. Indexing (PCR 2) was performed with 1 µl of locus-specific PCR product using TruSeq adapters (Illumina). Following bead purification, PCR product quality was assessed by electrophoresis, and quantified with a Qubit dsDNA High Sensitivity (HS) assay kit. The pooled library was then processed using the Ligation Sequencing Kit SQK-LSK114 (Oxford Nanopore Technologies) following the manufacturer’s recommendations. In brief, DNA repair and end-prep was performed, followed by adapter ligation for 10 min and cleanup using SPRIselect beads. The prepared library was loaded into a MinION Flow Cell (FLO-MIN114, Oxford Nanopore Technologies) following manufacturer’s recommendations. Sequencing was performed on the MinION Mk1B device (Oxford Nanopore Technologies). For long-read nanopore sequencing after four ssDNA fragment assembly, primers were designed to amplify a 5,927 bp (wild type) to 6,284 bp (Precise PA) amplicon using Kapa Long Range HotStart polymerase (Roche, 07961278001). Amplicons were purified with SPRIselect beads using a bead:sample ratio of 0.8:1, sequenced by Plasmidsaurus, and analysed using CRISPRLungo70. Raw sequencing reads were processed using CRISPRLungo (v0.1) for error filtering and alignment. To establish control datasets, simulated Nanopore sequencing data were generated using Badread71 under two conditions: a wild-type control and a Precise PA sequence control. Both controls were analysed with CRISPRLungo using two target sites as input.
Per-read mutations were extracted from the read_classification.txt output of CRISPRLungo for each control condition. Flap integration was assessed by examining mutations within a window of ±1 bp surrounding both the cleavage site and the flap end site. A read was considered to have undergone flap integration if at least one window contained a mutation in the wild-type control result while the corresponding window was mutation-free in the Precise PA control result. Indels between the two nick sites were further quantified using a window-based scoring function. Substitutions, insertions, and deletions overlapping between nick window were enumerated, and alignment identity was calculated as the number of matched bases divided by the total aligned bases. Of the two alignment results (wild-type reference and Precise PA reference), the one with higher alignment identity was selected. From the selected alignment, only mutations spanning within a 1 bp window around the cleavage site and flap site were used for mutation assessment.
Each read was subsequently assigned to one of the following categories using a custom Python script (see Code Availability). Reads in which the Precise PA reference provided the better alignment and no mutations were detected at the cleavage site, flap site, or junction regions were classified as {Precise_PA}. Among reads with confirmed flap integration, those with no accompanying indels were classified as {Flap_integration_without_indels}. The lengths of insertions and deletions at window were calculated for each read. Reads with net deletions were classified as {Flap integration_with deletion}, and reads with net insertions were classified as {Flap integration_with insertion}. For reads without flap integration, those with no mutations in any evaluated window were classified as {Unedited}, and those with only indels at the nick site were classified as {Indels_at_nick}. All remaining reads were classified as {Others}. Primers used in this study are provided in the Supplementary Information.
ddPCR
Genomic DNA was extracted and purified using the Monarch spin gDNA extraction kit (NEB, T3010L). For each ddPCR reaction, 25–50 ng of genomic DNA was used, and all conditions were performed in technical triplicates. The droplets were generated using a Bio-Rad QX200 AutoDG ddPCR system with ddPCR supermix (no dUTP) (Bio-Rad, 186-3025), and HindIII-HF was supplemented (NEB, R3104L) in each reaction. Following droplet generation, samples were amplified using the following conditions: 95 °C for 10 min, 40 cycles of 94 °C for 30 s, annealing (56–59 °C) for 60 s, and a final incubation at 98 °C for 10 min. Samples were then kept at 4 °C until analysis. Results were analysed using the QuantaSoft software (v1.7.4.0917), and the percentage of PA alleles harbouring the targeted transgene integration was determined as the ratio of PA allele relative to a genomic reference. Primers and probes used during this study are available in the Supplementary Information.
Flow cytometry and cell cycle analysis
The percentage of eGFP+ fluorescent cells was quantified using a BD LSRII flow cytometer, and 1 × 105 cells were analysed for each condition. Cells were cultured for 7 days (K562) or 10 days (Jurkat) post-nucleofection, and donor-only conditions were used as a negative control. For experiments using the PGK1 promoter, cells were cultured for 21 to 35 days to eliminate background fluorescence signal from non-integrated donor. For cell cycle analysis, cells were cultured in the presence or absence of 5 µM PD-0332991 24 h before and 72 h after nucleofection. For each nucleofection, 1 × 106 K562 cells were electroporated with an Amaxa 4D-nucleofector (Lonza) using the SF cell line nucleofection kit (pulse FF-120). The fold expansion was measured 3 days post-nucleofection by Trypan blue staining and manual counting using a haemocytometer. Cells were washed once with PBS, resuspended at 1 × 106 cells per ml in PBS supplemented with 10 µg ml−1 Hoechst 33342 (Sigma, B2261), and stained for 45 min at 37 °C in the dark, mixing every 15 min. After staining, cells were washed and resuspended in PBS, and 1 × 105 cells were analysed for each condition using a BD LSRII flow cytometer and BD FACSDiva v9.0 software.
For the CD3+ T cell proliferation assay, cells were labelled 24 h before nucleofection using the CellTrace Violet Cell Proliferation Kit (ThermoFisher Scientific, C34557) for 20 min in the dark at 37 °C, and the reaction was stopped using PBS supplemented with 2% BSA Stock Solution (Miltenyi Biotec, 130-091-376) according to the manufacturer’s instructions. Cells were washed, counted, and cultured at 1 × 106 cells per ml. T cells were stimulated or not with a CD3/CD28/CD2 activator for 24 h, electroporated (mock), and cultured for 72 h with 300 U ml−1 IL-2 (activated) or 1 ng ml−1 IL-7 and 1 ng ml−1 IL-15 (resting) before cell proliferation analysis. For cell cycle analysis, CD3+ T cells were counted, resuspended in ImmunoCult-XF T Cell Expansion Medium at 1 × 106 cells ml−1, and stained with 2 mg ml−1 of Hoechst 33342 (Millipore, B2261) for 45 min at 37 °C in the dark, mixing every 15 min. Pyronin Y (Sigma, 83200-10 G) was added to the cells to a final concentration of 5 mg ml−1 and incubated for further 45 min at 37 °C in the dark. After washing, cells were resuspended in PBS and flow cytometry was performed on a BD LSRFortesa flow cytometer and BD FACSDiva v9.0 software. Flow cytometric data visualization and analysis was performed using FlowJo (v10). Flow cytometry gating strategies used in this study are provided in Supplementary Fig. 10.
RNA extraction and quantitative real-time PCR
Total RNA was extracted from cells using the Quick-RNA Miniprep Plus Kit (Zymo Research, R1057). Complementary DNA (cDNA) was synthesized from 0.5 μg of total RNA using the iScript cDNA Synthesis Kit (Bio-Rad, 1708891). Quantitative real-time PCR was performed using 1/25 of the synthesized cDNA with SYBR Select Master Mix (Thermo Fisher Scientific, 4472908) on a QuantStudio 3 Real-Time PCR System (Thermo Fisher Scientific) using QuantStudio Design Analysis software 1.3. Relative gene expression was calculated using the \({2}^{-\Delta {C}_{{\rm{T}}}}\) method. Primer sequences are provided in the Supplementary Information.
Southern blotting
Genomic DNA was isolated using the Monarch spin gDNA extraction kit. 3-6 µg of gDNA was digested with BstEII-HF (NEB, R3162L) for 2 h at 37 °C and run on a 0.6% agarose gel followed by Southern blotting onto Hybond-N+ membrane (Amersham, RPN303B). The membrane was then UV crosslinked followed by pre-hybridization for 1 h at 42 °C in hybridization solution (DIG Easy Hyb) (Roche, 11603558001) supplemented with 100 µg ml−1 of denatured salmon sperm DNA (Invitrogen, 15632011). The pre-hybridization solution was discarded, and fresh pre-warmed hybridization solution supplemented with DIG-labelled probes at 10 ng ml−1 each was added followed by hybridization overnight at 42 °C. DIG-labelled probes were generated as previously described72. In brief, probes were synthesized by annealing the universal primer with the probe template followed by fill-in using Klenow Fragment (3′–5′ Exo-) (NEB, M0212M) and a dNTP mixture containing DIG-11 dUTP (Roche, 11573152910), blunting with T4 DNA polymerase (NEB, M0203S), and degradation of the template using lambda exonuclease (NEB, M0262S). After hybridization, the membrane was washed in 2× SSC 0.1% SDS for 5 min twice at room temperature, then washed in 0.2× SSC 0.1% SDS for 15 min twice at 50 °C. Detection was performed as described in the DIG wash and block buffer set (Roche, 11585762001). The membrane was then stripped by rinsing in water followed by washing twice in 0.2 M NaOH and 0.1% SDS for 15 min at 37 °C with constant agitation, followed by hybridization and detection as described above. DIG ladder was a mixture of DIG-labelled DNA Molecular Weight Marker III (Roche, 11669940910) and VII (Roche, 11218603910). Primer and probe sequences are provided in the Supplementary Information.
Donor-seq
GUIDE-seq and related methods8,18,73 were adapted to detect the genome-wide PA donor integration profile. Donor-seq library preparation was performed using Tn5 transposase assembled with pre-annealed adapters incubated at room temperature for one hour. Genomic DNA (100 ng) was tagmented using 1 µl of the transposome at 55 °C for 7 min. The reaction was stopped by adding 0.2% SDS (ThermoFisher Scientific, 15553027), and tagmented DNA was used for library amplification. Primary PCR amplification was performed with Platinum SuperFi PCR Master Mix (ThermoFisher Scientific, 12358050). Nested PCR amplification was performed using 1 µl of primary PCR product. Indexing (PCR 3) was performed using i5 primer (5′-AATGATACGGCGACCACCGAGATC-3′) and Illumina i7 TruSeq indexing primers. Following SPRI bead purification, PCR product quality was assessed by electrophoresis and TapeStation using a DS1000 High Sensitivity ScreenTape assay, and quantified with a Qubit dsDNA High Sensitivity (HS) assay kit. Donor-seq libraries were sequenced on an Illumina NovaSeq X system by the Harvard Biopolymers Core Facility.
For junction purity analysis, FASTQ files containing donor integration reads with UMIs encoded in the read headers were first processed using cutadapt to remove Illumina and Tn5 adaptor sequences (options:–overlap 10–error-rate 0.10 -q 20 -m 20; adaptor sequences: AGATCGGAAGAGCGTCGTGTAGGGAAAGAGTGT and CTGTCTCTTATACACATCT). Adaptor-trimmed reads were subsequently analysed using a custom Python pipeline (using edlib) to identify and trim donor insert sequences, allowing a maximum error rate of 0.1 with a minimum overlap of 10 bp, while enforcing exact matching of the terminal 5 bp of the insert to ensure precise junction definition. The extent of donor deletion at the insertion junction was quantified and appended to the read header. Trimmed reads were then aligned to the GRCh38 reference genome using Bowtie2 with the–very-sensitive-local option to enable partial and soft-clipped alignments. Using custom downstream scripts, alignment positions were compared with the expected cleavage site to calculate genome-side insertion and deletion events at the junction, followed by UMI-based deduplication to collapse PCR duplicates. All pipelines are publicly available at https://github.com/GuehoLab/DonorJunctionAnalysis.
For off-target nomination, Donor-seq data were analysed using a modified version of the Geneth’off GUIDE-seq pipeline74. First, we filtered for reads that contained the Donor-seq tag, allowing for errors. Next, we trimmed the Donor-seq tags and Illumina adapters from the ends of the reads. We then filtered the trimmed, paired-end reads for length, retaining pairs whose R1 read or R2 read exceeded 25 bp. Next, we mapped the reads to the GRCh38 reference genome via bowtie2, using the following parameters: -I 100, -X 1500,–dovetail,–no-mixed,–no-discordant. We retained reads that mapped to at least one region of the genome and whose primary alignment had a mapq score 20 or greater. We identified the integration site as the first base of the trimmed R2 read, as this base is immediately adjacent to the integrated Donor-seq tag. Finally, we deduplicated reads with the same UMI and integration site. Consistent with the original Geneth’off GUIDE-seq pipeline, we attempted to ‘rescue’ the R2 component of reads filtered out due to insufficient length or failure to align to the reference genome. In brief, we recovered the leftover reads, extracted the R2 component of the reads, filtered the reads on length, and aligned the single-end reads to the reference genome via bowtie, using the–no-unal parameter. We then retained reads with a sufficiently high mapq score, identified the integration site as the first base of the read, and deduplicated reads according to UMI and integration site. Overall, this procedure yielded a data frame whose rows corresponded to distinct bases and whose columns recorded the chromosome, coordinate, and number of UMIs of a given base. For guide-dependent off-target nomination, guides with more than four combined mismatches or nucleotide bulges in the spacer and NGG protospacer adjacent motif (PAM) sequence were excluded, and up to two mismatches in the 8-bp PAM-proximal seed region were tolerated. The off-target coordinates are provided in the Supplementary Information. The pipelines are publicly available at https://github.com/timothy-barry/genethoff-nf/tree/nature-revision.
Reporting summary
Further information on research design is available in the Nature Portfolio Reporting Summary linked to this article.
Data availability
Next generation sequencing data generated during this study are publicly accessible from the NCBI Sequence Read Archive database under accession PRJNA1418097. The human genome assembly GRCh38.p14 (NCBI RefSeq GCF_000001405.40) was used during this study. Source data are provided with this paper.
Code availability
All Donor-seq pipelines developed during this study are publicly available on Github at https://github.com/GuehoLab/DonorJunctionAnalysis and https://github.com/timothy-barry/genethoff-nf/tree/nature-revision. The custom CRISPResso2 and CRISPRLungo scripts for PA allele classification are available at https://github.com/Gue-ho/PA_Analysis.
References
Levesque, S. & Bauer, D. E. CRISPR-based therapeutic genome editing for inherited blood disorders. Nat. Rev. Drug Discov. 24, 907–925 (2025).
Article CAS PubMed PubMed Central Google Scholar
Ellis, G. I., Sheppard, N. C. & Riley, J. L. Genetic engineering of T cells for immunotherapy. Nat. Rev. Genet. 16, 103–107 (2021).
Google Scholar
Eyquem, J. et al. Targeting a CAR to the TRAC locus with CRISPR/Cas9 enhances tumour rejection. Nature 543, 113–117 (2017).
Article ADS CAS PubMed PubMed Central Google Scholar
Canver, M. C. et al. BCL11A enhancer dissection by Cas9-mediated in situ saturating mutagenesis. Nature 527, 192–197 (2015).
Article ADS CAS PubMed PubMed Central Google Scholar
Frangoul, H. et al. Exagamglogene autotemcel for severe sickle cell disease. N. Engl. J. Med. 390, 1649–1662 (2024).
Article CAS PubMed Google Scholar
Locatelli, F. et al. Exagamglogene autotemcel for transfusion-dependent β-thalassemia. N. Engl. J. Med. 390, 1663–1676 (2024).
Article CAS PubMed Google Scholar
Hustedt, N. & Durocher, D. The control of DNA repair by the cell cycle. Nat. Cell Biol. 19, 1–9 (2017).
Article CAS Google Scholar
Tsai, S. Q. et al. GUIDE-seq enables genome-wide profiling of off-target cleavage by CRISPR–Cas9 nucleases. Nat. Biotechnol. 33, 187–198 (2015).
Article ADS CAS PubMed Google Scholar
Cancellieri, S. et al. Human genetic diversity alters therapeutic gene editing off-target outcomes. Nat. Genet. 55, 34–43 (2023).
Article CAS PubMed Google Scholar
Kosicki, M. & Bradley, A. Repair of CRISPR–Cas9-induced double-stranded breaks leads to large deletions and complex rearrangements. Nat. Biotechnol. 36, 765–771 (2018).
Article ADS CAS PubMed PubMed Central Google Scholar
Leibowitz, M. et al. Chromothripsis as an on-target consequence of CRISPR–Cas9 genome editing. Nat. Genet. 53, 895–905 (2021).
Article CAS PubMed PubMed Central Google Scholar
Tsuchida, C. A. et al. Mitigation of chromosome loss in clinical CRISPR–Cas9-engineered T cells. Cell 186, 4567–4582.e20 (2023).
Article CAS PubMed PubMed Central Google Scholar
Zeng, J. et al. Gene editing without ex vivo culture evades genotoxicity in human hematopoietic stem cells. Cell Stem Cell 32, 191–208 (2025).
Article CAS PubMed Google Scholar
Komor, A. C., Kim, Y. B., Packer, M. S., Zuris, J. A. & Liu, D. R. Programmable editing of a target base in genomic DNA without double-stranded DNA cleavage. Nature 533, 420–424 (2016).
Article ADS CAS PubMed PubMed Central Google Scholar
Gaudelli, N. M. et al. Programmable base editing of A•T to G•C in genomic DNA without DNA cleavage. Nature 551, 464–471 (2017).
Article ADS CAS PubMed PubMed Central Google Scholar
Anzalone, A. V. et al. Search-and-replace genome editing without double-strand breaks or donor DNA. Nature 576, 149–157 (2019).
Article ADS CAS PubMed PubMed Central Google Scholar
Kim, D. Y. et al. Unbiased investigation of specificities of prime editing systems in human cells. Nucleic Acids Res. 48, 10576–10589 (2020).
Article CAS PubMed PubMed Central Google Scholar
Liang, S. et al. Genome-wide profiling of prime editor off-target sites in vitro and in vivo using PE-tag. Nat. Methods 20, 898–907 (2023).
Article CAS PubMed PubMed Central Google Scholar
Everette, K. A. et al. Ex vivo prime editing of patient haematopoietic stem cells rescues sickle-cell disease phenotypes after engraftment in mice. Nat. Biomed. Eng. 7, 616–628 (2023).
Article CAS PubMed PubMed Central Google Scholar
Chen, P. J. & Liu, D. R. Prime editing for precise and highly versatile genome manipulation. Nat. Rev. Genet. 24, 161–177 (2022).
Article CAS PubMed PubMed Central Google Scholar
Anzalone, A. V. et al. Programmable deletion, replacement, integration and inversion of large DNA sequences with twin prime editing. Nat. Biotechnol. 40, 731–740 (2022).
Article CAS PubMed Google Scholar
Yarnall, M. T. N. et al. Drag-and-drop genome insertion of large sequences without double-strand DNA cleavage using CRISPR-directed integrases. Nat. Biotechnol. 41, 500–512 (2022).
Article PubMed PubMed Central Google Scholar
Pandey, S. et al. Efficient site-specific integration of large genes in mammalian cells via continuously evolved recombinases and prime editing. Nat. Biomed. Eng. 9, 22–39 (2025).
Article CAS PubMed Google Scholar
Mukhametzyanova, L. et al. Activation of recombinases at specific DNA loci by zinc-finger domain insertions. Nat. Biotechnol. 42, 1844–1854 (2024).
Article CAS PubMed PubMed Central Google Scholar
Lampe, G. D. et al. Targeted DNA integration in human cells without double-strand breaks using CRISPR-associated transposases. Nat. Biotechnol. 42, 87–98 (2024).
Article CAS PubMed Google Scholar
Durrant, M. G. et al. Systematic discovery of recombinases for efficient integration of large DNA sequences into the human genome. Nat. Biotechnol. 41, 488–499 (2023).
Article CAS PubMed Google Scholar
Witte, I. P. et al. Programmable gene insertion in human cells with a laboratory-evolved CRISPR-associated transposase. Science 388, eadt5199 (2025).
Article CAS PubMed PubMed Central Google Scholar
Estes, B. J. G. et al. Development of circular AAV cargos for targeted seamless insertion with large serine integrases. Mol. Ther. Methods Clin. Dev. 33, 101490 (2025).
Article CAS PubMed PubMed Central Google Scholar
Perry, N. T. et al. Megabase-scale human genome rearrangement with programmable bridge recombinases. Science https://doi.org/10.1126/science.adz0276 (2025).
Article PubMed PubMed Central Google Scholar
Charlesworth, C. T., Hsu, I., Wilkinson, A. C. & Nakauchi, H. Immunological barriers to haematopoietic stem cell gene therapy. Nat. Rev. Immunol. 22, 719–733 (2022).
Article CAS PubMed PubMed Central Google Scholar
Decout, A., Katz, J. D., Venkatraman, S. & Ablasser, A. The cGAS–STING pathway as a therapeutic target in inflammatory diseases. Nat. Rev. Immunol. 21, 548–569 (2021).
Article CAS PubMed PubMed Central Google Scholar
Ferrari, S. et al. Choice of template delivery mitigates the genotoxic risk and adverse impact of editing in human hematopoietic stem cells. Cell Stem Cell 29, 1428–1444.e9 (2022).
Article CAS PubMed PubMed Central Google Scholar
Schiroli, G. et al. Precise gene editing preserves hematopoietic stem cell function following transient p53-mediated DNA damage response. Cell Stem Cell 24, 551–565 (2019).
Article CAS PubMed PubMed Central Google Scholar
Shy, B. R. et al. High-yield genome engineering in primary cells using a hybrid ssDNA repair template and small-molecule cocktails. Nat. Biotechnol. 41, 521–531 (2023).
Article CAS PubMed Google Scholar
Xie, K. et al. Efficient non-viral immune cell engineering using circular single-stranded DNA-mediated genomic integration. Nat. Biotechnol. 43, 1821–1832 (2024).
Article PubMed Google Scholar
Gibson, D. G. et al. Enzymatic assembly of DNA molecules up to several hundred kilobases. Nat. Methods 6, 343–345 (2009).
Article CAS PubMed Google Scholar
Jones, D. H. & Howard, B. H. A rapid method for recombination and site-specific mutagenesis by placing homologous ends on DNA using polymerase chain reaction. Biotechniques 10, 62–66 (1991).
CAS PubMed Google Scholar
Bubeck, P., Winkler, M. & Bautsch, W. Rapid cloning by homologous recombination in vivo. Nucleic Acids Res. 21, 3601–3602 (1993).
Article CAS PubMed PubMed Central Google Scholar
Watson, J. F. & García-Nafría, J. In vivo DNA assembly using common laboratory bacteria: a re-emerging tool to simplify molecular cloning. J. Biol. Chem. 294, 15271–15281 (2019).
Article CAS PubMed PubMed Central Google Scholar
Jiang, T., Zhang, X., Weng, Z. & Xue, W. Deletion and replacement of long genomic sequences using prime editing. Nat. Biotechnol. 40, 227–234 (2021).
Article CAS PubMed PubMed Central Google Scholar
Choi, J. et al. Precise genomic deletions using paired prime editing. Nat. Biotechnol. 40, 218–226 (2022).
Article CAS PubMed Google Scholar
Savage, S. A. et al. TINF2, a component of the shelterin telomere protection complex, is mutated in dyskeratosis congenita. Am. J. Hum. Genet. 82, 501–509 (2008).
Article CAS PubMed PubMed Central Google Scholar
Walne, A. J., Vulliamy, T., Beswick, R., Kirwan, M. & Dokal, I. TINF2 mutations result in very short telomeres: analysis of a large cohort of patients with dyskeratosis congenita and related bone marrow failure syndromes. Blood 112, 3594–3600 (2008).
Article CAS PubMed PubMed Central Google Scholar
Chen, P. J. et al. Enhanced prime editing systems by manipulating cellular determinants of editing outcomes. Cell 184, 5635–5652 (2021).
Article CAS PubMed PubMed Central Google Scholar
Yan, J. et al. Improving prime editing with an endogenous small RNA-binding protein. Nature 628, 639–647 (2024).
Article ADS CAS PubMed PubMed Central Google Scholar
Levesque, S. et al. Marker-free co-selection for successive rounds of prime editing in human cells. Nat. Commun. 13, 5909 (2022).
Article ADS CAS PubMed PubMed Central Google Scholar
Dacquay, L. C. et al. Dual inhibition of DNA-PK and Polϴ boosts precision of diverse prime editing systems. Nat. Commun. 16, 4290 (2025).
Article ADS CAS PubMed PubMed Central Google Scholar
Yan, X. Efficient precise integration of large DNA sequences with 3′-overhang dsDNA donors using CRISPR/Cas9. Proc. Natl Acad. Sci. USA 120, e2221127120 (2023).
Article Google Scholar
Sakuma, T., Nakade, S., Sakane, Y., Suzuki, K. I. T. & Yamamoto, T. MMEJ-Assisted gene knock-in using TALENs and CRISPR–Cas9 with the PITCh systems. Nat. Protoc. 11, 118–133 (2016).
Article CAS PubMed Google Scholar
Levesque, S., Cosentino, A., Verma, A., Genovese, P. & Bauer, D. E. Enhancing prime editing in hematopoietic stem and progenitor cells by modulating nucleotide metabolism. Nat. Biotechnol. 43, 534–538 (2025).
Article CAS PubMed Google Scholar
Roth, T. L. et al. Reprogramming human T cell function and specificity with non-viral genome targeting. Nature 559, 405–409 (2018).
Article ADS CAS PubMed PubMed Central Google Scholar
Zikry, T. M. et al. Cell cycle plasticity underlies fractional resistance to palbociclib in ER+/HER2− breast tumor cells. Proc. Natl Acad. Sci. USA 121, e2309261121 (2024).
Article CAS PubMed PubMed Central Google Scholar
Albanese, M. et al. Rapid, efficient and activation-neutral gene editing of polyclonal primary human resting CD4+ T cells allows complex functional analyses. Nat. Methods 19, 81–89 (2022).
Article CAS PubMed Google Scholar
Shin, J. J. et al. Controlled cycling and quiescence enables efficient HDR in engraftment-enriched adult hematopoietic stem and progenitor cells. Cell Rep. 32, 108093 (2020).
Article CAS PubMed PubMed Central Google Scholar
Lauridsen, F. K. B. et al. Differences in cell cycle status underlie transcriptional heterogeneity in the HSC compartment. Cell Rep. 24, 766–780 (2018).
Article CAS PubMed Google Scholar
Oedekoven, C. A. et al. Hematopoietic stem cells retain functional potential and molecular identity in hibernation cultures. Stem Cell Rep. 16, 1614–1628 (2021).
Article CAS Google Scholar
Hanlon, K. S. et al. High levels of AAV vector integration into CRISPR-induced DNA breaks. Nat. Commun. 10, 4439 (2019).
Article ADS PubMed PubMed Central Google Scholar
Suchy, F. P. et al. Genome engineering with Cas9 and AAV repair templates generates frequent concatemeric insertions of viral vectors. Nat. Biotechnol. 43, 204–213 (2024).
Article PubMed PubMed Central Google Scholar
Carnevale, J. et al. RASA2ablation in T cells boosts antigen sensitivity and long-term function. Nature 609, 174–182 (2022).
Article ADS CAS PubMed PubMed Central Google Scholar
Wei, J. et al. Targeting REGNASE-1 programs long-lived effector T cells for cancer therapy. Nature 576, 471–476 (2019).
Article ADS CAS PubMed PubMed Central Google Scholar
Chen, B. et al. Dynamic imaging of genomic loci in living human cells by an optimized CRISPR/Cas system. Cell 155, 1479–1491 (2013).
Article ADS CAS PubMed PubMed Central Google Scholar
Nelson, J. W. et al. Engineered pegRNAs improve prime editing efficiency. Nat. Biotechnol. 40, 402–410 (2021).
Article PubMed PubMed Central Google Scholar
Canny, M. D. et al. Inhibition of 53BP1 favors homology-dependent DNA repair and increases CRISPR–Cas9 genome-editing efficiency. Nat. Biotechnol. 36, 95–102 (2018).
Article CAS PubMed Google Scholar
Cong, L. et al. Multiplex genome engineering using CRISPR/Cas systems. Science 339, 819–823 (2013).
Article ADS CAS PubMed PubMed Central Google Scholar
Levesque, S. et al. Pharmacological control of CAR T cells through CRISPR-driven rapamycin resistance. Preprint at BioRxiv https://doi.org/10.1101/2023.09.14.557485 (2024).
Brinkman, E. K., Chen, T., Amendola, M. & Van Steensel, B. Easy quantitative assessment of genome editing by sequence trace decomposition. Nucleic Acids Res. 42, e168 (2014).
Article PubMed PubMed Central Google Scholar
Xu, L., Liu, Y. & Han, R. BEAT: a Python program to quantify base editing from Sanger sequencing. CRISPR J. 2, 223–229 (2019).
Article CAS PubMed PubMed Central Google Scholar
Conant, D. et al. Inference of CRISPR edits from Sanger trace data. CRISPR J. 5, 123–130 (2022).
Article CAS PubMed Google Scholar
Clement, K. et al. CRISPResso2 provides accurate and rapid genome editing sequence analysis. Nat. Biotechnol. 37, 215–226 (2019).
Article ADS Google Scholar
Hwang, G.-H. et al. Analysing long-read CRISPR experiments with CRISPRLungo. Nat. Biomed. Eng. https://doi.org/10.1038/s41551-026-01776-7 (2026).
Wick, R. Badread: simulation of error-prone long reads. J. Open Source Softw. 4, 1316 (2019).
Article ADS Google Scholar
Lai, T. P., Wright, W. E. & Shay, J. W. Generation of digoxigenin-incorporated probes to enhance DNA detection sensitivity. Biotechniques 60, 306–309 (2016).
Article CAS PubMed Google Scholar
Malinin, N. L. et al. Defining genome-wide CRISPR–Cas9 genome-editing nuclease activity with GUIDE-seq. Nat. Protoc. 16, 5592–5615 (2021).
Article CAS PubMed PubMed Central Google Scholar
Corre, G., Rouillon, M., Mombled, M. & Amendola, M. Advanced pipeline for CRISPR/Cas9 off-targets detection in Guide-seq and related integration-based assays. Preprint at BioRxiv https://doi.org/10.1101/2025.09.30.679427 (2025).
Download references
Acknowledgements
We thank G. Casirati, P. Genovese, S. Hallett, D. Liu, A. Li and C. Brendel for helpful discussions; S. A. Wolfe for sharing Tn5 transposase; J. P. Manis and the BCH Blood Donor Center for providing LRS blood cones; the HSCI-BCH Flow Cytometry Research Lab for technical support; and B. Liu, W. Xue and E. J. Sontheimer for communication of unpublished results.
Funding
D.E.B. was supported by the Doris Duke Foundation (2022092), the St. Jude Children’s Research Hospital Collaborative Research Consortium, the Harvard Stem Cell Institute and the National Institutes of Health (R01HG013618 and R01HL165061). S.A. was supported by the National Institutes of Health (R01DK107716). L.P. was supported by a Rappaport MGH Research Scholar Award 2024-2029. S.L. was supported by a Banting Postdoctoral Fellowship from the Canadian Institutes of Health Research, and a Next Generation of Scientists Award by the Cancer Research Society. N.K. is supported by an Overseas Research Fellowship from the Japan Society for the Promotion of Science. V.T. was supported by the German Research Foundation. V.A.C.S. was supported by a Postdoctoral Fellowship from the American Heart Association (25POST1377446). W.M. is supported by the National Institute of Health Research (NIH F30DK135340, T32GM007753 and T32GM144273), and a Medical Scientist Training Fellowship from Harvard Stem Cell Institute. HSPCs were obtained from Fred Hutch Cooperative Center of Excellence in Hematology (U54DK106829).
Ethics declarations
Competing interests
S.L. and D.E.B. have filed a patent application (WO 2025/038842 A1) covering PA technology. L.P. has financial interests in Edilytics Inc. Interests of L.P. were reviewed and are managed by Massachusetts General Hospital and Partners HealthCare in accordance with their conflict of interest policies. The other authors declare no competing interests.
Peer review
Peer review information
Nature thanks the anonymous reviewers for their contribution to the peer review of this work. Peer reviewer reports are available.
Additional information
Publisher’s note Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
Extended data figures and tables
Extended Data Fig. 1 Prime assembly enables exon recoding at the TINF2 dyskeratosis congenita cluster.
(a) Schematic of TINF2 dyskeratosis congenita (DC) cluster recoding using PA. (b) PA and imprecise allele quantification at TINF2 as determined by CRISPResso2 analysis from amplicon sequencing. K562 cells were electroporated with PA vectors and genomic DNA was harvested 3 days post-nucleofection. Data are plotted as mean ± s.e.m. from n = 3 independent biological replicates. (c) Schematic representation of five different ssDNA designs with varying overlap lengths. (d) Same as in (b) using different ssDNA overlap designs shown in (c). Data are plotted as mean ± s.e.m. from n = 3 independent biological replicates. (e) Same as in (b) using different flap lengths and ssDNA_v5 donors. Data are plotted as mean ± s.e.m. from n = 3 independent biological replicates. (f) Same as in (b) using HEK293T cells and the indicated concentration of each ssDNA_v5 donor. Where indicated, cells were cultured in the presence or absence of AZD7648 and PolQi1 for three days before genomic DNA was harvested. Data are plotted as mean ± s.e.m. from n = 3 independent biological replicates. (g) Percentage of precise PA and indels at the 5′ and 3′ junctions (in–out genotyping) as determined by CRISPResso2 analysis from amplicon sequencing. Genomic DNA from a single cell-derived K562 clone recoded at TINF2 exon 6 was used as a control. Data are plotted as mean ± s.e.m. from n = 3 independent biological replicates.
Source data
Extended Data Fig. 2 Characterization of prime assembly outcomes via amplicon sequencing.
(a) Schematic representation of the different types of outcomes observed after prime assembly as determined by amplicon sequencing. (b) Representative CRISPResso2 allele plot from an unedited K562 control. Representative allele plot is from one out of three independent biological replicates. (c) Representative CRISPResso2 allele plot from Fig. 1d with 16 pmol each ssDNA donor harboring phosphorothioate bonds. Alleles are aligned against their cognate reference sequence (unedited, precise PA, or flap integration) with 5 most abundant alleles per category shown in decreasing frequency order. Representative allele plot is from one out of three independent biological replicates. Red boxes illustrate insertions, and nucleotide substitutions are highlighted in bold.
Extended Data Fig. 3 Prime assembly enables precise site-specific transgene integration at the TRAC locus.
(a) Schematic representation of targeted transgene integration at the TRAC locus using PA. The primers used for in–out PCR-based genotyping are illustrated in orange. (b) Percentage of PA allele as determined by ddPCR. K562 cells were electroporated with PA vectors and 1.5 µg 3′-odsDNA donor targeting TRAC, and genomic DNA was harvested three days post-nucleofection. Data are plotted as mean ± s.e.m. from n = 3 independent biological replicates. Dotted lines indicate the average background level of the donor only controls. (c) Representative Sanger chromatograms from in–out PCR amplification of the prime assembly junctions at the TRAC locus from the samples in (b). The pegRNA spacers and cut sites, primer binding site (PBS), and the annealing regions between the 3′ flaps and donor overhangs are annotated. n = 3 independent biological replicates.
Source data
Extended Data Fig. 4 Prime assembly enables site-specific integration of up to 12 kb.
(a) Schematic representation of targeted integration of PuroR-EGFP transgenes ranging from 3.1 kb to 12.1 kb at the AAVS1 locus using PA. The primers used for in–out PCR-based genotyping are illustrated in orange. (b) Schematic representation of AAVS1-Exon1-PuroR chimeric mRNA splicing after targeted transgene integration at AAVS1. (c) Relative AAVS1-Exon1-PuroR chimeric mRNA expression as determined by RT-qPCR. K562 cells were electroporated with PA vectors and 2 µg dsDNA donor targeting AAVS1, and cultured in the presence of 1 µM AZD7648 and 1.5 µM PolQi1. Three days post-nucleofection, cells were cultured in the presence or absence of 1 µg/ml puromycin until all non-resistant cells were eliminated. Total RNA was harvested 21 days post-nucleofection. Data are plotted as mean ± s.e.m. from n = 3 independent biological replicates. (d) Representative Sanger chromatograms from in–out PCR amplification of the prime assembly junctions at the AAVS1 locus after 12.1 kb transgene integration from the samples in Fig. 2e and (c). The pegRNA spacers and cut sites, primer binding site (PBS), and the annealing regions between the 3′ flaps and donor overhangs are annotated. n = 3 independent biological replicates.
Source data
Extended Data Fig. 5 Donor-Seq integration profile reveals high genome-wide specificity with prime assembly.
(a) Percentage of precise junction reads from samples in Fig. 2b (IL2RG), Fig. 2d (TRAC) and Fig. 2g (AAVS1) as determined by Donor-Seq. Data are plotted as mean ± s.e.m. from n = 3 independent biological replicates. (b) Genome-wide donor integration profile from samples in Fig. 2b (IL2RG), Fig. 2d (TRAC) and Fig. 2g (AAVS1), as determined by Donor-Seq. Each plot represents the total number of deduplicated unique molecular identifier (UMI) reads from three independent biological replicates. (c) Total count of deduplicated UMI reads for each guide-dependent off-target identified by Donor-Seq from samples in (b). n = 3 independent biological replicates.
Source data
Extended Data Fig. 6 Prime assembly enables targeted assembly and integration of up to four dsDNA donors.
(a) Schematics of two split dsDNA donors integration at AAVS1 via prime assembly. The BstEII-HF restriction sites and the Southern blotting probes are illustrated. The AAVS1 probes (orange) target both the AAVS1-WT and the AAVS1-hPGK1-EGFP knock-in alleles, and the EGFP probes (blue) target the AAVS1-hPGK1-EGFP knock-in allele only. (b) Detection of site-specific transgene integration at AAVS1 using Southern blotting from samples in Fig. 2g. K562 cells were electroporated with PA vectors and 2 µg dsDNA donor or an equimolar ratio of split dsDNA donors targeting AAVS1, and cultured in the presence of 1 µM AZD7648 and 1.5 µM PolQi1 for three days. Genomic DNA was harvest 21 days post-nucleofection. Representative blot is from one out of three independent biological replicates. (c) Percentage of precise PA and indels at the flap and donor junctions as determined by CRISPResso2 analysis from amplicon sequencing. A plasmid encoding the expected allele was used as a control. Data are plotted as mean ± s.e.m. from n = 3 independent biological replicates. (d) Schematics and representative Sanger chromatograms from in–out PCR-based genotyping of the prime assembly junctions at the AAVS1 locus after four fragment assembly and integration from samples in Fig. 2i. All dsDNA donors are numbered from one to four. The primers used for the in–out PCR-based genotyping are illustrated in blue and orange. The pegRNA spacers and cut sites, primer binding site (PBS), and the annealing regions between the 3′ flaps and donor overhangs are annotated. Representative Sanger chromatograms are from one out of 3 independent biological replicates.
Source data
Extended Data Fig. 7 Benchmarking prime assembly against other gene targeting platforms using ddPCR.
(a) Schematic representation of site-specific transgene integration detection or WT allele disruption using ddPCR. Drop-off alleles include small indels, PA flap integration with or without indels, incomplete PASSIGE, and large deletions and rearrangements that disrupt the WT allele. (b) Percentage of precise integration and drop-off alleles as determined by ddPCR. K562 or Jurkat cells were electroporated with the indicated vector and donor targeting hPGK1-EGFP to AAVS1, EGFP to IL2RG, or EGFP to TRAC, and cultured for 21 days before collecting genomic DNA to eliminate background signal from non-integrated episomal donors. Where indicated, Jurkat cells were treated with 0.5 µM AZD7648 and 0.5 µM PolQi1 for three days post-nucleofection. Due to the presence of homology arms, the homology-directed repair (HDR) conditions do not discriminate between on-target integrations and episomal or off-target integration, and the dotted lines indicate the average background level of the donor only controls. Data are plotted as mean ± s.e.m. from n = 3 independent biological replicates.
Source data
Extended Data Fig. 8 Site-specific transgene integration at AAVS1 and IL2RG using single-stranded DNA donors.
(a) Schematic representation of targeted transgene integration at the AAVS1 locus using two different ssDNA donor designs and pegRNAs encoding for 32-nts flaps. (b) Percentage of EGFP+ cells as determined by flow cytometry. K562 cells were electroporated with PA vectors and 4 pmol of each ssDNA donor targeting AAVS1 or an equimolar concentration of 3’-odsDNA donor, and the percentage of EGFP+ cells was quantified 7 days post-nucleofection. Data are plotted as mean ± s.e.m. from n = 3 independent biological replicates. Dotted lines indicate the average background level of the donor only controls. (c) Percentage of EGFP+ cells as determined by flow cytometry. K562 cells were electroporated with PA vectors and 1.5 µg 3′-odsDNA donor (1X) targeting AAVS1 or the indicated molar ratio of ssDNA donors, and the percentage of EGFP-expressing cells was quantified 7 days post-nucleofection. Data are plotted as mean ± s.e.m. from n = 3 independent biological replicates. Dotted lines indicate the average background level of the donor only controls. (d) Same as in (c) with donors targeting IL2RG. Data are plotted as mean ± s.e.m. from n = 3 independent biological replicates. (e) Schematic representation of targeted transgene integration at the AAVS1 locus using the asymmetric v3 ssDNA donors. Stabilizing phosphorothioate bonds are illustrated with orange stars. (f) Percentage of EGFP+ cells as determined by flow cytometry. K562 cells were electroporated with PA vectors and the indicated concentration of ssDNA donors targeting AAVS1, and the percentage of EGFP+ cells was quantified 7 days post-nucleofection. Data are plotted as mean ± s.e.m. from n = 3 independent biological replicates. Dotted lines indicate the average background level of the donor only controls. (g) Same as in (e) for site-specific integration at IL2RG. (h) Percentage of EGFP+ cells as determined by flow cytometry. K562 cells were electroporated with PA vectors and 4 pmol of each ssDNA donor targeting IL2RG or an equimolar concentration of 3′-odsDNA donor, and the percentage of EGFP+ cells was quantified 7 days post-nucleofection. Data are plotted as mean ± s.e.m. from n = 3 independent biological replicates. Dotted lines indicate the average background level of the donor only controls.
Source data
Extended Data Fig. 9 Long-range nuclease prime assembly enables targeted chromosome arm deletion and megabase inversion.
(a) Schematic representation of targeted PuroR transgene integration and the deletion of the q arm of chromosome 19 using long-range nuclease prime assembly. (b) PCR-based genotyping of the expected PuroR transgene integration and 27 Mb deletion allele after nuclease prime assembly and puromycin selection with the primers illustrated in orange in (a). K562 cells were electroporated with nuclease PA vectors and 4 pmol of each ssDNA donor targeting chromosome 19, and cells were cultured with 1 µg/ml puromycin three days post-nucleofection until all non-resistant cells were eliminated. Representative gel is from one out of three independent biological replicates. (c) Schematic representation of targeted PuroR transgene integration and the deletion of genes present on the q arm of chromosome 19 using long-range nuclease prime assembly. (d) Percentage of PA allele and copy loss of genes present on chromosome 19 as determined by ddPCR, as described in (b). Data are plotted as mean ± s.e.m. from n = 3 independent biological replicates. (e) Schematic representation of targeted PuroR transgene integration and the 1 Mb inversion at chromosome 7 using ssDNA donors. (f) PCR-based genotyping of the expected PuroR transgene integration and 1 Mb inversion allele after prime assembly and puromycin selection, as described in (b). Representative gel is from one out of three independent biological replicates. (g) Same as in (b) after the installation of a 1 Mb inversion at chromosome 7. Data are plotted as mean ± s.e.m. from n = 3 independent biological replicates. (h) Genotype of single cell-derived K562 clones after the installation of a 1 Mb inversion at chromosome 7 using nuclease prime assembly and puromycin selection, or standard nuclease-driven inversion without selection, as determined by long-read nanopore sequencing. n = 47 and 48 single cell-derived clones for nuclease prime assembly and standard nuclease-driven inversion, respectively.
Source data
Extended Data Fig. 10 Prime assembly in G0/G1 resting CD3+ T cells using low concentrations of cytokines.
(a) Timeline for the cell cycle experiment. CD3+ T cells were stimulated for 24 h with a CD3/CD28/CD2 activator in the presence of 300 U/ml IL-2, or kept in resting culture conditions with 1 ng/ml IL-7 and IL-15. After 24 h, T cells were electroporated with PE7 mRNA, synthetic pegRNAs, and ssDNA donors (or mock electroporated), and cultured for three days with 300 U/ml IL-2 (activated) or 1 ng/ml IL-7 and IL-15 (resting). (b) Cell cycle progression as determined by flow cytometry. CD3+ T cells were activated or not for 24 h, electroporated (mock), and cultured for an additional three days. Cell cycle progression was monitored for four days. Data are plotted as mean ± s.e.m. from n = 4 independent biological replicates performed with CD3+ T cells from four healthy donors. (c) Schematic representation of targeted transgene integration at the TRAC locus using asymmetric v3 ssDNA donors. (d) Same as in (b) with CD3+ T cells electroporated with PE7 mRNA, synthetic pegRNAs, and ssDNA donors, as described in Fig. 5g. Cell cycle analysis was performed three days post-nucleofection. Data are plotted as mean ± s.e.m. from n = 3 independent biological replicates performed with CD3+ T cells from three healthy donors.
Source data
Supplementary information
Source data
Rights and permissions
Open Access This article is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License, which permits any non-commercial use, sharing, distribution and reproduction in any medium or format, as long as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons licence, and indicate if you modified the licensed material. You do not have permission under this licence to share adapted material derived from this article or parts of it. The images or other third party material in this article are included in the article’s Creative Commons licence, unless indicated otherwise in a credit line to the material. If material is not included in the article’s Creative Commons licence and your intended use is not permitted by statutory regulation or exceeds the permitted use, you will need to obtain permission directly from the copyright holder. To view a copy of this licence, visit http://creativecommons.org/licenses/by-nc-nd/4.0/.
Reprints and permissions
About this article
Cite this article
Levesque, S., Kawashima, N., Hwang, GH. et al. Targeted genomic integration and rearrangement using prime assembly. Nature (2026). https://doi.org/10.1038/s41586-026-11024-2
Download citation
Received:
Accepted:
Published:
Version of record:
DOI: https://doi.org/10.1038/s41586-026-11024-2