Proximity-guided graph learning reveals tumour-associated proximity antigens

Nature正文已收录本站

Main

The cell surface is a dynamic landscape where membrane protein organization regulates diverse biological processes, including immune synapse formation, epithelial junctional organization and receptor-mediated signalling1,5. Membrane protein proximity is a defining feature of immune receptors6, cytokine receptors7 and receptor tyrosine kinases (RTKs)3, many of which assemble into localized signalling hubs2.

In cancer, the cell surface proteome is remodelled through changes in protein abundance, trafficking, post-translational modifications, mutations and membrane organization4,8,9,10. These changes reshape surface protein spatial networks, disrupting normal physiology and promoting disease progression8,9,10,11,12. Notably, this protein-level spatial organization is not captured by gene or protein expression analyses, which measure abundance rather than proximity. Resolving membrane protein organization is therefore critical for contextualizing TAAs within functional neighbourhoods and guiding multispecific therapeutic design.

Despite the importance of spatial organization, scalable approaches for mapping surface protein microenvironments remain limited. We previously introduced complementary photocatalytic proximity-labelling technologies that covalently label proteins within a nanoscale radius to map membrane microenvironments13,14,15,16. In this Article, we integrate these orthogonal chemistries into an industrialized high-throughput microenvironment mapping (micromapping) workflow. Because labelling is governed by reactive intermediates with finite lifetimes and diffusion distances, these measurements capture local membrane neighbourhoods rather than only direct physical interactions. The workflow incorporates standardized procedures and cross-experiment normalization to quantitatively compare protein proximity across targets, cell systems and tumour contexts.

Using this platform, we mapped surface protein microenvironments across diverse cancer models. To extend these measurements beyond directly targeted proteins, we developed MetaMap, an analytical framework that defines spatial protein communities and infers conserved proximity relationships among non-targeted proteins. We then applied graph-based learning informed by protein proximity and abundance to identify candidate co-target pairs, which were subsequently prioritized using normal tissue expression, tumour-versus-normal expression, clinical proteomics and disease-relevant signalling features (Fig. 1). This approach prioritizes proximity-defined co-targets by revealing which TAAs are organized together on tumour cells, rather than being merely present.

Fig. 1: A data-driven framework for identifying TAPAs.

Industrialized microenvironment mapping (left) generates a global proximity network that captures local protein neighbourhoods across diverse cell systems. Integrating these proximity signatures with protein abundance enables detection of disease-associated protein communities and nomination of candidate co-targets (middle); pathway annotations, clinical proteomics and tumour-versus-normal expression further contextualize these candidates for biological interpretation and prioritization. Together, these analyses nominate TAPAs: surface antigens defined not only by abundance but also by disease-relevant proximity to tumour anchors (right), expanding therapeutic design towards spatial co-biology and actionable protein neighbourhoods. PC, photocatalyst; MS, mass spectrometry. The central network schematic is adapted from ref. 58, CC BY 4.0.

The distinction between spatial organization and expression is therapeutically important because TAAs remain a foundation of targeted cancer therapeutics, including antibody-drug conjugates (ADCs), T cell engagers (TCEs) and CAR-T cell therapies17,18, yet are often expressed in healthy tissues, where engagement can cause dose-limiting toxicity and a narrow therapeutic window. These limitations motivate a more context-aware approach to co-targeting, in which partner antigens are selected based on tumour-specific proximity relationships rather than expression alone. To test this strategy, we focused on RTKs, a clinically tractable family of TAAs (Supplementary Fig. 2), whose interaction networks make them well suited for proximity-informed co-target discovery. Previous studies mapped RTK intracellular interaction networks using affinity purification and enzyme-based proximity labelling in HEK293 cells19, leaving the extracellular membrane organization largely unexplored. Using this extracellularly anchored approach, we systematically resolved surface protein microenvironments of endogenously expressed RTKs across receptor families and tumour contexts.

Here we introduce the concept of TAPAs, proteins defined by disease-associated spatial proximity to canonical TAAs rather than by expression alone. Applying this concept to RTKs, we identified EGFR–CDCP1 as a TAA–TAPA pair that is spatially co-enriched across multiple tumour types and enhances tumour cell killing when co-engaged using ADC and TCE modalities. Together, this work presents a systematic surface protein proximity atlas and graph learning framework and establishes TAPAs as a new target class for multispecific therapeutic discovery.

Large-scale micromapping of RTKs

To establish a scalable workflow for micromapping, we profiled 12 RTKs spanning 10 structural families across 28 cancer cell systems with diverse tissue origins, antigen expression and activation states (Supplementary Figs. 3–6). Photocatalysts were delivered through secondary antibody conjugates using either of two orthogonal visible-light proximity-labelling systems: iridium/diazirine (Ir/Dz), which generates short-lived carbene intermediates for highly localized labelling, or riboflavin/biotin-tyramide (RFT/BT), which generates longer-lived phenoxy radicals that sample broader local membrane neighbourhoods20,21 (Fig. 2a and Supplementary Fig. 7). Because the spatial extent of labelling is determined by reactive intermediate lifetime and the geometry of photocatalyst placement22, the secondary antibody format positions photocatalysts further from the target, expanding the labelling footprint (Fig. 2a). The systems also differ in residue labelling preference, providing additional chemical complementarity20,21. Using this dual-chemistry workflow, we produced 248 proximity maps (Fig. 2a). To enable mapping at this scale, we standardized cell input and automated protein pulldown, digestion and tandem mass tag (TMT) labelling using a KingFisher Apex system. Across the dataset, intended RTK anchor proteins were consistently enriched, and strong concordance was observed between chemistries (Supplementary Fig. 8), supporting the reproducibility of systematic microenvironment mapping.

Fig. 2: Scalable micromapping of RTK surface neighbourhoods and benchmarking against the annotated protein interactome.

a, Overview of the industrialized micromapping workflow used to generate 248 RTK-anchored proximity maps across 12 RTKs and diverse cancer cell systems, combining antibody-guided localization with either of two orthogonal blue-light proximity-labelling chemistries (mid-range RFT/BT or short-range Ir/Dz), automated enrichment and sample processing, TMT multiplexing and liquid chromatography–tandem mass spectrometry (LC–MS/MS) quantification. t1/2, half-life. b, Global proximity network summarizing enriched protein neighbourhoods across the full dataset, revealing a densely connected landscape with receptor-specific structure. c, RTK interaction network highlighting recovered known3 RTK–RTK relationships alongside additional putative proximal partners. Targeted RTKs are shown in colour and non-targeted RTKs are shown in grey. d, Pairwise neighbourhood overlap across RTKs (using Jaccard similarity) showing shared and distinct microenvironment composition, including high-overlap receptor pairs. e, Benchmarking of RTK–RTK proximity-derived associations against external protein interaction databases (from STRING, CORUM, BioGRID and IntAct). RTK pairs are categorized as ‘targeted’ when at least one receptor in the pair was directly profiled by micromapping, or ‘non-targeted’ when neither receptor was directly profiled. ‘Recovered’, ‘missed’ and ‘novel’ denote externally annotated RTK–RTK pairs detected in our data, externally annotated pairs not detected in our data, and previously unannotated proximity-enriched RTK–RTK pairs, respectively. Bars summarize the distribution of recovered, missed and novel RTK–RTK interactions. f, Recovery of annotated RTK–RTK interactions by proximity mapping. Bars indicate the fraction of database-curated interactions from STRING, CORUM, BioGRID and IntAct detected in our dataset, stratified by targeted (77 out of 173 known interactions recovered (45%)) and non-targeted (18 out of 84 (21%)) categories.

Source data

To enable quantitative comparison across the proximity atlas, we developed a statistical processing pipeline based on median absolute deviation (MAD)-normalized t-statistics for each of the 248 datasets (Supplementary Fig. 9). This normalization compensates for differences in experimental dynamic range while preserving reproducible proximity relationships. High-confidence protein associations were defined using a stringent threshold (MAD-normalized t ≥ 2.0) (Supplementary Fig. 10), yielding interaction networks supported by consistent protein co-enrichment rather than experiment-specific variation.

Global connectivity analysis revealed both densely interconnected and receptor-specific neighbourhoods (Fig. 2b). Our platform recapitulated much of the known RTK interactome, including reported heterotypic RTK associations3 such as EGFR–MET, EGFR–HER2 and HER2–HER3, and also revealed previously unannotated proximal RTK pairs, including EGFR–PTK7, MET–EPHB2 and HER3–MST1R (Fig. 2c). Comparison with publicly disclosed therapeutic pipelines revealed that only three co-targeting combinations (EGFR×MET, EGFR×HER3 and HER2×HER3) are currently in clinical-stage evaluation for dual-targeting antibody-based modalities (Supplementary Fig. 2b), highlighting that most RTK–RTK proximity relationships identified by systematic mapping remain therapeutically unexplored.

Analysis of RTK neighbourhood similarity revealed distinct subsets of receptors, including HER2, EPHA2, EGFR, MET and PTK7, with highly overlapping microenvironments relative to other RTK pairs (Jaccard index 0.4–0.5; Fig. 2d and Supplementary Fig. 11). This pair-specific pattern supports selective spatial organization rather than random protein colocalization or preferential detection of abundant surface proteins (Supplementary Figs. 12 and 13). To benchmark against curated interaction databases, we classified RTK–RTK proximal pairs detected in our dataset based on whether at least one RTK in the pair was directly profiled by micromapping (‘targeted’) or whether neither RTK in the pair was directly profiled (‘non-targeted’). We compared these pairs with STRING23, CORUM24, BioGRID25 and IntAct26 databases, revealing substantial recovery of annotated RTK–RTK interactions and numerous previously unannotated proximity relationships (Fig. 2e). Recovered interactions accounted for nearly half of annotated targeted RTK pairs and extended to 18 out of 84 annotated interactions among non-targeted RTKs, demonstrating recovery of RTK relationships beyond directly profiled receptors (Fig. 2f). These results demonstrate that systematic proximity mapping captures structured membrane organization and enables exploration of previously unannotated protein associations.

MetaMap resolves non-targeted proximity

We explored whether large-scale proximity maps could be leveraged to infer spatial relationships among non-targeted surface proteins. To address this, we computed pairwise Spearman rank correlations across all micromaps, capturing proteins whose proximity signatures co-varied across RTK anchors, tumour cell models and experimental conditions. This correlation-based framework, termed MetaMap, defines reproducible spatial protein communities and infers non-targeted proximity relationships while reducing sensitivity to proteomic sampling variability (Fig. 3a). Because these associations are inferred from proximity-labelling signatures collected within a finite labelling radius, they represent reproducible spatial co-enrichment rather than direct physical interactions alone.

Fig. 3: MetaMap infers spatial protein communities from large-scale micromaps.

a, MetaMap workflow: micromaps are quality-filtered and normalized to reduce batch or efficiency effects, converted to protein-by-protein association matrices (Spearman rank correlation), and integrated to infer mutually proximal surface protein relationships beyond directly targeted antigens. b, Global MetaMap correlation heat map across proteins. The matrix positions of the 12 mapped RTK targets are indicated on the left and highlighted (yellow box). Edge annotations denote membrane proteins in red and non-membrane proteins in grey. c, Recovery of known protein–protein interactions across targeted RTK microenvironments. Each column represents a curated interaction pair, and rows indicate the RTK-anchored experiment in which the interaction was detected. d, Benchmarking MetaMap associations against curated interaction databases. Left, fold enrichment over background for non-targeted interactions: CORUM 6.52× (P = 0.002), STRING 4.94× (P = 0.002), BioGRID 2.72× (P = 0.002), IntAct 2.70× (P = 0.002) and combined 2.67× (P = 0.002); one-sided (upper-tail) permutation test, n = 500 permutations. In every case no random permutation reached the observed precision, so P = 0.002 is the minimum resolvable value. Right, precision across reference databases, with targeted interactions shown for comparison. e, Representative MetaMap community comprising integrin family members and cell adhesion molecules. Core proteins (coloured) include integrin alpha and beta subunits; extended neighbourhood shown as uncoloured nodes. Edges indicate a high-confidence correlation metric (Spearman ρ >0.4) and distinguish known, database annotated interactions (dashed) from potentially novel core–core (solid green) and extended network associations (solid red).

Source data

The resulting correlation structure reveals an organized surface interactome in which RTKs form distinct, reproducible heat map patterns (Fig. 3b). Protein abundance profile similarity was nearly identical between same- and different-cluster pairs, indicating that the clustering was not driven by protein abundance (Supplementary Fig. 14). Benchmarking demonstrated robust recovery of known direct and indirect interactions across nearly all RTKs (Fig. 3c). An exception was the FGFR family anchors (for example, FGFR2), which showed sparse recovery, probably reflecting more limited context sampling and distinct membrane organization across FGFR members (Supplementary Fig. 15). MetaMap-derived non-targeted associations matched curated interactions more frequently than size-matched random protein pairs (Fig. 3d).

Community detection identified coherent membrane protein assemblies within the MetaMap association network (Fig. 3e). One representative community comprised an integrin-centred adhesion network containing ITGA and ITGB family members27 together with canonical adhesion proteins including CD44, ALCAM and L1CAM that participate in established interactions28,29. Recovery of these adhesion components supports the ability of MetaMap to resolve biologically meaningful membrane neighbourhoods from correlated proximity signatures.

The same community contained recurrent proximity-defined associations absent from curated interaction databases. For example, MetaMap connected FAP, CD81, RECK and PAPPA to the core integrin network (Fig. 3e). Although these proteins have been implicated in adhesion-related biology30,31,32,33, their recurrent spatial co-association suggests unrecognized organization within integrin microenvironments. These analyses establish MetaMap as a correlation-based framework for defining spatial protein communities and inferring reproducible proximity relationships beyond directly targeted proteins. Therapeutic prioritization, however, requires combining spatial organization with biological context and tumour selectivity to identify clinically actionable TAA–TAPA pairs.

Graph learning for therapeutic co-pairs

Having established that large-scale micromapping captures reproducible, structured protein co-associations across the tumour cell surface, we next asked whether these spatial relationships could inform therapeutic co-target prioritization. We sought to identify co-target pairs defined by spatial organization rather than surface abundance alone by integrating proximity-derived associations with tumour-selective expression and biological context. EGFR serves as a stringent reference antigen because its broad surface distribution and therapeutic history provide a suitable test of proximity-informed co-target prioritization.

We applied graph-based machine learning models to directly labelled proximity data from all RTK-anchored micromaps (Supplementary Notes 1–9). We constructed a homogeneous protein graph per experiment using STRING-derived network connectivity together with experimentally measured proximity co-enrichment and protein expression (Fig. 4a). We trained three model architectures, a naive structure-only model (Node2Vec)34, a variational graph autoencoder (VGAE)35 and a graph attention network (GAT)36, to learn latent structure and generate EGFR-associated co-target predictions. Comparison of these model outputs revealed both shared and architecture-specific prioritization patterns (Fig. 4b). To visualize these results, we selected a representative subset of high-confidence predictions, with an emphasis on proteins recovered by VGAE and GAT, spanning well-established EGFR-associated proteins as well as additional proximity-defined candidates of biological interest. VGAE and GAT consistently recovered key EGFR signalling components, including HER2, HER3 and MET (Fig. 4b), supporting the biological validity of their learned representations. These outputs also included non-RTK surface proteins, among which CDCP1 was recovered by both VGAE and GAT as a prominent EGFR-associated prediction. Node2Vec produced noisier predictions and was less efficient at recovering known EGFR-associated relationships, reflecting its dependence on positional graph embeddings without access to the experimentally derived proximity and abundance features that inform GAT and VGAE.

Fig. 4: Graph-based learning identifies EGFR-associated proximal partners and highlights CDCP1.

a, Graph-based machine learning workflow. Heterogeneous graphs integrating STRING-derived connectivity with protein proximity and expression data are collapsed into homogeneous protein graphs to train Node2Vec, VGAE and GAT models for proximal protein prediction and TAA–TAPA prioritization. GCN, graph convolutional network. b, EGFR proximal protein predictions from Node2Vec, VGAE and GAT models. Dot size and colour indicate prediction score. Representative high-confidence candidates were consistently identified across GAT and VGAE models. High-confidence candidates were defined by model-specific prediction score thresholds (GAT ≥ 0.75 or VGAE ≥ 0.5) and recovered in at least 5 independent micromapping experiments. c, Left, GAT feature ablation showing average precision using STRING-derived graph structure alone (Structure), with protein expression (Protein exp), with proximity features (Prox), or with both features (Prox and exp). Right, cross-model performance with all features enabled; n represents the number of predictions per model. d, Multimodal annotation of ten representative high-confidence EGFR-associated proximal partners. Left, STRING high-confidence EGFR interaction status (STRING high conf) and drug availability (Has drug). Middle, log2 tumour-versus-normal expression differential (TCGA/GTEx) across five indications. Right, GTEx normal tissue expression. FC, fold change; TPM, transcripts per million. e, MetaMap-derived Spearman correlation coefficients for EGFR across all profiled proteins. f, Flow cytometry histograms of surface EGFR (blue), CDCP1 (green) and isotype control (grey) in keratinocytes, lung fibroblasts and HCC827 cells. g, EGFR-targeted micromaps showing enrichment of the anchor target EGFR and CDCP1 over isotype control in keratinocytes, lung fibroblasts and HCC827 cells. Box plots show peptide log2-transformed fold change values from EGFR-targeted micromapping experiments (n = 3 technical replicates). Centre lines indicate medians, boxes span the 25th to 75th percentiles and whiskers extend from minima to maxima; all individual data points are shown. h, Volcano plots of EGFR-anchored microenvironment maps in parental and osimertinib-resistant HCC827 cells, and dissociated SW48 xenograft tumour tissue (n = 3 technical replicates; P values from a two-sided moderated t-test with adjustment for multiple comparisons).

Source data

To determine the contributions of each data modality, we performed systematic feature-ablation analyses (Fig. 4c and Supplementary Fig. 16). Removal of either protein proximity or protein expression reduced model precision, whereas integrating both features consistently yielded the highest performance. Of note, proximity-derived information provided independent, non-redundant signal beyond protein expression or graph structure (that is, annotated protein interactions) alone. Training dynamics revealed rapid early learning, with GAT models achieving most performance gains within the first several epochs under a curriculum schedule that progressively introduced more challenging prediction tasks (Supplementary Fig. 17). This early convergence likely reflects two factors: learned attention weighting of edges by predictive relevance, and high signal to noise in the experimentally measured proximity and abundance of the input graphs. Beyond model performance, these proximity-informed outputs can be incorporated into orthogonal analyses to support biological hypothesis generation, as illustrated by the enrichment of proximity-prioritized surface protein pairs for DepMap coessentiality, including potentially novel pairs that are not represented in high-confidence interaction databases (Supplementary Fig. 18).

Building on these predictions, we integrated multiple independent annotation layers to contextualize the high-confidence EGFR-associated candidates in Fig. 4b, including evidence of EGFR association, clinical-stage therapeutic targeting and expression context. In this analysis, CDCP1 emerged as a proximity-defined EGFR-associated non-RTK, in contrast to canonical EGFR network components such as MET, HER2 and HER3 that dominate existing co-targeting strategies (Fig. 4d). CDCP1 has also been explored as a therapeutic target across antibody-based modalities, including ADCs, with limited programs in preclinical or early clinical development37,38, supporting its translational relevance.

Having identified CDCP1 as a high-confidence EGFR-associated candidate, we next sought to determine whether this relationship was consistent across experimental systems, disease contexts and orthogonal biological datasets. First, analysis of MetaMap correlation networks revealed strong and reproducible spatial co-enrichment between EGFR and CDCP1 across independent proximity-derived datasets (Fig. 4e), further supporting CDCP1 as a proximity-defined EGFR-associated protein. To extend this analysis, we generated 38 CDCP1-anchored micromaps spanning 19 tumour cell systems using both proximity-labelling chemistries (Supplementary Fig. 19). Across these diverse cellular contexts, CDCP1 exhibited robust co-enrichment of EGFR, mirroring reciprocal EGFR-anchored maps (Extended Data Fig. 1a,b). This relationship was also detected in independent micromaps generated using directly photocatalyst-conjugated targeting binders (Extended Data Fig. 1c–e). This spatial proximity was attenuated in normal epithelial and mesenchymal cells in which both EGFR and CDCP1 are expressed, indicating context-dependent spatial organization rather than simple co-expression (Fig. 4f,g). Across the broader cell line panel, EGFR and CDCP1 expression varied across cell systems, but did not correlate with proximity enrichment (Supplementary Fig. 20), indicating that the observed proximity relationship was not driven by EGFR or CDCP1 abundance.

Recent EGFR proximity-labelling studies have mapped EGFR-associated membrane neighbourhoods, although CDCP1 was not reported among the prominent EGFR-associated proteins identified in those datasets39,40,41. To determine whether CDCP1 could be detected in these models under our conditions, we performed additional RFT/BT EGFR-targeted micromapping in A549, A431 and H226 cells used in the prior studies. In each dataset, EGFR was significantly enriched as the targeted receptor, and CDCP1 was significantly co-enriched relative to isotype-matched controls (Supplementary Fig. 21), demonstrating its reproducible recovery across all three models. Differences in CDCP1 recovery across studies may reflect variations in EGFR binder format, labelling chemistry, enrichment workflows, mass spectrometry acquisition and analytical thresholds. By integrating diverse tumour cell-model sampling with reciprocal EGFR- and CDCP1-anchored micromaps, directly conjugated binder controls and graph-based contextualization, our approach identified and prioritized a recurrent EGFR–CDCP1 proximity relationship that was not evident in single-anchor or more narrowly sampled EGFR interactome studies.

In addition, the GAT model trained on the full RTK micromapping dataset predicted overlapping EGFR and CDCP1 proximal neighbourhoods that included surface and non-surface proteins, with enrichment of common signalling pathways supporting their participation in shared biological networks (Supplementary Fig. 22). Consistent with this shared organization, EGFR- and CDCP1-anchored micromaps independently enriched a conserved EGFR signalling module comprising GRB2, GAB1, SHC1 and SOS142 (Supplementary Fig. 23a–c). CDCP1 micromaps further highlighted proximity to proteins associated with known CDCP1 roles in tumour progression and tyrosine kinase inhibitor resistance, including SRC family kinases, PRKCδ and PRKCθ43,44,45 (Supplementary Fig. 23d). Beyond these previously characterized interactions, CDCP1 mapping identified additional proximal partners including PTPN1246, INPPL147 and the GTPase regulators ARHGEF548 and ASAP family proteins49 with roles in cell migration, invasion and metastasis (Supplementary Fig. 23d). These observations corroborate prior mechanistic and biochemical evidence that EGFR and CDCP1 associate in tumour cells45,50,51, a finding we further supported by immunoprecipitation of endogenous EGFR and recovery of CDCP1 (Supplementary Fig. 23e). Our proximity-mapping analyses extend these observations by defining EGFR–CDCP1 co-enrichment across diverse tumour contexts and spatial microenvironments.

Analysis of publicly available clinical proteomic datasets revealed increased CDCP1 protein abundance in tumours relative to matched normal tissues across multiple cancer types, a pattern independently supported by internal proteomic analyses (Supplementary Fig. 24). These findings align with prior reports of elevated CDCP1 expression across diverse tumour types, including settings associated with more aggressive disease38,50,52,53.

Prior mechanistic studies have shown that higher CDCP1 expression is associated with poorer progression-free survival in patients with EGFR mutations and increased CDCP1 expression in tyrosine kinase inhibitor-resistant contexts, including following osimertinib treatment44,50,54. To determine whether EGFR–CDCP1 spatial association is preserved under therapeutic pressure, we performed EGFR proximity mapping in parental and osimertinib-resistant HCC827 cellular states and observed that CDCP1 remains consistently enriched (Fig. 4h). In dissociated xenograft-derived tumour cells, EGFR proximity mapping demonstrated continued CDCP1 enrichment (Fig. 4h), indicating preservation of this spatial relationship within native tumour material. Finally, CDCP1 proximity mapping performed in the presence of EGF revealed increased EGFR–CDCP1 spatial association, consistent with prior studies reporting that EGF stimulation promotes CDCP1 upregulation and clustering55,56 (Supplementary Fig. 25). These computational, spatial, mechanistic and clinical analyses establish CDCP1 as a robust EGFR-associated TAPA for therapeutic co-targeting.

These results support a coherent, context-dependent spatial and biological relationship between EGFR and CDCP1 that extends beyond expression alone and provides the foundation for functional evaluation of CDCP1 as an EGFR-associated TAPA.

Functional evaluation of EGFR×CDCP1 co-targeting

Following prioritization of CDCP1 as a TAPA partner for EGFR, we investigated whether co-targeting this pair could enhance tumour-directed cytotoxicity using bispecific ADCs. We reasoned that spatial co-enrichment of EGFR and CDCP1 would increase productive dual engagement by a single bispecific molecule. Because CDCP1 exhibits more restricted normal tissue expression than EGFR (Fig. 4d), bispecific ADCs were designed with asymmetric target engagement, including substantially detuned apparent EGFR binding relative to CDCP1, to favour productive cis-engagement, internalization and payload delivery in cells where both receptors are co-expressed in close spatial proximity. This configuration implements logic-gated co-targeting that biases activity toward tumour cells co-expressing both receptors while minimizing uptake and cytotoxicity in normal cell systems.

We assembled a panel of EGFR×CDCP1 bispecific antibody tool molecules for ADC evaluation and selected a representative construct to assess proximity-based co-targeting (Supplementary Fig. 26). We assessed cellular uptake of the unconjugated EGFR×CDCP1 bispecific antibody across cancer cell lines spanning a range of EGFR and CDCP1 expression levels, alongside matched monovalent EGFR-only and CDCP1-only controls (Fig. 5a and Supplementary Fig. 26). The unconjugated bispecific antibody exhibited greater internalization than either monospecific control (Fig. 5b), consistent with cis-engagement of the two receptors, as supported by NanoBiT complementation and soluble ectodomain competition experiments showing that the bispecific antibody can concurrently engage EGFR and CDCP1 and bring the two receptors into close proximity at the tumour cell surface (Supplementary Fig. 27). These findings are consistent with multivalent engagement of EGFR and CDCP1 promoting enhanced receptor co-engagement and internalization, although receptor trafficking dynamics may also contribute57.

Fig. 5: Functional evaluation of EGFR×CDCP1 co-targeting through ADCs and TCEs.

a, Mechanism of action of a bispecific EGFR×CDCP1 ADC. Target binding induces internalization, lysosomal degradation and cytotoxic payload release. bsAb, bispecific antibody. b, Live-cell imaging of internalization of unconjugated EGFR×CDCP1 bispecific antibody and controls in EGFR and CDCP1-expressing cancer cell lines. Total integrated intensity relative to phase confluence (indicative of internalization) is plotted over time. Data are mean ± s.d. (n = 3 technical replicates). RCU, red calibrated unit. c, Relative viability of cancer cell lines treated with EGFR×CDCP1 ADC or matched DAR controls (Supplementary Fig. 26) measured by CellTiter-Glo and normalized to untreated cells. Data are mean ± s.d. (n = 4 technical replicates). d, Study design for SW48 xenograft. TV, tumour volume. e, Tumour growth of SW48 xenografts following treatment with indicated test articles. Data are mean tumour volume ± s.e.m. (n = 8 mice per group); arrows denote dosing day. Two-way analysis of variance (ANOVA) with Tukey’s multiple comparison; ***P = 0.0002. f, Study design for SW900 xenograft. g, Tumour growth of SW900 xenografts following treatment with indicated test articles. Dummy×Dummy indicates a non-targeting control (anti-Campylobacter jejuni and anti-respiratory syncytial virus (RSV)) with equivalent structure and molecular mass to the bispecific ADC. Data are mean tumour volume ± s.e.m. (n = 9 mice per group); arrows denote dosing day. Two-way ANOVA with Tukey’s multiple comparison; ****P < 0.0001. h, Mechanism of action of a trispecific EGFR×CDCP1×CD3 TCE. Bridging antigen-expressing tumour cells and CD3+ T cells induces cytotoxicity. i, T cell-dependent cellular cytotoxicity of trispecific EGFR×CDCP1×CD3 TCEs and controls. Indicated cancer cell lines were labelled with luciferase and incubated with peripheral blood mononuclear cells and antibodies. Viability was measured by luciferase signal. Data are mean ± s.d. (n = 3 technical replicates).

Source data

We next evaluated internalization of the EGFR×CDCP1 bispecific antibody in primary cervical and bronchial/tracheal epithelial cells expressing EGFR and CDCP1 at levels resembling those in tumour cells (Supplementary Fig. 28a). In contrast to cancer cells, minimal internalization was observed in primary cells (Supplementary Fig. 28b), indicating that efficient uptake depends on spatial colocalization rather than expression alone.

To assess cytotoxicity, antibodies were conjugated with valine-citrulline-linked monomethyl auristatin E (vcMMAE) using two ADC formats: a site-specific engineered cysteine conjugate at a drug-to-antibody ratio (DAR) of 2 (DAR2) or an interchain disulfide conjugate at a higher DAR (DAR ≈ 3) (Supplementary Figs. 26 and 29). In vitro and standard xenograft studies were performed using the site-specific DAR2 format. In vitro, the EGFR×CDCP1–MMAE (DAR2) ADC induced dose-dependent killing across multiple cancer cell lines and demonstrated superior potency relative to EGFR-only-MMAE (DAR2) or CDCP1-only-MMAE (DAR2) monovalent ADC controls (Fig. 5c). The increased potency of the bispecific ADC is consistent with its greater cellular uptake and supports a contribution from multivalent receptor engagement. Minimal cytotoxicity was observed in primary epithelial cells treated with either the bispecific or monospecific ADCs, consistent with their reduced internalization in these systems (Supplementary Fig. 28c). The bispecific ADC therefore showed greater potency in tumour cells than in primary epithelial cells, including a more than 500-fold shift in the most sensitive model.

We next investigated the in vivo activity of the EGFR×CDCP1 bispecific ADC in SW48 colorectal and BxPC3 pancreatic cancer xenograft models (Fig. 5d and Extended Data Figs. 2a and 3a,d). In the SW48 model, a single 3 mg kg−1 dose of the bispecific ADC produced more than 100% tumour growth inhibition (TGI) relative to the Fc-only control, comparable to a molar-matched EGFR-only control ADC (Fig. 5e). At 1 mg kg−1, the bispecific ADC maintained complete TGI, whereas the EGFR-only ADC produced only 43% TGI (Fig. 5e). The EGFR×CDCP1 ADC also improved overall survival without observable body weight loss (Extended Data Fig. 3b,c). Similar activity was observed in BxPC3 xenografts, where the bispecific ADC matched EGFR-only ADC activity at 3 mg kg−1 and, after re-dosing, further inhibited tumour growth and improved overall survival without observable body weight loss (Extended Data Figs. 2b and 3e,f).

To test whether in vivo activity depends on dual-antigen engagement, we implemented a dual-flank SW48 xenograft model in which mice carried parental tumours on one flank and matched CDCP1-knockdown (CDCP1-KD) tumours on the contralateral flank (Extended Data Fig. 2c and Supplementary Fig. 30). For this side-by-side comparison within the same mice, we used an interchain disulfide EGFR×CDCP1–MMAE ADC (DAR ≈ 3) as a functional tool reagent, with the same binding arms as the site-specific DAR2 construct (Supplementary Fig. 29). Mice received a single intravenous dose of EGFR×CDCP1–MMAE (DAR ≈ 3) or Fc-only-MMAE (DAR ≈ 3) control (3 mg kg−1 equivalent) using the dosing schedule shown (Extended Data Fig. 2c). The bispecific ADC produced marked suppression of parental tumour growth (97% TGI), whereas antitumour activity was substantially reduced in the CDCP1-KD flank (46% TGI; Extended Data Fig. 2d). By contrast, the Fc-only control ADC, which has an equivalent pharmacokinetic profile to the bispecific ADC (Supplementary Fig. 30), showed comparable tumour progression in both parental and CDCP1-KD tumours. These data demonstrate that antitumour activity depends on CDCP1 expression, supporting a dual-antigen-biased delivery mechanism rather than EGFR targeting alone. To compare coordinated versus independent targeting, we evaluated the EGFR×CDCP1 bispecific ADC alongside combined EGFR- and CDCP1-directed monospecific ADCs in an SW900 lung cancer xenograft model (Fig. 5f and Extended Data Fig. 3g). The bispecific ADC showed enhanced TGI and survival relative to the combination, consistent with a functional advantage of proximity-guided co-engagement (Fig. 5g and Extended Data Fig. 3h,i).

To assess whether the functional benefit of EGFR×CDCP1 could be extended across therapeutic modalities, we evaluated this pair in a trispecific TCE format using alternative EGFR- and CDCP1-binding arms combined with a CD3-binding arm in a tool molecule (Fig. 5h and Supplementary Fig. 31). We hypothesized that tumour-specific spatial proximity between EGFR and CDCP1 would promote more effective immune synapse formation when both antigens are co-engaged on the same cell. In T cell-dependent cellular cytotoxicity assays, the EGFR×CDCP1×CD3 trispecific TCE elicited greater tumour cell killing relative to EGFR×CD3 or CDCP1×CD3 controls across all three tumour cell lines tested, achieving more than 100-fold increased potency in one model (Fig. 5i). This enhanced activity is consistent with multivalent engagement of spatially proximal antigens to promote target-cell engagement, immune synapse formation and cytotoxic activity.

These studies demonstrate that TAPA-guided co-targeting of EGFR and CDCP1 translates across multiple therapeutic modalities, enabling logic-gated tumour selectivity and improved antitumour activity. Together with the computational and spatial analyses presented above, these functional data establish EGFR×CDCP1 as a therapeutically actionable TAA–TAPA pair and have informed the advancement of an EGFR×CDCP1 bispecific ADC towards clinical evaluation.

Discussion

This study demonstrates that scalable surface microenvironment mapping can resolve disease-relevant spatial organization on the tumour cell surface. By integrating proximity information with complementary molecular features, we define TAPAs, a class of co-targets distinguished by spatial context rather than expression alone. In this paradigm, therapeutic relevance emerges not only from target expression, but also from organization within disease-associated membrane neighbourhoods.

As a proof of concept, we identified CDCP1 as a TAPA that is spatially co-enriched with EGFR across multiple cancer cell systems, consistent with prior reports that EGFR and CDCP1 can associate in tumour contexts45,51, and evaluated EGFR×CDCP1 as a proximity-defined co-target pair. Although EGFR is broadly expressed in normal tissues, its spatial coupling with CDCP1, an antigen with more restricted normal tissue distribution (Fig. 4d), revealed a co-targeting axis that is not apparent from expression-based analyses alone. Dual engagement using bispecific ADCs and trispecific TCEs produced greater tumour cell killing and in vivo tumour control relative to single-antigen targeting, supporting spatially informed co-targeting as a strategy to improve therapeutic index. These functional effects should also be interpreted in the context of avidity, as bispecific formats increase effective binding through multivalent engagement. Accordingly, enhanced internalization and cytotoxicity are likely to reflect productive dual receptor engagement shaped by spatial organization, multivalent binding and receptor trafficking. Compared with administering separate ADCs, this integrated bispecific format may simplify dose optimization, coordinate target exposure and reduce overlapping toxicities.

The graph-based modelling framework deployed here captures higher-order relationships, including co-expression and indirect associations propagated through the network structure. Consistent with the underlying STRING-derived network, which integrates evidence from both direct physical interactions and broader functional associations (for example, co-expression, genomic context, experimental data and text mining), this approach is not intended to resolve direct pairwise interactions; rather, it identifies proteins that reside within the same spatially defined local network or membrane neighbourhood. More broadly, graph-based integration extends individual proximity measurements to infer higher-order spatial relationships among proteins that are not directly measured as pairs. We envision that the resulting maps and inferred relationships will provide resolution of surface protein organization that complements ongoing work in spatial transcriptomics, functional genomics, cell- and tissue-contextualized protein representation learning58, and emerging enzyme-based cell surface interactome mapping59.

These findings also position spatial co-enrichment as a biologically grounded organizing principle for next-generation multispecific therapeutic design. The proximity-mapping workflow is modular with respect to labelling chemistry and target class. We utilized complementary diazirine- and tyramide-based chemistries, and alternative proximity-labelling strategies could be used to tune labelling scale and spatial resolution. Although we focused on RTKs, the approach is broadly applicable across disease areas, including inflammatory, metabolic and neurological contexts, and to other classes of cell surface proteins, including G-protein coupled receptors, integrins, proteoglycans and proteins with non-canonical surface localization60, provided suitable targeting reagents are available. Differences in epitope accessibility, antigen density, labelling chemistry, experimental workflow and mass spectrometry acquisition may influence proximity detection and contribute to variation across studies. Future applications across additional target classes, tissues and disease settings will help define the extent to which disease-associated proximity relationships are conserved, context-dependent and therapeutically actionable.

Several limitations warrant discussion. Although this effort sampled multiple RTKs across diverse tumour cell systems, it does not fully resolve which spatial relationships are uniquely enriched in specific disease contexts versus those that reflect conserved membrane organization. Accordingly, TAPAs should be viewed as proximity-defined antigens whose enrichment and functional relevance vary across tumour lineage, genetic state and treatment context rather than as universally tumour-specific features. In addition, proximity labelling captures interactions within a finite spatial radius and temporal window, limiting resolution of fine-grained membrane organization and dynamic changes associated with receptor activation or trafficking.

Finally, because much of the mapping was performed in tumour cell systems that lack full microenvironmental complexity, stromal and immune interactions are incompletely represented. Extending this approach to matched normal tissues, treatment-altered states and primary tumours will enable comparisons of tumour-restricted versus conserved spatial architectures beyond those that are possible using expression- or genetics-based approaches alone. Such comparisons will require careful benchmarking given the lack of consensus across existing protein interaction databases and the absence of a unified ground truth.

Our integrated experimental and computational strategy provides a foundation for systematic co-target prioritization and a systems-level view of cell surface organization. Beyond oncology, TAPA discovery offers a generalizable strategy for translating disease-associated spatial relationships into proximity-guided therapeutic design that is applicable across diverse indications and therapeutic modalities. More broadly, just as expression profiling transformed target discovery, systematic mapping of protein proximity may define an additional organizational layer for understanding disease biology and guiding the design of next-generation multispecific therapeutics.

Methods

General cell culture

The following cancer cell lines were purchased from ATCC: A431 (CRL-1555), A549 (CRM-CCL-185), AGS (CRL-1739), BT549 (HTB-122), BXPC3 (CRL-1687), CACO2 (HTB-37), CAPAN2 (HTB-80), CAOV3 (HTB-75), HCC1954 (CRL-2338), HCC4006 (CRL-2871), HCC827 (CRL-2868), HCT116 (CCL-247), HPAC (CRL-2119), HS578T (HTB-126), HT29 (HTB-38), KatoIII (HTB-103), LS123 (CCL-255), MDAMB468 (HTB-132), NCIH1563 (CRL-5875), NCIH1650 (CRL-5883), NCIH1975 (CRL-5908), NCIH226 (CRL-5826), NCIH358 (CRL-5807), NCIH441 (HTB-174), NCIN87 (CRL-5822), PC3 (CRL-1435), SKBR3 (HTB-30), SKMES1 (HTB-58), SW48 (CCL-231) and SW900 (HTB-59). The following cells were purchased from Accegen: CALU1 (ABC-TC0110), EBC1 (ABC-TC0170) and MKN7 (ABC-TC0688).

The media for all cancer cells were purchased from ATCC and include RPMI-1640 (30-2001), DMEM (30-2002), EMEM (30-2003), F-12K (30-2004), IMDM (30-2005) and McCoy’s 5A (30-2007). All media were supplemented with FBS (10% or 20% final concentration as according to cell line supplier’s recommendation, Thermo Scientific, 10082147) and penicillin/streptomycin (1% final concentration from a 100× stock, Thermo Scientific, 15140163). BT549 cells were supplemented with insulin (Thermo Scientific, 12585-014) at 0.023 U ml−1. HS578T was supplemented with insulin at 0.01 mg ml−1. Cell culture media were filter-sterilized via 0.2 µm Nalgene Rapid-Flow Sterile Disposable Filter Units with PES Membrane, 500 ml capacity (Thermo Scientific, 566-0020) or 1,000 ml capacity (Thermo Scientific, 567-0020). Cells were grown in manufacturer’s recommended media with the following exceptions: MDA-MB-468 cells were grown in DMEM, HPAC in IMDM, SW48 in McCoy’s 5A and SW900 in RPMI-containing medium.

Primary bronchial/tracheal epithelial cells (PCS-300-010), epidermal keratinocytes (PCS-200-011) and cervical epithelial cells (PCS-480-011) were purchased from ATCC. Primary bronchial/tracheal epithelial cells were grown in airway epithelial cell basal medium (ATCC, PCS-300-030) supplemented with a bronchial epithelial cell growth kit (ATCC, PCS-300-040). Primary epidermal keratinocytes were grown in dermal cell basal medium (ATCC, PCS-200-030) supplemented with keratinocyte growth kit (ATCC, PCS-200-040). Primary cervical epithelial cells were grown in cervical epithelial cell basal medium (ATCC, PCS-480-032) supplemented with cervical epithelial growth kit (ATCC, PCS-480-042).

All cells were grown at 37 °C with 5% CO2 in 75 cm2 (Corning, 430641U) or 150 cm2 (Corning, 430825) vented cap sterile cell culture flasks or 15 cm plates (Thermo Scientific, 150468).

Method for generating the osimertinib-resistant HCC827 cells

HCC827 cells were cultured in 500 nM osimertinib. The medium and osimertinib were replaced every 3–4 days. Once the cells began to double in 500 nM osimertinib (around day 35) the concentration was increased to 1 µM. Cells were allowed to expand and then frozen back. Resistance to osimertinib was confirmed in a dose–response assay.

General synthetic information

Riboflavin tetraacetate (RFT) photocatalyst utilized in these studies (Supplementary Fig. 7) was synthesized as previously described14. The iridium photocatalyst and biotin-containing diazirine probe (Ar-PEG3-Biotin) utilized in these studies (Supplementary Fig. 7) was synthesized as previously described13.

Preparation of secondary antibody–photocatalyst conjugate

A 350 µl aliquot of polyclonal goat anti-mouse IgG (Millipore, AP124) or polyclonal goat anti-rabbit IgG (Millipore, AP132) was combined with 50 µl of 1 M sodium bicarbonate buffer (pH 8.5, Thermo Scientific, J60408.AK) in a low protein binding tube. Four microlitres of 100 mM azidobutyric acid NHS ester (prepared in DMSO, Broadpharm, BP-22526) was added and the reaction mixture was incubated for 1.5 h at room temperature in the dark. After 1.5 h, an additional 4 µl of 100 mM azidobutyric acid NHS ester linker was added, and the sample was incubated for 1.5 h at room temperature in the dark. In the meantime, a Zeba Spin desalting column (2 ml column, 40,000 MWCO, Thermo Scientific, 87769 or A57762) was prepared by first removing the storage solution (centrifuged 2,000g for 3 min at 4 °C). The column was then primed by three washes of 1 ml 50 mM Tris pH 7.5 and spun at 2,000g for 5 min at 4 °C after each addition except for the last spin which was extended to 10 min. After the second incubation of antibody and azidobutyric acid NHS ester, the sample was buffer exchanged into 50 mM Tris pH 7.5 using the Tris primed Zeba Spin desalting column and spun at 2,000g for 2 min at 4 °C. The antibody–azide conjugate was transferred to a new tube for click chemistry azide-alkyne cycloaddition of the photocatalyst using the Click-iT Protein Reaction Buffer Kit (Thermo Scientific, C10276). Fifteen microlitres of photocatalyst (from 5 mM stock in DMSO) was added to the antibody–azide conjugate and mixed. Fifteen microlitres of copper sulfate and 15 µl of additive 1 were added and mixed, and the reaction was incubated for 1.5 min at room temperature. Following incubation, 30 µl of additive 2 was added to the reaction mixture and incubated in the dark for 30 min at room temperature. After, the sample was buffer exchanged into DPBS (Thermo Scientific, 14190144) using a Zeba Spin desalting column primed in the same manner as above except with DPBS. In some instances, the conjugated antibodies were spun at 16,000g for 10 min at 4 °C to remove precipitates. The final protein concentration of the secondary antibody–photocatalyst conjugate was determined using the Pierce BCA Protein Assay kit (Thermo Scientific, 23227) according to the manufacturer’s instructions. The photocatalyst concentration was determined by measuring the absorbance (350 nm for Ir and 450 nm for RFT) and compared to a standard curve consisting of known free photocatalyst concentrations. The micromolar concentration of photocatalyst was divided by the micromolar concentration of antibody to determine the antibody:photocatalyst ratio. A ratio of 1:8 was routinely obtained.

Preparation of direct antibody–photocatalyst conjugates

CDCP1 and EGFR-targeted micromapping using directly conjugated antibody–photocatalyst conjugates was performed using EGFR×Fc and CDCP1×Fc binders as well as a non-binding anti-RSV×Fc control (Supplementary Fig. 32). All antibodies were produced in ExpiCHO-S cells (Thermo Scientific, A29127) in ExpiCHO expression medium (Thermo Scientific, A2910001). These were purified using an AKTA pure with Unicorn v7.10 SP1 software. All samples were first purified using PrismA HiTrap affinity capture (Cytiva, 17549854), neutralized with 1 M NaPi pH 7.0 and followed by size-exclusion chromatography (SEC) purification over a Superdex 200 increase (Cytiva, 28990944) in PBS (Fisher Scientific, SH3025602). Four antibodies were produced in-house; an anti-EGFR (cetuximab, sequence obtained from US patent 7,060,80861), anti-CDCP1 clone 41A9 (sequence obtained from US patent 11,702,48162) anti-RSV (sequence obtained from https://go.drugbank.com/drugs/DB00110) and an antibody fragment referred to as ‘Fc’ which is comprised of a human IgG1 antibody sequence from just above the hinge (E215) to the C terminus (G446). Each of these carried mutations in the CH3 domain specific for conducting Fab arm exchange63. The Fc antibody was directly conjugated with azide by buffer exchanging antibody into PBS using Zeba desalting columns. 10% v/v of 7.5% NaHCO3 was added to adjust the pH (Hyclone, sh30033.01). 3-Azidopropanoic acid NHS ester (Broadpharm, BP-23800) was prepared to 50 mM stock solution in DMSO. 20 molar equivalents was used as a challenge ratio and added to the solution for 90 min on a shaker platform protected from light. Sample was buffer exchanged into PBS using Zeba desalting columns to remove excess azide. Fab arm exchange was then performed on the azide conjugated Fc with each of the targeting arms (EGFR, CDCP1 and RSV). Equimolar amounts of each antibody were added to the reaction and adjusted to a 25 mM concentration of 2-mercaptoethylamine-HCl (2-MEA, Sigma-Aldrich, 30078). The samples were incubated at 37 °C for 2 h, then buffer exchanged with PBS to remove 2-MEA. This azide labelling of the Fc approach allows for high-throughput production of azide conjugated bispecifics while ensuring there is no disruption to the binding sites on the targeting arm.

Two hundred microlitres of 10 µM targeting arm×Fc-azide antibody was covalently conjugated to the photocatalyst using the Click-iTTM Protein Reaction Buffer Kit and purified using a Zeba Spin desalting column as described above. An isotype control (anti-RSV) was prepared in parallel using the same protocol. The final protein concentration of the direct antibody–photocatalyst conjugate was determined using the Pierce BCA Protein Assay Kit - Reducing Agent Compatible (Thermo Scientific, 23250) according to the manufacturer’s instructions. The photocatalyst concentration was determined as described above. Ratios of 1:9 and 1:5 were routinely observed for Ir and RFT, respectively.

Flow cytometry analysis of target proteins

Cells were collected with Accutase (BioLegend, 423201), centrifuged at 800g for 5 min at 4 °C, resuspended in DPBS, counted using a Beckman Coulter Vi-CELL XR Cell Viability Analyzer and aliquoted at 200,000 cells per well in a 96-well U-bottom plate (Fisher, 351177) in 100 µl DPBS. One hundred microlitres of DPBS containing LIVE/DEAD Fixable Near-IR Dead Cell Stain (Thermo Scientific, L10119) at a 1:1,000 dilution was added to the cells and then incubated for 30 min at room temperature in the dark. The cells were centrifuged at 800g for 5 min at 4 °C and washed once with 200 µl fluorescence-activated cell sorting (FACS) buffer (DPBS containing 2% FBS and 2 mM EDTA). Cells were resuspended in 100 µl containing 25 nM primary antibody (see ‘List of primary antibodies used for microenvironment mapping’) and incubated for 30 min at 4 °C. One hundred microlitres of FACS buffer was added and the cells were centrifuged to remove the primary. The cells were then washed once in 200 µl FACS buffer before resuspension in 100 µl of FACS buffer containing a 1:250 dilution of PE secondary (goat anti-mouse BioLegend 405307 or donkey anti-rabbit BioLegend 406421) for 30 min at 4 °C. One hundred microlitres of FACS buffer was added and the cells were centrifuged to remove the secondary. The cells were then washed once in 200 µl FACS buffer before resuspension in 50 µl fixation buffer (BioLegend, 420801) for 30 min at room temperature in the dark. FACS buffer (150 µl) was added and the cells were centrifuged. The cells were then washed once in 200 µl FACS buffer before resuspension in 150 µl FACS buffer. Samples were then run on a BD LSRFortessa X-20 with BD FACSDiva software v9.2 or kept at 4 °C overnight for next day analysis. Data were analysed using FlowJo, v10 (Supplementary Fig. 5c).

Flow cytometry competitive binding assay

NCI-H1975 cells were collected with Accutase, centrifuged at 800g for 5 min at 4 °C, resuspended in DPBS, counted using a Vi-CELL XR Cell Viability Analyzer and aliquoted at 2 million cells each to low protein binding tubes. Cells were spun down and resuspended in 500 µl cold DPBS. Five micrograms of primary antibody was next added: for EGFR, cetuximab antibody (prepared as described above, obtained from US patent 7,060,80861) or human isotype antibody (BioLegend, 403502) was used; for CDCP1, CDCP1 antibody (BioLegend, 324002) or mouse isotype antibody (BD, 556648) was used. For the competitive sample, 25 µg competing antibody was added immediately before the primary. In the case of EGFR competition, EGFR (BD, 555996) was used; for CDCP1 competition, CDCP1 41A9 (prepared as described above, obtained from US patent 11,702,48162) was used. Samples were incubated for 1.5 h at 4 °C on a rotisserie. Cells were spun down at 800g for 5 min at 4 °C and washed once with 1 ml cold DPBS before resuspension in 500 µl cold DPBS containing Zombie Green Viability dye (1:500 dilution from stock prepared according to manufacturer’s instructions, BioLegend, 423111). Secondary antibody specific for the primary, but not the competing antibody was added—for EGFR competition, goat anti-Human APC (1:100 dilution, R&D Systems, F0135); for CDCP1 competition, goat anti-Mouse APC (1:500 dilution, Thermo Scientific, A865). Samples were incubated for 30 min at 4 °C on a rotisserie. Cells were spun down at 800g for 5 min at 4 °C and washed once with 1 ml cold DPBS before resuspension in 500 µl cold DPBS and transferred to FACS tubes. Samples were then run on a BD Accuri C6 Plus with CSampler Plus. Data were acquired with the BD Accuri CSampler Plus software v1.0.34.1 and analysed using FlowJo, v10 using the previously described gating strategy except the Viability gate utilized the FITC-A channel and the primary antibody was detected with the APC-A channel (Supplementary Fig. 33).

Micromapping of surface proteins in live cells for LC–MS/MS analysis

For two antibody-based micromapping, cells were collected using Accutase and centrifuged at 800g for 5 min at 4 °C before being resuspended in DPBS and counted using a Vi-CELL XR Cell Viability Analyzer. For each condition, 10–20 million cells were aliquoted into each low protein binding tube, with experiments performed in triplicate as independently processed technical replicates (n = 3). The cells were spun down and resuspended in 1 ml cold DPBS containing a targeting primary antibody or an appropriate isotype control (see Supplementary Fig. 34 for a list of primary antibodies used for micromapping) at 1 µg antibody for every million cells and incubated at 4 °C for 30 min on a rotisserie (for example, 10 million cells were resuspended in 1 ml cold DPBS and 10 µg of primary antibody was added). After the primary antibody incubation, cells were spun down and washed twice with 1 ml cold DPBS. The cells were then resuspended in 1 ml cold DPBS containing the secondary antibody conjugated to photocatalyst at 1 µg antibody for every million cells and incubated at 4 °C for 30 min on a rotisserie. Cells were then spun down and washed twice with 1 ml cold DPBS. The cells were resuspended in 1 ml cold DPBS containing 250 µM biotin probe (biotin-tyramide from ApexBio, A80111000). The samples were put in a bio-photoreactor64 for 2 min (for RFT photocatalyst samples) or 3 min (for Ir photocatalyst samples) and irradiated at full intensity. The samples were then spun down and washed twice with 1 ml cold DPBS. The cells were then lysed in one of the two following methods: (1) 1 ml membrane permeabilization buffer (MEM-PER Plus Membrane Fractionation Kit, Thermo Scientific, 89842) with 1× protease inhibitor tablet (Sigma-Aldrich, 4693159001). The samples were then allowed to rotate for 20 min at 4 °C before centrifugation at 16,000g for 15 min at 4 °C. The supernatant was discarded and the membrane pellet was resuspended in 300 µl RIPA (Thermo Scientific, 89901) with 1% sodium dodecyl sulfate (SDS, prepared from a 20% stock, Quality Biological, 351-066-101) and 1× protease inhibitor tablet, sonicated and boiled at 95 °C for 5 min. 1 ml RIPA buffer was added and the lysate was sonicated again to homogenize. (2) Alternatively, the cells were lysed in 1 ml RIPA buffer containing 1× protease inhibitor tablet and 1:1,000 dilution of benzonase (Sigma-Aldrich, 70664-3) and incubated for 15 min at 4 °C on a rotisserie. After lysis, the protein concentrations were measured by BCA assay and stored at −80 °C until the bead enrichment.

Micromapping with directly conjugated photocatalyst antibody was carried out as described above with the following changes. Instead of primary and secondary antibody, an antibody directly conjugated with photocatalyst was added at a molar amount corresponding to 1 µg of full-length antibody for every 1 million cells. A directly photocatalyst-conjugated anti-RSV antibody was used as the isotype control. Lysis was carried out only using the RIPA buffer containing 1× protease inhibitor tablet and 1:1,000 dilution of benzonase.

Following labelling, biotinylated proteins were enriched using streptavidin magnetic beads using either a manual or automated (KingFisher Apex, Thermo Scientific) workflow. In all cases, enrichment followed a common binding and wash procedure, with samples processed either as bead-retained material or by competitive biotin elution, as specified below.

For bead enrichment, 100 or 250 µl of streptavidin magnetic beads (Thermo Scientific, 88817) were washed twice with 1 ml RIPA. Equal protein amounts were added to the beads and the differences in volume were made up with RIPA buffer. The tubes were allowed to rotate 3 h at room temperature. The beads were collected on a magnetic rack and washed three times with 1 ml DPBS containing 1% SDS. Next, the beads were washed three times with 1 ml DPBS containing 1 M NaCl (prepared from a 5 M stock, Research Products International, S24600-500.0) followed by three washes with 1 ml DPBS containing 10% ethanol (200 proof stock, Fisher, BP2818-500) and then one wash with RIPA buffer.

After washing, enriched proteins were processed using one of two workflows. For on-bead digestion, the beads were carried forward directly for downstream proteomic processing. For direct protein elution, proteins were released from the beads by boiling at 95 °C for 10 min in 4× Laemmli buffer (Bio-Rad, 1610747) supplemented with 20 mM DTT (RPI, D11000-10.0) and 25 mM biotin (Sigma-Aldrich, 14400-1G).

Automated bead enrichment was performed using a KingFisher system with the same binding and wash sequence, with wash volumes reduced to 900 µl. In this format, the final RIPA wash was replaced by release of beads into 180 µl of 0.2 M HEPES buffer (pH 8.5; prepared from 0.5 M stock, Thermo Scientific, J63218.AK). Plates were sealed and samples stored at −80 °C prior to proteomic analysis.

Protein extraction and digestion for LC–MS/MS analysis

For direct protein elution, proteins released from the beads were precipitated with trichloroacetic acid and washed with ice-cold acetone. Dried protein pellets were reduced and alkylated by resuspending in 4 M urea with 5 mM tris(2-carboxyethyl) phosphine (TCEP) and 20 mM chloroacetamide (CAA) and incubating at 20 °C for 20 min. Proteins were digested by adding an equal volume of 100 mM Tris-Cl, pH 8.5 with 200 ng lysyl endopeptidase (lysC, FUJIFILM Wako, 125-05061) and shaking overnight at 30 °C then adding an equal volume of 50 mM Tris-Cl, pH 8.5, with 200 ng trypsin (Promega, V5111) and shaking for 6 h at 37 °C. For a subset of micromapping experiments, directly eluted proteins were processed using a previously described method13. For on-bead digestion, bead-enriched lysates were reduced and alkylated on the beads with 5 mM TCEP and 20 mM CAA for 20 min. The beads were then washed with 0.2 M HEPES, pH 8.5 buffer before a 4-h digestion with lysC and overnight digestion with trypsin as above. Digested peptides were pipetted off the beads and one 0.2 M HEPES wash of the beads was performed and added to the digested peptides to recover any remaining peptides from the beads. This protocol was initially performed in microcentrifuge tubes using a magnet rack and manual pipetting and bead mixing. It was later adapted to a 96-well automated workflow using the KingFisher system. For the automated workflow, the separate digestions with lysC and trypsin were combined into a single four-hour step performed on the KingFisher with mixing at 37 °C, with no noticeable effect on digestion efficiency. Following digestion, peptides were labelled with 250 µg tandem mass tag (TMTpro; Thermo Scientific, A52045) isobaric reagents for 2 h at room temperature. Labelling efficiency was checked by pooling 5 µl from each sample within a single plex. For the full experiment, all samples were quenched with hydroxylamine (0.5%) and pooled. Mixes were acidified with TFA (2%) and desalted with SPE-C18 columns (Waters Sep-Pak). Desalted TMT mixes were dried down in a SpeedVac concentrator (Thermo Scientific) and fractionated using the high pH reverse-phase peptide fractionation kit (Pierce) into 13 fractions (5%-50% acetonitrile in 0.1% triethylamine). Every fourth fraction was combined to make 4 pools (1-5-9-13, 2-6-10, 3-7-11, 4-8-12) which were then dried down in a SpeedVac and resuspended in 5% formic acid for LC–MS/MS analysis.

Liquid chromatography

All mass spectrometry experiments used the same liquid chromatography configuration. Peptides were separated on a Vanquish Neo UHPLC system (Thermo Fisher Scientific) operating in direct injection mode at a flow rate of 350 nl min−1 through a 25 cm × 75 µm Aurora Ultimate C18 capillary column (IonOpticks) maintained at 60 °C using a PRSO-V2 column oven (Sonation). Mobile phase A consisted of 0.1% formic acid in water and mobile phase B consisted of 0.1% formic acid in 80% acetonitrile. All data were acquired using an Orbitrap Eclipse Tribrid mass spectrometer (Thermo Fisher Scientific) running Xcalibur v4.0.4084.22, equipped with a nanospray ionization source operating in positive ion mode (spray voltage 2,400 V; ion transfer tube temperature 275 °C).

TMT–SPS–MS3 acquisition of micromap samples

Labelled peptides were separated using an 84-min gradient: 8% to 28% B over 76 min, 28% to 38% B over 5 min, and 38% to 55% B over 3 min, followed by a column wash at 100% B. The mass spectrometer was operated in data-dependent acquisition mode using a 3-s cycle time. Full MS1 survey scans were acquired in the Orbitrap (resolution = 120,000; scan range = 400–1,600 m/z; AGC target = 4 × 105; maximum injection time = 246 ms). Precursor ions were selected from 425–1,600 m/z for fragmentation, restricted to charge states 2–5, with a minimum intensity threshold of 5,000 counts. Dynamic exclusion was applied for 40 s after a single observation using a ±10 ppm mass tolerance window, excluding isotope peaks and limiting selection to one charge state per precursor within each cycle.

Selected precursors were isolated with a 1.0 m/z quadrupole window and fragmented by collision-induced dissociation (CID; NCE = 32%; activation time = 10 ms; activation Q = 0.25) in the linear ion trap (rapid scan rate; AGC target = 2 × 104; maximum injection time = 50 ms). MS2 spectra were searched in real time using the instrument-embedded Real-Time Search65 (RTS) module against a reviewed human protein sequence database with common contaminants. The RTS search used fully tryptic digestion (Trypsin/P) with up to one missed cleavage, static modifications of TMTpro (+304.207 Da) on lysine residues and peptide N termini and carbamidomethylation (+57.021 Da) on cysteine, and a variable modification of oxidation (+15.995 Da) on methionine, with a maximum search time of 40 ms per spectrum. Peptides were accepted for MS3 triggering using charge state-dependent XCorr thresholds (z = 2: XCorr ≥ 1.5, dCn ≥ 0.05; z = 3: XCorr ≥ 2.0, dCn ≥ 0.05; z = 4: XCorr ≥ 2.5, dCn ≥ 0.10; z = 5: XCorr ≥ 3.0, dCn ≥ 0.10) within 10 ppm precursor mass tolerance.

Confidently identified precursors triggered synchronous precursor selection (SPS) MS3 scans, in which up to 8 MS2 product ions were co-isolated using multi-notch quadrupole isolation66 (MS2 isolation window = 2.0 m/z; MS3 quadrupole isolation window = 1.3 m/z) and fragmented by HCD (NCE = 45%). TMTpro reporter ions were detected in the Orbitrap (resolution = 50,000; scan range = 110–400 m/z; AGC target = 5 × 105; maximum injection time = 1,000 ms).

TMT–SPS–MS3 data analysis of micromap samples

Raw files were processed using Thermo Proteome Discoverer (v3.0.1). Peptide and protein identification used SEQUEST HT with INFERYS rescoring and Percolator validation against the canonical human proteome (UniProtKB UP000005640, release 2024_08; 20,420 entries) supplemented with a common contaminants database and an equal-size concatenated reversed-sequence decoy database. Searches required fully tryptic cleavage (Trypsin/P) with up to two missed cleavages, a precursor mass tolerance of 20 ppm, and a fragment ion tolerance of 1.005 Da. TMTpro (+304.207 Da) was set as a static modification on lysine residues and peptide N termini; carbamidomethylation (+57.021 Da) on cysteine was set as a static modification; and oxidation (+15.995 Da) on methionine was set as a variable modification. PSM confidence was assessed by Percolator using INFERYS-derived spectral prediction features in a semi-supervised SVM framework, and results were filtered to 1% FDR (strict) and 2% FDR (relaxed) at both the peptide and protein levels. Protein grouping was performed using the strict parsimony principle.

TMTpro reporter ion quantification was performed on Orbitrap MS3 spectra using signal-to-noise (S/N) ratios integrated within a 15 ppm window using the most confident centroid method. Quantification spectra were required to meet the following quality thresholds: average reporter ion S/N ≥ 15, precursor co-isolation interference <50%, and SPS ion mass matches ≥65%. Only TMTpro-labelled peptides were used for quantification; unique and razor peptides were included. A minimum channel occupancy of 25% was required per quantification spectrum. Isotopic impurity correction was applied using lot-specific correction factors supplied by the manufacturer. Protein abundances were calculated by summing peptide-level S/N values and normalized to total peptide amount within each plex.

Experiments were multiplexed using TMTpro 18-plex reagents. Each individual plex comprised between 6 and 18 samples — one isotype control condition and one to five antibody-targeted conditions, with three technical replicates per condition. Quantitative comparisons were made between each targeted condition and the isotype control within the same plex.

Tables containing RTK micromap enrichment values (log2 fold change and associated statistical metrics) are provided in Supplementary Table 1.

Sample preparation for whole-cell and patient tumour proteomic analysis

Whole-cell extracts were generated for cells listed in Supplementary Fig. 4a. Three technical replicates of 5 × 105 cells were lysed in parallel in 96-well plates. Before aliquoting, washed cell pellets were resuspended in ice-cold PBS at 1.1 × 107 cells per ml and 45 µl was transferred to the lysis tube or plate. Cells were pre-treated for 10 min on ice with Pierce Universal Nuclease (125 U total, 88701) and lysed by the addition of 50 µl of 2× lysis buffer (100 mM HEPES pH 8.5, 1% SDS, 2% sodium deoxycholic acid, 2 mM MgCl2, 100 mM NaCl, 10 mM TCEP, 40 mM CAA, 1× protease inhibitor). The cells were incubated with mixing at 37 °C for 10 min. 10 µl of 20% SDS was added to bring the final concentration to >2% and the samples were mixed for an additional 10 min at 70 °C. Protein concentrations were measured using the Pierce reducing agent compatible BCA kit (Thermo Scientific, 23250).

Frozen human tumour and normal tissue samples analysed in Supplementary Fig. 24 were obtained from Audubon Bioscience (see Supplementary Table 2 for a list of patient information for tissue samples). Samples were cut on ice with a clean scalpel, blotted dry to remove excess liquid, weighed, and stored at −80 °C. For whole-cell extract analysis, aliquots of approximately 10 mg were lysed in microcentrifuge tubes using the procedure described above for cultured cells with the following alterations. Prior to lysis, tissue samples were washed 4 times with ice-cold PBS to remove blood and preincubated with 20 µl of nuclease solution before adding 100 µl of 1× lysis buffer and grinding with a plastic SpiralPestle (RPI, 2999017) driven by a rotary pestle motor (Cole-Parmer, EW-44468-25) for 30 s. The pestle was washed with an additional 100 µl of lysis buffer to remove residual lysate and the samples were mixed at 37 °C for 10 min before the addition of 1/10 volume 20% SDS and a final mixing incubation at 70 °C for 10 min. Samples were cleared by centrifugation and the supernatant was transferred to fresh tubes. Protein concentrations were measured by BCA analysis as previously described. Larger aliquots (around 30 mg) of tissue samples were also processed to isolate plasma membrane fractions using the Minute Plasma Membrane fractionation kit (Invent, SM-005) per the manufacturer’s protocol. For all cultured cell and tissue samples, we digested 10 µg of total protein using the SP3 method67 with lysC and trypsin. The resulting peptides were desalted using Supel Swift HLB dispersive tips (DPX Technologies, DPX170501), eluted with 70% acetonitrile, dried in a SpeedVac, and resuspended in 5% formic acid.

Proteome profiling via DIA acquisition

Peptides (200–500 ng) from whole-cell or patient tumour and normal samples were loaded and separated using a 93-min gradient: 5% B held for 3 min, increasing to 28% B over 73 min, then 28% to 36% B over 7 min, and 36% to 50% B over 4 min, followed by a column wash at 100% B.

The mass spectrometer was operated in data-independent acquisition (DIA) mode. MS1 survey scans were acquired in the Orbitrap (resolution = 60,000; scan range = 380–985 m/z; AGC target = 4 × 105; maximum injection time = 100 ms). Each survey scan was followed by 60 DIA MS2 scans using fixed 10 m/z isolation windows with 1 m/z overlap and instrument-optimized window placement, covering a precursor mass range of 380–980 m/z. MS2 fragment ions were detected in the Orbitrap across a scan range of 145–1,450 m/z (resolution = 15,000; HCD NCE = 30%; AGC target = 1 × 105; maximum injection time = 40 ms) in centroid mode. The total DIA cycle time was 3 s.

DIA data analysis of whole-cell proteomes

DIA raw files were analysed using Spectronaut (v20.5.260227; Biognosys AG) in directDIA+ (Deep) mode against the canonical human proteome (UniProtKB UP000005640, release 2024_08; 20,420 entries) supplemented with a custom contaminants database (38 entries). The Pulsar search engine was used with fully specific tryptic digestion (Trypsin/P), up to two missed cleavages, and peptide lengths of 7–52 residues. Fixed modifications included carbamidomethylation of cysteine; variable modifications included oxidation of methionine and acetylation of protein N termini, with a maximum of two variable modifications per peptide. Decoys were generated using a mutated sequence strategy with a dynamic decoy limit set to 10% of the library size. Mass tolerances for both MS1 and MS2 were determined dynamically for each run using Spectronaut’s automated calibration. Retention time calibration used deep learning-assisted iRT regression with non-linear local regression and automated iRT source assignment.

Precursor identification was controlled at 1% q-value at both the run and experiment levels; protein group identification was controlled at 1% q-value at the experiment level and 5% at the run level, with single-hit proteins assessed using a stratified FDR rule. Protein inference was performed using the IDPicker algorithm. Quantification was based on MS2 peak areas extracted within dynamically determined retention time and mass tolerance windows. Protein abundances were computed by summing the mean MS2 area of the top 3 precursors per peptide sequence, using up to the top 5 peptides per protein group. Interference correction was applied using only confidently identified peptides, requiring a minimum of 2 MS1 and 3 MS2 observations, with multi-channel interferences excluded. Cross-run normalization used Spectronaut’s automatic normalization strategy anchored exclusively to human proteome FASTA entries, excluding contaminant proteins. Quantities were reported at the protein group level across the full experiment.

Cell copy number quantification using proteomic ruler method

Absolute protein copy numbers used for analyses in Supplementary Figs. 13 and 20 were calculated using the proteomic ruler approach68. For each cell line, the total proteomic signal was summed across all quantified proteins, and a separate sum was computed over a fixed set of histone proteins. The histone signal fraction was calculated as the histone signal divided by the total proteomic signal. The cellular DNA mass was determined as 6.5 pg × (ploidy/2), with ploidy values obtained from DepMap (OmicsGlobalSignatures.csv) (DepMap, Broad (2026). DepMap Public 26Q1. Dataset. https://depmap.org) and a diploid state (ploidy = 2) assumed for cell lines lacking ploidy. The total protein mass per cell was derived as the histone mass divided by the histone signal fraction. The mass of each individual protein was then computed as its fractional share of the total signal (protein signal/total signal) multiplied by the total protein mass per cell. Finally, protein masses were converted to copy numbers per cell by dividing the protein mass by its molecular weight and multiplying by Avogadro’s number.

Corresponding protein abundance measurements for cells listed in Supplementary Fig. 4a are provided in Supplementary Table 3.

EGFR immunoprecipitation and mass spectrometry analysis

Five million cells were aliquoted to low protein binding tubes and washed once in cold DPBS. Cells were lysed in 550 µl 40 mM HEPES pH 7.4, 150 mM NaCl, 1× Halt Protease and Phosphatase Inhibitor Cocktail, EDTA-free (Thermo Scientific, 78441), 2 mM EDTA, 1% (w/v) freshly prepared n-dodecyl β-D-maltoside detergent and vortexed for 10 s. Samples were kept on ice for 30 min and vortexed for 10 s every 10 min. Lysates were then centrifuged at 16,000g for 15 min at 4 °C. Five hundred microlitres of supernatant was transferred to a KingFisher deepwell plate (Thermo Scientific, 95040450B). Five micrograms of primary antibody (biotinylated EGFR clone 528 (Santa Cruz Biotechnology, sc-120b), EGFR clone EGFR.1 (BD, 555996) or Cetuximab (R&D Systems, MAB9577-100)) or isotype (biotinylated mouse isotype (Thermo Scientific, 13-4714-85), mouse isotype (BD, 556648) or human isotype (Biolegend, 403502)), was added to each well before transferring to the Kingfisher Apex. Lysates with the primary antibody were mixed for 30 min at 4 °C. Fifty microlitres of streptavidin magnetic beads (Thermo Scientific, 88817) or protein A/G magnetic beads (Thermo Scientific, 88802) were mixed with 550 µl Pierce IP Lysis Buffer (Thermo Scientific, 87787) and subsequently washed twice with 900 µl Pierce IP Lysis Buffer. Preincubated lysates were then mixed with the magnetic beads for 1 h at 4 °C before washing the beads four times with 900 µl Pierce IP Lysis Buffer. Proteins were eluted from the beads by mixing with 50 µl Pierce IgG Elution Buffer (Thermo Scientific, 21004) for 10 min at room temperature twice. The two eluate fractions were pooled and mixed with 15 µl 1 M Tris pH 7.4 to neutralize. Samples were stored at −80 °C prior to proteomic analysis.

The eluted immunocomplexes were processed and analysed using the same methods described for cell extracts above. The samples were first reduced and alkylated with TCEP and chloroacetamide. Proteins were collected, washed, digested with lysC and trypsin, and desalted using the SP3 method HLB dispersive tips. Samples were analysed using a 30 min data-independent acquisition method on the Orbitrap Eclipse with all other parameters the same as for cell extracts and searched in Spectronaut using the same schema with all runs grouped together to enable match between runs. Cross-run normalization was omitted, as it relies on higher sample uniformity than present in immunoprecipitations with different antibodies and systems.

Analysis of RTK expression and signalling-state diversity across mapped cell lines

Cancer Cell Line Encyclopedia (CCLE) reverse-phase protein array (RPPA) phospho-signalling heat map and RPPA data were obtained from the CCLE69, using the CCLE_RPPA_20181003 level-4 matrix and the CCLE_RPPA_Ab_info_20181226 antibody annotation table. Of the 28 study cell lines, 26 were present in CCLE RPPA (BT549 and CACO2 had no RPPA data and were excluded). A 16-antibody panel was selected to span RTK-direct phospho-epitopes (for example, EGFR pY1068/pY1173, HER2 pY1248, HER3 pY1289 and MET pY1235) and core downstream effectors of the RAS–MAPK, PI3K–AKT–mTOR and STAT axes. The _Caution suffix on raw CCLE antibody names denotes the MD Anderson RPPA Core validation tier; it was retained as a validation annotation and stripped from display labels. Each antibody was z-scored across the 26 matched lines, and CCLE mRNA abundance [log2(TPM + 1)] for the 12 study RTKs was likewise z-scored across the same lines. The two matrices were displayed as stacked heat maps sharing a common cell line axis, with columns ordered by Ward hierarchical clustering of the combined z-score matrix and the colour scale (RdBu_r) clipped at ±2.5z. Analyses were performed in Python 3.10+ using pandas ≥2.1, NumPy ≥1.26, SciPy ≥1.11, and matplotlib/seaborn.

Bioinformatic analysis of micromapping experiments

Primary bioinformatic analysis of LC–MS/MS data was performed in the R statistical computing environment. Peptide-level abundance data were used to identify the number of peptides corresponding to each protein. To normalize for loading differences, peptide abundance was normalized to the summed total abundance for each sample; these totals were averaged, and individual normalized values were rescaled by this average. Peptide-level data were then merged to protein-level data by calculating the median of all peptides assigned to a given protein.

Differential protein enrichment was subsequently assessed using the limma R package (v3.64.3). log2-transformed summed intensity values were used to calculate statistical significance, reporting t-statistics and P values. A linear model was fit to each protein and an empirical Bayes procedure moderated the per protein residual variances by shrinking them toward a common prior, yielding a two-sided moderated t-statistic and corresponding P value for the target-versus-control contrast. To account for multiple hypothesis testing, P values were adjusted using the Benjamini–Hochberg procedure. Results were visualized using Seaborn, generating volcano plots that display log2 fold change against the negative log10-transformed adjusted P values.

Proximity interaction networks

Global RTK proximity network was constructed using the NetworkX Python library ≥3.2, where edges connected experimental target proteins to enriched proteins. We applied permissive enrichment thresholds (log2FC > 0.1, q-value < 0.2) to define edges, ensuring the inclusion of lower-affinity or transient interactions. The resulting network topology was exported in GEXF format and spatialized in Gephi using the ForceAtlas2 layout algorithm (LinLog mode).

The RTK–RTK interaction network in Fig. 2c was constructed in NetworkX from previously reported3, literature-curated RTK heterointeractions, augmented with high-confidence intra-family interactions from STRING (experimental score > 0.7). By contrast, analyses in Fig. 2e,f were evaluated against a broader, more inclusive reference set integrating multiple interaction databases (STRING23, CORUM24, BioGRID25, IntAct26), enabling assessment against less stringent but more comprehensive interaction annotations. Metrics were calculated as described in ‘Evaluation metrics’. The computational graph was exported to Gephi for final formatting.

Overlap and similarity analysis

To systematically compare RTK interactomes across targets, we generated a two-dimensional array of pairwise comparisons. Overlaps were visualized as area-proportional Venn diagrams using matplotlib-venn, with circle sizes weighted by the total number of enriched proteins. The Jaccard index was computed to quantify set similarity (J = |A ∩ B|/|A ∪ B|). To facilitate interpretation, conditions were ordered by Euclidean distance, placing phenotypes with the greatest overlap in proximity.

Score normalization

Raw t-statistics were normalized using robust z-score transformation applied independently within each experiment. The robust z-score was computed as:

$${z}_{i}=\frac{{t}_{i}-{\rm{median}}(t)}{\text{MAD}(t)}$$

where ti is the moderated t-statistic for protein i from the limma differential enrichment analysis, median(t) is the median of the moderated t-statistics across all proteins quantified in that experiment, and MAD is median absolute deviation of the same set of t-statistics, scaled by 1.4826 for consistency with the standard deviation of a normal distribution [scipy.stats.median_abs_deviation with scale = ‘normal’]. This per experiment normalization equalizes distributional differences across experiments (loud versus quiet targeted protein effects), enabling fair cross-experiment comparisons.

Hit calling

Proteins were flagged as enriched (hits) if their normalized score exceeded a threshold of 2.0 MAD units:

$${\mathrm{hit}}_{i}=1[{z}_{i} > 2.0]$$

An alternative statistical criterion required both log2FC > 1.0 and Benjamini–Hochberg adjusted P value ≤ 0.05.

Pairwise correlation analysis for MetaMap

Spearman rank correlations were computed between all protein pairs across experiments using pairwise-complete observations. Correlations were computed only for pairs with at least ten co-observations (experiments where both proteins had non-missing scores). The correlation coefficient ρ was computed as the Pearson correlation of ranked values:

$$\rho =\frac{\sum _{i}({R}_{x,i}-{\bar{R}}_{x})({R}_{y,i}-{\bar{R}}_{y})}{\sqrt{\sum _{i}{({R}_{x,i}-{\bar{R}}_{x})}^{2}\sum _{i}{({R}_{y,i}-{\bar{R}}_{y})}^{2}}}$$

Where Rx,i denotes the rank of protein x in experiment i among valid observations. Ties were resolved using average ranking. P values were computed analytically using the t-distribution approximation:

$$t=\rho \sqrt{\frac{n-2}{1-{\rho }^{2}}},p=2\times P(T > |t|),T \sim {t}_{n-2}$$

Multiple testing correction was performed using the Benjamini–Hochberg procedure at FDR α = 0.05 with monotonicity enforcement.

Network clustering

Protein modules were identified using the Louvain community detection algorithm [NetworkX] on weighted undirected graphs constructed from significant correlations. Edge weights corresponded to Spearman ρ values. Parameters: correlation threshold ≥0.3 for edge inclusion, resolution parameter = 1.0, minimum cluster size = 5 proteins, random seed = 42 for reproducibility.

Community clusters, such as the one depicted in Fig. 3e, were constructed by drawing edges between core community proteins with high correlation (ρ > 0.4). Additional edges connect highly correlated proteins of interest to core members. Edges are annotated as documented in reference databases or as potentially novel.

Assessment of abundance-profile similarity among MetaMap clusters

To test whether community clusters could be explained by shared protein abundance rather than by the proximity signal from which they were derived, all unordered pairs of proteins within the dataset that mapped to a cluster and to a cell line proteomics measurement were enumerated. Proteomics intensities (copies per cell, ploidy-corrected) were log10-transformed, and pairwise abundance-profile similarity was computed as the Spearman correlation of log10 copies across cell lines, requiring at least 5 overlapping cell lines. Each pair was labelled as same-cluster or different-cluster, and assigned a MetaMap proximity signal by joining to the precomputed Spearman edges table; only significant edges were used, and pairs without a stored edge were treated as missing rather than zero.

Marginal differences in abundance Spearman and MetaMap ρ between same-cluster and different-cluster pairs were tested with the two-sided Mann–Whitney U statistic (scipy.stats.mannwhitneyu), with the rank-biserial effect size r = 2U/(n1 × n2) − 1 reported alongside (equivalent to 2·AUROC − 1), and the Kolmogorov–Smirnov D statistic for distributional shape (scipy.stats.ks_2samp).

Reference interaction databases

Protein–protein interaction references were obtained from STRING (v12)23, BioGRID (human physical interactions)25, IntAct (human molecular interactions)26 and CORUM (manually curated protein complexes, v5.1)24. For STRING interactions, only high-confidence associations (combined score ≥ 0.7) were retained. CORUM complexes were expanded to all pairwise protein combinations within each complex (minimum complex size: 2 proteins). Interactions were normalized such that protein_a ≤ protein_b alphabetically to ensure consistent pair representation.

To evaluate whether correlated protein pairs are enriched for known interactions, we compared predicted pairs against a random baseline. For each reference database, we calculated precision as defined above. We then generated 500 random samples of protein pairs, matched in size to the predicted set, drawn from the same protein universe. Fold enrichment was computed as the ratio of predicted precision to mean background precision. Statistical significance was assessed using an empirical P value: the proportion of random samples achieving precision equal to or greater than the predicted precision, with a + 1 correction applied to both numerator and denominator to avoid zero P values (P = (k + 1)/(n + 1), where k is the count of random samples greater than or equal to predicted and n = 500 iterations).

Visualization

Hierarchical clustering used Ward linkage with Euclidean distance on correlation matrices. Kernel density estimates used bandwidth adjustment factor 4.5 for score distributions.

Differential gene expression analysis

Differential gene expression analysis was performed for each protein-indication pairing using DESeq270 implemented via pyDESeq2 (v0.4). Gene expression data were obtained from the UCSC Toil Recompute harmonized dataset71, which reprocesses TCGA tumour samples and GTEx normal tissue samples through a unified RNA-seq pipeline to eliminate batch effects arising from different sequencing protocols.

For each indication (primary site), expression profiles from primary tumour samples (TCGA) were compared against normal tissue samples (GTEx). Statistical significance was assessed using the Wald test. P values were adjusted for multiple testing using the Benjamini–Hochberg procedure to control the false discovery rate (FDR) at α = 0.05. A minimum of three samples per group was required for analysis.

Results are reported as log2 fold change (tumour versus normal) with corresponding adjusted P values.

DepMap coessentiality analysis

Protein coessentiality was analysed using the DepMap Public 26Q1 Chronos gene–effect matrix72 (CRISPRGeneEffect.csv) (DepMap, Broad (2026). DepMap Public 26Q1. Dataset. https://depmap.org) and was limited to the 22 experimental cell lines from our proximity experiments that were present in this release (CACO2, CAPAN2, HCC4006, MKN7, NCIH1563, SW900 were not in the DepMap file). A protein was considered essential in a given cell line if its Chronos gene effect score was ≤−0.5. A protein pair was considered coessential if both proteins were essential in at least one shared cell line. Surface proteins were annotated using the SURFY surfaceome predictions (Surfaceome Label = “surface”)73. Known protein–protein interactions were defined using STRING interactions with a combined_score ≥ 700. Protein pairs predicted by the GAT model were labelled as novel if they were absent from the STRING ≥ 700 interaction set.

For coessentiality enrichment analysis, the ‘All’ group was defined as the full background set of protein pairs evaluated in the GAT scoring space. The ‘GAT all’ group was defined as the subset of All pairs predicted by GAT in at least five micromap experiments as potential proximal partners. The ‘Surface×Surface’ group was defined as the subset of All pairs in which both proteins were annotated as surface proteins by SURFY73. The ‘GAT Surface×Surface’ group was defined as the subset of Surface×Surface pairs predicted by GAT in at least five micromap experiments. In Supplementary Fig. 18, the coessential fraction for each group was calculated as the number of coessential protein pairs divided by the total number of protein pairs in that group. Dashed lines mark the All or Surface×Surface background coessential fraction used for comparison.

EGFR×CDCP1 predicted overlap analysis

We collected all GAT-predicted EGFR and CDCP1 proximal partners that were recurrently predicted in at least five independent micromapping experiments, yielding 894 EGFR partners and 388 CDCP1 partners, of which 385 were shared. These three sets were visualized as a Venn diagram in which each region was annotated with the fraction of its members classified as cell surface proteins according to the SURFY in silico surfaceome predictions73, with region fill shaded in proportion to this surface fraction.

For pathway-level interpretation, the 385 shared partners were submitted to the Reactome AnalysisService API74 with results restricted to Homo sapiens. From the enriched pathways, we selected a curated set of nine RTK and signal-transduction pathways. For each pathway, we report the percentage of its annotated proteins recovered by the shared EGFR and CDCP1 predicted partners, with each bar partitioned by the fraction of recovered proteins classified as surface versus non-surface.

Clinical proteomic analysis of CDCP1

Clinical proteomic datasets were obtained from the Proteomics Data Commons (PDC) and from proteomic analyses performed in-house on human tumour and normal tissue samples. PDC datasets were used for the cross-cancer comparisons shown in Supplementary Fig. 24a, whereas the in-house datasets were used for the NSCLC and CRC comparisons shown in Supplementary Fig. 24b,c.

For the PDC analysis, all studies containing a proteomics file within the protein assembly data type were included, from which relevant cancer indications were selected for analysis. Gene symbols from each proteomics file were mapped to UniProt75 accessions by first attempting to match against UniProt primary gene symbols. When no primary symbol match was found, gene symbols were matched to protein synonyms to maximize coverage.

Protein-level abundances were retained at single-sample resolution together with their study, cancer indication, and tissue-type annotations. Comparisons were restricted to primary tumour versus adjacent normal tissue; metastatic and unannotated tissue categories were excluded. A study × indication combination was analysed only if both tissue types were represented and each group contained at least three samples. To avoid mixing quantification platforms and study-specific normalization schemes within a single comparison, each indication in Supplementary Fig. 24a was drawn from one study. Abundance values are reported on a log2 scale as provided by the source study.

For each indication, CDCP1 abundance was compared between normal and tumour samples using a two-sided Mann–Whitney U test. Samples were treated as independent; donor-matched pairs were not modelled as paired, so reported P values are conservative with respect to the paired subset. Tumour to normal fold change was computed on the log2 scale as the ratio of the group geometric means.

$$\mathrm{FC}={2}^{(\mathrm{mean}[{\log }_{2}\mathrm{tumour}]-\mathrm{mean}[{\log }_{2}\mathrm{normal}])}$$

For the in-house analysis, plasma membrane proteomic datasets were generated from normal and malignant human NSCLC and CRC tissue samples. Mass spectrometry signal intensities derived using Spectronaut were normalized to protein molecular mass (in Da). CDCP1 intensities were retained at single-sample resolution, and normal and malignant samples were compared using a two-sided Mann–Whitney U test. Samples were treated as independent.

Cell starvation and EGF stimulation for microenvironment mapping

CAOV3 cells were grown in the manufacturer’s recommended medium. Two sets of cells were washed twice with DPBS and starved in medium lacking FBS for 24 h. The medium was replaced in both sets with one of the sets supplemented with 30 ng ml−1 EGF (Thermo Scientific, PHG0311L) for an additional 24 h. Cells were then collected for western blot analysis and micromapping.

For Western blot analysis, equal numbers of cells were lysed in RIPA with 1× protease inhibitor tablet and 1:1,000 dilution of benzonase (Sigma-Aldrich, 70664-3). Protein concentrations were measured by BCA assay and normalized. Samples were mixed with 4× Laemmli buffer (Bio-Rad, 1610747) with β-mercaptoethanol (Thermo Scientific, AC125472500) and stored at −80 °C. Before loading, samples were boiled at 95 °C for 5 min and run 1 h at 180 V on a 12% TGX Criterion gel (Bio-Rad, 5671045). The iBright Prestained Protein Ladder (Thermo Scientific, LC5615) was included as a protein ladder. Gel was washed briefly in water before transfer using an iBlot2 Gel Transfer device and PVDF stacks (Thermo Scientific, IB24001). PVDF blots were blocked in 3% bovine serum albumin (Sigma, A7906-100G) in 1× TBST (diluted from 20× stock, Boston BioProducts, IBB-181X-4L) for at least 1 h. Blocking buffer was replaced and primary antibodies for EGFR (1:1,000 dilution, Cell Signaling Technology, 4267S), CDCP1 (1:1,000 dilution, Cell Signaling Technology, 4115S), and ß-actin (1:5,000 dilution, Thermo Scientific, MA5-15739) were added. The next day, blots were washed three times 5 min each with 1× TBST before addition of goat anti-mouse IRDye 680RD (1:5,000 dilution, Li-COR, 926-68070) and goat anti-rabbit IRDye 800CW (1:5,000 dilution, Li-COR, 926-32211) in blocking buffer for 1 h. Blots were washed three times 5 min each with 1× TBST and two times briefly with water before imaging on a Li-COR Odyssey CLx v2.2. Uncropped Western blots are provided in Supplementary Fig. 1.

Western blot analysis of EGFR phosphorylation under micromap conditions

A549 cells were grown in manufacturer’s recommended medium. Subconfluent cells were washed twice with DPBS and starved in medium lacking FBS for 24 h. Cells were then collected in serum free medium and counted. Ten million cells were aliquoted to low protein binding tubes and spun down 800g for 5 min at 4 °C. For pre-treatment with antibodies, cell pellets were resuspended in 1 ml serum free medium with 10 µg human isotype or cetuximab (R&D Systems, MAB9577-100) and incubated at 37 °C for 15 min. To stimulate EGFR phosphorylation, EGF (Thermo Scientific, PHG0311L) was diluted 1:100 in serum free medium and then 2 µl was added to cells (20 ng ml−1 final concentration) for 15 min at 37 °C. Cells were washed twice in cold DPBS and then resuspended in 1 ml cold DPBS. For post-EGF stimulation antibody treatment, 10 µg human isotype (Biolegend, 403502), cetuximab (R&D Systems, MAB9577-100), mouse isotype or mouse anti-EGFR (clone EGFR.1) primary antibodies were added to cells. Samples were incubated at 4°C for 30 min on a rotisserie. Cells were then spun down and washed twice with 1 ml cold DPBS before being resuspended in 1 ml cold DPBS and incubated at 4°C for 30 min on a rotisserie. Cells were then spun down and washed twice with 1 ml cold DPBS before lysing in 1 ml RIPA buffer containing 1× Halt Protease and Phosphatase Inhibitor Cocktail, EDTA-free (Thermo Scientific, 78441) and 1:1,000 dilution of benzonase and incubated for 15 min at 4°C on a rotisserie. After lysis, the protein concentrations were measured by BCA assay and samples were stored at -80 °C. Equal protein amounts were mixed with 4× Laemmli buffer with β-mercaptoethanol and boiled at 95 °C for 5 min before running on SDS-PAGE as previously described.

For Western blot analysis of EGFR phosphorylation, 20 µg lysate was run 60 min at 150 V on a 12% TGX Criterion gel (Bio-Rad, 5671044). Gel was transferred using an iBlot2 Gel Transfer device and nitrocellulose stacks (Thermo Scientific, IB23001). Nitrocellulose membranes were blocked using Intercept Blocking Buffer (Li-COR, 927-60001) for 2 h. Blocking buffer was replaced and primary antibodies for Total EGFR (1:5,000 dilution, Cell Signaling Technology 2239S), phospho-EGFR (Tyr1068) (1:1,000 dilution, Cell Signaling Technology, 2234S) and β-actin (1:5,000 dilution, Thermo Scientific, MA5-15739) were diluted in 1:2 Intercept Blocking Buffer:H2O and incubated overnight at 4 °C. The next day, blots were washed four times 10 min each with 1× TBST (diluted from 20× stock, Boston BioProducts, IBB-181X-4L) before addition of goat anti-mouse IRDye 680RD (1:15,000 dilution) and goat anti-rabbit IRDye 800CW (1:15,000 dilution) in Intercept Blocking Buffer for 1 h. Blots were washed four times 10 min each with 1× TBST and two times briefly with water before imaging on a Li-COR Odyssey CLx v2.2. Uncropped Western blots are provided in Supplementary Fig. 1.

NanoBiT complementation assay for EGFR×CDCP1 co-engagement

The NanoBiT protein–protein interaction detection system (Promega)76 was adapted to EGFR and CDCP1 by fusing LgBiT or SmBiT to the C terminus of each full-length receptor and cloning the receptor-NanoBiT fusion proteins into pcDNA3.4 mammalian expression vectors. Expression plasmids were generated for every combination of receptor (EGFR/CDCP1) and NanoBiT component (LgBiT/SmBiT), such that the NanoBiT system could be applied in two arrangements: (1) co-expressing CDCP1-LgBiT and EGFR-SmBiT; or (2) co-expressing EGFR-LgBiT and CDCP1-SmBiT (see Supplementary Table 4 for DNA and protein sequences of each construct).

Expi293 cells (Gibco, A14635) were transiently transfected with NanoBiT components using 293fectin (Gibco, 12347019). For each transfection, transfection mixtures were prepared by mixing 6 µg 293fectin and 3 µg of each expression plasmid (6 µg total DNA) in 200 µl Opti-MEM (Gibco, 31985070) and incubating the mixture for 20 min at room temperature. Each transfection mixture was then added to 3 ml of Expi293 cells cultured at 106 cells per ml in Expi293 Expression Medium (Gibco, A1435101). Cells were then cultured overnight at 37 °C, 8% CO2, 80% humidity, and shaking at 125 rpm (24 mm throw).

The following day, Expi293 cells were diluted to 0.1 × 106 cells per ml, mixed with a 1× concentration of Nano-Glo Vivazine Substrate (Promega, N2580), transferred to a white, opaque 96-well plate, and cultured for another hour. Cells were then treated with 100 nM of antibody or an equivalent volume of PBS and transferred to a plate reader to monitor luminescence over 90 min at 1 min intervals.

For each condition, raw luminescence values, \({\rm{RLU}}(t)\), were first normalized to their initial time point to obtain a time-normalized luminescence, \({\mathrm{RLU}}_{\mathrm{rel}}(t)\)

$${\mathrm{RLU}}_{\mathrm{rel}}(t)=\frac{\mathrm{RLU}(t)}{\mathrm{RLU}(t=0)}$$

RLUrel(t) was then background-subtracted using the mean time-normalized luminescence of the PBS control, RLUrel,PBS(t), to obtain the final normalized luminescence, RLUnorm(t):

$${\mathrm{RLU}}_{\mathrm{norm}}(t)={\mathrm{RLU}}_{\mathrm{rel}}(t)-{\mathrm{RLU}}_{\mathrm{rel},\mathrm{PBS}}(t)$$

Concurrent binding assay

To assess simultaneous EGFR and CDCP1 engagement by the bispecific antibody, a concurrent binding assay was performed using a soluble ectodomain detection format adapted from a previous method for evaluating cross-arm binding by cell-bound bispecific antibodies77. Parental SW48 and SW48 CDCP1-KD cells were incubated with serial dilutions of the anti-EGFR×CDCP1 bispecific antibody (bsAb), starting at 500 nM, for 1 h at 4 °C. Following antibody incubation, cells were washed with FACS buffer and incubated with PE-conjugated rat anti-human IgG Fc antibody (BioLegend, 410708) in the presence of either 50 nM biotinylated soluble EGFR ectodomain (Acro Biosystems, EGR-H82E3) or 50 nM biotinylated soluble CDCP1 ectodomain (Acro Biosystems, CD1-H82E4) for 30 min at 4 °C. Cells were subsequently washed and stained with streptavidin–APC (BioLegend, 405207) for an additional 30 min at 4 °C. After a final wash, samples were analysed using a BD LSRFortessa X-20 flow cytometer with BD FACSDiva v9.2. Data were processed using FlowJo software v10, and mean fluorescence intensity (MFI) values were used to quantify binding. Fluorescence intensities were normalized, and binding curves were generated by plotting the APC-to-PE fluorescence ratio against antibody concentration.

Internalization, ADC cytotoxicity, and TCE cytotoxicity assays

For internalization assays, cells (SW48, BXPC3, PC3, bronchial/tracheal epithelial cells and cervical epithelial cells) were grown in manufacturer’s recommended media and were collected and plated at 20,000 cells per well in a tissue culture treated 96-well plate. Cells were allowed to adhere at 37 °C overnight. The following day, test and control antibodies were prepared in a 3:1 molar ratio of Incucyte Human Fabfluor-pH Antibody Labeling Dye (Sartorius, 4722). Labelling reactions were incubated in the dark for 15 min at 37 °C. Immediately following incubation, 50 µl of 2× labelled antibody solution was added to each well containing 50 µl of culture medium to achieve the final concentration (2.5 nM for SW48, 10 nM for all other cell lines). Plates were transferred to an Incucyte S3/SX1 live-cell imaging system (Sartorius) and imaged every 45 min for 24 h in both phase contrast and red fluorescence channels (10× objective, 400 ms exposure). Imaging began 30 min after placement in the incubator to allow condensation to dissipate. Images were analysed using Incucyte software (v2023A). Phase segmentation was performed using the AI Confluence default settings, while red fluorescence was analysed using Top-Hat segmentation (radius = 30 µm; threshold = 0.2 RCU). Total red integrated intensity per well was normalized to phase confluence and exported for visualization in GraphPad Prism (v10.4.1).

For cytotoxicity assays, cells (SW48, BXPC3, PC3, bronchial/tracheal epithelial cells and cervical epithelial cells) were collected and plated in their supplier’s recommended media at 300 cells per well in 45 μl per well in tissue culture treated white 384-well plates (Corning 3570). Cells were allowed to adhere at 37 °C overnight, and then an ADC dilution plate was prepared by making serial dilutions of 10× concentrated ADCs in PBS in a 96-well plate. An automated liquid handler was used to move 5 μl of these ADCs into each well of the 384-well assay plates to get n = 4 of each condition. Plates were returned to the incubator for five days. To read out cell viability, 10 μl of CellTiter-Glo (Promega) was added to each well, plates were incubated for 30 min at room temperature, and end-point luminescence was read out using a plate reader. To calculate viability, all luminescence values were normalized to the average of untreated wells (antibody concentration of 0) for a given plate and cell type. To calculate half-maximal inhibitory concentration (IC50) values, a four-parameter curve fit was applied.

To measure T cell-dependent cellular cytotoxicity, selected cell lines (SW48, HCC827 and PC3) were stably infected with a lentiviral construct expressing a codon-optimized firefly luciferase driven by EF1α and selected with blasticidin to generate cell lines that could be monitored for viability. For the assay, luciferized cells were collected and plated in their supplier’s recommended media at 20,000 cells per well in 96-well tissue culture treated plates. That same day, peripheral blood mononuclear cells (PBMCs) were thawed and put in culture in RPMI medium with 10 ng ml−1 IL-12 (Gibco 200-12H). Both types of cells were incubated at 37 °C overnight. The following day, PBMCs were collected, counted, and plated in fresh medium with no IL-12 at 200,000 cells per well in the assay plate with the tumour cells. TCE antibodies were added in triplicate and assay plates were returned to the incubator for two days, then read out by adding an equal volume of Steady-Glo (Promega) and reading out on a luminescence plate reader. Luminescence signal was normalized to wells with PBMCs but no antibodies added to determine tumour cell cytotoxicity.

Generation of SW48 CDCP1-KD cell line

SW48 CDCP1-KD cells were generated using CRISPR–Cas9–mediated gene disruption. A synthetic dual–nuclear localization signal (sNLS)–SpCas9 nuclease (Aldevron) was combined with a pooled set of CDCP1-targeting CRISPR guide RNAs obtained from Synthego (UGAACCUCCCCAAAAGGACA, GACCGAUCUGCCUCAGGCGA and CCAGAUGAAAGUUCUGUUGA). Guide selection followed the manufacturer’s design criteria for targeting early coding exons to promote functional knockdown. Ribonucleoprotein (RNP) complexes were prepared following general guidelines provided by Integrated DNA Technologies (IDT) for Cas9–gRNA complex assembly. After treatment of SW48 parental cells with the assembled complexes, cells were expanded under standard culture conditions to allow outgrowth of edited populations. Following recovery, bulk-edited cultures were screened for loss of CDCP1 surface expression by flow cytometry. CDCP1-KD cells were enriched by fluorescence-activated cell sorting (BD FACSAria), ensuring a stable CDCP1-KD cell line suitable for downstream analysis.

In vivo cell line-derived xenograft studies

Female CrTac:NCr-Foxn1<nu> (NCr nude) or NOD.Cg-Prkdcscid Il2rgtm1Sug/JicTac (NOG) mice, age 5 to 8 weeks, were obtained from Taconic Biosciences. All mice were housed with up to five mice per cage in individually ventilated cages in a pathogen-free animal facility. Mice were maintained under artificial lighting (12 h) in a controlled ambient temperature of 68 to 79 °F (20–26 °C), and relative humidity between 30 and 70%. Mice were acclimated for at least five days before the experiments. All animal procedures were conducted in accordance with, and with approval of, the policy of the Bloodworks NW Research Institutional Animal Care and Use Committee.

Tumour cells (2 × 106 to 10 × 106 per 100 µl) were implanted subcutaneously into the right hind flank of female mice in a 1:1 mixture of base medium and matrigel (Corning, 356234). In the dual-flank model, parental tumour cells were implanted in the right hind flank and CDCP1-KD tumour cells were implanted in the contralateral flank.

Tumours were measured by digital callipers, and tumour volume (TV) was calculated as TV = (L × W × W)/2, where L (length) is the longest measurement (in mm) and W (width) is perpendicular to L. When tumours reached ~100–150 mm3, mice were randomized by TV into study groups using Benchling In Vivo Software v0.0.8 (n = 8–10 mice per group). Researchers were not blinded to the test articles administered to each treatment group.

The ADCs were administered at the indicated dose in 100 µl formulation buffer (20 mM histidine, 8% sucrose, pH 5.5) via tail vein injection. Researchers were not blinded to the test articles administered to each treatment group. Tumours and body weight (BW) were measured twice weekly until control-treated tumours reached 2000 mm3, at which point TGI and mean body weight change were evaluated. TGI was calculated as [1 − (T − T0)/(C − C0)] × 100, where T is mean tumour volume of the treatment group at the end of the study, T0 is mean tumour volume of the treatment group at the start of treatment, C is mean tumour volume of the control group at the end of the study, and C0 is mean tumour volume of the control group at the start of treatment. TGI statistics were performed on GraphPad Prism using a two-way ANOVA with Tukey’s all groups comparison. Animals were promptly euthanized upon reaching the 2,000 mm3 tumour volume limit (maximum volume permitted by IACUC). Animals were euthanized early if they lost >20% of their initial body weight, had severe tumour ulceration, or became moribund.

Treated mice continued to be monitored for long-term responses until day 80, or when tumours reached 800 mm3 (survival end-point). A Kaplan–Meier analysis was performed in which a survival event was considered a tumour volume exceeding 800 mm3. Survival statistics were performed on GraphPad Prism using logrank (Mantel–Cox) test. P values of less than 0.05 were considered significant (*P < 0.05; **P < 0.01; ***P < 0.001; ****P < 0.0001).

Serum pharmacokinetic analysis

Female NCr nude mice were implanted with SW48 tumour cells as described above. When tumours reached ~150 mm3, mice were randomized by tumour volume into study groups (n = 10 mice per group) and ADCs were administered as described above. Blood was collected under anaesthesia via retro-orbital bleed using a non-heparinized microcapillary tube (Fisherbrand) to obtain 100 µl in Microvette Serum Gel tubes (Sarstedt) at 4 time points (1, 24, 72 and 168 h post-dose; n = 3 to 4 mice per time point). Serum was isolated and frozen at −80 °C for pharmacokinetic analysis.

Serum samples were thawed from −80 °C and diluted 1:100 in 1× PBS. Five microliters of each serum sample was plated in each well of a 96-well ½ Area AlphaPlate (Revvity). Standard curves were prepared by diluting untreated mouse serum 1:100 in 1× PBS and performing an 8-point, 5-fold serial dilution of the test article in untreated serum starting at a concentration of 5,000 ng ml−1. Five microliters of each standard curve sample was plated in duplicate on the same 96-well ½ Area AlphaPlate as their respective serum samples. Target antibody (Biotinylated anti-MMAE) was prepared in 1× immunoassay buffer (Revvity) at 0.25 µg ml−1 for total ADC detection. Twenty microliters of target antibody solution was added to each well. Samples were shaken on a plate shaker for 1 min at 100 revolutions per min (rpm) and incubated for 1 h at room temperature. Anti-human immunoglobulin G (IgG) Acceptor beads (Revvity) were prepared at a final concentration of 20 µg ml−1 in 1× immunoassay buffer, and 12.5 µl of solution was added to each well of the 96-well plate. Samples were shaken on a plate shaker for 1 min at 100 rpm and incubated for 1 h at room temperature. Streptavidin donor beads (Revvity) were prepared at a final concentration of 20 µg ml−1 in 1× immunoassay buffer, and 12.5 µl of solution was added to each well of the 96-well plate. Samples were covered with aluminium foil, shaken on a plate shaker for 1 min at 100 rpm, and incubated for 30 min at room temperature. After incubation, luminescence was measured on the BioTek Synergy Neo2 plate reader using Gen5 version 3.12 software.

Raw relative light units (RLUs) for each set of pharmacokinetic samples were converted into nM analyte concentrations by fitting to their respective dose–response curves in GraphPad Prism. Values in the standard curve that fell above the hook effect were omitted before fitting RLUs to the curve. Analyte concentration values were then scaled by the dilution factor at which the serum samples were diluted in 1× PBS.

Cell line-derived xenograft tumour dissociation for photocatalytic proximity labelling

Female NCr nude mice were implanted with SW48 tumour cells as described above. When tumours reached ~500 mm3, tumours were excised with scissors and forceps, placed in RPMI-1640 medium (Gibco, A10491-01) and kept on ice. Single cell suspensions of SW48 tumour cells were prepared by chopping tumour tissue with a razor blade and digesting with an enzyme cocktail supplied in Miltenyi’s Tumor Dissociation Kit, Human (130-095-929). Approximately 500 mg of tumour tissue was incubated with enzyme cocktail in gentleMACS C tubes (Miltenyi, 130-096-334) on a gentleMACS Dissociator (Miltenyi) following the manufacturer’s protocol. Cell suspensions were washed with RPMI-1640 medium, passed through a 70 µm MACS SmartStrainer (Miltenyi, 130-110-916), washed with PBS and kept on ice. Viable cell counts were obtained, and cells were frozen in CryoStor CS10 medium (Stem Cell Technologies, 07930) until use.

Terminal tissue collection and immunohistochemistry

Tumours from untreated mice were dissected with skin attached and fixed in 10% neutral buffered formalin (Sigma-Aldrich, HT501128) for 24 h, then transferred to 70% ethanol (Decon Labs, 8601). Tumours were paraffin embedded, sectioned, and analysed by immunohistochemistry at Acepix Biosciences using validated immunohistochemistry assays for CDCP1 and EGFR. Specimens were embedded tumour side down (excess skin was trimmed if needed). If the tumour was large, it was bisected from the centre of the tumour and both pieces were embedded on the same block cut side down so sections were collected from the centre of the tumour. Specimens were incubated with anti-CDCP1 (Cell Signaling Technologies, CST4115) or anti-EGFR (Abcam, ab227642) for 1 h at 0.42 and 0.0667 mg ml−1, respectively.

Antibody and ADC production

EGFR×CDCP1 bispecifics, TCE trispecifics, and their monovalent and bivalent controls were transfected in ExpiCHO-S cells (Thermo Scientific, A29127) in ExpiCHO expression medium (Thermo Scientific, A2910001). Proteins were purified on a Cytiva ÄKTA Pure with Unicorn v7.10 SP1 software. All samples were first purified using PrismA HiTrap affinity capture (Cytiva, 17549854), neutralized with 1 M NaPi pH 7.0, and polished by SEC purification over a Superdex 200 Increase column (Cytiva, 28990944) in PBS (Fisher Scientific, SH3025602).

All bispecifics were generated using knob-in-hole approach78. These molecules contain CH3 mutations that preferentially drive heavy chain heterodimer formation. Light chain swapping is not a concern because the molecules are asymmetric, containing only one Fab arm. Monospecific, bivalent controls did not utilize these mutations and are typical of IgG1 Fc. The Dummy×Dummy bispecific is comprised of the RSV Fab (see ‘Preparation of direct antibody–photocatalyst conjugates’) as well as an anti-Campylobacter jejuni VHH with UniParc ID UPI000456342C.

Purity was determined by analytical SEC on an Agilent 1260 Bioinert system with OpenLab CDS Chemstation Edition software Rev.C01.10 using an AdvanceBio SEC 300 Å, 4.6 ×150 mm, 2.7 µm LC column (Agilent, PL1580-3301) with an AdvanceBio SEC 300 Å, 4.6 ×50 mm, 2.7 µm guard column (Agilent, PL1580-1301).

Interchain conjugate DAR was determined by analytical HIC on an Agilent 1260 Bioinert system using an AdvanceBio HIC, 4.6 ×100 mm (Agilent, 685975-908). A 20 min gradient was used to elute the product from 1.2 M ammonium sulfate (Thermofisher, J64419.A3), 50 mM NaPi pH 7.0 (Thermofisher, J63791.AP) to 50 mM NaPi, 20% Isopropanol (Thermofisher, 383910025).

Identity and drug-to-antibody ratio (DAR) analyses were performed on a Waters BioAccord LC–MS system operating with waters_connect v4.1 and Intact Mass v1, using a BioResolve RP mAb Polyphenyl Column (Waters, 186008945).

Binding functionality was performed using a Sartorius Octet BLI (Sartorius, RH16) with Octet BLI Discovery Version 13.0.3.26 by capturing antibodies to AHC2 sensors (Sartorius, 18-5142) and incubating in target proteins of interest; EGFR, CDCP1 and CD3 (ACRObiosystems, EGRH5222, CD1H52H6, CDEH5223 respectively).

Conjugates utilizing engineered cysteines were generated following a protocol similar to that previously described for site-specific Thiomab conjugation79. Antibodies and bispecifics were first buffer exchanged into 100 mM Tris, 1 mM EDTA pH 8.0 at concentrations above 5 mg ml−1. DTT (Thermo Scientific, A39255) was added at 80 molar equivalents to fully reduce interchain disulfides and engineered cysteines, either 3 h at room temp or overnight at 4 °C. The product was buffer exchanged into 20 mM Tris pH 7.5 using Zeba desalting columns (Thermo Scientific, 89892) to remove DTT and released cysteine and glutathione.

DHAA (Thermo Scientific, 250930050) was added at 15 molar equivalents and incubated at room temperature for one hour or until interchain disulfides reform. DHAA was removed by desalting into PBS using a Zeba column. 10% v/v DMSO (Sigma, D2653) was added to prepare for the linker-payload addition. Maleimide-VC-PAB-MMAE (Medchemexpress, HY-15575) was solubilized in DMSO at 5 mM. Four molar equivalents of maleimide-VC-PAB-MMAE were added and incubated for at least one hour at room temperature or overnight at 4 °C.

Material was run over a preparative SEC column, Superdex 200 Increase, to remove aggregates and free linker-payload. The final material was buffer exchanged and concentrated using Amicon Ultra concentrators (Millipore, UFC903024) to remove remaining free linker-payload, formulate it into the appropriate buffer, and reach the desired concentration.

For non-site-specific conjugates utilizing interchain disulfides, the antibodies and bispecifics were buffer exchanged into 100 mM Hepes pH 7.0 using Zeba columns. TCEP (Thermo Scientific, 77720) was added at 3 molar equivalents to the solution and mixed for 90 min at 25 °C with mixing. 10% v/v DMSO was added to the solution and mixed. Maleimide-VC-PAB-MMAE was solubilized in DMSO at 5 mM. 6 molar equivalents of maleimide-VC-PAB-MMAE were added and mixed at room temperature for at least 60 min. Material was run over a preparative SEC column, Superdex 200 Increase to remove aggregates and free linker-payload. The final material was buffer exchanged and concentrated using Amicon Ultra to also remove free linker-payload, formulate it into the appropriate buffer, and reach the desired concentration.

Reporting summary

Further information on research design is available in the Nature Portfolio Reporting Summary linked to this article.

Data availability

All data reported in this study are available in the main text, Supplementary Information, Supplementary Tables, Source Data files, and the public repositories described below. The proximity enrichment and protein abundance datasets generated in this study have also been deposited at Zenodo80 (https://doi.org/10.5281/zenodo.21678744). Mass spectrometry proteomics data have been deposited to ProteomeXchange under the dataset identifier PXD078009. This study also used the following publicly available datasets: STRING v12.0 (9606.protein.links.v12.0.txt.gz; https:/string-db.org), CORUM v5.1 (humanComplexes.txt; https:/mips.helmholtz-munich.de/corum/), IntAct release 250 (human.txt, from human.zip; https:/ftp.ebi.ac.uk/pub/databases/intact/2025-08-08/psimitab/species/human.zip), BioGRID v4.4.245 (BIOGRID-ORGANISM-Homo_sapiens-4.4.245.tab3.txt; https://thebiogrid.org), DepMap Public 26Q1 (CRISPRGeneEffect.csv, OmicsGlobalSignatures.csv; https:/depmap.org/portal/data_page/), and the predicted human surfaceome from SURFY (dataset S1 from ref. 73). Source data are provided with this paper.

Code availability

Code used for computational analyses and protein-neighbourhood modelling described in this study are available at Zenodo81 (https://doi.org/10.5281/zenodo.21679069).

References

  1. Belardi, B., Son, S., Felce, J. H., Dustin, M. L. & Fletcher, D. A. Cell–cell interfaces as specialized compartments directing cell function. Nat. Rev. Mol. Cell Biol. 21, 750–764 (2020).

    Article  CAS  PubMed  Google Scholar 

  2. Garcia-Parajo, M. F., Cambi, A., Torreno-Pina, J. A., Thompson, N. & Jacobson, K. Nanoclustering as a dominant feature of plasma membrane organization. J. Cell Sci. 127, 4995–5005 (2014).

    Article  PubMed  PubMed Central  Google Scholar 

  3. Paul, M. D. & Hristova, K. The RTK interactome: overview and perspective on RTK heterointeractions. Chem. Rev. 119, 5881–5921 (2019).

    Article  CAS  PubMed  Google Scholar 

  4. Leth-Larsen, R., Lund, R. R. & Ditzel, H. J. Plasma membrane proteomics and its application in clinical cancer biomarker discovery. Mol. Cell. Proteomics 9, 1369–1382 (2010).

    Article  CAS  PubMed  PubMed Central  Google Scholar 

  5. Hui, E. Cis interactions of membrane receptors and ligands. Annu. Rev. Cell Dev. Biol. 39, 391–408 (2023).

    Article  CAS  PubMed  PubMed Central  Google Scholar 

  6. Li, M. & Yu, Y. Innate immune receptor clustering and its role in immune regulation. J. Cell Sci. 134, jcs249318 https://doi.org/10.1242/jcs.249318 (2021).

    Article  CAS  PubMed  PubMed Central  Google Scholar 

  7. Salavessa, L. et al. Cytokine receptor cluster size impacts its endocytosis and signaling. Proc. Natl Acad. Sci. USA 118, e2024893118 (2021).

    Article  CAS  PubMed  PubMed Central  Google Scholar 

  8. Huang, L. & Muthuswamy, S. K. Polarity protein alterations in carcinoma: a focus on emerging roles for polarity regulators. Curr. Opin. Genet. Dev. 20, 41–50 (2010).

    Article  CAS  PubMed  PubMed Central  Google Scholar 

  9. Mollinedo, F. & Gajate, C. Lipid rafts as signaling hubs in cancer cell survival/death and invasion: implications in tumor progression and therapy. J. Lipid Res. 61, 611–635 (2020).

    Article  CAS  PubMed  PubMed Central  Google Scholar 

  10. Sharifi Tabar, M., Francis, H., Yeo, D., Bailey, C. G. & Rasko, J. E. J. Mapping oncogenic protein interactions for precision medicine. Int. J. Cancer 151, 7–19 (2022).

    Article  CAS  PubMed  PubMed Central  Google Scholar 

  11. Kuzmanov, U. & Emili, A. Protein–protein interaction networks: probing disease mechanisms using model systems. Genome Med. 5, 37 (2013).

    Article  CAS  PubMed  PubMed Central  Google Scholar 

  12. Casaletto, J. B. & McClatchey, A. I. Spatial regulation of receptor tyrosine kinases in development and cancer. Nat. Rev. Cancer 12, 387–400 (2012).

    Article  CAS  PubMed  PubMed Central  Google Scholar 

  13. Geri, J. B. et al. Microenvironment mapping via Dexter energy transfer on immune cells. Science 367, 1091–1097 (2020).

    Article  ADS  CAS  PubMed  PubMed Central  Google Scholar 

  14. Oslund, R. C. et al. Detection of cell-cell interactions via photocatalytic cell tagging. Nat. Chem. Biol. 18, 850–858 (2022).

    Article  CAS  PubMed  Google Scholar 

  15. Bechtel, T. J. et al. Proteomic mapping of intercellular synaptic environments via flavin-dependent photoredox catalysis. Org. Biomol. Chem. 21, 98–106 (2023).

    Article  CAS  Google Scholar 

  16. Hope, T. O. et al. Targeted proximity-labelling of protein tyrosines via flavin-dependent photoredox catalysis with mechanistic evidence for a radical-radical recombination pathway. Chem. Sci. 14, 7327–7333 (2023).

    Article  CAS  PubMed  PubMed Central  Google Scholar 

  17. Jain, M. D., Abramson, J. S. & Ansell, S. M. Easy as ABC: managing toxicities of antibody–drug conjugates, bispecific antibodies, and CAR T-cell therapies. Am. Soc. Clin. Oncol. Educ. Book 45, e473916 (2025).

    Article  PubMed  Google Scholar 

  18. Okpasuo, O. J. et al. The evolving landscape of antibody-based cancer therapies: from monospecific to multi-specific and beyond. Crit. Rev. Oncol. Hematol. 217, 105037 (2026).

    Article  PubMed  Google Scholar 

  19. Salokas, K. et al. Physical and functional interactome atlas of human receptor tyrosine kinases. EMBO Rep. 23, e54041 (2022).

    Article  CAS  PubMed  PubMed Central  Google Scholar 

  20. Leung, K. K., Schaefer, K., Lin, Z., Yao, Z. & Wells, J. A. Engineered proteins and chemical tools to probe the cell surface proteome. Chem. Rev. 125, 4069–4110 (2025).

    Article  CAS  PubMed  PubMed Central  Google Scholar 

  21. Bechtel, T. J., Reyes-Robles, T., Fadeyi, O. O. & Oslund, R. C. Strategies for monitoring cell–cell interactions. Nat. Chem. Biol. 17, 641–652 (2021).

    Article  CAS  PubMed  Google Scholar 

  22. Bartholow, T. G. et al. Photoproximity labeling from single catalyst sites allows calibration and increased resolution for carbene labeling of protein partners in vitro and on cells. ACS Cent. Sci. 10, 199–208 (2024).

    Article  CAS  PubMed  Google Scholar 

  23. Szklarczyk, D. et al. The STRING database in 2023: protein–protein association networks and functional enrichment analyses for any sequenced genome of interest. Nucleic Acids Research 51, D638–D646 (2022).

    Article  Google Scholar 

  24. Giurgiu, M. et al. CORUM: the comprehensive resource of mammalian protein complexes—2019. Nucleic Acids Res. 47, D559–D563 (2018).

    Article  Google Scholar 

  25. Oughtred, R. et al. The BioGRID interaction database: 2019 update. Nucleic Acids Res. 47, D529–D541 (2018).

    Article  Google Scholar 

  26. del Toro, N. et al. The IntAct database: efficient access to fine-grained molecular interaction data. Nucleic Acids Res. 50, D648–D653 (2021).

    Google Scholar 

  27. Hynes, R. O. Integrins: bidirectional, allosteric signaling machines. Cell 110, 673–687 (2002).

    Article  CAS  PubMed  Google Scholar 

  28. Ma, L., Pan, Q., Sun, F., Yu, Y. & Wang, J. Cluster of differentiation 166 (CD166) regulates cluster of differentiation (CD44) via NF-κB in liver cancer cell line Bel-7402. Biochem. Biophys. Res. Commun. 451, 334–338 (2014).

    Article  ADS  CAS  PubMed  Google Scholar 

  29. Dippel, V. et al. Influence of L1-CAM expression of breast cancer cells on adhesion to endothelial cells. J. Cancer Res. Clin. Oncol. 139, 107–121 (2013).

    Article  CAS  PubMed  Google Scholar 

  30. Chen, W. et al. Fibroblast activation protein (FAP)+ cancer-associated fibroblasts induce macrophage M2-like polarization via the fibronectin 1–integrin α5β1 axis in breast cancer. Oncogene 44, 2396–2412 (2025).

    Article  CAS  PubMed  Google Scholar 

  31. Chang, Y. & Finnemann, S. C. Tetraspanin CD81 is required for the alpha v beta5-integrin-dependent particle-binding step of RPE phagocytosis. J. Cell Sci. 120, 3053–3063 (2007).

    Article  CAS  PubMed  PubMed Central  Google Scholar 

  32. Miki, T. et al. The β1-integrin-dependent function of RECK in physiologic and tumor angiogenesis. Mol. Cancer Res. 8, 665–676 (2010).

    Article  CAS  PubMed  Google Scholar 

  33. Yu, M., Wang, J., Liu, S., Wang, X. & Yan, Q. Novel function of pregnancy-associated plasma protein A: promotes endometrium receptivity by up-regulating N-fucosylation. Sci. Rep. 7, 5315 (2017).

    Article  ADS  PubMed  PubMed Central  Google Scholar 

  34. Grover, A. & Leskovec, J. Node2Vec: scalable feature learning for networks. In Proc. 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining 855–864 (Association for Computing Machinery, 2016).

  35. Kipf, T. N. & Welling, M. Variational graph auto-encoders. Preprint at https://doi.org/10.48550/arXiv.1611.07308 (2016).

  36. Veličković, P. et al. Graph attention networks. In Proc. 6th International Conference on Learning Representations (OpenReview.net, 2018).

  37. Yan, W. et al. HY0001a: A novel antibody-drug conjugate (ADC) targeting CUB domain containing protein 1 (CDCP1). Cancer Res. 85, 4256–4256 (2025).

    Article  Google Scholar 

  38. Um, Y. J. et al. CDCP1-targeting ADC outperforms standard therapies in Ras-mutant pancreatic cancer. Mol. Ther. Oncol. 33, 201024 (2025).

    Article  CAS  PubMed  PubMed Central  Google Scholar 

  39. Yamada, K. et al. Proximity extracellular protein–protein interaction analysis of EGFR using AirID-conjugated fragment of antigen binding. Nat. Commun. 14, 8301 (2023).

    Article  ADS  PubMed  PubMed Central  Google Scholar 

  40. Tong, F., Zhou, W., Janiszewska, M. & Seath, C. P. Multiprobe photoproximity labeling of the EGFR interactome in glioblastoma using red-light. J. Am. Chem. Soc. 147, 9316–9327 (2025).

    Article  ADS  CAS  PubMed  PubMed Central  Google Scholar 

  41. Lin, Z. et al. Temporal photoproximity labeling of ligand-activated EGFR neighborhoods using MultiMap. Nat. Chem. Biol. 22, 192–204 (2026).

    Article  CAS  PubMed  Google Scholar 

  42. Wee, P. & Wang, Z. Epidermal growth factor receptor cell proliferation signaling pathways. Cancers 9, 52 https://doi.org/10.3390/cancers9050052 (2017).

    Article  CAS  PubMed  PubMed Central  Google Scholar 

  43. Khan, T., Kryza, T., Lyons, N. J., He, Y. & Hooper, J. D. The CDCP1 signaling hub: a target for cancer detection and therapeutic intervention. Cancer Res. 81, 2259–2269 (2021).

    Article  CAS  PubMed  Google Scholar 

  44. Murakami, Y. et al. AXL/CDCP1/SRC axis confers acquired resistance to osimertinib in lung cancer. Sci. Rep. 12, 8983 (2022).

    Article  ADS  CAS  PubMed  PubMed Central  Google Scholar 

  45. Law, M. E. et al. CUB domain-containing protein 1 and the epidermal growth factor receptor cooperate to induce cell detachment. Breast Cancer Res. 18, 80 (2016).

    Article  PubMed  PubMed Central  Google Scholar 

  46. Sun, T. et al. Activation of multiple proto-oncogenic tyrosine kinases in breast cancer via loss of the PTPN12 phosphatase. Cell 144, 703–718 (2011).

    Article  CAS  PubMed  PubMed Central  Google Scholar 

  47. Müller, S. & Jücker, M. The functional roles of the Src homology 2 domain-containing inositol 5-phosphatases SHIP1 and SHIP2 in the pathogenesis of human diseases. Int. J. Mol. Sci. 25, 5254 https://doi.org/10.3390/ijms25105254 (2024).

    Article  CAS  PubMed  PubMed Central  Google Scholar 

  48. Komiya, Y. et al. The Rho guanine nucleotide exchange factor ARHGEF5 promotes tumor malignancy via epithelial-mesenchymal transition. Oncogenesis 5, e258 (2016).

    Article  CAS  PubMed  PubMed Central  Google Scholar 

  49. Song, S. et al. ASAP1 promotes epithelial to mesenchymal transition by activating the TGFβ pathway in papillary thyroid cancer. Cancer Med. 14, e71075 (2025).

    Article  CAS  PubMed  PubMed Central  Google Scholar 

  50. Qi, X. et al. CDCP1: A promising diagnostic biomarker and therapeutic target for human cancer. Life Sci. 301, 120600 (2022).

    Article  CAS  PubMed  Google Scholar 

  51. Gusenbauer, S., Vlaicu, P. & Ullrich, A. HGF induces novel EGFR functions involved in resistance formation to tyrosine kinase inhibitors. Oncogene 32, 3846–3856 (2013).

    Article  CAS  PubMed  Google Scholar 

  52. Gough, M. et al. Receptor CDCP1 is a potential target for personalized imaging and treatment of poor outcome HER2+, triple-negative, and metastatic ER+/HER2− breast cancers. Clin. Cancer Res. 31, 1504–1519 (2025).

    Article  CAS  PubMed  Google Scholar 

  53. Zhao, N. et al. CUB domain-containing protein 1 (CDCP1) is a target for radioligand therapy in castration-resistant prostate cancer, including PSMA null disease. Clin. Cancer Res. 28, 3066–3075 (2022).

    Article  CAS  PubMed  PubMed Central  Google Scholar 

  54. Karachaliou, N. et al. Common co-activation of AXL and CDCP1 in EGFR-mutation-positive non-small cell lung cancer associated with poor prognosis. EBioMedicine 29, 112–127 (2018).

    Article  PubMed  PubMed Central  Google Scholar 

  55. Qi, X. et al. Mechanistic insights into CDCP1 clustering on non-small-cell lung cancer membranes revealed by super-resolution fluorescent imaging. iScience 26, 106103 (2023).

    Article  ADS  CAS  PubMed  PubMed Central  Google Scholar 

  56. Dong, Y. et al. The cell surface glycoprotein CUB domain-containing protein 1 (CDCP1) contributes to epidermal growth factor receptor-mediated cell migration. J. Biol. Chem. 287, 9792–9803 (2012).

    Article  CAS  PubMed  PubMed Central  Google Scholar 

  57. Zhao, F. et al. Hijacking extracellular targeted protein degrader-drug conjugates for enhanced drug delivery. J. Am. Chem. Soc. 147, 39912–39925 (2025).

    Article  ADS  CAS  PubMed  PubMed Central  Google Scholar 

  58. Li, M. M. et al. Contextual AI models for single-cell protein biology. Nat. Methods 21, 1546–1557 (2024).

    Article  CAS  PubMed  PubMed Central  Google Scholar 

  59. Floyd, B. M. et al. Mapping the nanoscale organization of the human cell surface proteome reveals new functional associations and surface antigen clusters. Preprint at bioRxiv https://doi.org/10.1101/2025.02.12.637979 (2025).

  60. Delaveris, C. S. et al. Autophagolysosomal exocytosis inverts Src kinase onto the cell surface in cancer. Science 391, eaec1778 (2026).

    Article  CAS  PubMed  Google Scholar 

  61. Goldstein, N. I., Giorgio, N. A., Jones, S. T. & Saldanha, J. W. Humanized anti-EGF receptor monoclonal antibody. US patent 7,060,808 (2006).

  62. Ren, H. et al. Anti-cub domain-containing protein 1 (CDCP1) antibodies, antibody drug conjugates, and methods of use thereof. US patent 11,702,481 (2023).

  63. Labrijn, A. F. et al. Controlled Fab-arm exchange for the generation of stable bispecific IgG1. Nat. Protoc. 9, 2450–2463 (2014).

    Article  CAS  PubMed  Google Scholar 

  64. Bissonnette, N. B. et al. Design of a multiuse photoreactor to enable visible-light photocatalytic chemical transformations and labeling in live cells. ChemBioChem. 21, 3555–3562 (2020).

    Article  CAS  PubMed  PubMed Central  Google Scholar 

  65. Erickson, B. K. et al. Active instrument engagement combined with a real-time database search for improved performance of sample multiplexing workflows. J. Proteome Res. 18, 1299–1306 (2019).

    Article  CAS  PubMed  PubMed Central  Google Scholar 

  66. McAlister, G. C. et al. MultiNotch MS3 enables accurate, sensitive, and multiplexed detection of differential expression across cancer cell line proteomes. Anal. Chem. 86, 7150–7158 (2014).

    Article  ADS  CAS  PubMed  PubMed Central  Google Scholar 

  67. Hughes, C. S. et al. Single-pot, solid-phase-enhanced sample preparation for proteomics experiments. Nat. Protoc. 14, 68–85 (2019).

    Article  CAS  PubMed  Google Scholar 

  68. Wiśniewski, J. R., Hein, M. Y., Cox, J. & Mann, M. A “proteomic ruler” for protein copy number and concentration estimation without spike-in standards. Mol. Cell. Proteomics 13, 3497–3506 (2014).

    Article  PubMed  PubMed Central  Google Scholar 

  69. Ghandi, M. et al. Next-generation characterization of the Cancer Cell Line Encyclopedia. Nature 569, 503–508 (2019).

    Article  ADS  CAS  PubMed  PubMed Central  Google Scholar 

  70. Love, M. I., Huber, W. & Anders, S. Moderated estimation of fold change and dispersion for RNA-seq data with DESeq2. Genome Biol. 15, 550 (2014).

    Article  PubMed  PubMed Central  Google Scholar 

  71. Vivian, J. et al. Toil enables reproducible, open source, big biomedical data analyses. Nat. Biotechnol. 35, 314–316 (2017).

    Article  ADS  CAS  PubMed  PubMed Central  Google Scholar 

  72. Dempster, J. M. et al. Chronos: a cell population dynamics model of CRISPR experiments that improves inference of gene fitness effects. Genome Biol. 22, 343 (2021).

    Article  PubMed  PubMed Central  Google Scholar 

  73. Bausch-Fluck, D. et al. The in silico human surfaceome. Proc. Natl Acad. Sci. USA 115, E10988–E10997 (2018).

    Article  ADS  CAS  PubMed  PubMed Central  Google Scholar 

  74. Milacic, M. et al. The Reactome Pathway Knowledgebase 2024. Nucleic Acids Res. 52, D672–D678 (2023).

    Article  Google Scholar 

  75. Consortium, T. U. UniProt: the Universal Protein Knowledgebase in 2025. Nucleic Acids Res. 53, D609–D617 (2024).

    Article  Google Scholar 

  76. Dixon, A. S. et al. NanoLuc complementation reporter optimized for accurate measurement of protein interactions in cells. ACS Chem. Biol. 11, 400–408 (2016).

    Article  ADS  CAS  PubMed  Google Scholar 

  77. Mazor, Y. et al. Enhanced tumor-targeting selectivity by modulating bispecific antibody binding affinity and format valence. Sci. Rep. 7, 40098 (2017).

    Article  ADS  CAS  PubMed  PubMed Central  Google Scholar 

  78. Ridgway, J. B., Presta, L. G. & Carter, P. Knobs-into-holes’ engineering of antibody CH3 domains for heavy chain heterodimerization. Protein Eng. 9, 617–621 (1996).

    Article  CAS  PubMed  Google Scholar 

  79. Tumey, L. N. Antibody-Drug Conjugates: Methods and Protocols, 1st edn (Springer, 2020).

  80. Scandore, C. & Guernsey, J. Proximity-guided graph learning reveals tumor-associated proximity antigens - dataset. Zenodo https://doi.org/10.5281/zenodo.21678744 (2026).

  81. Scandore, C. & Guernsey, J. Proximity-guided graph learning reveals tumor-associated proximity antigens - software. Zenodo https://doi.org/10.5281/zenodo.21679069 (2026).

Download references

Acknowledgements

The authors thank the members of InduPro for helpful discussions and technical support, and Y. Zheng for assistance with scientific illustrations.

Funding

This work received no external funding.

Author information

Author notes

  1. These authors contributed equally: Cody Scandore, Clare F. Malone, Christopher K. May

Authors and Affiliations

  1. InduPro, Cambridge, MA, USA

    Cody Scandore, Clare F. Malone, Christopher K. May, Jeff Guernsey, Noah Dephoure, Rebecca A. Howell, Kendall R. Johnson, Lydia Vignale, Tali Vittum, Emma Dawson, Francesca Nardi, Quynh Ton, Heath E. Klock, Ertan Eryilmaz, Rob C. Oslund & Olugbeminiyi O. Fadeyi

  2. InduPro, Seattle, WA, USA

    Anna K. de Regt, Hayley Ma, Ben Setter, Carol L. Farr, Sophia Romero, Tsadik Habtetsion, Brian Woodruff, Martin Mathay, Julia Swanson, Mikaela Rusnak, Payam E. Farahani, Robert W. Gene, Jason Misurelli, Zach Caldwell, Hengyu Xu, Michael Hornsby, Marc A. Gavin, Pamela M. Holland & Scott A. Lesley

Authors

  1. Cody Scandore
  2. Clare F. Malone
  3. Christopher K. May
  4. Anna K. de Regt
  5. Jeff Guernsey
  6. Hayley Ma
  7. Noah Dephoure
  8. Ben Setter
  9. Rebecca A. Howell
  10. Kendall R. Johnson
  11. Carol L. Farr
  12. Sophia Romero
  13. Lydia Vignale
  14. Tali Vittum
  15. Emma Dawson
  16. Tsadik Habtetsion
  17. Francesca Nardi
  18. Brian Woodruff
  19. Martin Mathay
  20. Julia Swanson
  21. Mikaela Rusnak
  22. Quynh Ton
  23. Payam E. Farahani
  24. Robert W. Gene
  25. Jason Misurelli
  26. Zach Caldwell
  27. Hengyu Xu
  28. Michael Hornsby
  29. Marc A. Gavin
  30. Heath E. Klock
  31. Ertan Eryilmaz
  32. Pamela M. Holland
  33. Scott A. Lesley
  34. Rob C. Oslund
  35. Olugbeminiyi O. Fadeyi

Contributions

O.O.F. and R.C.O. conceived of the work. C.K.M., N.D., R.A.H., K.R.J., L.V., T.V. and E.D. designed and performed micromapping and proteomic experiments. C.S. and J.G. designed and executed bioinformatic and machine learning analysis. C.S., C.F.M., J.G., F.N., R.C.O. and O.O.F. conceived and developed the multimodal target prioritization workflow. B.S., C.L.F., B.W., M.M., J.S., Q.T., P.E.F., R.W.G., J.M., Z.C., H.X., M.H. and H.E.K. designed, engineered, characterized, and generated ADCs and TCEs. C.F.M., A.K.d.R., H.M., S.R., T.H., F.N., M.R. and M.A.G. designed, executed and performed in vitro and in vivo ADC and/or TCE evaluation. E.E., P.M.H. and S.A.L. provided insight and direction for experimental design. O.O.F., R.C.O. and C.S. wrote the manuscript with input from all authors.

Corresponding authors

Correspondence to Rob C. Oslund or Olugbeminiyi O. Fadeyi.

Ethics declarations

Competing interests

All authors were/are employed by InduPro during the experimental planning, execution and/or preparation of this manuscript.

Peer review

Peer review information

Nature thanks Kevin Leung, James Wells, and the other, anonymous reviewers for their contribution to the peer review of this work. Peer reviewer reports are available.

Additional information

Publisher’s note Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.

Extended data figures and tables

Extended Data Fig. 1 Reciprocal EGFR-CDCP1 enrichment is reproduced across micromapping formats, binders and cell systems.

a) Schematic of targeted photocatalytic micromapping of EGFR or CDCP1 and capture of the reciprocal proximal protein. b) Enrichment of CDCP1 in EGFR-targeted micromaps and EGFR in CDCP1-targeted micromaps, shown as log2 fold change values relative to an isotype-matched non-specific IgG control. Each point represents an individual micromapping measurement using riboflavin/biotin-tyramide chemistry (RFT), iridium/diazirine chemistry (Ir), or the combined dataset (ALL). In all box plots, the center line denotes the median (50th percentile), box bounds denote the 25th and 75th percentiles, and whiskers extend to the minimum and maximum values. All individual data points are overlaid. Numbers in parentheses above each box plot indicate the sample size (n, independent measurements). c) Schematic of EGFR- or CDCP1-targeted micromapping using directly conjugated monovalent antibody–photocatalyst conjugates. d) Volcano plots from two independent EGFR-targeted micromapping experiments in HCC827 lung adenocarcinoma cells using directly conjugated cetuximab-photocatalyst conjugates. Cetuximab recognizes a different EGFR epitope from the antibody used in Fig. 2a and Supplementary Figs. 5 and 8 (Supplementary Fig. 33). EGFR and CDCP1 are highlighted in orange and green, respectively, and were significantly co-enriched in each experiment (P < 0.05 and log2 fold change > 1.5; n = 3 technical replicates per experiment; P values from a two-sided moderated t-test with adjustment for multiple comparisons). e) Volcano plots from two independent CDCP1-targeted micromapping experiments in NCIH1650 lung adenocarcinoma cells using directly conjugated CDCP1 41A9-photocatalyst conjugates. CDCP1 41A9 recognizes a different CDCP1 epitope from the antibody used in Supplementary Fig. 19 (Supplementary Fig. 33). EGFR and CDCP1 are highlighted in orange and green, respectively, and were significantly co-enriched in each experiment (P < 0.05 and log2 fold change > 0.6; n = 3 technical replicates per experiment; P values from a two-sided moderated t-test with adjustment for multiple comparisons).

Source data

Extended Data Fig. 2 Study design and tumor growth in BxPC3 and SW48 xenograft studies.

a) Study design for the BxPC3 xenograft study. b) Tumor growth of BxPC3 xenografts following treatment with the indicated test articles. Data are mean tumor volume ± s.e.m. (n = 10 mice per group); arrows denote dosing days. c) Study design for the SW48 dual-flank xenograft study. d) Tumor growth of SW48 parental and CDCP1-KD xenograft tumors treated with the indicated test articles. Data are mean tumor volume ± s.e.m. (n = 8 mice per group). Two-way ANOVA with Tukey’s multiple comparison; *** P = 0.0006. DAR2, drug-antibody ratio of 2; DAR~3, drug-antibody ratio of approximately 3; i.v., intravenous; KD, knockdown; TV, tumor volume; wk, week.

Source data

Extended Data Fig. 3 Tumor antigen expression, survival and body weight in xenograft efficacy studies.

a) Representative EGFR and CDCP1 expression assessed by immunohistochemistry (IHC) in SW48 xenograft tumors. Images are representative of three independent tumors. b) Kaplan-Meier survival curves for mice shown in Fig. 5e (n = 8 mice per group). A tumor volume of 800 mm3 was defined as the survival endpoint. Statistical significance was assessed using a logrank (Mantel-Cox) test; **** P < 0.0001. c) Mean body weight ± s.e.m. for mice shown in Fig. 5e (n = 8 mice per group). d) Representative EGFR and CDCP1 expression assessed by IHC in BxPC3 xenograft tumors. Images are representative of three independent tumors. e) Kaplan-Meier survival curves for mice shown in Extended Data Fig. 2b (n = 10 mice per group). A tumor volume of 800 mm3 was defined as the survival endpoint. Statistical significance was assessed using a logrank (Mantel-Cox) test; *** P < 0.001. f) Mean body weight ± s.e.m. for mice shown in Extended Data Fig. 2b (n = 10 mice per group). g) Representative EGFR and CDCP1 expression assessed by IHC in SW900 xenograft tumors. Images are representative of three independent tumors. h) Kaplan-Meier survival curves for mice shown in Fig. 5g (n = 9 mice per group). A tumor volume of 800 mm3 was defined as the survival endpoint. Statistical significance was assessed using a logrank (Mantel-Cox) test; **** P < 0.0001. i) Mean body weight ± s.e.m. for mice shown in Fig. 5g (n = 9 mice per group). DAR2, drug-antibody ratio of 2.

Source data

Supplementary information

Supplementary Information (download PDF )

This file contains 36 Supplementary Figures, nine Supplementary Notes, and Supplementary References. The figures cover cell-system selection and controls, large-scale micromapping and downstream bioinformatic analyses, EGFR–CDCP1-focused micromapping and proteomics, ADC and TCE binder characterization, concurrent binding, CDCP1-knockout cell characterization, and Fc pharmacokinetics. The notes detail graph learning methods and computational pseudocode.

Reporting Summary (download PDF )

Supplementary Table 1 (download XLSX )

RTK micromap proximity data. This workbook contains protein-level results from 248 RTK micromapping experiments, including target and proximal protein identifiers, map ID, labelling chemistry, cell line, log2 fold change, and negative log10-transformed P values.

Supplementary Table 2 (download XLSX )

Tissue sample patient information for proteomic analysis. This workbook contains de-identified patient information from commercially obtained tissue samples used for proteomic analysis including vendor-reported sex/gender, race and ethnicity, donor age at collection, diagnosis, histology, anatomical site, disease stage, smoking history and primary versus recurrent specimen status.

Supplementary Table 3 (download XLSX )

DIA proteomics abundance data. This workbook contains data-independent acquisition proteomic abundance measurements across the profiled cell lines, including protein identifiers, mass-normalized intensity values, and estimated copies per cell calculated using the proteomic ruler method.

Supplementary Table 4 (download XLSX )

NanoBiT expression plasmids. This workbook contains DNA and protein sequences for EGFR and CDCP1 NanoBiT constructs.

Supplementary Data (download ZIP )

Source data for Supplementary Figures.

Peer Review File (download PDF )

Source data

About this article

Check for updates. Verify currency and authenticity via CrossMark

Cite this article

Scandore, C., Malone, C.F., May, C.K. et al. Proximity-guided graph learning reveals tumour-associated proximity antigens. Nature (2026). https://doi.org/10.1038/s41586-026-11003-7

Download citation

  • Received:

  • Accepted:

  • Published:

  • Version of record:

  • DOI: https://doi.org/10.1038/s41586-026-11003-7