Main
Cooperation is widespread across social species2,3 and is foundational to human society4. Among the most common strategies for organizing cooperative behaviour are leadership and followership, roles that allow individuals to exert asymmetric influence over one another in pursuit of shared goals1,5,6,7,8,9,10,11. However, a robust and ethologically grounded paradigm for studying leader–follower dynamics is still lacking in mice, posing a critical barrier to harnessing modern neuroscience tools in this model organism12,13. Consequently, the neural mechanisms by which social roles are instantiated as distinct behavioural strategies and shape coordinated decision-making remain largely unclear14.
In contrast to discrete turn-based games15,16,17, naturalistic social interactions unfold continuously, requiring moment-to-moment coordination18. Moreover, an animal’s underlying goals and strategies are often unobservable, particularly during reciprocal interactions between multiple individuals. Although inverse reinforcement learning (IRL) provides a principled framework for inferring internal models from observed behaviour19,20, extending IRL to multi-agent biological systems introduces substantial complexity. These challenges have limited our understanding of how neural circuits construct role-specific models of a social partner and implement the differential influence that characterizes leader–follower dynamics.
To address these questions, we developed a cooperative foraging paradigm in mice and a multi-agent IRL (MAIRL) framework. We find that stable leader and follower roles emerge spontaneously through reciprocal interaction, and that the mPFC encodes not only leading versus following but also a role-specific, egocentric social value map of the partner that mirrors MAIRL-inferred goals. These findings reveal how leader–follower dynamics are instantiated as asymmetric prefrontal representations, linking role-based decision-making to egocentric partner encoding at the cellular resolution.
Emergent leader–follower roles
In the cooperative foraging paradigm, pairs of water-restricted, same-sex cage mates freely interact in a square arena to obtain water rewards (Fig. 1a and Supplementary Video 1). Each arena wall features a reward zone with two ports, enabling simultaneous access. A trial begins when either mouse crosses a central initiation point, triggering random activation of two of four reward zones indicated by light-emitting diode (LED) lights. To earn a reward, both mice must navigate to the same active zone and make concurrent nose pokes into the two ports. Trials end in error if mice poke concurrently in different active zones (mismatch error) or either mouse pokes in an inactive zone (unrewarded error), resulting in no reward and a brief timeout (Methods). All but 1 of the 89 pairs reached the training criterion of 80% correct trials across three consecutive sessions. Female mice learned faster and performed slightly better during the criterion sessions than male mice (Fig. 1b and Extended Data Fig. 1a–f). Notably, with training, mouse pairs significantly improved the synchronization of their nose pokes at the reward ports, even though precise timing was not required, as one mouse could wait for the other by holding its nose poke (Extended Data Fig. 1g).
a, Task schematic. The diagram was created using BioRender; Wu, H. https://BioRender.com/hwyq5u6 (2026). b, Learning curves showing the cumulative proportion of pairs reaching the criterion; female mice were faster than male mice (one session per day). c, Representative session of a well-trained mouse pair, mouse 1 (m1) and mouse 2 (m2). Right, trials initiated (init.) and led per mouse. d–f, An example pair: trials led and initiated by m2, and the cooperation rate across training days (d), the cooperation rate versus leader asymmetry (e) and initiator asymmetry (f). n = 15 sessions. g,h, Cooperation rate scales with leader asymmetry (g) and, more weakly, initiator asymmetry (h). Each point represents one session; the thin lines indicate individual-pair fits; the thick line shows the combined fit. n = 681 sessions, 40 pairs. i, Leader asymmetry at the criterion predicts learning. Each point represents one pair (40 pairs). j,k, The within-session cooperation rate is higher on leader-led trials than on follower-led trials (j; 673 sessions) and on initiator-initiated trials than on responder-initiated trials (k; 676 sessions), both from 40 pairs. The dashed line indicates unity. l, The proportion of trials led by the leader in each well-trained pair (n = 678–4,766 trials per pair; 32 pairs). The lighter points represent individual sessions. Data are mean ± s.e.m. m, Trials led versus initiated by each mouse, dissociating leader–follower from initiator–responder roles; leaders initiate no more than chance (n = 32 pairs, P = 0.45). n, The proportions of trials led and initiated by the dominant mouse do not differ from chance (initiation, P = 0.24; leadership, P = 0.29; n = 22 pairs); dominance does not predict the joint leader–initiator distribution (P > 0.12). o, Well-trained leaders that are dominant or subordinate; dominance does not predict leadership at any stage (n = 14 leaders, P > 0.17). p, Trials led in the original versus swapped pairing; 12 out of 14 leaders maintained their role (P = 0.0065). q, The cumulative proportion of mice reaching the criterion; swapped pairs learn faster than the original pairs (n = 16 pairs). r, Trials led by m1 versus m2 when two leaders or two followers are paired. The arrows link the original to new roles, with one mouse reorganizing to restore a leader–follower split in 13 out of 16 pairs. Statistical analysis was performed using Kolmogorov–Smirnov tests (b and q), Pearson correlation (e, f and i), linear mixed-effects (g, h, j and k), linear regression (i), binomial tests (l–p) and Wilcoxon signed-rank tests (m and n); tests were two-tailed except for one-tailed binomial in p. #0.05 ≤ P < 0.10, *P < 0.05, **P < 0.01, ***P < 0.001; NS, not significant. The same conventions are used throughout. Details are provided in Supplementary Table 1.
Source data
Pairs developed two independent role structures: the mouse initiating more trials in a session is the initiator, and its partner is the responder; the mouse arriving first at the reward zone in most trials is the leader, its partner is the follower (Methods, Fig. 1c and Extended Data Fig. 1h–j). A follower still leads, and a responder still initiates, on a minority of trials. Leaders and initiators emerged in 90% (36 out of 40) and 87.5% (35 out of 40) of well-trained pairs, respectively (two-tailed binomial tests, α = 0.01). Leadership defined by arrival order closely aligned with an independent measure based on the timing of initial orientation towards the chosen reward zone (Extended Data Fig. 1k). During learning, leader asymmetry (absolute difference in the proportion of trials led by the two mice) strongly correlated with performance (Fig. 1d–h and Extended Data Fig. 1l–p), suggesting that leader–follower differentiation facilitates cooperation. Moreover, leader asymmetry at the criterion predicted the speed of learning across pairs (Fig. 1i). Cooperation rate was higher in trials led by the leader and initiated by the initiator (Fig. 1j,k and Extended Data Fig. 1q,r), indicating a behavioural advantage from role adherence. Nonetheless, performance also improved substantially in follower-led and responder-initiated trials, suggesting flexible accommodation of deviations from the session-level role structure (Extended Data Fig. 2a,b). Once established, these roles remained stable across sessions (Fig. 1l). Although some initiators appeared less likely to lead (Fig. 1d and Extended Data Fig. 1l), these two roles were overall dissociable (Fig. 1m and Extended Data Fig. 2c,d). Social roles were not predicted by performance metrics during early solo foraging (Extended Data Fig. 2e–g). We further tested dominance status using the tube test21, warm spot test22,23 and reward competition24,25. Neither leadership nor initiatorship correlated with dominance at well-trained or earlier stages (Fig. 1n,o and Extended Data Fig. 3a–u), indicating that these roles constitute distinct social structures independent of hierarchy.
To test whether social roles persist across partners, we swapped two well-trained pairs such that the leader of one pair was teamed with the follower of another. In both sexes, most mice maintained their original roles in the new pairings, demonstrating that leader–follower identity is a stable construct carried across contexts (Fig. 1p and Extended Data Fig. 4a,b). Some new pairs reached the criterion immediately after swapping, while others required adaptation time, although less than during initial learning (Fig. 1q and Extended Data Fig. 4c,d). When two leaders or two followers were paired, one mouse adopted the complementary role to enable cooperation (Fig. 1r and Extended Data Fig. 4e,f). Dominance status did not predict social roles or leader–follower switching in new pairings (Extended Data Fig. 4g–k). Together, these results underscore the social nature of the task: leader–follower identities are stable across partners yet flexible when required.
Well-trained mice adopted stereotyped movement patterns, or behavioural motifs, identified through empirical observation and manual annotation (Methods, Fig. 2a,b, Extended Data Fig. 5a,b and Supplementary Videos 2–5). These motifs exhibited consistent and distinct movement signatures across behavioural metrics (Extended Data Fig. 6a–g) and corresponded closely to structure revealed by unsupervised behavioural classification26 (Fig. 2c and Extended Data Fig. 7a–e). All motifs involve pair-level interaction—three (track, sharp turn, join) are annotated with respect to the actor, defined as the mouse performing the action, whereas the synchronized travel motif reflects a joint behaviour that is not assigned to either individual. For example, for the track motif, a mouse slows to monitor its partner and align arrival times, and for the sharp turn motif, it abruptly redirects to rejoin its partner. Animal pairs exhibited diverse motif profiles with no consistent sex differences (Fig. 2d and Extended Data Fig. 7f,g). Leaders primarily used the join motif when following, while followers displayed a more diverse repertoire, increasing join or sharp turn motifs as following demands increased (Extended Data Fig. 7h–s). These results suggest that mouse pairs develop individual, role-specific strategies that flexibly adapt to changing demands for cooperation.
a, Representative behavioural motifs. Top, video frames; the actor mouse is shown in orange (none for synchronized travel (sync)). Bottom, trajectories from one session. Scale bar, 10 cm. b, Motif sequences from example sessions. c, Unsupervised LISBET embeddings across sessions, coloured by manual motif labels (as in b). UMAP, uniform manifold approximation and projection. d, Motif use across 20 pairs (2–4 sessions each), ordered by hierarchical clustering. e, Choices on self-initiated trials, solo versus cooperative. θ, heading (HD), the angle of the mouse’s neck-to-nose axis relative to the east. f,g, P(north), the probability of choosing north, versus heading at trial onset, conditioned on partner heading or solo, for followers (f) and leaders (g) on self-initiated trials. For f–i, n = 8 pairs. h,i, Partner presence reduces leader and follower sensitivity to their own heading (h). An east/south-facing partner decreases, and a west/north-facing partner increases, follower north-zone bias, while a west/north-facing partner increases leader bias (i). Data are coefficient ± s.e. n = 4,991–5,129 (follower) and 3,881–3,963 (leader) trials per condition. j, Choices on partner-initiated trials, conditioned on partner heading. k, The P(north) for leaders (top) and followers (bottom) across starting positions: partner facing north (left) or east (right). Scale bar, 8 cm. For k–r, n = 38 pairs. l, ΔP(north), the difference in P(north) between the partner facing north and east, for leaders and followers. m, Both leaders and followers choose north more when their partner faces north (n = 36 bins). n,o, P(north) versus distance difference to east and north zones (E–N), conditioned on partner heading, for leader (n) and follower (o) choice; the points indicate bins; curves represent linear fits (n = 36 bins, 4 conditions). Leader P(north) is higher when the partner faces north/west than south/east. Follower P(north) ranks are as follows: west, north, east, south. p, Spatial variables for predicting choice. q, The accuracy in predicting leader and follower choice from both animals’ heading difference (Δθ = θEast – θNorth) and distance difference (Δd = dEast – dNorth) exceeds shuffled controls (P < 1.0 × 10−4, 10,000 iterations). Data are mean ± s.d. across 100 folds. r, Decision weights from q. The lines represent individual pairs. Data are mean ± s.e.m. All weights differ from zero except for follower distance weight for leader choice, and differ between roles. s, Two-agent MARL simulation: each agent (a1, a2) observes its own and partner’s state and active zones and selects actions through a Q module. t, Asymmetry in adopting initiator (left) and leader (right) roles. The brown circles and lines represent significant pairs (183 each; n = 200 simulated pairs). Statistical analysis was performed using logistic mixed-effects regression (h and i), paired t-tests (m and r), one-sample t-tests (r), repeated-measures analysis of variance (ANOVA) with Tukey–Kramer post hoc test (n and o), bootstrap versus shuffled (q) and binomial against 0.5 (t, α = 0.01); tests were two-tailed except for the bootstrap test in q. Details are provided in Supplementary Table 1. The diagrams in e, j and p were created using BioRender; Wu, H. https://BioRender.com/hwyq5u6 (2026).
Source data
Asymmetric yet reciprocal role dynamics
We next examined how a mouse makes foraging decisions based on its social role. We first compared its reward zone choices in self-initiated trials during cooperative foraging to a solo control (Fig. 2e), in which the mouse foraged alone and was rewarded at either active zone. Trials with adjacent active zones were rotated to a common north-east configuration (Methods), and we examined how trial-by-trial choice depended on the mouse’s heading, conditioned on the partner’s heading or position at trial onset.
In solo foraging, mice in both leader and follower roles showed clear heading-dependent choices, for example, animals facing north at the trial onset were more likely to choose north (Fig. 2f,g (black curves)). In cooperative foraging, this dependence was preserved but was strongly modulated by the partner’s heading. Followers, in particular, showed reduced sensitivity to their own heading and increased bias towards the leader’s (Fig. 2f,h,i). Leaders were also influenced by the follower’s heading, showing smaller but highly significant reductions in sensitivity and corresponding biases (Fig. 2g–i). Similar effects were observed when considering the partner’s position (Extended Data Fig. 8a–d). These results indicate that both roles integrate the partner’s spatial information into decision-making, with followers showing greater sensitivity to the leader.
We next examined a mouse’s decisions on partner-initiated trials, in which it could begin from any location in the arena at the trial onset (Fig. 2j). After applying the same rotation procedure, mice in both the leader and follower roles exhibited a clear spatial gradient in choice preference, favouring the closer reward zone (Fig. 2k). This preference was strongly modulated by the partner’s heading, for example, a mouse was more likely to choose north when its partner faced north at trial onset, particularly when it was at similar distances from the north and east zones (Fig. 2k–m). The partner’s heading consistently influenced both leader and follower decisions across all directions in both sexes (Fig. 2n,o and Extended Data Fig. 8e–n).
Finally, to quantify the relative influence of spatial variables within a single framework, we used logistic regression to fit the decisions of each role using four predictors: each animal’s position and heading difference relative to the two active zones at trial onset (Fig. 2p). This regression reliably predicted zone choice for both animals (Fig. 2q). In particular, the leader’s heading and position exerted a stronger influence on the follower’s choice than its own spatial variables (Fig. 2r). Conversely, the follower’s heading, but not position, significantly influenced the leader’s decision. Similar patterns were observed in both sexes (Extended Data Fig. 8o–r) and across all spatial configurations (Extended Data Fig. 8s,t). Together, these results reveal an asymmetric yet reciprocal influence between the two roles during cooperative decision-making.
Spontaneous role emergence in simulation
To examine how coordinated strategies and asymmetric roles might arise from shared incentives alone, we first modelled cooperative behaviour using forward multi-agent reinforcement learning (MARL). Each agent independently learned a value function from experience using a curriculum mirroring mouse training, without prior biases or explicit role assignments (Fig. 2s and Methods). All simulated agent pairs learned the task, achieving accuracy comparable to or exceeding that of mice (Extended Data Fig. 9a–c). Despite identical initializations, behavioural asymmetries reliably emerged across pairs (Extended Data Fig. 9d). During testing, 183 out of 200 pairs established statistically significant initiators and 183 consistent leaders (Fig. 2t). As in mice, fewer initiators than responders were also leaders, but this difference was not significant, indicating that initiatorship and leadership are dissociable (Extended Data Fig. 9e). These results suggest that asymmetric social roles can self-organize through interaction alone, probably driven by stochastic fluctuations in learning that are amplified by coordination demands, echoing the emergence of stable roles among genetically identical mice.
The mPFC supports role-specific cooperation
The frontal cortex is implicated in social cognition across primates27,28 and rodents29,30,31,32, and both the mPFC and orbitofrontal cortex (OFC) encode cooperation-related variables33,34, yet their causal contributions to social role dynamics remain unclear. We examined the causal roles of the mPFC and OFC using the inhibitory designer receptor exclusively activated by designer drug (DREADD) hM4D(Gi), validated by in vivo recordings (Fig. 3a and Extended Data Fig. 10a–c). We inactivated mainly female pairs, given similar behaviour across sexes and the re-pairing constraints of post-surgical males (Methods). Bilateral inactivation of the mPFC in both animals, but not the OFC, disrupted cooperation (Fig. 3b and Extended Data Fig. 10d,e), driven by an increase in mismatch errors rather than changes in motivation or movement (Extended Data Fig. 10f–h). mPFC inactivation also diminished leader asymmetry (Fig. 3c) without affecting initiator asymmetry (Fig. 3d). Cooperation declined on both leader- and follower-led trials (Fig. 3e,f), suggesting a general impairment in coordination. Only followers showed a significant reduction in reaction time (Fig. 3g,h and Methods), possibly reflecting less deliberation. Logistic regression revealed that during inactivation, followers placed greater decision weights on their own heading and less on leader position, whereas leader strategies remained unchanged (Fig. 3i). The impairment was similar when inhibiting predominantly excitatory neurons (Extended Data Fig. 11a,b), and performance remained intact in the controls (Extended Data Fig. 11c,d).
a, mPFC chemogenetic inactivation. For a–m, n = 11 pairs unless otherwise noted. The thin lines/points represent individual data; the thick lines/bars show the median. For decision-weight panels (i, k and m), data are coefficient ± s.e. CNO, clozapine N-oxide; Ctrl, control. b–h, Inactivating the mPFC in both roles reduces the cooperation rate (b), leader asymmetry (c), leader-led (e) and follower-led (f) cooperation rate, and follower reaction time (g), but not initiator asymmetry (d) or leader reaction time (h). i, Both-role inactivation decreases leader distance weight and increases follower heading weight in follower decisions (6,798 trials). j, Selective inactivation of the leader mPFC does not affect the cooperation rate, leader asymmetry or initiator asymmetry. k, Leader mPFC inactivation decreases follower distance weight in both leader and follower decisions (6,691 trials). l, Selective inactivation of follower mPFC decreases cooperation rate and trends towards lower leader asymmetry, with no effect on initiator asymmetry. m, Follower mPFC inactivation decreases leader heading weight and increases follower heading weight in follower decisions, and increases follower heading weight in leader decisions (6,805 trials). n, Non-social-stimulus tracking. n = 8 (tracking mice), n = 32 (social followers) and n = 10 (non-social inactivation). The bars show the median. o, The non-social tracking correct rate. The dashed line indicates the correct rate on source cooperative trials (0.93). p, The proportion of trials mice follow versus lead the non-social stimulus. q, The following probability for tracking mice versus social followers. r,s, The correct rate of tracking mice versus social followers, on following (r) or leading (s) trials. t, mPFC inactivation in the non-social task does not affect the correct rate or following probability. u,v, Example follower (u) and leader (v) neurons. Left, trial-by-trial ΔF/F (recorded mouse leads (L)/follows (F)). Purple/green ticks indicate leader/follower arrival. Right, data are mean ± s.e.m. n = 4 (leaders) and n = 4 (followers). w, The reward zones chosen when leading or following. Data are mean ± s.e.m. No effect of zone, arrival order or interaction. x, The arrival-order selectivity index (SI) for follower (6,995) and leader (8,395) neurons. Shading shows selective neurons. More follower neurons prefer leading, and more leader neurons prefer following (both P < 1.1 × 10−27). y, The arrival-order SI over time per selective neuron in followers (1,399) and leaders (1,216), grouped by follow preferring (F) or lead preferring (L). z, Logistic-regression decoding of arrival order. Data are mean ± s.e.m. across animals; shuffled pools for both groups. auROC, area under the receiver operating characteristic curve. Statistical analysis was performed using Wilcoxon signed-rank tests (b–h, j, l, o, p and t), logistic mixed-effects regression (i, k and m), Wilcoxon rank-sum tests (q–s and x), repeated-measures ANOVA (w) and z-test for proportions (x) (all two-tailed); and permutation tests for selectivity (x and y); α = 0.05. Details are provided in Supplementary Table 1. The diagrams in a and n were created using BioRender; Wu, H. https://BioRender.com/hwyq5u6 (2026).
Source data
To dissect role-specific contributions, we selectively inactivated the mPFC in either the leader or the follower. Inactivation of the mPFC in leaders did not significantly impair task performance or disrupt the role dynamics (Fig. 3j and Extended Data Fig. 11e). Leaders showed reduced weighting of follower distance in decision-making, paralleled by a similar reduction in weighting of self-distance by non-inactivated followers (Fig. 3k), possibly indicating a compensatory adjustment. By contrast, inactivation of the follower mPFC reduced cooperation rates, with a trend towards impaired role dynamics (Fig. 3l and Extended Data Fig. 11f). Inactivated followers relied more on their own heading and less on partner cues (Fig. 3m), as seen during the inactivation of both roles. Notably, the non-inactivated leaders increased sensitivity to the follower’s heading, suggesting compensatory adjustments to maintain coordination.
We next used wireless optogenetics to achieve greater temporal precision in mPFC silencing (Extended Data Fig. 11g–l). Optogenetic inactivation of both animals throughout the trials impaired task performance (Extended Data Fig. 11m), consistent with chemogenetic results, but did not disrupt leadership or initiatorship (Extended Data Fig. 11n,o). Both roles showed reduced weighting of position and heading cues (Extended Data Fig. 11p). Brief inactivation for 1–2 s from the trial onset was sufficient to impair performance, highlighting a critical early decision period (Extended Data Fig. 11q). However, silencing either role alone did not impair behaviour (Extended Data Fig. 11r), suggesting that an intact mPFC in either animal can sustain coordination with a perturbed partner. These results corroborate the essential role of the mPFC in cooperation and suggest that it confers resilience against transient disruption in the partner.
To test whether leader–follower dynamics require conspecific interaction, we developed a non-social-stimulus-tracking task in which a mouse tracked a moving, inanimate stimulus replaying cooperative foraging trajectories, with the reward contingent on both mice arriving at the same reward zone (Fig. 3n and Methods). The performance plateaued at 75.0% correct, below both the cooperative foraging training criterion and the 93.4% correct rate on the original trials from which the trajectories were drawn (Fig. 3o and Extended Data Fig. 12a). As the inanimate stimulus could not respond to the mouse, animals favoured a fixed following strategy (Fig. 3p), but they still followed less consistently than the followers in cooperative foraging (Fig. 3q and Extended Data Fig. 12b). The correct rate was comparable when mice followed the non-social stimulus and when followers followed their social partner (Fig. 3r). By contrast, performance was markedly worse when mice occasionally led the stimulus than when followers led in cooperative foraging, probably because the inanimate stimulus cannot reciprocally adjust like a social partner (Fig. 3s; compare with Extended Data Fig. 2a). Mice also had longer reaction times during non-social tracking (Extended Data Fig. 12c). These results demonstrate that a simple, fixed following strategy cannot reproduce the performance or bidirectional role dynamics observed in cooperative foraging. Importantly, bilateral chemogenetic inactivation of the mPFC did not affect the correct rate, following probability or movement speed in the non-social task (Fig. 3t and Extended Data Fig. 12d), indicating that the mPFC is specifically required for cooperating with conspecifics rather than tracking generic moving stimuli.
mPFC encodes leading versus following
To examine how social roles are represented in the mPFC, we recorded mPFC activity using one-photon calcium imaging from either the leader or the follower during cooperative foraging (Extended Data Fig. 12e). Many mPFC neurons encoded task-relevant signals including choice (Extended Data Fig. 12f,g), with a greater proportion of choice-selective neurons in leaders (Extended Data Fig. 12h). Population decoding of zone choice was consistently more accurate in leaders, suggesting more-robust choice representations in leadership roles (Extended Data Fig. 12i).
Notably, individual mPFC neurons encoded leading versus following, distinguishing whether the recorded mouse arrived first or second on a trial (Fig. 3u,v and Extended Data Fig. 12j,k). Such neurons were found in both roles, some preferring trials in which the animal followed, others when it led. To rule out confounding factors related to arrival order, we conducted several control analyses. First, reward zone and port choices did not differ between trials when mice arrived first or second, despite individual port preferences (Fig. 3w and Extended Data Fig. 12l). Second, approach speed and trajectories to the reward zones were similar regardless of the arrival order (Extended Data Fig. 12m,n). Third, although mice that arrive first wait longer at the reward zone, only around 20% of arrival-order-selective neurons exhibited ramping activity correlated with wait time, suggesting possible modulation by reward expectation, and this proportion did not differ from the overall population (Extended Data Fig. 12o). After regressing out speed and position contributions and stratifying trials by port choice, 20.0% of follower neurons and 14.5% of leader neurons were selective for arrival order (Fig. 3x). Notably, both groups showed a bias for the less frequent trial type, for example, more neurons in followers were selectively active when they led (Fig. 3x,y). A logistic regression decoder classified arrival order well above chance, with higher accuracy using follower mPFC activity (Fig. 3z). Together, these results suggest that the mPFC not only encodes momentary arrival order, but also embeds information about the animal’s role identity.
An egocentric partner map in the mPFC
Given that cooperative decisions depend on self and partner positions, we examined whether and how the mPFC represents these variables. Notably, we found that mPFC neurons predominantly encode partner position in egocentric rather than allocentric coordinates (Fig. 4a–g and Extended Data Fig. 13a–c). We term each neuron’s preferred egocentric partner location its social receptive field, for example, proximal front left for the neuron in Fig. 4a–d. Quantification using a set of three criteria (Methods) revealed that 8.3% of neurons encoded partner position in an egocentric reference frame, approximately four times more than those tuned to partner position in allocentric coordinates or egocentric coordinates not aligned with heading (Fig. 4e). These proportions were similar between leaders and followers, indicating that spatial encoding of self and partner is preserved across roles. At the population level, decoders trained on mPFC activity predicted both partner distance and egocentric angle (Fig. 4h–k and Extended Data Fig. 14a,b). Decoding accuracy for partner distance did not differ between roles (Fig. 4i), but partner angle decoding was significantly more accurate in followers, consistent with their greater need to orient towards leaders (Fig. 4k). Using CEBRA35, a contrastive learning framework, we projected mPFC population activity into a low-dimensional latent space to examine how partner distance and egocentric angle are organized. These embeddings revealed clear structure for both variables, significantly more organized than control embeddings trained with time-shifted labels (Fig. 4l–o). These findings corroborate the decoding results and suggest that partner information is organized into a continuous, egocentric representation in the mPFC.
a–d, The response of an example neuron across reference frames (inset) for self and partner position: allocentric self (a), allocentric partner (b), egocentric partner unaligned with HD (c), and egocentric partner aligned with HD (d). Scale bar, 10 cm (a and b), 20 cm (c and d). e, The proportion of neurons tuned to each representation. Points represent individual mice; the bars show the median values. Egocentric partner HD aligned tuning does not differ between roles (n = 6 per group) but exceeds egocentric unaligned and allocentric partner tuning (n = 12 mice). f,g, Example follower (f) and leader (g) neurons with egocentric partner tuning (as in d). Scale bar, 20 cm. h–k, Cross-validated confusion matrices for decoding partner distance (h) and angle (j) in an example leader (left) and follower (right). The decoding accuracy for partner distance does not differ between roles (i); the accuracy for partner angle is higher in followers (k). n = 6 per group. Points represent mice; the bars show the median values; the dashed line indicates chance. l–o, CEBRA multi-session embeddings. Example mouse using data (left) or time-shifted (right) partner distance (l) and partner angle (n) labels. Both distance (m) and angle (o) data are more organized than shifted controls. n = 8 animals. p, Partner-distance tuning in leaders and followers. The heat maps show the normalized activity per neuron sorted by preferred distance; the histograms show the preferred-bin distribution; the dashed lines show chance (1/7). Both groups over-represent near distances (leaders, 126 out of 268 neurons in the two proximal bins; followers, 250 out of 480 neurons). FR, firing rate. q, Partner-angle tuning as in p. Followers show a central peak (215 out of 644 neurons in the three centre bins); leaders do not (70 out of 311 neurons). r, Partner-distance tuning after retraining under the mismatch rule. The mirrored histograms show the original match rule (as in p). Leader tuning is unchanged; follower tuning shifts towards longer distances. s, Partner-angle tuning after mismatch retraining; both roles shift from frontal to rear angles. t–y, The probability of role deviation (t,u) or mismatch error (w,x) as a function of trial-by-trial activity in follower (t and w) or leader (u and x) neurons tuned to frontal (x axis) and rear (y axis) partner positions. Each pixel indicates the outcome probability when frontal responses fall below the x-axis value and rear responses exceed the y-axis value. Logistic-regression coefficients predicting role deviation (v) and mismatch error (y). Data are regression coefficient ± s.e. Trials were pooled across 82 sessions (6 leaders, 6 followers). Statistical analysis was performed using Wilcoxon rank-sum tests (e, i and k), Wilcoxon signed-rank tests (e), Kolmogorov–Smirnov tests (r and s) and logistic regression (t–y) (all two tailed); and Wilcoxon signed-rank tests (m and o) and binomial tests (p and q) (all one tailed). Details are provided in Supplementary Table 1. The diagrams in a–d were created using BioRender; Wu, H. https://BioRender.com/hwyq5u6 (2026).
Source data
We next examined the tuning of these egocentric partner-selective neurons. In both roles, mPFC neurons selective for partner distance showed a strong preference for close-range positions (Fig. 4p). While the tuning profile for partner angle was relatively uniform in leaders, follower neurons showed a pronounced bias towards angles near zero degrees (Fig. 4q), consistent with the follower’s need to orient towards and follow the leader. The asymmetric and skewed tuning suggests that mPFC neurons may reflect the behavioural value associated with specific partner positions. To test this, we first imaged mPFC activity in water-restricted mice interacting outside the task context, which significantly reduced the proportion of egocentric partner-selective neurons (Extended Data Fig. 14c). We then retrained a subset of well-trained pairs on a reversed mismatch rule, in which mice were rewarded when entering different active zones (Methods and Extended Data Fig. 14d). Notably, preferred partner distance in followers shifted towards longer ranges (Fig. 4r), while partner angle tuning decreased around 0° and increased near ±180° in both roles (Fig. 4s), consistent with the strategy of avoiding close frontal proximity under the mismatch rule. Together, these results demonstrate that partner representations in the mPFC are not static spatial maps but flexibly reflect their social value in a role-dependent manner.
Although these neurons’ representations are otherwise stable, small variations in their activity (Extended Data Fig. 14e) predicted trial-level role deviation under the match rule: followers were more likely to lead when neurons tuned to frontal partner positions responded more weakly and rear-tuned neurons responded more strongly (Fig. 4t). By contrast, leaders were more likely to follow when activity in both front- and rear-tuned neurons was elevated, with front-tuned neurons exerting a stronger influence (Fig. 4u,v). These fluctuations also predicted mismatch errors: error rates increased when rear-tuned neurons in followers responded more strongly and when front-tuned leader neurons were less active (Fig. 4w–y). Together, these findings reveal that trial-by-trial fluctuations in mPFC partner representations predict both role deviation and coordination errors.
Inferred value aligns with mPFC activity
We developed a MAIRL36 algorithm, to examine mouse latent value functions, or internal goal representations, that guide cooperative decision-making. MAIRL decomposes the joint value function into marginal and interaction maps, reducing the computational complexity of multi-agent inference (Methods). Based on observed animal trajectories, we inferred the best-fitting social value functions comprising separate components for self and partner position as well as partner distance and angle (Fig. 5a). Both leaders and followers primarily valued their own position and partner distance, with followers exhibiting significantly greater log-likelihood gain for the latter, consistent with their need to maintain proximity (Fig. 5b). Neither role placed value on the partner’s allocentric position. Followers showed further improvements in model fit when partner angle was included, whereas leaders exhibited less sensitivity to this variable. To confirm that these inferred values capture behaviour, we predicted stepwise decisions from the value functions and reward zone choices from simulated trajectories; both were above chance for leaders and followers (Fig. 5c).
a, Inference schematic. The diagram was created using BioRender; Wu, H. https://BioRender.com/hwyq5u6 (2026). b, Model-fit improvement as the log-likelihood (LL) gain for an example and all pairs (n = 6 per role). The bars show the mean. The points represent individual mice. The fractions on the charts represent the number of mice per role showing a significant improvement as each component is added (α = 0.01). c, Behavioural predictions (n = 6 per role). Left, the log-likelihood per decision. Right, the accuracy for reward zone prediction from simulated trajectories; models with partner angle cannot be simulated (Methods). The bars show the mean values; the points represent individual mice; the dashed line represents chance. d, Inferred value functions on correct trials for followers (top) and leaders (bottom). Left, example self-position value map. Middle, the value at the central initiation point, east and north ports, and the mean of all other locations. The bars show the mean values; and the points represent individual mice. Right, partner-distance value. Followers value proximal partner locations more than leaders. n = 6 per role. Scale bar, 10 cm. e, The partner-distance value conditioned on partner angle in followers peaks at close range when the partner is in front (left) and diminishes when the partner is behind (right). n = 6 mice. f, Correlation between log-likelihood gain from adding partner distance (left) or angle (right) and decoding accuracy for that term (n = 12 leaders and followers combined). The dashed line indicates the significance threshold for log-likelihood gain (α = 0.01). g, The R2 score for predicting the inferred composite value from mPFC population activity. Each point represents one session. n = 21 leader, 22 follower sessions. The P value was calculated from a circularly shifted null (Methods). The dashed line represents α = 0.01. h, The difference in inferred partner-distance value between control and chemogenetic inactivation for followers (left) and leaders (right). Followers show a decreased proximal value after inactivation. n = 11 mice. i, Inferred follower self-position value map in the control condition and during inactivation. Scale bar, 10 cm. j, Simulations using inferred value functions, substituting only the inactivation-derived partner-distance value into the control set, reproduce the cooperative impairment. n = 11 mice. Statistical analysis was performed using χ2 tests of nested models (b) and permutation tests (g) (all one tailed); and one-sample t-tests (c and e), paired t-tests (d, h and j), two-sample t-tests (d) and Pearson correlation (f) (all two tailed). Details are provided in Supplementary Table 1.
Source data
The composite value functions comprise maps for self position and partner distance; the self-position map displayed hotspots at the initiation point and reward ports, consistent with their task relevance (Fig. 5d (left and middle)). By contrast, the partner distance map revealed a prominent peak in followers, but not leaders, when the partner was nearby (Fig. 5d (right)). This peak was significantly stronger when the leader was positioned in front than behind (Fig. 5e), reflecting a key follower strategy of maintaining close, front-facing proximity to the leader. These value representations were diminished during error trials and under the mismatch rule (Extended Data Fig. 15a,b). Notably, the improvement in model fit resulting from including partner angle correlated with its decodability from mPFC population activity across animals, and a similar trend was observed for partner distance (Fig. 5f). Moreover, the inferred value functions could be directly decoded from mPFC population activity in all follower recordings and a subset of leader sessions (Fig. 5g and Extended Data Fig. 15c), suggesting more-robust value encoding in the follower mPFC. These results establish a quantitative correspondence between MAIRL-inferred value functions and mPFC population representations.
To directly link mPFC inactivation to altered social value representations, we fit the MAIRL model to trajectories from chemogenetic inactivation sessions. This analysis revealed a significant reduction in inferred value for proximal partner distance in followers, along with modest changes in self-position value, consistent with broad suppression of mPFC activity (Fig. 5h,i and Extended Data Fig. 15d). Importantly, simulations based on these inactivation-derived value functions recapitulated the observed impairment in cooperative behaviour, when using altered partner-distance values (Fig. 5j), but not self-position values (paired t-test, t = 1.53, P = 0.16; n = 11 animals). Together, these findings show that MAIRL-inferred value functions are decodable from mPFC population activity and selectively disrupted by its inactivation, consistent with a dynamic, role-specific social value map.
Discussion
Our results reveal that leadership in mice reflects a structured asymmetry in social influence during joint decision-making. Although leaders are operationally defined by arrival time, they exert a disproportionate impact on their partner’s trajectories and target choice, consistent with evolutionary definitions of the role8. However, this influence is not unidirectional: leaders also accommodate followers, performing robustly on follower-led trials, and follower-mPFC inactivation drives a compensatory rise in leader sensitivity to partner cues. The lower performance and loss of mPFC dependence when mice tracked an inanimate stimulus further highlights the bidirectional interactions between social partners. Consistent with this, mPFC neurons in both roles encode leader–follower dynamics and egocentric partner position, and a subset of leaders assign value to partner distance and angle. Thus, leadership is an asymmetric yet reciprocal partnership.
In social species, leadership often emerges from the interaction between individual predispositions and social context1,37. However, in our task, leadership and initiatorship are not predicted by baseline locomotor and learning metrics, nor are they associated with dominance hierarchy. Initiators were not consistently followers despite often being positioned farther from the active zones at trial onset, further indicating that role emergence is shaped by cooperative demands rather than preexisting individual differences. Consistent with this interpretation, forward MARL modelling showed that social roles can self-organize among initially identical agents, underscoring the capacity of shared incentives alone to drive behavioural asymmetry. What biases individuals towards particular roles, whether stochastic fluctuations during learning or subtle predispositions38,39 that are not captured by our measures, remains an important open question.
Multiple lines of evidence confirm that the leader–follower dynamics capture genuine social interaction. First, the strong reciprocity between leader and follower roles reflects dynamic mutual influence rather than unilateral control. Second, learning rates differed across sexes despite similar role dynamics, consistent with sex-dependent influences on social behaviour40,41,42. Third, roles persisted when animals were paired with new partners yet adapted to individual idiosyncrasies, indicating stable but flexible social identities. Fourth, pairs exhibited diverse role-specific motif repertoires, revealing both behavioural flexibility and social individuality. Finally, neural encoding of trial-by-trial leadership dynamics and partner position in leaders and followers suggests that both roles maintain internal models of their partner, reinforcing the fundamentally social nature of this interaction.
We show that the mPFC has a critical role in leader–follower dynamics, a social dimension that, to our knowledge, has not previously been addressed in mechanistic studies of social cognition in animals14,27,30,34,43,44. Although chemogenetic and optogenetic perturbations differ in spatial coverage and temporal duration, their effects converge on showing impaired cooperative performance and reciprocal adjustments between partners. However, chemogenetic inactivation reduced leader–follower asymmetry, whereas optogenetic silencing did not. This pattern is consistent with sustained inactivation altering behavioural strategy, while transient silencing primarily affects trial-by-trial execution. Supporting this interpretation, MAIRL analyses showed that chemogenetic inactivation reduced the inferred partner-distance value in followers. Complementing these causal findings, the mPFC encodes trial-by-trial leading versus following in a role-dependent and asymmetric manner, potentially supporting credit assignment45 and role differentiation. Moreover, neurons preferentially encode the less-frequent trial type within each role, suggesting that the mPFC tracks role identity as a stable reference against which moment-to-moment arrival order is evaluated.
Our paradigm enabled us to examine how the brain constructs internal models of a social partner depending on one’s social role. In contrast to previous studies identifying social spatial representations in the hippocampus46,47,48 and prelimbic cortex49, our recordings reveal that the mPFC encodes partner position predominantly in egocentric coordinates, with a fourfold enrichment over other spatial reference frames. Critically, these representations are not static spatial maps: they reorganize when task rules change, diminish outside the task context, and predict both role deviation and error probability on a trial-by-trial basis. Thus, in contrast to scalar common currency value signals in striatal circuits50,51, the mPFC constructs a value-structured cognitive map of egocentric partner position.
MAIRL provides a principled framework for disentangling the latent value components that guide cooperative decision-making. By decomposing joint value into interpretable features, such as self-position and partner distance and angle, the model avoids enumerating the full multi-agent state space and, instead, captures the compact variables that are most relevant for coordination. This decomposition is not only computationally efficient, but may reflect a broader neural strategy for managing the representational complexity of social interaction52, where identity, history, context and spatial relationships can rapidly expand the space of possible states53. In our task, the inferred value functions were dynamic, role-dependent and decodable from mPFC population activity, establishing a direct link between computational inference and neural representation. More broadly, MAIRL offers a generalizable approach for formalizing latent cognitive variables in complex social behaviour, providing a path to test how brains achieve dimensionality reduction in multi-agent settings.
Together, our paradigm and MAIRL provide an experimentally tractable and computationally interpretable approach for studying the neural basis of leader–follower dynamics. This integrates behavioural strategies, latent social values and neural population activity within a single framework, revealing how the mPFC supports role-based cooperation. Our work establishes a foundation for dissecting how social roles emerge, adapt and guide collective behaviour. By extending this approach to larger groups, richer social contexts and more sophisticated inference problems54, future studies can test how brains navigate the complexity of group interaction and may inform both models of social dysfunction and biologically inspired multi-agent artificial intelligence.
Methods
Animals
Wild-type C57BL/6J female and male mice (2–6 months old; Jackson Laboratory) were used as experimental animals. Mice were housed under a reversed 12 h–12 h light–dark cycle (lights off at 07:00 and on at 19:00) in a temperature- and humidity-controlled environment (18–23 °C, 40–60% humidity), with ad libitum access to food and water. Training began when mice were approximately 2–3 months of age. All pairs were age-matched at the start of training. During training and experiments, mice were water restricted but maintained at 80–90% of their initial body weights. Experiments were conducted during the dark cycle, and each training session lasted for about 1 h, during which mice received 0.5–1.5 ml of water from the task. Animals received supplemental water as necessary to maintain their body weights. All procedures complied with the National Institutes of Health Guide for the Care and Use of Laboratory Animals and were approved by the Icahn School of Medicine at Mount Sinai Institutional Animal Care and Use Committee.
Behavioural apparatus
The training arena was an 18 inch × 18 inch white-acrylic square chamber. Four reward zones were positioned at the centre of each arena wall, each with two adjacent water delivery ports. Reward ports were 3D-printed using white and transparent resins, each incorporating an infrared beam for nose-poke detection and a white LED to indicate its availability. Water was delivered through a stainless-steel tube within each port, controlled by a solenoid valve (Lee, LHDB0533418H). An initiation detector, fitted with an infrared LED and a reflective sensor, was positioned at the centre of the arena to detect trial initiation when approached by mice. Analogue signals from all ports were digitized through the Arduino Nano and acquired through a DAQ system (National Instruments, USB-6001). Behavioural control was implemented in LabVIEW (National Instruments, 2014) to manage trial timing, LED signalling, water delivery and data logging. Mouse behaviour was recorded at 30 fps using an overhead camera (Teledyne FLIR, BFS-U3-16S2M-CS).
Video tracking
We used the machine-learning-based tools SLEAP (v1.3.3)55 and Ensemble Kalman Smoother (EKS) (v0.0.0)56 to annotate the nose, neck and torso positions of both mice in each video frame, enabling precise tracking of their movement and posture. To identify individual mice, we shaved a small patch of back fur on one mouse in each pair. Each mouse was annotated with three key points: nose (tip of the nose), neck (centre of the neck) and torso (centre of mass), which were connected to form a skeleton. Approximately 4,200 video frames from 10 mouse pairs were manually annotated for training a SLEAP model, which was then used to automatically estimate the key points of other pairs. To optimize tracking accuracy, we additionally applied EKS—a post-processing method that refines pose estimation outputs by smoothing several model predictions, resulting in more robust tracking. Custom MATLAB scripts (MathWorks, 2024b) were used to align the behaviour data from LabVIEW with the video frames based on the LED onset and offset.
Behavioural definitions
Mouse position was defined using its neck position, which was also used to calculate speed and acceleration. The reward zone was defined as a semicircle with a 10 cm radius centred at the midpoint between each pair of reward ports (Extended Data Fig. 1h). Arrival was defined as the first frame in which a mouse’s neck position entered a reward zone. Reaction time was defined as the time from trial onset to arrival at the chosen reward zone.
Cooperative foraging task training
Before training, mice were water restricted to 1 ml per day for at least 3 days and habituated to the experimenter and the behavioural set-up. Mice underwent one training session per day, progressing through four main stages.
Stage 1: light-guided water retrieval
Mice were trained to associate illuminated ports with water rewards within a single reward zone, separated from the rest of the arena by an acrylic barrier. On each trial, the left or right port in that reward zone was randomly illuminated; poking the illuminated port triggered water delivery and turned off the light. Mice advanced to stage 2 after 3 days of training.
Stage 2: learning to initiate trials
Mice were given access to the entire arena, including all four reward zones and the central initiation point. Trials began when a mouse was detected by the sensor at the initiation point, triggering illumination of both ports in a single active reward zone. Mice were required to poke an illuminated port within 15 s to receive a water reward; poking an inactive port terminated the trial and was followed by a timeout (6–10 s). Mice that reached 80% correct progressed to stage 3.
Stage 3: cooperative foraging shaping
Two same-sex cage mates (not necessarily siblings) that completed stage 2 were introduced to the arena together. Either mouse could initiate a trial, illuminating both reward ports within a single reward zone. Mice were required to each poke an illuminated port in the same zone within 15 s to receive a reward. The two mice may arrive at different times, and one can wait for the other, but the reward is delivered only when both mice nose-poke simultaneously. Poking an inactive port by either mouse terminated the trial and triggered a timeout (6–10 s). Mice that reached 80% correct advanced to stage 4a.
Stage 4a: cooperative foraging training
On each trial, either mouse could initiate by crossing a central initiation point, which standardized the position of that mouse at trial onset. The partner’s position was unconstrained and naturally variable, enabling assessment of how one animal’s spatial state influences the other’s decision-making. After trial initiation, two of the four reward zones were randomly illuminated as active zones. To receive a water reward, both mice were required to choose the same reward zone, each nose poking into one of its two ports (the match rule). As in stage 3, the two mice may arrive at different times, and one can wait for the other, but reward is delivered only when both mice nose-poke simultaneously. We refer to this joint action by two agents towards a shared goal as cooperation.
Trials were terminated without reward if the two mice chose ports from different active zones (mismatch error) or if either mouse poked any port from an inactive zone (unrewarded error). All error trials were followed by a timeout (6–10 s) added to the inter-trial interval (ITI, 3–8 s) before the next trial. Omitted trials are defined as failure to respond within 15 s in early stage 4a or within 8 s in late stage 4a. The correct rate is defined as the number of correct trials divided by all trials excluding omitted trials. Criterion performance is defined as 80% correct across three consecutive sessions, and pairs are excluded if they do not reach criterion after 50 days of training.
Social role assignment
Once a pair reached the training criterion (≥80% correct rate for three consecutive sessions), we assigned each mouse as either a leader or follower, and independently as an initiator or responder. To determine leadership, we computed the proportion of trials led by one mouse in a session and compared it to chance (0.5) using a two-tailed binomial test (α = 0.01). If the proportion was significantly biased, the mouse that led more trials was designated the leader; otherwise, no leader was assigned. The same procedure was applied to assign initiators and responders. Leader asymmetry was defined as the absolute difference between the proportion of trials led by each mouse; initiator asymmetry was defined analogously for trial initiation. Importantly, these classifications reflect role biases over sessions rather than fixed identities: leaders may follow and responders may initiate on a small subset of trials. Unrewarded errors and omitted trials were excluded from analysis, as leader and follower roles could not be assigned in these trials, and both occurred infrequently (0.5% and 1.2%, respectively, during the well-trained stage; Extended Data Fig. 1b,c). Accordingly, we evaluated mouse performance in these analyses using correct and mismatch trials only and defined a cooperation rate as the proportion of correct trials relative to the sum of correct and mismatch trials.
Manual annotation of behavioural motifs
Trial start and end times were extracted from the behavioural recordings to segment the continuous video into individual trials. A 30-frame pre-trial buffer was appended to each segment to provide temporal context. Segmented trials were concatenated into a single video (5–15 min in duration) for manual annotation. Videos were then annotated frame by frame by trained observers using Avidemux (v.2.8.1). For each identified behavioural motif, the start and end frame numbers were recorded. We then mapped these indices to corresponding frames in the original video and generated Gantt charts for visual inspection. Motifs were quality-checked and excluded if they exceeded a predefined duration threshold or occurred during the ITI.
The following four social behavioural motifs were defined and manually annotated. All behavioural motifs reflect pairwise interactions, but the track, sharp turn and join motifs were attributed to the actor mouse that performed the action. The synchronized travel motif was not assigned to either mouse, as it involved joint, symmetric behaviour.
Synchronized travel
The two mice move in parallel, maintaining similar linear and angular velocities and a consistent close distance throughout the trajectory. They arrive at the same reward zone in a coordinated, synchronized manner.
Track
One mouse, while initially approaching an active reward zone, slows down and turns its head to monitor the movement of its more distant partner. After detecting the partner’s approach, it adjusts its timing to allow both animals to arrive at the reward ports nearly simultaneously.
Sharp turn
Each mouse initially moves towards a different active reward zone. During the approach, one mouse abruptly changes its trajectory by more than 90°, redirecting its movement towards the reward zone selected by the partner.
Join
One mouse initiates movement towards an active zone independently. The partner mouse, initially stationary or undecided, subsequently aligns its trajectory with the first mouse after observing its movement. Both animals then travel together towards the same reward zone.
Partner swapping
Groups of four same-sex cage mates were randomly assigned to two dyads for cooperative foraging training. Once both pairs reached the training criterion and the social roles were assigned, we swapped partners either by pairing the leader from one dyad and the follower from the other (Fig. 1p,q and Extended Data Fig. 4a–d), or by pairing both leaders or both followers (Fig. 1r and Extended Data Fig. 4e–k). If both mice in a new pair were shaved, we marked one with black dye (Stoelting, 50450) for tracking with SLEAP. We then resumed training with the new pairs until they again reached the training criterion.
Solo foraging control (stage 2b)
Solo foraging was typically performed after stage 2 to assess the reward zone preferences of individual mice in the absence of a partner. After initiation, two reward zones were randomly illuminated, and the mouse received water by poking any port from the active zones. Poking an inactive port ended the trial without a reward and was followed by a timeout.
Cooperative foraging with rule reversal (stage 4d)
In stage 4d, task contingency was reversed: mice were rewarded for choosing different active zones (the mismatch rule). Choosing the same reward zone terminated the trial without reward. All other task parameters remained unchanged. A subset of mouse pairs well-trained in stage 4a was retrained in stage 4d.
Non-social stimulus tracking task
For the non-social-stimulus-tracking task (Fig. 3n), the behavioural arena was identical to that used for cooperative foraging, except that the opaque floor was replaced with a transparent base. A projector positioned beneath the arena projected a visual stimulus onto the floor through a mirror and filters. Before training, mice were water restricted and habituated to the experimenter and apparatus as described above. Training consisted of four stages.
Stage 1: single-zone reward acquisition
Mice were trained to obtain water rewards from a single active reward zone. On each trial, a black elliptical visual stimulus (3 × 6 cm) appeared at the centre of the arena and remained stationary for 0.5 s before moving towards the reward zone along a straight trajectory with small jitter orthogonal to the direction of motion. After reaching the reward zone, the stimulus remained at the port for 5 s while rotating. A nose poke at one of the corresponding reward ports within this response window triggered water delivery. Failure to respond was recorded as an error and followed by a 6 s timeout. The ITI was 3 s. Mice typically advanced to the next stage after approximately 3 days of training.
Stage 2: initiation training
Mice were trained to voluntarily initiate trials by crossing the central initiation point. After initiation, the visual stimulus appeared at the centre for 0.5 s before moving towards the single active reward zone as in stage 1. A correct trial required a nose poke at one of the corresponding reward ports within a 5 s response window. Incorrect responses were followed by a 6 s timeout. Mice that reached 80% correct progressed to stage 3.
Stage 3: two-choice tracking
Two adjacent reward zones were made available. After trial initiation, the stimulus appeared at the centre for 0.5 s and then moved to one of the two active reward zones, selected randomly. Mice were required to track the moving stimulus and poke one of the corresponding reward ports within a 5 s response window to receive a reward. Incorrect trials were followed by a 6 s timeout. Mice typically advanced after approximately 5–6 days of training.
Stage 4: trajectory-based stimulus tracking
This stage approximated the spatial and temporal structure of the cooperative foraging task. All four reward zones were made available. Instead of moving along a straight path, the stimulus replayed trajectories drawn from a library of 20 well-trained cooperative foraging sessions. The library included both leader and follower trajectories from trials initiated by the animal, ensuring that each replayed trajectory began at the centre of the arena. On each trial, one trajectory was randomly selected (50% chance from the leader pool and 50% from the follower pool) and replayed from the centre towards the corresponding reward port. The reward was delivered if the mouse poked the same reward zone towards which the stimulus was directed within the response window; otherwise, the trial was terminated and followed by a 6 s timeout. Training continued until performance plateaued, defined as less than 5% variation in the correct rate across three consecutive sessions.
Dominance hierarchy assays
To determine hierarchy, we used three independent and well-established assays: the tube test57, warm spot test22,23 and reward competition test24,25. These assays probe competitive interactions under different behavioural contexts to provide convergent measures of dominance hierarchy.
Tube test
The tube test was performed in a round-robin design with groups of four cage mates to assess dominance hierarchy. In a subset of animals, tube tests were repeated at multiple training stages (stage 2/2b, stage 3, early stage 4a and well-trained stage 4a) to assess the relationship between leader–follower roles and dominance hierarchy at each stage. In some cases, the tube test was performed within the two-animal pair only, which was sufficient for assessing the relationship between dominance hierarchy and leader/follower role assignment. During the test, mice were placed at opposite ends of a transparent acrylic tube (12 inches long, 1.25 inches in diameter) and encouraged to enter. As only one mouse could pass through the tube at a time, the mouse that forced the other to retreat was designated the dominant individual. In the round-robin design, each four-animal group completed six rounds of pairwise testing, and testing continued daily until a stable hierarchy was observed for four consecutive days.
Warm-spot test
Experiments were conducted in a covered transparent acrylic arena (1 ft × 0.5 ft × 0.5 ft; length × width × height). A black circular plastic platform (2.5 cm in radius) was positioned in one corner of the arena to serve as the warm spot. A small USB heating pad placed beneath the platform provided localized warmth. Cages containing four cage mates were placed on ice together with the test arena for 30 min before testing. Mice were then transferred to the arena and marked individually for identification. Behaviour was recorded for 20 min. The time each mouse spent on the warm spot was quantified using BORIS (v.9.2.3)58.
Reward competition test
The reward competition test assessed dominance during direct competition for a limited resource. As mice were well trained on the cooperative foraging task, they readily poked illuminated reward ports without additional training. In this test, a single reward port from the cooperative foraging apparatus was used, and two mice competed for access to the port, with reward availability signalled by LED illumination. Each session consisted of 20 trials with a 3 s ITI. After a 2 day habituation period, the mice were tested in two sessions (one session per day). For cages containing four mice, all pairwise combinations were tested, yielding six matches per day in a randomized order.
Stability criteria
In the tube test, stable social rankings were achieved by design through daily testing for up to 2 weeks, until a consistent hierarchy was established. For the warm-spot test, social rank was considered stable if the ranking, determined by time spent on the warm spot, was identical across two consecutive days. All tested groups reached stability within 2 days, except for one instance in which the relative ranking of two mice reversed. For the reward competition test, stability was defined as one mouse winning at least 60% of all trials summed across the two test sessions25.
Elo score calculation
Dominance hierarchies were quantified using a sequential Elo rating system, originally developed for ranking chess players59 and later adapted to estimate dominance hierarchies in animal groups25,60. The Elo score provides a dynamic, continuous estimate of an individual’s dominance strength within its group and is updated sequentially after each contest. All individuals were initialized with an Elo score of 1,000. After each dyadic interaction between individuals A and B, with current ratings, \({R}_{{\rm{A}}}\) and \({R}_{{\rm{B}}}\), we first calculated the expected probability that A would win:
$${E}_{{\rm{A}}}=\frac{1}{1+{10}^{({R}_{{\rm{B}}}-{R}_{{\rm{A}}})/400}}$$
(1)
Ratings were then updated sequentially according to:
$${R}_{{\rm{A}}}^{{\prime} }={R}_{{\rm{A}}}+K({S}_{{\rm{A}}}-{E}_{{\rm{A}}})$$
(2)
where \({R}_{{\rm{A}}}^{{\prime} }\) is the updated score, SA equals 1 for a win, 0 for a loss and 0.5 for a tie. The parameter K, which determines the sensitivity of rating updates, was set to 20. The rating of individual B, \({R}_{{\rm{B}}}^{{\prime} }\), was updated analogously using \({S}_{{\rm{B}}}=1-{S}_{{\rm{A}}}\) and the corresponding expected probability \({E}_{{\rm{B}}}\).
For the tube test and reward competition test, outcomes were derived from repeated round-robin dyadic contests within each group. The warm spot assay did not involve direct dyadic encounters; instead, the total occupancy time for each individual was ranked within the group, and pairwise outcomes were inferred such that the individual with greater occupancy time was assigned a win in each dyadic comparison. These inferred outcomes were incorporated into the same sequential Elo framework. Final Elo scores were used as continuous measures of relative dominance within each assay and were used to calculate cross-assay correlations in dominance ranking.
Virus
The following viral vectors were purchased from Addgene: AAV8-hSyn-hM4D(Gi)-mCherry (7 × 1012 viral genomes (vg) per ml; 50475-AAV8), AAV8-CaMKIIα-hM4D(Gi)-mCherry (2 × 1012 vg per ml; 50477-AAV8), AAV8-hSyn-mCherry (1 × 1013 vg per ml; 114472-AAV8), AAV9-syn-jGCaMP8m-WPRE (1 × 1013 vg per ml; 162375-AAV9), AAV9-Syn-GCaMP6f-WPRE-SV40 (7 × 1012 vg per ml; 100837-AAV9) and AAV5-CAG-GFP (7 × 1012 vg per ml; 37825-AAV5). AAV8-CaMKIIα-KALI1-eYFP (1.8 × 1013 vg per ml; GVVC-AAV-292) was obtained from the Stanford University Virus Core.
Stereotaxic surgery
Mice were anaesthetized with a cocktail of ketamine (100 mg per kg) and xylazine (10 mg per kg) and secured in a stereotaxic frame (Kopf Instruments, Model 940). Anaesthesia was maintained with 1–1.5% isoflurane throughout the procedure. Viral injections were performed using a glass capillary connected to a nanoinjector (Drummond Scientific, Nanoject III) at a rate of 1–2 nl s−1. Stereotaxic coordinates were determined according to the Paxinos and Franklin mouse brain atlas. After surgery, mice received either a single dose of the long-acting analgesic buprenorphine (3.25 mg per kg, Ethiqa XR) or carprofen (5 mg per kg) for three consecutive days.
Chemogenetic inactivation of the mPFC was performed using AAV8-hSyn-hM4D(Gi)-mCherry or AAV8-CaMKIIα-hM4D(Gi)-mCherry, with AAV8-hSyn-mCherry as the control. Viral injections were administered at four sites per hemisphere to cover the entire mPFC (bregma coordinates: anteroposterior (AP), +1.98 mm; mediolateral (ML), ±0.45 mm; dorsoventral (DV), −1.85 mm/−1.45 mm, 200 nl per depth; AP, +1.18 mm; ML, ±0.45 mm; DV, −1.10 mm, 300 nl; AP, +0.50 mm; ML, ±0.45 mm; DV, −0.80 mm, 300 nl; AP, −0.20 mm; ML, ±0.45 mm; DV, −0.80 mm, 300 nl). For chemogenetic inactivation of the OFC, 300 nl of AAV8-hSyn-hM4D(Gi)-mCherry was injected bilaterally at AP, +2.46 mm; ML, ±1.00 mm; DV, −1.75 mm.
For optogenetic inactivation of the mPFC, AAV8-CaMKIIα-KALI1-eYFP was used, with AAV5-CAG-GFP as the control. Viral injections were performed at two sites per hemisphere (300 nl per site): AP, +1.98 mm; ML, ±0.45 mm; DV, −1.85 mm/−1.45 mm; and AP, +1.18 mm; ML, ±0.45 mm; DV, −1.10 mm. After injections, bilateral optical fibres (Amuza, TeleLCD-Y-2.0-500-0.9) were implanted at AP, +1.98 mm; ML, ±0.45 mm; DV, −0.80 mm. Optical fibres were secured with black dental acrylic (Lang Dental Manufacturing, Ortho-Jet) to firmly attach the fibre and block external light.
To implant a GRIN (gradient index) lens for miniscope recording, we performed a craniotomy over the mPFC at AP, +1.98 mm; ML, +0.45 mm with a 1-mm-radius window. Brain tissue above the mPFC was aspirated with a 27-gauge blunt or bent needle while continuously irrigating with cortex buffer to preserve tissue integrity. The resulting cavity was shaped to be as close to cylindrical as possible. We then injected 600 nl of AAV9-Syn-jGCaMP8m-WPRE or AAV9-Syn-GCaMP6f-WPRE-SV40 virus unilaterally at AP, +1.98 mm; ML, +0.45 mm; DV, −1.70 mm. After injection, a 1-mm-diameter, 4-mm-length GRIN lens (Inscopix) was positioned at the injection site and implanted 1.45 mm below the brain surface. The lens was secured with cyanoacrylate adhesive (Loctite) and further protected with a layer of low-toxicity silicone adhesive (World Precision Instruments, Kwik-Sil). Finally, dental acrylic was applied to stabilize the implant and cover the remaining exposed skull.
After 2–3 weeks of viral expression, mice were reanaesthetized and placed back into the stereotaxic frame. The overlying dental cement was carefully drilled off to expose the implanted GRIN lens. A Miniscope61 pre-mounted on a fixed baseplate was positioned above the lens, and the field of view was monitored in real time. The Miniscope was gradually lowered until the imaging field was visible. Once we observed the optimal field of view, the baseplate was secured with cyanoacrylate and dental cement. To minimize interference from external light, an outer layer of black dental cement was applied. A protective cap was then attached with screws, and mice were allowed to recover with postoperative analgesia. Viral expression and implant placement were verified histologically in all animals, and only animals with correct targeting were included in the analyses. Calcium imaging as well as inactivation during cooperative foraging were performed primarily in female pairs, because animals were single-housed after surgery to protect implants and incisions during recovery, and adult males could not be reliably re-paired. The non-social-stimulus-tracking task, which tests individual mice and does not require re-pairing included both sexes.
Chemogenetics
For chemogenetic manipulation, we injected the viral vectors into untrained mice and allowed at least 4 weeks of viral expression while simultaneously training them on the cooperative foraging task. We administered the DREADD agonist clozapine N-oxide (CNO; 5 mg per kg; Tocris, 4936) dissolved in saline through intraperitoneal injection 30 min before the start of the session. We performed the control session on the day before the inactivation session, when we injected mice with an equal volume of saline to control for potential effects of liquid intake and handling. To control for potential non-specific effects of CNO, we administered the same dose to a separate control group expressing mCherry and evaluated their behaviour in the cooperative foraging task.
Wireless optogenetics
We used a wireless optogenetic system (Amuza, Teleopto) for mPFC inactivation. After completing behavioural training, virus injection and fibre implantation surgery, mice were given 2–3 weeks for viral expression and post-surgical recovery. Only mice that reached the training criterion were included in the optogenetic experiments, all conducted within 4 weeks after surgery. Before the stimulation session, animals were fitted with dummy receivers for 3 days during training to acclimatize to the implants and receivers.
For stimulation, a lightweight receiver was connected to the implanted optical fibres. Stimulation signals were generated by a task control system (National Instruments, NI-6001) as transistor–transistor logic (TTL) pulses and transmitted through BNC cables to the optogenetic control unit. The behavioural set-up enabled precise inactivation across different task periods, including throughout the trial and 1 s or 2 s from trial onset. During stimulation sessions, light pulses were delivered on one-third of the trials according to a pseudorandomized schedule. To prevent intensity fluctuations caused by battery discharge, we limited stimulation intensity to one-third of the maximum output of the implanted LED (590 nm, ~0.8 mW per hemisphere). Optogenetic inhibition in both or a single animal was achieved by independently activating the receiver in each mouse.
One-photon calcium imaging
Post-surgical mice with clear imaging fields were selected and trained in the cooperative foraging task. Calcium imaging was performed during well-trained stage 4a of the training protocol. Each day, one mouse underwent calcium imaging while the other was fitted with a dummy scope of equal weight to control for potential impact of the Miniscope attachment. In each recording session, neural activity was recorded from only one animal, either the leader or the follower. Imaging data were acquired using the Miniscope DAQ system and synchronized with behavioural video recordings at 30 fps. A TTL signal generated by the Miniscope DAQ was transmitted through a trigger cable (6-pin GPIO Hirose Connector Cable, FLIR) to the behaviour-recording camera, ensuring frame-by-frame alignment. We used Suite2p62 for region-of-interest identification and extraction of calcium transients, and Cascade63 to infer spike rate from the calcium signals. To compute fluorescence changes (ΔF/F), we first subtracted 70% of the local neuropil signal from the fluorescence signal of each cell to obtain the raw trace. To estimate a dynamic baseline, we first smoothed the raw trace using a Gaussian filter with s.d. σ = 6.7 s to reduce high-frequency noise and then applied a two-step filtering process: first, a moving minimum filter with a window size of 500 s was applied to the smoothed trace to track the lower envelope; a moving maximum filter of the same window size was then applied to this result to obtain a conservative estimate of the baseline. ΔF/F was computed by normalizing the raw trace to this dynamic baseline. Analyses of selectivity for choice and leading versus following included 6 animals expressing GCaMP6f and 2 animals expressing GCaMP8m. Spatial selectivity analyses included 6 animals expressing GCaMP6f and 6 animals expressing GCaMP8m. Mice were excluded from calcium imaging analysis if they show incorrect viral targeting, insufficient viral expression, a low-quality field of view or low-quality calcium signals.
Electrophysiological recording
To validate mPFC inactivation, extracellular single-unit recordings were performed acutely in head-fixed animals that expressed the appropriate effector genes. Four animals were used for validating chemogenetic inactivation and three animals for optogenetic inactivation. Recordings used 64-channel silicon probes (Cambridge Neurotech, Acute H3 probe) stereotaxically targeted to the mPFC using Bregma coordinates: AP, 1.78–2.22 mm; ML, 0.45 mm; DV, 1.85–2.30 mm. Before recording, a craniotomy (0.3–0.8 mm in diameter) was performed over the target area, and the recording depth was determined using micromanipulator readings. Data were acquired on the Intan RHD recording system (Intan Technologies) and digitized at 25 kHz. After each recording session, the brain surface was protected with silicone gel (Dow Corning, 3-4680) and Kwik-Sil (World Precision Instruments). Voltage signals were high-pass filtered and automatically sorted using Kilosort364. Spike clusters were then manually curated using the Phy GUI65 to merge spikes from the same units and exclude noise and poorly isolated units. Recording sites were visualized by staining the probes with Vybrant DiO (Invitrogen, V22886) or Vybrant DiI (Invitrogen, V22885) and verified histologically after recording.
To validate chemogenetic inactivation, we first recorded a baseline level of activity from animals expressing AAV8-hSyn-hM4D(Gi)-mCherry. We then administered CNO (5 mg per kg, intraperitoneal) while continuing to record, to capture the effect of inactivation. To validate inactivation using KALI-1, we used animals expressing AAV8-CaMKIIα-KALI1-eYFP in the mPFC. An optical fibre (RWD, 0.50 NA, 200 μm diameter) was positioned approximately 0.5 mm above the brain surface, producing an illuminated area with an estimated radius of around 0.3 mm for 590 nm illumination. To prevent collision with the fibre, the recording probe was inserted at a 10-degree angle. The brain surface was illuminated at varying light intensities, with each power level tested 10 times within a recording session (2 s of illumination per light pulse, 30 s between pulses). For validation of the wireless optogenetic device, an optical fibre (Amuza) was implanted at a 10-degree angle, secured 0.5 mm below the brain surface with dental cement. A recording probe was inserted at a 10-degree angle to avoid collision and positioned to record from mPFC neurons beneath the illuminated area.
Unsupervised behavioural classification
We computed LISBET (v0.3.0) embeddings, a self-supervised representation of behaviour that captures body kinematics in a high-dimensional latent space26. Using the pretrained LISBET model, we generated 64-dimensional embeddings for each video frame. As body keypoints were tracked with SLEAP but the pretrained LISBET model was trained on DeepLabCut (DLC) coordinates, we generated pseudo-DLC keypoints by inferring the corresponding DLC body points from the SLEAP-derived coordinates. To prevent embeddings from capturing task-irrelevant information (for example, which of the four reward zones was chosen), we restricted the analysis to correct trials and rotated all trials to a common configuration with the chosen reward zone in the north. Embeddings were computed across all sessions from all mouse pairs using a temporal window of 20 video frames. The target frame included a context of 20 frames into the past.
To quantify the organization of LISBET embeddings with respect to human-defined behavioural motifs, each datapoint was assigned to a cluster corresponding to its manually annotated motif label. We then computed a modified silhouette score for each datapoint \(i\):
$${s}_{{i}}=\frac{{a}_{i}-{b}_{i}}{{a}_{i}}$$
(3)
where ai is the mean distance between point i and all points assigned to other clusters, and bi is the mean distance between point i and all other points within the same cluster. This metric resembles the standard silhouette score, with two modifications that increase sensitivity to structured but non-clustered data: (1) ai is defined as the mean distance to all points in all other clusters, rather than only to the closest other cluster as in the standard formulation; and (2) the denominator is ai, rather than max(ai,bi). For each mouse pair, silhouette scores were averaged across all frames and sessions to obtain a single summary value.
Spatial decision-making analysis
To examine how a mouse’s decision is influenced by its partner on trials initiated by the mouse of interest, we compared the mouse’s zone choice in the cooperative foraging task to a solo foraging control. For this analysis, all adjacent trial types were rotated to align the active zones to a north-east configuration. Trials longer than 5 s (10.0% of all trials), unrewarded errors (0.5%; Extended Data Fig. 1b) and omitted trials (1.2%; Extended Data Fig. 1c) were excluded.
To visualize the mouse’s choice (Fig. 2f,g and Extended Data Fig. 8a,b), we performed the following logistic regression:
$$\begin{array}{l}\mathrm{Logit}({P}_{\mathrm{North}})\,=\,{\beta }_{0}+{\beta }_{1}{\theta }_{\mathrm{East}}+{\beta }_{2}{\theta }_{\mathrm{North}}+{\beta }_{3}{I}_{\mathrm{Partner}}\\ \,+\,{\beta }_{4}{\theta }_{\mathrm{East}}{I}_{\mathrm{Partner}}+{\beta }_{5}{\theta }_{\mathrm{North}}{I}_{\mathrm{Partner}}\end{array}$$
(4)
where PNorth denotes the probability that the mouse of interest selects the north zone. θEast and θNorth are the mouse’s absolute head angles relative to the east and north directions, respectively. IPartner indicates whether the partner is facing a given reward zone (Fig. 2f,g) or whether the partner is located in a given reward zone (Extended Data Fig. 8a,b), depending on the condition. The β terms are fitted coefficients: β0 captures baseline bias towards the north zone in the solo control (in log-odds units); β1 and β2 reflect the animal’s sensitivity to heading information in the solo control (with β1 > 0, β2 < 0, in well-trained animals); β3 quantifies the shift in bias attributable to the partner; and β4 and β5 reflect changes in heading sensitivity in the presence of the partner. To visualize the effect of a partner in a given state (for example, facing or located in the east), the regression included only solo trials and cooperative trials matching that state; cooperative trials with the partner in other headings or locations were excluded.
To quantify changes in heading sensitivity and directional bias induced by the presence of the partner (Fig. 2h,i and Extended Data Fig. 8c,d), we first computed a single heading-difference variable:
$${\Delta \theta =\theta }_{\mathrm{East}}-{\theta }_{\mathrm{North}}$$
(5)
where larger values indicate greater alignment with the north than the east reward zone. Using the same IPartner coding, we then fitted the following model separately for leaders and followers within each partner condition:
$$\mathrm{Logit}({P}_{\mathrm{North}})={\beta }_{0}+{\beta }_{1}\Delta \theta +{\beta }_{2}{I}_{\mathrm{Partner}}+{\beta }_{3}\Delta \theta {I}_{\mathrm{Partner}}$$
(6)
Similar to equation (4), β2 quantifies the partner-induced changes in choice bias for the north zone; β3 estimates the partner-induced changes in sensitivity to the mouse’s own heading. Here IPartner = 1 for cooperative trials and 0 for solo foraging, with trials selected per partner state as described for equation (4).
To examine the influence of spatial variables on both leader and follower decision-making across all trials in a single framework (Fig. 2p–r and Extended Data Fig. 8o–r), we first calculated the differences in each animal’s heading and distance relative to the two active zones at trial onset, and standardized these values to the range of [0,1] before model fitting:
$${\Delta \theta =\theta }_{\mathrm{East}}-{\theta }_{\mathrm{North}}$$
(7)
$${\Delta d=d}_{\mathrm{East}}-{d}_{\mathrm{North}}$$
(8)
We then used the following equations to fit the leader and follower’s decision, respectively:
$$\begin{array}{l}\mathrm{Logit}({P}_{\mathrm{North}}^{\mathrm{Leader}})\,=\,{\beta }_{0}+{\beta }_{1}\Delta {\theta }^{\mathrm{Leader}}+{\beta }_{2}\Delta {d}^{\mathrm{Leader}}\\ \,+\,{\beta }_{3}\Delta {\theta }^{\mathrm{Follower}}+{\beta }_{4}\Delta {d}^{\mathrm{Follower}}\end{array}$$
(9)
$$\begin{array}{l}\mathrm{Logit}({P}_{\mathrm{North}}^{\mathrm{Follower}})\,=\,{\beta }_{0}+{\beta }_{1}\Delta {\theta }^{\mathrm{Leader}}+{\beta }_{2}\Delta {d}^{\mathrm{Leader}}\\ \,+\,{\beta }_{3}\Delta {\theta }^{\mathrm{Follower}}+{\beta }_{4}\Delta {d}^{\mathrm{Follower}}\end{array}$$
(10)
Although the equations are structurally identical, separate β coefficient sets were fit for leaders and followers. We predicted the leader’s and follower’s choice using these logistic fits (Fig. 2q and Extended Data Fig. 8o,q). The GLM was fit to a training set, and the model performance was evaluated on a held-out test set as the proportion of correctly predicted choices. We used 90% of the data for training and 10% for testing, repeated across 100 random splits.
To assess statistical significance, we compared model performance against a shuffled null distribution. For each fold, a shuffled version of the data was created by randomly permuting choice labels, and the model was trained and tested using the same predictor variables. Model accuracy on true versus shuffled data was compared across folds using a bootstrap procedure (10,000 iterations) to estimate the P value for the observed difference in mean accuracy. This procedure was applied separately to predict leader and follower choices, yielding role-specific estimates.
To examine the impact of chemogenetic and optogenetic inactivation on heading sensitivity and choice bias (Fig. 3i,k,m and Extended Data Fig. 11p), we used the following logistic regression to fit the leader or follower’s decision:
$$\begin{array}{l}\mathrm{Logit}({P}_{\mathrm{North}})\,=\,{\beta }_{0}+{\beta }_{1}\Delta {\theta }^{\mathrm{Leader}}+{\beta }_{2}\Delta {d}^{\mathrm{Leader}}+{\beta }_{3}\Delta {\theta }^{\mathrm{Follower}}\\ \,+\,{\beta }_{4}\Delta {d}^{\mathrm{Follower}}+{\beta }_{5}{I}_{\mathrm{Inact}}+{\beta }_{6}\Delta {\theta }^{\mathrm{Leader}}{I}_{\mathrm{Inact}}+{\beta }_{7}\Delta {d}^{\mathrm{Leader}}{I}_{\mathrm{Inact}}\\ \,+\,{\beta }_{8}\Delta {\theta }^{\mathrm{Follower}}{I}_{\mathrm{Inact}}+{\beta }_{9}\Delta {d}^{\mathrm{Follower}}{I}_{\mathrm{Inact}}\end{array}$$
(11)
where β6 to β9 quantify the changes in sensitivity to leader and follower’s heading and distance induced by inactivation. Trials were pooled across sessions and animals to increase the statistical power.
Choice selectivity analysis
To identify neurons selective for the animal’s reward zone choice (Extended Data Fig. 12f–h), we computed mean calcium responses within the 1-s time window before trial end across trials for each neuron. We then used a one-way ANOVA to test for significant differences in response across the four choice conditions separately for each neuron, without correction for multiple comparisons. Neurons significant at α = 0.05 were classified as choice selective.
To quantify the strength of choice tuning across neurons, we computed a variance-based selectivity index (SI) defined as:
$$\mathrm{SI}=\frac{\text{s.d.}({\mu }_{{i}})}{\bar{\mu }}$$
(12)
where μi is the mean response of the neuron for each choice condition i, s.d.(μi) is the s.d. across those means and \(\overline{\mu }\) is the grand mean. This SI captures how much a neuron’s activity deviates across choice conditions relative to its overall activity level.
To assess whether neuronal population activity encoded zone choice, we trained a support vector machine decoder to classify choice labels from trial-aligned calcium signals in each session. Neural activity was extracted from ΔF/F traces in a window spanning 2 s before and 2 s after trial end (corresponding to 60 pre- and 60 post-alignment frames). To reduce noise, we smoothed the traces by computing the mean responses in non-overlapping, three-frame time bins. At each time bin, we trained a linear support vector machine (LIBLINEAR implementation66) to classify which of the four reward zones the animal chose based on randomly sampled subsets of 100 simultaneously recorded neurons. Trials were randomly split, stratified by choice, into training (90%) and test (10%) sets across 50 cross-validation folds. The classification accuracy was computed as the mean correct rate on held-out trials. This procedure was repeated across time bins to obtain a temporal profile of decoding performance. All analyses were performed separately for leader and follower roles and for each session and then averaged across all available sessions from each animal.
Selectivity for leading versus following
To identify neurons selective for leading versus following (Fig. 3u,v,x), we computed mean ΔF/F responses within a 1 s time window after the arrival of the recorded mouse for each trial. Trials in which both animals arrived in the same video frame were excluded. For each neuron, responses were grouped by whether the recorded animal was leading or following on a given trial. For each neuron, we calculated an area under the receiver operating characteristic curve (auROC) comparing calcium activity on trials where the recorded animal led versus followed, after regressing out speed and position contributions. The auROC was linearly scaled to the range of [−1, 1] by computing SI = 2 × (auROC − 0.5). Positive SI values indicate stronger activity when leading, and negative values indicate greater activity when following. To assess statistical significance, we generated a null distribution of SI values using stratified label shuffling. Specifically, to control for potential spatial choice confounds, we permuted the arrival-order labels within each port choice category (left versus right) independently. This procedure was repeated 1,000 times per neuron, and a P value was calculated as the fraction of shuffled SIs of which the absolute magnitude exceeded that of the observed SI. Neurons of which the observed absolute SI exceeded the 95th percentile of the shuffled distribution were classified as selective for leading versus following.
To quantify neural selectivity for trial-by-trial leading versus following over time (Fig. 3y), we computed a SI for each neuron within the peri-arrival window. ΔF/F traces were aligned to the arrival of the recorded mouse, and a fixed time window (−2 s to +3 s, corresponding to 60 pre- and 90 post-alignment frames) was extracted for analysis. Trials in which both animals arrived in the same video frame were excluded. For each time point and neuron, we calculated the selectivity using the auROC-based SI as described above, yielding a time-by-neuron matrix of SI values. To visualize population dynamics, we plotted the SI heat map across neurons sorted by the timing of their peak selectivity.
To further examine how mPFC population activity encodes leading versus following (Fig. 3z), we trained a logistic regression decoder to classify whether the recorded animal led or followed on a given trial, based on randomly sampled subsets of 100 simultaneously recorded neurons. We excluded trials with ties in arrival times and aligned calcium traces to the arrival times of the recorded animal. For each neuron, ΔF/F signals were binned into non-overlapping windows (3-frame bins). For each time bin, we trained a logistic regression model (LIBLINEAR, L2-regularized) on 90% of trials and tested on the remaining 10%, stratified by arrival-order labels. If a test set lacked representation from either class, we incrementally increased the holdout proportion until both classes were included. This procedure was repeated over 50 cross-validation folds. Decoding performance was quantified as the mean auROC across all test folds. To establish significance, we compared decoding performance against a null distribution generated from random permutations of arrival-order labels. For each permutation, we applied the same decoding pipeline and recorded the corresponding auROC.
Spatial selectivity analysis
To identify neurons selective for allocentric or egocentric positions of the self and the partner (Fig. 4a–g and Extended Data Fig. 13a–c), we computed deconvolved spike rate maps in four spatial reference frames: (1) the position of the recorded mouse in the arena (allocentric self); (2) the partner’s position in the arena (allocentric partner); (3) the partner’s position in egocentric coordinates without heading (HD) alignment (egocentric partner unaligned with HD); and (4) the partner’s position aligned to the heading of the recorded mouse (egocentric partner aligned with HD). All maps were constructed using 5 × 5 cm spatial bins. We included all frames in a session but excluded the frames in which either mouse was within any reward zone (10 cm from any reward ports; Extended Data Fig. 1h) to rule out reward-related confounds. For each neuron, we quantified spatial selectivity using three criteria. Only neurons passing all three criteria in a given reference frame were classified as selective for that spatial variable.
Spatial information content
Spatial information was calculated67 as:
$$I=\sum _{i}{p}_{i}\frac{{\lambda }_{i}}{\lambda }{\log }_{2}\left(\frac{{\lambda }_{i}}{\lambda }\right)$$
(13)
where λi is the mean spike rate in the ith spatial bin, λ is the overall mean spike rate, and pi is the occupancy probability of the ith bin. Significance was assessed by circularly shifting the deconvolved spike rate (n = 100 permutations) relative to the position labels. Cells with spatial information exceeding the 95th percentile of the null distribution were considered significant.
Spatial coherence
We computed the mean correlation between each spatial bin and the mean of its eight neighbours, across all bins. Cells with coherence values above the 95th percentile of a circularly shifted null distribution were deemed to be coherent.
Within-session stability
To assess tuning stability within a session, we computed the Spearman correlation between rate maps generated from the first and second halves of the session. Neurons were considered to be consistent if their correlation exceeded the 95th percentile of a null distribution obtained by circularly shifting spike times.
Decoding partner distance and angle
To decode partner distance and egocentric partner angle on a frame-by-frame basis (Fig. 4h–k), we applied logistic regression classifiers to the population activity. The partner distance was discretized into 5-cm bins spanning 0–35 cm, and the egocentric partner angle was divided into 15 uniform bins (−180° to 180°, positive when the partner is on the left). Neurons from multiple imaging sessions of the same animals were pooled after binning. To balance class sizes, an equal number of frames were randomly sampled without replacement from each bin across sessions. Sampled frames were then sorted by their original timestamps to preserve their temporal order for decoding. Deconvolved spike rates from all included neurons were extracted for each frame, and the corresponding distance or angle bin index was used as the class label.
To assess decoding accuracy, we implemented multi-class logistic regression using the LIBLINEAR solver (L2-regularized). A fivefold cross-validation was used, in which the data were evenly split into five non-overlapping segments. In each fold, four segments were used for training and the remaining segment for testing, such that all data segments were used once as the test set. Decoding was repeated 10 times with 300 randomly selected neurons per repetition.
For each cross-validation fold and repetition, we recorded the predicted class probabilities for test frames and averaged them across repetitions to form a confusion matrix, representing the probability of predicting each distance or angle bin given the ground truth. As a performance metric, we computed the proportion of predictions matching the true distance label (Fig. 4i), or the proportion of predictions falling within ±1 bin of the true angle label, circularly defined (Fig. 4k). To assess statistical significance, final decoding accuracy was compared to a null distribution obtained by decoding shuffled data generated by circularly shifting the deconvolved spike traces. Confusion matrices were also generated for the shuffled data (Extended Data Fig. 14a,b).
CEBRA-behaviour multi-session embedding
We trained each CEBRA (v.0.5.0rc1) embedding on a single behavioural feature as a continuous label (such as the partner distance or angle), using 80% of trials pooled across all sessions from a single animal for training, and evaluated the model performance on the remaining 20% of held-out trials. To assess statistical significance, each embedding was compared to a temporally shuffled control, in which the behavioural variable was shifted by half the length of each session across both training and test sets. CEBRA embeddings were configured with a batch size of 500, a constant temperature of 0.01, a hidden layer size of 32, an output dimensionality of 3, a learning rate of 0.0003 and a cosine distance metric. The delta conditional distribution was used, and training was run for 1,000 iterations using the offset10-model architecture.
To quantify the structure of CEBRA embeddings, we discretized the behaviour variable of interest into quintiles containing an equal number of points. We then defined a modified silhouette score for each datapoint i as:
$${s}_{{i}}=\frac{{a}_{i}-{b}_{i}}{{a}_{i}}$$
(14)
where ai is the mean distance between point i and all points in different quintiles, and bi is the mean distance between point i and all other points within the same quintile. This metric resembles the standard silhouette score, with two modifications that increase sensitivity to structured but non-clustered data: (1) ai is defined as the mean distance to all points in all other quintiles, rather than only to the closest other quintile as in the standard formulation; and (2) the denominator is ai, rather than max(ai,bi). For each mouse, silhouette scores were averaged across all frames and sessions to obtain a single summary value.
Partner distance and angle tuning
To characterize neuronal tuning to the partner’s egocentric position, we constructed tuning curves for partner distance and angle using calcium imaging data collected across animals and sessions (Fig. 4p–s). Neurons were first selected based on significant tuning to egocentric partner position, identified through 2D egocentric spatial maps as described above (Fig. 4e). For distance tuning, we further selected only neurons of which the partner-distance spatial information exceeded the 95th percentile of the circularly shifted null distribution. For angle tuning, we retained only neurons of which the Rayleigh vector length exceeded the 95th percentile of the circularly shifted null distribution. For each selected neuron, the tuning curve was normalized to its peak firing rate to allow comparison across neurons. The resulting matrices of normalized tuning curves were sorted by each neuron’s peak bin location (distance or angle) and visualized as heat maps. The population distribution of peak distances or angles was summarized using histograms.
Trial-by-trial analysis of partner position-selective activity
To determine whether moment-to-moment fluctuations in partner position-selective neurons predict behavioural outcomes, we quantified neural activity within the 2 s window preceding the recorded animal’s arrival at the reward zone (Extended Data Fig. 14e). Spatial selectivity was defined using the three criteria described above. For each neuron, the social receptive field (SRF) was defined in egocentric coordinates as the region in which partner presence elicits greater than half-maximal activity. Neurons were classified as front-tuned (partner angle −60° to 60° relative to the neck–nose axis) or rear-tuned (all other angles) according to the angular location of their SRF centre.
For each neuron and trial, activity was averaged across frames within the 2-s pre-arrival window during which the partner occupied the neuron’s SRF. Trials without SRF occupancy during this time window were excluded for that neuron. Trial-level responses were converted into percentile ranks within each neuron to express activity as relative trial-by-trial fluctuations independent of absolute firing-rate differences. We examined three behavioural outcomes: (1) whether a follower led on that trial; (2) whether a leader followed; and (3) whether the trial resulted in a mismatch error. To quantify the relationship between neural fluctuations and behavioural outcomes, we fit separate logistic regression models for leaders and followers. In each model, the binary behavioural outcome (for example, leading versus following; mismatch versus correct) was predicted from the trial-level percentile activity of a given neuronal population (for example, front-tuned neurons in leaders). Thus, each model estimated how fluctuations in the activity of a defined neuronal subgroup modulated the log-odds of the behavioural outcome on a trial-by-trial basis.
MARL modelling
We developed a forward MARL model to simulate cooperative behaviour in a spatial foraging task similar to the mouse paradigm. The environment was discretized into an 11 × 11 square grid in which two agents, represented by their spatial coordinates, learned to navigate towards and jointly occupy randomly activated reward zones. The global state at each time step was defined as:
$$s=(\{{p}_{1},{p}_{2},\ldots ,{p}_{N}\},\tau ),$$
(15)
where pi ∈ Z2 denotes the spatial coordinates of agent i, and τ encodes the state of the reward ports (pre-activation versus post-activation). There are n = 2 agents in this simulation.
Each agent i independently learned a state-value function Qi(s), which was updated using temporal-difference learning:
$${Q}_{i}(s)\leftarrow (1-\alpha ){Q}_{i}(s)+\alpha [{r}_{i}(t)+{\gamma }\mathop{\max }\limits_{a\in {A}^{N}}{Q}_{i}({S}^{{\prime} })]$$
(16)
where α is the learning rate, γ is the discount factor, and ri(t) is the reward received at time t. Action selection followed an epsilon-greedy policy over the joint action set AN, where A = {stay, up, down, left, right}. With probability ϵ, a uniformly random action was selected; otherwise, agents evaluated possible joint actions and selected the move leading to the state s′ that maximized Qi(s′). Invalid moves (off-grid) were masked out.
Trials began with agents randomly placed on the grid. Training consisted of two stages: (1) initiation, where agents navigated towards the centre of the arena to activate two reward zones; and (2) cooperative foraging, where both agents were trained to simultaneously arrive at the same active reward zone. Successful coordination yielded a reward for each agent; penalties were applied for entering inactive zones (unrewarded error), choosing different active targets (mismatch error) or exceeding a maximum trial duration without reward (omission). Each movement incurred a small energy cost. Training continued until convergence, with environment resets upon reward collection, errors or omissions.
MAIRL modelling
Pre-processing of foraging trajectories
To pre-process the foraging trajectories, each mouse was represented as a massless point anchored at its neck position. The arena was discretized into a 10 × 10 grid with 5 cm spatial bins, and each animal’s position was assigned to the nearest grid point. To minimize artefacts introduced by discretizing the arena, we excluded time steps where both animals remained stationary. Each animal had nine possible actions: stay, up, down, left, right, up-left, up-right, down-left and down-right.
Model set-up
We modelled the cooperative foraging task as a Markov decision process involving two agents with shared objectives. This Markov decision process is defined by the tuple \(\langle {\mathcal{S}},{\mathcal{A}},{\mathcal{P}},r,\gamma \rangle \), where \({\mathcal{S}}=\,{{\mathcal{S}}}_{1}\times {{\mathcal{S}}}_{2}\) represents the joint environmental state formed as the Cartesian product of individual state sets \({{\mathcal{S}}}_{i}\); in our task, each agent occupies one of the 100 discrete spatial locations, yielding 10,000 possible joint states. \({\mathcal{A}}={{\mathcal{A}}}_{1}\times {{\mathcal{A}}}_{2}\) denotes the set of joint actions, with each agent having nine valid options, resulting in 81 total joint actions. \({\mathcal{P}}=P({s}^{{\prime} }|s,a):\,{\mathcal{S}}\times {\mathcal{A}}\to {\mathcal{\Delta }}({\mathcal{S}})\) defines the transition dynamics, where Δ represents a probability distribution over \({\mathcal{S}}\). We assumed deterministic transitions such that a specific joint action a always leads to a predetermined next state s′. The reward function \({r}:{\mathcal{S}}{\mathbb{\to }}{\mathbb{R}}\) maps the current joint state s to a scalar value. γ ∈ [0, 1] is the discount factor for future rewards.
We then estimated the value map that maximizes the likelihood of the observed trajectory. Following the IRL framework20,68,69, given \(\langle {\mathcal{S}},{\mathcal{A}},{\mathcal{P}},\,\gamma \rangle \) and N trials of both agents’ trajectories D = {ζ1, ζ2, …, ζN}, we inferred the unknown reward functions r such that P(D|r) is maximized. Each trajectory ζi consists of independent joint state-action pairs \({\zeta }_{{i}}=\{({s}_{{t}},{a}_{{t}}){\}}_{t=0}^{T}\). Consequently, the posterior probability of observing expert trajectory ζi can be calculated under the specific policy π derived from value function r, as shown in the following equation:
$$P({\zeta }_{{\rm{i}}}|r)=\mathop{\prod }\limits_{t=0}^{T}\pi ({a}_{{t}}|{s}_{{t}};{r})p({s}_{{t}})$$
(17)
Therefore, the core of this maximization problem is the parameterization of the joint policy function π(at|st;r), using a probabilistic modelling of social interactions based on MARL. Note that policy function π could be derived from value function v using the value iteration algorithm, and both share the same underlying parameterization.
Parameterization of joint value function
To infer the reward function \({r}:{\mathcal{S}}{\mathbb{\to }}{\mathbb{R}}\) in the two-agent foraging task, we must estimate \(|{\mathcal{S}}|=\mathrm{10,000}\) parameters. This is computationally demanding in a multi-agent setting, as the size of the joint state space \(|{\mathcal{S}}|\) grows exponentially with the number of agents. To address this, we implemented a value decomposition approach, allowing the joint value function to be expressed as a sum of marginal and interaction maps, as defined by:
$$r({s}_{1},{s}_{2})=\alpha m({s}_{1})+\alpha n({s}_{2})+\beta \phi (d({s}_{1},{s}_{2}))$$
(18)
where si is agent i’s current location, d denotes the Dijkstra distance between s1 and s2 (that is, the shortest number of steps required to travel between two locations), and α, β are linear coefficients describing the weights of the value map functions. This approach drastically reduces the number of parameters from 10,000 to 212. Numerical analyses and experiments36 demonstrated that this value decomposition achieves negligible reconstruction error relative to the original joint value function, when joint rewards are sparse. Consequently, our objective is to learn α, β, m, n, ϕ to maximize \({\prod }_{t=0}^{T}\pi ({a}_{t}|{s}_{t};r)\) over the observed state-action pairs {st,at}.
Maximum entropy policy formulation
To facilitate the inference problem, we adopted a differentiable maximum-entropy policy formulation to derive the policy from the joint value function69
$$\pi ({a|s})=\frac{\exp Q(s,a)}{{\sum }_{{a}^{{\prime} }{\mathscr{\in }}{\mathcal{A}}}\exp Q(s,{a}^{{\prime} })}$$
(19)
where Q(s,a) is a soft Q-function obtained by performing soft value iteration:
$$Q(s,a)=r(s,a)+\gamma \sum _{{s}^{{\prime} }}P({s}^{{\prime} }{|s},a)\log \,\left(\sum _{{a}^{{\prime} }{\mathscr{\in }}{\mathcal{A}}}\exp Q({s}^{{\prime} },{a}^{{\prime} })\right)$$
(20)
Note that a temperature term is not needed in the softmax function, as it is absorbed into the absolute value of Q and subsequently r, which is the variable we aim to estimate.
Inference algorithm
The inference procedure was detailed previously36. Our objective is to learn α, β, m, n, ϕ to maximize \({\prod }_{t=0}^{T}\pi ({a}_{t}|{s}_{t};r)\) over N observed trials. We assume all parameters follow a Gaussian prior with known variance and zero mean, with the weight coefficients following \(\alpha \sim {\mathcal{N}}(0,{{\sigma }}_{0}^{2})\), and \(\beta \sim {\mathcal{N}}(0,{\sigma }_{0}^{2})\). Incorporating priors on the map functions m, n, ϕ is equivalent to adding an L2 regularizer with coefficients λ1 and λ2 to their entries. The parameter optimization procedure could be written as:
$${\alpha }^{* }={\mathrm{argmax}}_{\alpha }\sum _{({s}_{t},{a}_{t})\in D}\log P({a}_{t}|{s}_{t})-\frac{1}{2{{\sigma }}_{0}^{2}}{\alpha }^{2}$$
(21)
$${\beta }^{* }={{\rm{argmax}}}_{\beta }\sum _{({s}_{t},{a}_{t})\in D}\log \,P({a}_{t}|{s}_{t})-\frac{1}{2{\sigma }_{0}^{2}}{\beta }^{2}$$
(22)
$$\begin{array}{c}({m}^{* },{n}^{* },{{\phi }}^{* })=\,{{\rm{argmax}}}_{m,n,{\phi }}\sum _{({s}_{t},{a}_{t})\in D}\log P({a}_{t}|{s}_{t})\\ \,-{{\lambda }}_{1}{{||m||}}^{2}-{{\lambda }}_{1}{{||n||}}^{2}-{{\lambda }}_{2}{{||}\phi {||}}^{2}\end{array}$$
(23)
We used coordinate ascent to iteratively update the weights α, β and the maps m, n, ϕ while holding the other parameters fixed. Separate learning rates η1 and η2 were used for updating the weights and the maps, respectively.
We used 80% of the trials for model fitting and the remaining 20% to calculate the test log-likelihood as model evidence. Three random searches were initialized with different seeds and the one with the highest test log-likelihood was used. The preset hyperparameters were learning rates η1 = 0.1, η2 = 0.005 and regularization strengths λ1 = 5, λ2 = 1 selected for interpretability of the recovered value functions of the physical locations (that is, m(s1) and n(s2)). Without appropriate regularization, the recovered value maps tended to exhibit randomly activated regions. The exact values of the regularization parameters were not critical, provided that they imposed sufficient sparsity constraints. All of the other hyperparameters were searched over the following ranges: γ ∈ {0.70, 0.90, 0.99} and \({\sigma }_{0}^{2}\in \{0.02,\,0.2,\,1.0\}\). Based on test set log-likelihood and convergence speed, we chose \(\gamma =0.90\), \({\sigma }_{0}^{2}=1.0.\)
Estimation of individual trajectories
In the above sections, we maximize the likelihood of observing the joint action pair, P(a|s) = P(a1,a2|s1,s2), which incorporates information from both agents. Given the asymmetric behaviour observed in the cooperative foraging task, it is also informative to estimate the individual value functions that govern each animal’s decision-making. In this section, subscripts refer to individual agents, and the time index t is omitted for clarity. Assuming that animals have independent control over their decisions, we separate the probability into
$$P({a}_{1},{a}_{2}|{s}_{1},{s}_{2})=P({a}_{1}|{s}_{1},{s}_{2})P({a}_{2}|{s}_{1},{s}_{2})$$
(24)
This formulation enables us to model the decision-making process of an individual animal. Without loss of generality, we aimed to identify the reward function \({r}^{1}({s}_{1},{s}_{2})\) for animal 1 such that \(P({a}_{1}|{s}_{1},{s}_{2};{r}^{1})\) is maximized. Following the policy derivation described in the earlier section, this reward function \({r}^{1}\) defines a joint policy \(\pi ({a}_{1},{a}_{2}|{s}_{1},{s}_{2};{r}^{1})\), from which we derived the marginal policy:
$$\pi ({a}_{1}|{s}_{1},{s}_{2};{r}^{1})=\sum _{{a}_{2}}\pi ({a}_{1}|{s}_{1},{s}_{2};{r}^{1})$$
(25)
The remaining steps in the inference procedure followed accordingly, using the marginalized policy in place of the joint policy.
Incorporating egocentric partner angle
The partner’s angle θ was calculated as described in the previous sections. The angle was then discretized into the front (−90° to 90°) or the rear (90° to 180° and −180° to −90°) fields. Consequently, the discretized θ took the value of either 0 or 1. Angle information was provided to the Markov chain as input. We therefore maximized P(a|s,θ) given r(s|θ). To further simplify the dependence on angle, we allowed the decomposed joint value function to depend on \(\theta \) only through the interaction term:
$$r(s|\theta )=\alpha m({s}_{1})+\alpha n({s}_{2})+\beta \phi (d|\theta )$$
(26)
Model selection and comparison
We focused on different parameterizations of the decomposed joint value function for model selection (Fig. 5 and Supplementary Table 2). These models were nested in complexity. For each model, we estimated its fit after convergence of the inference procedure described in the ‘Inference algorithm’ section. Model comparisons between nested models were conducted using a χ2 test, with degrees of freedom equal to the difference in the number of parameters. The baseline model (model with the allocentric position of self) was additionally compared to a constant prediction model, whose log-likelihood was computed as the number of decision pairs times \(\log \frac{1}{9}\).
Simulation of foraging trajectory to predict reward zone choice and performance
To intuitively assess model fitting performance, we simulated foraging trajectories using the inferred joint value function and predicted the reward zone chosen by the animals. To simulate the trajectory for a mouse in each trial, we initialized the mouse at the same starting position as in the observed data. At each time step, the animal’s action was sampled from its marginalized policy \(\pi ({a}_{1}|{s}_{1},{s}_{2};{r}^{1})\), derived from the inferred joint value function. The selected action determined the animal’s next location, while its partner’s location was taken directly from the observed trajectory. This enabled us to simulate individual decisions in a two-agent foraging set-up. The simulation terminated when the animal entered either active reward zone. The prediction was considered to be correct if the simulated choice matched the observed choice (Fig. 5c). A trial was scored as successful cooperation if the animal’s simulated choice matched that of its partner (Fig. 5j). This process was repeated for all the trials in each session, and the correct prediction rate was calculated for each model. Note that this procedure was not applied to models incorporating partner angle. Whereas egocentric partner angle is directly available from the observed trajectories for likelihood estimation, it cannot be determined during trajectory simulation.
Decoding of inferred value from population neural activity
To determine whether mPFC population activity encodes inferred value functions from our model, we performed linear decoding of value estimates \(r({s}_{t})\in {{\mathbb{R}}}^{1\times T}\) from population activity \({N}_{t}\in {{\mathbb{R}}}^{N\times T}\), on a session-by-session basis. The dataset was randomly split into 80% of the time steps for training and 20% for testing. To prevent over-representation of repetitive datapoints, consecutive frames with identical inferred total value were removed, which typically occurred during prolonged stationary periods (speed <2 cm s−1). Model performance was quantified using the R2 score on the test set. To assess statistical significance, we generated a null distribution by applying circular time shifts to the neural data and repeating the decoding procedure 1,000 times (Extended Data Fig. 15c). The P value was calculated by fitting a Gaussian distribution to the null R2 scores and computing the percentile rank of the observed R2 within this distribution. The minimum P value was capped at 10−16 for numerical precision.
Statistics and reproducibility
Behavioural, chemogenetic, optogenetic, calcium imaging and electrophysiological validation experiments were performed in independent cohorts of animals, and all quantitative analyses were based on the full datasets described in the figures and Supplementary Table 1. Reproducibility is reflected by the number of independent biological units included in each analysis (animals, animal pairs, recording sessions or neurons), which are reported throughout the Article. No statistical methods were used to predetermine sample size. Sample sizes were chosen based on previous studies and our experience with similar behavioural, perturbation and calcium imaging experiments, and are consistent with those commonly used in the field. Animals were randomly assigned to experimental groups whenever applicable. For manipulation experiments, animals were randomly allocated to treatment conditions before behavioural testing. Experimenters were blinded to experimental conditions when possible. They were not blinded in inactivation experiments, in which animal identities had to be tracked to assign the correct social roles, or in calcium imaging experiments, in which only one animal can be imaged at a time.
For representative behavioural examples (Figs. 1c and 2a,b), example sessions and video frames were selected from datasets representative of the behavioural patterns observed across all analysed animal pairs. For representative neuronal activity examples (Figs. 3u,v and 4a–d,f,g), examples were selected from datasets obtained across all recorded animals and illustrated effects quantified at the population level. For representative histological images and validation experiments (such as viral expression, implant targeting and electrophysiological validation), similar results were observed across the animals included in the corresponding analyses. Behavioural analyses in Fig. 1 drew on nested subsets of the full training cohort depending on data requirements (for example, availability of complete video and trial-level data, number of well-trained sessions or reliable role assignment); the exact number of pairs is reported in each figure panel. No experiments were excluded from analysis except according to predefined criteria described in the Methods.
Use of large language models
Large language models (ChatGPT, OpenAI; Claude, Anthropic) were used to assist with language editing and refinement of the Article text. All scientific content, interpretations and conclusions were developed by the authors.
Reporting summary
Further information on research design is available in the Nature Portfolio Reporting Summary linked to this article.
Data availability
The data are available at Zenodo70 (https://doi.org/10.5281/zenodo.20723195). Source data are provided with this paper.
Code availability
The code for analysis in this paper is available at GitHub (https://github.com/HerbertWuLab/Yuan_manuscript_2026).
References
King, A. J., Johnson, D. D. & Van Vugt, M. The origins and evolution of leadership. Curr. Biol. 19, R911–916 (2009).
Article CAS PubMed Google Scholar
Wilson, E. O. Sociobiology: The New Synthesis (Belknap Press of Harvard Univ. Press, 1975).
Axelrod, R. & Hamilton, W. D. The evolution of cooperation. Science 211, 1390–1396 (1981).
Article ADS MathSciNet CAS PubMed Google Scholar
Tomasello, M. Why We Cooperate (MIT Press, 2009).
Couzin, I. D., Krause, J., Franks, N. R. & Levin, S. A. Effective leadership and decision-making in animal groups on the move. Nature 433, 513–516 (2005).
Article ADS CAS PubMed Google Scholar
Nagy, M., Akos, Z., Biro, D. & Vicsek, T. Hierarchical group dynamics in pigeon flocks. Nature 464, 890–893 (2010).
Article ADS CAS PubMed Google Scholar
Van Vugt, M., Hogan, R. & Kaiser, R. B. Leadership, followership, and evolution: some lessons from the past. Am. Psychol. 63, 182–196 (2008).
Article PubMed Google Scholar
Smith, J. E. et al. Leadership in Mammalian societies: emergence, distribution, power, and payoff. Trends Ecol. Evol. 31, 54–66 (2016).
Article PubMed Google Scholar
Boesch, C. Cooperative hunting roles among taï chimpanzees. Hum. Nat. 13, 27–46 (2002).
Article PubMed Google Scholar
Hermalin, B. E. Toward an economic theory of leadership: leading by example. Am. Econ. Rev. 88, 1188–1206 (1998).
Google Scholar
Stander, P. E. Cooperative hunting in lions: the role of the individual. Behav. Ecol. Sociobiol. 29, 445–454 (1992).
Article Google Scholar
Luo, L., Callaway, E. M. & Svoboda, K. Genetic dissection of neural circuits: a decade of progress. Neuron 98, 256–281 (2018).
Article CAS PubMed PubMed Central Google Scholar
Carandini, M. & Churchland, A. K. Probing perceptual decisions in rodents. Nat. Neurosci. 16, 824–831 (2013).
Article CAS PubMed PubMed Central Google Scholar
Rilling, J. K. & Sanfey, A. G. The neuroscience of social decision-making. Annu. Rev. Psychol. 62, 23–48 (2011).
Article PubMed Google Scholar
Wang, H. & Kwan, A. C. Competitive and cooperative games for probing the neural basis of social decision-making in animals. Neurosci. Biobehav. Rev. 149, 105158 (2023).
Article PubMed PubMed Central Google Scholar
von Neumann, J., Morgenstern, O. & Rubinstein, A. Theory of Games and Economic Behavior (60th Anniversary Commemorative Edition) (Princeton Univ. Press, 1944).
Yu, J. H., Napoli, J. L. & Lovett-Barron, M. Understanding collective behavior through neurobiology. Curr. Opin. Neurobiol. 86, 102866 (2024).
Article CAS PubMed PubMed Central Google Scholar
Erlich, J. C. et al. Mice dynamically adapt to opponents in competitive multi-player games. Preprint at bioRxiv https://doi.org/10.1101/2025.02.14.638359 (2025).
Arora, S. & Doshi, P. A survey of inverse reinforcement learning: challenges, methods and progress. Artif. Intell. 297, 103500 (2021).
Article MathSciNet Google Scholar
Ashwood, Z., Jha, A. & Pillow, J. W. Dynamic inverse reinforcement learning for characterizing animal behavior. Adv. Neural Inform. Process. Syst. 35, 29663–29676 (2022).
Article Google Scholar
Wang, F. et al. Bidirectional control of social hierarchy by synaptic efficacy in medial prefrontal cortex. Science 334, 693–697 (2011).
Article ADS CAS PubMed Google Scholar
Zhou, T. et al. History of winning remodels thalamo-PFC circuit to reinforce social dominance. Science 357, 162–168 (2017).
Article ADS CAS PubMed Google Scholar
LeClair, K. B. et al. Individual history of winning and hierarchy landscape influence stress susceptibility in mice. eLife https://doi.org/10.7554/eLife.71401 (2021).
Padilla-Coreano, N. et al. Cortical ensembles orchestrate social competition through hypothalamic outputs. Nature 603, 667–671 (2022).
Article ADS CAS PubMed PubMed Central Google Scholar
Cum, M. et al. A multiparadigm approach to characterize dominance behaviors in CD1 and C57BL6 male mice. eNeuro https://doi.org/10.1523/ENEURO.0342-24.2024 (2024).
Chindemi, G., Girard, B. & Bellone, C. LISBET: a machine learning model for the automatic segmentation of social behavior motifs. Preprint at https://arxiv.org/abs/2311.04069 (2023).
Amodio, D. M. & Frith, C. D. Meeting of minds: the medial frontal cortex and social cognition. Nat. Rev. Neurosci. 7, 268–277 https://doi.org/10.1038/nrn1884 (2006).
Article CAS PubMed Google Scholar
Apps, M. A., Rushworth, M. F. & Chang, S. W. The anterior cingulate gyrus and social cognition: tracking the motivation of others. Neuron 90, 692–707 (2016).
Article CAS PubMed PubMed Central Google Scholar
Yizhar, O. et al. Neocortical excitation/inhibition balance in information processing and social dysfunction. Nature 477, 171–178 (2011).
Article ADS CAS PubMed PubMed Central Google Scholar
Chang, S. W., Gariepy, J. F. & Platt, M. L. Neuronal reference frames for social decisions in primate frontal cortex. Nat. Neurosci. 16, 243–250 (2013).
Article CAS PubMed Google Scholar
Yamamuro, K. et al. A prefrontal-paraventricular thalamus circuit requires juvenile social experience to regulate adult sociability in mice. Nat. Neurosci. 23, 1240–1252 (2020).
Article PubMed PubMed Central Google Scholar
Li, S. W. et al. Frontal neurons driving competitive behaviour and ecology of social groups. Nature 603, 661–666 (2022).
Article ADS CAS PubMed PubMed Central Google Scholar
Jiang, M. et al. Evolution and neural representation of mammalian cooperative behavior. Cell Rep. 37, 110029 (2021).
Article CAS PubMed Google Scholar
Jiang, M. et al. Neural basis of cooperative behavior in biological and artificial intelligence systems. Science 391, eadw8151 (2026).
Article CAS PubMed PubMed Central Google Scholar
Schneider, S., Lee, J. H. & Mathis, M. W. Learnable latent embeddings for joint behavioural and neural analysis. Nature 617, 360–368 (2023).
Article ADS CAS PubMed PubMed Central Google Scholar
Chen, Y., Cheng, Y., Kwak, M., Radulescu, A. & Wu, H. Z. Unveiling value functions in social cognition with multi-agent inverse reinforcement learning. Preprint at bioRxiv https://doi.org/10.1101/2024.10.09.617461 (2026).
Van Vugt, M. Evolutionary origins of leadership and followership. Pers. Soc. Psychol. Rev. 10, 354–371 (2006).
Article PubMed Google Scholar
Rands, S. A., Cowlishaw, G., Pettifor, R. A., Rowcliffe, J. M. & Johnstone, R. A. Spontaneous emergence of leaders and followers in foraging pairs. Nature 423, 432–434 (2003).
Article ADS CAS PubMed Google Scholar
Solié, C. et al. Dopaminergic mechanisms of dynamical social specialization. Nature 654, 163–172 (2026).
Eagly, A. H. Sex Differences in Social Behavior: A Social-role interpretation 1st edn (Psychology Press, 1987).
Dulac, C., O’Connell, L. A. & Wu, Z. Neural control of maternal and paternal behaviors. Science 345, 765–770 (2014).
Article ADS CAS PubMed PubMed Central Google Scholar
Wei, D., Talwar, V. & Lin, D. Neural circuits of social behaviors: Innate yet flexible. Neuron 109, 1600–1620 (2021).
Article CAS PubMed PubMed Central Google Scholar
Haroush, K. & Williams, Z. M. Neuronal prediction of opponent’s behavior during cooperative social interchange in primates. Cell 160, 1233–1245 (2015).
Article CAS PubMed PubMed Central Google Scholar
Shi, W., Meisner, O. C., Jadi, M. P., Chang, S. W. C. & Nandy, A. S. Canonical decision computations underlie behavioral and neural signatures of cooperation in primates. Neuron 114, 2657–2667.e5 (2026).
Witkowski, P. P., Park, S. A. & Boorman, E. D. Neural mechanisms of credit assignment for inferred relationships in a structured world. Neuron 110, 2680–2690 (2022).
Article CAS PubMed Google Scholar
Danjo, T., Toyoizumi, T. & Fujisawa, S. Spatial representations of self and other in the hippocampus. Science 359, 213–218 (2018).
Article ADS CAS PubMed Google Scholar
Omer, D. B., Maimon, S. R., Las, L. & Ulanovsky, N. Social place-cells in the bat hippocampus. Science 359, 218–224 (2018).
Article ADS CAS PubMed Google Scholar
Zhang, X. et al. Multiplexed representation of others in the hippocampal CA1 subfield of female mice. Nat. Commun. 15, 3702 (2024).
Article ADS CAS PubMed PubMed Central Google Scholar
Murugan, M. et al. Combined social and spatial coding in a descending projection from the prefrontal cortex. Cell 171, 1663–1677 (2017).
Article CAS PubMed PubMed Central Google Scholar
Rangel, A., Camerer, C. & Montague, P. R. A framework for studying the neurobiology of value-based decision making. Nat. Rev. Neurosci. 9, 545–556 (2008).
Article CAS PubMed PubMed Central Google Scholar
Veselic, S. et al. A cognitive map for value-guided choice in the ventromedial prefrontal cortex. Cell 188, 3259–3273 (2025).
Article CAS PubMed PubMed Central Google Scholar
Radulescu, A., Shin, Y. S. & Niv, Y. Human representation learning. Annu. Rev. Neurosci. 44, 253–273 (2021).
Article CAS PubMed PubMed Central Google Scholar
Wittmann, M. K. et al. Basis functions for complex social decisions in dorsomedial frontal cortex. Nature 641, 707–717 (2025).
Frith, U. & Frith, C. D. Development and neurophysiology of mentalizing. Philos. Trans. R. Soc. Lond. B. 358, 459–473 (2003).
Article Google Scholar
Pereira, T. D. et al. SLEAP: A deep learning system for multi-animal pose tracking. Nat. Methods 19, 486–495 (2022).
Article CAS PubMed PubMed Central Google Scholar
Biderman, D. et al. Lightning Pose: improved animal pose estimation via semi-supervised learning, Bayesian ensembling and cloud-native open-source tools. Nat. Methods 21, 1316–1328 (2024).
Article CAS PubMed PubMed Central Google Scholar
Fan, Z. et al. Using the tube test to measure social hierarchy in mice. Nat. Protoc. 14, 819–831 (2019).
Article CAS PubMed Google Scholar
Friard, O. & Gamba, M. BORIS: a free, versatile open-source event-logging software for video/audio coding and live observations. Methods Ecol. Evol. 7, 1325–1330 (2016).
Article Google Scholar
Elo, A. E. The Rating of Chessplayers, Past and Present. (ARCO, 1978).
Albers, P. C. H. & de Vries, H. Elo-rating as a tool in the sequential estimation of dominance strengths. Anim. Behav. 61, 489–495 (2001).
Article Google Scholar
Cai, D. J. et al. A shared neural ensemble links distinct contextual memories encoded close in time. Nature 534, 115–118 (2016).
Article ADS CAS PubMed PubMed Central Google Scholar
Pachitariu, M. et al. Suite2p: beyond 10,000 neurons with standard two-photon microscopy. Preprint at bioRxiv https://doi.org/10.1101/061507 (2017).
Rupprecht, P. et al. A database and deep learning toolbox for noise-optimized, generalized spike inference from calcium imaging. Nat. Neurosci. 24, 1324–1337 (2021).
Article CAS PubMed PubMed Central Google Scholar
Pachitariu, M., Sridhar, S., Pennington, J. & Stringer, C. Spike sorting with Kilosort4. Nat. Methods 21, 914–921 (2024).
Article CAS PubMed PubMed Central Google Scholar
Rossant, C. et al. Spike sorting for large, dense electrode arrays. Nat. Neurosci. 19, 634–641 (2016).
Article CAS PubMed PubMed Central Google Scholar
Fan, R.-E., Chang, K.-W., Hsieh, C.-J., Wang, X.-R. & Lin, C.-J. LIBLINEAR: A library for large linear classification. J. Mach. Learn. Res. 9, 1871–1874 (2008).
Google Scholar
Skaggs, W., McNaughton, B., Gothard, K. & Markus, E. An information-theoretic approach to deciphering the hippocampal code. Adv. Neural Inform. Process. Syst. 5, 1030–1037 (1993).
Google Scholar
Abbeel, P. & Ng, A. Y. Apprenticeship learning via inverse reinforcement learning. In Proc. Twenty-First International Conference on Machine Learning 1 (ACM Press, 2004).
Ziebart, B. D., Maas, A. L., Bagnell, J. A. & Dey, A. K. Maximum entropy inverse reinforcement learning. In Proc. Twenty-Third AAAI Conference on Artificial Intelligence 1433–1438 (AAAI Press, 2008).
Cheng, Y. et al. Data for ‘Asymmetric prefrontal representations for leader–follower dynamics’. Zenodo https://doi.org/10.5281/zenodo.20723195 (2026).
Download references
Acknowledgements
We thank K. Deisseroth for the AAV-CaMKIIα-KALI1-eYFP construct; L. Paninski for advice on behavioural analysis; S. J. Russo, D. Cai, T. Shuman, B. Sweis, H. Morishita, C. A. Duan, X. H. Hou, X. Wu, Z. Pennington, A. Baggetta, Y. Zaki, Z. Wick, P. Philipsberg, F. K. Chiang, L. Li, R. D.-d. Cuttoli, H. Li and I. Orsolic for advice on experiments; W. Janssen, B. Wu, E. Sullaway, C. Shamsu and N. Tillison for assistance; and P. J. Kenny, P. Rudebeck, M. N. Shadlen and C. Dulac for comments on the manuscript.
Funding
This work was supported by the Alfred P. Sloan Foundation, the Chan Zuckerberg Initiative, the Simons Foundation, the Esther A. & Joseph Klingenstein Fund and the Friedman Brain Institute (H.Z.W.); a Seaver Fellowship (Y. Cheng); a Swartz Postdoctoral Fellowship (Y. Chen); the Air Force Office of Scientific Research (FA9550-22-1-0337) (N.R.); and the NIMH (MH132653) (T.P.).
Ethics declarations
Competing interests
H.Z.W. is listed as an inventor on a patent applied for by the Icahn School of Medicine at Mount Sinai that covers the cooperative foraging paradigm described here. The other authors declare no competing interests.
Peer review
Peer review information
Nature thanks Andry Andrianarivelo, Camilla Bellone, Steve Chang and the other, anonymous, reviewer(s) for their contribution to the peer review of this work. Peer reviewer reports are available.
Additional information
Publisher’s note Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
Extended data figures and tables
Extended Data Fig. 1 Training progression and emergence of social roles.
a, Learning curves for female (red, n = 62 pairs) and male (blue, n = 27 pairs) pairs; each line, one pair. The dashed line indicates the 80% correct training criterion. b, Unrewarded error rates are low during the learning (median = 0.016, IQR = 0.011) and well-trained (median = 0.0051, IQR = 0.0096) stages, and decrease significantly (two-tailed Wilcoxon signed-rank test, ***P = 4.0 × 10−6; n = 40 pairs). Points, individual pairs; line, median. c, Omission rates are low during the learning (median = 0.015, IQR = 0.015) and well-trained (median = 0.012, IQR = 0.018) stages and do not differ (P = 0.39; n = 40 pairs). Points, individual pairs; line, median. d–f, Correct rates for female and male mouse pairs in the first (d), second (e), and third (f) sessions at the training criterion. Females perform slightly better than males on the second and third sessions. n = 62 female, 27 male pairs. Two-tailed Wilcoxon rank-sum test: P = 0.25 (d), *P = 0.020 (e), *P = 0.017 (f). Bars, mean ± SEM. g, Time difference in nose poke at the reward port between partners on correct trials, first training day versus first criterion day. Points, individual pairs; black line, median (two-tailed Wilcoxon signed-rank test, **P = 0.0026; n = 15 pairs). h, Definition of reward zones in the arena. The diagram was created using BioRender; Wu, H. https://BioRender.com/hwyq5u6 (2026). i,j, Probability that a trial is led by the leader as a function of the leader’s and follower’s distance to (i) or heading relative to (j) the chosen reward zone at trial onset. k, Proportion of trials led by leaders defined by arrival time correlates with the proportion of trials in which that same animal first oriented toward the chosen reward zone (Pearson correlation, R = 0.71, P = 3.3 × 10−64, n = 412 sessions from 40 pairs). Each point, one session from a mouse pair; thin lines, linear fits for individual pairs; thick line, all sessions combined. l–n, Role differentiation and learning in a second example pair with faster learning than the pair shown in Fig. 1d–f. l, Cooperation rate alongside the proportion of trials led (solid lines) or initiated (dashed lines) by one mouse across training days. m, Correlation between leader asymmetry and cooperation rate. Pearson correlation: R = 0.96, P = 1.3 × 10−4; n = 8 sessions. n, Correlation between initiator asymmetry and cooperation rate. R = 0.74, P = 0.034. Each point, one session. o, p, Correlation between role asymmetry and cooperation rate during learning across mouse pairs, with data grouped by sex (red, females; blue, males). Each point represents one session from a mouse pair. Thin lines indicate linear fits for individual pairs, and the thick line indicates the fit across all sessions combined. Leader asymmetry is strongly correlated with cooperation rate in both females and males (o, females, linear mixed-effects model, P = 1.9 × 10−22; n = 29 pairs, 479 sessions; males, P = 1.3 × 10−14; n = 11 pairs, 202 sessions; P = 0.77 for the effect of sex), whereas initiator asymmetry shows a weaker correlation with greater variability across pairs (p, females, linear mixed-effects model, P = 0.049, n = 29 pairs, 479 sessions; males, P = 0.057; n = 11 pairs, 202 sessions; P = 0.25 for the effect of sex). q,r, Cooperation rate in trials led (q) or initiated (r) by each mouse, grouped by sex. Each point represents one session from a mouse pair. Lines indicate linear fits for individual pairs, and the dashed line indicates unity. Performance is significantly higher in trials led by leaders (q, females, linear mixed-effects model, P = 2.2 × 10−9, n = 29 pairs, 473 sessions; males, P = 2.9 × 10−10; n = 11 pairs, 200 sessions; P = 0.45 for the effect of sex) and in trials initiated by initiators (r; females, linear mixed-effects model, P = 0.0055, n = 29 pairs, 475 sessions; males, P = 0.0092; n = 11 pairs, 201 sessions; P = 0.61 for the effect of sex).
Source data
Extended Data Fig. 2 Learning, role differentiation, and solo foraging performance.
a, Learning performance stratified by trial arrival order. Left, cooperation rate across training in leader-led and follower-led trials (n = 40 pairs); each line, one pair (purple, leader-led; green, follower-led). Middle and right, cooperation rates increase significantly from the first training day to the first criterion day in both leader-led (middle; two-tailed Wilcoxon signed-rank test, P = 6.0 × 10−5) and follower-led trials (right; P = 6.0 × 10−5); thin lines, individual pairs; thick line, median. b, Learning performance stratified by trial initiation. Left, cooperation rate across training in initiator-initiated and responder-initiated trials (n = 40 pairs). On some days, one mouse did not initiate any trials, resulting in missing data points for that condition. Middle and right, cooperation rates increase significantly from the first training day to the first criterion day in initiator-initiated (middle; P = 6.0 × 10−5) and responder-initiated trials (right; P = 8.0 × 10−5). c,d, Leaders and followers adopt initiator or responder roles across pairs, grouped by sex; only pairs with reliable leader and initiator roles are included (two-tailed binomial test against 0.5, α = 0.01). Each point, one mouse. The proportions of leader-initiated trials do not differ from chance in either females (c, two-tailed Wilcoxon signed-rank test against 0.5, P = 0.57, n = 24 pairs) or males (d, P = 0.38, n = 8 pairs). e, Early solo foraging (Stage 2) metrics (correct rate, mean speed, 95th percentile speed, reaction time in correct trials, trials to training criterion, and correct trials completed per session) are compared between sexes. Sex differences are observed only for correct trials completed per session, consistent with females’ lower body weight and correspondingly reduced water requirement under water restriction (two-tailed Wilcoxon rank-sum test, P = 0.0021, n = 16 females and 16 males); other metrics show no sex differences (correct rate, P = 0.75; mean speed, P = 0.44; 95th percentile speed, P = 0.61; reaction time, P = 0.46; trials to criterion, P = 0.16). Bars, mean ± SEM; points, individual mice. f, Solo foraging metrics do not predict future leader-follower roles (two-tailed Wilcoxon signed-rank test, correct rate, P = 0.43, n = 8 female and 8 male leaders, and 8 female and 8 male followers; mean speed, P = 0.38; 95th percentile speed, P = 0.84; reaction time, P = 0.53; trials to criterion, P = 0.28; correct trials completed per session: females, P = 0.33; males, P = 0.25; n = 16 leader-follower pairs, including 8 female and 8 male pairs). Bars, median; points, individual mice. g, Solo foraging metrics do not predict future initiator-responder roles (correct rate, P = 0.12, n = 8 female and 8 male initiators, and 8 female and 8 male responders; mean speed, P = 0.96; 95th percentile speed, P = 0.38; reaction time, #P = 0.098; trials to criterion, P = 0.92; correct trials completed per session: females, P = 0.33; males, P = 0.25).
Source data
Extended Data Fig. 3 Dominance hierarchies are dissociable from leadership and initiatorship roles.
a, Stability of dominance relationships assessed using the tube test in a round-robin design. Left, representative results from one group of four mice. Right, results pooled across 12 groups. Each point indicates the mean rank of animals within each rank category defined on the previous day. b, Stability of dominance relationships assessed using the tube test, shown separately for female groups (left, n = 6) and male groups (right, n = 6). c, Tube test outcomes for leaders and followers on the criterion day. Each cell indicates the number of leader-follower pairs (two pairs per group of four animals) with the corresponding combination of leader and follower wins. The win rate does not differ between leaders and followers (two-tailed Wilcoxon signed-rank test, P = 0.23, n = 23 total leader-follower pairs). d, Same analysis as in c, shown separately for female (left) and male (right) leader-follower pairs (two-tailed Wilcoxon signed-rank tests: females, P = 0.20, n = 12 pairs; males, P = 0.71, n = 11 pairs). e,j,o, Schematics of the three social hierarchy assays: tube test (e), warm spot test (j), and reward competition test (o). f,k,p, Proportion of pairwise comparisons that are stable in each assay, shown separately for females and males (3 female and 3 male groups, each consisting of 4 animals; 6 pairwise comparisons per group). Stability criteria are described in Methods. In the tube test, all groups reach stable rankings by design. g,l,q, Agreement of dominance measures across assays. g, Cross-assay agreement is evaluated only for pairwise dominance relationships that are stable in both assays being compared (reward competition vs. tube test; tube test vs. warm spot; reward competition vs. warm spot). l, Time spent in the warm spot plotted as a function of tube-test rank (n = 6 groups of 4 animals). Bars, mean ± SEM. q, Percentage of trials won in the pairwise reward competition test for dominant (dom) and subordinate (sub) mice, as determined by the tube test (n = 6 groups with 6 pairwise comparisons per group). Bars, mean ± SEM. h,m,r, Correlation matrices of Elo dominance scores across assays for all mice (h, n = 24 animals in 6 groups of 4), females (m, n = 12 animals), and males (r, n = 12 animals). Pearson correlation coefficients are shown within each matrix (h, all mice: reward competition vs. tube test, R = 0.70; warm spot vs. tube test, R = 0.86; warm spot vs. reward competition, R = 0.62; m, females: reward competition vs. tube test, R = 0.56; warm spot vs. tube test, R = 0.92; warm spot vs. reward competition, R = 0.63; r, males: reward competition vs. tube test, R = 0.84; warm spot vs. tube test, R = 0.80; warm spot vs. reward competition, R = 0.61). i,n,s, The proportion of trials led or initiated by the dominant animal, as defined by each corresponding dominance assay, does not differ from chance (two-tailed Wilcoxon signed-rank tests against 0.5; tube test: leadership, P = 0.97; initiation, P = 0.18; warm spot: leadership, P = 0.15; initiation, P = 0.73; reward competition: leadership, P = 0.57; initiation, P = 0.97; n = 12 dominant animals for all assays). Resp, responder. t, Number of dominant and subordinate mice (determined by tube test across training stages) identified as leaders in the well-trained stage. Left: females; right: males. The proportions do not differ from chance at any training stage (two-tailed binomial test against 0.5, P > 0.17 for all four stages in females, P > 0.37 in males). No significant changes in proportions are observed across stages (chi-square test, females: χ²(3, n = 36) = 1.00, P = 0.80; males: χ²(3, n = 20) = 0.66, P = 0.88). u, Number of dominant and subordinate mice identified as initiators across training stages within the same cohort. Left: all animals; middle: females; right: males. The proportions do not differ from chance at any stage (two-tailed binomial test against 0.5, P > 0.17 in all animals, P > 0.17 in females, P = 1.00 in males). No significant changes in proportions are observed across stages (all animals: χ²(3, n = 56) = 2.58, P = 0.46; females: χ²(3, n = 36) = 3.00, P = 0.39; males: χ²(3, n = 20) = 0.61, P = 0.90). The diagrams in e, j and o were created using BioRender; Wu, H. https://BioRender.com/hwyq5u6 (2026).
Source data
Extended Data Fig. 4 Social roles are preserved across partner swaps and reorganize independently of dominance hierarchy.
a,b, Leader-follower roles are consistently preserved after swapping partners: females, 8/9 pairs (a); males, 4/5 pairs (b). Only pairs with significant leaders (two-tailed binomial test against 0.5, α = 0.01) in original pairing are included. Each point, one animal. c,d, Swapped pairs reach the training criterion faster than original pairs in both sexes, reaching significance in females (c, two-tailed Kolmogorov-Smirnov test, P = 0.0069, n = 10 pairs) and trending in males (d, P = 0.077, n = 6 pairs). e,f, Role reorganization following pairing of two original leaders or two original followers. Each arrow represents a swapped pairing. Open circles indicate the proportions led by each mouse in their respective original pairs; arrowheads, the newly formed pair. Convergence on the diagonal indicates role adjustment to re-establish new leaders and followers in 7 of the 10 new female pairs (e) and 6 of the 6 male pairs (f), with no significant sex difference in proportions (two-proportion z-test, P = 0.14). g,h, Dominance hierarchy does not predict social role assignment during original training (g) or after same-role pairing (h). The proportion of trials led or initiated by the dominant mice does not differ from chance (two-tailed Wilcoxon signed-rank tests against 0.5, original training: trials led, P = 0.57, trials initiated, P = 0.30; same-role pairing: trials led, P = 0.61, trials initiated, P = 0.11; n = 12 animals). Each point represents an animal, with females shown in red and males in blue. i, Among animals that switch leader-follower roles after same-role pairing, the number that are dominant (determined by the tube test). Bars, total number of switching animals; shading, the dominant subset. The proportion does not differ from chance (0.5), whether all animals are combined (binomial test, P = 0.39, n = 12 animals) or analysed by category (P = 0.69). j,k, The proportion of trials led during original training does not predict which animal assumes the leader role after same-role pairing when two leaders are paired (j; two-tailed Wilcoxon signed-rank test, P = 0.95, n = 8 leader dyads) or when two followers are paired (k; P = 0.15, n = 8 follower dyads). Thin lines, individual pairs; thick line, median.
Source data
Extended Data Fig. 5 Representative frames and trajectory overlays of behavioural motifs.
a, Representative video frames (233-ms bins) illustrating each behavioural motif. b, Overlaid neck trajectories across all instances of each motif within a session. Each trace represents one trial. Orange traces denote the actor (the mouse performing the tracking, joining, or sharp turn action); no actor is defined for the synchronized travel motif. Frames and overlays are representative of motifs observed across 50 sessions from 20 pairs.
Extended Data Fig. 6 Distinct kinematic dynamics across behavioural motifs.
a, Schematic illustrating how inter-animal distance was computed. b, Temporal profiles of inter-animal distance for the four behavioural motifs, showing distinct motif-specific dynamics. Black lines indicate the mean across instances and grey shading indicates the standard deviation (same conventions in d and f). Sync, n = 682; track, n = 359; sharp turn, n = 1,395; join, n = 3,853; pooled across 50 sessions from 20 pairs; same sample sizes in d, f, and g. c, Schematic illustrating the calculation of head angle difference between the two animals. d, Temporal profiles of head angle difference for each motif. e, Schematic illustrating the neck speed. f, Temporal profiles of neck speed for each motif. g, Joint trajectories of neck speed (x-axis) and inter-animal distance (y-axis) for each behavioural motif, revealing distinct motif-specific structure in feature space. The thick line indicates the mean trajectory, and the thin lines show 1% of randomly selected individual trajectories. Green and red points denote the starting and ending positions of the mean trajectory, respectively. The diagrams in a, c and e were created using BioRender; Wu, H. https://BioRender.com/hwyq5u6 (2026).
Source data
Extended Data Fig. 7 Behavioural motifs validated by unsupervised classification and their relationship to sex and following behaviour.
a–e, High-dimensional LISBET embeddings of mouse body-part kinematics during behavioural motifs, visualized in 2D by UMAP and coloured by human-defined motif labels. Embeddings are computed without access to the human labels. a, Example mouse pair. Left, LISBET embeddings coloured by human-defined motif labels; right, the same embeddings with temporally shifted labels as control. b, LISBET embeddings for all mouse pairs, shown in the same format as in a. c, Modified silhouette scores of LISBET embeddings with aligned and shuffled motif labels across mouse pairs (one-tailed Wilcoxon signed-rank test, P = 9.5 × 10−7; n = 20 pairs). Each point represents the average score for one mouse pair across all frames and sessions; higher values indicate stronger organization of embeddings by motif labels (see Methods). d,e, LISBET embeddings for female (d) or male (e) pairs, shown in the same format as in a. f, Principal component analysis (PCA) of motif usage across pairs. Each point represents one pair and is coloured by sex. The first two PCs account for 46.0% and 35.2% of the total variance, respectively. Male and female pairs are intermingled in PC space, without significant clustering (one-tailed Wilcoxon signed-rank test comparing same- and opposite-sex distances in PC space, P = 0.68; n = 20 pairs). g, Behavioural motif usage by sex, quantified as the proportion of each motif in total motif instances across both leaders and followers. Overall, females and males show similar motif usage: sync (P = 0.25), follower join (P = 0.59), follower sharp (P = 0.88), follower track (*P = 0.025), leader join (P = 0.13), leader sharp (P = 1.00), and leader track (P = 0.19) (two-tailed Wilcoxon rank-sum tests, n = 11 female pairs, 9 male pairs). Bars, mean; points, individual pairs. h–k, Correlation between motif usage and following behaviour in both sexes. n = 50 sessions from 20 pairs, Pearson correlation. Each point represents one session; same conventions in l–s. h, Join usage in leaders (R = 0.83, P = 7.1 × 10−14); i, Join usage in followers (R = 0.30, P = 0.033); j, Sharp usage in followers (R = 0.27, P = 0.056); k, Track usage in followers (R = 0.13, P = 0.36). l–o, Correlation between motif usage and following behaviour in females (n = 32 sessions from 11 female pairs): l, Join usage in leaders (R = 0.83, P = 6.1 × 10−9); m, Join usage in followers (R = 0.11, P = 0.55); n, Sharp usage in followers (R = 0.45, P = 0.0090); o, Track usage in followers (R = 0.060, P = 0.76). p–s, Correlation between motif usage and following behaviour in males (n = 18 sessions from 9 male pairs): p, Join usage in leaders (R = 0.87, P = 2.1 × 10−6); q, Join usage in followers (R = 0.61, P = 0.0073); r, Sharp usage in followers (R = 0.010, P = 0.97); s, Track usage in followers (R = 0.30, P = 0.23).
Source data
Extended Data Fig. 8 Partner-dependent modulation of spatial decision-making in leaders and followers.
a, P(north) versus the follower’s heading at trial onset. Points, proportion of north choices in 45° heading bins; curves, logistic fits conditioned on partner position (east/south/west/north) or solo foraging. b, Same analysis for mice in the leader role; format as in a. c–d, Beta estimates from logistic regression showing how partner position affects the mouse of interest’s sensitivity to its own heading (c) and choice bias (d); positive values indicate bias toward north, negative toward east. c, When the mouse of interest is the follower, partner position significantly reduces its sensitivity (all ***P < 1.6 × 10−35; n = 8 pairs, 4,517–4,894 trials) and vice versa (east: **P = 0.0028; south, west, north: ***P < 6.7 × 10−5; n = 8 pairs, 3,536–3,624 trials). Bars, logistic-regression coefficient ± s.e. d, Leader in the east and south decreases follower bias, while in west and north increases it (all ***P < 1.5 × 10−20). Follower in the east, south or west significantly increases leader bias (all ***P < 1.1 × 10−5). e–i (females) and j–n (males), Analysis of a mouse’s choice on partner-initiated trials, conditioned on partner heading. e, Choice of the mouse of interest as leader (top) or follower (bottom). Spatial bins represent P(north) across the mouse’s initial positions within the foraging arena. Left and right columns are conditioned on the partner (white cartoon mouse) facing north or east, respectively. f, Top, Leader ΔP(north), calculated as P(north | partner facing north) - P(north | partner facing east). Bottom, follower ΔP(north). g, Both leaders and followers are more likely to choose north on partner-initiated trials when the partner faces north. Each point denotes a spatial bin from the left plots. Top, leader P(north), paired t-test, P = 0.015, n = 36 bins; north/east: 4,959/4,976 trials pooled across 28 pairs. Bottom, follower P(north), P = 3.4 × 10−9; north/east: 2,986/3,035 trials. h, Leader P(north) plotted against the difference in its distance to the east and north zones, conditioned on partner heading. Each point is the P(north) within a spatial bin; curves show linear fits. P(north) is higher when the partner faces west or north than east or south (repeated-measures ANOVA across four conditions, followed by Tukey-Kramer post hoc test, P = 8.7 × 10−4, 0.067, 8.2 × 10−7, 0.0072 for east vs. west, east vs. north, south vs. west, south vs. north, respectively; P > 0.19 for east vs. south and west vs. north). i, Follower P(north) plotted against the difference in its distance to the east and north zones, conditioned on partner heading. P(north) ranks highest to lowest when the partner faces west, north, east, and south (all pairwise comparisons P < 2.4 × 10−8). j,k, As in e,f, for males. l–n, As in g–i, for males. l, Leader P(north), paired t-test, P = 9.5 × 10−5 (north/east: 1,368/1,390 trials, 10 pairs); follower, P = 8.4 × 10−6 (1,275/1,401). m, Repeated-measures ANOVA + Tukey-Kramer post hoc, all P < 0.0019 (east/south vs. west/north), P > 0.32 (east vs. south, west vs. north). n, Follower ranks west > north > east > south (all pairwise P < 4.9 × 10−5). o–p (females) and q–r (males), Logistic regression analysis of choice behaviour based on the spatial variables from both animals. o, Logistic regression using Δθ and Δd from both animals to predict the choice of the mouse of interest as leader or follower. Cross-validated predictions outperform shuffled controls (bootstrap test, P < 1.0 × 10−4, 10,000 iterations). Bars, mean ± SD across cross-validation folds. p, Beta coefficients for the predictions. Each point represents an individual mouse, and lines connect animals from the same pair. Bars and error bars indicate mean ± SEM. All groups are significantly different from zero except follower’s distance weight (Δd) for leader choice (one-sample t-test, ***P < 2.1 × 10−6; n = 28 pairs, 486–4,258 trials per pair). Coefficients also differ significantly between leaders and followers (paired t-test; all P < 1.6 × 10−7). q,r, As in o,p, for males. Cross-validated > shuffled (bootstrap test, P < 1.0 × 10−4). All coefficients differ from zero except follower Δd for leader choice (one-sample t-test, **P < 0.0018; n = 10 pairs, 473–3,671 trials), and between roles (paired t-test, all P < 0.0010). s,t, Beta coefficients from the logistic regression model examining the impact of spatial variables on each animal’s choice. Regression is performed separately for each spatial configuration, without rotation of the arena. N & E, E & S, S & W, W & N; N, north; E, east; S, south; W, west. Format as in p. s, coefficients for leader choice. t, coefficients for follower choice. There is no significant effect of trial configuration on any coefficient, indicating that rotational alignment does not introduce systematic biases. Repeated-measures ANOVA for leader choice: leader heading, F(3,111) = 0.93, P = 0.43; leader distance, F(3,111) = 0.74, P = 0.53; follower heading, F(3,111) = 0.50, P = 0.68; follower distance, F(3,111) = 0.32, P = 0.81. For follower choice: leader heading, F(3,111) = 0.67, P = 0.57; leader distance, F(3,111) = 0.75, P = 0.53; follower heading, F(3,111) = 0.053, P = 0.98; follower distance, F(3,111) = 0.065, P = 0.98.
Source data
Extended Data Fig. 9 Emergence of coordinated behaviour and asymmetric social roles in simulation.
a, Cumulative reward over 1 × 106 timesteps during initiation training and testing, where both agents are rewarded whenever either enters the central initiation zone. In training (left), extensive early exploration transitions to exploitation, producing a convex-upward increase in cumulative reward. Testing (right) shows linear increase under exploitation only. Grey lines indicate individual pairs; red lines indicate the mean with a thin band of SEM. b, Similar trend in cooperation training, where agents are rewarded only when arriving at the same target zone. However, early mismatches drive cumulative reward below zero in training (left), resulting in a U-shaped recovery as coordination is learned over more than 5 × 107 timesteps. Grey lines indicate individual pairs; red lines indicate the mean with a thin band of SEM. c, Correct rate across cooperative training. Thin grey lines show each pair’s proportion of correct trials binned into 100 evenly spaced blocks, revealing a sigmoidal rise from near-zero to ceiling performance. d, Role-adoption asymmetry during cooperative training. Brown circles with connecting lines show each pair’s proportions of trials initiated and led by agent 1 versus agent 2 for pairs with significant bias (initiate 192/200, lead 175/200; n = 200 pairs; two-tailed binomial test, α = 0.01); grey lines without circles near the centre indicate pairs without significant bias. e, Proportion of trials initiated (x axis) and led (y axis) by individual agents during testing. Only agents with statistically significant bias in both initiation and leadership are shown and tallied by quadrant. Each point, one agent. The proportion of leaders who are also initiators does not differ significantly from chance (two-tailed binomial test against 0.5, P = 0.31).
Source data
Extended Data Fig. 10 Region-specific disruption of cooperative behaviour by mPFC but not OFC inactivation.
a, Histological verification of hM4D(Gi) expression in the mPFC at anterior-posterior (AP) coordinates +1.98, +1.18, +0.50, and −0.20 mm relative to bregma. Scale bar, 500 μm. b, mPFC population firing rates from example recording sessions following CNO (left, n = 49 units) or saline (right, n = 62 units) injection. Black line shows the mean; grey shading indicates the SEM. c, mPFC mean firing rates during the 0–5 min and 25–30 min time windows after CNO (left) or saline (right) injections. Neural activity significantly decreases after CNO (mean ± SEM; paired t-test, ***P = 8.3 × 10−6; n = 82 neurons from 2 animals), but not after saline injection (P = 0.30; n = 190 neurons from 2 animals). d, Histological validation of hM4D(Gi) expression for OFC inactivation at the anterior-posterior (AP) coordinate of +2.46 mm relative to bregma. e, Bilateral OFC inactivation in both leaders and followers did not affect cooperation rate (two-tailed Wilcoxon signed-rank test, P = 0.16, n = 6 pairs), leader asymmetry (P = 0.16), or initiator asymmetry (P = 0.44). Thin lines, individual data points; thick line, median; same conventions in f,g. f, mPFC inactivation in both leaders and followers increases the mismatch rate (two-tailed Wilcoxon signed-rank test, **P = 0.0029; n = 11 pairs) and decreases the omission rate (*P = 0.037), but does not significantly affect other behavioural measures, including unrewarded error rate (P = 0.19), number of trials completed per session (P = 0.72), and proportion of trials initiated by initiators (P = 0.32) or responders (P = 0.32). g, mPFC inactivation in both leaders and followers does not significantly affect mean inter-animal distance (P = 0.21), leader and follower mean speed (P = 0.10 and P = 0.76, respectively), leader and follower 95th percentile speed (P = 0.24 and #P = 0.083, respectively), or proportion of stationary time (leader, P = 0.21; follower, P = 0.70). h, Lack of correlation between cooperation rate and movement speed of leaders and followers, with both variables normalized to the control session the day before the inactivation session. Each point, one animal. No significant correlations are found for mean speed (leaders: Pearson correlation, R = 0.061, P = 0.86, n = 11 pairs; followers: R = 0.090, P = 0.79) or 95th percentile speed (leaders: R = 0.18, P = 0.59; followers: R = 0.25, P = 0.46). The diagram in d was created using BioRender; Wu, H. https://BioRender.com/hwyq5u6 (2026).
Source data
Extended Data Fig. 11 Chemogenetic validation and optogenetic mPFC inactivation during cooperation.
a, Schematic and histological validation of CaMKIIα-hM4D(Gi) expression in the mPFC at the same anterior-posterior (AP) coordinates as the hSyn-hM4D(Gi). b, Bilateral mPFC inactivation in both leaders and followers using the CaMKIIα promoter to preferentially target excitatory neurons decreases cooperation rate (two-tailed Wilcoxon signed-rank test, *P = 0.031, n = 6 pairs) and leader asymmetry (*P = 0.031), while increasing the mismatch rate (*P = 0.031). Inactivation does not significantly affect the unrewarded error rate (P = 0.81) or initiator asymmetry (P = 1.00). Thin lines, individual data points; thick line, median; same conventions in d–f and m–r. c,d, In mCherry-expressing controls, CNO administration in both leaders and followers does not affect cooperation rate (two-tailed Wilcoxon signed-rank test, P = 0.62, n = 4 pairs), leader-led cooperation rate (P = 0.12), follower-led cooperation rate (P = 1.00), leader asymmetry (P = 0.88), or initiator asymmetry (P = 0.38). e,f, Additional behavioural metrics from hSyn-hM4D(Gi)-driven mPFC inactivation (n = 11 pairs). e, mPFC inactivation in leaders does not significantly affect the omission rate (two-tailed Wilcoxon signed-rank test, P = 0.77), the proportion of trials initiated by the initiator (P = 0.76), or the proportion of trials initiated by the responder (P = 0.76). f, mPFC inactivation in followers does not significantly affect omission rate (#P = 0.067), the proportion of trials initiated by the initiator (P = 0.17), or the proportion of trials initiated by the responder (P = 0.17). g, Schematic of wireless optogenetic mPFC inactivation. h, Histological validation of KALI-1 expression and optic fibre implant positions in the mPFC at the AP coordinate of +1.98 mm relative to bregma. i, Spike raster of an example mPFC neuron expressing KALI-1 under wired 590 nm LED illumination across power intensities (0.3–3.2 mW). j, mPFC mean firing rate across light intensities from an example recording session (n = 66 units). k, Normalized mPFC mean firing rate under LED illumination across power intensities (mean ± SEM; n = 135 units from 2 animals). l, Normalized mPFC mean firing rate under wireless LED illumination (mean ± SEM; n = 44 units from 1 animal). m–p, Optogenetic inactivation in both roles throughout the trial impairs cooperation rate (m) but not leader (n) or initiator asymmetry (o) (n = 8 pairs) and decreases follower heading weight in leader decisions, and decreases leader heading, leader distance, and follower heading weights in follower decisions (p; n = 1,188 trials from 8 pairs). q, Cooperation rate with optogenetic inactivation for the first 2 s (left) or 1 s (right) of the trial (n = 7 pairs). r, Selective inactivation of leader (left) or follower (right) mPFC throughout the trial does not affect cooperation rate (n = 8 pairs). The diagrams in a, c, e and g were created using BioRender; Wu, H. https://BioRender.com/hwyq5u6 (2026).
Source data
Extended Data Fig. 12 Non-social tracking performance and mPFC coding of choice and arrival order.
a, Learning performance in the non-social stimulus tracking task (n = 8 mice; 4 females and 4 males). Points indicate mean ± SEM. There is no significant sex difference in correct rate (two-tailed Wilcoxon rank-sum test, P = 0.77). Correct rate is higher when the mouse follows than leads the non-social stimulus (two-tailed Wilcoxon signed-rank test, P = 0.0078; n = 8 mice). Correct rate is also higher when the non-social stimulus replays the leader than follower mouse trajectory (P = 0.0078; n = 8 mice). Thin lines, individual data points; thick line, median; same below. b, The proportion of trials in which the mouse follows the non-social stimulus is higher when the non-social stimulus replays the leader than follower mouse trajectory (P = 0.0078, two-tailed Wilcoxon signed-rank test; n = 8 mice). c, Reaction times are longer in the non-social stimulus tracking task than in the cooperative foraging trials from which mouse trajectories are drawn (two-tailed Wilcoxon signed-rank test, P = 0.0078; n = 8 mice). Reaction times do not differ between females and males (two-tailed Wilcoxon rank-sum test, P = 0.49; n = 4 females and 4 males) but are longer when the mouse follows than leads the non-social stimulus (P = 0.0078; n = 8 mice). d, Mean speed and 95th percentile speed are not affected by mPFC inactivation (two-tailed Wilcoxon signed-rank test; mean speed, P = 0.92; 95th percentile speed, P = 1.00; n = 10 mice). e, Histological validation of GCaMP6f expression and GRIN lens position (yellow rectangle) at the AP coordinate of +1.98 mm relative to bregma. f,g, Example choice-selective neurons recorded from a leader (f) and a follower (g). Left: trial-by-trial calcium activity (ΔF/F). Each row is a trial, grouped by the reward zone chosen by the mouse and sorted by trial duration. The white and yellow ticks indicate the onset and offset of the trials, respectively. Right: Curves and shading show mean ΔF/F ± SEM across trials. h, Choice selectivity indices for leaders (left) and followers (right). Shading, selective neurons. i, SVM decoding of choice. Mean ± SEM across animals (4 leaders, 4 followers; 4–11 sessions per animal); shuffled pools both groups. j,k, Example arrival order-selective neurons from a follower (j) and a leader (k). Left, trial-by-trial ΔF/F grouped by arrival order (L, recorded mouse leads; F, follows). Purple and green ticks, leader and follower arrival. Right, mean ΔF/F ± SEM. l, Left-port choice when leading versus following does not differ (n = 8 animals); dashed line, unity. m, Arrival speed when leading or following. Each line, one animal. n, Arrival trajectories when leading or following. Lines, individual arrival trajectories. Inset shows plotting area (red square) relative to the full arena. o, Proportion of arrival-order-selective neurons modulated by reward expectation does not differ from overall population (two-tailed Wilcoxon signed-rank test, n = 8 animals).
Source data
Extended Data Fig. 13 Spatial tuning of mPFC neurons across allocentric and egocentric reference frames.
a, Neural activity of the same example neuron shown in Fig. 4a–d, overlaid on movement trajectory across spatial reference frames: allocentric view of self and partner, and egocentric view of the partner (either unaligned or aligned with self heading). b,c, The same follower (b) and leader (c) neurons shown in Fig. 4f,g, displayed across all reference frames. Example neurons are representative of the spatially tuned mPFC populations quantified in Fig. 4 (n = 12 animals, 82 sessions; total 18,690 neurons recorded).
Extended Data Fig. 14 Decoding controls, task specificity, rule-reversal behaviour, and trial-by-trial neural activity.
a, Control decoding accuracy of partner angle using circularly time-shifted labels. Left plots: Confusion matrices show the probabilities of correct predictions for leaders and followers. Right plots: Decoders trained with true labels significantly outperform those trained with shifted labels (one-tailed Wilcoxon signed-rank test; leaders, *P = 0.016; followers, *P = 0.016; n = 6 animals per group). Each point represents one animal. b, Control decoding accuracy of egocentric partner distance using circularly time-shifted labels. Format as in a. Decoders trained with true labels significantly outperform the shifted labels (leaders, *P = 0.016; followers, *P = 0.016). c, Proportion of neurons selective for egocentric partner position is significantly lower in animals freely interacting outside the task context (two-tailed Wilcoxon rank-sum test, **P = 0.0023, n = 5 free-run animals, 12 task-engaged animals). Bars, mean; points, individual animals. d, Behavioural performance under rule reversal (reward on mismatch; mean ± SEM; n = 3 pairs, 10 sessions each). Dashed line indicates the 80% correct training criterion. e, Trial-by-trial response of an example neuron tuned to frontal partner position. Left, the social receptive field (SRF) for egocentric partner position, defined in egocentric coordinates as the region in which partner presence elicits greater than half-maximal activity; the white contour marks the SRF boundary, and the star indicates the centre. Right, trial-by-trial responses of the same neuron in the 2-s window preceding the arrival of the recorded mouse at the reward zone; each trace represents one trial, and the green segment indicates periods when the partner occupies the SRF. The diagram in e was created using BioRender; Wu, H. https://BioRender.com/hwyq5u6 (2026).
Source data
Extended Data Fig. 15 Neural encoding and task-dependent modulation of inferred value functions.
a, Inferred value functions in error trials for leaders (top row) and followers (bottom row); format as in Fig. 5d. b, Inferred value functions under the mismatch rule from leaders (top row) and followers (bottom row). Followers no longer assign higher value to proximal partner locations than leaders (paired t-test, P = 0.63; n = 3 pairs). c, Coefficient of determination (R2) from linear regression predicting the MAIRL-inferred value function using population neural activity: leader (top, permutation test, P = 0.39), follower (bottom, P = 1.1 × 10−8). Null distributions were generated by circularly shifting value labels 1,000 times; P-values reflect the probability of the observed R2 under a Gaussian fit to the null distribution. d, Inferred value functions under the match rule in control (top) and chemogenetic inactivation (bottom) sessions. In leaders, chemogenetic inactivation does not alter inferred values at the centre, east, or north locations (paired t-tests: centre, P = 0.16; east, P = 0.13; north, P = 0.37; n = 11 animals). In followers, inactivation significantly reduces inferred value at the centre and north locations (centre, P = 0.011; north, P = 0.0092), but not at the east location (P = 0.36; n = 11 animals). In both leaders and followers, and under both conditions, inferred values at the initiation and reward locations are significantly higher than at other locations (paired t-tests, all P ≤ 0.0068).
Source data
Supplementary information
Reporting Summary (download PDF )
Supplementary Table 1 (download XLSX )
Complete statistical reports for all figure panels.
Supplementary Table 2 (download XLSX )
Multi-agent inverse reinforcement learning (MAIRL) model formulation under different hypotheses.
Peer Review File (download PDF )
Supplementary Video 1 (download MP4 )
A pair of mice performing the cooperative foraging task. Shown are three consecutive trials in which the two well-trained mice coordinate their movements to arrive together at the same active reward zone, illustrating the stable leader and follower roles that emerge with training. Playback speed, 1×.
Supplementary Video 2 (download MP4 )
Sync motif. In the synchronized travel (sync) motif, the pair move in parallel, maintaining similar linear and angular velocities and a consistent close distance, and arrive at the same active reward zone together. Playback speed, 0.5×.
Supplementary Video 3 (download MP4 )
Track motif. One mouse, while approaching an active reward zone, slows and turns its head to monitor its more distant partner. On detecting the partner’s approach, it adjusts its timing so that both animals arrive at the reward ports nearly simultaneously. Playback speed, 0.5×.
Supplementary Video 4 (download MP4 )
Sharp turn motif. The two mice initially head toward different active reward zones. During the approach, one mouse abruptly redirects its trajectory by more than 90°, turning toward the zone selected by its partner and restoring coordinated movement. Playback speed, 0.5×.
Supplementary Video 5 (download MP4 )
Join motif. One mouse moves toward an active reward zone on its own. The partner, initially stationary or undecided, aligns its trajectory after observing this movement, and both animals travel together to the same active reward zone. Playback speed, 0.5×.
Source data
Rights and permissions
Open Access This article is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License, which permits any non-commercial use, sharing, distribution and reproduction in any medium or format, as long as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons licence, and indicate if you modified the licensed material. You do not have permission under this licence to share adapted material derived from this article or parts of it. The images or other third party material in this article are included in the article’s Creative Commons licence, unless indicated otherwise in a credit line to the material. If material is not included in the article’s Creative Commons licence and your intended use is not permitted by statutory regulation or exceeds the permitted use, you will need to obtain permission directly from the copyright holder. To view a copy of this licence, visit http://creativecommons.org/licenses/by-nc-nd/4.0/.
Reprints and permissions
About this article
Cite this article
Cheng, Y., Chen, Y., Kwak, M. et al. Asymmetric prefrontal representations for leader–follower dynamics. Nature (2026). https://doi.org/10.1038/s41586-026-10900-1
Download citation
Received:
Accepted:
Published:
Version of record:
DOI: https://doi.org/10.1038/s41586-026-10900-1