Main
Personality is defined by relatively stable patterns of thinking, feeling and behaving that vary across individuals1,2. Decades of research have demonstrated that human personality traits can be meaningfully and efficiently summarized by five relatively independent broad factors known as the Big Five: extraversion (which encompasses traits such as sociability, assertiveness and energy level), agreeableness (compassion, respectfulness and trust), conscientiousness (organization, productiveness and responsibility), neuroticism (anxiety, depression and volatility) and openness to experience (imagination, curiosity and aesthetic sensitivity)6,7. Given personality’s relevance to labour market outcomes8,9,10, physical health10,11,12, mental health13,14 and human demography15,16,17, understanding the molecular genetics of personality is important for developing biopsychosocial models of human behaviour and life outcomes.
Here we report meta-analytic GWAS results from the Revived Genomics of Personality Consortium (ReGPC), a collaboration across 46 cohorts incorporating between 611,037 and 1.14 million participants per Big Five trait. We also provide a frequently asked questions document that explains the rationale and results of this study in everyday language (https://osf.io/crh9d). This effort increases the number of significant loci relative to most recent work18 from 3 to 131 for conscientiousness, from 4 to 39 for agreeableness, from 8 to 126 for openness to experience, from 14 to 258 for extraversion and from 224 to 706 for neuroticism. The additive effects tagged by the common genetic variants (single-nucleotide polymorphism (SNP) heritability, h2SNP) account for an estimated 4.8–16.2% of total personality variation across traits and models. This is consistent with previous GWAS of personality3,19,20,21,22 and less than estimates of h2 from classic twin and pedigree studies and molecular genetic methods that index effects of rare variants and genetic interactions, which range from around 40% to 60%18,23,24,25. Critically, all methods find that personality is substantially influenced by a person’s environment.
Our analyses reveal that genetic associations with personality traits are generalizable and minimally confounded, with widespread external correlates. Although personality traits often vary on average across groupings of people26,27, we find that genetic associations with each trait are highly similar across four Western country clusters, reporter perspective (self versus close other), age and measurement instrument. Similarly, polygenic indices (PGIs) constructed from each primary GWAS predict personality consistently across five independent cohorts. Biological characterization implicates trait-differentiated molecular and cellular systems of greater complexity than previously theorized28,29. Within-family analyses indicate that, in contrast to GWAS of other behavioural traits4,5,30, genetic associations with personality traits are only minimally confounded by population stratification, dynastic effects and assortative mating. Finally, we characterize personality’s wide-ranging genetic correlates and consequences across socially relevant behaviours and life outcomes ranging from psychopathology and physical health to research participation and residential mobility, highlighting a central role of personality in understanding the human experience.
Population-level meta-analysis
We performed population-level GWAS meta-analysis of each Big Five trait in genetic similarity-stratified and genetic similarity-combined cohorts of participants with European-like (EUR-like) (kcohorts = 46) and African-like (AFR-like) (kcohorts = 10) genomes, as measured by genetic similarity to reference panels (Supplementary Tables 1–3 and Supplementary Notes 1 and 2). We jointly analysed results using inverse variance-weighted meta-analysis (IVW), supplemented with multi-ancestry meta-analysis (MAMA)31, resulting in pooled associations for around 10 million SNPs for each Big Five trait.
We identified 1,260 approximately independent (r2 < 0.10 within a 250 kb range) genome-wide-significant (GWS; P < 5 × 10−8) SNPs associated with the Big Five traits, an increase of 3 to 43 times over the most recent findings for each trait (Table 1 and Supplementary Tables 4–13). Of the 1,260 lead SNPs, 824 (65%) were found in novel 250 kb loci that had not previously been associated with variation in the respective trait. Manhattan plots of IVW results are displayed in Fig. 1a–e. Demonstrating the independence between the Big Five, 82% of genomic loci containing a GWS SNP were associated with only a single trait (Supplementary Table 14), and the average absolute genetic correlation among Big Five traits among participants with EUR-like genomes was 0.19 (Fig. 1f (below the diagonal)), similar to phenotypic correlations32 (Fig. 1f (above the diagonal)).
a–e, Manhattan plots of genetic associations for each Big Five personality trait: extraversion (a; n = 662,617), agreeableness (b; n = 611,037), conscientiousness (c; n = 641,167), neuroticism (d; n = 1,136,711) and openness to experience (e; n = 611,985). The points correspond to two-sided test P values of GWAS regressions. Genome-wide significance is denoted by a red line corrected for multiple comparisons at P = 5 × 10−8. A selection of genes containing or nearby the most significant lead SNPs are annotated in each panel (Supplementary Tables 30–34). f, Correlations among the Big Five personality traits. Genetic correlations estimated using LDSC among participants with EUR-like genomes are below the diagonal; meta-analytic phenotypic correlations from ref. 32 are above the diagonal. Standard errors are shown in parentheses. A, agreeableness; C, conscientiousness; E, extraversion; N, neuroticism; O, openness to experience.
Full size table
Genetic effects were highly congruent when splitting EUR-like GWAS cohorts into two halves, re-estimating the GWAS in each half and estimating genetic correlations (rg) between the two (mean rg = 0.94) (Supplementary Table 15). Individual SNP effects were extremely small, reflecting the highly distributed genetic architecture of personality. For example, the median association estimate among extraversion lead SNPs, corrected for the winner’s curse33, is 0.009 s.d. per effect allele (corresponding to scoring at the 50.35th versus the 50th percentile in extraversion). On average, the estimated effective number of independently associated common SNPs for each trait was 16,180 (Supplementary Table 17). We provide additional GWAS results, including genetic ancestry-stratified analyses and comparisons illustrating similarity between IVW and MAMA estimates, in Extended Data Fig. 1 and Supplementary Tables 16–20.
Characterizing SNP heritability
Among participants with EUR-like genomes, SNP heritability (h2SNP) estimated for the Big Five traits using linkage disequilibrium score regression (LDSC)34 ranged from 4.8% (s.e. = 0.2%) for agreeableness to 9.3% (s.e. = 0.3%) for extraversion (Table 1, Supplementary Table 4 and Supplementary Note 3). Importantly, these SNP heritability estimates from GWAS meta-analysis index genetic effects that are consistent across contributing cohorts. To allow for variability in genetic effects across cohorts, we conducted a random effects meta-analysis of cohort-specific h2SNP estimates, which indicated an average h2SNP of 8.6% (s.e. = 0.6%) across traits (ranging from 7.4% for agreeableness to 10.6% for extraversion; Table 1 and Supplementary Table 21), with significant variability across cohorts (mean τ = 3.6%). Random response error by the participants cannot systematically relate to their genome35,36. Accordingly, we found that personality measures with greater reliability (lower random response error) tended to be more heritable (b = 6.7%, s.e. = 0.7%; Extended Data Fig. 2). In this analysis, the expected h2SNP for a measure of typical (median) reliability (α = 0.81) ranged from 9.3% for agreeableness (s.e. = 0.6%) to 13.3% for extraversion (s.e. = 0.6%), and h2SNP completely disattenuated for measurement error ranged from 10.8% for agreeableness (s.e. = 0.9%) to 15.8% for extraversion (s.e. = 0.9%; Table 1).
To further characterize the generalizability of genetic associations with personality, we examined the concordance of genetic signal across geography, age, veteran status, measurement instrument and reporter perspective (Table 1 and Extended Data Fig. 3). Genetic effects were similar but not identical across four western country clusters (USA, continental Europe, Nordic and UK–Australia, mean rg = 0.86, mean s.e. = 0.15), three age groups (young (≤25 years), middle (25–64 years) and older (65 years and older), mean rg = 0.80, mean s.e. = 0.18), between the Million Veteran Program and other, primarily non-veteran cohorts (mean rg = 0.82, mean s.e. = 0.04), and across five personality measurement instruments (mean rg = 0.85, mean s.e. = 0.07). Additional characterization of genetic architecture across measurement instruments using genomic structural equation modelling37 confirmed that genetic effects plausibly operate at the level of broad cross-instrument latent factors, with only one locus showing significantly heterogenous effects across measurement instruments (Supplementary Tables 22–24 and Extended Data Fig. 4). Notably, genetic associations with agreeableness were less consistent across cohorts (Table 1), explaining in part why agreeableness exhibited lower heritability than other traits in the meta-analytic GWAS. In the Estonian Biobank, in which the personality of the participants was assessed both by their self-report (n = 73,983) and by reports by close others (n = 20,269), we found strong genetic overlap between rater perspectives (mean rg = 0.84, mean s.e. = 0.12), indicating that the genetic architecture of personality is not an epiphenomenon of self-perception. In sex-stratified analyses of neuroticism in the UK Biobank cohort, X-chromosome-linked h2SNP did not differ between male individuals (n = 168,989; h2SNP,X = 0.23%; s.e. = 0.04%) and female individuals (n = 198,139; h2SNP,X = 0.18%; s.e. = 0.03%; Pdifference = 0.33). The dosage compensation ratio (\(\hat{{\rm{\gamma }}}\) = 1.26, s.e. = 0.30) was intermediate between no compensation (0.5) and full compensation (2.0) but was estimated relatively imprecisely. Genetic effects were correlated near-unity across sex (rg = 0.96; 95% confidence interval (CI) = 0.81–1.10).
Biological follow-up of GWAS signals indicated that enriched gene sets intersected across the Big Five (mean enrichment rank-order ρ = 0.72; Extended Data Fig. 5), providing evidence for trait-overlapping molecular and cellular systems in personality neurobiology despite only modest genetic correlations (Fig. 1f). Consistent with theories of personality development that emphasize the prefrontal cortex38,39, genetic associations for each Big Five trait, except for agreeableness, were enriched in genes expressed in the prefrontal cortex (among these, top lead SNPs implicate RCE1, FOXP2 and SEMA6D, indicated in Fig. 1; Supplementary Tables 25–34). All traits demonstrated strong enrichment in protein-truncating variant-intolerant gene sets specifically expressed in neurons (such as ARNTL, TCF4 and NEGR1; Fig. 1). This pattern suggests personality-relevant variants are under negative selection, which appears inconsistent with evolutionary theories that posit balancing selection mechanisms as responsible for maintaining genetic variation in personality40,41. Neurobiological theories of the Big Five attribute dopaminergic aetiology to variation in openness to experience and extraversion, and serotonergic aetiology to variation in conscientiousness, neuroticism and agreeableness28,29. Based on sets of genes differentially expressed in specific brain cell types, both in mouse and human post mortem tissues, we found little support for these hypotheses across trait-stratified tests: neither the effect size nor P value placed serotonin or dopamine among the most enriched neuron types (Supplementary Tables 25–29), suggesting that their prominence in the literature relative to other types of neurons is not warranted.
Polygenic prediction of personality
We predicted personality traits in five independent cohorts using PGIs. PGI weights were constructed using our EUR-like discovery GWAS with SBayesR42. In all Big Five traits in all five cohorts, PGIs significantly predicted their respective phenotypic personality trait among participants with EUR-like genomes (meta-analytic β ranged from 0.097 (0.005) for agreeableness to 0.189 (0.005) for extraversion; Fig. 2a, Supplementary Tables 35 and 36 and Supplementary Note 4), representing an improvement in prediction over past research for extraversion, conscientiousness and openness to experience3,43. Highly consistent estimates across cohorts indicates a limited role of birth year, nationality and ascertainment method in prediction accuracy.
a, Standardized β point estimates for regression model prediction of phenotypic personality traits among participants with EUR-like and AFR-like genomes in five cohorts using PGI weights derived from the EUR-like population-level GWAS. Numerical results are provided in Supplementary Table 35. HRS, Health and Retirement Study. b, Polygenic prediction of external phenotypes among participants with EUR-like genomes in the Add Health cohort (n = 5,111). Point estimates for regression-model prediction of continuous traits are standardized β values depicted by circles, and estimates for binary traits are logistic β values depicted with diamonds. For binary traits, the percentage of affirmative responses is noted in parentheses. The dagger symbol (†) indicates questions administered to only a subset of participants (Supplementary Information). Numerical results are provided in Supplementary Table 39. c, Meta-analytic correlations between PGIs for each pair of family members, disattenuated for measurement error using structural equation models and scaled to represent the inferred genetic similarity among parents. The double dagger symbol (‡) indicates that educational attainment estimates come from ref. 47, estimated across spouses in the MoBa cohort. Numerical results are provided in Supplementary Table 40. d, Genetic correlations with accelerometer activity, which indexes physical movement across a 24 h period, in the UK Biobank. nEUR-like ≈ 95,000. See ref. 55 for complete details. e, Polygenic associations with census area of residence at age 50 years and older in the UK Biobank, controlling for birthplace in addition to sex, age, array and genetic principal components. nEUR-like = 419,261. Only associations that are significant at P < 0.01 after application of Benjamini–Hochberg false discovery rate correction and that are sign-concordant within sibship are plotted. In a and b, the error bars depict 95% CIs. *P < 0.01, determined using a two-sided regression test with no correction for multiple comparisons.
We also predicted personality traits among participants with AFR-like genomes in the Add Health and Health and Retirement Study cohorts, again using weights constructed from the EUR-like GWAS. PGI prediction is expected to substantially decrease when there are differences in genetic ancestry between discovery and target samples44. Nevertheless, PGIs significantly predicted extraversion (meta-analytic β = 0.099 (0.016)), neuroticism (β = 0.071 (0.017)) and openness to experience (β = 0.047 (0.017)) at P < 0.01; Fig. 2a and Supplementary Tables 35 and 36), providing the first evidence for significant PGI prediction of personality traits beyond individuals with EUR-like genomes.
Associations with behaviours and outcomes
We quantified the relevance of personality genetics to relevant behaviours and life outcomes across tests of genetic correlation, PGI prediction and Mendelian randomization (MR) applied to EUR-like GWAS data (Supplementary Note 5). Genetic correlation and PGI prediction results applied to AFR-like GWAS data are reported in Supplementary Note 5 and Extended Data Fig. 6.
Genetic correlations
Genetic correlations estimated using LDSC indicate widespread genetic sharing between personality traits and health-relevant daily behaviours. Conscientiousness in particular was genetically correlated with reduced substance use (across substances, mean rg = −0.21), greater sports frequency (rg = 0.35) and greater preference for low-calorie foods (rg = 0.17), as well as the likely downstream consequences of these behaviours: fewer spells in the hospital (rg = −0.21), healthier ageing (rg = 0.14) and lower body mass index (BMI; rg = −0.16) (Supplementary Tables 37 and 38). Personality was also genetically correlated with fluctuations in accelerometer-measured behaviour across the day, with openness to experience linked to increased night-time activity (rg ≈ 0.25) and conscientiousness to increased activity during the day (rg ≈ 0.30; Fig. 2d).
We found substantial genetic correlations between the Big Five and each of ten psychiatric disorders and four transdiagnostic psychiatric disorder factors (Fig. 3). Genetic profiles associated with neuroticism were associated with increased risk for all forms of psychopathology (mean cross-disorder rg = 0.43), particularly internalizing disorders (factor rg = 0.73), whereas genetic profiles associated with agreeableness displayed cross-cutting protective associations (mean rg = −0.20). We also identified broad patterns of positive genetic associations between openness and psychopathology (mean rg = 0.20), which are notably discrepant from the small and inconsistent associations commonly reported in phenotypic research13,14. Many other personality–psychopathology associations were specific and differentiated: extraversion genetics were associated with lower risk for internalizing disorders (factor rg = −0.08) but greater risk for neurodevelopmental disorders (factor rg = 0.18), and conscientiousness genetics were associated with lower risk for neurodevelopmental disorders (factor rg = −0.24) but greater risk for compulsive disorders (factor rg = 0.13). These findings provide strong support for close connections between personality traits and psychiatric disorders posited in psychiatric nosologies13,14.
Letters corresponding to personality traits represent point estimates of the genetic correlation. Error bars represent the 95% CIs. Associations are depicted in grey when not significant based on a two-sided test threshold of P < 0.01 with no correction for multiple testing. Associations were estimated using LDSC among participants with EUR-like genomes. Descriptions of each phenotype and numerical results are provided in Supplementary Tables 37 and 38. ADHD, attention deficit hyperactivity disorder; HDL, high-density lipoprotein; LDL, low-density lipoprotein.
Genetic correlations between personality traits and survey behaviour in the UK Biobank indicated links between personality and research participation (Fig. 3). In particular, genetic variants associated with openness to experience and agreeableness were also associated with completing optional questionnaires (mean rg = 0.22 and 0.11, respectively), whereas genetic variants associated with neuroticism were associated with fewer questionnaire responses (rg = −0.23) and more ‘I don’t know’ and ‘prefer no response’ answers (mean rg = 0.32) after beginning a questionnaire. The genetics of personality may therefore have a widespread, underacknowledged influence on sample composition and responses in published research that relies on survey data. In Supplementary Note 5.1, we further describe genetic correlation results from each of these nine outcome clusters.
PGI associations with outcomes
Genetic variants inherited at conception are expected to partially predict outcomes later in life through nondeterministic transactions with the environment45,46. We examined PGI prediction of interviewer–rater qualities and delinquency in Add Health—a representative sample of young US adults (n = 5,110). We found that personality PGIs predicted how participants were perceived by their in-person interviewer during data collection (Fig. 2b and Supplementary Table 39). Those with PGIs reflecting lower neuroticism and higher extraversion, agreeableness and openness to experience were more often perceived to have a more attractive personality, and those with PGIs reflecting higher agreeableness and conscientiousness were perceived to be more well groomed. In a subsample of at-risk Add Health participants, personality PGIs were associated with a greater likelihood of delinquent behaviour and intergenerational contact with the prison system (Fig. 2b and Supplementary Table 39): high polygenic propensity for neuroticism and low polygenic propensity for agreeableness had small-to-moderate associations with school suspension, running away from home and ever pulling a knife or gun on another person.
Our sociodemographic analysis in the UK Biobank (n = 419,261) linked personality PGIs to the characteristics of one’s residence in later adulthood, controlling for place of birth to reflect residential mobility. Genetic profiles indicative of higher openness to experience predicted greater likelihood of moving into cosmopolitan professional areas and away from suburbia and challenged neighbourhoods by middle–older adulthood (age 50 years), whereas genetic profiles indicative of higher conscientiousness predicted greater likelihood of moving away from cosmopolitan areas and into affluent suburban communities (Fig. 2e). We present additional residential mobility analyses focused on urban/rural migration in Extended Data Fig. 7. These analyses substantiate theoretical associations between personality and life outcomes that have yet to be submitted to GWAS in large samples. Importantly, PGI prediction of behaviours and outcomes from personality-relevant genetic variants was small in magnitude. These PGIs cannot predict an individual’s future behaviour.
We also quantified personality similarity between spouses over recent generations. Inferred genetic correlations between parents, estimated using error-corrected PGI similarity36 meta-analysed across spouses, sibling pairs and cousin pairs from UK Biobank (npairs = 187,703) and deCODE (ndyads = 133,358), indicated minimal assortative mating for each of the Big Five traits (estimated parental rg = 0.01–0.04; Fig. 2c and Supplementary Table 40). Over past generations, parents were only minimally more genetically similar in their personality traits, on average, than would be expected by chance, and far less genetically similar than in educational attainment (rg = 0.37)47,48. Phenotypic assortative mating estimates between deCODE parents (npairs = 5,317) were stronger in magnitude but remained small for all traits besides openness to experience (rs = 0.07–0.12; openness r = 0.18; Extended Data Fig. 8 and Supplementary Table 41), consistent with previous estimates48.
Mendelian randomization
To test causality and directionality in associations between personality traits and biobehavioural outcomes, we applied MR tests to a subset of 13 outcomes most likely to comport with core assumptions of this method (Supplementary Tables 42 and 43). We found plausible causal effects for 33 exposure–outcome pairings, where effects directionally replicated across MR estimators and multiple 95% CIs excluded 0 (Extended Data Fig. 9). Twenty-four analyses indicated effects of personality on biobehavioural outcomes, providing evidence that personality traits affect physical health—for example, extraversion increased COVID infection risk, conscientiousness decreased BMI and reduced likelihood of smoking initiation and frequency of spells in hospital, and neuroticism increased the likelihood of scoring lower on a healthy ageing/longevity factor. Furthermore, consistent with theoretical perspectives that life experiences affect personality development1,26, we found nine plausibly causal effects of biobehavioural outcomes on personality, encompassing broad multi-trait effects of educational attainment, increased BMI and smoking initiation on average levels of personality. Reassuringly, we found no evidence for MR associations between personality and three negative controls (birthweight, number of sisters and number of brothers) that would only be associated with one’s own personality through intergenerational confounding (residual gene–environment correlation). These tests provide valuable information on the potential causes and consequences of personality traits, which to date have been especially challenging to obtain using other methods49. Nevertheless, future work should further triangulate the MR-based causal inferences reported here with additional methods and data that permit strong causal inference50.
Evaluating familial confounding
Polygenic prediction within families
We used data from the Netherlands Twin Register (NTR) (nfamilies = 2,956) and Twins Early Development Study (TEDS) (nfamilies = 4,751) to predict personality traits within dizygotic twin pairs. These models predict personality from twin differences in PGIs constructed from the primary EUR-like population GWAS. By holding family environment and parent genotype constant, and leveraging the randomness of intergenerational genetic transmission, these within-family comparisons substantially reduce genetic confounding that may be present in population-level GWAS4,5,51,52 (Supplementary Note 6).
For all Big Five traits in both cohorts, within-twin pair PGI differences predicted the respective personality trait, with indistinguishable magnitude from population-level PGI prediction. The average within-family effect (β), meta-analysed across the NTR and TEDS cohorts and averaged across the Big Five, was β = 0.159, compared with β = 0.165 for the population-level effect (a prediction ratio of 96%; Fig. 4b and Supplementary Tables 35 and 36). This similarity suggests negligible familial confounding in population-level genetic associations with personality. This result contrasts starkly with the notable within-family attenuation found for other behavioural traits, such as educational attainment (mean within/population predictive ratio of around 60%)4,5,53, cognitive ability (around 80%)4,5 and externalizing psychopathology (about 75%)22.
a, Population-level and within-family PGI prediction in two cohorts: NTR and TEDS. b, The total PGI prediction for parent–offspring in the deCODE cohort (significance is indicated by asterisks below the corresponding bar), decomposed into directly transmitted effects (effects of the proband’s PGI; darker bars) and familial confounds (effects of the proband’s parents’ PGIs; lighter bars). The 95% CIs and significance asterisks for estimated familial confounding are shown (Supplementary Table 44). The ratio of direct effects to direct plus confounding effects is shown at the bottom of the corresponding bar. Comparison estimates for educational attainment (EA) come from a PGI constructed from the ref. 53 GWAS of educational attainment with the deCODE and 23&Me cohorts held out. In a and b, point estimates are standardized β values from regressions predicting phenotypic personality traits using PGI weights derived from the EUR-like population-level GWAS. c, h2SNP estimates were determined using with LDSC from the within-family GWAS (wit), presented alongside full EUR-like population-level GWAS (full pop; Table 1) and matched population-level (pop) GWAS in cohorts that contributed to within-family GWAS (Supplementary Table 46). The rg value below each pair of traits indicates their genetic correlation estimated using LDSC, with deviation from rg = 1 tested using a nested χ2 model in genomic structural equation modelling. The P values above each pair of traits indicate the results of a test for differences in h2SNP between population and within-family estimates. For a–c, the error bars depict the 95% CIs and standard errors are shown in parentheses. P values were calculated using two-sided tests with no correction for multiple testing. NS, not significant (P ≥ 0.05); *P < 0.05, ***P < 0.001.
We further tested for familial confounding using a complementary method that is expected to be less biased by potential effects that siblings may have on one another52. We used parent–offspring data from deCODE (n = 34,506) to decompose population-level PGI prediction into direct effects (genetic variants passed down from parent to offspring) and familial confounds (parental genetic variants not transmitted to the offspring, population stratification and assortative mating). Across the Big Five, direct effects accounted for an average of 96.2% of variance in population-level PGI prediction (Fig. 4 and Supplementary Table 44), indicating that genetic associations with personality operated nearly exclusively through direct transmission from parent to offspring. The identicality of these results to those of within-twin PGI analyses indicates negligible confounding from effects that siblings have on one another. Only for conscientiousness, non-transmitted genetic confounds were significantly predictive of offspring personality (explaining 8.1% of the total predicted effect, indirect effect B = 0.015, 95% CI = 0.003–0.027, P = 0.01). Non-transmitted confounds did not differ in magnitude across mothers and fathers (Supplementary Table 45). By contrast, for educational attainment, only 75.1% of variance was explained by direct effects (indirect effect B = 0.069, 95% CI = 0.055–0.083, P = 3.3 × 10−21; Supplementary Table 44). After controlling for familial confounds, PGI prediction of extraversion by the extraversion PGI (B = 0.207) was similar in magnitude to PGI prediction of educational attainment by the educational attainment PGI53 (B = 0.208). Overall, these tests rule out large or moderate effects of familial confounding in genetic associations with personality; they indicate that such effects are negligible to small in magnitude.
Within-family genome-wide analyses
As a test of familial confounding at the level of individual SNPs5,52, we conducted within-family GWAS for each Big Five trait. We assembled data across 12 contributing EUR-like cohorts (n = 31,544–50,725 across traits) that ascertained genetic data among family members, using best-practice quality control5. For neuroticism, we combined these data with published within-family GWAS estimates (Supplementary Table 46 and Extended Data Fig. 10).
LDSC-estimated h2SNP values from within-family GWAS ranged from 7.6% (s.e. = 1.6%) for agreeableness to 13.4% (3.2%) for openness to experience, evincing comparable magnitudes to those reported for the complete meta-analyses of population GWAS (Fig. 4 and Supplementary Table 47). To maximize comparability between within-family and population-level GWAS, we also estimated h2SNP using meta-analysis of population-level GWAS among the matched subset of 12 cohorts who contributed to the within-family GWAS (Fig. 4). For each trait, LDSC within-family intercepts were lower than population-level intercepts, suggesting that these models accounted for residual confounding in population-level models.
h2SNP comparisons across models, using P < 0.05 as a significance threshold to maximize identification of potential differences, indicated that, for only openness to experience, matched population-level h2SNP estimates were significantly greater than within-family h2SNP estimates (19.4%, 95% CI = 15.8–22.9% versus 13.4%, 95% CI = 7.1–19.6%, P = 0.02). This observed difference was driven largely by an inflated population-level h2SNP estimate for openness to experience in the Estonian Biobank sample (Supplementary Table 46). By contrast, previously reported comparisons for other social and behavioural traits, such as educational attainment, depression and household income, indicate that LDSC h2SNP is attenuated by up to 50% in within-family data5. LDSC-estimated genetic correlations between within-family and population-level GWAS summary data averaged rg = 0.86 and were significantly less than 1.00 only for neuroticism (rg = 0.75, 95% CI = 0.61–0.90, P < 0.001) and agreeableness (rg = 0.79, 95% CI = 0.61–0.97, P = 0.03). These results indicate strong correspondence in genetic effect estimates between population-level and within-family analyses, but we note generally wider CIs, indicating a need for additional within-family GWAS data collection. Further comparisons across non-overlapping cohorts indicated that population-level and within-family effect sizes for significant and suggestively associated SNPs (P < 1 × 10−5) were highly similar in magnitude, demonstrating that uncontrolled confounding in population-level GWAS is small in magnitude and that genetic associations with personality are replicable (Supplementary Note 6.7 and Supplementary Table 47).
Genetic correlations between these within-family Big Five summary statistics and the 85 population-level external GWAS phenotypes were similar but not identical to fully population-level genetic correlations (mean correlation among rg estimates = 0.83; Supplementary Table 48 and Extended Data Fig. 10). Fully within-family genetic correlations between these summary data and 18 phenotypes from ref. 4 were also highly similar to population-level genetic correlations, despite notably less power (mean r between rg estimates = 0.87; Supplementary Table 49 and Extended Data Fig. 10).
Discussion
In a consortium effort assembling data from 46 cohorts covering 611,037 to 1.14 million participants, we conducted highly powered GWAS of each of the Big Five personality traits, establishing a rigorous basis for inference in personality genomics and highlighting the fundamental role of personality in the human experience.
Genetic and environmental factors, along with their correlations and interactions, contribute to variation in personality23,45,46. We estimated h2SNP at 4.8–9.3% using population-level data and 7.6–13.4% using within-family data. As expected from psychometric theory, heritability estimates were higher when personality was measured using more reliable instruments, revealing a predictable trade-off between reliability (longer tests) and larger samples that should factor into future data collection. Notably, our h2SNP estimates, which index only effects of common genetic variants, are markedly lower than the 40–60% heritability estimates produced based on twin and family methods18,23 and molecular genetic methods that more completely incorporate effects of rare variants and genetic interactions24,25.
Comparisons of population-level and within-family associations indicated that genetic associations with personality are relatively unconfounded by assortative mating or by environmental factors that are shared across family members, including uncontrolled population stratification and dynastic effects51. This result contrasts with those obtained for other social science traits, such as educational attainment, income, externalizing behaviour and cognitive performance4,5,22,51,53, for which genetic associations are substantially attenuated in within-family analyses compared to population-level analyses. However, confounding was not entirely absent, and even small confounding effects can be meaningful. For agreeableness and neuroticism, within-family and population GWAS effect estimates were genetically correlated at less than unity. For openness to experience, h2SNP from within-family GWAS was significantly lower than from matched-sample population GWAS (but not full-population GWAS), and for conscientiousness, non-transmitted parental alleles had a small but significant contribution to PGI prediction. Finally, highly powered analyses of assortative mating identified effects that were significant but very small in magnitude. Collection of additional within-family data will provide further leverage to accurately and precisely identify loci indexing direct genetic effects. Importantly, even direct genetic effects do not determine an individual’s personality. Rather, genetic associations act probabilistically across large samples of people54, and may interact with and operate through environmental experiences45.
Genetic associations with personality largely generalized across different regions of Western geography, age groups, self-rated versus other-rated report, military service versus general population samples, and measurement instrument. This suggests that widespread genetic inference is possible, even though the mechanisms linking genetic variation to personality trait differences are numerous, dynamic and complex. One key departure from this overall pattern pertains to agreeableness, which exhibited lower genetic correlations across geography and rater instrument, suggesting that genetic effects relating to this trait may be more varied across groups and/or measurements, and inference must be more contextualized. In the future, analyses stratified by sex, more finely by age, and by more expansive sets of ancestral and cultural backgrounds, will further refine inferences about the differentiation of personality genetics across environments.
Functional genomic analyses and biological annotation indicated that, despite their only modest genetic intercorrelations, genetic associations with each of the Big Five personality traits occur through overlapping molecular and cellular systems. Single-cell gene expression data integrated with our GWAS indicated that there is no straightforward reductive mapping from broad, multifaceted personality traits to specific neuron types or neurotransmitters. These analyses did not fully align with pre-existing theories that specifically link serotonergic and dopaminergic neuronal processes to personality28,29. However, the availability of these high-powered GWAS combined with ever-expanding access to temporal and spatial brain gene expression data will enable future developmental and system-specific analysis of personality aetiology.
Much of the scientific value of studying personality stems from its applicability to nearly all aspects of human life6,7,8,9,10,11,12,13,14,15,16,17. We extend this tenet by linking heritable variation in personality traits to core mental and physical health outcomes, labour market performance and reproduction, as well as specific behaviours such as survey response, lifespan residential mobility and impressions made on others. Causal inference in these associations is especially valuable, given that personality traits are highly stable over time26 and resistant to simple experimental manipulation49. MR results implicate personality traits as both potential causes and consequences of a range of health conditions. Given personality’s pervasive relevance, researchers across scientific fields will benefit from incorporating personality traits when predicting and explaining biopsychosocial outcomes.
Methods
Cohorts and measures
We incorporated personality trait data from 46 cohorts participating in the ReGPC. Cohorts were included that provided data on genotyped participants with EUR-like and AFR-like genomes who had completed a validated multi-item personality inventory compatible with the Big Five personality trait taxonomy. Descriptive statistics for each cohort are provided in Supplementary Tables 1 and 2. Descriptions of each cohort are provided in Supplementary Note 7. A PRISMA flowchart for this meta-analysis is available at OSF (https://osf.io/sh8gr). These analyses were not pre-registered. As this meta-analysis used de-identified summary-level data, it was deemed exempt from ethics board approval by the University of Texas at Austin Institutional Review Board (STUDY00001941); ethics approval for individual cohorts that collected data and contributed to the meta-analysis is presented in Supplementary Note 7. All of the participants consented to the research, and this research was performed in accordance with all relevant guidelines and regulations.
Population-level meta-analysis
For each trait, in each cohort, we conducted population-level genome-wide association analyses among participants with EUR-like and AFR-like genomes56 across the 22 chromosomes and the X chromosome, following a standard operating procedure (https://osf.io/rsc49). In each cohort, we excluded SNPs with in-sample minor allele frequency (MAF) < 1%, call rate < 95%, deviation from Hardy–Weinberg Equilibrium (P < 1 × 10−5), poor imputation quality (INFO < 0.40) and those with poor clustering on visual inspection of intensity plots. Participants were excluded if they had low overall call rates (<95%), excess autosomal heterozygosity or homozygosity, were duplicated samples, had gender inconsistent with sex (which is often indicative of mislabelled demographic data57) or had chromosomal abnormalities. Association analyses were conducted in each cohort by regressing personality trait score on each biallelic 1000 Genomes 3v5 EUR or AFR SNP58 with PLINK2 (ref. 59), PLINK60, GCTA61, BOLT-LLM62, FastGWA63 or proprietary software, using mixed models to account for relatedness when appropriate. These analyses included controls for age, age2, birth cohort, sex and genetic principal components, to address potential lifespan and birth cohort differences in personality traits1,26. Some groups contributed separate results for potentially overlapping cohorts of participants. In these cases, we meta-analysed estimates while accounting for dependency of their estimation errors. We harmonized and directionally aligned effects using EasyQC64 and checked that effects were directionally consistent across cohorts by estimating genetic correlations with LDSC34 (conducted on HapMap3 SNPs with an INFO threshold of ≥0.90) in GenomicSEM37.
For each Big Five trait, we then conducted IVW meta-analysis of GWAS estimates across all of the contributing cohorts, separately for participants with EUR-like and AFR-like genomes, in R (v.4.3.1)65. For each meta-analysed SNP, we re-estimated n using observed meta-analytic MAF and standard error. This produced 10 sets of summary statistics (five traits × two ancestral groups). We quantified the number of genome-wide significant lead SNPs at P < 5 × 10−8 that were linkage disequilibrium (LD)-independent of other lead SNPs nearby on the genome (within a 250 kb window) using FUMA66. We also identified lead SNPs that were LD-independent of all other lead SNPs on the chromosome regardless of genomic distance. Ancestry-stratified Manhattan plots of meta-analytic genome-wide associations are presented in Extended Data Fig. 1.
We next conducted transancestry IVW meta-analyses to synthesize data across the k = 10 cohorts that included participants with both EUR-like and AFR-like genomes, and the additional k = 36 cohorts including solely participants with AFR-like genomes. As around 95% of participants were of EUR-like ancestry, and because stronger linkage disequilibrium among EUR-like compared with AFR-like superpopulations67 means that using the EUR reference panel probably leads to conservative estimation of the number of lead SNPs, we initially identified candidate lead SNPs using the 1000 Genomes 3v5 EUR reference panel. Using information on SNP LD from the 1000 Genomes 3v5 EUR and AFR reference panels, we estimated transancestry sample size-weighted LD matrices for each chromosome across all candidate transancestry lead SNPs and all lead SNPs from the ancestry-stratified analyses, which we used to identify additional LD-independent lead SNPs. To supplement transancestry IVW results, we also conducted transancestry meta-analyses using MAMA31.
We tested the consistency of the genetic signal across cohorts by splitting the 46 EUR-like discovery cohorts into two balanced halves, conducting IVW meta-analyses separately in each half and using LDSC to test the concordance of genetic signal across split halves. We identified novel lead SNPs in the current GWAS relative to past GWAS as those that were outside the same 250 kb locus as or LD-independent of a significant SNP in past GWAS of the Big Five. Finally, we quantified similarity in genetic signal among the Big Five by identifying overlapping 250 kb genetic loci across pairs of traits and by estimating genetic correlations between pairs of Big Five traits using LDSC.
Characterizing SNP heritability
We estimated h2SNP for each personality trait using LDSC. As LDSC requires ancestrally homogenous data, we applied this method separately to GWAS summary statistics of participants with EUR-like genomes, AFR-like genomes and, for comparison purposes, EUR-like genomes in the subset of 10 cohorts that ascertained participants with both EUR-like and AFR-like genomes. We estimated polygenicity using stratified LD fourth moments regression68, and we estimated within-trait transancestry genetic correlations using POPCORN69.
We additionally applied LDSC to GWAS data on each individual cohort and then meta-analysed h2SNP estimates across cohorts using fixed-effects and random-effects meta-analyses, with the R package metafor70. To quantify the effect of measurement error of the personality measure on h2SNP, we conducted meta-regressions, where h2SNP in a cohort was regressed on the Cronbach’s α reliability of its personality measure (reliabilities are shown in Supplementary Table 1). For each trait, we also used LDSC to correlate genetic effects across each pair of cohorts with positive h2SNP estimates (1,781 total pairings) and used fixed-effects meta-analysis to obtain an average estimate for this cross-cohort genetic correlation.
To quantify the consistency of genetic signal in each of the Big Five across groups, we sorted EUR-like cohorts by five grouping variables with substantial variation and n > 10,000 per group: geographical region (US, UK–Australia, continental Europe and Nordic), age group (most participants <25 years, 25–64 years and 65 years and older), rater perspective (self-report and other-report, with data for these comparisons contributed exclusively from the Estonian Biobank cohort), military veteran status (MVP cohort participants and participants from all 45 other cohorts) and measurement instrument family (Big Five Inventory71,72,73,74, NEO Personality Inventory75,76,77,78,79, International Personality Item Pool inventories80,81, 100 Nuances of Personality Inventory82 and Eysenck Personality Inventory83,84). We next estimated genetic correlations across groups using bivariate LDSC. Finally, we compared X-chromosomal associations with neuroticism across male and female individuals in the UK Biobank cohort, estimating h2SNP, cross-sex rg and the dosage compensation ratio85,86 on sex-stratified summary data of unrelated male (n = 168,856) and female (n = 197,927) individuals with EUR-like genomes.
We applied GenomicSEM to formally model the genetic factor structure of the different measurement instruments and to compute the QSNP statistic, which quantifies the extent to which SNP associations with each instrument varied beyond what would be expected under the factor model87. We specified a five-factor model in which each instrument served as an indicator of its respective Big Five trait and cross-loadings were added based on an iterative data-driven approach. For each SNP, we computed the QSNP statistic by comparing a model in which each factor was regressed on the individual SNP to one which all specific measurement instruments were regressed on the individual SNP.
To identify annotations enriched for h2SNP, we submitted each of the five EUR-like GWAS summary statistics to MAGMA88, including SNPs with INFO ≥ 0.80. The gene sets included in this analysis came from previous work89. We then indexed the similarity of enrichment across the Big Five using pairwise rank correlations of enrichment effect sizes from these analyses.
Polygenic prediction of personality
We generated PGI weights for each Big Five trait by applying SBayesR42 to EUR-like GWAS estimates for the set of 1.3 million HapMap3 SNPs. PGI weights for each of the Big Five were then used to construct PGIs in five cohorts: the National Longitudinal Study of Adolescent to Adult Health (Add Health), German Socioeconomic Panel (GSOEP), Health and Retirement Study (HRS), Netherlands Twin Register (NTR) and Twins Early Development Study (TEDS). Each cohort ascertained participants with EUR-like genomes, and Add Health and HRS also ascertained participants with AFR-like genomes. As HRS, NTR and TEDS also contributed to population-level GWAS, we held data from each respective cohort out when estimating PGI weights in that cohort, to prevent overlap in training and prediction sets.
We used each of these PGIs to predict the corresponding Big Five trait in each cohort. We estimated population-level effects using ordinary least squares regression for cohorts of unrelated participants (Add Health, GSOEP and HRS) and linear regression with cluster-robust standard errors across families to account for relatedness in cohorts of twins (NTR and TEDS). These regressions controlled for age, age2, sex, 10 ancestral principal components and genotyping batch effects, with significance set at P < 0.01. We also generated PGI weights both among a broader set of 2.9 million SNPs and, using PRS-CSx90, we compared this prediction to prior estimates. As we describe further in Supplementary Note 4, these changes did not improve prediction, so we focus on analyses using SBayesR and the set of 1.3 million HapMap SNPs.
Associations with behaviours and outcomes
We quantified genetic correlations between the Big Five and 85 external phenotypes that had previously been submitted to GWAS, among participants with EUR-like genomes, using LDSC. For 13 of these phenotypes, GWAS data were available to also estimate these genetic correlations among participants with AFR-like genomes. We set the significance threshold at P < 0.01 for these analyses. Outcomes and their source GWAS are presented in Supplementary Table 37.
We also used PGIs to estimate associations between the Big Five and 19 phenotypic behaviours and outcomes that have not been submitted to GWAS and could therefore not be analysed with LDSC. To do this, we adapted the method described in the ‘Polygenic prediction of personality’ section above to predict each behaviour or outcome from each of the Big Five PGIs across participants. In residential mobility PGI analyses, we included statistical controls for place of birth (so that effects reflect residential movement), median home prices (to control for affordability), age, sex and five principal components. We further filtered results to only those that directionally replicated within family to reduce potential confounding (see the ‘Evaluating familial confounding’ section below) and applied the false-discovery rate correction91.
We quantified assortative mating in personality over recent generations using a method that disattenuates PGIs for measurement error, enabling the estimation of genetic correlations across family members. Specifically, we estimated two independent sets of PGI weights for each participant, one from each split-half GWAS (see the ‘Population-level meta-analysis’ section above). We then estimated a structural equation model in the R package lavaan92 that used the two PGIs as indicators of a latent, measurement error-free PGI for each participant36. We correlated these PGIs across pairs of spouses, siblings and cousins in the UK Biobank and deCODE cohorts (a path diagram of this latent correlation model is provided in Extended Data Fig. 8b). PGI correlations between members of spousal dyads in this model directly estimate genetic assortative mating. For siblings and cousin dyads, we transformed PGI correlations to be on the same scale as spouses. We then used IVW meta-analysis to synthesize these estimates across relationship types and the two cohorts.
We conducted MR analyses to examine putatively causal bidirectional effects between personality traits and external phenotypes, using the R package TwoSampleMR93. For these tests, we selected 13 external behaviours and outcomes from the broader pool of 85 on the basis of previous personality theory (Supplementary Note 5.5). We also selected three phenotypes to serve as negative controls (birthweight, number of sisters and number of brothers), as these cannot be directly caused by an individual’s personality. For each phenotype, we conducted MR tests using three estimators chosen for their complementarity and robustness to common violations of MR assumptions94: weighted mode with Steiger filtering95 weighted median with Steiger filtering96 and MR-CAUSE97. To identify replicable effects in the context of multiple tests, we considered an effect to be significant only if estimates were directionally consistent across all three MR methods with unadjusted 95% confidence or credibility intervals that excluded 0 across at least two methods.
Evaluating familial confounding
To evaluate familial confounding in population-level PGIs, we compared population-level PGI prediction to within-family PGI prediction: the prediction of phenotypic variation in personality traits from variation in genotype, controlling for the genotypes of one’s family members. For both the NTR and TEDS cohorts, we constructed PGI weights from the primary EUR-like population GWAS meta-analysis, excluding the respective twin cohort used for PGI prediction. We then estimated within-family PGI associations among dizygotic twin pairs, controlling for family level fixed effects, as well as sex, age and age2. We pooled estimates across the two cohorts and compared them to pooled population-level PGI estimates to obtain an index of residual confounding in population-level PGIs.
We then estimated within-family PGI prediction in the deCODE cohort, in which parent–offspring units were each genotyped. In this cohort, we predict an offspring’s personality trait from their PGI, controlling for their mother’s PGI, their father’s PGI, sex, age, age2 and their county of birth. For each trait, we obtained estimates of the population-level effect (the total prediction) and the indirect genetic confounding (the contribution to prediction made by parental PGI scores, both together and separately by mother and father), and we estimated the direct genetic effect by subtracting the indirect genetic confounding from the population-level effect. To provide a point of comparison, we repeated this within-family PGI prediction using the phenotype of educational attainment. Educational attainment PGI weights were estimated in SBayesR according to the procedure used for the Big Five, using GWAS summary data from a previous study53, excluding the 23&Me and deCODE cohorts.
We also conducted within-family GWAS analyses on EUR-like participant cohorts (k = 12) that ascertained genomic and personality data among family members. In each cohort, analysts estimated population-level and within-family effects for each SNP on each Big Five trait using either snipar30 or a pipeline developed previously4. For these analyses, we conducted stringent quality control, including only SNPs with imputation INFO ≥ 0.99 and correcting downwardly biased h2SNP estimates in small samples (Supplementary Note 6.3), and we pooled effects across cohorts with IVW meta-analysis. For neuroticism, we incorporated additional previously published summary data4. For each SNP, sample size was re-estimated using the meta-analytic standard error and 1000 Genomes 3v5 MAF, capped at the 75th percentile of the summed observed sample size across contributing cohorts. This method produced within-family summary statistics for each of the Big Five.
To enable direct comparison, we re-estimated population-level GWAS across these 12 within-family cohorts using the same method as in the ‘Population-level meta-analysis’ section above, restricted to SNPs with INFO ≥ 0.99 and with sample size estimated using 1000 Genomes 3v5 panel MAF rather than meta-analytic sample MAF. We then submitted these within-family GWAS summary statistics and matched population GWAS summary statistics to LDSC.
To confirm that within-family analyses reduced confounding, we compared population-level LDSC intercepts to within-family LDSC intercepts34. We used nested model comparison tests in genomic structural equation modelling to test the equivalence of h2SNP across matched population-level and within-family GWAS, and to test whether genetic correlations between population-level and within-family GWAS were less than unity98. We also compared the average SNP effect sizes between population-level and within-family GWAS, among SNPs with relatively strong statistical signal (P < 1 × 10−5) that were >250 kb apart on the genome.
Using LDSC, we estimated genetic correlations between the Big Five within-family GWAS summary data and population-level GWAS summary data from the 85 external behaviours or outcomes. We compared these correlations to the corresponding estimates fully based on population-level GWAS. We also estimated genetic correlations entirely based on within-family GWAS (for which summary data were available for 18 external behaviours or outcomes)4; these estimates were compared with the corresponding estimates based completely on population-level GWAS.
Reporting summary
Further information on research design is available in the Nature Portfolio Reporting Summary linked to this article.
Data availability
Summary statistics for GWAS of each Big Five trait, excluding 23andMe participants, are available at OSF99 (https://doi.org/10.17605/OSF.IO/HGNSM). This includes EUR-like, AFR-like and trans-ancestry meta-analytic population-level summary statistics, polygenic index weights and EUR-like within-family summary statistics, neuroticism summary statistics both including and excluding UK Biobank participants, and split-half and country-, age-, perspective- and questionnaire-stratified summary statistics. Other data used in this study cannot be made public and are available, within the constraints of the relevant regulations and data use agreements, on request.
Code availability
Code for all analyses performed this study is available at OSF99 (https://doi.org/10.17605/OSF.IO/HGNSM).
References
Roberts, B. W. & Yoon, H. J. Personality psychology. Ann. Rev. Psychol. 73, 489–516 (2022).
Article Google Scholar
McCrae, R. R. & John, O. P. An introduction to the five-factor model and its applications. J. Pers. 60, 175–215 (1992).
Article CAS PubMed Google Scholar
Gupta, P. et al. A genome-wide investigation into the underlying genetic architecture of personality traits and overlap with psychopathology. Nat. Hum. Behav. 8, 2235–2249 (2024).
Article PubMed PubMed Central Google Scholar
Howe, L. J. et al. Within-sibship genome-wide association analyses decrease bias in estimates of direct genetic effects. Nat. Genet. 54, 581–592 (2022).
Article CAS PubMed PubMed Central Google Scholar
Tan, T. et al. Family-GWAS reveals effects of environment and mating on genetic associations. Preprint at medRxiv https://doi.org/10.1101/2024.10.01.24314703 (2024).
John, O. P., Naumann, L. P. & Soto, C. J. in Handbook of Personality: Theory and Research 3rd edn (eds John, O. P. et al.) 114–158 (Guilford Press, 2008).
Matthews, G., Deary, I. J. & Whiteman, M. C. Personality Traits (Cambridge Univ. Press, 2009).
Anni, K., Vainik, U. & Mõttus, R. Personality profiles of 263 occupations. J. Appl. Psychol. 110, 481–511 (2024).
Borghans, L., Golsteyn, B. H., Heckman, J. J. & Humphries, J. E. What grades and achievement tests measure. Proc. Natl Acad. Sci. USA 113, 13354–13359 (2016).
Article ADS CAS PubMed PubMed Central Google Scholar
Roberts, B. W., Kuncel, N. R., Shiner, R., Caspi, A. & Goldberg, L. R. The power of personality: the comparative validity of personality traits, socioeconomic status, and cognitive ability for predicting important life outcomes. Perspect. Psychol. Sci. 2, 313–345 (2007).
Article PubMed PubMed Central Google Scholar
Hampson, S. E., Goldberg, L. R., Vogt, T. M. & Dubanoski, J. P. Mechanisms by which childhood personality traits influence adult health status: educational attainment and healthy behaviors. Health Psychol. 26, 121–125 (2007).
Article PubMed PubMed Central Google Scholar
Moffitt, T. E. et al. A gradient of childhood self-control predicts health, wealth, and public safety. Proc. Natl Acad. Sci. USA 108, 2693–2698 (2011).
Article ADS CAS PubMed PubMed Central Google Scholar
Kotov, R., Gamez, W., Schmidt, F. & Watson, D. Linking “big” personality traits to anxiety, depressive, and substance use disorders: a meta-analysis. Psychol. Bull. 136, 768–821 (2010).
Article PubMed Google Scholar
Widiger, T. A. et al. Personality in a hierarchical model of psychopathology. Clin. Psychol. Sci. 7, 77–92 (2019).
Article Google Scholar
Antonoplis, S. & John, O. P. Who has different-race friends, and does it depend on context? Openness (to other), but not agreeableness, predicts lower racial homophily in friendship networks. J. Pers. Soc. Psychol. 122, 894–919 (2022).
Article PubMed Google Scholar
Jokela, M. Selective residential mobility and social influence in the emergence of neighborhood personality differences: Longitudinal data from Australia. J. Res. Pers. 86, 103953 (2020).
Article Google Scholar
Duckitt, J. & Sibley, C. G. Personality, ideology, prejudice, and politics: a dual-process motivational model. J. Pers. 78, 1861–1894 (2010).
Article PubMed Google Scholar
Vukasović, T. & Bratko, D. Heritability of personality: a meta-analysis of behavior genetic studies. Psychol. Bull. 141, 769 (2015).
Article PubMed Google Scholar
Sanchez-Roige, S. et al. CADM2 is implicated in impulsive personality and numerous other traits by genome-and phenome-wide association studies in humans and mice. Transl. Psychiatry 13, 167 (2023).
Article CAS PubMed PubMed Central Google Scholar
Lo, M. T. et al. Genome-wide analyses for personality traits identify six genomic loci and show correlations with psychiatric disorders. Nat. Genet. 49, 152–156 (2017).
Article CAS PubMed Google Scholar
Nagel, M. et al. Meta-analysis of genome-wide association studies for neuroticism in 449,484 individuals identifies novel genetic loci and pathways. Nat. Genet. 50, 920–927 (2018).
Article CAS PubMed Google Scholar
Karlsson Linnér, R. et al. Multivariate analysis of 1.5 million people identifies genetic associations with traits related to self-regulation and addiction. Nat. Neurosci. 24, 1367–1376 (2021).
Article PubMed PubMed Central Google Scholar
Keller, M. C., Coventry, W. L., Heath, A. C. & Martin, N. G. Widespread evidence for non-additive genetic variation in Cloninger’s and Eysenck’s personality dimensions using a twin plus sibling design. Behav. Genet. 35, 707–721 (2005).
Article PubMed Google Scholar
Markel, G. et al. Nature, nurture, and socioeconomic outcomes: new evidence from sib pairs and molecular genetic data. Preprint at SSRN https://doi.org/10.2139/ssrn.5225447 (2025).
Yengo, L. et al. Within-family heritability estimates for behavioural and disease phenotypes from 500,000 sibling pairs of diverse ancestries. Preprint at medRxiv https://doi.org/10.1101/2025.09.17.25336022 (2025).
Bleidorn, W. et al. Personality stability and change: a meta-analysis of longitudinal studies. Psychol. Bull. 148, 588–619 (2022).
PubMed Google Scholar
Ebert, T. et al. Are regional differences in psychological characteristics and their correlates robust? Applying spatial-analysis techniques to examine regional variation in personality. Perspect. Psychol. Sci. 17, 407–441 (2022).
Article PubMed Google Scholar
DeYoung, C. G., Grazioplene, R. G. & Allen, T. A. in Handbook of Personality: Theory and Research 4th edn (eds John, O. P. & Robins, R. W.) 193–216 (Guilford Press, 2022).
Depue, R. A. & Collins, P. F. Neurobiology of the structure of personality: dopamine, facilitation of incentive motivation, and extraversion. Behav. Brain Sci. 22, 491–517 (1999).
Article CAS PubMed Google Scholar
Young, A. I. et al. Mendelian imputation of parental genotypes improves estimates of direct genetic effects. Nat. Genet. 54, 897–905 (2022).
Article CAS PubMed PubMed Central Google Scholar
Turley, P. et al. Multi-ancestry meta-analysis yields novel genetic discoveries and ancestry-specific associations. Preprint at bioRxiv https://doi.org/10.1101/2021.04.23.441003 (2021).
Park, H. H. et al. Meta-analytic five-factor model personality intercorrelations: eeny, meeny, miney, moe, how, which, why, and where to go. J. Appl. Psychol. 105, 1490 (2020).
Article PubMed Google Scholar
Forde, A., Hemani, G. & Ferguson, J. Review and further developments in statistical corrections for Winner’s Curse in genetic association studies. PLoS Genet. 19, e1010546 (2023).
Article CAS PubMed PubMed Central Google Scholar
Bulik-Sullivan, B. K. et al. LD score regression distinguishes confounding from polygenicity in genome-wide association studies. Nat. Genet. 47, 291–295 (2015).
Article CAS PubMed PubMed Central Google Scholar
Spearman, C. Correlation calculated from faulty data. Br. J. Psychol. 3, 271 (1910).
Google Scholar
Tucker-Drob, E. M. Measurement error correction of genome-wide polygenic scores in prediction samples. Preprint at bioRxiv https://doi.org/10.1101/165472 (2017).
Grotzinger, A. D. et al. Genomic structural equation modelling provides insights into the multivariate genetic architecture of complex traits. Nat. Hum. Behav. 3, 513–525 (2019).
Article PubMed PubMed Central Google Scholar
Casey, B. J. & Caudle, K. The teenage brain: self control. Curr. Dir. Psychol. Sci. 22, 82–87 (2013).
Article PubMed PubMed Central Google Scholar
Romer, D., Reyna, V. F. & Satterthwaite, T. D. Beyond stereotypes of adolescent risk taking: placing the adolescent brain in developmental context. Dev. Cogn. Neurosci. 27, 19–34 (2017).
Article PubMed PubMed Central Google Scholar
Penke, L. & Jokela, M. The evolutionary genetics of personality revisited. Curr. Opin. Psychol. 7, 104–109 (2016).
Article Google Scholar
Zietsch, B. P. Genomic findings and their implications for the evolutionary social sciences. Evol. Hum. Behav. 45, 106596 (2024).
Article Google Scholar
Lloyd-Jones, L. R. et al. Improved polygenic prediction by Bayesian multiple regression on summary statistics. Nat. Commun. 10, 5086 (2019).
Article ADS PubMed PubMed Central Google Scholar
Alemu, R. et al. An updated polygenic index repository: expanded phenotypes, new cohorts, and improved causal inference. Preprint at bioRxiv https://doi.org/10.1101/2025.05.14.653986 (2025).
Wang, Y. et al. Polygenic prediction across populations is influenced by ancestry, genetic architecture, and methodology. Cell Genom. 3, 100408 (2023).
Article CAS PubMed PubMed Central Google Scholar
Scarr, S. & McCartney, K. How people make their own environments: a theory of genotype → environment effects. Child Dev. 54, 424–435 (1983).
CAS PubMed Google Scholar
Tucker-Drob, E. M. How do individual experiences aggregate to shape personality development? Eur. J. Pers. 31, 570–571 (2017).
Torvik, F. A. et al. Modeling assortative mating and genetic similarities between partners, siblings, and in-laws. Nat. Commun. 13, 1108 (2022).
Article ADS CAS PubMed PubMed Central Google Scholar
Horwitz, T. B., Balbona, J. V., Paulich, K. N. & Keller, M. C. Evidence of correlations between human partners based on systematic reviews and meta-analyses of 22 traits and UK Biobank analysis of 133 traits. Nat. Hum. Behav. 7, 1568–1583 (2023).
Article PubMed PubMed Central Google Scholar
Grosz, M. P., Rohrer, J. M. & Thoemmes, F. The taboo against explicit causal inference in nonexperimental psychology. Perspect. Psychol. Sci. 15, 1243–1255 (2020).
Article PubMed PubMed Central Google Scholar
Bailey, D. H. et al. Causal inference on human behaviour. Nat. Hum. Behav. 8, 1448–1459 (2024).
Article PubMed Google Scholar
Nivard, M. G. et al. More than nature and nurture, indirect genetic effects on children’s academic achievement are consequences of dynastic social processes. Nat. Hum. Behav. 8, 771–778 (2024).
Article PubMed PubMed Central Google Scholar
Veller, C. & Coop, G. M. Interpreting population-and family-based genome-wide association studies in the presence of confounding. PLoS Biol. 22, e3002511 (2024).
Article CAS PubMed PubMed Central Google Scholar
Okbay, A. et al. Polygenic prediction of educational attainment within and between families from genome-wide association analyses in 3 million individuals. Nat. Genet. 54, 437–449 (2022).
Article CAS PubMed PubMed Central Google Scholar
Madole, J. W. & Harden, K. P. Building causal knowledge in behavior genetics. Behav. Brain Sci. 46, e182 (2023).
Article Google Scholar
Grotzinger, A. D. et al. Genetic architecture of 11 major psychiatric disorders at biobehavioral, functional genomic and molecular genetic levels of analysis. Nat. Genet. 54, 548–559 https://doi.org/10.1038/s41588-022-01057-4 (2022).
National Academies of Sciences, Engineering, and Medicine. Using Population Descriptors in Genetics and Genomics Research: A New Framework for an Evolving Field https://doi.org/10.17226/26902 (National Academies, 2023).
Turner, S. et al. Quality control procedures for genome-wide association studies. Curr. Protoc. Hum. Genet. 68, 1.19.1–1.19.18 (2011).
Google Scholar
The 1000 Genomes Project Consortium. A global reference for human genetic variation. Nature 526, 68–74 (2015).
Article Google Scholar
Chang, C. C. et al. Second-generation PLINK: rising to the challenge of larger and richer datasets. Gigascience 4, s13742-015 (2015).
Article Google Scholar
Purcell, S. et al. PLINK: a tool set for whole-genome association and population-based linkage analyses. Am. J. Hum. Genet. 81, 559–575 (2007).
Article CAS PubMed PubMed Central Google Scholar
Yang, J., Lee, S. H., Goddard, M. E. & Visscher, P. M. GCTA: a tool for genome-wide complex trait analysis. Am. J. Hum. Genet. 88, 76–82 (2011).
Article CAS PubMed Google Scholar
Loh, P. R. et al. Efficient Bayesian mixed-model analysis increases association power in large cohorts. Nat. Genet. 47, 284–290 (2015).
Article CAS PubMed PubMed Central Google Scholar
Jiang, L. et al. A resource-efficient tool for mixed model association analysis of large-scale data. Nat. Genet. 51, 1749–1755 (2019).
Article CAS PubMed Google Scholar
Winkler, T. W. et al. Quality control and conduct of genome-wide association meta-analyses. Nat. Protoc. 9, 1192–1212 (2014).
Article PubMed PubMed Central Google Scholar
R Core Team. R: a language and environment for statistical computing (R Foundation for Statistical Computing, 2023).
Watanabe, K., Taskesen, E., Van Bochoven, A. & Posthuma, D. Functional mapping and annotation of genetic associations with FUMA. Nat. Commun. 8, 1826 (2017).
Article ADS PubMed PubMed Central Google Scholar
Reich, D. E. et al. Linkage disequilibrium in the human genome. Nature 411, 199–204 (2001).
Article ADS CAS PubMed Google Scholar
O’Connor, L. J. et al. Extreme polygenicity of complex traits is explained by negative selection. Am. J. Hum. Genet. 105, 456–476 (2019).
Article PubMed PubMed Central Google Scholar
Brown, B. C., Ye, C. J., Price, A. L. & Zaitlen, N. Transethnic genetic-correlation estimates from summary statistics. Am. J. Hum. Genet. 99, 76–88 (2016).
Article CAS PubMed PubMed Central Google Scholar
Viechtbauer, W. Conducting meta-analyses in R with the metafor package. J. Stat. Softw. 36, 1–48 (2010).
Article Google Scholar
John, O. P., Donahue, E. M. & Kentle, R. The Big Five Inventory (Univ. of California, 1991).
Zakrisson, I. Big Five Inventory (BFI): Utprövning för Svenska Förhållanden (Mid Sweden Univ., 2010).
Mullins-Sweatt, S. N., Jamerson, J. E., Samuel, D. B., Olson, D. R. & Widiger, T. A. Psychometric properties of an abbreviated instrument of the five-factor model. Assessment 13, 119–137 (2006).
Article PubMed Google Scholar
Rammstedt, B., Kemper, C. J., Klein, M. C., Beierlein, C. & Kovaleva, A. A short scale for assessing the big five dimensions of personality: 10 item big five inventory (BFI-10). Methods Data Anal. 7, 17 (2013).
Google Scholar
Costa, P. T. & McCrae, R. R. Professional Manual for the NEO-PI-R and NEO-FFI (Psychological Assessment Resources, 1992).
Bjornsdottir, G. et al. Psychometric properties of the Icelandic NEO-FFI in a general population sample compared to a sample recruited for a study on the genetics of addiction. Pers. Indiv. Differ. 58, 71–75 (2014).
Article Google Scholar
Ostendorf, F. & Angleitner, A. NEO-Persönlichkeitsinventar nach Costa und McCrae, Revidierte Fassung (NEO-PI-R) (Hogrefe, 2004).
Terracciano, A. The Italian version of the NEO PI-R: conceptual and empirical support for the use of targeted rotation. Pers. Indiv. Differ. 35, 1859–1872 (2003).
Article Google Scholar
Hoekstra, H. A., Ormel, J. & De Fruyt, F. Persoonlijkheidsvragenlijsten NEO-PI-R & NEO-FFI (Swets & Zeitlinger, 1996).
Goldberg, L. R. The development of markers for the Big-Five factor structure. Psychol. Assess. 4, 26–42 (1992).
Article Google Scholar
Goldberg, L. R. in Personality Psychology in Europe Vol. 7 (eds Mervielde, I. et al.) 7–28 (Tilburg Univ. Press, 1999).
Henry, S. & Mõttus, R. The 100 nuances of personality: development of a comprehensive, non-redundant personality item pool. OSF https://doi.org/10.17605/OSF.IO/TCFGZ (2023).
Eysenck, H. J. & Eysenck, S. B. G. The Eysenck Personality Inventory (Hodder & Stoughton, 1964).
Eysenck H. J. & Eysenck, S. B. G. Manual of the Eysenck Personality Questionnaire: (EPQ-R Adult) (Educational and Industrial Testing Service, 1994).
Sidorenko, J. et al. The effect of X-linked dosage compensation on complex trait variation. Nat. Commun. 10, 3009 (2019).
Article ADS PubMed PubMed Central Google Scholar
Lee, J. J. et al. Gene discovery and polygenic prediction from a genome-wide association study of educational attainment in 1.1 million individuals. Nat. Genet. 50, 1112–1121 (2018).
Article CAS PubMed PubMed Central Google Scholar
Clapp Sullivan, M. L. et al. Beyond the factor indeterminacy problem using genome-wide association data. Nat. Hum. Behav. 8, 205–218 (2024).
Article PubMed PubMed Central Google Scholar
De Leeuw, C. A., Mooij, J. M., Heskes, T. & Posthuma, D. MAGMA: generalized gene-set analysis of GWAS data. PLoS Comput. Biol. 11, e1004219 (2015).
Article PubMed PubMed Central Google Scholar
Akingbuwa, W. A., Hammerschlag, A. R., Bartels, M., Nivard, M. G. & Middeldorp, C. M. Ultra-rare and common genetic variant analysis converge to implicate negative selection and neuronal processes in the aetiology of schizophrenia. Mol. Psychiatry 27, 3699–3707 (2022).
Article CAS PubMed PubMed Central Google Scholar
Ruan, Y. et al. Improving polygenic prediction in ancestrally diverse populations. Nat. Genet. 54, 573–580 (2022).
Article CAS PubMed PubMed Central Google Scholar
Benjamini, Y. & Hochberg, Y. Controlling the false discovery rate: a practical and powerful approach to multiple testing. J. R. Stat. Soc. B 57, 289–300 (1995).
Article MathSciNet Google Scholar
Rosseel, Y. lavaan: an R package for structural equation modeling. J. Stat. Softw. 48, 1–36 (2012).
Article Google Scholar
Hemani, G. et al. The MR-Base platform supports systematic causal inference across the human phenome. eLife 7, e34408 (2018).
Article PubMed PubMed Central Google Scholar
Sanderson, E. et al. Mendelian randomization. Nat. Rev. Methods Primers 2, 6 (2022).
Article CAS PubMed PubMed Central Google Scholar
Hartwig, F. P., Davey Smith, G. & Bowden, J. Robust inference in summary data Mendelian randomization via the zero modal pleiotropy assumption. Int. J. Epidemiol. 46, 1985–1998 (2017).
Article PubMed PubMed Central Google Scholar
Bowden, J., Davey Smith, G., Haycock, P. C. & Burgess, S. Consistent estimation in Mendelian randomization with some invalid instruments using a weighted median estimator. Genet. Epidemiol. 40, 304–314 (2016).
Article PubMed PubMed Central Google Scholar
Morrison, J., Knoblauch, N., Marcus, J. H., Stephens, M. & He, X. Mendelian randomization accounting for correlated and uncorrelated pleiotropic effects using genome-wide summary statistics. Nat. Genet. 52, 740–747 (2020).
Ennis, G. et al. Genomic taxometric analysis of negative emotionality and major depressive disorder highlights a gradient of genetic differentiation across the severity spectrum. Preprint at medRxiv https://doi.org/10.1101/2025.01.30.25321336 (2025).
Schwaba, T., Tucker-Drob, E., Nivard, M. G., Clapp Sullivan, M. L. & Abdellaoui, A. Data and code for ‘Robust inference and correlates from genetic associations with personality’. OSF https://doi.org/10.17605/OSF.IO/HGNSM (2026).
Download references
Acknowledgements
Acknowledgements for individual cohorts are provided in Supplementary Note 7.
Funding
T.S., W.A.A., M.G.N., M.L.C.S., C.M.W. and E.M.T.-D. were supported by NIMH R01MH120219. E.M.T.-D., M.L.C.S., C.M.W., Y.D. and J.d.l.F. were supported by NIA R01AG073593. W.A.A. is also supported by a Dutch Ministry of Education, Culture and Sciences (OCW) Talent and Early Development Grant. M.G.N., E.C.C., G.H. and G.D.S. are members of the MRC Integrative Epidemiology Unit at the University of Bristol, which is supported by the Medical Research Council and the University of Bristol (MC_UU_00032/1 and MC_UU_00032/7). J.D.T. was supported by NIMH K01MH141330. W.D.H. and C.X. were supported by a Career Development Award from the Medical Research Council (MRC) (MR/T030852/1) for the project titled “From genetic sequence to phenotypic consequence: genetic and environmental links between cognitive ability, socioeconomic position, and health”. T.K. was funded by the European Research Council (ERC Consolidator Grant awarded under the Horizon Europe framework, grant agreement 101087395. F.S. was supported by the Hector foundation II and by a 2023 NARSAD Young Investigator Grant (31537) from the Brain & Behavior Research Foundation with support from the Families for Borderline Personality Disorder Research. R.C. was supported by the Jacobs Foundation grant no. 2023-1510-00 and the Research Council of Norway (grant no. 325245). H.A. was supported by the research council of Norway (274611, 324620). F.A.T. was supported by Research Council of Norway (262700). E.C.C. was supported by the research council of Norway (RCN; 274611) and the South-Eastern Norway Regional Health Authority (HSØ; 2021045). J.L. received financial support from the Strategic Research Council (SRC) established within the Academy of Finland (decision number 352700). K.I. and U. Vainik have been funded by Estonian Research Council’s personal research funding start-up grants PSG656 and PSG759. R.M. has been funded by Estonian Research Council’s team grant PRG2190. U. Võsa and T. Esko have been funded by Estonian Research Council’s team grant PRG1291. A.A. is supported by the Amsterdam UMC Fellowship. T.T. and L.F. were supported in part by the National Institute on Aging Intramural Research Program. A.G. was supported by the European Research Council (GEPSI 946647). Acknowledgements for cohort funding are included in the Supplementary Information.
Ethics declarations
Competing interests
The authors declare no competing interests.
Peer review
Peer review information
Nature thanks Matthew Keller, Xuanyu Lyu, Aysu Okbay and Brent Roberts for their contribution to the peer review of this work. Peer reviewer reports are available.
Additional information
Publisher’s note Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
Extended data figures and tables
Extended Data Fig. 1 Visualization of ancestry-specific meta-analyses.
Panels A-E depict Manhattan plots among meta-analysed cohorts of participants with EUR-like genomes for the traits of Extraversion (N = 620,193), agreeableness (N = 567,556), conscientiousness (N = 598,047), neuroticism (N = 1,088,604), and openness to experience (N = 568,489). See also Summary Table S4. Points correspond to two-sided test p-values of GWAS regressions. Genome wide significance is denoted with a red line corrected for multiple comparisons at p = 5 × 10−8. p-values smaller than this threshold are considered significant. The blue line indicates p = 5 × 10−5. Panels F-J depict Manhattan plots among meta-analysed cohorts of participants with AFR-like genomes for the same traits: Extraversion (N = 43,201), agreeableness (N = 43,482), conscientiousness (N = 43,120), neuroticism (N = 48,110), and openness to experience (N = 43,498). See also Summary Table S4. Panels K-O depict similarity between (on the X axis) transancestry inverse-variance weighted meta-analytic GWAS Z-statistics and (on the Y axis) multi-ancestry meta-analysis GWAS Z-statistics for participants with EUR-like genomes boosted by data from participants with AFR-like genomes. Each dot represents a single SNP association with a Big Five trait, and the red line depicts an intercept of 0 and a slope of 1, visualizing the null hypothesis that effects are equivalent in magnitude between methods. See also Summary Table S17.
Extended Data Fig. 2 Meta-regressions of SNP heritability against scale reliability across cohorts.
Note: The regression line depicts the association between scale reliability (Cronbach’s Alpha) and h2SNP, weighted by inverse standard error of h2SNP estimate. Circles are scaled by inverse standard errors, such that larger circles represent more precise estimates. On the x-axis, Alpha values were re-scaled by subtracting 1 from all alpha values, such that an alpha of 0.0 represents perfect scale reliability. In Panel A, the regression is depicted for a multilevel model in which h2SNP is estimated for each trait, nested within each cohort. In panels B-F, the regression is depicted separately for each Big Five trait. See also Summary Tables S1 and S21.
Extended Data Fig. 3 LDSC Genetic correlations stratified by groups: geographic regions, age groups, veteran status, rater perspective, and measurement instrument.
Panel A depicts associations between and among four geographic regions: US = United States. UKAUS = United Kingdom and Australia. ConEU = Continental Europe. Heritability of each trait in each group is presented on the diagonal. Red squares denote genetic correlations among measures of the same trait across groups. Panel B depicts correlations between and among young, medium, and old age groups. Young = Mean cohort age ±1.5 SD ≤ 25. Mid = Mean cohort age ±1.5 SD = 26–64. Older = Mean cohort age ±1.5 SD ≥ 65. Heritability of each trait in each group is presented on the diagonal. Red squares denote genetic correlations among measures of the same trait across groups. Panel C depicts correlations between and among the Million Veterans Program sample, which is composed of US military veterans, and the other 45 ReGPC cohorts, which have few military veterans. Vet = Million Veterans Program. Nonvet = The Remaining 45 ReGPC cohorts. Heritability of each trait in each group is presented on the diagonal. Panel D depicts genetic correlations between and among self-rated and other-rated personality trait measurement in the Estonian Biobank. Heritability of each trait in each group is presented on the diagonal. Panel E depicts genetic correlations between and among personality measurement instruments. BFI = Big Five Inventory. NEO = NEO Inventory. IPIP = International Personality Item Pool. 100NP = 100 Nuances of Personality. EPQ = Eysenck Personality Questionnaire. Heritability of each trait measured by each questionnaire is presented on the diagonal. Red squares denote genetic correlations among measures of the same trait across questionnaires.
Extended Data Fig. 4 Multivariate analysis of genetic signal across measurement instruments.
Panel A depicts a confirmatory five-factor model of measurement-instrument stratified GWAS of the Big Five, conforming to simple structure. Standardized estimates are displayed. BFI = Big Five Inventory. IPIP = International Personality Item Pool. NEO = NEO Inventory. EPQ = Eysenck Personality Questionnaire. 100NP = 100 Nuances of Personality. Panel B depicts the residual genetic correlation matrix for this confirmatory simple structure model. Red squares denote residual genetic correlations among measures of the same trait. To account for violations of simple structure, Panel C depicts the final factor model of measurement-instrument stratified GWAS of the Big Five, allowing for cross-loadings. Panel D depicts the residual genetic correlation matrix from this final factor model. Panels E-I depict Manhattan plots for each latent Big Five personality trait depicted in the Panel C measurement model and panel J depicts a Manhattan plot of QSNP from measurement-instrument stratified GWAS of the model in Panel C. In these panels, points correspond to two-sided test p-values of GWAS regressions. Genome wide significance is denoted with a red line, corrected for multiple comparisons at p = 5 × 10−8. P-values smaller than this threshold are considered significant. Note that the Y-axis scale for neuroticism is scaled differently than the other panels. See also Summary Table S23. Panels K-O plot SNP associations with each individual Big Five measure as a function of that measure’s loading for a representative genome-wide significant hit with low QSNP for each trait. Each point depicts the relation between a SNP’s association with an instrument (Beta) and the instrument’s loading on the factor model in Panel C, with bars corresponding to the standard error of measurement. These SNPs are annotated in Panels E-I. The comparatively steep slope of the line corresponding to factor for which the SNP is a hit illustrates the specificity of the SNP effect on the factor, and the tight scatter of the points around their expectations is illustrative of the plausibility of assumption that the genetic variant operates on the individual measures via the factors (i.e. low QSNP). Both betas and factor loadings are standardized with respect to the genetic variance of the corresponding measure. Points are coloured according to the trait on which the loading corresponds to. Solid points are primary loadings. Unfilled points are cross-loadings (from a measure of another trait). The superimposed lines have slopes equal to the estimated SNP effect on the respective trait and represent the model-based expectations. N for each SNP are presented in Summary Table S24. Panel P plots SNP associations with different extraversion and neuroticism measures as a function of factor loading for a genome-wide-significant QSNP that is in high LD with Extraversion and Neuroticism hits (annotated in panel J). Each point depicts the relation between a SNP’s association with an instrument (Beta) and the instrument’s loading on the factor model in Panel C, with bars corresponding to the standard error of measurement. The diffuse scatter of the points around their expectations is illustrative of the implausibility of assumption that the genetic variant operates on the individual measures via the factors (i.e. high QSNP). Solid points are primary loadings. Unfilled points are cross-loadings (from a measure of another trait). The superimposed lines have slopes equal to the estimated SNP effect on the respective trait and represent the model-based expectations. N for each SNP are presented in Summary Table S24. Panels Q and R depict corresponding plots for extraversion and neuroticism hits (annotated in panels E and H, respectively). Each point depicts the relation between a SNP’s association with an instrument (Beta) and the instrument’s loading on the factor model in Panel C, with bars corresponding to the standard error of measurement. Both betas and loadings are standardized with respect to the genetic variance of the corresponding measure. Points are coloured according to the trait on which the loading corresponds to. Solid points are primary loadings. Unfilled points are cross-loadings (from a measure of another trait). The superimposed lines have slopes equal to the estimated SNP effect on the respective trait and represent the model-based expectations.
Extended Data Fig. 5 Gene-set heritability enrichment analyses.
Panels A-E depict enrichment for each of the Big Five traits. Each point estimate depicts the standardized enrichment beta; bars correspond to standard errors. PI = Protein-Truncating Intolerant genes. An X indicates the intersection of two gene sets. Number of genes in each set are in parentheses. Associations with a black star are significant after false-discovery-rate correction91. Panel F depicts pairwise rank correlations of enrichment betas across the Big Five. PI = Protein-truncating Intolerant genes. An X indicates the intersection of two gene sets. Larger points have larger weights, scaled by the inverse product of the standard errors of the gene set association betas. Gene sets with significantly differing enrichment across pairs of traits are labelled. See also Supplementary Tables S25–S29.
Extended Data Fig. 6 Genetic correlations and PGI associations among participants with AFR-like genomes.
Panel A depicts Genetic correlations between personality and health, psychopathology, and substance use, estimated using LDSC. Letters correspond to point estimates of genetic correlation,. Note. E = Extraversion. A = Agreeableness. C = Conscientiousness. N = Neuroticism. O = Openness to Experience. Error bars depict 95% confidence intervals. Associations not significant with two-sided test p < 0.05, with no further correction for multiple testing, are depicted in grey. Full numeric results are provided in Supplementary Table S38. Panel B depicts correlations among the Big Five in participants with AFR-like genomes. Note: Ext = Extraversion. Agr = Agreeableness. Con = Conscientiousness. Neu = Neuroticism. Ope = Openness to experience. Numbers in parentheses depict standard errors. Genetic correlations estimated using LDSC. Panel C depicts prediction of outcomes in the Add Health dataset using polygenic index weights derived from the EUR-like GWAS applied to participants with AFR-like genomes (N = 1,753). Note: Point estimates for effect sizes of continuous outcomes (labelled with a parenthesized C) are in standardized units, whereas associations with binary outcomes (labelled with a parenthesized B, with percentage affirmative answers) are depicted in units of logistic beta. Error bars depict 95% confidence intervals. No associations are significant at two-sided test p < 0.05, with no further correction for multiple testing. Full numeric results are provided in Supplementary Table S39.
Extended Data Fig. 7 Polygenic prediction of residential mobility.
Panel A depicts a heatmap of residential preferences among UK Biobank participants with standard controls, based on the Z-statistics (right-hand bar) of the association between Big Five PGI and 22 area classifications. Associations are estimated controlling for 5 principal components, sex, age, genotyping array, and a fixed effect for Middle-layer Super Output Area geographic region at birth. Panel B depicts a heatmap of residential preferences with additional controls for median home prices. This heatmap based on the Z-statistics (right-hand bar) of the association between Big Five PGI and 22 area classifications. Associations are estimated with controls in Panel A, plus additional control for local median home prices in 1995, 2000, 2005, and 2010. Panel C depicts PGI associations with moving or remaining in urban and rural areas among participants in the UK Biobank, stratified by gender. Point estimates are in standardized units. Error bars represent 95% confidence intervals.
Extended Data Fig. 8 Assortative Mating.
The left side of the figure depicts phenotypic correlations between mothers and fathers in the deCODE cohort (N = 5,317 unique mate pairs). Mates were defined as individuals known to have children together in the deCODE genealogical database. Each square shows the correlation coefficient with standard error in parenthesis. See also Supplementary Tables S40. The right side of the figure depicts a structural equation model for estimating genetic assortative mating among partners. In the path diagram of the latent PGI model, the two PGIs estimated from split-half discovery GWAS data are depicted as PGI 1 and PGI 2. The factor loadings are constrained to be equal to one another, as denoted by an = sign, and the within-PGI cross-sibling PGI residual associations are allowed to correlate. Factor variances are constrained to 1 to identify the metric of the genetic factors.
Extended Data Fig. 9 Mendelian Randomization.
The figure on top depicts the causal diagram assumed by Mendelian Randomization (MR) approaches. The dashed lines with the red X represent paths assumed to be absent by traditional MR. The three approaches to MR that were implemented in this paper use different approaches to estimating causal effects of the exposure on the outcome using multiple independent SNPs, such that they are robust to violations of these standard assumptions by a subset of SNPs. These approaches are described in the Supplementary Text. The table depicts results of MR tests that indicate putatively causal effects of personality traits on outcomes, or of outcomes on personality traits. Effects are bolded if the 95% confidence interval (in parentheses) excludes zero. For each pairing of exposure and outcome, we consider effects significant if confidence intervals exclude zero for at least two of the three estimators. Note: MR-CAUSE = Mendelian Randomization Causal Analysis Using Summary Effect estimates. Study Part.: = Study participation. IDK = “I don’t know” response. Summary statistics for outcomes are described in Supplementary Table S37. Numbers in parentheses indicate 95% confidence intervals (for weighted median and weighted mode tests) or 95% credible intervals (for MR-CAUSE tests). Bolded numbers indicate 95% confidence/credible interval excludes 0. Results presented here only include exposure-outcome pairs where the 95% credible interval/confidence interval for at least 2 out of 3 tests excludes zero, and all three tests produce estimates in the same direction. See also Supplementary Tables S42, S43.
Extended Data Fig. 10 Within-Family GWAS results.
Panels A-E depict Manhattan plots among meta-analysed cohorts of participants with EUR-like genomes for the traits of Extraversion (N = 47,630), agreeableness (N = 31,607), conscientiousness (N = 31,544), neuroticism (N = 55,637), and openness to experience (N = 32,216). In these panels, points correspond to two-sided test p-values of GWAS regressions. Genome wide significance is denoted with a red line corrected for multiple comparisons at p = 5 × 10−8. P-values smaller than this threshold are considered significant. The blue line indicates p = 5 × 10−5. See also Supplementary Table S47. Panels F-J depict the correspondence between (on the X axis) population-level genetic correlations between personality traits and external outcomes (from Fig. 3 in the main text) and (on the Y axis) genetic correlations where personality trait data come from within-family GWAS and external outcome data come from population-level GWAS. Each of 85 personality-outcome correlations are plotted for each of the Big Five. The dashed line with a slope of 1 and an intercept of 0 depicts equality in correlations across the two estimates. Genetic correlations were estimated using LDSC. See also Supplementary Table S48. Panels K-O depict the correspondence between (on the X axis) population-level genetic correlations between personality traits and external outcomes, using population-level GWAS data (from Fig. 3 in the main text) and (on the Y axis) within-family genetic correlations between personality traits and external outcomes, using within-family GWAS data. Within-family GWAS data of external outcomes come from Howe and colleagues (2022). For each outcome, 95% confidence intervals are presented stretching horizontally, for inferential uncertainty in population-level estimates, and vertically, for inferential uncertainty in within-family estimates. Outcomes are coloured if the within-family genetic correlation is significant at p < 0.05 and grey if nonsignificant. The dashed line with a slope of 1 and an intercept of 0 depicts equality in correlations across the two estimates. Genetic correlations were estimated using LDSC. See also Supplementary Table S49. Panel P depicts LDSC genetic correlations across the Big Five for the full population-level meta-analytic sample, and Panel Q depicts LDSC genetic correlations across the Big Five for the within-family meta-analytic sample. Numbers in parentheses represent standard errors.
Supplementary information
Rights and permissions
Open Access This article is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License, which permits any non-commercial use, sharing, distribution and reproduction in any medium or format, as long as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons licence, and indicate if you modified the licensed material. You do not have permission under this licence to share adapted material derived from this article or parts of it. The images or other third party material in this article are included in the article’s Creative Commons licence, unless indicated otherwise in a credit line to the material. If material is not included in the article’s Creative Commons licence and your intended use is not permitted by statutory regulation or exceeds the permitted use, you will need to obtain permission directly from the copyright holder. To view a copy of this licence, visit http://creativecommons.org/licenses/by-nc-nd/4.0/.
Reprints and permissions
About this article
Cite this article
Schwaba, T., Clapp Sullivan, M.L., Akingbuwa, W.A. et al. Robust inference and correlates from genetic associations with personality. Nature (2026). https://doi.org/10.1038/s41586-026-10992-9
Download citation
Received:
Accepted:
Published:
Version of record:
DOI: https://doi.org/10.1038/s41586-026-10992-9