
Article 25 February 2022
Chenyu Li, Shola M. Richards, Ghazal Quinn, Amin Abedini, Minyan Zhu, Tanya Verma, Samer Mohandes, Rebecca Pitts, Vesna Barros, Xiazi Qiu, Taehwan Shin, Joseph J. Loureiro, Nancy Finkel, Aditya Surapaneni, Josef Coresh, Morgan E. Grams, Anil Karihaloo, Hongzhe Li, Anurag Verma, Marylyn Ritchie, Daniel J. Rader, Penn Medicine BioBank, William F. Dietrich, Lori L. Jennings & Katalin Susztak
Abstract
Individuals of African ancestry carrying APOL1 (apolipoprotein L1) high-risk genotypes face a markedly increased risk of kidney failure, yet tools to identify those individuals likely to progress to chronic kidney disease are lacking. Here we profiled plasma proteomes of 851 Penn Medicine BioBank participants of African ancestry (285 males and 566 females) with APOL1 high-risk genotypes and preserved estimated glomerular filtration rate (eGFR) (≥60 ml min−1 1.73 m−2). Using elastic net Cox regression adjusted for age, sex, eGFR and albuminuria, we derived a nine-protein APOL1 Proteomic Risk Score (APRS) that predicts a composite outcome of ≥40% eGFR decline, kidney failure or death. APRS achieved a time-dependent area under the receiver operating characteristic curve (tAUC) of 86.5%, outperforming the Kidney Failure Risk Equation (66.1%) and polygenic risk scores, with 10-year event rates of 62.5% versus 3.3% across risk quintiles. External validation in Atherosclerosis Risk in Communities and UK Biobank cohorts confirmed robust accuracy (tAUC 82–85%) and consistent performance across demographic and clinical subgroups. Plasma levels of APRS component proteins correlated with kidney tissue fibrosis and tubular injury pathways, indicating strong biological plausibility. By enabling early and accurate prediction of disease progression in APOL1 high-risk individuals, APRS bridges the gap between genetic susceptibility and clinical translation. This scalable and biologically informed approach provides a precision medicine framework for early intervention and may accelerate development of APOL1-targeted therapies to reduce kidney disease disparities.
Similar content being viewed by others
Article 25 February 2022

Article Open access02 January 2025

Article Open access30 November 2023
Kidney failure (also known as end-stage kidney disease (ESKD)) is a life-threatening condition that requires dialysis or kidney transplantation for survival and imposes enormous global and societal costs. Worldwide1,2, chronic kidney disease (CKD) is estimated to affect more than 800 million people, and, in the United States alone, over 800,000 people are living with kidney failure. Medicare expenditures exceeded $52 billion in 2021, with per-person costs more than twice those without kidney failure3,4,5.
The burden of kidney failure is disproportionately high among African ancestry individuals, who develop kidney failure at nearly four times the rate of European ancestry individuals6. This disparity reflects social determinants of health, unequal access to care and genetic susceptibility. Variants in APOL1 (refs. 7,8), discovered in 2010, are among the strongest genetic risk factors for kidney failure. An estimated 4–5 million African Americans9,10 and tens of millions worldwide carry the high-risk genotype (two APOL1 risk alleles, G1 and/or G2)11,12. Although most high-risk carriers remain disease free, an estimated one in five progresses to kidney failure—substantially higher than in individuals with zero or one risk allele—making APOL1 a critical driver of racial disparities.
Therapies targeting APOL1 biology, such as the investigational inhibitor inaxaplin13,14, are now emerging and hold promise for preventing kidney failure in high-risk individuals. However, their use is constrained by the inability to identify which carriers are most likely to progress before CKD develops. Despite the major personal and economic burden of kidney failure, current prognostic tools remain inadequate. Clinical equations such as the Kidney Failure Risk Equation (KFRE)15 perform well only after CKD is established16, when much of the damage is irreversible. Genetic approaches, including polygenic risk scores (PRSs) and known APOL1 modifiers17, provide only modest discrimination and are not clinically actionable.
Plasma proteomics offers a potential solution. Protein levels are tightly regulated, reflect dynamic biology and can reveal subclinical injury not captured by standard measures18,19,20,21,22,23. We, therefore, performed broad-scale plasma proteomic profiling in APOL1 high-risk individuals with preserved eGFR (≥60 ml min−1 1.73 m−2) to develop and validate a prognostic biomarker score in two external cohorts. This approach directly addresses the gap in early risk prediction and provides a framework for precision medicine aimed at reducing disparities in APOL1-associated kidney disease.
Results
Baseline characteristics and outcomes
Of the 57,170 participants in the Penn Medicine BioBank (PMBB) who underwent exome sequencing, 1,310 carried the APOL1 high-risk (G1/G1, G2/G2 or G1/G2) genotype. After excluding 109 with prior kidney transplantation and 88 with kidney failure, 1,113 participants remained eligible for analysis (Fig. 1), including 262 with eGFR <60 ml min−1 1.73 m−2. As shown in Table 1, participants with eGFR ≥60 ml min−1 1.73 m−2 were younger (mean age, 49.2 ± 15.1 years versus 57.7 ± 15.8 years) and more often female (66.5% versus 51.5%). The mean eGFR in this group was 90.6 ± 17.0 ml min−1 1.73 m−2, and the median urine albumin−creatinine ratio (UACR) was 17.2 mg g−1 (interquartile range (IQR), 11−31). Hypertension (62.4% versus 85.1%), diabetes (27.9% versus 37.4%) and cardiovascular disease (9.2% versus 14.5%) were more common in the eGFR <60 group, whereas the protective p.N264K allele was more common in the eGFR ≥60 group than in the eGFR <60 group (4.7% versus 1.9%)17,24.
Fig. 1: Study design and analysis.
Data from the PMBB were used to integrate clinical, genetic and proteomic information over a 10-year follow-up period. Primary outcomes included mortality, diagnosis of ESKD, a ≥40% decline in eGFR, long-term dialysis and kidney transplantation. Proteomic profiling for ARIC validation was performed on visit 2 biospecimens.
Table 1 Baseline characteristics of the study participantsOver 10 years of follow-up, the composite outcome (≥40% eGFR decline, kidney failure or death) occurred in 18.0% of those with baseline eGFR ≥60 ml min−1 1.73 m−2 and in 55.3% of those with eGFR <60 ml min−1 1.73 m−2. These findings demonstrate a substantial event burden even before CKD is clinically apparent, underscoring the need for improved risk prediction.
Proteomic profiling and biomarker selectionBaseline plasma samples were profiled with SomaScan, quantifying 7,549 proteoforms. In participants with eGFR ≥60 ml min−1 1.73 m−2, 2,161 proteoforms were significantly associated with outcomes (Benjamini−Hochberg-adjusted P < 0.01; Extended Data Fig. 1). Because many proteins were highly correlated with each other and with baseline eGFR or UACR, we sought a panel of markers that predict progression independent of routinely measured clinical parameters. To develop and validate such a panel, the eGFR ≥60 group was randomly partitioned, with 80% of participants being used for marker selection and the remaining 20% reserved for independent testing (Supplementary Table 1). Elastic net Cox regression with cross-validation reduced redundancy and identified a nine-protein signature that independently predicted risk. To identify potential substitutes, we examined correlations between the nine proteins and the broader proteomic dataset. Although some proteins were highly correlated, suitable substitutes were difficult to identify for several markers (Extended Data Fig. 2). Within the nine-protein panel itself, correlations were weak in the eGFR ≥60 group (Pearsonʼs correlation coefficient, R, 0.1−0.4) and strong (0.3−0.7) when eGFR <60 ml min−1 1.73 m−2 (Extended Data Fig. 3).
To understand and improve the potential biological plausibility of our biomarkers as risk predictors, we examined the associations between circulating proteins and kidney tissue pathology in an independent cohort of human kidney samples (n = 474 for RNA sequencing (RNA-seq) and n = 325 for proteomics)23,25,26,27. Because fibrosis is an independent and strong predictor of kidney failure, we focused on this pathological feature and found that four of the nine proteins were associated with interstitial fibrosis in human kidney tissue (P < 0.01 at both the mRNA and protein levels; Extended Data Fig. 4 and Supplementary Fig. 1). This suggests that the panel may captured biologically relevant injury pathways not reflected by conventional clinical measures.
Development of the APRSHaving established the prognostic potential of nine individual markers that demonstrated prognostic discrimination in the eGFR ≥60 group (tAUC >67%; Extended Data Table 1), we next sought to integrate the nine proteins. These proteins were combined with age, sex, eGFR and UACR, using the same clinical variables as in the KFRE but reestimating their coefficients in our cohort, to generate the APOL1 proteomic risk score (APRS). The APRS substantially outperformed the KFRE in participants with eGFR ≥60 ml min−1 1.73 m−2 (Fig. 2a and Table 2). The KFRE was included as a comparator because it is the most widely validated clinical prediction tool for kidney outcomes, although its discrimination is known to diminish in individuals with normal eGFR. By contrast, in participants with eGFR <60 ml min−1 1.73 m−2, where the KFRE performs strongly, the APRS showed similar discrimination (Fig. 2b), with much of its variance explained by eGFR (R2 = 0.54; Extended Data Fig. 5, which shows R correlations between eGFR and APRS for any eGFR ≥60 and <60 groups). For eGFR <60 group, the R2 is 0.54.
Fig. 2: Proteomic risk prediction in APOL1 high-risk individuals.
a, tAUC and hazard ratios for risk prediction models and individual biomarkers among participants with eGFR ≥60 ml min−1 1.73 m−2 (n = 680). The APRS, conventional markers including eGFR (per 5 ml min−1 1.73 m−2), UACR (per doubling), the KFRE and individual biomarkers are shown for comparison. Shaded lines represent 95% confidence intervals. b, Similar analysis among participants with baseline eGFR <60 ml min−1 1.73 m−2 (n = 262). c, Cumulative event curves over follow-up (n = 851). Solid lines represent cumulative incidence estimates and shaded areas represent 95% confidence intervals, stratified by APRS above the 80th percentile versus below the 20th percentile in participants with eGFR ≥60 ml min−1 1.73 m−2. Corresponding hazard ratios were estimated using Cox proportional hazards models with two-sided Wald tests. d, Hazard ratios (solid lines) with 95% confidence intervals (shaded areas) by quintile for APRS and KFRE for composite event in participants with eGFR ≥60 ml min−1 1.73 m−2 (n = 851). Δ hazard ratios were estimated by comparing regression coefficients within a Cox proportional hazards model using a two-sided Wald test.
Table 2 tAUC over 10 years for predicting composite outcomes in different cohorts using various predictive modelsRisk stratification by quintiles revealed clear and consistent separation of outcomes: 10-year cumulative incidence ranged from 3.3% in the lowest APRS quintile to 62.5% in the highest (P = 4.73 × 10−23; Fig. 2c). Across every quintile, APRS provided better discrimination than KFRE, with a significantly steeper gradient of risk (Δ hazard ratio = 2.68; P = 1.45 × 10−24; Fig. 2d). Subgroup analyses showed that APRS effect size was robust across strata defined by age, sex, hypertension and diabetes (Supplementary Fig. 2). Interaction testing confirmed significant effect modification by APOL1 genotype (interaction hazard ratio = 1.35; P = 5.11 × 10−3), indicating that the prognostic strength of APRS was particularly pronounced in APOL1 high-risk compared to low-risk individuals (Extended Data Table 2 and Supplementary Fig. 3).
Performance of the APRSTo further assess its prognostic value, we examined the performance of the APRS in the eGFR ≥60 test cohort.
In this group, APRS achieved a tAUC of 86.5% for the composite outcome, 85.7% for mortality and 88.1% for kidney events (Fig. 3a,b, Extended Data Fig. 5 and Extended Data Table 3). At 5 years, APRS showed a sensitivity of 76.3% and a specificity of 86.7%, with concordance indexes (C-indexes) of 82.0% for the composite outcome, 84.4% for kidney outcomes and 78.5% for mortality (Fig. 3c). Decision curve analysis demonstrated greater net clinical benefit than either a treat-all or a treat-none strategy across a range of thresholds (Fig. 3d). Assuming an effective treatment (that is, inaxaplin) that reduces the risk of the composite endpoint by 27%13,14, APRS (≥95th percentile) nearly halved the number needed to treat (NNT; 4.9 with APRS versus 8.4 with KFRE and 23.9 with CKD PRS), underscoring its potential clinical utility (Extended Data Fig. 6a). Because mortality may act as a competing risk for kidney events, we further evaluated model performance using competing-risk analyses. Discrimination remained similar when death was treated as a competing event, with tAUC values of 86.5% for the composite outcome, 87.3% for kidney events and 82.6% for mortality (Extended Data Fig. 6b). Incorporating competing risks did not materially improve predictive performance, indicating that the final model retains robust discrimination in the presence of competing events.
Fig. 3: Comparative discrimination and clinical utility of four risk scores in APOL1 high-risk individuals with eGFR above 60 ml min−1 1.73 m−2.
a, tAUC plotted as a function of baseline eGFR (ml min−1 1.73 m−2), illustrating how discriminative accuracy varies across levels of kidney function (n = 433). b, tAUC plotted over follow-up time (years), showing how discrimination for each model evolves longitudinally (n = 171). c, AUC for 5-year event prediction for each risk score (n = 171). d, Decision curve analysis over 5 years across clinical decision thresholds, with reference lines for ‘treat all’ and ‘treat none’; NNTs were calculated among individuals above the 95th percentile of each score distribution (n = 171). In a and b, solid lines represent median tAUC estimates, whereas, in c and d, lines represent bootstrap mean estimates; shaded ribbons indicate 95% confidence intervals in all panels.
To benchmark APRS against existing tools (scores), we compared its performance to the KFRE, to the Chronic Renal Insufficiency Cohort (CRIC) proteomic score28 and to CKD PRS29. The CRIC proteomic score is notable as it contains 65 proteins derived from patients with established CKD (eGFR <60 ml min−1 1.73 m−2) yet shares only three proteins with our nine-protein APRS panel (Supplementary Table 2). In our primary target population with eGFR ≥60 ml min−1 1.73 m−2, APRS achieved higher discrimination than KFRE (tAUC 86.5% versus 66.2%) and outperformed the CRIC proteomic score (79.0%). APRS maintained this level of performance in the eGFR <60 stratum (262 participants and 145 events) where CRIC and KFRE were originally developed, performing similarly to KFRE (84.2% versus 82.4%) and CRIC (83.2%). PRS performed poorly across both eGFR strata (tAUC 58.5% and 53.8%), consistent with previous observations of limited predictive value for static genetic measures in diverse populations.
External validation of the APRSAPRS showed robust and reproducible performance across training, test and external cohorts (Table 2). In the PMBB training cohort of participants with eGFR ≥60 ml min−1 1.73 m−2, discrimination was strong (tAUC, 86.7%) and remained similar in the independent PMBB test set, consistently outperforming KFRE and CRIC score. Among APOL1 low-risk African ancestry participants in the PMBB, APRS provided substantially better discrimination (tAUC, 73.1%) than KFRE (60.1%) and CRIC score (69.3%). Similarly, among APOL1 low-risk European ancestry participants in the PMBB, APRS outperformed KFRE (66.9% versus 58.9%) but was similar to CRIC score (66.7%).
External validation in independent population-based cohorts (Extended Data Table 4) confirmed these findings. The Atherosclerosis Risk in Communities (ARIC) study included 314 APOL1 high-risk African American participants with preserved kidney function at baseline, representing a community-dwelling population with mean age of 55.8 years and mean eGFR of 96.2 ml min−1 1.73 m−2. Only 18 participants in this subgroup had eGFR below 60 ml min−1 1.73 m−2. In this validation set, APRS maintained strong discrimination with a tAUC of 82.2% over 10 years, with 43 composite events observed during follow-up. When restricted to participants with preserved eGFR (≥60 ml min−1 1.73 m−2), discrimination was modestly attenuated (tAUC, 77.5%), likely reflecting the absence of UACR measurements in ARIC; nevertheless, APRS substantially outperformed CRIC score in this setting (52.0%). The UK Biobank (UKBB) enrolled 204 APOL1 high-risk participants of African ancestry with similar baseline characteristics (mean age 52.2 years and mean eGFR 87.1 ml min−1 1.73 m−2, with nine participants having eGFR below 60 ml min−1 1.73 m−2). In this cohort, APRS showed strong discrimination with a tAUC of 84.7%. Despite differences in recruitment era, geographic location and healthcare systems across cohorts, APRS showed consistent performance beyond the academic medical center setting of PMBB. We additionally evaluated APRS, CRIC score and KFRE in participants without APOL1 high-risk genotypes across ancestries (Table 2). Although all three models retained some discriminatory ability in APOL1 low-risk populations, performance was uniformly attenuated compared to APOL1 high-risk groups, with APRS remaining similar to CRIC score and consistently outperforming KFRE.
DiscussionWe developed and validated a plasma proteomic risk score that substantially improves prediction of kidney outcomes in individuals carrying high-risk APOL1 genotypes. By integrating nine protein biomarkers (SPON1, SUMO2, EPHA10, REG3A, WFDC2, LYZ, MMP7, NPPB and CILP2) with limited clinical covariates, APRS markedly outperformed established clinical equations and genetic risk scores, especially among individuals with eGFR ≥60 ml min−1 1.73 m−2. This addresses a critical unmet need: once CKD is clinically apparent, patients already face elevated risks of cardiovascular disease, mineral and bone disorders and premature death, and progression to kidney failure can be rapid. Although APOL1 high-risk genotypes are among the strongest genetic predictors of kidney failure, genotype information alone has not been actionable for clinical decision-making. APRS overcomes this limitation by translating static genetic risk into a dynamic, clinically usable prediction tool.
Existing risk prediction tools have important limitations. The KFRE, although widely validated and accurate in advanced CKD, performs poorly when eGFR is normal. PRSs capture inherited susceptibility but remain static and ancestry dependent. Proteomic profiling, by contrast, provides a dynamic readout of ongoing biology18,19, integrating the cumulative influence of APOL1 risk variants and environmental exposures. In our study, more than 80% of participants had preserved kidney function without albuminuria—a group rarely included in prior risk prediction studies—underscoring both the novelty and the clinical relevance of this approach30,31,32,33. The APRS was developed and validated to address the critical unmet need in APOL1 high-risk individuals, achieving a tAUC of 86.5% in this group with preserved eGFR. As demonstrated in our full cross-population testing, this performance is substantially superior to the predictive value retained in APOL1 low-risk carriers of African ancestry (tAUC up to 80.2%) and European ancestry (tAUC up to 66.9%). APRS also performed similarly to KFRE in established CKD, suggesting that the biomarker panel captures both shared pathways of CKD progression and APOL1-specific mechanisms. These findings indicate that aptamer-based proteomics is an effective framework for risk prediction; although clinical implementation is most feasible and cost-effective in the enriched APOL1 high-risk group, the approach could also be extended to develop analogous models in other populations.
Proteins incorporated in APRS are associated with pathways plausibly involved in kidney injury that may not be captured by conventional measures in individuals with preserved function. In the preserved GFR group, APRS showed weak correlations with eGFR, and kidney tissue expression of these proteins was associated with fibrosis severity, suggesting that the signature may relate to early tissue damage before functional decline becomes apparent. Even individuals with preserved eGFR and low albuminuria remain at risk of progression, indicating that reliance on UACR alone could miss early disease signals. Specific proteins point to biologically plausible hypotheses for future testing: MMP7 has been reported as a marker of tubular injury and fibrosis23; WFDC2 correlates with interstitial fibrosis and rapid decline34,35; and LYZ is associated with fibroblast proliferation and tubular cell senescence36. Together, these findings generate hypotheses that APRS may reflect tubular stress, immune activation and extracellular matrix remodeling, which are implicated in APOL1-mediated nephropathy37,38,39,40. However, these associations are correlative; whether these proteins actively drive disease progression or reflect altered renal clearance due to reduced nephron number requires experimental validation. Thus, although APRS predicts risk, its biological interpretation remains hypothesis generating rather than mechanistically proven.
The clinical implications of APRS are substantial. First, APRS may enable earlier identification of high-risk individuals for intensified surveillance long before CKD is detected by standard measures. Second, it could guide use of emerging APOL1-targeted therapies such as inaxaplin, by identifying those most likely to benefit and by serving as a pharmacodynamic readout of treatment effect13,14. APRS nearly halved NNT, making it a more efficient tool for targeting interventions. Third, serial measurements may permit longitudinal monitoring of risk trajectories in clinical practice. Finally, APRS could transform trial design by enriching enrollment with high-risk individuals, thereby increasing event rates, reducing sample size and accelerating therapeutic development. Conversely, low APRS values could provide reassurance for carriers unlikely to progress, potentially avoiding unnecessary interventions. Notably, given the disproportionate burden of CKD and kidney failure among individuals of African ancestry, early risk stratification with APRS offers a precision medicine strategy to help narrow longstanding health disparities. In this context, the reliance of APRS on aptamer-based proteomic profiling may further support its translational potential, as this technology offers advantages in simplicity, scalability and cost-effectiveness compared to traditional antibody-based protein assays such as ELISA. Nevertheless, low APRS values should not be interpreted as a substitute for standard clinical monitoring, and APRS should be considered an adjunct to established risk assessment approaches. Furthermore, prospective decision impact studies will be required to determine whether APRS-guided strategies meaningfully improve clinical outcomes before routine clinical implementation.
Several limitations should be acknowledged. Because this was an observational study, residual confounding cannot be excluded, and limitations inherent to electronic health record (EHR) codes, including potential misclassification, incomplete capture of clinical events and variable coding practices across sites, may have affected outcome ascertainment. External validation in ARIC and UKBB confirmed robust performance, but event counts in these cohorts were modest, leading to wide confidence intervals. Cross-platform differences (SomaScan versus Olink) required harmonization, underscoring the need for assay standardization. Finally, although aptamer technology is scalable and cost-effective relative to antibody-based assays, technical complexity may limit near-term clinical deployment. Ongoing improvements in proteomic platforms and decreasing costs are likely to reduce these barriers. Prospective trials will be required to establish how best to integrate the APRS into everyday clinical care.
The APRS offers a valuable early risk stratification framework for individuals with high-risk APOL1 genotypes. With targeted therapies such as inaxaplin already advancing through clinical trials, the roadmap to translate this finding into clinical utility involves several complementary steps. Independent validation in prospective cohorts represents an important next phase to confirm efficacy in real-world settings. Concurrently, the APRS could serve as a subject enrichment tool for these therapeutic trials, potentially improving efficiency and reducing costs. Finally, adapting the current proteomic methodology into a simplified, high-throughput clinical assay would facilitate broader accessibility for routine risk assessment. In conclusion, a plasma proteomic risk score enables accurate and early prediction of adverse events in APOL1 high-risk individuals, before CKD is clinically evident. By transforming APOL1 genetic risk into a clinically actionable prediction tool, the APRS provides a precision medicine framework to support early intervention and reduce longstanding racial disparities in kidney failure.
Methods
PMBB cohort
The PMBB is a large academic biobank that recruits participants from the University of Pennsylvania Health System, with recruitment beginning in 2008. To date, the PMBB has enrolled over 250,000 participants—approximately 30% of whom are from non-European ancestries—and around 57,170 individuals have undergone whole-exome sequencing41. Demographic information, medical history (via International Classification of Disease (ICD) codes), medication use and clinical assessments were extracted from the EHR. Laboratory tests, including serum creatinine, blood urea nitrogen and electrolyte levels at baseline, were also obtained. The eGFR was calculated using the creatinine-based 2021 Chronic Kidney Disease Epidemiology Collaboration (CKD-EPI) equation42.
All participants provided written informed consent for the use of their biospecimens, genetic data and EHR data for research. Genomic DNA samples were transferred to the Regeneron Genetics Center and stored at –80 °C until sample preparation. Whole-exome sequencing was performed with reads mapped to Genome Reference Consortium Build 38 (GRCh38); samples failing quality metrics (for example, low sequencing coverage) were excluded, as described previously. Ancestry was estimated by exome data using principal component analysis. A set of high-quality, common single-nucleotide polymorphisms overlapping with HapMap3 was extracted. Principal components were first calculated for HapMap3 samples and then used as the reference space, onto which all study samples were projected. To classify ancestry, a kernel density estimation approach was applied to the joint distributions of the first four principal components for each HapMap3 population. For each sample, normalized likelihoods of belonging to each reference population were obtained. Reference populations were then assigned if likelihoods exceeded prespecified thresholds. Based on these assignments, samples were grouped into one of the following ancestral classes: African, European, East Asian, South Asian or Admixed American. Subsequently, the APOL1 was analyzed in all African ancestry participants with available exome sequence data41.
Individuals were classified into high-risk APOL1 genotype groups based on the presence of the G1 (G1a and G1b) and G2 risk alleles. Participants with two risk alleles (G1/G1, G2/G2 or G1/G2) were classified as high-risk11,12, whereas those with zero or one risk allele (G0/G0, G0/G1 or G0/G2) were classified as low-risk. We included all participants aged ≥18 years with available high-risk APOL1 genotype and at least one subsequent follow-up record or documented clinical event (for example, mortality and dialysis). Exclusion criteria included known ESKD or with kidney transplantation at baseline, missing or ambiguous APOL1 data and excessive missing data in key clinical covariates. To minimize reverse causation, clinical diagnoses (for example, hypertension, diabetes and cardiovascular disease) were required to have been recorded in the EHR prior to the blood draw. Laboratory values used as baseline covariates were extracted from tests performed within a prespecified window of 2 months before to 1 month after the plasma draw; for each participant, we used the single measurement closest to the draw date to reduce temporal misalignment. For UACR, we used the most recent measurement obtained within the 2 years preceding the plasma draw. When only a urine protein–creatinine ratio or a dipstick protein result was available, we converted these to estimated UACR using the conversion equations from ref. 43. This work was conducted under University of Pennsylvania Institutional Review Board (IRB)-approved protocols (815796, 813913, 855821 and 857403; study protocol).
Outcomes and follow-upThe primary outcomes were defined as a composite of kidney events (long-term dialysis, ESKD diagnosis, kidney transplantation or a ≥40% decline in eGFR from baseline) and all-cause mortality. Mortality was included given its clinical importance and as a critical competing risk for kidney failure44. For APOL1 high-risk individuals with eGFR ≥60 ml min−1 1.73 m−2, the mean follow-up duration was 7.18 ± 2.89 years, with a mean time to event of 6.51 ± 3.07 years. Outcome data were obtained through medical record review and ICD-10 codes. Participants were censored at the time of death, loss to follow-up or the end of a 10-year follow-up period, whichever occurred first.
Proteomic profilingPMBB plasma samples were collected at baseline and subjected to proteomic profiling using the SomaScan platform (SomaLogic), which employs modified aptamers (SOMAmers) to probe the human proteome22. SomaScan v.4.1 pools 7,524 SOMAmers to probe 6,386 distinct proteins (https://menu.somalogic.com/). The SomaScan assay is a highly multiplexed, sensitive and reproducible proteomic technology that has been extensively validated and used in numerous clinical studies23,30,31,32,33,45,46,47,48. The SomaScan assay is based on the use of modified DNA aptamers, called SOMAmers, which are designed to bind specific protein targets with high affinity and specificity22. Each SOMAmer is uniquely tagged with a DNA barcode that allows for quantification using a custom DNA microarray. In brief, the diluted plasma samples were incubated with a pool of SOMAmers, and, subsequently, the SOMAmer−protein complexes were captured on streptavidin-coated beads, washing away unbound proteins and SOMAmers. The bound SOMAmers were then eluted and hybridized to a custom DNA microarray containing complementary sequences to the SOMAmer barcodes. The microarrays were scanned using a SureScan Dx Microarray Scanner (Agilent Technologies), and the fluorescence intensity of each SOMAmer was quantified as a measure of the relative abundance of its corresponding protein target.
Raw data were processed using SomaScan Data Analysis Software (SomaLogic) to generate relative fluorescence units (RFUs) for each SOMAmer. The RFUs were then normalized using a set of internal calibrator samples to adjust for any assay-specific biases and to ensure comparability across different plates and runs. Rigorous quality control measures were implemented throughout the proteomic profiling process to ensure data integrity and reliability. These included the use of internal calibrator samples, quality control metrics for sample and assay performance and the exclusion of any samples or SOMAmers that failed to meet predefined quality criteria49. The log2-transformed RFUs were then used for subsequent statistical analyses and model development.
UKBB cohortThe UKBB is a large-scale, community-based cohort comprising over 500,000 participants aged 40–69 years at recruitment (2006–2010) from 22 assessment centers across the United Kingdom. For validation purposes, we included all participants who underwent plasma proteomic profiling50. This subset is representative of the overall UKBB population. Among these, 1,171 individuals of African descent were identified—204 with APOL1 high-risk. APOL1 genotypes were imputed (using Data-Field 21007) based on the TOPMed R2 reference panel after phasing with Eagle v.2.4 and converting from GRCh37 to GRCh38 via LiftOver51. ESKD was identified using data_coding_19 codes N185 and N180 in fields 41270 and 41280 and using READV3_CODE mapped to ICD-10 codes N185 and N180 in Table 1060; hypertension was captured using data_coding_19 with the regular expression ‘I1[012345]’ in fields 41270 and 41280 and using READV3_CODE mapped to ICD-10 with the same expression in Table 1060; diabetes was defined using data_coding_19 with the regular expression ‘E1[01234]’ in fields 41270 and 41280 and using READV3_CODE mapped to ICD-10 with the same expression in Table 1060; kidney transplant history was ascertained using data_coding_19 code Z940 in fields 41270 and 41280 and using READV3_CODE mapped to ICD-10 code Z940 in Table 1060; dialysis treatment was identified through procedure_concept entries matching ‘[Dd]ialysis’ in Table 936 and TERMV3_DESC entries matching ‘[Dd]ialysis’ in Table 1060; mortality outcomes were obtained from death register data in Table 1058; serum creatinine measurements were drawn from measurement_concept_id 37392176 and 3020564 in Table 931, from repeat-measure identifiers p30700_i.* and p23478_i.* in field 30700 and from TERMV3_DESC entries matching ‘[Ss]erum [Cc]reatinine$’ in Table 1060; and cystatin C was obtained via measurement_concept_id 3030366 in Table 931. Olink measurements were log2 transformed and normalized to harmonize scales. Missing proteins in Olink were imputed by mapping overlapping proteins to the SomaScan scale and predicting absent targets with a multi-output penalized regression trained on SomaScan data (statistical analysis protocol). This study was conducted under application number 273810. The validation from the UKBB for this project was approved by the University of Pennsylvania IRB (protocol 855821).
ARIC studyThe ARIC study enrolled 15,792 participants (aged 45–65 years) from four US communities (Washington County, Maryland; Forsyth County, North Carolina; Jackson, Mississippi; and Minneapolis, Minnesota) between 1987 and 1989, with follow-up visits every 3 years initially and subsequently at varying intervals52. We included 314 African American participants with APOL1 high-risk genotyping (performed using TaqMan assays for G1 and G2)8,53 who attended visit 2 (approximately 3 years after baseline) with eGFR≥60 ml min−1 1.73 m−2; proteomic profiling on these visit 2 biospecimens was performed using the SomaScan 5K assay. To ensure comparability with the PMBB, we truncated follow-up at 10 years from baseline. This study was conducted under application number MP4524. The validation from the ARIC for this project was approved by the University of Pennsylvania IRB (protocol 855821). All validation analyses were independently performed by statisticians from different institutions.
Human kidney samplesKidney tissue samples (n = 474 for RNA-seq and n = 325 for proteomics) for this study were procured from surgical nephrectomies, ensuring that only the normal parts of the tissue, specifically those at least 2 cm from any cancerous lesions, were used for analysis. An honest broker deidentified the samples and collected corresponding clinical information, such as age, race, sex and diabetes and hypertension status, in addition to creatinine values. The eGFR was subsequently determined using the latest CKD-EPI equations42. Kidney samples were formalin fixed, paraffin embedded and stained with periodic acid–Schiff. Whole-tissue imaging was conducted using the Aperio system, which is a platform that digitizes and assists in analyzing pathology slides. Samples were scored in an unbiased manner by a specialized renal pathologist23,25,26. The use of these samples and data was approved by the University of Pennsylvania IRB under the category of ‘exempt’, negating the need for informed consent due to the deidentified nature of the study samples.
Kidney tissue RNA-seq and data processingRNA isolation, sequencing and analysis were performed as previously published27. Total RNA was isolated from kidney tissue using the RNeasy Mini Kit (Qiagen) according to the manufacturer’s instructions, including the DNase digestion step. RNA quality was assessed by Agilent Bioanalyzer 2100. The cDNA library was prepared using NEBNext Ultra II RNA Library Prep Kit for Illumina. Then, cDNA libraries were sequenced on an Illumina NovaSeq 6000 platform using the NovaSeq PE150 protocol. Adaptor and lower-quality bases were trimmed with TrimGalore (v.0.4.5). Reads were aligned to the human genome (hg19) using STAR (v.2.7.3a). Gene and isoform expression levels of transcripts per million were estimated using RSEM (v.1.3.0).
Kidney tissue proteomics and data processingKidney tissue samples were snap frozen, cryopulverized (CryoMill; Retsch) and lysed using T-PER extraction reagent supplemented with protease inhibitors (Roche Diagnostics). Protein concentrations were determined via bicinchoninic acid assay (Thermo Fisher Scientific). Similar to plasma proteomics, the tissue proteomic profiling was performed using the SomaScan v.4.1 platform (SomaLogic), which uses slow off-rate modified DNA aptamers (SOMAmers) to quantify protein targets. To ensure data consistency, quality control measures included hybridization controls, pooled calibrators and buffer-only replicates for monitoring background signals and batch effects. Data underwent normalization to correct for within-run hybridization variability, followed by intrastudy median normalization.
PRSA previously published genome-wide PRS for CKD29 was applied. For each individual, the PRS was calculated as the weighted sum of risk alleles, with weights corresponding to the effect sizes reported in the original study. In brief, the PRS was computed as , where M is the number of variants with non-zero weights and βi is the effect size for variant i. To enable genome-wide PRS computation, we used imputed genotype data from the PMBB (Release 2.0), which includes participants genotyped using the Illumina Global Screening Array (Freeze 2.0). Genotypes were phased with Eagle2 and imputed using Minimac4 on the Michigan TOPMed r3 reference panel (GRCh38). The imputed dataset was filtered to retain high-confidence variants (average R2 > 0.3 or directly genotyped in either batch) and minor allele frequency > 0.01, ensuring compatibility with published PRS weights. PRS values were computed in PLINK 2.0 using the overlapping variant set between the published PRS and the PMBB-imputed genotypes. The PRS analysis was included primarily as a benchmark to contextualize the incremental predictive value of the APRS.
CRIC proteomics modelFor benchmarking, we applied the previously published CRIC proteomic risk score, which was developed using SomaScan profiling of 65 plasma proteins in patients with established CKD28. The original coefficients were applied to our SomaScan data after log2 transformation and harmonization of units. For each participant, a CRIC score was calculated as the weighted sum of these proteins. The performance of the CRIC score was evaluated using tAUC and C-index for the composite outcome in the PMBB.
Marker selection and model constructionTo identify the optimal combination of proteins for outcome prediction, we implemented an elastic net penalized Cox proportional hazards modeling framework based on the random generation of candidate protein panels. This strategy is conceptually related to random subspace approaches54 and ensemble feature selection methods55,56 but differs in that resampling is performed on the feature space rather than on individuals, and each candidate protein panel is evaluated independently. This was implemented through the following sequential steps: (1) initial protein screening and candidate panel generation by identifying proteins showing significant univariate associations (Benjamini−Hochberg-adjusted P < 0.05) with the outcome (top 20%) and then randomly sampling 40–90% of these proteins without replacement for each candidate panel; (2) elastic net modeling with grid search, fitting an elastic net Cox model for each sampled protein panel and conducting a comprehensive grid search across the mixing parameter α (ranging from 0.1 to 0.9 in 0.1 increments) and the regularization strength λ (spanning 100 logarithmically spaced values); (3) performance evaluation via eight-fold cross-validation using the mean C-index and large-scale combinatorial exploration of approximately one million unique candidate protein panels to thoroughly explore the vast combinatorial space of protein interactions and identify the optimal combination; (4) intermediate feature selection by stability, retaining proteins with a selection frequency of at least 30% across candidate panels to avoid multiple testing problems and reduce noise, ensuring that only consistently predictive proteins were retained; and (5) final model construction and refinement using the intermediate protein set with another round of elastic net penalized Cox proportional hazards modeling to eliminate highly correlated proteins and optimize model parameters based on the most promising candidates identified in the initial screening and evaluation process. The final proteomic signature was defined as the panel of proteins and corresponding penalty parameters that achieved the highest average cross-validated C-index across all iterations.
The model risk score is: , where are the coefficients and are the protein markers. The APRS formula is:
Model performance evaluationWe evaluated the performance of the predictive model using a comprehensive set of metrics designed for survival analysis. The evaluation function takes as input the true survival outcomes from the training and test sets, along with the predicted risk scores and the timepoints at which the predictions are made. The function first prepares the data by aligning the predicted risk scores with the true survival outcomes and applying inverse probability of censoring weights (IPCW) to account for censoring in the test set. The IPCW for each sample i at time is calculated as: where is the estimated probability of being uncensored at time given the covariates . The samples are then sorted by descending risk score at each timepoint. Let be the number of samples and m be the number of timepoints. For each timepoint , we define: : true positives, : false positives, : true negatives, : false negatives. Next, the function calculates various performance metrics at each timepoint, including: AUC: where is the true-positive rate (sensitivity) and is the false-positive rate (1 − specificity) at time . Model sensitivity, specificity F1 score, accuracy, Matthews correlation coefficient (MCC), positive predictive value (PPV) and negative predictive value (NPV) were calculated as: , , F1t where Precisiont , and Recallt , Accuracyt , MCCt , and . These metrics are computed by considering the predicted risk scores as a binary classifier at each timepoint, with the threshold determined by the point that maximizes the sum of sensitivity and specificity. Finally, if multiple timepoints are evaluated, the function computes the mean of each performance metric weighted by the probability of survival at each timepoint:
where is the performance metric at time and is the estimated probability of survival at time . The evaluation function returns a dictionary containing the performance metrics at each timepoint and, if applicable, the weighted mean of each metric across all timepoints. All performance metrics are reported on a 0–100% scale for simplicity.
Statistical analysisContinuous variables are presented as median ± s.d. or IQR, and categorical variables are presented as frequencies and percentages. The independent t-test was used for normally distributed continuous variables, the Mann−Whitney U-test for non-normally distributed continuous variables and Pearson’s χ2 test for categorical variables. Missing data were handled across the entire PMBB dataset. Variables with more than 15% missingness were excluded. Remaining missing values were imputed using multiple imputation by chained equations; imputation was performed separately within the training and test sets to avoid information leakage. Survival analyses were conducted using the Kaplan−Meier method, and differences in survival probabilities among the study groups were assessed using the log-rank test. The proportional hazards assumption was assessed using covariate-specific and global tests based on Schoenfeld residuals (Grambsch–Therneau), together with graphical evaluation of log–minus–log survival plots and plots of scaled Schoenfeld residuals versus time. Where evidence of non-proportionality was observed, time-dependent effects were modeled by introducing interactions between the covariate and functions of time (for example, log(time)) or by fitting stratified Cox models. The predictive performance of the developed models was evaluated using a comprehensive set of metrics, including AUC, C-index, specificity, sensitivity, F1 score, precision, recall, accuracy, FPR, FNR, MCC, PPV and NPV. To systematically identify proteins correlated with the nine APRS proteins, we computed multiple similarity metrics between each APRS protein vector and all other proteins measured by SomaScan. In the radar plot, proteins from APRS were ordered using a greedy traversal of the core–core correlation matrix, iteratively selecting the most strongly correlated unvisited protein. Non-APRS protein angles were determined by a weighted circular mean of APRS protein angles, with weights derived from their absolute correlations to APRS proteins.
The APOL1 high-risk with eGFR ≥60 ml min−1 1.73 m−2 group was divided into an 80% training set and a 20% test set using stratified sampling. In the training set, an elastic net Cox model was applied to select protein markers. Model performance was evaluated using eight-fold cross-validation to minimize overfitting, and predictive accuracy was assessed with AUC and tAUC at various timepoints57. The final model included nine proteins plus age, sex, baseline eGFR and log2(UACR) and was fitted using an elastic net Cox model. Sample size was estimated using the Schoenfeld method for the Cox proportional hazards model and an events per variable (EPV) approach (with EPV set at 15). The EPV method yielded the required sample size of 640 participants. By contrast, based on a two-sided significance level of 0.05, an expected 30% marker-positive rate, an anticipated event rate of 25% and a target hazard ratio of 1.648—assuming median dichotomization of the model score—a total of 685 participants was determined to provide an effective statistical power of approximately 85%. Net benefits from decision curve analysis were calculated by weighting false positives and false negatives under the assumption of equal misclassification costs58. Feature stability was checked by bootstrap enumeration. NNT was calculated as the reciprocal of the absolute risk reduction, where absolute risk reduction = control event rate × relative risk reduction (RRR = 0.27, reported for inaxaplin)13. Two-tailed P values less than 0.05 were considered statistically significant. All analyses were performed using Python 3.9.13 and R 4.3.2.
Reporting summaryFurther information on research design is available in the Nature Portfolio Reporting Summary linked to this article.
Data availabilityThe datasets analyzed in this study are not publicly available owing to participant privacy and data use agreements but may be accessed through application to the PMBB (https://pmbb.med.upenn.edu/). Access requires approval by the PMBB data access committee and execution of the PMBB data use agreement. Requests for proteomic data or verification analyses may be directed to the corresponding author (ksusztak@pennmedicine.upenn.edu). An initial response is generally provided within approximately 2 weeks. ARIC data may be requested from the ARIC Data Coordinating Center by obtaining study approval, executing a data and materials distribution agreement and submitting a data request form to aricdata@unc.edu. Alternatively, ARIC data are available through the NHLBI BioLINCC (https://biolincc.nhlbi.nih.gov/) repository and the dbGaP (https://dbgap.ncbi.nlm.nih.gov/beta/study/phs000280.v9.p3) subject to their application procedures. Review timelines typically require approximately 4–8 weeks. For datasets derived from the UKBB, data access is governed by the UKBB’s established policies. Researchers must apply through the UKBB Access Management System, available at https://www.ukbiobank.ac.uk/, outlining the purpose of the intended use. Applications are reviewed by the UKBB, and decisions are generally provided within approximately 4–6 weeks. Source data are provided with this paper.
ReferencesRomagnani, P. et al. Chronic kidney disease. Nat. Rev. Dis. Primers 3, 17088 (2017).
GBD Chronic Kidney Disease Collaboration. Global, regional, and national burden of chronic kidney disease, 1990−2017: a systematic analysis for the Global Burden of Disease Study 2017. Lancet 395, 709–733 (2020).
Liyanage, T. et al. Worldwide access to treatment for end-stage kidney disease: a systematic review. Lancet 385, 1975–1982 (2015).
Levin, A. et al. Global kidney health 2017 and beyond: a roadmap for closing gaps in care, research, and policy. Lancet 390, 1888–1917 (2017).
National Institute of Diabetes and Digestive and Kidney Diseases. Kidney disease statistics for the United States. NIDDK https://www.niddk.nih.gov/health-information/health-statistics/kidney-disease (2024).
Muntner, P. et al. Racial differences in the incidence of chronic kidney disease. Clin. J. Am. Soc. Nephrol. 7, 101–107 (2012).
Genovese, G. et al. Association of trypanolytic APOL1 variants with kidney disease in African Americans. Science 329, 841–845 (2010).
Parsa, A. et al. APOL1 risk variants, race, and progression of chronic kidney disease. N. Engl. J. Med. 369, 2183–2196 (2013).
Dummer, P. D. et al. APOL1 kidney disease risk variants: an evolving landscape. Semin. Nephrol. 35, 222–236 (2015).
Kruzel-Davila, E., Wasser, W. G., Aviram, S. & Skorecki, K. APOL1 nephropathy: from gene to mechanisms of kidney injury. Nephrol. Dial. Transplant. 31, 349–358 (2016).
Kopp, J. B. et al. APOL1 genetic variants in focal segmental glomerulosclerosis and HIV-associated nephropathy. J. Am. Soc. Nephrol. 22, 2129–2137 (2011).
Limou, S., Nelson, G. W., Kopp, J. B. & Winkler, C. A. APOL1 kidney risk alleles: population genetics and disease associations. Adv. Chronic Kidney Dis. 21, 426–433 (2014).
Egbuna, O. et al. Inaxaplin for proteinuric kidney disease in persons with two APOL1 variants. N. Engl. J. Med. 388, 969–979 (2023).
Gbadegesin, R. & Lane, B. Inaxaplin for the treatment of APOL1-associated kidney disease. Nat. Rev. Nephrol. 19, 479–480 (2023).
Tangri, N. et al. Multinational assessment of accuracy of equations for predicting risk of kidney failure a meta-analysis. JAMA 315, 164–174 (2016).
Grams, M. E. et al. The Kidney Failure Risk Equation: evaluation of novel input variables including eGFR estimated using the CKD-EPI 2021 equation in 59 cohorts. J. Am. Soc. Nephrol. 34, 482–494 (2023).
Gupta, Y. et al. Strong protective effect of the APOL1 p.N264K variant against G2-associated focal segmental glomerulosclerosis and kidney disease. Nat. Commun. 14, 7836 (2023).
Aebersold, R. & Mann, M. Mass-spectrometric exploration of proteome structure and function. Nature 537, 347–355 (2016).
Mischak, H., Delles, C., Vlahou, A. & Vanholder, R. Proteomic biomarkers in kidney disease: issues in development and implementation. Nat. Rev. Nephrol. 11, 221–232 (2015).
Kammer, M. et al. Integrative analysis of prognostic biomarkers derived from multiomics panels helps discrimination of chronic kidney disease trajectories in people with type 2 diabetes. Kidney Int. 96, 1381–1388 (2019).
Cisek, K., Krochmal, M., Klein, J. & Mischak, H. The application of multi-omics and systems biology to identify therapeutic targets in chronic kidney disease. Nephrol. Dial. Transplant. 31, 2003–2011 (2016).
Gold, L. et al. Aptamer-based multiplexed proteomic technology for biomarker discovery. PLoS ONE 5, e15004 (2010).
Hirohama, D. et al. Unbiased human kidney tissue proteomics identifies matrix metalloproteinase 7 as a kidney disease biomarker. J. Am. Soc. Nephrol. 34, 1279–1291 (2023).
Narjoz, C., Vinh-Hoang-Lan Julie, T., Rabant, M., Karras, A. & Pallet, N. Diagnostic yield of APOL1 p.N264K variant screening in daily practice. Kidney Int. Rep. 9, 1916–1918 (2024).
Quinn, G. Z. et al. Renal histologic analysis provides complementary information to kidney function measurement for patients with early diabetic or hypertensive disease. J. Am. Soc. Nephrol. 32, 2863–2876 (2021).
Hirohama, D. et al. The proteogenomic landscape of the human kidney and implications for cardio-kidney-metabolic health. Nat. Med. 31, 3917−3929 (2025).
Sheng, X. et al. Mapping the genetic architecture of human traits to cell types in the kidney identifies mechanisms of disease and potential treatments. Nat. Genet. 53, 1322–1333 (2021).
Dubin, R. F. et al. Proteomics of CKD progression in the chronic renal insufficiency cohort. Nat. Commun. 14, 6340 (2023).
Khan, A. et al. Genome-wide polygenic score to predict chronic kidney disease across ancestries. Nat. Med. 28, 1412–1420 (2022).
Jiang, L. et al. Prospective observational study on biomarkers of response in pancreatic ductal adenocarcinoma. Nat. Med. 30, 749–761 (2024).
Helgason, H. et al. Evaluation of large-scale proteomics for prediction of cardiovascular events. JAMA 330, 725–735 (2023).
Williams, S. A. et al. Plasma protein patterns as comprehensive indicators of health. Nat. Med. 25, 1851–1857 (2019).
Niewczas, M. A. et al. A signature of circulating inflammatory proteins and development of end-stage renal disease in diabetes. Nat. Med. 25, 805–813 (2019).
LeBleu, V. S. et al. Identification of human epididymis protein-4 as a fibroblast-derived mediator of fibrosis. Nat. Med. 19, 227–231 (2013).
Schmidt, I. M. et al. Plasma proteomics of acute tubular injury. Nat. Commun. 15, 7368 (2024).
Ren, Y., Yu, M. J., Zheng, D. N., He, W. F. & Jin, J. Lysozyme promotes renal fibrosis through the JAK/STAT3 signal pathway in diabetic nephropathy. Arch. Med. Sci 20, 233–247 (2024).
Beckerman, P. et al. Transgenic expression of human APOL1 risk variants in podocytes induces kidney disease in mice. Nat. Med. 23, 429–438 (2017).
Wu, J. et al. APOL1 risk variants in individuals of African genetic ancestry drive endothelial cell defects that exacerbate sepsis. Immunity 54, 2632–2649 (2021).
Wu, J. et al. The key role of NLRP3 and STING in APOL1-associated podocytopathy. J. Clin. Invest. 131, e136329 (2021).
McNulty, M. T. et al. A glomerular transcriptomic landscape of apolipoprotein L1 in Black patients with focal segmental glomerulosclerosis. Kidney Int. 102, 136–148 (2022).
Verma, A. et al. The Penn Medicine BioBank: towards a genomics-enabled learning healthcare system to accelerate precision medicine in a diverse population. J. Pers. Med. 12, 1974 (2022).
Inker, L. A. et al. New creatinine- and cystatin C−based equations to estimate GFR without race. N. Engl. J. Med. 385, 1737–1749 (2021).
Sumida, K. et al. Conversion of urine protein−creatinine ratio or urine dipstick protein to urine albumin−creatinine ratio for use in chronic kidney disease screening and prognosis: an individual participant-based meta-analysis. Ann. Intern. Med. 173, 426–435 (2020).
Perkovic, V. et al. Canagliflozin and renal outcomes in type 2 diabetes and nephropathy. N. Engl. J. Med. 380, 2295–2306 (2019).
Ferkingstad, E. et al. Large-scale integration of the plasma proteome with genetics and disease. Nat. Genet. 53, 1712–1721 (2021).
Lehallier, B. et al. Undulating changes in human plasma proteome profiles across the lifespan. Nat. Med. 25, 1843–1850 (2019).
Deo, R. et al. Proteomic cardiovascular risk assessment in chronic kidney disease. Eur. Heart J. 44, 2095–2110 (2023).
Xu, Y. et al. An atlas of genetic scores to predict multi-omic traits. Nature 616, 123–131 (2023).
Kim, C. H. et al. Stability and reproducibility of proteomic profiles measured with an aptamer-based platform. Sci. Rep. 8, 8382 (2018).
Sun, B. B. et al. Plasma proteomic associations with genetics and health in the UK Biobank. Nature 622, 329–338 (2023).
Taliun, D. et al. Sequencing of 53,831 diverse genomes from the NHLBI TOPMed Program. Nature 590, 290–299 (2021).
Wright, J. D. et al. The ARIC (Atherosclerosis Risk In Communities) study: JACC Focus Seminar 3/8. J. Am. Coll. Cardiol. 77, 2939–2959 (2021).
Rosenberg, A. et al. Surrogate endpoints in apolipoprotein L1−associated kidney disease: evaluation in three cohorts. Clin. J. Am. Soc. Nephrol. 20, 23−30 (2025).
Ho, T. K. The random subspace method for constructing decision forests. IEEE Trans. Pattern Anal. Mach. Intell. 20, 832–844 (1998).
Meinshausen, N. & Bühlmann, P. Stability selection. J. R. Stat. Soc. B 72, 417–473 (2010).
Yin, Q. Y., Li, J. L. & Zhang, C. X. Ensembling variable selectors by stability selection for the Cox model. Comput. Intell. Neurosci. 2017, 2747431 (2017).
Kamarudin, A. N., Cox, T. & Kolamunnage-Dona, R. Time-dependent ROC curve analysis in medical research: current methods and applications. BMC Med. Res. Methodol. 17, 53 (2017).
Vickers, A. J. & Elkin, E. B. Decision curve analysis: a novel method for evaluating prediction models. Med. Decis. Making 26, 565–574 (2006).
This work was supported by National Institutes of Health (NIH), National Institute of Diabetes and Digestive and Kidney Diseases (NIDDK) grants R01 DK076077, R01 DK087635 and R01 DK105821 (to K.S.). The ARIC study has been funded in whole or in part with federal funds from the National Heart, Lung, and Blood Institute (NHLBI), NIH, Department of Health and Human Services, under contract numbers 75N92022D00001, 75N92022D00002, 75N92022D00003, 75N92022D00004 and 75N92022D00005. SomaLogic conducted SomaScan assays in exchange for use of ARIC data, which was supported in part by NIH/NHLBI grants R01 HL134320 and R01 DK124399. T.V. was a summer research student from Riverdale Country School at the time of this work; her listed affiliation reflects where the research was conducted. We acknowledge the PMBB for providing data and thank the patients of Penn Medicine who consented to participate in this research program. We would also like to thank the PMBB team and Regeneron Genetics Center for providing genetic variant data for analysis. The PMBB is approved under IRB protocol 813913 and is supported by the Perelman School of Medicine at the University of Pennsylvania, by a gift from the Smilow family and by the National Center for Advancing Translational Sciences of the NIH under Clinical and Translational Science Awards number UL1TR001878. PMBB members include D. J. Rader, M. D. Ritchie, J. Weaver, N. Naseer, G. Sirugo, A. Poindexter, J. Dever, A. Harvey, S. Linn, N. Srivastava, M. Livingstone, F. Vadivieso, S. DerOhannessian, T. Tran, J. Stephanowski, S. Santos, N. Haubein, J. Dunn, A. Verma, C. Morse Kripke, M. Risman, R. Judy, C. Wollac, S. S. Verma, S. Damrauer, Y. Bradford, S. Dudek, T. Drivas and Z. Rodriguez. The authors thank the staff and participants of the ARIC study for their important contributions. This work was conducted in collaboration with Novartis and the Susztak laboratory.
Author information
Authors and Affiliations
Renal, Electrolyte and Hypertension Division, Department of Medicine, Perelman School of Medicine, University of Pennsylvania, Philadelphia, PA, USA
Chenyu Li, Ghazal Quinn, Amin Abedini, Minyan Zhu, Tanya Verma, Samer Mohandes & Katalin Susztak
Penn/CHOP Kidney Innovation Center, University of Pennsylvania, Philadelphia, PA, USA
Chenyu Li, Ghazal Quinn, Amin Abedini, Minyan Zhu, Tanya Verma, Samer Mohandes & Katalin Susztak
Institute for Diabetes, Obesity, and Metabolism, Perelman School of Medicine, University of Pennsylvania, Philadelphia, PA, USA
Chenyu Li, Ghazal Quinn, Amin Abedini, Minyan Zhu, Tanya Verma, Samer Mohandes & Katalin Susztak
Department of Genetics, Perelman School of Medicine, University of Pennsylvania, Philadelphia, PA, USA
Chenyu Li, Ghazal Quinn, Amin Abedini, Minyan Zhu, Tanya Verma, Samer Mohandes, Marylyn Ritchie, Daniel J. Rader & Katalin Susztak
Novartis Biomedical Research, Novartis Campus, Basel, Switzerland
Shola M. Richards & Vesna Barros
Novartis Biomedical Research, Cambridge, MA, USA
Rebecca Pitts, Xiazi Qiu, Taehwan Shin, Joseph J. Loureiro, Nancy Finkel, William F. Dietrich & Lori L. Jennings
Division of Precision Medicine, New York University School of Medicine, New York, NY, USA
Aditya Surapaneni, Josef Coresh & Morgan E. Grams
Cardiorenal Division, R&ED, Novo Nordisk, Malov, Denmark
Anil Karihaloo
Department of Biostatistics, Epidemiology and Informatics, University of Pennsylvania, Philadelphia, PA, USA
Hongzhe Li
Division of Translational Medicine and Human Genetics, Department of Medicine, Perelman School of Medicine, University of Pennsylvania, Philadelphia, PA, USA
Anurag Verma
Division of Informatics, Department of Biostatistics, Epidemiology and Informatics, University of Pennsylvania, Philadelphia, PA, USA
Anurag Verma
Division for Biomedical Informatics and AI, Department of Public Health Sciences, Medical University of South Carolina, Charleston, SC, USA
Marylyn Ritchie
Department of Medicine, Perelman School of Medicine, University of Pennsylvania, Philadelphia, PA, USA
Daniel J. Rader
Concept and design: K.S., L.L.J., W.F.D., N.F., J.J.L., S.M.R., A.A., G.Q. and C.L. Paper drafting: C.L. and K.S. Statistical analysis: C.L., S.M.R., S.M., V.B., H.L. and K.S. Data collection and interpretation: S.M.R., N.F., R.P., X.Q., T.S., M.Z., T.V., L.L.J., A.V., M.R., D.J.R., C.L. and K.S. External validation: A.S., J.C., M.E.G. and A.K. Paper revision: C.L., S.M.R., V.B., R.P., X.Q., T.S., N.F., J.J.L. and K.S. All authors critically reviewed and agreed to the submission of the final paper.
Corresponding author Ethics declarationsCompeting interestsThe laboratory of K.S., including C.L., G.Q., A.A., M.Z., T.V. and S.M., receives research support from Gilead, Novo Nordisk, Novartis, GlaxoSmithKline, Boehringer Ingelheim, Regeneron, Genentech and Calico Life Sciences. J.C. is a scientific advisor receiving fees from SomaLogic and Healthy.io. S.M.R., R.P., V.B., X.Q., T.S., J.J.L., N.F., W.F.D. and L.L.J. are employees and stockholders of Novartis. A.K. is an employee of Novo Nordisk. The remaining authors declare no competing interests.
Peer reviewPeer review informationNature Medicine thanks Daniel Gale, Yan Sun and the other, anonymous, reviewer(s) for their contribution to the peer review of this work. Peer reviewer reports are available. Primary Handling Editor: Michael Basson, in collaboration with the Nature Medicine team.
Additional informationPublisher’s note Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.