
Article Open access10 March 2026
Esperanza Elías-Cabot, Sara Romero-Martín, José Luis Raya-Povedano, Alejandro Rodríguez-Ruiz & Marina Álvarez-Benito
Abstract
Artificial intelligence (AI) systems have been demonstrated to improve the accuracy of screening mammograms. Here this prospective, paired, noninferiority clinical trial evaluated whether AI could safely reduce workload by excluding low-risk exams from radiologist reading. Between March 2022 and January 2024, 31,301 women were included in the trial and underwent routine mammograms. Two reading strategies were applied in parallel: standard double-blind reading and partially autonomous AI-supported screening, where cases classified by AI as low risk were assessed as normal and the rest were double read with AI support. The primary outcomes were radiologist workload, cancer detection rate and recall rate. In the AI strategy, radiologist workload was 63.6% lower; the cancer detection rate was 15.2% higher (95% confidence interval 6.6%, 24.4%), increasing from 6.3 of 1,000 to 7.3 of 1,000, P < 0.001; and the recall rate was not noninferior and was 14.8% higher (95% confidence interval 9.0%, 20.6%). Subanalyses by modality highlighted a similar workload reduction in digital mammography (−62.1%) and digital breast tomosynthesis (−65.5%). However, in digital mammography, the cancer detection rate increased by 1.6 of 1,000 and the recall rate by 1.3%, while both remained stable in digital breast tomosynthesis. These results demonstrate the feasibility of a partially automated AI workflow in breast cancer screening, avoiding human reading of studies classified as low risk. ClinicalTrials.gov: NCT04849776.
Similar content being viewed by others
Article Open access10 March 2026

Article Open access07 January 2025

Article Open access07 March 2025
Main
Breast cancer screening programs with digital mammography (DM) have been established for decades, with proven benefits in reducing breast cancer-related mortality1. Still, screening programs face clinical challenges in terms of both missed cancers and false positive recalls2. These can be improved by the use of digital breast tomosynthesis (DBT) and/or double reading, but at the cost of an ever-increasing reading workload3. In light of population aging, recommendations for starting screening at younger ages and the increasing shortage of qualified radiologists in some areas, alternative approaches using artificial intelligence (AI) have been proposed to maintain or even improve screening quality without increasing radiologists’ workload4.
Over the last few years, several retrospective studies have shown that radiologists improved their cancer detection accuracy when using an AI system as concurrent reading support5,6,7,8. Furthermore, AI systems implemented as a stand-alone solution have been found to achieve human-like performance in simulated retrospective studies9,10,11,12, classifying exams according to cancer probability and resulting in simulations of AI-based breast cancer strategies that suggested improved clinical outcomes4,13,14,15,16,17. Several prospective studies have been designed to overcome the limitations of retrospective studies and to address and quantify the real-life impact of AI in breast cancer screening in which radiologists interact with AI. The first results from these studies have recently been published18,19,20,21,22. Some of these studies showed that, compared to standard double reading, AI-supported double reading21 or using AI as a third reader22 can improve mammography screening metrics. Other prospective studies18,19,20 demonstrate that the use of AI to replace a reader in double-reading screening with DM is safe and effective, leading to a significantly reduced workload. A common trend in the outcomes of many retrospective and prospective studies is the high performance accuracy of AI in identifying large numbers of low-risk breast screening exams with a very high negative predictive value. As a consequence, the radiologist’s reading could be completely avoided for such low-AI-risk exams, thus effectively introducing autonomous AI reading and partially automated screening. This has been simulated retrospectively in two large-scale studies with DBT population screening, demonstrating a negative predictive value of AI of 99.91% (ref. 17) and 99.97% (ref. 4) for the screening exams classified by AI as lowest risk.
The Artificial Intelligence in Breast Cancer Screening Program in Córdoba (AITIC) clinical trial is a prospective, paired, noninferiority, accuracy study (NCT04849776) that aims to confirm prospectively the results obtained in these studies by evaluating whether an AI system can be used to omit human reading completely in a large proportion of the screening exams classified as low risk for cancer, with noninferior results in the cancer detection rate (CDR) and recall rate (RR) compared to the standard strategy. This trial has a paired design, where each mammogram (either DM or DBT, depending on resource availability) has two reading strategies: the standard of care (double reading without AI support) and the AI-based intervention strategy (double reading with AI support if AI risk is higher than a predetermined threshold; alternatively, the exam will be automatically considered as normal).
This paired, noninferiority trial prospectively evaluates the use of an AI system in a real-world population-based breast cancer screening program for the autonomous and automated reading of a large number of examinations without requiring radiologist involvement in terms of workload, CDR and RR. This study includes DBT, which is increasingly used and recommended in screening programs23.
Results
Women included in the trial and data characteristics
From 15 March 2022 to 11 January 2024, 33,171 consecutive women participated in the breast cancer screening program. On arrival at the screening site, they were informed of the study and offered the opportunity to participate on a voluntary basis. A total of 31,856 women agreed to participate in this trial (AITIC, ClinicalTrials.gov ID NCT04949776) and signed the informed consent form (the patient information sheet and informed consent form are included in Zenodo24. Of these, 555 women (1.7%) were excluded because of the different reasons explained in Fig. 1, resulting in 31,301 women being included (17,333 DM and 13,968 DBT).
Fig. 1: Study profile.
Flowchart describing the inclusion of the study population, the women excluded and the different reasons for exclusion.
Table 1 shows the characteristics of the participants included in the study and data recorded. The median age was 59 years old (interquartile range 64 to 54 years). This was the first screening for 11.3% of the participants. In terms of breast density, the distribution was as follows: A: 20.6%; B: 46.5%; C: 27.6%; and D: 5.3%. There were 252 screen-detected cancers between both strategies.
Table 1 Characteristics of the patients included in the study and data recorded Primary outcomesTable 2 shows the results obtained for the main outcomes (workload, CDR and RR) for the entire sample. The calculation is based on the patients who were included in the study. Extended Data Figs. 1 and 2 show the noninferiority plots for CDR and RR.
Table 2 Screening performance outcomes in total, for each strategyThe standard strategy resulted in 62,602 radiologists screening readings and 198 screen-detected cancers. The CDR and RR were 6.3 of 1,000 (95% confidence interval (CI) 5.5 of 1,000, 7.2 of 1,000) and 4.8% (95% CI 4.6%, 5.0%), respectively. The AI strategy involved reading only 36.4% (11,384) of all screening exams, which resulted in 22,768 radiologist readings and 228 screen-detected cancers. The remaining studies (19,917 women with low-AI-risk exams) were classified as normal in this strategy without human reading. The CDR and RR were 7.3 of 1,000 (95% CI 6.3 of 1,000, 8.2 of 1,000) and 5.5% (95% CI 5.3%, 5.8%), respectively. For more details, see Table 2.
When comparing the results between the standard strategy and the AI strategy, there was a reduction in the workload of −63.6% (95% CI −64.2, −63.1 (−39,834 readings)); an increase in the CDR of 15.2% (95% CI 6.6%, 24.4%), which is an absolute difference of 1.0 of 1,000 (95% CI 0.4 of 1,000, 1.5 of 1,000; P < 0.001); and an increase in the RR of 14.8% (95% CI 9.0%, 20.6%), which is an absolute difference of 0.7% (95% CI 0.4%, 1.0%). Therefore, the CDR in the AI strategy was noninferior and statistically superior to that of the standard strategy while RR was not noninferior to the standard strategy. See Table 2 for more details.
Secondary outcomes and screening performance outcomes by modalityThe positive predictive value (PPV) was similar in both strategies: 13.19% (95% CI 11.50%, 14.90%) in the standard strategy and 13.23% (95% CI 11.63%, 14.83%) in the AI strategy (absolute difference of 0.04% (95% CI 2.3%, 2.4%)).
The false positive rate (FPR) was higher in the AI strategy (4.8% (95% CI 4.5%, 5.0%)) than in the standard strategy (4.2% (95% CI 4.0%, 4.4%)), with an absolute difference of 0.6% (95% CI 0.3%, 0.9%).
On analyzing the results according to imaging modality (DM and DBT), in addition to a significant workload reduction (−62.1% and −65.5%, respectively), the results in terms of CDR and RR were different in each modality. In the AI strategy with DM, the CDR was 33.7% (95% CI 19.8%, 50.5%), higher than in the standard strategy (absolute difference of 1.6 of 1,000; 95% CI 0.9%, 2.4%). In addition, the RR was 28.2% (95% CI 20.2%, 36.3%), higher than in the standard strategy (absolute difference of 1.3%; 95% CI 1.0%, 1.8%). In the AI strategy with DBT, the CDR and RR were similar to the standard strategy. A relative difference of 0.9% (95% CI −10%, 11.8%), which is an absolute difference of 0.1 of 1,000 in the CDR, and a relative difference of −2.4% (95% CI −10.6%, 5.7%), which is an absolute difference of −0.1% in the RR, were observed. See Table 3 and Extended Data Figs. 1 and 2 for more details.
Table 3 Screening performance outcomes for each strategy in the subgroup analysis by modality (DM and DBT separately) Characteristics of the cancers detectedIn total, 252 cancers were detected (189 invasive and 63 carcinomas in situ): 198 via the standard strategy (158 invasive (79.8%) and 40 carcinomas in situ (20.2%)) and 228 cancers via the AI strategy (174 invasive (76.3%) and 54 carcinomas in situ (23.6%)). Thus, the AI strategy detected 10.1% more invasive carcinomas (95% CI 1.8%, 18.7%); 35% more carcinomas in situ (95% CI 7.6%, 60.3%); and a higher proportion of grade I invasive carcinomas (30.2% (95% CI 13.0%, 48.3%)), T1 (13.5% (95% CI 3.2%, 24.1%)) and N0 invasive carcinomas (15.6% (95% CI 6.1%, 25.4%)) than the standard strategy, without observing significant differences in the molecular profile.
With DM, a total of 122 cancers were diagnosed (89 invasive and 33 carcinomas in situ). The AI strategy detected more invasive carcinomas, carcinomas in situ, grade I invasive carcinomas, and T1 and N0 invasive carcinomas than the standard strategy. The AI-based strategy showed an increase in the detection of luminal A tumors (44.4% (95% CI −6.5%, 88.3%)).
With DBT, a total of 130 cancers were diagnosed (100 invasive and 30 carcinomas in situ). No differences were found between the histopathological characteristics of cancers detected in both strategies in this modality.
Table 4 presents the overall characteristics of cancers detected by each strategy and modality.
Table 4 Histological characteristics of detected cancers Cancers missed for each strategyTwenty-four cancers were detected only by the standard strategy, of which 23 of 24 (96%) were only recalled by one of the two radiologists. Of these, 11 cancers were identified by AI as low risk (scores 1–7) and were automatically labeled as normal in the AI strategy (9 of these low-risk cancers were DBT exams and 2 were DM exams). The other 13 cancers received scores of 8–10, but were not identified as cancer nor recalled by the human readers involved as part of the AI strategy. On the other hand, there were 54 screen-detected cancers only detected by the AI strategy, most of them (34 of 54 (63%)) were recalled by both radiologists.
Extended Data Table 1 shows the scores assigned by the AI system to the cancers detected. For more details, see Extended Data Tables 2–5.
SafetyNo adverse events were reported in this study.
DiscussionThe results from this prospective, paired, noninferiority clinical trial demonstrate that in a prospective setting it is safe to use AI to identify screening exams that can be automatically labeled as normal and avoid radiologist reading in a population screening program that uses both DM and DBT exams.
Overall, the AI strategy, which entailed radiologists reading only 36% of total exams, led to a 15.2% increase in the CDR from 6.3 of 1,000 to 7.3 of 1,000 and a 14.8% increase in the RR from 4.8% to 5.5%, while maintaining the PPV. Therefore, the CDR in the AI strategy was considered noninferior and statistically superior to that of the standard strategy and RR was not considered noninferior. Notably, cancer detection increased for both carcinomas in situ and invasive carcinomas, as the AI strategy detected 14 more carcinomas in situ and 16 more invasive carcinomas than the standard strategy. These results confirm that the high AI accuracy in identifying screening exams that are very likely normal together with a high cancer detection accuracy that provides effective concurrent decision support could be used to safely implement a partially autonomous AI screening workflow.
Interestingly, these prospective results are in line with early retrospective simulations investigating the use of such an AI strategy in which exams that were highly likely to be normal could be automatically assessed as normal4,25,26,27. All found a potential safe reduction (no decrease in sensitivity or specificity) in the screening workload of 70%, 63%, 72% and 63%, respectively, in four distinct screening populations across Europe. Of note, our results include DBT screening, which is known to take at least twice as long to read as DM screening exams3. Therefore, the suggested safe workload reduction of 64% would be most impactful on radiologists’ time in DBT screening programs.
Other prospective studies18,19,20,21,22 have investigated the introduction of AI in DM-based breast cancer screening. PRAIM21 was a multicenter study that compared the performance of AI-supported double reading (voluntarily chosen by radiologists) to standard double reading. It found an increase in the CDR of 17.6% in the AI group. It also analyzed a fictitious scenario in which studies triaged as normal by AI were not read by the radiologists. It showed that, in addition to a 56.7% reduction in the workload, the CDR remained 16.7% higher and the RR was statistically superior and 15% lower in the AI group than in the control group. On the other hand, MASAI18,19 and ScreenTrustCAD20 introduced AI within the scope of reducing the workload. These studies were focused on totally or partially replacing the double reading of screening exams by single reading. Interestingly, MASAI, which maintained double reading of the most suspicious exams with AI support (similar to our approach for the high-AI-risk exams) also reported an increase in the CDR (+20%) and a workload reduction of 44% without a negative impact on false positives. In a different design, ScreenTrustCAD investigated a direct replacement of the second radiologist reading by AI for all cases, without any AI support for the single reading, and found an impact on the CDR of +4%, a 50% workload reduction and a 21% increase in false positives prior to arbitration.
The benefits of double reading of high-AI-risk exams with AI support and the replacement of double reading by single reading in low-risk exams are also supported by the results of the observational study after AI implementation in the Capital Region of Denmark (increase in the CDR of 17%, reduction in false positives of 32% and reduction in workload of 33%)28 using the same AI system as MASAI and AITIC.
Our results show a 14.8% increase in RR in the AI strategy. The study conducted by Ng et al.22 showed 0.16–0.30% additional recalls. Two other prospective studies (MASAI and ScreenTrustCAD)18,19,20 also showed an increase in the number of abnormal studies in the AI strategy, but both studies, unlike ours, held consensus meetings to assess these abnormal studies, which resulted in fewer women being recalled. The design of our screening program, which is without consensus, is associated with more recalls and may be more difficult to compare with studies conducted in screening programs with consensus. In addition, there are other factors in our study design that may also increase recalls, such as the fact that, in the AI reading strategy, AI-supported human double reading of cases with scores of 8–10 is performed, in contrast to other prospective studies where AI-supported double reading was only performed in cases with a score of 1018,19 or the AI information was not available during human reading or was on demand20,21. Nevertheless, it is not possible to know the false negatives of nonrecalled women following the consensus of these studies.
Although the main objective of our study was not to compare results between DM and DBT, when looking at the impact of AI by modality (DM and DBT), in addition to the significant workload reduction (−62.1% and −65.5%, respectively), the results in terms of CDR and RR were different in each modality.
In the AI strategy with DM, the CDR was 33.7% higher than in the standard strategy and the RR was 28.2% higher than in the standard strategy. In the AI strategy with DBT, the CDR and RR were similar to the standard strategy. A relative difference of 0.9% in the CDR and a relative difference of −2.4% in the RR were observed. The AI strategy with DM shows a substantial increase in the detection rate of both invasive carcinomas and carcinomas in situ as well as in the percentage of G1, T1 and N0 invasive carcinomas compared to the standard strategy. In addition, an increase in the detection of luminal A tumors was observed. MASAI19 shows similar results in terms of increased detection of invasive carcinomas in the AI strategy (24.4%), especially due to small tumors and those with lymph-node negative results, which should entail a positive impact on patient prognosis and lower aggressiveness and treatment costs. The increase in detection of carcinomas in situ in the AI with DM strategy is slightly higher than that observed in MASAI (66.7% and 51.1%, respectively). However, these authors showed greater carcinoma in situ detection (57.9%) in the AI strategy in their previous results18 and hypothesize that the radiologists’ experience in working with AI and the feedback from the care of recalled women has probably influenced the later data. As in the MASAI series19, a significant percentage of the carcinomas in situ detected by the AI strategy (9 of 12) were intermediate or high grade. The RR was not shown to be noninferior. However, taking into account the increase in detection of 1.6 of 1,000, it can be claimed that this increase in recalls of 1.3% is acceptable.
In tomosynthesis, no impact on the detection or RR was demonstrated. Noninferiority was not able to be demonstrated, probably due to the sample size and threshold of the objective (noninferiority of −5%). The high screening performance with DBT in this program even without AI should be noted, likely due to the extensive experience of the radiologists reading DBT (median of 5 years of experience reading DBT screening tests). AI systems may perform differently for DBT than for DM and may be more accurate for DM when compared to the average radiologist performance9. However, the results in the detection and RRs achieved by the AI-DBT strategy (8.1 of 1,000 and 4.6%, respectively) are superior to those achieved by the AI-DM strategy (6.6 of 1,000 and 6.2%, respectively). Nevertheless, the significant reduction in workload achieved (65.5%) is the most important aspect in favor of incorporating AI in screening programs using DBT. Thus, the strengths of this work are that it is a prospective, paired study that aimed to provide information on the use of AI in a real-life screening setting with DBT, allowing for assessing nonhuman reading of a significant percentage of studies classified as low risk by AI.
One of the major barriers for clinical adoption of our proposed strategy is certainly the ethical challenges arising from having most screening mammograms read only by AI and automatically assessed as normal without radiologists’ involvement. This could lead to cancers being missed (11 cancers in our study), but, as demonstrated in our study, there are more cancers being missed by radiologists when not using AI for decision support in screening (54 cancers in our study). Furthermore, among the cancers with a score <8 missed by the AI strategy, most (91%) were only detected by one of the two readers in the standard strategy, indicating their potential subtle appearance. Nevertheless, the ethical implications of relying on AI for initial reads versus radiologists must be carefully considered to ensure patient safety and trust in screening processes. If there is no radiologist review of a significant proportion of exams, additional screening quality assurance processes such as automated mammography image quality control and continuous postmarket surveillance of AI performance are necessary steps before implementing such an autonomous AI screening workflow for the very likely normal screening exams.
Our study has the limitation of being a single-site investigation with a screening workflow of double reading without arbitration by expert radiologists in breast imaging and breast screening who have several years of experience in the use of AI. Further studies are needed to validate these findings in other settings. In addition, mammography images were acquired using four devices, all from a single vendor. Results were also obtained with one breast AI system; extrapolation to other AI systems must be evaluated further, given the significant differences in AI accuracy, robustness and user interface across commercial systems. The sample size was calculated to include all studies for DM and DBT together, reflecting our usual practice, and not for DM and DBT separately. Consequently, the results in DBT have been less consistent. Finally, as it is a paired design, it is not possible to know the interval carcinoma rate for each strategy.
Further research is needed on the legal and safety considerations involved in preventing the reading of low-risk studies, as well as on the performance of AI in DBT.
In conclusion, AI triage and AI-supported screening that excludes low-risk mammograms from radiologist reading could be a safe and effective screening strategy that would allow for a substantial reduction in reading workload in breast cancer screening programs without negatively affecting cancer detection or the PPV of the recalled studies.
Fig. 2: Flowchart describing the paired study design of the AITIC trial indicating the primary and secondary endpoints.
Each participant’s exam was read using two strategies: standard strategy (double human reading without AI support) and the AI-supported strategy, with double human reading with AI support only for cases classified by the AI system with a score of 8 to 10. In this strategy, cases classified by the AI system with a score of 1 to 7 were considered low risk (approximately 70% of the studies) and would be automatically classified as normal.
MethodsEthics approvalAll patients provided written informed consent before enrollment. This prospective clinical trial was compliant with the Health Insurance Portability and Accountability Act and the design was preregistered at https://classic.clinicaltrials.gov/ct2/show/NCT04949776. In March 2021, the study protocol and the informed consent received a favorable ruling from the Institutional Review Board (IRB) at Reina Sofía University Hospital of Córdoba Research Ethics Committee (IRB No 4932).
Clinical trial design and patientsThis is a prospective paired noninferiority experimental study.
The individual deidentified participant dataset, data dictionary defining each field and study protocol are publicly accessible via Zenodo at https://doi.org/10.5281/zenodo.17625633 (ref. 24).
This clinical trial was carried out in the Córdoba Breast Cancer Screening Unit, Córdoba, Spain, as part of the Andalusian screening program in Spain. Women participants aged 50 to 71 years old (including women who reach that age in the year of appointment) are invited to participate in the screening program at 2-year intervals (ClinicalTrials.gov ID: NTC04949776).
All women participating in the screening program from 15 March 2022 to 11 January 2024 were eligible and invited to participate in this clinical trial. They were asked to sign an informed consent form. For screening purposes, sex and age data were defined by reference to the population census. All women aged 50 to 71 years, regardless of socioeconomic status, were considered.
After signing informed consent, women were excluded from the clinical trial if they had symptoms or signs of suspected breast cancer; if they had breast prostheses; or if their images could not be processed by the AI system (for example, because of the presence of visible breast implants or because of erroneous PACS image transfer).
To investigate the hypothesis and to account for ethical aspects, this clinical trial had a paired design (Fig. 2). Each participant received both reading strategies:
The standard of care: double human reading without AI support.
The AI strategy: double human reading with AI support (by two additional radiologists) only for cases classified by the AI system with a score of 8 to 10 (approximately 30% of the studies most likely to have cancer). Cases classified by the AI system with a score of 1 to 7 were considered low risk (approximately 70% of the studies) and automatically classified as normal.
Each participant underwent a standard four-view DM or DBT exam based on resource availability, according to the standard of care. There were no rules governing imaging modality selection. The assignment of women to the different types of mammography equipment was not changed for this study. Images were acquired using four devices: three DM devices (Lorad Selenia, Hologic) and one DBT device (3Dimensions, Hologic).
All exams were independently double read without consensus or arbitration. Screening readings were randomly assigned according to availability of the nine dedicated breast radiologists in the hospital’s radiology department (3 to 21 years of experience in breast cancer screening), regardless of whether they were readings with or without AI assistance and regardless of the reader’s experience. The mammogram offered to the reader was the oldest one performed that had not been read by that same reader. In each reading, the radiologist noted the breast density (measured subjectively according to the BI-RADS classification31), the presence of findings and their location in the breast, the assigned BI-RADS category (1 to 5) and the decision whether or not to recall. Women were recalled if any reader decided to recall.
Women received a single result (recall or no recall) once all readings were completed. If recalled, additional exams (special mammography views, tomosynthesis, contrast-enhanced mammography and/or ultrasonography) and, when required, biopsy were performed at the women’s referral unit. The screening program did not include ultrasonography or other additional examinations for women with high breast density. No adverse events were reported in this study.
This clinical trial used a commercially available AI system for breast cancer detection, Transpara (version 1.7 ScreenPoint Medical), which had been used for previous research studies by the same group of radiologists involved in this study.
The AI system’s performance, generalizability and clinical utility has been previously investigated in over 30 other peer-reviewed scientific publications. Most significantly, researchers have shown that this system can achieve stand-alone breast cancer detection performances comparable to radiologists10, that radiologists become more accurate when using such an AI system for concurrent decision support reading mammograms, leading to higher screening CDRs18,19, and that it is safe and effective to implement in breast cancer screening to replace the need for the double reading18,19.
The AI system automatically detects and classifies regions suspicious of breast cancer in radiological images that adhere to the input compatibility criteria of the system and with standard Digital Imaging and Communications in Medicine. This version of the system can be used in combination with DM and DBT images acquired with systems from all major mammography equipment manufacturers (Siemens Healthineers, Hologic, General Electric, Giotto, Planmed, Fujifilm). The system can analyze any number of standard (craniocaudal, mediolateral oblique) and nonstandard (exaggerated, lateral) views per breast. Images of women with breast implants are not compatible unless the implant has been displaced during compression.
The AI system outputs per exam are (1) explainable image overlay markings for regions suspicious of breast cancer per view, alongside a suspiciousness score per region between 1 and 100 and an indication of the finding morphological type, and (2) an exam-based indication of the risk that the exam contains a visible lesion suspicious of breast cancer that requires additional follow-up (three categories: 1–7 low risk, 8–9 intermediate risk, 10 elevated risk).
The system uses deep convolutional neural networks as backbone algorithms to analyze DM and DBT images and detect lesions suspicious for breast cancer. Convolutional neural networks are state-of-the-art machine learning tools for image classification. Different convolutional neural networks are trained with different architectures and training datasets for malignancy detection of different morphological subtypes and taking a variable number of images as input.
The AI system was developed using a large-scale heterogeneous curated database of more than 15 million two-dimensional and three-dimensional X-ray breast images acquired in real-world breast cancer screening programs and diagnostic clinics. The data originated from over 15 sites across 10 countries across North America, Europe and Asia, including over 15,000 malignant cases (pathology proven). Negative cases were confirmed through multiple years of follow-up.
The AI system software version and operating points used in the present evaluation were established before the clinical trial. The choice of a score of 7 as the cutoff point for indicating that human reading was not necessary was based on previous studies conducted by this research team and other published research4,5,25. None of the clinical data used in this clinical trial was used in any aspect of algorithm development.
The human readers in the AI strategy had available the information provided by the AI system (marks according to the type of lesion with a score according to the probability of having cancer and global score of the study), but the radiologist was the person who ultimately made the decision to recall a woman or not.
Data on the procedures, outcomes, participants and detected cancers were retrieved from the medical records. Race and ethnicity were not individually recorded, but the majority of the target population was Caucasian.
SafetyThe study protocol was reviewed and approved by the institutional ethics committee, which determined that the study posed minimal risk to participants.
Safety monitoring focused on nonphysical risks potentially associated with the use of AI-based image analysis, including data integrity, patient privacy and the potential for diagnostic error. All mammographic images were fully anonymized prior to analysis, and data handling complied with applicable data protection regulations. Any unexpected incidents related to data processing, algorithm malfunction or breaches of confidentiality were predefined as reportable events and were monitored throughout the study period.
OutcomesThe primary outcomes were workload, CDR and RR calculated globally. Workload was computed as the absolute number of radiologist readings in each strategy. A cancer detected during the diagnostic work-up after a recall and confirmed by pathology was considered the gold standard.
The CDR was estimated as the number of cancers detected per 1,000 screening examinations for each strategy. Cases that were not recalled or that were recalled with malignancy not demonstrated on complementary imaging tests or biopsies were considered negative. Follow-up lasted for a minimum of 180 days after screening mammography. For each case, a strategy was positive if one or both readings included referral.
As secondary objectives, PPV of recalls and FPR were assessed globally. A subgroup analysis by modality was performed and workload, CDR, RR, PPV of recalls and FPR were calculated for DM and for DBT separately. The PPV of recalls was calculated as the number of cancers diagnosed per 100 women recalled for each strategy. The FPV was calculated as the number of women recalled who were not finally diagnosed with cancer for each strategy per 100 negative screening examinations.
Sample sizeFor the sample size calculation, the method by Connor32 for paired studies was used, taking into account the sample size necessary to demonstrate noninferiority (5% margin) in terms of sensitivity for cancer detection via McNemar’s z-test (one-sided). An alpha (type I error estimate) value of 5% and a beta (type II error estimate) of 20% (representing a power for the study of 80%, 1 − beta) was used. In addition, we used the parameters of an assumed cancer incidence of 6 of 1,000 (value of cancer incidence when AI is not used in screening based on initial experience and more recently verified in ref. 5). We also calculated how many cancers would be detected by the standard strategy alone and how many would be detected by the AI strategy proposed in this study using the database and methods used in ref. 4 but using Transpara software version 1.7.0 (as opposed ref. 4, which used version 1.6.0, which has inferior algorithms). In the results obtained, 1.8% of cancers were detected using the standard strategy but not using the proposed strategy with AI, while the opposite occurred in 9% of cancers. With these parameters, the sample size needed was defined as 27,000 women in order to demonstrate the noninferiority of the AI strategy over the standard strategy globally, including both DM and tomosynthesis studies.
The three coprimary endpoints were CDR, RR and workload. It is worth mentioning that no formal multiplicity correction was applied between them for the sample size calculation. Readers are encouraged to interpret the findings in the context of these coprimary outcomes.
Statistical analysesThe hypothesis raised in this work is that a partially autonomous AI-supported screening strategy (that is, no human reading in cases with low-AI risk (AI strategy)) would allow for reducing the workload of a screening program that uses both DM and DBT exams compared to the standard-of-care reading without demonstrating inferiority in terms of the CDR and RR of the program. Descriptive analyses of the variables considered in the study were carried out in the complete dataset of women. Categorical variables were expressed as absolute and relative frequencies, while continuous variables were expressed as median and interquartile range. To compare the AI-assisted strategy with the standard-of-care strategy, we evaluated the CDR and RR, as well as secondary metrics including the PPV of recall and FPR. For all variables, percentages were calculated for each strategy, along with corresponding 95% CI. Differences between AI and standard strategies were assessed using Wald statistics with or without the Bonett–Laplace adjustment33, as appropriate.
For the primary outcomes (CDR and RR), we conducted noninferiority testing using the following approach: reference values for the CDR and RR were identified from previously published noninferiority studies and screening benchmarks in the literature. Based on clinical consensus, a 5% relative tolerance margin over those values was considered acceptable for declaring noninferiority. For the CDR (where higher values are better), a noninferiority margin was defined as a 5% relative reduction from the reference value. For the RR (where lower values are generally preferred), the noninferiority margin was set as a 5% relative increase from the reference value.
These thresholds were used as the minimum acceptable performance levels for the AI-assisted strategy. One-sided 97.5% CIs were established for the difference in proportions (AI standard) for the CDR and RR. These were derived from two-sided 95% CIs. Noninferiority was declared if the lower limit of the one-sided CI exceeded the predefined noninferiority margin. Visual representation of these thresholds is shown in Extended Data Figs. 1 and 2.
If noninferiority was demonstrated for a given outcome, we proceeded to test for superiority using a paired exact McNemar test. For PPV and FPR, noninferiority was not formally tested. Instead, standard two-sided hypothesis tests were conducted: a z-test was used for comparing PPV (assuming independent data) and a paired McNemar test was used for comparing FPR.
Reporting summaryFurther information on research design is available in the Nature Portfolio Reporting Summary linked to this article.
Data availabilityIndividual deidentified participant dataset, data dictionary defining each field, protocol and informed consent form-patients information sheet are publicly accessible via Zenodo at https://doi.org/10.5281/zenodo.17625633 (ref. 24). Access ends 10 years following article publication.
Code availabilityThe code for training and developing the evaluated AI algorithm (Transpara version 1.7, ScreenPoint Medical) is part of a proprietary system. The commercial AI system is available for external research evaluation collaborations for researchers who provide a relevant and methodologically sound proposal. Proposals should be directed to alejandro.rodriguezruiz@screenpointmed.com and will be answered within 1 month. We provide a technical description of the backbone algorithm of the AI system in the Methods. The software used to perform the statistical analyses in this study was R, version 4.4.2. The main open-source libraries used for the analyses were: ggplot2 (for graphic purposes) and PropCIs (for confidence intervals estimation).
Caswell-Jin, J. L. et al. Analysis of breast cancer mortality in the US—1975 to 2019. JAMA 331, 233–241 (2024).
Lee, C. I. et al. National performance benchmarks for screening digital breast tomosynthesis: update from the Breast Cancer Surveillance Consortium. Radiology 307, e222499 (2023).
Romero Martín, S. et al. Prospective study aiming to compare 2D mammography and tomosynthesis + synthesized mammography in terms of cancer detection and recall. From double reading of 2D mammography to single reading of tomosynthesis. Eur. Radiol. 28, 2484–2491 (2018).
Raya-Povedano, J. L. et al. AI-based strategies to reduce workload in breast cancer screening with mammography and tomosynthesis: a retrospective evaluation. Radiology 300, 57–65 (2021).
Elías-Cabot, E., Romero-Martín, S., Raya-Povedano, J. L., Brehl, A. K. & Álvarez-Benito, M. Impact of real-life use of artificial intelligence as support for human reading in a population-based breast cancer screening program with mammography and tomosynthesis. Eur. Radiol. 34, 3958–3966 (2024).
van Winkel, S. L. et al. Impact of artificial intelligence support on accuracy and reading time in breast tomosynthesis image interpretation: a multi-reader multi-case study. Eur. Radiol. 31, 8682–8691 (2021).
Rodríguez-Ruiz, A. et al. Detection of breast cancer with mammography: effect of an artificial intelligence support system. Radiology 290, 305–314 (2019).
Pacilè, S. et al. Improving breast cancer detection accuracy of mammography with the concurrent use of an artificial intelligence tool. Radiol. Artif. Intell. 2, e190208 (2020).
Romero-Martín, S. et al. Stand-alone use of artificial intelligence for digital mammography and digital breast tomosynthesis screening: a retrospective evaluation. Radiology 302, 535–542 (2022).
Rodriguez-Ruiz, A. et al. Stand-alone artificial intelligence for breast cancer detection in mammography: comparison with 101 radiologists. J. Natl Cancer Inst. 111, 916–922 (2019).
Salim, M. et al. External evaluation of 3 commercial artificial intelligence algorithms for independent assessment of screening mammograms. JAMA Oncol. 6, 1581–1588 (2020).
McKinney, S. M. et al. International evaluation of an AI system for breast cancer screening. Nature 577, 89–94 (2020).
Larsen, M., Aglen, C. F., Hoff, S. R., Lund-Hanssen, H. & Hofvind, S. Possible strategies for use of artificial intelligence in screen-reading of mammograms, based on retrospective data from 122,969 screening examinations. Eur. Radiol. 32, 8238–8246 (2022).
Shoshan, Y. et al. Artificial intelligence for reducing workload in breast cancer screening with digital breast tomosynthesis. Radiology 303, 69–77 (2022).
Lång, K. et al. Identifying normal mammograms in a large screening population using artificial intelligence. Eur. Radiol. 31, 1687–1692 (2021).
Rodriguez-Ruiz, A. et al. Can we reduce the workload of mammographic screening by automatic identification of normal exams with artificial intelligence? A feasibility study. Eur. Radiol. 29, 4825–4832 (2019).
Dahlblom, V., Dustler, M., Tingberg, A. & Zackrisson, S. Breast cancer screening with digital breast tomosynthesis: comparison of different reading strategies implementing artificial intelligence. Eur. Radiol. 33, 3754–3765 (2023).
Lång, K. et al. Artificial intelligence—supported screen reading versus standard double reading in the Mammography Screening with Artificial Intelligence trial (MASAI): a clinical safety analysis of a randomised, controlled, non-inferiority, single-blinded, screening accuracy study. Lancet Oncol. 24, 936–944 (2023).
Hernström, V. et al. Screening performance and characteristics of breast cancer detected in the Mammography Screening with Artificial Intelligence trial (MASAI): a randomised, controlled, parallel-group, non-inferiority, single-blinded, screening accuracy study. Lancet Digit. Health 7, e175–e183 (2025).
Dembrower, K. et al. Artificial intelligence for breast cancer detection in screening mammography in Sweden: a prospective, population-based, paired-reader, non-inferiority study. Lancet Digit. Health 5, e703–e711 (2023).
Eisemann, N. et al. Nationwide real-world implementation of AI for cancer detection in population-based mammography screening. Nat. Med. 31, 917–924 (2025).
Ng, A. Y. et al. Prospective implementation of AI-assisted screen reading to improve early detection of breast cancer. Nat. Med. 29, 3044–3049 (2023).
European Guidelines for Breast Cancer Screening. European Commission (May 2023); https://cancer-screening-and-care.jrc.ec.europa.eu/en/ecibc/european-breast-cancer-guidelines?topic=65&usertype
Elías-Cabot, E., Romero-Martín, S., Raya-Povedano, J. L., Rodríguez-Ruiz, A. & Álvarez-Benito, M. Artificial Intelligence in Breast Cancer Screening Program in Cordoba (AITIC). Zenodo https://doi.org/10.5281/zenodo.17625633 (2025).
Lauritzen, A. D. et al. An artificial intelligence-based mammography screening protocol for breast cancer: outcome and radiologist workload. Radiology 304, 41–49 (2022).
Dembrower, K. et al. Effect of artificial intelligence-based triaging of breast cancer screening mammograms on cancer detection and radiologist workload: a retrospective simulation study. Lancet Digit. Health 2, e468–e474 (2020).
Leibig, C. et al. Combining the strengths of radiologists and AI for breast cancer screening: a retrospective analysis. Lancet Digit. Health 4, e507–e519 (2022).
Lauritzen, A. D. et al. Early indicators of the impact of using AI in mammography screening for breast cancer. Radiology 311, e232479 (2024).
Brierley, J., Gospodarowicz, M. D. & Wittekind, C. T. TNM Classification of Malignant Tumors International Union Against Cancer (Wiley, 2017).
WHO Classification of Tumours. Breast Tumours Vol. 2 (WHO, 2019).
D’Orsi C. J. et al. ACR BI-RADS Atlas, Breast Imaging Reporting and Data System (American College of Radiology, 2013).
Connor, R. J. Sample size for testing differences in proportions for the paired-sample design. Biometrics 43, 207–211 (1987).
Bonett, D. G. & Price, R. M. Adjusted Wald confidence interval for a difference of binomial proportions based on paired data. J. Educ. Behav. Stat. 37, 479–488 (2012).
We thank the IT department of the Hospital Universitario Reina Sofía (Córdoba), in particular to J. Antonio D. Osuna, C. G. Cortés, and R. M. Eslava for their support in the installation of the AI software in the real-world setting of the screening program. Thanks to A. I. C. Luna for her support in maintaining the study databases. Thanks to the women who participated in the study. Thanks to the SEDIM Foundation for its support of breast imaging research. In March 2023, EEC received for this study a grant (20,000 euros) from the SEDIM Foundation. The funder had no role in study design, data collection and analysis, decision to publish or preparation of the manuscript.
Author informationAuthors and AffiliationsMaimónides Biomedical Research Institute of Córdoba (IMIBIC), Córdoba, Spain
Esperanza Elías-Cabot, Sara Romero-Martín, José Luis Raya-Povedano & Marina Álvarez-Benito
Breast Cancer Unit, Department of Diagnostic Radiology, Reina Sofía University Hospital, Córdoba, Spain
Esperanza Elías-Cabot, Sara Romero-Martín, José Luis Raya-Povedano & Marina Álvarez-Benito
University of Córdoba, Córdoba, Spain
Esperanza Elías-Cabot, Sara Romero-Martín, José Luis Raya-Povedano & Marina Álvarez-Benito
ScreenPoint Medical BV, Nijmegen, The Netherlands
Alejandro Rodríguez-Ruiz
E.E.-C. and M.A.-B. conceptualized the study and wrote the study protocol with input from S.R.-M., J.L.R.-P. and A.R.-R. E.E.-C., S.R.-M. and J.L.R.-P. contributed to clinical data collection. E.E.-C. and M.A.-B. contributed to data analysis. S.R.-M., J.L.R.-P. and A.R.-R. contributed to data interpretation. E.E.-C. wrote the first draft of the report. All authors reviewed the paper and provided important intellectual content. All authors discussed the results and contributed to the final paper. All authors read and approved the manuscript.
Corresponding author Ethics declarationsCompeting interestsThe study was not funded by Industry. A.R.R. is an employee of ScreenPoint Medical. The other authors declare no competing interests.
Peer reviewPeer review informationNature Medicine thanks Ritse Mann and the other, anonymous, reviewer(s) for their contribution to the peer review of this work. Primary Handling Editors: Ming Yang and Lorenzo Righetto, in collaboration with the Nature Medicine team.
Additional informationPublisher’s note Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
Extended dataExtended Data Fig. 1 Non-inferiority plots for CDR.The horizontal axis is measured in percentage units of cancer detection. The dots of the horizontal lines represent the relative difference in CDR between both strategies. The whiskers represent the lower limits of the confidence intervals. The dashed vertical red line represents the non-inferiority threshold. Green indicates ‘non-inferiority achieved’ and blue indicates ‘non-inferiority not confirmed’. Relative differences in CDR between the AI strategy and standard strategy for the whole population (N = 31301): 15.2 (6.6, 24.4), for DM (N = 17333): 33.7 (19.8, 50.5), and for DBT (N = 13968): 0.9 (−10.0, 11.8). DM: digital mammography; DBT: digital breast tomosynthesis; CDR: cancer detection rate; AI: artificial intelligence.
Extended Data Fig. 2 Non-inferiority plots for RR.The horizontal axis is measured in percentage units of recall. The dots of the horizontal lines represent the relative difference in RR between both strategies. The whiskers represent the lower limits of the confidence intervals. The dashed vertical red line represents the non-inferiority threshold. Green indicates ‘non-inferiority achieved’ and blue indicates ‘non-inferiority not confirmed’. Relative difference in RR between the AI strategy and standard strategy for the whole population (N = 31301): 14.8 (9.0, 20.6), for DM (N = 17333): 28.2 (20.2, 36.3), and for DBT (N = 13968): −2.4 (−10.6, 5.7). DM: digital mammography; DBT: digital breast tomosynthesis; RR: recall rate; AI: artificial intelligence.
Extended Data Table 1 Score assigned by the artificial intelligence system to the detected cancers Extended Data Table 2 Cancers detected and not detected by both strategies (all studies) Extended Data Table 3 Cancers detected and not detected by both strategies (Digital Mammography studies) Extended Data Table 4 Cancers detected and not detected by both strategies (Digital Breast Tomosynthesis studies) Extended Data Table 5 Characteristics of cancers with scores 1–7 (low risk), not detected by the AI strategy Supplementary informationReporting Summary (download PDF )Rights and permissionsOpen Access This article is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License, which permits any non-commercial use, sharing, distribution and reproduction in any medium or format, as long as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons licence, and indicate if you modified the licensed material. You do not have permission under this licence to share adapted material derived from this article or parts of it. The images or other third party material in this article are included in the article’s Creative Commons licence, unless indicated otherwise in a credit line to the material. If material is not included in the article’s Creative Commons licence and your intended use is not permitted by statutory regulation or exceeds the permitted use, you will need to obtain permission directly from the copyright holder. To view a copy of this licence, visit http://creativecommons.org/licenses/by-nc-nd/4.0/.