
Article Open access11 July 2024
Abstract
Background
Next-generation sequencing (NGS) tests are integral to oncology care. To address the need for clinical and NGS data management, interpretation, and reporting, we developed iCatalog for the multi-institutional Individualized Cancer Therapy 2/Genomic Assessment Informs Novel Therapy Consortium (GAIN) pediatric precision oncology (PO) study.
MethodsWe designed iCatalog as a secure, web-based clinical decision support application that stores and integrates clinical, specimen, and molecular data from multiple sources at the patient level. The knowledge base (KB) and centralized patient/test database are intended to manage information for the 825 patients expected to enroll in the GAIN study. User permissions and access are controlled. Gene- and variant-level interpretation is facilitated through linked external resources and an internal KB that can be updated during application use. iCatalog generates editable, study-specific patient reports for each molecular test.
ResultsLaunched to support the GAIN study, iCatalog integrates genomic data from eight NGS platforms, generates 1002 clinical interpretation reports, and stores data for 1194 tests involving 777 patients with pediatric solid tumors across 133 diagnoses. The KB contains pediatric cancer-specific curations, authored by the research team, spanning 581 genes and 2659 variants (including 2146 single-nucleotide variants and insertions-deletions, 235 copy-number variants, 278 structural variants).
ConclusionsiCatalog is a robust tool designed and proven to support a PO study. It integrates clinical and genomic data to facilitate the clinical interpretation and reporting of variants identified through NGS testing while maintaining a pediatric-specific KB generated during the study. As a scalable, modular platform, iCatalog can accelerate clinical decision-making and elevate PO insights across studies.
Plain language summaryiCatalog is a web-based application developed to support a study evaluating the benefit of tumor profiling in the care of young cancer patients. iCatalog combines information used to determine patient care implications of tumor profiling. Patient (e.g., age, cancer diagnosis), sample (e.g., identification number, date), and genomic data (e.g., mutations, deletions, amplifications) are merged. iCatalog links genomic results to knowledge repositories and auto-generates reports, allowing experts to easily create patient-specific reports for use in the clinic. On this platform, 1002 reports were generated detailing the therapeutic and diagnostic implications of tumor profiling. Pediatric-specific knowledge on 581 genes was generated. As tumor sequencing becomes standard of care, tools reducing barriers to interpreting results for patient care become more important.
Similar content being viewed by others
Article Open access11 July 2024

Article Open access21 January 2026

Article Open access15 May 2026
Introduction
Genomic profiling of patient tumors using clinical next-generation sequencing (NGS) tests has significantly advanced precision oncology (PO), making it a fundamental component of the precision medicine workflow. A recognized gap between NGS results and the implementation of PO care lies in the variant interpretation and assessment of potential clinical impact. Many platforms and databases (DBs) exist to streamline the interpretation workflow and comply with industry standards for clinical genomics. These systems include knowledge bases (KBs) designed to store clinically relevant variant information and facilitate variant filtering, annotation, and interpretation1,2,3,4,5,6,7. In addition, PO research requires managing patient and genomic data and tracking the clinical significance of each alteration within the context of the patient’s presentation. Though the development of DBs, KBs, and data management tools has increased exponentially with the rise of genomic technology, only a limited number of systems fulfilled our needs for the clinical interpretation of molecular variants for a cohort of patients with rare cancers when launching a large pediatric oncology consortium study in 20158.
The Individualized Cancer Therapy 2/Genomic Assessment Informs Novel Therapy Consortium (GAIN; NCT02520713) study is a multi-institutional, prospective, observational cohort study led by the Dana-Farber Cancer Institute (DFCI), enrolling pediatric patients with solid tumors. It aimed to evaluate the clinical impact of molecular testing in patients with newly diagnosed high-risk and recurrent solid tumors. Targeted NGS was performed on patient tumor samples, test results were returned to oncologists, and data on patient treatment and outcomes were collected8. The study design included a clinical interpretation step where genomic variants in the tumor were assessed for functional impact and clinical significance in the diagnostic and therapeutic domains. We also included a hereditary risk domain to flag possible or confirmed germline variants, as patients received tumor-only sequencing and many also had germline testing. A clinical interpretation report was provided to the enrolling oncologist summarizing this interpretation. Study methods and results, including genomic variants identified, pathway analysis, and clinical impact of molecular testing, have previously been published8.
The 13 GAIN consortium members identified the need for a tool to manage large genomic and clinical datasets, support clinical interpretation, store clinical significance data related to study objectives, generate reports, and create a pediatric cancer-specific KB within a multi-institutional research study framework. Considering the limitations of systems available in 2015 and the need for a sophisticated PO tool for rare pediatric cancers, the University of Chicago (UChicago) and DFCI determined that creating a new custom system was the best path forward. Here, we describe iCatalog, which features unique capabilities including a streamlined, interoperable approach to NGS data storage, management, interpretation, and reporting across institutions. iCatalog supports multi-institutional academic studies and addresses gaps, including the limited pediatric focus in existing platforms and the lack of ability to capture genomic variant interpretations for research objectives and patient care. The platform is optimized for pediatric solid tumors to support a population with limited targeted therapeutic options. Its extensive use in the GAIN study supports additional potential utilization as a collaborative platform for future PO research.
MethodsGap assessmentiCatalog was built to support the GAIN study, launched in 2015 and targeting accrual of 825 patients. The first 192 tumor samples from 164 patients had clinical interpretation reports generated in Microsoft Word. Given the large patient cohort, this method proved inefficient and did not provide the data management infrastructure needed for the study. Hence, a team comprised of a curation scientist, study principal investigator, molecular pathologist, programmer, and research staff assessed the features needed for a tool to meet the project objectives.
Secure and compliant multi-institutional data sharingAs the lead site of the multi-institutional GAIN study, DFCI developed a data governance framework to ensure secure, ethical, and legally compliant handling of patient data. The study was conducted according to the principles of the Declaration of Helsinki (2013 version) and the International Conference on Harmonization Guidelines for Good Clinical Practice; it is approved by the institutional review board (IRB) of each of the participating institutions [Columbia University IRB, University of Chicago Biological Sciences Division/University of Chicago Medical Center IRB, University of Colorado Multiple IRB (COMIRB), Children’s National Hospital IRB, Children’s Hospital of Philadelphia IRB, Einstein IRB, Memorial Sloan Kettering Cancer Center IRB, Nationwide Children’s IRB, Seattle Children’s IRB, University of California San Francisco IRB, University of Utah IRB, UT Southwestern IRB] with the Dana-Farber/Harvard Cancer Center IRB serving as the lead IRB. Each participant provided consent and, when relevant, ascent for participation. The study protocol at each study site outlined specific procedures for secure data storage, access, and use8,9. A Memorandum of Understanding (MOU) between DFCI and each participating institution defines and governs (1) consortium aims, (2) management, including the use of the GAIN study’s DB (iCatalog), (3) data and materials use, including approval from a committee representing each site, (4) data ownership, (5) material and data transfer, and (6) intellectual property.
As part of data use, transfer, and access, the data governance structure is as follows.
Data is stored in the iCatalog DB hosted at DFCI and UChicago.
Access to data and materials beyond the study’s primary aims requires GAIN Data Use Committee approval.
PHI deposited into iCatalog is owned and accessible by the depositing institution. Only study investigators, staff at the lead site, and staff managing iCatalog have access to case-level data for all participants.
Principal Investigators (PIs) from GAIN institutions may submit data requests to the Committee. Non-participating institutions or industry may submit requests only if they are collaborating with a PI from a GAIN participating institution.
Following approval of a data request, the data is shared with the requesting institution or company. Data transfer agreements are required for data sharing when not covered by the protocol or MOU.
The team developing and maintaining the study’s data systems has expertise in data governance, managing PHI, and information security.
DesignBased on the study needs assessment, the UChicago Center for Research Informatics and DFCI investigators and research staff collaborated to identify specifications for iCatalog. The system was based on SIMPL, a molecular pathology data management tool developed at UChicago, and adapted to meet our specific needs10. Review of sequencing quality, variant calling, and normalization is performed before data are uploaded into iCatalog, as described in our prior publication8. CSV files containing variant calls and associated data from testing laboratories are uploaded to iCatalog. These files include NGS QC metrics such as sequencing depth and variant allele fraction, which are reviewed during curation.
iCatalog was intended to have two separate interacting components, a KB and a DB, and support the generation of study-specific reports. The KB, decoupled from patient-level data, contains gene curations with information across five domains and variant curations with an assigned classification of variant pathogenicity (Fig. 1). The KB content can be created using links facilitating access to external resources or edited during case review or as a stand-alone activity.
Fig. 1: System overview.
iCatalog ingests patient, specimen, and test data as well as gene and variant annotations via Web API calls, and genomic sequencing data via batch upload. The back-end (server-side) houses the relational database management system (MySQL) to store all uploaded and user-generated data. The front-end (client-side) is a web interface that allows users across collaborating hospitals to view and edit the knowledge base and generate case-specific reports.
Other DBs and systems were used to store clinical information, specimen-level metadata, pathology reports, and to track samples, as previously described9. Transfer of select information from these external systems into the iCatalog DB occurs via API. Data transferred to the iCatalog DB to facilitate clinical interpretation includes patient identifiers, demographics, enrolling diagnosis, specimen variables (procedure date, viable tumor percentage), and test information (requisition number, result date, report date). The enrollment and test diagnoses, along with the pathology report and molecular data, are reviewed to assign an integrated diagnosis with its corresponding ICD-O code. Genomic data from NGS tests include single-nucleotide variant (SNV) annotations, including c., p., and allelic fraction, and copy-number variant (CNV) annotations, including integer copy-number values for amplifications. These fields are listed in Supplementary Data 1-2. Clinical interpretation of variants and reports generated in iCatalog is also stored in the DB.
System ProgrammingiCatalog is a web-based application built on the Django Web Framework running on Linux. Programming was conducted using jQuery 3.3.1 (http://www.cs.ubc.ca/labs/spl/projects/jquery) and Bootstrap 3.3 (https://getbootstrap.com), which includes HyperText Markup Language (HTML), Cascading Style Sheets (CSS), and JavaScript. The Django version 5.0 Web Framework, running on Python 3.12, supports website development and maintenance. Prerequisites for Python modules, such as various open Lightweight Directory Access Protocol (LDAP)-related and DB-related packages, were installed. The Python virtual environment enables version sandboxing, the latest version of Python, and other required packages, including Django. The open-source MySQL software version 8.0.26 (Oracle Corporation, Redwood City, CA) was selected as the DB management system to store all user-generated data (Fig. 1). MySQL is a mature and widely used system known for its strong performance, scalability, and broad hosting compatibility. It integrates smoothly with Django and benefits from robust tooling and community support.
Hardware and Operating SystemsHardware requirements are modest. The system runs on Windows/IIS (Microsoft, Redmond, WA) or Linux, and Nginx (https://www.nginx.com) web server environments. Supplementary Fig. 1 shows the system architecture and software environment. Nginx directs dynamic requests to a reverse proxy web server, Gunicorn version 20.0.0 (http://gunicorn.org). It runs on virtual machines (VMs) behind a secure firewall within a large HIPAA-compliant computing cluster maintained by UChicago.
SecurityThe system adheres to the organization’s Information Technology security policies, which are based on the National Institute of Standards and Technology (NIST) 800-53 Cybersecurity Framework (https://www.nist.gov/cyberframework). Data is stored within a secure institutional infrastructure protected by authentication protocols. User login is authenticated through one or more institutional LDAP servers or local user accounts stored in iCatalog. iCatalog also incorporates Secure Sockets Layer encryption. LDAP allows a straightforward connection to existing hospital user verification systems for access, security, and password management. The system administrator (Super Admin) is responsible for maintaining security and functional updates to iCatalog. The administrator creates user accounts and ensures user access is governed by institutional policies and enforced through group, institution, and role-based access controls, ensuring that users have access only to data relevant to their permissions (Fig. 2). iCatalog logs all user interactions with the system, enabling retrospective evaluation and continuous management of user activity. This feature helps troubleshoot problems or errors. The DB is backed up nightly to redundant tape drive systems housed within the Center for Research Informatics. Altogether, these measures ensure data confidentiality and integrity, controlled access across participating institutions, and compliance with IRB-approved protocols and applicable regulatory standards.
Fig. 2: iCatalog supports multi-institution data sharing and reporting.
The iCatalog user hierarchy diagram illustrates user permissions based on role and group, and access to PHI based on institution. The Lead Team is responsible for overall DB management. User Administrators in the Lead Team add users, assign institutions and roles, and manage users across all institutions. The Super Admin and the Lead Team work together to optimize the system for study-specific needs and have access to the entire DB and KB. Collaborating sites have access to the complete KB but can only view and interact with the genomic and clinical data of participants at their institution. KB: knowledge base; DB: database; PI: Principal Investigator.
Analysis of the GAIN study database and knowledge base in iCatalogDB content, KB entries, and reports generated for the GAIN study between April 2017 (system launch) and December 2024 were analyzed. Regarding DB content, we describe the total number of patients, patient diagnoses, and genomic reports. The proportion of variants with clinical impact in the diagnostic, therapeutic, and cancer risk (hereditary) domains was calculated. Historical data in iCatalog helped assess the count of therapeutic implications assigned and signed off, by drug class over time, starting with the first report signed off in 2016.
The iCatalog KB content is reported descriptively. The total number of gene curations across five domains and variant curations, as well as the number of versions stored for each entry, are reported. We calculated the number of unique diagnoses and gene-association entries and queried the KB for curated and non-curated unique SVs, SNVs/Indels, and CNVs. A curated variant is a variant with interpretation content, including a classification (benign, likely benign, pathogenic, likely pathogenic, variant of uncertain significance), entered into the KB by a study team member (usually a staff scientist). Non-curated variants are stored as part of the genomic data, but do not have content entered in the KB. The distribution of variant classifications is reported for curated variants in genes with at least three likely pathogenic or pathogenic variants.
ResultsNeeds assessment identified key features for iCatalogThe iCatalog web application was designed to fulfill the specific needs of the national GAIN study (Table 1). Given the multi-institutional nature of the study, role and institution-based access to PHI, as described in the methods, was required. Workflows, case assignments, and permissions for system capabilities (such as report sign-off) were restricted. Curators, such as staff scientists with genomics training, are responsible for drafting the report. After drafting, the case is assigned to an oncologist, often the study principal investigator (PI), for review and sign-off. Only oncologists have sign-off permissions. Clinical research coordinators who are trained on patient privacy, protocol compliance, and study operations access reports after sign-off for distribution (Fig. 2).
Table 1 Key Features Implemented in iCatalogIn clinical practice, genomic data are often generated from multiple tumor-only targeted NGS tests, sometimes of different types. This was true for patients enrolled in this study. Variant data coming from six sequencing laboratories performing a total of six DNA-based and two RNA-based NGS tests can be uploaded to iCatalog (Supplementary Data 3). The system can be further customized to accept additional tests. iCatalog accelerates the variant review process and report generation, meeting our needs. For example, it provides a dashboard (Supplementary Fig. 2) and displays a patient page that provides easy access to all tests linked to the patient (Fig. 3A). On the test page, users can select or deselect any variant as a curation candidate for further review. Annotations from external sources, such as OncoKB, BayesDel, and REVEL, facilitate the selection of curation candidates and their interpretation11,12,13,14,15. In addition, rules were implemented to automatically select variants for curation, such as selected copy number amplifications but not gains, reducing the manual process (Fig. 3B). Gene and variant curation pages interface with the KB and external resources through open APIs (Supplementary Data 4 lists all external resources). KB curations can be created if they do not already exist, or existing information can be edited to reflect new evidence (Fig. 3C). Gene and variant curation summaries are automatically generated using KB curations for all variant candidates. These editable and individualized curation summaries, linked to the patient test, are stored in the DB. For example, a clinical trial description can be removed if the patient does not meet the age requirement. The curation summaries are displayed on the clinical impact page, providing the necessary information for efficient and accurate selection of clinical impact (Supplementary Fig. 3).
Fig. 3: iCatalog workflow.
A Clinical research coordinators oversee the input of clinical, sample, and genomic data. Clinical and sample data are linked from a patient and sample-tracking DB (GAINCast). B Following data upload, users can view all genomic information and automated annotations to select variants for curation. C Curation scientists and oncologists can create, select, and update KB entries before associating a curation with a variant in the test being interpreted. D Variant diagnostic, hereditary risk, and therapeutic impact are indicated with limited option entries. E Curations and clinical impact assignments are included in an individualized report and displayed on a review page for editing. Case-level interpretations are linked to a specific test and stored in the DB, not the KB. F A PDF report is generated and stored after final sign-off. KB: knowledge base; DB: database. Examples of users include clinical research coordinators (CRCs), curators, and oncologists.
iCatalog facilitates and stores variant-specific clinical impact interpretations across three domains, making them available later for analysis. For each variant, the user navigates to a page where the clinical impacts in the domains of diagnosis, hereditary risk, and therapy are entered. Clinical impact entries have limited response options, optimizing data for future analysis. In addition, prior clinical impact entries in the therapeutic domain for the same gene and variant type can be searched and viewed. Clinical impact entries are stored in the patient DB and do not alter the KB (Fig. 3D).
iCatalog auto-generates modifiable reports with case-specific curation summaries, containing gene curations, variant curations, and clinical impact statements, accessible on the review page, enabling users to review and edit (Fig. 3E). Before the oncologist signs off, they can also update discrete clinical interpretation entries as needed. The sign-off step generates a final report in PDF format. The final report cannot be modified, but an addendum can be added (Fig. 3F, Supplementary Data 13).
Additional features of iCatalog are case status tracking and data exploration. From the moment tests are entered to the point the report is sent to the treating physician and patient, the system tracks the report status and records any modifications. It logs and timestamps all changes and assigns a report status summarizing pending tasks, such as pending curation. Since the system serves as both a DB and a KB, it stores curations and clinical impact assignments at the individual case level. Clinical and genomic data, as well as data generated during curation, such as clinical impact selections, can be queried through pre-built reports or exported as CSV files (Supplementary Data 5-6).
DeploymentiCatalog was established by UChicago in a development environment. The production site launched in April 2017 to support the GAIN study. By this time, the team had already written and signed off on 192 Microsoft Word reports, but over 500 more patients were expected to enroll in the study. UChicago made the site available to select GAIN members at collaborating sites. Additional modifications to the source code were made on an ongoing basis to customize the system as desired.
The UChicago team installed the system in a desktop test environment connected to their LDAP system. The initial steps included obtaining internal review board (IRB) approval for the study, inter-institutional agreements for collaborative development, defining user roles, designating the appropriate manager of role assignment, and identifying privacy and security requirements. We then (a) implemented firewall rules (for interconnection between the iCatalog web server, the DB server, users’ computers, and storage systems), (b) configured file permissions, and (c) determined a backup strategy and methodology.
The generalized steps to iCatalog setup included: (i) cloning the iCatalog source code and establishing version control as desired; (ii) installing operating system-specific prerequisites for Python modules; (iii) configuring a DB for iCatalog to use; (iv) setting up a Python virtual environment; (v) updating the configuration, DB, and LDAP files with site-specific settings.
Once the configuration was complete, iCatalog was run using simple Python scripts to populate the DB, create an administrative user, and start the web application. Choosing and deploying web servers, considering factors such as server load, redundancy requirements, ease of configuration, and other relevant aspects, requires considerable planning and configuration effort. One standard solution is to use the fast, lightweight Nginx web server as a front-end, which serves static files. This gateway interface server relays dynamic requests to iCatalog.
GAIN databaseAs of December 31, 2024, iCatalog stores 1194 test results from 777 patients with 133 unique diagnoses. iCatalog has been used to generate 1002 clinical interpretation reports for 650 patients and stores an additional 192 reports written in Microsoft Word before the launch of iCatalog. Patient tumor samples were sequenced with one of the six DNA- or two RNA-based sequencing panels. Thirty-four percent (263/777) of patients had sequencing data from more than one test, and 15% (113/777) had sequencing data from more than one test type. In these cases, iCatalog harmonized data across test types, including the OncoPanel and Children’s Hospital of Philadelphia (CHOP) Somatic DNA panel. iCatalog also allows the study team to view sequencing results over time. The median interval between the first and last test report from patients with two or more test results is 1.1 years (range, 0–7.9 years) (Table 2). Upon uploading molecular data, we used automated variant annotations and linked external resources to filter and select variants. Thus, the KB, while expansive, is limited to curations of variants more likely to have a functional impact.
Table 2 Database and Knowledge Base Generated for the GAIN Study (n = 777)After selecting and reviewing variant curation candidates across all patients, 1805 out of 5657 variants reviewed were associated with and assigned at least one hereditary, diagnostic, or therapeutic impact (Fig. 4A; Supplementary Data 7). At least one clinical impact was assigned to variants in 315 genes. Specifically, 163, 184, and 140 genes were associated with diagnostic, therapeutic, or hereditary clinical implications, respectively (Supplementary Data 8). Furthermore, as expected in a pediatric cohort, 67% of diagnostic implications were associated with structural variants, while therapeutic and hereditary implications were more evenly distributed across all variant types (Fig. 4B). Upon oncologist sign-off, iCatalog records timestamps for each therapeutic implication assigned to variants on case-specific genomic interpretation reports (Fig. 4C; Supplementary Data 9). The sign-off timeline reveals the emergence of new drug classes such as RAF inhibitors16, and failed targeted therapy approaches such as direct EWSR1::FLI1-targeting agents17,18. By preserving this data, iCatalog showcases the evolution of knowledge from therapeutic trials over time. Detailed information on the cohort, the clinical impact of sequencing, and a summary of the clinical interpretation for a subset of the patient cohort has been reported previously8.
Fig. 4: Overview of the GAIN study DB in iCatalog.
Number A and corresponding percentage B of all SNVs/Indels, CNVs, and SVs with a clinical impact. C Timeline showing the count of therapeutic implications by drug class. Only included classes associated with therapeutic impact at least 20 times. Bubble size and color represent the monthly count for each class, while bars on the right show the total count for each drug class. SNV: single-nucleotide variant; Indel: insertion-deletion; CNV: copy-number variant; SV: structural variant.
Knowledge base generated for the GAIN studyiCatalog generated a robust pediatric solid tumor KB. Curation was performed for 2659 of 7053 (38%) variants and 581 of 683 (85%) genes identified through NGS testing in the GAIN study. Gene curation is divided into five domains, and curation versions for each domain are tracked in iCatalog. Analysis of curation versions across each domain revealed that 10% (58/581) of genes had more than one curation version for each of the five domains. The KB stores curations for 396 germline-phenotype and cancer-risk associations, 497 gene-to-all cancer associations, 250 specific pediatric cancer diagnoses-gene associations, and 224 gene-to-therapy associations (Fig. 5A; Supplementary Data 10). Of the genes associated with a specific cancer diagnosis, 88 had a KB entry for more than one cancer diagnosis (Fig. 5B). The low rate of pediatric cancer-specific curations (43%; 250/581 genes) is due to the limited pediatric relevance seen in the genes included in standard tumor profiling tests. As expected, there is a lower rate of genes with therapeutic curations (39%; 224/581 genes) as the list of genes with therapeutic significance is limited by the availability of precision therapies (Fig. 5A-B; Supplementary Data 10). Notably, the therapeutic implications domain had the most versions as it is the most dynamic. Across these gene curation domains, 4528 unique references were cited during curation.
Fig. 5: Overview of the GAIN study gene and variant KB in iCatalog.
A Number of unique genes curated across five domains and the corresponding number of curation versions stored. B Number of unique genes with at least one pediatric cancer-specific curation. C Circle plot showing the total curated and non-curated variant distribution [% (n)] across variant types. D Distribution [% (n)] of assigned variant classifications for curated variants. E Stacked bar graph of genes, with at least three likely pathogenic or pathogenic variants, showing the distribution of variant classifications, including variants not classified (light gray). SNV: single-nucleotide variant; Indel: insertion-deletion; CNV: copy-number variant; SV: structural variant.
The KB currently contains curation entries for 2659 of the 7053 unique variants identified (38%). Specifically, it includes an entry for 34% (2146/6348) of SNVs/Indels, 57% (235/410) of CNVs (two-copy deletions and amplifications) marked for curation, and 94% (278/295) of SVs (Fig. 5C; Supplementary Data 11). Unlike CNVs and SVs, which were primarily curated, most SNVs/Indels were not manually curated because automated variant annotations indicated they were likely benign. Of those curated, 20% (782/2146) of variants were classified as likely pathogenic or pathogenic, and 83% (1778/2146) of variants were classified as variants of uncertain significance, demonstrating the need for better approaches to determining the functional impact of rare variants (Fig. 5D). Genes with the highest number of pathogenic or likely pathogenic alterations in this cohort of pediatric solid tumor patients were TP53 (n = 113), DICER1 (n = 20), RB1 (n = 19), and STAG2 (n = 18) (Fig. 5E).
DiscussioniCatalog addresses major challenges in interpreting the clinical impact of variants identified in targeted NGS DNA and RNA tests for pediatric oncology patients. It automates several time-consuming steps, including annotating variants to determine oncogenicity, providing access to external genomics DB at the point of use, applying previously developed knowledge, and creating editable reports. Additionally, it provides access to different data types (clinical and genomic) from multiple sources (including multiple testing labs) over time in integrated patient and test-level views. Collaboration is facilitated through role-based access from numerous institutions. Importantly, iCatalog collects and tracks clinical interpretations made in reference to diagnostic implications, germline risk associations, and therapeutic approaches.
Numerous open-access tools for PO facilitate variant annotation and offer supplementary knowledge essential for clinical interpretation. However, most of these tools do not manage data and workflows or allow customization of pathogenicity and clinical expertise to individual cases19. Additionally, existing PO solutions have limitations for use in pediatric care and research. For example, the open-access Cancer Genome Interpreter requires queries to be restricted to a specific cancer diagnosis20. This approach has a reasonable rationale: oncogenesis, treatment response, and resistance mechanisms may differ by tissue type. However, when dealing with rare pediatric cancers, preclinical and clinical evidence related to the patient’s specific cancer may not exist, so it is essential to be able to query genomic evidence and knowledge in a diagnosis-agnostic manner.
To our knowledge, iCatalog is the first tool and KB specifically developed for a pediatric solid tumor PO study. Importantly, it provides a model system for case-specific curation, enabling queries across various systems and flexibility for curating in the context of numerous rare diseases. A recently presented academically developed open-access tool, PORI (Platform for Oncogenomic Reporting and Interpretation), developed for the Canadian Personalized OncoGenomics Program, takes a similar approach and has some of the iCatalog capabilities7. However, since it was designed for adult cancers, the system relied more on external knowledge and data, and its use did not generate a rich pediatric PO KB. We provide a comparison of our KB with others used in the field (Supplementary Data 12). The comparison highlights its unique characteristics, including its ability to capture diagnosis-agnostic therapeutic evidence and pediatric-specific content lacking in others, as well as to store information on hereditary risk, often stored in a database separate from somatic findings.
Limitations of iCatalog include the limited set of clinical variables currently stored and linked to each tumor specimen. This is intentional as the clinical variables are meant to facilitate variant interpretation, not to support all aspects of PO studies. While iCatalog was used to generate genomic interpretation reports that support clinical discussions, it is not a validated tool for formal sign-off of molecular laboratory reports. Further clinical validation would be required to use iCatalog in these settings. Additionally, iCatalog does not process genomic data from raw data files. However, there is also potential to enhance the system by automating the harmonization of KBs, including both internally developed and external KBs. Compared with platforms such as cBioPortal and the St. Jude data portal, iCatalog offers fewer analytical and visualization capabilities. Notably, our KB is currently not publicly available. We plan to revisit this once the necessary dedicated effort, infrastructure, and funding are in place.
ConclusionWhile initially designed for a prospective pediatric precision oncology study, iCatalog is flexible and customizable. It can be further customized to support germline variant curation and store additional clinical variables relevant to other cancer types. We modified the iCatalog code for two additional use cases, leveraging its adaptability to meet specific research needs. We adapted the code for use within DFCI to support multiple ongoing research projects, as well as the Participant Engagement and Cancer Genomic Sequencing Network Count Me In Osteosarcoma (osproject.org) and Leiomyosarcoma (lmsproject.org) Projects. Major adjustments included decoupling from the existing sample-tracking system, adapting to ingest whole-exome sequencing data, and converting to a project-based system. We believe that other institutions needing a precision medicine tool with such capabilities may benefit from implementing or adapting iCatalog.
Data availabilityData and additional supporting information are available in supplemental tables. The source data for Fig. 4A-B is in Supplementary Data 7 and 8, and Fig. 4C is in Supplementary Data 9. The source data for Fig. 5A-B is in Supplementary Data 10, and Fig. 5C is in Supplementary Data 11. Data not included in the supplemental files are available in a prior publication by Church et al.8 or can be provided upon request to Katherine Janeway (katherine_janeway@dfci.harvard.edu).
Code availabilityiCatalog was developed through a collaborative effort between the Dana-Farber Cancer Institute (DFCI) and the University of Chicago, with the University of Chicago retaining administrative control over the release of the source code. The codes for the original iCatalog and research versions are hosted in a private Bitbucket repository and available upon request. Interested parties may request access by contacting the co-first author and software developer at the University of Chicago (Wenjun Kang, wkang2@bsd.uchicago.edu) or the corresponding author (Katherine Janeway, katherine_janeway@dfci.harvard.edu).
ReferencesJiang, W. et al. Tri(c)DB: an integrated platform of knowledgebase and reporting system for cancer precision medicine. J. Transl. Med. 21, 885 (2023).
Yu, Y. et al. PreMedKB: an integrated precision medicine knowledgebase for interpreting relationships between diseases, genes, variants and drugs. Nucleic Acids Res. 47, D1090–D1101 (2019).
Dahary, D. et al. Genome analysis and knowledge-driven variant interpretation with TGex. BMC Med Genomics 12, 200 (2019).
de Bruijn, I. et al. Genome nexus: a comprehensive resource for the annotation and interpretation of genomic variants in cancer. JCO Clin. Cancer Inf. 6, e2100144 (2022).
Borchert, F. et al. Knowledge bases and software support for variant interpretation in precision oncology. Brief Bioinform. https://doi.org/10.1093/bib/bbab134 (2021).
Tamborero, D. et al. The Molecular Tumor Board Portal supports clinical decisions and automated reporting for precision oncology. Nat. Cancer 3, 251–261 (2022).
Reisle, C. et al. A platform for oncogenomic reporting and interpretation. Nat. Commun. 13, 756 (2022).
Church, A. J. et al. Molecular profiling identifies targeted therapy opportunities in pediatric solid cancer. Nat. Med. 28, 1581–1589 (2022).
Harris, M. H. et al. Multicenter feasibility study of tumor molecular profiling to inform therapeutic decisions in advanced pediatric solid tumors: the individualized cancer therapy (iCat) study. JAMA Oncol. 2, 608–615 (2016).
Kang, W. et al. System for informatics in the molecular pathology laboratory: an open-source end-to-end solution for next-generation sequencing clinical data management. J. Mol. Diagn. 20, 522–532 (2018).
Suehnholz, S. P. et al. Quantifying the expanding landscape of clinical actionability for patients with cancer. Cancer Discov. 14, 49–65 (2024).
Chakravarty, D. et al. OncoKB: a precision oncology knowledge base. JCO Precis Oncol. https://doi.org/10.1200/PO.17.00011 (2017).
Ioannidis, N. M. et al. REVEL: an ensemble method for predicting the pathogenicity of rare missense variants. Am. J. Hum. Genet. 99, 877–885 (2016).
Tian, Y. et al. REVEL and BayesDel outperform other in silico meta-predictors for clinical variant classification. Sci. Rep. 9, 12752 (2019).
Feng, B. J. PERCH: a unified framework for disease gene prioritization. Hum. Mutat. 38, 243–251 (2017).
Hashem, O. et al. An overview of RAF kinases and their inhibitors (2019-2023). Eur. J. Med Chem. 275, 116631 (2024).
Povedano, J. M. et al. TK216 targets microtubules in Ewing sarcoma cells. Cell Chem. Biol. 29, 1325–1332 e4 (2022).
Meyers, P. A. et al. Open-label, multicenter, phase I/II, first-in-human trial of TK216: a first-generation EWS::FLI1 fusion protein antagonist in ewing sarcoma. J. Clin. Oncol. 42, 3725–3734 (2024).
Gazola, A. A., Lautert-Dutra, W., Archangelo, L. F., Reis, R. B. D. & Squire, J. A. Precision oncology platforms: practical strategies for genomic database utilization in cancer treatment. Mol. Cytogenet. 17, 28 (2024).
Nakken, S. et al. Personal cancer genome reporter: variant interpretation report for precision oncology. Bioinformatics 34, 1778–1780 (2018).
Funding for this study was provided by the Precision For Kids Pan Mass Challenge Team, the 4 C’s Fund, Lamb Family Fund, C&S Wholesale Grocers, and C&S Charities.
Author information
Author notes
Navin R. Pinto
Present address: University of Colorado Anschutz Medical Campus, Boulder, CO, USA
These authors contributed equally: Wenjun Kang, Lorena Lazo de la Vega.
University of Chicago, Chicago, IL, USA
Wenjun Kang, Mark A. Applebaum, Mengjie Chen, Julie A. Johnson & Samuel Volchenboum
Dana-Farber/Boston Children’s Cancer and Blood Disorders Center, Boston, MA, USA
Lorena Lazo de la Vega, Hannah Comeau, Ergina Agastra, Ellen Sukharevsky, Evelina Ceca, Laura Corson, Joseph White, Alanna J. Church & Katherine A. Janeway
Broad Institute of MIT and Harvard, Cambridge, MA, USA
Lorena Lazo de la Vega, Ellen Sukharevsky, Evelina Ceca, Alanna J. Church & Katherine A. Janeway
Primary Children’s Hospital, Salt Lake City, UT, USA
Luke D. Maese
University of Utah Huntsman Cancer Institute, Salt Lake City, UT, USA
Luke D. Maese
Children’s National Hospital, Washington, DC, USA
AeRang Kim
George Washington University School of Medicine and Health Sciences, Washington, DC, USA
AeRang Kim
Children’s Hospital at Montefiore, Bronx, NY, USA
Daniel A. Weiser
Albert Einstein College of Medicine, Bronx, NY, USA
Daniel A. Weiser
Comer Children’s Hospital, Chicago, IL, USA
Mark A. Applebaum
Nationwide Children’s Hospital, Columbus, OH, USA
Susan I. Colace
Ohio State University College of Medicine, Columbus, OH, USA
Susan I. Colace
Seattle Children’s Hospital, Seattle, WA, USA
Navin R. Pinto
University of Washington, Seattle, WA, USA
Navin R. Pinto
Harvard Medical School, Boston, MA, USA
Alanna J. Church & Katherine A. Janeway
W.K.: validation, software development, resources, visualization; L.L.D.L.V.: data curation, writing—original draft, visualization, formal analysis of knowledge base, refinement of user interface; H.C.: resources, visualization, refinement of user interface, project administration; E.A.: data curation, visualization, refinement of user interface; L.D.M.: investigation, data curation; A.K.: investigation, data curation; E.S.: visualization, refinement of user interface; E.C.: resources, supervision, project administration; L.C.: conception, refinement of user interface, data curation; J.W.: resources, refinement of user interface; D.A.W.: investigation, data curation; M.A.A.: investigation, data curation; S.I.C.: investigation, data curation; M.C.: resources, supervision; J.A.J.: resources, supervision; S.V.: conception, funding acquisition; N.R.P.: investigation, data curation, refinement of user interface; A.J.C.: conception; K.A.J.: investigation, conception and refinement of user interface visualization, validation, data curation, funding acquisition; All authors contributed to revisions and approved the final manuscript for publication.
Corresponding author Ethics declarationsCompeting interestsW.K., L.L.D.L.V., H.C., E.A., L.D.M., A.K., E.S., E.C., L.C., J.W., D.A.W., M.A.A., S.I.C., S.V., N.R.P., A.J.C. have no conflicts. M.C. consults for Tellic and Impetus. J.A.J. has received travel support from MDClone to attend scientific meetings. K.A.J. consults for Recordati and receives research funding from AstraZeneca.
Peer reviewPeer review informationCommunications Medicine thanks Yaqiong Jin and the other, anonymous, reviewer(s) for their contribution to the peer review of this work. A peer review file is available.
Additional informationPublisher’s note Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.