Welcome to the Data Gathering Repository for the Index of NCI Studies (INS)!
This repository uses a curated list of National Cancer Institute (NCI) programs to automate the gathering of information about grants, projects and their associated research outputs. This process uses publicly available resources from the NIH RePORTER API, NIH iCite, NCBI E-utilities (through BioPython), and NCBI dbGaP APIs along with curated and submitted data.
The data gathered here are compatible with the INS Data Model.
The finalized output files for each release are available as a .zip on the Latest Release page. These TSV files are loaded into the INS database and match the information available when browsing the site.
git clone https://github.com/CBIIT/INS-Data.git
cd INS-Data
uv venv .venv
.venv\Scripts\activate # Windows. On macOS/Linux: source .venv/bin/activate
uv pip install -r requirements.txt
uv pip install -e .Before running, you will need:
- A CSV of NCI programs from ODS (step 3)
- An iCite database snapshot, ~10GB (step 4)
- An NCBI API key in a
.envfile (step 5)
Then run the full pipeline:
python main.pyFor detailed setup, prerequisites, and optional steps, see How to Use this Repository.
The INS Data Gathering workflow consists of the following steps designed to run together in order:
- Gather Programs
- Gather Grants
- Gather Projects
- Gather Publications
- Gather GEO Datasets
- Gather SRA Datasets
- Process CEDCD Cohorts
- Package Data
- Validate Data
The workflow is supported by additional, independent steps run as needed:
- Process CTD^2 Datasets
- Gather dbGaP Datasets
- Merge Curated dbGaP Datasets
- Generate dbGaP Subset Clones
- Process Curated Sources (DCEG Cohorts, NCCR)
- Curate Resources
This step gathers programs curated by the NCI Office of Data Sharing (ODS).
Programs represent a coherent assembly of plans, project activities, and supporting resources contained within an administrative framework, the purpose of which is to implement an organization's mission or some specific program-related aspect of that mission.
Within INS, programs are the highest-level node. Programs contain projects.
All program processing is handled in the gather_program_data.py module and can be run as an independent process with the command:
python modules/gather_program_data.py-
Process Qualtrics Input
- Processes the curated CSV of Programs received from the NCI Office of Data Sharing (ODS). This CSV is an export of survey results from the Qualtrics survey tool. Each Program in this export includes a curation of associated Notices of Funding Opportunities (NOFOs) or Awards.
- NOTE: Though similar terms, 'Award' is used throughout this documentation to refer to the Award values provided within the Qualtrics CSV, while 'Grant' is used to refer to the grants gathered from NIH RePORTER and used in downstream processing.
-
Validate provided NOFOs and Awards
- Compares each NOFO or Award to expected string formatting patterns to check for validity
- Generates versioned
invalidAwardReportandinvalidNofoReportin thereportsdirectory if any unexpected patterns are found - Prompts user to review and correct any issues. There are two ways to make corrections:
- Make manual changes within the raw qualtrics file
qualtrics_output_{version}_raw.csvand save asqualtrics_output_{version}_manual_fix.csv. ChangeQUALTRICS_TYPEtomanual_fixwithinconfig.py. - Add
suggested_fixand optionalcommentcolumn(s) to theinvalidAwardReportorinvalidNofoReportcsv and save to thedata/reviewed/{version}/directory with the_reviewedsuffix. The validation step will automatically check to see if this file exists and make any suggested changes specified within. If any invalid values still remain after this fix, aninvalidAwardReport_corrected.csvorinvalidNofoReport_corrected.csvis generated in thereportsdirectory.
- Make manual changes within the raw qualtrics file
-
Validate Program names and generate IDs
- Checks for duplicates or unexpected combinations in Program Names and Program Acronyms
- If duplicates are found, then a manual fix must be made to the qualtrics input file to avoid downstream issues.
- Save the fixed file as
qualtrics_output_{version}_manual_fix.csv. ChangeQUALTRICS_TYPEtomanual_fixwithinconfig.py.
- Generates a
program_idfrom the Program Acronym
- Checks for duplicates or unexpected combinations in Program Names and Program Acronyms
-
Generate clean Program file
- Saves intermediate
program.csvin versioneddata/01_intermediate/directory for reference and downstream use - The fields expected are defined and can be modified in
config.py
- Saves intermediate
This step gathers grants from the NIH RePORTER. Only grants associated with programs (above) are gathered.
Grants are financial assistance mechanisms providing money, property, or both to an eligible entity to carry out an approved project or activity. A grant is used whenever the NIH Institute or Center anticipates no substantial programmatic involvement with the recipient during performance of the financially assisted activities.
Within INS, a grant is usually an annual support to a multi-year project.
- 5U24CA209999-02 | Monitoring tumor subclonal heterogeneity over time and space (FY2017)
- 5U24CA209999-03 | Monitoring tumor subclonal heterogeneity over time and space (FY2018)
- 5U10CA031946-23 | Cancer and Leukemia Group B (FY2004)
- 1P50CA217691-01A1 | Emory University Lung Cancer SPORE (FY2019)
All grants processing is handled within the gather_grant_data.py module, except for the statistics step handled with the build_program_project_stats.py module. They can be run as independent processes with the commands:
python modules/gather_grant_data.py
python modules/build_program_project_stats.py-
Get grants data from NIH RePORTER API
- This process takes approximately 5-10 minutes for ~70 programs
- For each Key Program, this queries the NIH RePORTER API to gather a list of all associated extramural grants along with descriptive data for each grant.
- The NOFOs (e.g.
RFA-CA-21-038;PAR21-346) and/or Awards (e.g.1 U24 CA274274-01;P50CA221745;3U24CA055727-26S1) provided for each Key Program are used as the query. - The following exclusion are also applied within the query:
- Subprojects are excluded
- Grants prior to fiscal year 2000 are excluded
- Grants receiving no funding from NCI are excluded
-
Process grants data
- Reformats the data received from the NIH RePORTER API for use within INS.
- Removes extraneous fields and rename fields to match the existing INS data model.
- Flattens nested JSON structures. In particular, the PI, PO, and agency funding fields have this structure.
- Formats names to standardize capitalization
- The fields expected are defined and can be modified in
config.py
- Reformats the data received from the NIH RePORTER API for use within INS.
-
Save grants data for each Program
- Adds the associated Program ID to each grant
- Combines grants data from all programs and store as a versioned
grant.csvwithin thedata/01_intermediate/directory.
-
Generate program and project statistics
- Builds reports useful for testing and validation but not intended for ingestion into the site
grantsStatsByProgram.csvgroups grants data by Key Program and aggregates counts of grants, projects, searched values, and earliest fiscal yearsharedProjectsByProgramPair.csvlists pairs of Key Programs and counts of projects that are associated with both
- Statistics are handled within the
build_program_project_stats.pymodule
- Builds reports useful for testing and validation but not intended for ingestion into the site
This step derives projects from the grants data gathered from NIH RePORTER. All projects are associated with at least one program.
Projects are the primary unit of collaborative research effort, sometimes also known as core projects or parent projects. Projects receive funding from awards to conduct research and produce outputs, often across multiple years.
Within INS, projects are organized as a grouping of grants with identical Activity Code, Institute Code, and Serial Number.
- U24CA209999 | Monitoring tumor subclonal heterogeneity over time and space
- U10CA031946 | Cancer and Leukemia Group B
All projects processing is handled within the gather_project_data.py module and can be run as an independent process with the command:
python modules/gather_project_data.py-
Aggregate grants into projects
-
Using the intermediate grants file
grant.csvas input, this process builds a table of all projects and fills it with information pulled from grants for each project. -
Each project field is populated using the appropriate aggregation type:
- Newest/Oldest: Some fields should be populated with values from either the newest (e.g. project title, abstract) or oldest (e.g. project start date) grant.
- Identical: Some fields are identical across all grants within a project (e.g. organization and location details). For consistency, the project value is populated with the value from the newest non-supplement grant.
- List: Some fields are populated by gathering and listing all values from all grants within a project (e.g. opportunity number)
-
Newest/oldest grants are decided using the award notice date if available. If not available, fiscal year is used.
-
Because grant supplements are not always representative of the project, supplements are ignored when pulling newest/oldest values when non-supplement grants are available within a project.
-
-
Validate values and add Program IDs
- Validates that all values that should be identical between grants within a project are identical. Any mismatches are reported in
mismatchedProjectValueReport.csvin thereportsdirectory. - Adds Program IDs to projects. To support the data model, any project associated with more than one program will appear multiple times (once for each program).
- Stores project data in the versioned
project.csvwithin thedata/01_intermediate/directory.
- Validates that all values that should be identical between grants within a project are identical. Any mismatches are reported in
This step gathers publication data from NIH RePORTER, NCBI PubMed, and NIH iCite. All publications are associated with at least one project.
Publications gathered for INS only include articles represented in PubMed with a unique PubMed ID. Projects associated with the CCDI program are excluded from the publications workflow.
All publication processing is handled within the gather_publication_data.py module and can be run as an independent process with the command:
python modules/gather_publication_data.py-
Get associated PMIDs from the NIH RePORTER API
- This process takes approximately 45 minutes for ~2500 project IDs
- For each project in
project.csv, this queries the NIH RePORTER API to gather a list of all associated PubMed IDs (PMIDs)- For information on how NIH RePORTER links projects to publications, see their FAQ
- Note that the project-to-publication link is many-to-many. A single project can be associated with multiple publications, and a single publication can be associated with multiple projects
- These are stored in the intermediate "checkpoint" file
projectPMIDs.csv. For subsequent data gathering runs on the same start date (version), this file can be used instead of gathering PMIDs again.
-
Gather select PubMed information for each PMID
- NOTE: This process can take several hours to gather information and is highly dependent upon the number of PMIDs gathered. The rate is approximately 15,000-20,000 publications per hour and it is recommended to run this outside of peak hours (9am - 5pm EST)
- Uses the BioPython Entrez package to access the PubMed API via Entrez e-Utilities and query PMIDs for PubMed information.
- The following fields are pulled from PubMed:
- Title
- Authors
- Publication date
- NOTE: Publication dates are inconsistent within the PubMed data. When month and/or date cannot be identified, they will default to January and/or 1st. (e.g.
2010 Decwould be interpreted as2010-12-01and2010as2010-01-01). Publication year is never estimated.
- NOTE: Publication dates are inconsistent within the PubMed data. When month and/or date cannot be identified, they will default to January and/or 1st. (e.g.
- Because of the long processing time, checkpoint files are saved periodically in a temporary
temp_pubmed_chunkfilesdirectory within the versioneddata/01_intermediate/directory.- The default length of each checkpoint file is 2000 rows, but this can be changed in config.py with
PUB_DATA_CHUNK_SIZE - Whenever this workflow is run for the same start date (version), any existing checkpoint files are all loaded together and the unique PMIDs within are accounted for. Each run will check for any missing PMIDs and restart the data gathering wherever it left off. This allows the publications workflow to be stopped and restarted without problems.
- The default length of each checkpoint file is 2000 rows, but this can be changed in config.py with
-
Gather iCite information for each PMID
- Uses the most recent NIH iCite bulk download (zipped CSV) to access iCite data for each PMID
- This manual download process must be completed before starting the automated workflow
- The downloaded
icite_metadata.zipshould be stored in a versioned (e.g.2023-11/) directory in thedata/00_input/icite/directory- Note: iCite files are ~10GB and are not stored under git control
- The download takes 2-4 hours depending upon download speed.
- Note: The iCite API was explored, but operates much slower (3-4x) than the PubMed API.
- The downloaded
- The following fields are pulled from iCite:
- Title
- Authors
- Publication Year
- Citation Count
- Relative Citation Ratio (RCR)
- Saves checkpoint file
icitePMIDData.csvcontaining all iCite data fields for PMIDs of interest
-
Combine and clean PubMed and iCite data for each PMID
- Adds the unique metrics from iCite (Citation Count & RCR) to the PubMed data for each PMID
- Checks for any missing values in the PubMed information and fill in with iCite information where available
- When values are conflicting, PubMed is used as the default. iCite values are only used when PubMed value is completely missing
- Saves checkpoint file
mergedPMIDData.csvcontaining all merged PubMed and iCite data for PMIDs of interest - Cleans the publications data by removing rows with the following issues:
- Publication year before 2000
- No publication information for any fields from PubMed or iCite
- Stores a report of removed publications in a versioned
removedPublicationsReport.csvwithin thereports/directory - Stores the intermediate publication output
publication.csvin the versioneddata/01_intermediate/directory
This step gathers Gene Expression Omnibus (GEO) Dataset data from NCBI GEO using NCBI E-utilities. All GEO datasets are associated with at least one publication and at least one project.
All GEO dataset processing is handled within the gather_geo_data.py module and can be run as an independent process with the command:
python modules/gather_geo_data.py-
Map PMIDs to GEO IDs
- Reads the publication data (with PMIDs and project info) and uses NCBI E-utilities to find all GEO dataset IDs (GSE/GDS) linked to each PMID.
- Saves a mapping of PMIDs to GEO IDs as a checkpoint file for reuse.
-
Gather GEO Metadata
- For each unique GEO ID, retrieves metadata from NCBI using the ESummary API, including accession, title, description, sample count, assay method, and FTP links.
- Saves the raw metadata as a JSON checkpoint file.
-
Extract FTP Metadata
- For each GEO dataset, connects to the NCBI GEO FTP server and parses the series matrix files to extract additional metadata fields, such as contributors.
- Saves the extracted FTP metadata as a JSON checkpoint file.
-
Combine and Format Data
- Merges GEO metadata, FTP metadata, and project/program information into a single DataFrame.
- Adds standardized columns (e.g., URLs, UUIDs, type, repository, and empty fields for downstream compatibility).
- Removes non-dataset accessions and ensures only GSE/GDS datasets are included.
-
Save Final Output
- Saves the processed
geo_datasets.csvfile in the intermediate data directory for downstream use and validation.
- Saves the processed
This step gathers Sequence Read Archive (SRA) dataset data from NCBI SRA using NCBI E-utilities. All SRA datasets are associated with at least one publication and at least one project.
All SRA dataset processing is handled within the gather_sra_data.py module and can be run as an independent process with the command:
python modules/gather_sra_data.py-
Map PMIDs to SRA IDs
- Reads the publication data (with PMIDs and project info) and uses NCBI E-utilities to find all SRA experiment IDs (accessions) linked to each PMID.
- Processes PMIDs in configurable batches (default 2000) to enable resumption and handle large datasets efficiently.
- Saves batch mapping files and tracks failed PMIDs for retry in a second pass.
-
Map SRA IDs to Study IDs
- For each unique SRA experiment ID, retrieves the parent study ID (SRP/ERP) using NCBI E-utilities.
- Handles both NCBI SRA (SRP) and European Nucleotide Archive (ERP) study accessions.
- Creates a study-centric view by aggregating all PMIDs associated with each unique study.
-
Gather SRA Study Metadata
- For each unique study ID, pulls metadata from NCBI E-utilities.
- Parses and formats metadata fields like study title, abstract, and investigator.
-
Link to NCI Programs
- Traces INS associations backwards from SRA -> PMID -> Project -> Program
- Derives associated NCI Division/Office/Center (DOC) and program(s) for each dataset.
- Combines PMID information from both the SRA metadata and the publication-based mapping.
-
Generate UUIDs and Format Output
- Adds standardized columns like UUIDs, URLs, type, repository name for downstream handling.
- Applies validation to remove records missing required fields.
-
Save Final Output
- Saves the processed
sra_datasets.csvfile in the intermediate data directory for downstream use and validation. - Generates comprehensive reports of failed PMID lookups with error details for troubleshooting.
- Saves the processed
This step processes cohort metadata from the NCI Cancer Epidemiology Descriptive Cohort Database (CEDCD). These cohorts are included as datasets in INS and are not enriched by APIs or other external resources.
All CEDCD cohort processing is handled within the gather_cedcd_data.py module and can be run as an independent process with the command:
python modules/gather_cedcd_data.py-
Load CEDCD cohort metadata
- Reads the input CSV file of cohort metadata received from the CEDCD team.
- Loads all cohort entries, including multiple versions of the same cohort if present.
-
Filter to newest cohort versions
- Keeps only the newest version of each unique cohort title, ensuring that only the most up-to-date information is included.
-
Clean and standardize data
- Removes newline and hidden return characters from all string fields.
- Cleans up PI values and standardizes column names and values for INS compatibility.
-
Add INS-specific fields
- Adds required columns such as URLs, UUIDs, type, repository, and empty fields for downstream compatibility.
- Sets hard-coded values for fields like
primary_diseaseanddataset_docas needed.
-
Save intermediate output
- Saves the processed cohort data as
cedcd_datasets.csvin the intermediate data directory for downstream use and final packaging.
- Saves the processed cohort data as
This independent step gathers NCI-supported datasets from the NCBI dbGaP. These are not necessarily associated with any of the Programs, Grants, Projects, or Publications gathered in the main workflow.
Datasets within INS represent open-access or controlled data released as an output of an NCI-supported effort. INS datasets are sourced from dbGaP study metadata as well as several other repositories and curated sources described below.
- phs002790 | Childhood Cancer Data Initiative (CCDI): Molecular Characterization Initiative
- phs002153 | Genomic Characterization CS-MATCH-0007 Arm S1
All Dataset gathering steps are handled within the gather_dbgap_data.py module.
This module is not included in the main.py workflow and must be run independently:
python modules/gather_dbgap_data.pyNOTE: This module is intended to automate the effort to gather a large number of datasets, but the results must undergo manual review and curation. Because of this, the datasets module is not intended to be run as regularly as the main workflow.
-
Load input CSV of NCI-supported dbGaP studies
- CSV is retrieved from the dbGaP Advanced Search by filtering for IC: NCI and downloading resulting list
- This list has phs study accessions of interest along with some descriptive metadata
-
Enrich dataset data with dbGaP API resources
- Use the dbGaP Study Metadata API to gather additional study metadata for each phs
- Use the dbGaP SSTR API to gather additional study metadata for each phs
-
Enrich dataset data with NCI administrative information
- Use a file received from the NCI Office of Data Sharing (ODS) that monitors dbGaP data submissions as input
- Associate Grant Program Administrators (GPAs) and NCI Divisions, Offices, and Centers (DOCs) with each phs
- Lookup tables associating GPAs and DOCs were built using public information from the following resources:
-
Manually review and curate final datasets file
- Combine and clean all dataset inputs and export as a CSV for curation
- Missing, erroneous, or unstructured source data is refined by expert curation
- The curated file is validated and ready for INS data loading
This independent step merges a previously curated dbGaP datasets file with a newly gathered dbGaP datasets file. It preserves hand-curated values from previous releases while incorporating any newly added studies from the latest dbGaP search results.
All merge processing is handled within the merge_curated_dbgap.py module and can be run as an independent process with the command:
python modules/merge_curated_dbgap.pyNOTE: This module should be run after gather_dbgap_data.py has produced a new gathered output and a curated file from a previous release is available.
-
Load old curated and new gathered datasets
- Loads the previously curated dbGaP TSV (from the prior release cycle) and the newly gathered dbGaP TSV
- Uses
dataset_source_id(phs accession) as the merge key
-
Merge datasets
- Rows present in the old curated file are kept as-is, preserving hand-edited values and existing UUIDs
- Rows present only in the new file are appended as new studies requiring curation
- Rows present only in the old file are retained with a warning, as they may have been removed from the latest dbGaP search results but should not be silently dropped
- Regenerates deterministic UUID5 values on old curated rows to ensure consistency with the current scheme
-
Detect and review title changes
- Identifies studies where the title differs between the old and new datasets
- On the first run, exports a review CSV (
merge_title_review.csv) for the user to indicate whether to accept the new title or keep the old one - On subsequent runs, if a reviewed CSV is found, the user's title decisions are applied automatically
-
Save merged output
- Saves the merged file as
dbgap_datasets_merged.tsvin the versioneddata/01_intermediate/dbgap/directory - The merged file is ready for manual curation before final packaging
- Saves the merged file as
This independent step creates cloned dataset entries for dbGaP studies that are also available through other NCI data repositories, such as the Cancer Research Data Commons (CRDC) or the NCI Genomic Data Commons (GDC).
For each dbGaP study tagged with a storage distribution, this step clones the row from the final curated dbGaP output and relabels it with the target repository name. All other field values (including the dataset_source_id and dataset_source_url, which still point to dbGaP) are preserved unchanged. A new deterministic UUID5 is generated for each clone.
All subset clone processing is handled within the gather_dbgap_subset_clones.py module and can be run as an independent process with the command:
python modules/gather_dbgap_subset_clones.pyNOTE: This module should be run after package_output_data.py has produced the final curated clean dbGaP TSV. It does not modify the source file.
-
Load subset CSVs
- For each configured subset key (e.g.
CRDC,CTDC,GDC), loads the corresponding subset CSV fromdata/00_input/dbgap/to get the set of relevant phs accessions - Multiple subset keys can map to the same target repository (e.g. both
CRDCandCTDCsubsets produce clones labelled "CRDC")
- For each configured subset key (e.g.
-
Clone and relabel matching rows
- Filters the curated clean dbGaP output to rows matching the subset phs accessions
- Replaces
dataset_source_repowith the target repository name - Generates a new deterministic UUID5 for each clone based on the target repo and phs accession
- When the same phs accession appears in multiple subsets that map to the same repo, only one clone is produced
-
Save subset clones output
- Saves the cloned dataset entries as
dbgap_subset_clones.tsvin the versioneddata/02_output/dbgap/directory
- Saves the cloned dataset entries as
This independent step processes curated datasets and file metadata from the Cancer Target Discovery and Development (CTD^2) Network. These datasets are included in INS based on curated input files and are not enriched by APIs or other external resources.
All CTD^2 processing is handled within the gather_ctd2_data.py module and can be run as an independent process with the command:
python modules/gather_ctd2_data.pyNOTE: CTD^2 input data is not expected to change frequently. This module should only be run when dataset or file metadata needs to be updated.
-
Process CTD^2 dataset metadata
- Loads a curated CSV of CTD^2 dataset metadata from
data/00_input/ctd2/ - Generates deterministic UUID5 values for each dataset based on source repository, title, and description
- Saves the intermediate output as
ctd2_datasets.csvin the versioneddata/01_intermediate/ctd2/directory
- Loads a curated CSV of CTD^2 dataset metadata from
-
Process CTD^2 file metadata
- Loads a curated CSV of file metadata from
data/00_input/ctd2/ - Generates deterministic UUID5 values for each file based on file name, file type, and access level
- Loads a curated CSV of file metadata from
-
Map files to datasets
- Maps each file to a dataset using the curated
download_file_linksfield - Validates the mapping by checking for unmapped files, datasets with no mapped files, and files mapped to more than one dataset
- Saves the processed file metadata (with dataset UUIDs) as
ctd2_filedata.csvin the versioneddata/01_intermediate/ctd2/directory
- Maps each file to a dataset using the curated
Some datasets in INS are sourced from curated or submitted input files. These sources do not have an automated gathering pipeline — their curated TSVs are placed directly in the data/01_intermediate/ directory and processed during the Package Data step. Deterministic UUID5 values are generated for each dataset during packaging.
Cohort datasets from the NCI Division of Cancer Epidemiology and Genetics (DCEG) are included in INS as curated datasets. The curated TSV is placed in a versioned data/01_intermediate/dceg_cohorts/ directory and the DCEG_COHORTS_VERSION in config.py should be updated to match.
Datasets from the NCI Center for Cancer Research (NCCR) Data Platform are included in INS as curated datasets. The curated TSV is placed in a versioned data/01_intermediate/nccr/ directory and the NCCR_VERSION in config.py should be updated to match.
This independent step manages the curation and packaging of NCI research resources for INS. Resources represent tools, datasets, databases, and other research aids curated for the cancer research community.
Resource curation is managed through a curated TSV file at data/01_intermediate/resources/resources_{date}.tsv. After edits are made to this file (adding, updating, or removing resources), the packaging module processes it into the final output.
All resource packaging is handled within the package_resources.py module and can be run as an independent process with the command:
python modules/package_resources.py-
Read curated resources TSV
- Automatically finds the most recent
resources_{date}.tsvindata/01_intermediate/resources/ - Handles common Excel encoding artifacts (BOM, cp1252 fallback, whitespace)
- Automatically finds the most recent
-
Auto-correct formatting
- Ensures the
typecolumn is set toresourcefor every row - Deduplicates and alphabetically sorts values within semicolon-separated fields (tool type, research area, etc.)
- Ensures the
-
Generate deterministic UUIDs
- Stamps each row with a deterministic UUID5 based on
resource_source_id - Checks for duplicate UUIDs (critical error if found)
- Stamps each row with a deterministic UUID5 based on
-
Validate data quality
- Checks for duplicate
resource_source_idvalues, malformed IDs, blank required fields, and non-ASCII characters - Generates an issue report distinguishing auto-corrected issues from those requiring manual curation
- Checks for duplicate
-
Save outputs
- Writes the processed TSV to
data/02_output/resources/resources_{date}_processed.tsv - Writes a detailed issue report (with field completion summary and facet value counts) to
reports/resources/resources_{date}_issue_report.txt
- Writes the processed TSV to
This step standardizes, validates, and exports all data gathered from other steps of the process.
All final data packaging steps are handled within the package_output_data.py module and can be run as an independent process with the command:
python modules/package_output_data.py-
Finalize columns in output files
- Adds a
typecolumn and fills with appropriate values required for data loading(e.g.programorproject) - Specifies columns and ordering for all files. These can be configured within
config.py
- Adds a
-
Standardize characters
- Normalizes any non-standard characters within data gathered from sources
- Characters were normalized using NFKC and then converted to ASCII before saving as UTF-8
-
Perform special handling steps
- Removes any publications with a publication date more than 365 days before the associated project start date
- This helps to address the over-citation issue occasionally encountered when authors cite inappropriate grant support. See this 2023 Office of Extramural Research article for more details on the issue.
- Any publications removed in this process are saved in a
removedEarlyPublications.csvin thereports/../packagingReportsdirectory. - The 365-day buffer can be modified within
config.py
- Removes any publications with a publication date more than 365 days before the associated project start date
-
Validate and save final outputs
- Validates that files do not contain duplicates or other inconsistencies
- NOTE: Nodes with many-to-many relationships (e.g.
publications) will have duplicate records where all values are identical except for the linking column (e.g.project.project_id). This is intentional.
- NOTE: Nodes with many-to-many relationships (e.g.
- Validates that all list-like columns are separated by semicolons
- Generates a list of enumerated values for each column specified in
config.py. These can be used to update the INS Data Model before loading. - Saves all final outputs as TSV files in the
data/02_outputdirectory. These are ready for INS data loading.
- Validates that files do not contain duplicates or other inconsistencies
This step builds a Data Validation Excel file with expected data values. It can be used to validate that data is loaded and displayed correctly in the INS UI.
All Data Validation Generation steps are handled within the build_validation_file.py module and can be run as an independent process with the command:
python modules/build_validation_file.py-
Load data from finalized TSV output files
- This reads final output TSVs to generate the validation file, but does not change or save the output TSVs.
-
Build "Report Info"
- Each step will build a dataframe ready for export as a separate tab within the Excel output
- Each tab within the Excel will represent a different type of information or validation
- The first "report_info" tab lists basic versioning and timestamp information
- This tab can be used to verify that the data version matches the data validation version
-
Build "Single Program Results"
- Creates a table where each unique program is listed with total counts of associated projects, grants, and publications
- Count of projects, grants, and publications should match counts on the INS Program Details Page or on the Explore page when a single program is selected within the filter
-
Build "Single Project Results"
- Creates a table where each unique project is listed with total counts of associated grants and publications
- Count of grants and publications should match counts on the INS Project Details Page
-
Format and save all tables to Excel as separate tabs
- Excel file is saved in the versioned
reports/directory with a timestamp noting when it was generated
- Excel file is saved in the versioned
Note: The raw output files for each release are also available as a .zip in the Latest Release
-
Go to the data/02_output directory.
-
Open the folder named with the most recent date. This notes the date that the Programs file was received from ODS.
-
Open the folder named with the most recent gathered date. This notes the date that the data gathering pipeline was run.
For example, the files within
data/02_output/2024-01-04/gathered-2024-01-10/used a Program file received on January 4th, 2024 as the input into the gathering process run on January 10th, 2024. -
Download the TSV files within this directory. These represent the latest, finalized data intended for loading into INS.
-
Clone the repo to your local machine
- Option 1: Use the built-in GitHub cloning method or a tool like GitHub desktop
- Option 2: Open the command terminal and navigate to the desired destination. Run the following:
git clone https://github.com/CBIIT/INS-Data.git
-
Set up the environment
- Create a virtual environment using uv (recommended) or standard Python:
# Option 1: uv (recommended) uv venv .venv # Option 2: standard Python python -m venv .venv
- Activate the environment:
# Windows .venv\Scripts\activate # macOS/Linux source .venv/bin/activate
- Install dependencies (replace
uv pipwithpipif not using uv):
uv pip install -r requirements.txt uv pip install -e .- If you add a new dependency during development, install it and add the pinned version to
requirements.txt:
uv pip install <package-name> # Then manually add or update <package-name>==<version> in requirements.txt
Only top-level packages are listed — uv resolves transitive dependencies automatically at install time.
-
Add or update the input CSV from ODS
- If necessary, update the Qualtrics CSV received from ODS
- Rename and place it in the
data/00_input/qualtrics/folder- Name should be in the format
qualtrics_output_{version}_{type}.csv(e.g.qualtrics_output_2023-07-19_raw.csv)
- Name should be in the format
- If the Qualtrics CSV is updated, also update the values for
QUALTRICS_VERSIONandQUALTRICS_TYPEinconfig.pyto match the Qualtrics CSV as needed.
-
Download the most recent iCite Database Snapshot
- The file is a ~10GB zipped CSV and can take 2-4 hours to manually download
- Place the file in a versioned raw data directory (i.e.
data/00_input/icite/{version}/icite_metadata.zip) - Update the value of
ICITE_VERSIONinconfig.pyto match the version directory name (e.g.2023-11)
-
Get an NCBI API Key
- Follow NCBI instructions to register for an NCBI account and get an API key
- Create a new file in the INS-Data root directory named
.envand fill with the following:
NCBI_EMAIL = <your.name@nih.gov> NCBI_API_KEY = <API Key from NCBI>- Replace the values above (without <>) with your email and key
- NOTE: Because
.envis listed in.gitignore, this file will not be added to GitHub. - Never commit API keys to GitHub. Keep them on your local.
- Failure to add a valid API key here will triple the time required for the PubMed API data gathering process.
-
Add or update the input CSV from CEDCD
- Reach out to the Cancer Epidemiology Descriptive Cohort Database (CEDCD) team to ask if an updated CSV is available
- If so, add the csv to a versioned input
00_input/cedcd/folder - Update the
CEDCD_VERSIONinconfig.py
-
Run the pipeline
- In the command terminal, run the main workflow from the INS-Data root directory with:
python main.py
-
This will run all steps of the workflow in order and save all output files in locations defined in
config.py -
OPTIONAL - Instead of
main.py, modules can be run as independent processes with the following commands:python modules/gather_program_data.py python modules/gather_grant_data.py python modules/build_program_project_stats.py python modules/gather_project_data.py python modules/gather_publication_data.py python modules/gather_geo_data.py python modules/gather_sra_data.py python modules/gather_cedcd_data.py python modules/package_output_data.py python modules/build_validation_file.py
-
NOTE: If running modules independently, ensure that necessary output files from preceding modules already exist for the same start date (version)
-
Gather/Curate dbGaP Datasets (optional)
- On the NCBI dbGaP Advanced Search, filter for "NIH Institute: NCI" and download the CSV with the "Save Results" button. Store this in
data/00_input/dbgap/study_{DBGAP_CSV_VERSION}.csvwhereDBGAP_CSV_VERSIONis the date of download - Update the
DBGAP_CSV_VERSIONinconfig.py - Obtain an updated GPA study list from a point of contact at the NCI Office of Data Sharing and store as
data/00_input/dbgap/gpa_tables/gpa_study_table.csv - If needed, manually update the
gpa_doc_lookup_table.csvin the same folder - Run the dbGaP gathering module:
python modules/gather_dbgap_data.py
- Merge the new gathered output with the previously curated dbGaP file:
python modules/merge_curated_dbgap.py
- Review the merge output and the title review CSV if generated. Manually curate the merged file as needed and save as
dbgap_datasets_merged_curated.tsvin the versioneddata/01_intermediate/dbgap/directory. - Run the data packaging module, which will auto-detect and use the curated version if available:
python modules/package_output_data.py
- If applicable, place subset CSVs (e.g.
CRDC_subset_{version}.csv,GDC_subset_{version}.csv) indata/00_input/dbgap/and generate cloned entries for other repositories:
python modules/gather_dbgap_subset_clones.py
- On the NCBI dbGaP Advanced Search, filter for "NIH Institute: NCI" and download the CSV with the "Save Results" button. Store this in
-
Process CTD^2 Datasets (optional)
- This step is only needed when CTD^2 dataset or file metadata changes
- Place curated CSVs (
ctd2_datasets_{version}.csvandctd2_filedata_{version}.csv) indata/00_input/ctd2/ - Update the
CTD2_VERSIONinconfig.py - Run the CTD^2 module:
python modules/gather_ctd2_data.py
-
Add or update curated source files (optional)
- DCEG Cohorts: Place the curated TSV as
dceg_datasets_curated.tsvin a versioneddata/01_intermediate/dceg_cohorts/directory. UpdateDCEG_COHORTS_VERSIONinconfig.py. - NCCR Data Platform: Place the curated TSV as
nccr_datasets_curated.tsvin a versioneddata/01_intermediate/nccr/directory. UpdateNCCR_VERSIONinconfig.py. - Curated sources are processed automatically during the Package Data step.
- DCEG Cohorts: Place the curated TSV as
Tests are located in the tests/ directory and use pytest. Each test file corresponds to a module in modules/.
Run all tests with coverage and verbose output:
pytest --cov=modules -vRun a single test file:
pytest tests/test_gather_project_data.py -vSome tests are marked with @pytest.mark.live_api and make real calls to external APIs (NIH RePORTER, PubMed, GEO, dbGaP, SRA) to verify they are reachable and returning expected response schemas. These run by default with pytest and take a few seconds.
To run only the live API tests:
pytest -m live_api -vTo skip them (e.g., when offline):
pytest -m "not live_api"The test suite includes checks that NCBI_API_KEY and NCBI_EMAIL are set in your .env file. These are grouped with the live_api marker and run by default locally. If these variables are missing, those tests will fail with a message pointing to the setup instructions. They are skipped in CI where .env is not available.
During the PMID gathering step, you may see an error of the form
Received a 500 error for <project_id>. Retrying after 2 seconds. Attempt <n>/5
This error is caused by issues with the NIH RePORTER API. When we previously had this issue, we sent a message to NIH RePORTER support, and they fixed the problem.
During the PubMed data gathering step, you may see an error of the form
Error fetching information for PMID <pmid>: list index out of range
This error can be safely ignored. The out of range list indices come from NIH RePORTER returning old PMID's. These PMID's are generally from before the year 2000 and would be ignored later in the script even if they were gathered.
If the NIH RePORTER API is unresponsive, it could be because the API is receiving too many requests at a time. It could be worth trying to run this data gathering script during a time when the API is likely to receive fewer requests, such as overnight.

