In the box

One CSV and the FHIR data dictionary that describes it.

Clinical Extract writes data-dictionary.json with every study.csv: one entry per column with its description, the FHIR element it came from and an example value. Your IRB, honest broker and analysts all work from this one file.

1 One file

Each column's meaning, FHIR source and example value.

study.csv has one row per patient in the cohort and one column per approved variable. data-dictionary.json sits next to it and records, for every column, what it holds, which FHIR element it came from and what a value looks like. Two columns can share a FHIR element, as the stage group and the ER status do; the LOINC code in each description tells them apart. Reviewers can trace any cell to its FHIR resource, and analysts can read each definition without asking the data team.

study.csv and its dictionaryRows 1 to 4 and seven of the 24 columns of study.csv, a 212-row synthetic breast cancer study, with each column joined by a line to its entry in data-dictionary.json below it. Each entry has a name, a description, a fhirSource and an example. patient_ref: Patient resource ID on your FHIR server, from Patient.id, example eKx3p9Qa. birth_year: year of birth, from Patient.birthDate, example 1961. dx_code: primary cancer diagnosis, ICD-10-CM, from Condition.code, example C50.412. dx_date: date of diagnosis, from Condition.onsetDateTime, example 2024-03-14. stage_group: AJCC clinical stage group, LOINC 21908-9, from Observation.valueCodeableConcept, example IIA. er_status, highlighted: estrogen receptor status, LOINC 16112-5, from Observation.valueCodeableConcept, example Positive. first_chemo: first chemotherapy agent ordered, RxNorm, from MedicationRequest.medicationCodeableConcept, example paclitaxel. The CSV rows read: eKx3p9Qa, 1961, C50.412, 2024-03-14, IIA, Positive, paclitaxel; b7Tq2LmR, 1954, C50.911, 2024-05-02, IIIB, Negative, docetaxel; Vn4c8WzE, 1970, C50.212, 2024-06-21, IB, Positive, and an empty first_chemo; Hm4r1XoP, 1948, C50.511, 2024-08-09, IIB, Positive, doxorubicin.study.csv   212 rows, 24 columns: rows 1 to 4, 7 columns shownrowpatient_refbirth_yeardx_codedx_datestage_grouper_statusfirst_chemo1eKx3p9Qa1961C50.4122024-03-14IIAPositivepaclitaxel2b7Tq2LmR1954C50.9112024-05-02IIIBNegativedocetaxel3Vn4c8WzE1970C50.2122024-06-21IBPositive4Hm4r1XoP1948C50.5112024-08-09IIBPositivedoxorubicindata-dictionary.jsonone entry per columnnamepatient_refbirth_yeardx_codedx_datestage_grouper_statusfirst_chemodescriptionPatientresource IDon your FHIRserverYear ofbirthPrimary cancerdiagnosis,ICD-10-CMDate ofdiagnosisAJCC clinical stagegroup (LOINC21908-9)Estrogen receptorstatus (LOINC16112-5)First chemotherapy agentordered (RxNorm)fhirSourcePatient.idPatient.birthDateCondition.codeCondition.onsetDateTimeObservation.valueCodeableConceptObservation.valueCodeableConceptMedicationRequest.medicationCodeableConceptexampleeKx3p9Qa1961C50.4122024-03-14IIAPositivepaclitaxel
study.csv and its dictionaryRows 1 to 4 and seven of the 24 columns of study.csv, a 212-row synthetic breast cancer study, with each column joined by a line to its entry in data-dictionary.json below it. Each entry has a name, a description, a fhirSource and an example. patient_ref: Patient resource ID on your FHIR server, from Patient.id, example eKx3p9Qa. birth_year: year of birth, from Patient.birthDate, example 1961. dx_code: primary cancer diagnosis, ICD-10-CM, from Condition.code, example C50.412. dx_date: date of diagnosis, from Condition.onsetDateTime, example 2024-03-14. stage_group: AJCC clinical stage group, LOINC 21908-9, from Observation.valueCodeableConcept, example IIA. er_status, highlighted: estrogen receptor status, LOINC 16112-5, from Observation.valueCodeableConcept, example Positive. first_chemo: first chemotherapy agent ordered, RxNorm, from MedicationRequest.medicationCodeableConcept, example paclitaxel. The CSV rows read: eKx3p9Qa, 1961, C50.412, 2024-03-14, IIA, Positive, paclitaxel; b7Tq2LmR, 1954, C50.911, 2024-05-02, IIIB, Negative, docetaxel; Vn4c8WzE, 1970, C50.212, 2024-06-21, IB, Positive, and an empty first_chemo; Hm4r1XoP, 1948, C50.511, 2024-08-09, IIB, Positive, doxorubicin.study.csv  212 rows, 24 columnsdata-dictionary.json  one entry eachSeven columns: rows 1 to 4, then the entrypatient_refeKx3p9Qa, b7Tq2LmR, Vn4c8WzE, Hm4r1XoPdescriptionPatient resource ID on your FHIRserverfhirSourcePatient.idexampleeKx3p9Qabirth_year1961, 1954, 1970, 1948descriptionYear of birthfhirSourcePatient.birthDateexample1961dx_codeC50.412, C50.911, C50.212, C50.511descriptionPrimary cancer diagnosis,ICD-10-CMfhirSourceCondition.codeexampleC50.412dx_date2024-03-14, 2024-05-02, 2024-06-21, 2024-08-09descriptionDate of diagnosisfhirSourceCondition.onsetDateTimeexample2024-03-14stage_groupIIA, IIIB, IB, IIBdescriptionAJCC clinical stage group (LOINC21908-9)fhirSourceObservation.valueCodeableConceptexampleIIAer_statusPositive, Negative, Positive, PositivedescriptionEstrogen receptor status (LOINC16112-5)fhirSourceObservation.valueCodeableConceptexamplePositivefirst_chemopaclitaxel, docetaxel, (empty), doxorubicindescriptionFirst chemotherapy agent ordered(RxNorm)fhirSourceMedicationRequest.medicationCodeableConceptexamplepaclitaxel
Figure 1. study.csv and its dictionary. Every column in study.csv has an entry in data-dictionary.json that says what it holds, which FHIR element it came from and what a value looks like. Two columns can share a FHIR element, as stage_group and er_status do; the LOINC code in the description tells them apart. Seven of 24 columns shown; every value is synthetic.

2 The entry

The four parts of an entry.

The key
The CSV column name: er_status. Lowercase, underscores, stable across runs.
description
The meaning and its code system: "Estrogen receptor status (LOINC 16112-5)".
fhirSource
The FHIR element the values are read from: Observation.valueCodeableConcept.
example
A sample value in the column's form: "Positive", "2024-03-14", "C50.412".
Listing 1. Two entries of data-dictionary.json for the synthetic breast cancer study: the same FHIR element, told apart by the LOINC code in the description.
{
  "er_status": {
    "description": "Estrogen receptor status (LOINC 16112-5)",
    "fhirSource": "Observation.valueCodeableConcept",
    "example": "Positive"
  },
  "stage_group": {
    "description": "AJCC clinical stage group (LOINC 21908-9)",
    "fhirSource": "Observation.valueCodeableConcept",
    "example": "IIA"
  }
}

3 The IRB

A research data dictionary the IRB reads before the export.

The dictionary exists before the export runs: it is the approved variable list, one entry per variable, and the IRB application and the honest broker's review quote the file the analyst receives. A variable the protocol did not approve has no entry and no column. The informatics team page shows how an honest broker reviews the extract. The steps between the token and the file are on the how-it-works page; the file rules are in the output format.

4 REDCap and your EDC

From data-dictionary.json to a REDCap data dictionary.

A REDCap data dictionary is a CSV with one row per field: variable name, form name, field type, field label, choices. Every entry in data-dictionary.json carries the variable name (the key) and the label (the description), so the REDCap dictionary is a row per entry, with the field type set from the example, and study.csv imports as the records. An EDC that takes a CSV with a field map works the same way.

5 Questions

Questions about the files

What is a FHIR data dictionary?

A dictionary whose entries point at FHIR elements: for every column, the description with its code system, the FHIR element the values are read from (fhirSource) and an example value. data-dictionary.json is one, keyed by the CSV column names.

What is a research data dictionary?

The document that defines every variable in a study dataset: its name, meaning, type, units or code system, and origin. The IRB and the analyst both need it, and Clinical Extract writes it with the CSV, as part of the same export.

Can I load the CSV into REDCap?

A REDCap data dictionary is a CSV with a row per field (variable name, form, field type, field label, choices). data-dictionary.json carries the variable name and the label for every column, so you build the REDCap dictionary with a row per entry and import the values from study.csv.

What is fhirSource?

The FHIR element a column’s values are read from, written as resource and path: Condition.code, Observation.valueCodeableConcept, MedicationRequest.medicationCodeableConcept. Two columns can share a source (ER status and stage group are both Observation values); the LOINC code in the description tells them apart.

Are the values codes or text?

Coded values arrive as their display text with the code system named in the description (LOINC 16112-5: Positive), dates as ISO 8601, and identifiers as the FHIR resource id. Each column’s rule is in its entry.

Next

See Clinical Extract run on one of your studies.

Tell us which EHR you run and what the study or registry needs. We reply within one business day to set a meeting time.

Request a demo