Documentation

Output format: study.csv, data-dictionary.json and the manifest.

Clinical Extract writes three files into the study folder: the CSV, the dictionary that describes every column of it, and the bulk export manifest the EHR returned for the run. This page sets out the rules for each file.

1 The three files

One folder per study.

A run writes study-{id}/ with study.csv, data-dictionary.json, the NDJSON files it downloaded (one per resource type, one resource per line) and the manifest. Figure 1 shows the shape of each file with an excerpt from the synthetic breast cancer study.

Output filesA ruled table of the three files in the study folder study-0142/, with each file’s shape, fields and an excerpt. study.csv, 212 rows and 24 columns: CSV, UTF-8, a header row, then one row per patient in the study cohort; one column per approved variable, among them patient_ref, birth_year, gender, dx_code, dx_date, stage_group, er_status, pr_status, her2_status and first_chemo; the excerpt shows the header row and the row eKx3p9Qa, 1961, female, C50.412, 2024-03-14, IIA, Positive, Positive, Negative, paclitaxel. data-dictionary.json, 24 entries, highlighted: a JSON object with one entry per CSV column, keyed by the column name, each with description (meaning and code system), fhirSource (the FHIR element the value is read from) and example (a sample value); the excerpt is the er_status entry: description Estrogen receptor status (LOINC 16112-5), fhirSource Observation.valueCodeableConcept, example Positive. manifest.json, 4 output files: the bulk export manifest the EHR returned for the run, with transactionTime (the time the data is as of), request (the kickoff URL), requiresAccessToken, output (one entry per NDJSON file: type, url and count) and error (OperationOutcome files); the excerpt shows transactionTime 2026-09-28T14:02:11Z, the request URL [base]/Group/{cohort-id}/$export with _type Patient, Condition, Observation and MedicationRequest, requiresAccessToken true, one output entry written out, type Observation, url Observation_1.ndjson, count 48310, a note that 3 more entries follow for Patient, Condition and MedicationRequest, and an empty error list.study-0142/  the study folderFileShapeFieldsExcerptstudy.csv212 rows, 24 columnsCSV, UTF-8; a headerrow, then one rowper patient in thecohortone column per approved variable:patient_ref, birth_year, gender,dx_code, dx_date, stage_group,er_status, pr_status,her2_status, first_chemo, ...patient_ref,birth_year,gender,dx_code,  dx_date,stage_group,er_status,pr_status,  her2_status,first_chemoeKx3p9Qa,1961,female,C50.412,2024-03-14,IIA,  Positive,Positive,Negative,paclitaxeldata-dictionary.json24 entriesJSON object; oneentry per CSVcolumn, keyed by thecolumn namekey  the CSV column namedescription  meaning, code systemfhirSource  the FHIR elementexample  a sample value{  "er_status": {    "description": "Estrogen receptor status      (LOINC 16112-5)",    "fhirSource":      "Observation.valueCodeableConcept",    "example": "Positive"  }}manifest.json4 output filesJSON; the bulkexport manifest theEHR returned for theruntransactionTime  data as ofrequest  the kickoff URLrequiresAccessTokenoutput[]  type, url, counterror[]  OperationOutcome files{  "transactionTime": "2026-09-28T14:02:11Z",  "request": "[base]/Group/{cohort-id}/    $export?_type=Patient,Condition,    Observation,MedicationRequest",  "requiresAccessToken": true,  "output": [    { "type": "Observation",      "url": ".../Observation_1.ndjson",      "count": 48310 }    ... 3 more: Patient, Condition,      MedicationRequest  ],  "error": []}
Output filesA ruled table of the three files in the study folder study-0142/, with each file’s shape, fields and an excerpt. study.csv, 212 rows and 24 columns: CSV, UTF-8, a header row, then one row per patient in the study cohort; one column per approved variable, among them patient_ref, birth_year, gender, dx_code, dx_date, stage_group, er_status, pr_status, her2_status and first_chemo; the excerpt shows the header row and the row eKx3p9Qa, 1961, female, C50.412, 2024-03-14, IIA, Positive, Positive, Negative, paclitaxel. data-dictionary.json, 24 entries, highlighted: a JSON object with one entry per CSV column, keyed by the column name, each with description (meaning and code system), fhirSource (the FHIR element the value is read from) and example (a sample value); the excerpt is the er_status entry: description Estrogen receptor status (LOINC 16112-5), fhirSource Observation.valueCodeableConcept, example Positive. manifest.json, 4 output files: the bulk export manifest the EHR returned for the run, with transactionTime (the time the data is as of), request (the kickoff URL), requiresAccessToken, output (one entry per NDJSON file: type, url and count) and error (OperationOutcome files); the excerpt shows transactionTime 2026-09-28T14:02:11Z, the request URL [base]/Group/{cohort-id}/$export with _type Patient, Condition, Observation and MedicationRequest, requiresAccessToken true, one output entry written out, type Observation, url Observation_1.ndjson, count 48310, a note that 3 more entries follow for Patient, Condition and MedicationRequest, and an empty error list.study-0142/  the study folderstudy.csv  212 rows, 24 columnsCSV, UTF-8; a header row, then one row perpatient in the cohortFIELDSone column per approved variable:patient_ref, birth_year, gender, dx_code,dx_date, stage_group, er_status, pr_status,her2_status, first_chemo, ...EXCERPTpatient_ref,birth_year,gender,dx_code,  dx_date,stage_group,er_status,pr_status,  her2_status,first_chemoeKx3p9Qa,1961,female,C50.412,2024-03-14,IIA,  Positive,Positive,Negative,paclitaxeldata-dictionary.json  24 entriesJSON object; one entry per CSV column, keyedby the column nameFIELDSkey  the CSV column namedescription  meaning, code systemfhirSource  the FHIR elementexample  a sample valueEXCERPT{  "er_status": {    "description": "Estrogen receptor status      (LOINC 16112-5)",    "fhirSource":      "Observation.valueCodeableConcept",    "example": "Positive"  }}manifest.json  4 output filesJSON; the bulk export manifest the EHRreturned for the runFIELDStransactionTime  data as ofrequest  the kickoff URLrequiresAccessTokenoutput[]  type, url, counterror[]  OperationOutcome filesEXCERPT{  "transactionTime": "2026-09-28T14:02:11Z",  "request": "[base]/Group/{cohort-id}/    $export?_type=Patient,Condition,    Observation,MedicationRequest",  "requiresAccessToken": true,  "output": [    { "type": "Observation",      "url": ".../Observation_1.ndjson",      "count": 48310 }    ... 3 more: Patient, Condition,      MedicationRequest  ],  "error": []}
Figure 1. Output files. Every run writes study.csv, data-dictionary.json with one entry per CSV column, and the bulk export manifest the EHR returned, so each file can be traced to the request and the moment it came from. Values are synthetic.

2 study.csv

The CSV rules.

  • Encoding and shape. UTF-8, comma-separated, RFC 4180 quoting; a header row, then one row per patient in the cohort.
  • Columns. One per approved variable, in the dictionary's order; patient_ref first, the FHIR Patient id on your server.
  • Values. Coded values as their display text with the code system in the dictionary description; dates as ISO 8601 (2024-03-14); numbers unformatted; an empty cell where the record holds no value.
  • Stability. The same column name means the same thing in every run of the study; a changed variable list is a new dictionary and a new file.

3 data-dictionary.json

The dictionary schema.

A JSON object with one entry per CSV column, keyed by the column name. Each entry has three fields: description (the meaning and its code system), fhirSource (the FHIR element the values are read from, as resource and path) and example (a sample value in the column's form). Keys are lowercase with underscores.

Listing 1. The dictionary schema: one entry per column, three fields each.
{
  "<column name>": {
    "description": "<meaning, with the code system>",
    "fhirSource": "<Resource.element[.path]>",
    "example": "<a sample value in the column's form>"
  }
}

4 The manifest

Kept as the EHR returned it.

The manifest is the HL7 Bulk Data Access status response: transactionTime (the time the data is as of), request (the kickoff URL with its _type), requiresAccessToken, output (one entry per NDJSON file: type, url, count) and error (OperationOutcome files). Clinical Extract keeps it with the files, unchanged, so each file traces to the request and the moment it came from. The exchange that produces it is on the bulk export guide.

5 Questions

Questions about the files

Why JSON for the dictionary and CSV for the data?

The data is a table, and every analysis tool opens a CSV. The dictionary is a record per column with named fields, and JSON keeps those fields named, keyed by the column, without a second header row to parse.

Can we choose the column names?

They are the keys of the approved variable list: the dictionary carries whatever names the protocol uses, lowercase with underscores, stable across runs.

What does a column hold when a patient has several results?

Each column’s rule (which result, which window, which encounter) is in its description, so the same column means the same thing in every row. A protocol that wants every result gets one column per occurrence, numbered.

Is the manifest changed in any way?

No. The manifest is the JSON the EHR’s status endpoint returned when the export completed, kept with the files so each file traces to the request and the time the data is as of.

Next

See Clinical Extract run on one of your studies.

Tell us which EHR you run and what the study or registry needs. We reply within one business day to set a meeting time.

Request a demo