Documentation
Output format: study.csv, data-dictionary.json and the manifest.
Clinical Extract writes three files into the study folder: the CSV, the dictionary that describes every column of it, and the bulk export manifest the EHR returned for the run. This page sets out the rules for each file.
1 The three files
One folder per study.
A run writes study-{id}/ with study.csv, data-dictionary.json, the NDJSON files it downloaded (one per resource type, one resource per line) and the manifest. Figure 1 shows the shape of each file with an excerpt from the synthetic breast cancer study.
2 study.csv
The CSV rules.
- Encoding and shape. UTF-8, comma-separated, RFC 4180 quoting; a header row, then one row per patient in the cohort.
- Columns. One per approved variable, in the dictionary's order;
patient_reffirst, the FHIR Patient id on your server. - Values. Coded values as their display text with the code system in the dictionary description; dates as ISO 8601 (
2024-03-14); numbers unformatted; an empty cell where the record holds no value. - Stability. The same column name means the same thing in every run of the study; a changed variable list is a new dictionary and a new file.
3 data-dictionary.json
The dictionary schema.
A JSON object with one entry per CSV column, keyed by the column name. Each entry has three fields: description (the meaning and its code system), fhirSource (the FHIR element the values are read from, as resource and path) and example (a sample value in the column's form). Keys are lowercase with underscores.
{ "<column name>": { "description": "<meaning, with the code system>", "fhirSource": "<Resource.element[.path]>", "example": "<a sample value in the column's form>" } }
4 The manifest
Kept as the EHR returned it.
The manifest is the HL7 Bulk Data Access status response: transactionTime (the time the data is as of), request (the kickoff URL with its _type), requiresAccessToken, output (one entry per NDJSON file: type, url, count) and error (OperationOutcome files). Clinical Extract keeps it with the files, unchanged, so each file traces to the request and the moment it came from. The exchange that produces it is on the bulk export guide.
5 Questions
Questions about the files
Why JSON for the dictionary and CSV for the data?
The data is a table, and every analysis tool opens a CSV. The dictionary is a record per column with named fields, and JSON keeps those fields named, keyed by the column, without a second header row to parse.
Can we choose the column names?
They are the keys of the approved variable list: the dictionary carries whatever names the protocol uses, lowercase with underscores, stable across runs.
What does a column hold when a patient has several results?
Each column’s rule (which result, which window, which encounter) is in its description, so the same column means the same thing in every row. A protocol that wants every result gets one column per occurrence, numbered.
Is the manifest changed in any way?
No. The manifest is the JSON the EHR’s status endpoint returned when the export completed, kept with the files so each file traces to the request and the time the data is as of.
Next
See Clinical Extract run on one of your studies.
Tell us which EHR you run and what the study or registry needs. We reply within one business day to set a meeting time.