Bulk FHIR guide

Bulk FHIR export, from kickoff to CSV.

Clinical Extract runs a study-scoped bulk FHIR export from your EHR. This guide covers each request, from the kickoff to a CSV your analysts can open.

This guide shows every kickoff, status and manifest header in full.

1 Definition

What is bulk FHIR?

Bulk FHIR is the HL7 FHIR Bulk Data Access specification, once nicknamed Flat FHIR: an asynchronous $export operation that returns a population's records as NDJSON files, one FHIR resource per line, instead of one REST call per resource. The FHIR bulk data API has three parts: a kickoff request, a status URL to poll, and the files listed in a manifest when the export is done.

Every certified EHR has it. The certification criterion ONC §170.315(g)(10) requires certified EHR technology to export data for a group of patients through HL7 FHIR Bulk Data Access, authorized with SMART Backend Services, and under 45 CFR 170.404(b)(3) developers had to make it available to their customers by December 31, 2022. So a FHIR bulk data export is the standard way to get a cohort's coded records out of any certified EHR, and it is the way Clinical Extract works: it runs the FHIR bulk data access API for the study cohort's Group, then turns the NDJSON into study.csv and a data dictionary.

A FHIR export from a cloud FHIR store (a Google Cloud or Azure bulk export) uses the same operation; this guide is about the EHR's own endpoint, which is where the chart lives.

2 Three kinds of export

System, patient and Group exports, and the one a study needs.

A system-level export ($export on the server) takes everything. A patient-level export (Patient/$export) takes every patient. A FHIR group export (Group/{id}/$export) takes the members of one Group, and that is the export the certification criterion requires and the one a study needs. Clinical Extract runs the Group export and adds _type to request only the resource types the study reads, so the request covers only the cohort and its variables.

Three kinds of exportThe same EHR drawn three times as a grid: rows are resource types, columns are patients. The rows in the Patient compartment are Patient, Encounter, Condition, Observation, MedicationRequest, Procedure, DiagnosticReport and Immunization; the rows outside it are Practitioner, Organization, Location and Medication. a, system level, GET [base]/$export: every cell is exported, every resource on the server, patient data or not. b, patient level, GET [base]/Patient/$export: every patient column of the compartment rows is exported, and referenced resources outside the compartment are returned at the server’s option. c, group level, GET [base]/Group/{cohort-id}/$export, highlighted: only the columns of the Group’s members are exported, 212 of 125,000 patients.aSystem levelGET [base]/$exportbPatient levelGET [base]/Patient/$exportcGroup levelGET [base]/Group/{cohort-id}/$exportIN THE PATIENT COMPARTMENTOUTSIDE ITPatientEncounterConditionObservationMedicationRequestProcedureDiagnosticReportImmunizationPractitionerOrganizationLocationMedicationticks: the Group’s memberscolumns: patientsevery resource on the server,patient data or notall 125,000 patientsthe compartment of everypatientall 125,000 patientsthe compartments of theGroup’s members only212 of 125,000 patientsexportedexported: the study cohortreferenced, at the server’s optionnot exported
Three kinds of exportThe same EHR drawn three times as a grid: rows are resource types, columns are patients. The rows in the Patient compartment are Patient, Encounter, Condition, Observation, MedicationRequest, Procedure, DiagnosticReport and Immunization; the rows outside it are Practitioner, Organization, Location and Medication. a, system level, GET [base]/$export: every cell is exported, every resource on the server, patient data or not. b, patient level, GET [base]/Patient/$export: every patient column of the compartment rows is exported, and referenced resources outside the compartment are returned at the server’s option. c, group level, GET [base]/Group/{cohort-id}/$export, highlighted: only the columns of the Group’s members are exported, 212 of 125,000 patients.aSystem levelGET [base]/$exportIN THE COMPARTMENTOUTSIDE ITPatientEncounterConditionObservationMedicationRequestProcedureDiagnosticReportImmunizationPractitionerOrganizationLocationMedicationevery resource on the server,patient data or notall 125,000 patientsbPatient levelGET [base]/Patient/$exportIN THE COMPARTMENTOUTSIDE ITPatientEncounterConditionObservationMedicationRequestProcedureDiagnosticReportImmunizationPractitionerOrganizationLocationMedicationthe compartment of every patientall 125,000 patientscGroup levelGET [base]/Group/{cohort-id}/$exportIN THE COMPARTMENTOUTSIDE ITPatientEncounterConditionObservationMedicationRequestProcedureDiagnosticReportImmunizationPractitionerOrganizationLocationMedicationticks: the Group’s membersthe compartments of the Group’s members only212 of 125,000 patientsexportedexported: the study cohortreferenced, at the server’s optionnot exported
Figure 1. Three kinds of export. The same EHR, three kickoff requests: a system-level export takes everything on the server, a patient-level export takes every patient, and a Group export takes the study cohort alone. Clinical Extract runs the Group export and adds _type to request only the resource types the study needs. The grid is schematic; the patient counts are illustrative.

3 Kickoff to files

The FHIR bulk export request pattern and its headers.

The kickoff is one GET with Prefer: respond-async and Accept: application/fhir+json. The server answers 202 Accepted at once, with a Content-Location header naming the status URL. Clinical Extract polls that URL; while the export runs the server answers 202 with Retry-After (and, on Epic, X-Progress), and Clinical Extract waits the interval Retry-After specifies. When the export is done the status URL answers 200 with the manifest: transactionTime (the time the data is as of), request (the kickoff URL), requiresAccessToken, output (one entry per NDJSON file: its type, its URL, its line count) and error (OperationOutcome files). Each file is fetched with Accept: application/fhir+ndjson and the bearer token, and the status URL is deleted once the files are safe in your study folder.

Kickoff to filesTwo panels. Panel a, the async pattern in four exchanges between Clinical Extract and your EHR. 1, kickoff: GET /Group/{cohort-id}/$export with Prefer: respond-async; the EHR answers 202 Accepted with Content-Location: {status-url}. 2, poll, repeated until 200 OK: GET {status-url}; the EHR answers 202 Accepted with Retry-After: 30 and X-Progress: 75% complete. 3, complete: GET {status-url}; the EHR answers 200 OK, Content-Type: application/json, and the body is the manifest in panel b. 4, download, once per output entry: GET {output url} with Accept: application/fhir+ndjson; the EHR answers 200 OK with one resource per line. Panel b, the manifest beside the files it lists. The manifest holds transactionTime 2026-09-28T14:02:11Z, annotated: no resource in the files is newer than this; request, the kickoff URL with _type=Patient,Condition,Observation,MedicationRequest; requiresAccessToken true, annotated: each file GET sends the bearer token; the output array, highlighted, with one entry per file giving type, url and count; and an empty error array, annotated: OperationOutcome files, if any; none here. An arrow runs from each output url to its NDJSON file, whose lines are numbered from 1 to the entry’s count, each line one resource of the entry’s type: Patient_1.ndjson, 212 lines; Condition_1.ndjson, 1,684 lines; Observation_1.ndjson, 48,310 lines; MedicationRequest_1.ndjson, 3,905 lines.aThe async patternfour exchanges, left to rightClinical Extract to your EHRyour EHR to Clinical Extract1KICKOFFGET /Group/{cohort-id}/$exportPrefer: respond-async202 AcceptedContent-Location: {status-url}2POLLrepeats until 200 OKGET {status-url}202 AcceptedRetry-After: 30X-Progress: 75% complete3COMPLETEGET {status-url}200 OKContent-Type: application/jsonthe body is the manifest, b4DOWNLOADonce per output entryGET {output url}Accept: application/fhir+ndjson200 OKone resource per linebThe manifest and the files it liststhe 200 OK body; each output entry is one file{  "transactionTime": "2026-09-28T14:02:11Z",  "request": "[base]/Group/{cohort-id}/$export?_type=Patient,Condition,    Observation,MedicationRequest",  "requiresAccessToken": true,  "output": [    {      "type": "Patient",      "url": ".../Patient_1.ndjson",      "count": 212    },    {      "type": "Condition",      "url": ".../Condition_1.ndjson",      "count": 1684    },    {      "type": "Observation",      "url": ".../Observation_1.ndjson",      "count": 48310    },    {      "type": "MedicationRequest",      "url": ".../MedicationRequest_1.ndjson",      "count": 3905    }  ],  "error": []}no resource in the files is newer than thiseach file GET sends the bearer tokenOperationOutcome files, if any; none hereTHE NDJSON FILES, IN YOUR ENVIRONMENTPatient_1.ndjson212 lines1{"resourceType":"Patient","id":"eKx3p9Qa",...}...212{"resourceType":"Patient",...}Condition_1.ndjson1,684 lines1{"resourceType":"Condition","id":"c41a08",...}...1,684{"resourceType":"Condition",...}Observation_1.ndjson48,310 lines1{"resourceType":"Observation","id":"9f2d71",...}...48,310{"resourceType":"Observation",...}MedicationRequest_1.ndjson3,905 lines1{"resourceType":"MedicationRequest","id":"m7b203",...}...3,905{"resourceType":"MedicationRequest",...}
Kickoff to filesTwo panels. Panel a, the async pattern in four exchanges between Clinical Extract and your EHR. 1, kickoff: GET /Group/{cohort-id}/$export with Prefer: respond-async; the EHR answers 202 Accepted with Content-Location: {status-url}. 2, poll, repeated until 200 OK: GET {status-url}; the EHR answers 202 Accepted with Retry-After: 30 and X-Progress: 75% complete. 3, complete: GET {status-url}; the EHR answers 200 OK, Content-Type: application/json, and the body is the manifest in panel b. 4, download, once per output entry: GET {output url} with Accept: application/fhir+ndjson; the EHR answers 200 OK with one resource per line. Panel b, the manifest beside the files it lists. The manifest holds transactionTime 2026-09-28T14:02:11Z, annotated: no resource in the files is newer than this; request, the kickoff URL with _type=Patient,Condition,Observation,MedicationRequest; requiresAccessToken true, annotated: each file GET sends the bearer token; the output array, highlighted, with one entry per file giving type, url and count; and an empty error array, annotated: OperationOutcome files, if any; none here. An arrow runs from each output url to its NDJSON file, whose lines are numbered from 1 to the entry’s count, each line one resource of the entry’s type: Patient_1.ndjson, 212 lines; Condition_1.ndjson, 1,684 lines; Observation_1.ndjson, 48,310 lines; MedicationRequest_1.ndjson, 3,905 lines.aThe async patternfour exchanges, top to bottomClinical Extract to your EHRyour EHR to Clinical Extract1KICKOFFGET /Group/{cohort-id}/$exportPrefer: respond-async202 AcceptedContent-Location: {status-url}2POLLrepeats until 200 OKGET {status-url}202 AcceptedRetry-After: 30X-Progress: 75% complete3COMPLETEGET {status-url}200 OKContent-Type: application/jsonthe body is the manifest, b4DOWNLOADonce per output entryGET {output url}Accept: application/fhir+ndjson200 OKone resource per linebThe manifest and its fileseach output entry is one file{  "transactionTime": "2026-09-28T14:02:11Z",  "request": "[base]/Group/{cohort-id}/$export    ?_type=Patient,Condition,Observation,    MedicationRequest",  "requiresAccessToken": true,  "output": [    {      "type": "Patient",      "url": ".../Patient_1.ndjson",      "count": 212    },    {      "type": "Condition",      "url": ".../Condition_1.ndjson",      "count": 1684    },    {      "type": "Observation",      "url": ".../Observation_1.ndjson",      "count": 48310    },    {      "type": "MedicationRequest",      "url": ".../MedicationRequest_1.ndjson",      "count": 3905    }  ],  "error": []}THE NDJSON FILES, IN YOUR ENVIRONMENTPatient_1.ndjson212 lines1{"resourceType":"Patient",...}...212{"resourceType":"Patient",...}Condition_1.ndjson1,684 lines1{"resourceType":"Condition",...}...1,684{"resourceType":"Condition",...}Observation_1.ndjson48,310 lines1{"resourceType":"Observation",...}...48,310{"resourceType":"Observation",...}MedicationRequest_1.ndjson3,905 lines1{"resourceType":"MedicationRequest",...}...3,905{"resourceType":"MedicationRequest",...}
Figure 2. Kickoff to files. (a) A bulk export runs asynchronously: the kickoff returns at once with a status URL, Clinical Extract polls it at the interval Retry-After sets, and the completed status returns the manifest. (b) The manifest is the export’s table of contents. Each output entry names one NDJSON file: its url is the file, its type is the resourceType of every line, and its count is the number of lines. Headers and manifest fields follow the Bulk Data Access specification; the URLs, IDs, header values, timestamp and counts are illustrative.
Listing 1. The kickoff and the first status poll. Every value in braces is what your EHR returns; the headers are the specification's own.
GET {fhir-base}/Group/{cohort-id}/$export
    ?_type=Patient,Condition,Observation,MedicationRequest
Accept: application/fhir+json
Prefer: respond-async
Authorization: Bearer {access token}

202 Accepted
Content-Location: {status-url}

GET {status-url}
202 Accepted
Retry-After: 120
X-Progress: {what the server reports}
Listing 2. The completed status response: the manifest for a 212-patient study, one NDJSON file per requested type. Counts are illustrative.
{
  "transactionTime": "2026-09-28T14:02:11Z",
  "request": "{fhir-base}/Group/{cohort-id}/$export?_type=Patient,Condition,Observation,MedicationRequest",
  "requiresAccessToken": true,
  "output": [
    { "type": "Patient", "url": "{file-url-1}", "count": 212 },
    { "type": "Condition", "url": "{file-url-2}", "count": 1684 },
    { "type": "Observation", "url": "{file-url-3}", "count": 48310 },
    { "type": "MedicationRequest", "url": "{file-url-4}", "count": 3905 }
  ],
  "error": []
}

4 NDJSON to CSV

Turning NDJSON lines into CSV rows.

Each line of an NDJSON file is one complete FHIR resource. Clinical Extract joins it to its patient's row by subject.reference, uses the LOINC code (or the RxNorm, ICD-10-CM or SNOMED CT code) to pick the column, writes the coded value into the cell, and records where the column came from in data-dictionary.json: the description with its code system, the FHIR element, an example. The result is study.csv, with a row for each patient in the cohort and a column for each approved variable, and every column described. The steps between the token and the file are on the how-it-works page.

NDJSON to CSVPanel a: line 1 of Observation_1.ndjson, one JSON object on one line: resourceType Observation, id 9f2d71, status final, category laboratory, code LOINC 16112-5 Estrogen receptor [Interpretation] in Tissue, subject Patient/eKx3p9Qa, effectiveDateTime 2024-03-20, valueCodeableConcept SNOMED CT 10828004 Positive. Panel b: study.csv, row 1 of 212, with the four parts of line 1 behind its cells drawn out and each traced to its column. subject.reference Patient/eKx3p9Qa joins the row: patient_ref eKx3p9Qa. The LOINC code 16112-5 picks the er_status column, and the display Positive, highlighted, fills its cell. effectiveDateTime 2024-03-20 dates the result: er_date 2024-03-20. Panel c: the data-dictionary.json entry for each column. patient_ref: Patient resource ID on your FHIR server, fhirSource Patient.id, example eKx3p9Qa. er_status: Estrogen receptor status (LOINC 16112-5), fhirSource Observation.valueCodeableConcept, example Positive. er_date: Date of the estrogen receptor result, fhirSource Observation.effectiveDateTime, example 2024-03-20.aObservation_1.ndjsonline 1 of 48,310: one resource, one line, wrapped to fit{"resourceType":"Observation","id":"9f2d71","status":"final","category":[{"coding":[{"system":"http://terminology.hl7.org/CodeSystem/observation-category","code":"laboratory"}]}],"code":{"coding":[{"system":"http://loinc.org","code":"16112-5","display":"Estrogen receptor [Interpretation] in Tissue"}]},"subject":{"reference":"Patient/eKx3p9Qa"},"effectiveDateTime":"2024-03-20","valueCodeableConcept":{"coding":[{"system":"http://snomed.info/sct","code":"10828004","display":"Positive"}]}}1bstudy.csvrow 1 of 212, and the part of line 1 behind each cellJOINS THE ROW"subject":{"reference":"Patient/eKx3p9Qa"}PICKS THE COLUMN"code":{"coding":[{"system":"http://loinc.org","code":"16112-5",...}]}FILLS THE CELL"valueCodeableConcept":{"coding":[{"system":"http://snomed.info/sct","code":"10828004","display":"Positive"}]}DATES THE RESULT"effectiveDateTime":"2024-03-20"patient_refeKx3p9QaPatient/eKx3p9Qa gives the rower_statusPositivethe coded displayer_date2024-03-20the result datecdata-dictionary.jsonthe entry for each column"patient_ref": {descriptionPatient resource ID on yourFHIR serverfhirSourcePatient.idexampleeKx3p9Qa}"er_status": {descriptionEstrogen receptor status(LOINC 16112-5)fhirSourceObservation.valueCodeableConceptexamplePositive}"er_date": {descriptionDate of the estrogen receptorresultfhirSourceObservation.effectiveDateTimeexample2024-03-20}
NDJSON to CSVPanel a: line 1 of Observation_1.ndjson, one JSON object on one line: resourceType Observation, id 9f2d71, status final, category laboratory, code LOINC 16112-5 Estrogen receptor [Interpretation] in Tissue, subject Patient/eKx3p9Qa, effectiveDateTime 2024-03-20, valueCodeableConcept SNOMED CT 10828004 Positive. Panel b: study.csv, row 1 of 212, with the four parts of line 1 behind its cells drawn out and each traced to its column. subject.reference Patient/eKx3p9Qa joins the row: patient_ref eKx3p9Qa. The LOINC code 16112-5 picks the er_status column, and the display Positive, highlighted, fills its cell. effectiveDateTime 2024-03-20 dates the result: er_date 2024-03-20. Panel c: the data-dictionary.json entry for each column. patient_ref: Patient resource ID on your FHIR server, fhirSource Patient.id, example eKx3p9Qa. er_status: Estrogen receptor status (LOINC 16112-5), fhirSource Observation.valueCodeableConcept, example Positive. er_date: Date of the estrogen receptor result, fhirSource Observation.effectiveDateTime, example 2024-03-20.aObservation_1.ndjsonline 1 of 48,310, wrapped to fit{"resourceType":"Observation","id":"9f2d71","status":"final","category":[{"coding":[{"system":"http://terminology.hl7.org/CodeSystem/observation-category","code":"laboratory"}]}],"code":{"coding":[{"system":"http://loinc.org","code":"16112-5","display":"Estrogen receptor [Interpretation] in Tissue"}]},"subject":{"reference":"Patient/eKx3p9Qa"},"effectiveDateTime":"2024-03-20","valueCodeableConcept":{"coding":[{"system":"http://snomed.info/sct","code":"10828004","display":"Positive"}]}}1b  study.csv and c  its dictionaryrow 1 of 212, 3 of 24 columnsJOINS THE ROW"subject":{"reference":"Patient/eKx3p9Qa"}Patient/eKx3p9Qa gives the rowpatient_refeKx3p9Qa"patient_ref": {descriptionPatient resource ID on your FHIRserverfhirSourcePatient.idexampleeKx3p9Qa}PICKS THE COLUMN"code":{"coding":[{"system":"http://loinc.org","code":"16112-5",...}]}FILLS THE CELL"valueCodeableConcept":{"coding":[{"system":"http://snomed.info/sct","code":"10828004","display":"Positive"}]}the coded displayer_statusPositive"er_status": {descriptionEstrogen receptor status (LOINC16112-5)fhirSourceObservation.valueCodeableConceptexamplePositive}DATES THE RESULT"effectiveDateTime":"2024-03-20"the result dateer_date2024-03-20"er_date": {descriptionDate of the estrogen receptorresultfhirSourceObservation.effectiveDateTimeexample2024-03-20}
Figure 3. NDJSON to CSV. Each line of an NDJSON file is one complete FHIR resource. Clinical Extract joins it to its patient’s row by subject.reference, uses the LOINC code to pick the column, writes the coded value into the cell, and records where every column came from in data-dictionary.json. Three of the study’s 24 columns shown; the resource IDs and values are synthetic.

5 Real-world throughput

Export time for a study and for a whole population.

Jones et al. (JAMIA, 2024) measured (g)(10) bulk export at five sites. Epic sites ran at 502 to 2,827 resources per minute, and Cerner above 8,000 per minute; it took the sites 2 to 119 days (mean 65) from submitting cohort criteria to the first successful bulk request. At those rates, export time grows with the number of resources requested: a 212-patient study cohort exports in under three hours at Epic's slowest reported rate and in minutes at Cerner's, while a whole population takes weeks. Clinical Extract therefore exports only the study Group, and the cohort builder fixes the cohort and its count before the first request. Source: Jones et al. J Am Med Inform Assoc. 2024, doi:10.1093/jamia/ocae040.

Real-world throughputTwo bar charts. Panel a, reported bulk export throughput in resources per minute: Epic sites ran at 502 to 2,827, and Cerner, now Oracle Health, ran above 8,000. Panel b, on a log time scale, the export time Epic’s range implies as the cohort grows: 212 patients, the study cohort, 54,111 resources by its manifest, 19 minutes to 1.8 hours, highlighted; then, at 400 resources per patient, 500 patients, 200,000 resources, 71 minutes to 6.6 hours; 5,000 patients, 2 million resources, 11.8 hours to 2.8 days; 25,000 patients, 10 million resources, 2.5 to 13.8 days; 125,000 patients, the whole population, 50 million resources, 12 to 69 days.aThroughput at Epic and Cerner sitesresources per minute02,5005,0007,50010,000Epic range across Epic sitesEpic sites: 502 to 2,827 resources per minute502 to 2,827Cerner (now Oracle Health)Cerner: above 8,000 resources per minuteabove 8,000bExport time by cohort size, at Epic’s rangelog scale10 min1 h1 day1 wk1 mo212 patients the study cohort, 54,111 resources212 patients, 54,111 resources: 19 min to 1.8 h19 min to 1.8 h500 patients 200,000 resources500 patients, 200,000 resources: 71 min to 6.6 h71 min to 6.6 h5,000 patients 2 million resources5,000 patients, 2 million resources: 11.8 h to 2.8 days11.8 h to 2.8 days25,000 patients 10 million resources25,000 patients, 10 million resources: 2.5 to 13.8 days2.5 to 13.8 days125,000 patients the whole population, 50 million resources125,000 patients, 50 million resources: 12 to 69 days12 to 69 days
Real-world throughputTwo bar charts. Panel a, reported bulk export throughput in resources per minute: Epic sites ran at 502 to 2,827, and Cerner, now Oracle Health, ran above 8,000. Panel b, on a log time scale, the export time Epic’s range implies as the cohort grows: 212 patients, the study cohort, 54,111 resources by its manifest, 19 minutes to 1.8 hours, highlighted; then, at 400 resources per patient, 500 patients, 200,000 resources, 71 minutes to 6.6 hours; 5,000 patients, 2 million resources, 11.8 hours to 2.8 days; 25,000 patients, 10 million resources, 2.5 to 13.8 days; 125,000 patients, the whole population, 50 million resources, 12 to 69 days.aThroughput at Epic and Cerner sitesresources per minute05,00010,000Epic range across Epic sitesEpic sites: 502 to 2,827 resources per minute502 to 2,827Cerner (now Oracle Health)Cerner: above 8,000 resources per minuteabove 8,000bExport time at Epic’s rangelog scale10 min1 h1 day1 mo212 patients the study cohort212 patients, 54,111 resources: 19 min to 1.8 h19 min to 1.8 h500 patients500 patients, 200,000 resources: 71 min to 6.6 h71 min to 6.6 h5,000 patients5,000 patients, 2 million resources: 11.8 h to 2.8 days11.8 h to 2.8 days25,000 patients25,000 patients, 10 million resources: 2.5 to 13.8 days2.5 to 13.8 days125,000 patients the whole population125,000 patients, 50 million resources: 12 to 69 days12 to 69 days
Figure 4 as a table
SeriesValueBasis
Epic sites, throughput502 to 2,827 resources per minuteReported, Jones et al.
Cerner (now Oracle Health), throughputabove 8,000 resources per minuteReported, Jones et al.
212 patients, the study cohort, 54,111 resources19 min to 1.8 hArithmetic on Epic’s range, the study’s manifest count
500 patients, 200,000 resources71 min to 6.6 hArithmetic on Epic’s range, 400 resources per patient
5,000 patients, 2 million resources11.8 h to 2.8 daysArithmetic on Epic’s range, 400 resources per patient
25,000 patients, 10 million resources2.5 to 13.8 daysArithmetic on Epic’s range, 400 resources per patient
125,000 patients, the whole population, 50 million resources12 to 69 daysArithmetic on Epic’s range, 400 resources per patient
Figure 4. Real-world throughput. (a) Bulk export throughput reported at Epic and Cerner sites. (b) Export time at Epic’s reported range grows with the number of resources: a 212-patient study cohort exports in under three hours, and the whole population takes weeks. Panel b is illustrative arithmetic: the study cohort at its manifest count, 54,111 resources (Listing 2), and the larger cohorts at 400 resources per patient: 200,000 resources for a 500-patient study, 50 million for a whole population of 125,000. Source for (a): Jones et al. J Am Med Inform Assoc. 2024. doi:10.1093/jamia/ocae040.

6 By EHR

The same API on every certified EHR, in each vendor's terms.

Each EHR page shows the registration, the activation and the export in that vendor's own vocabulary, with a dated fact table and the specimens: Epic, Oracle Health, eClinicalWorks, athenahealth, MEDITECH, NextGen and Veradigm EHR. How each one exposes the API is the hub. For eClinicalWorks, the $export exchange request by request is in our developer guide for eClinicalWorks.

7 Questions

Questions about bulk FHIR

What does bulk data mean here?

Many patients’ resources at once, as files. Instead of one request per patient and resource, the bulk data access operation returns NDJSON files, one FHIR resource per line, for a whole population, a Group of patients, or the system.

What is the FHIR bulk API?

The FHIR bulk API is the $export operation from HL7 Bulk Data Access: a kickoff, a status URL to poll, and NDJSON files listed in a manifest. Clinical Extract runs it for the study cohort’s Group.

What is Flat FHIR?

The nickname the Bulk Data Access specification carried in its early drafts: "flat" because NDJSON files hold one resource per line with no bundle around them. The specification is HL7 FHIR Bulk Data Access; the operation is $export.

What is a FHIR Group export?

GET Group/{id}/$export: the bulk export scoped to the members of one Group resource. It is the export every certified EHR must offer, and the one Clinical Extract runs, with the study cohort as the Group.

How long does a bulk FHIR export take?

It scales with the number of resources requested. Jones et al. measured 502 to 2,827 resources per minute at Epic sites and above 8,000 at Cerner; a 212-patient study cohort exports in under three hours at Epic’s reported rates (Figure 4), and a whole population takes weeks.

Do we need a data warehouse first?

No. The bulk FHIR API reads the EHR’s own store, and a study-scoped export lands in your study folder as a CSV and its data dictionary. A study export does not need a warehouse, which is how informatics teams run study pulls without the warehouse queue.

Next

See Clinical Extract run on one of your studies.

Tell us which EHR you run and what the study or registry needs. We reply within one business day to set a meeting time.

Request a demo