Bulk FHIR guide
Bulk FHIR export, from kickoff to CSV.
Clinical Extract runs a study-scoped bulk FHIR export from your EHR. This guide covers each request, from the kickoff to a CSV your analysts can open.
This guide shows every kickoff, status and manifest header in full.
1 Definition
What is bulk FHIR?
Bulk FHIR is the HL7 FHIR Bulk Data Access specification, once nicknamed Flat FHIR: an asynchronous $export operation that returns a population's records as NDJSON files, one FHIR resource per line, instead of one REST call per resource. The FHIR bulk data API has three parts: a kickoff request, a status URL to poll, and the files listed in a manifest when the export is done.
Every certified EHR has it. The certification criterion ONC §170.315(g)(10) requires certified EHR technology to export data for a group of patients through HL7 FHIR Bulk Data Access, authorized with SMART Backend Services, and under 45 CFR 170.404(b)(3) developers had to make it available to their customers by December 31, 2022. So a FHIR bulk data export is the standard way to get a cohort's coded records out of any certified EHR, and it is the way Clinical Extract works: it runs the FHIR bulk data access API for the study cohort's Group, then turns the NDJSON into study.csv and a data dictionary.
A FHIR export from a cloud FHIR store (a Google Cloud or Azure bulk export) uses the same operation; this guide is about the EHR's own endpoint, which is where the chart lives.
2 Three kinds of export
System, patient and Group exports, and the one a study needs.
A system-level export ($export on the server) takes everything. A patient-level export (Patient/$export) takes every patient. A FHIR group export (Group/{id}/$export) takes the members of one Group, and that is the export the certification criterion requires and the one a study needs. Clinical Extract runs the Group export and adds _type to request only the resource types the study reads, so the request covers only the cohort and its variables.
_type to request only the resource types the study needs. The grid is schematic; the patient counts are illustrative. 3 Kickoff to files
The FHIR bulk export request pattern and its headers.
The kickoff is one GET with Prefer: respond-async and Accept: application/fhir+json. The server answers 202 Accepted at once, with a Content-Location header naming the status URL. Clinical Extract polls that URL; while the export runs the server answers 202 with Retry-After (and, on Epic, X-Progress), and Clinical Extract waits the interval Retry-After specifies. When the export is done the status URL answers 200 with the manifest: transactionTime (the time the data is as of), request (the kickoff URL), requiresAccessToken, output (one entry per NDJSON file: its type, its URL, its line count) and error (OperationOutcome files). Each file is fetched with Accept: application/fhir+ndjson and the bearer token, and the status URL is deleted once the files are safe in your study folder.
output entry names one NDJSON file: its url is the file, its type is the resourceType of every line, and its count is the number of lines. Headers and manifest fields follow the Bulk Data Access specification; the URLs, IDs, header values, timestamp and counts are illustrative. GET {fhir-base}/Group/{cohort-id}/$export
?_type=Patient,Condition,Observation,MedicationRequest
Accept: application/fhir+json
Prefer: respond-async
Authorization: Bearer {access token}
202 Accepted
Content-Location: {status-url}
GET {status-url}
202 Accepted
Retry-After: 120
X-Progress: {what the server reports} { "transactionTime": "2026-09-28T14:02:11Z", "request": "{fhir-base}/Group/{cohort-id}/$export?_type=Patient,Condition,Observation,MedicationRequest", "requiresAccessToken": true, "output": [ { "type": "Patient", "url": "{file-url-1}", "count": 212 }, { "type": "Condition", "url": "{file-url-2}", "count": 1684 }, { "type": "Observation", "url": "{file-url-3}", "count": 48310 }, { "type": "MedicationRequest", "url": "{file-url-4}", "count": 3905 } ], "error": [] }
4 NDJSON to CSV
Turning NDJSON lines into CSV rows.
Each line of an NDJSON file is one complete FHIR resource. Clinical Extract joins it to its patient's row by subject.reference, uses the LOINC code (or the RxNorm, ICD-10-CM or SNOMED CT code) to pick the column, writes the coded value into the cell, and records where the column came from in data-dictionary.json: the description with its code system, the FHIR element, an example. The result is study.csv, with a row for each patient in the cohort and a column for each approved variable, and every column described. The steps between the token and the file are on the how-it-works page.
subject.reference, uses the LOINC code to pick the column, writes the coded value into the cell, and records where every column came from in data-dictionary.json. Three of the study’s 24 columns shown; the resource IDs and values are synthetic. 5 Real-world throughput
Export time for a study and for a whole population.
Jones et al. (JAMIA, 2024) measured (g)(10) bulk export at five sites. Epic sites ran at 502 to 2,827 resources per minute, and Cerner above 8,000 per minute; it took the sites 2 to 119 days (mean 65) from submitting cohort criteria to the first successful bulk request. At those rates, export time grows with the number of resources requested: a 212-patient study cohort exports in under three hours at Epic's slowest reported rate and in minutes at Cerner's, while a whole population takes weeks. Clinical Extract therefore exports only the study Group, and the cohort builder fixes the cohort and its count before the first request. Source: Jones et al. J Am Med Inform Assoc. 2024, doi:10.1093/jamia/ocae040.
Figure 4 as a table
| Series | Value | Basis |
|---|---|---|
| Epic sites, throughput | 502 to 2,827 resources per minute | Reported, Jones et al. |
| Cerner (now Oracle Health), throughput | above 8,000 resources per minute | Reported, Jones et al. |
| 212 patients, the study cohort, 54,111 resources | 19 min to 1.8 h | Arithmetic on Epic’s range, the study’s manifest count |
| 500 patients, 200,000 resources | 71 min to 6.6 h | Arithmetic on Epic’s range, 400 resources per patient |
| 5,000 patients, 2 million resources | 11.8 h to 2.8 days | Arithmetic on Epic’s range, 400 resources per patient |
| 25,000 patients, 10 million resources | 2.5 to 13.8 days | Arithmetic on Epic’s range, 400 resources per patient |
| 125,000 patients, the whole population, 50 million resources | 12 to 69 days | Arithmetic on Epic’s range, 400 resources per patient |
6 By EHR
The same API on every certified EHR, in each vendor's terms.
Each EHR page shows the registration, the activation and the export in that vendor's own vocabulary, with a dated fact table and the specimens: Epic, Oracle Health, eClinicalWorks, athenahealth, MEDITECH, NextGen and Veradigm EHR. How each one exposes the API is the hub. For eClinicalWorks, the $export exchange request by request is in our developer guide for eClinicalWorks.
7 Questions
Questions about bulk FHIR
What does bulk data mean here?
Many patients’ resources at once, as files. Instead of one request per patient and resource, the bulk data access operation returns NDJSON files, one FHIR resource per line, for a whole population, a Group of patients, or the system.
What is the FHIR bulk API?
The FHIR bulk API is the $export operation from HL7 Bulk Data Access: a kickoff, a status URL to poll, and NDJSON files listed in a manifest. Clinical Extract runs it for the study cohort’s Group.
What is Flat FHIR?
The nickname the Bulk Data Access specification carried in its early drafts: "flat" because NDJSON files hold one resource per line with no bundle around them. The specification is HL7 FHIR Bulk Data Access; the operation is $export.
What is a FHIR Group export?
GET Group/{id}/$export: the bulk export scoped to the members of one Group resource. It is the export every certified EHR must offer, and the one Clinical Extract runs, with the study cohort as the Group.
How long does a bulk FHIR export take?
It scales with the number of resources requested. Jones et al. measured 502 to 2,827 resources per minute at Epic sites and above 8,000 at Cerner; a 212-patient study cohort exports in under three hours at Epic’s reported rates (Figure 4), and a whole population takes weeks.
Do we need a data warehouse first?
No. The bulk FHIR API reads the EHR’s own store, and a study-scoped export lands in your study folder as a CSV and its data dictionary. A study export does not need a warehouse, which is how informatics teams run study pulls without the warehouse queue.
Next
See Clinical Extract run on one of your studies.
Tell us which EHR you run and what the study or registry needs. We reply within one business day to set a meeting time.