How it works

How Clinical Extract turns your EHR's export into a CSV.

Clinical Extract connects to your EHR once as a read-only backend app, exports only the study cohort, keeps only the approved variables, and returns study.csv and data-dictionary.json. The sections below follow each request in order.

1 The sequence

Authorize, kick off, poll, download, flatten.

Authorize. Clinical Extract signs a JWT with a private key that stays in your key store and posts it to your EHR's token endpoint with grant_type=client_credentials; the token that comes back carries system/Group.read for the cohort and one read scope per exported resource type. The protocol is SMART Backend Services; our guide to SMART on FHIR covers it in full.

Kick off. One GET Group/{cohort-id}/$export with _type set to the resource types the protocol names and Prefer: respond-async. The EHR answers 202 Accepted with the status URL.

Poll. Clinical Extract polls the status URL, waiting the interval each Retry-After header gives, until it answers 200 with the manifest.

Download. Each NDJSON file the manifest lists is fetched with the bearer token into your study folder, one resource per line.

Flatten. The approved variables are picked out of the resources into study.csv, one row per patient, and each column gets an entry in data-dictionary.json. The bulk export guide lists every header and manifest field.

The bulk export sequenceA sequence diagram in three lanes: your environment, Clinical Extract and your EHR, with time running down, in five phases. Authorize: 1, Clinical Extract reads its RS384 private signing key from your key store. 2, it signs a JWT with iss and sub set to its client_id, aud set to the token URL, exp five minutes out and a unique jti. 3, it posts to the token URL with grant_type=client_credentials, client_assertion_type=urn:ietf:params:oauth:client-assertion-type:jwt-bearer, the signed JWT as client_assertion, and scope system/Patient.read, system/Group.read, system/Condition.read, system/Observation.read and system/MedicationRequest.read. 4, the EHR returns 200 OK with an access_token, token_type bearer, expires_in 300. Kick off, highlighted: 5, GET /Group/{cohort-id}/$export with _type=Patient,Condition,Observation,MedicationRequest; the EHR replies 202 Accepted with a Content-Location status URL. Poll: 6, GET {status-url}, repeated at the pace of Retry-After; the EHR replies 202 Accepted until the export is done, then 200 OK with the JSON manifest. Download: 7, GET {output-url} once per file in the manifest; the EHR replies 200 OK with NDJSON. Deliver: 8, Clinical Extract flattens the NDJSON and keeps the approved variables, for example Condition.code in ICD-10-CM to dx_code, Observation LOINC 16112-5 to er_status and MedicationRequest RxNorm to first_chemo. 9, it writes study.csv, 212 rows, and data-dictionary.json, one entry per column, into your environment.Your environmentkey store,study folderClinical ExtractYour EHRcertifiedbulk FHIR APIAUTHORIZE1signing keyRS384 private key,from your key store2sign the JWT, RS384iss = sub = {client_id}aud = {token-url}exp: 5 minutes out, jti: unique3POST {token-url}grant_type=client_credentialsclient_assertion_type=urn:ietf:params:oauth:  client-assertion-type:jwt-bearerclient_assertion={signed JWT}scope=system/Patient.read      system/Group.read      system/Condition.read      system/Observation.read      system/MedicationRequest.read4200 OKaccess_token, token_type: bearer,expires_in: 300KICK OFF5GET /Group/{cohort-id}/$export?_type=Patient,Condition,Observation,       MedicationRequestreply: 202 Accepted, Content-LocationPOLL6GET {status-url}repeats at the pace of Retry-Afterreply: 202 Accepted, then 200 OK       with the JSON manifestDOWNLOAD7GET {output-url}once per file in the manifestreply: 200 OK, NDJSONDELIVER8flatten, keep the approved variablesdx_code      Condition, ICD-10-CMer_status    Observation, LOINC 16112-5first_chemo  MedicationRequest, RxNorm9study.csv, 212 rowsdata-dictionary.json,one entry per column
The bulk export sequenceA sequence diagram in three lanes: your environment, Clinical Extract and your EHR, with time running down, in five phases. Authorize: 1, Clinical Extract reads its RS384 private signing key from your key store. 2, it signs a JWT with iss and sub set to its client_id, aud set to the token URL, exp five minutes out and a unique jti. 3, it posts to the token URL with grant_type=client_credentials, client_assertion_type=urn:ietf:params:oauth:client-assertion-type:jwt-bearer, the signed JWT as client_assertion, and scope system/Patient.read, system/Group.read, system/Condition.read, system/Observation.read and system/MedicationRequest.read. 4, the EHR returns 200 OK with an access_token, token_type bearer, expires_in 300. Kick off, highlighted: 5, GET /Group/{cohort-id}/$export with _type=Patient,Condition,Observation,MedicationRequest; the EHR replies 202 Accepted with a Content-Location status URL. Poll: 6, GET {status-url}, repeated at the pace of Retry-After; the EHR replies 202 Accepted until the export is done, then 200 OK with the JSON manifest. Download: 7, GET {output-url} once per file in the manifest; the EHR replies 200 OK with NDJSON. Deliver: 8, Clinical Extract flattens the NDJSON and keeps the approved variables, for example Condition.code in ICD-10-CM to dx_code, Observation LOINC 16112-5 to er_status and MedicationRequest RxNorm to first_chemo. 9, it writes study.csv, 212 rows, and data-dictionary.json, one entry per column, into your environment.Your environment  key store, study folderClinical ExtractYour EHR  certified bulk FHIR APIAUTHORIZE1signing keyRS384 private key,from your key store2sign the JWT, RS384iss = sub = {client_id}aud = {token-url}exp: 5 minutes out, jti: unique3POST {token-url}grant_type=client_credentialsclient_assertion_type=  urn:ietf:params:oauth:  client-assertion-type:jwt-bearerclient_assertion={signed JWT}scope=system/Patient.read  system/Group.read  system/Condition.read  system/Observation.read  system/MedicationRequest.read4200 OKaccess_token, expires_in: 300token_type: bearerKICK OFF5GET /Group/{cohort-id}/$export?_type=Patient,Condition,  Observation,MedicationRequestreply: 202 Accepted,  Content-LocationPOLL6GET {status-url}repeats at the pace of Retry-Afterreply: 202 Accepted, then 200 OK  with the JSON manifestDOWNLOAD7GET {output-url}once per file in the manifestreply: 200 OK, NDJSONDELIVER8flatten, keep approved variablesdx_code: Condition, ICD-10-CMer_status: Observation,  LOINC 16112-5first_chemo: MedicationRequest,  RxNorm9study.csv, 212 rowsdata-dictionary.json,one entry per column
Figure 1. The bulk export sequence, as Clinical Extract runs it. It authorizes with SMART Backend Services, kicks off one Group/{cohort-id}/$export scoped to the study cohort and the resource types the protocol names, polls the status URL, downloads each NDJSON file listed in the manifest, and writes the CSV and its data dictionary. The token carries system/Group.read for the cohort and one read scope per exported _type. Every kickoff, status and manifest header is set out in full on the bulk FHIR export page. The 212-row count is illustrative.

2 What stays where

Every record, file and credential stays inside your environment.

The EHR sits on the boundary of your environment, on-premises or vendor-hosted. Clinical Extract runs inside it, on your server or in your cloud tenant. The private signing key never leaves your key store. The NDJSON, the CSV and the dictionary land in your study folder, and the token carries read scopes only. Only Clinical Extract software comes into your environment. We never receive patient data. The security page covers the scopes, the data path and your controls.

What stays whereA matrix of what a study run touches against where it stays. Columns: your EHR, on-premises or vendor-hosted, sits on the edge of a dashed boundary labelled your environment; Clinical Extract, on your server or cloud tenant, and the study folder, in your storage, sit inside it. A fourth column, Releases, sits outside it. Patient data: patient records for every patient stay in the EHR; the study cohort, Group/{cohort-id}, stays in the EHR; NDJSON export files come from the EHR and stay with Clinical Extract; study.csv, 212 rows, and data-dictionary.json, one entry per column, come from Clinical Extract and stay in the study folder. Credentials: the RS384 private signing key stays with Clinical Extract in your key store; the client_id and public key stay registered with the EHR; the short-lived access token comes from the EHR and stays with Clinical Extract. Software: Clinical Extract releases arrive as source with a Dockerfile, are built into a container in your environment and run there. Every other row is marked none held in the Releases column: releases come in, and we never receive data. The boundary edge between the study folder and Releases is highlighted: patient data stays inside it.YOUR ENVIRONMENTYour EHRon-premises orvendor-hostedClinical Extractyour server orcloud tenantStudy folderyour storageReleasessoftware releasesin; no data backIN A STUDY RUNwhere it stayswhere it comes fromnone heldPATIENT DATAPatient recordsevery patientStudy cohortGroup/{cohort-id}NDJSON export filesone resource per linestudy.csv212 rowsdata-dictionary.jsonone entry per columnCREDENTIALSPrivate signing keyRS384, in your key storeclient_id and public keybackend app registrationAccess tokenshort-lived, read scopesSOFTWAREClinical Extract releasessource and Dockerfilepatient data stays inside
What stays whereA matrix of what a study run touches against where it stays. Columns: your EHR, on-premises or vendor-hosted, sits on the edge of a dashed boundary labelled your environment; Clinical Extract, on your server or cloud tenant, and the study folder, in your storage, sit inside it. A fourth column, Releases, sits outside it. Patient data: patient records for every patient stay in the EHR; the study cohort, Group/{cohort-id}, stays in the EHR; NDJSON export files come from the EHR and stay with Clinical Extract; study.csv, 212 rows, and data-dictionary.json, one entry per column, come from Clinical Extract and stay in the study folder. Credentials: the RS384 private signing key stays with Clinical Extract in your key store; the client_id and public key stay registered with the EHR; the short-lived access token comes from the EHR and stays with Clinical Extract. Software: Clinical Extract releases arrive as source with a Dockerfile, are built into a container in your environment and run there. Every other row is marked none held in the Releases column: releases come in, and we never receive data. The boundary edge between the study folder and Releases is highlighted: patient data stays inside it.YOUR ENVIRONMENTYour EHRon-premises orvendor-hostedSTAYS HEREPatient recordsevery patientStudy cohortGroup/{cohort-id}client_id and public keybackend app registrationClinical Extractyour server or cloud tenantSTAYS HERENDJSON export filesone resource per line, from your EHRPrivate signing keyRS384, in your key storeAccess tokenshort-lived, read scopes, from yourEHRClinical Extract releasessource and Dockerfile, from ReleasesStudy folderyour storageSTAYS HEREstudy.csv212 rows, from Clinical Extractdata-dictionary.jsonone entry per column, from ClinicalExtractReleasessoftware releases in; no data backpatient data stays insideHOLDSno patient data, no credentials
Figure 2. What stays where. Every record, file and credential in a study run stays in your EHR or inside your environment, and the private signing key stays in your key store. The 212-row study is illustrative.

3 The data dictionary

One entry per column, traced to its FHIR element.

Every column in study.csv has one entry in data-dictionary.json, keyed by the column name, with a description (the meaning and its code system), the FHIR element the values come from (fhirSource) and an example value. A reviewer can trace any cell back to the resource it came from. Every column described shows the file in full.

Anatomy of a data-dictionary entryThree listings side by side. Left, one Observation from the NDJSON export, abridged: resourceType Observation, status final, code LOINC 16112-5, valueCodeableConcept coded SNOMED CT 10828004 with display Positive, subject Patient/eKx3p9Qa. Center, data-dictionary.json with the er_status entry expanded between collapsed neighbors: description "Estrogen receptor status (LOINC 16112-5)", fhirSource "Observation.valueCodeableConcept", example "Positive". Right, column 7 of 24 of study.csv, headed er_status, its first four values: Positive, Negative, Positive, Positive, of 212 rows. Numbered links: 1, the entry key er_status is the column header. 2, the description names LOINC 16112-5, the code on the Observation. 3, highlighted, fhirSource names the element the value is read from, the Observation’s valueCodeableConcept. 4, the example Positive is a value as it appears in the column. Four notes below explain each field.OBSERVATION, NDJSONone line, abridged and indentedDATA-DICTIONARY.JSON24 entries, one per columnSTUDY.CSVcolumn 7 of 24{"resourceType": "Observation", "status": "final", "code": {"coding": [{   "system": "http://loinc.org",   "code": "16112-5"}]}, "valueCodeableConcept": {   "coding": [{   "system": "http://snomed.info/sct",   "code": "10828004",   "display": "Positive"}]}, "subject": {"reference":   "Patient/eKx3p9Qa"}}{  "patient_ref": { ... },  ...  "er_status": {1    "description": "Estrogen receptor status (LOINC 16112-5)",2    "fhirSource": "Observation.valueCodeableConcept",3    "example": "Positive"4  },  "pr_status": { ... },  ...}er_statusPositiveNegativePositivePositive... 212 rows1keyThe column name. It is theheader of the column instudy.csv, character forcharacter.2descriptionWhat the column holds, inwords, with the code thatselects it: LOINC 16112-5,Estrogen receptor[Interpretation] in Tissue.3fhirSourceThe FHIR element every value isread from: valueCodeableConcepton the Observation that carriesthat LOINC code.4exampleA value as it appears in thecolumn. Positive is the displayof SNOMED CT 10828004.
Anatomy of a data-dictionary entryThree listings side by side. Left, one Observation from the NDJSON export, abridged: resourceType Observation, status final, code LOINC 16112-5, valueCodeableConcept coded SNOMED CT 10828004 with display Positive, subject Patient/eKx3p9Qa. Center, data-dictionary.json with the er_status entry expanded between collapsed neighbors: description "Estrogen receptor status (LOINC 16112-5)", fhirSource "Observation.valueCodeableConcept", example "Positive". Right, column 7 of 24 of study.csv, headed er_status, its first four values: Positive, Negative, Positive, Positive, of 212 rows. Numbered links: 1, the entry key er_status is the column header. 2, the description names LOINC 16112-5, the code on the Observation. 3, highlighted, fhirSource names the element the value is read from, the Observation’s valueCodeableConcept. 4, the example Positive is a value as it appears in the column. Four notes below explain each field.OBSERVATION, NDJSONone line, abridged{"resourceType": "Observation", "status": "final", "code": {"coding": [{   "system": "http://loinc.org",   "code": "16112-5"}]},2 "valueCodeableConcept": {3   "coding": [{   "system": "http://snomed.info/sct",   "code": "10828004",   "display": "Positive"}]}, "subject": {"reference":   "Patient/eKx3p9Qa"}}DATA-DICTIONARY.JSON24 entries"er_status": {1  "description":2   "Estrogen receptor status    (LOINC 16112-5)",  "fhirSource":3   "Observation.valueCodeableConcept",  "example": "Positive"4}STUDY.CSVcolumn 7 of 24er_status1PositiveNegativePositive4Positive... 212 rows1keyThe column name. It is the header of thecolumn in study.csv, character forcharacter.2descriptionWhat the column holds, in words, with thecode that selects it: LOINC 16112-5,Estrogen receptor [Interpretation] inTissue.3fhirSourceThe FHIR element every value is read from:valueCodeableConcept on the Observation thatcarries that LOINC code.4exampleA value as it appears in the column.Positive is the display of SNOMED CT10828004.
Figure 3. Anatomy of a data-dictionary entry. Every column in study.csv has one entry in data-dictionary.json, keyed by the column name, with a description, the FHIR element the values come from and an example value. The fhirSource field traces each value back to its FHIR element. The row values and resource IDs are synthetic.

4 Next

Related pages.

Next

See Clinical Extract run on one of your studies.

Tell us which EHR you run and what the study or registry needs. We reply within one business day to set a meeting time.

Request a demo