For research informatics

Clinical research informatics without the warehouse queue.

Clinical Extract gives investigators study-scoped extracts from your EHR's certified bulk FHIR API. Their requests no longer wait behind warehouse builds.

Each extract holds only the study cohort and the variables the protocol approved, in one CSV with a data dictionary.

1 From request to file

Four steps from a research data request to a study export.

The investigator's approved variable list; the cohort, built from coded criteria or matched from a list; the honest broker's review of the data dictionary; the export. The Group export narrows the export to the study cohort and its resource types, and the minimum-necessary cut keeps only the approved variables, so identifiers the protocol did not approve never reach study.csv.

EHR data for research arrives coded (ICD-10-CM, LOINC, RxNorm, SNOMED CT) in study.csv, and data-dictionary.json defines every column.

From request to fileFour numbered steps. Step 1, request: the study request for study-0142 names the cohort (ICD-10-CM C50.x, diagnosed 2024, stage II or III) and 24 approved data elements, among them birth year, gender, diagnosis code and date, AJCC stage group, ER, PR and HER2 status, and first chemotherapy. Step 2, minimum necessary, highlighted: Clinical Extract kicks off GET /Group/{cohort-id}/$export with _type=Patient,Condition,Observation,MedicationRequest and Prefer: respond-async, so the NDJSON export holds 212 patients and 4 resource types. A cut line then separates the exported NDJSON elements from the study.csv columns. Patient.name, Patient.address, Patient.telecom and Patient.identifier (MRN) are not approved and stop at the line. Patient.id becomes patient_ref, Patient.birthDate birth_year, Patient.gender gender, Condition.code dx_code, Condition.onsetDateTime dx_date, Observation LOINC 21908-9 stage_group, 16112-5 er_status, 16113-3 pr_status, 48676-1 her2_status, and MedicationRequest.medicationCodeableConcept (RxNorm) first_chemo. Step 3, honest broker: the broker reads data-dictionary.json, for example the er_status entry with its description Estrogen receptor status (LOINC 16112-5), fhirSource Observation.valueCodeableConcept and example Positive, and checks that all 24 columns are on the approval, that every column has a fhirSource, and that the 212 rows are the study cohort. Step 4, analyst: study.csv, 212 rows and 24 columns, with data-dictionary.json, 24 entries.1REQUEST2COHORT AND VARIABLES3HONEST BROKER4ANALYSTPROTOCOLstudy-0142IRB approvedCOHORTICD-10-CM C50.xdiagnosed 2024stage II or IIIDATA ELEMENTS24 approved, among thembirth year, genderdiagnosis code and dateAJCC stage groupER, PR, HER2 statusfirst chemotherapyGET /Group/{cohort-id}/$export  ?_type=Patient,Condition,    Observation,MedicationRequestPrefer: respond-asyncNDJSON for 212 patients, 4 resource typesMINIMUM NECESSARYEXPORTED NDJSONSTUDY.CSVPatient.name.address.telecom.identifier (MRN).idpatient_ref.birthDatebirth_year.gendergenderCondition  ICD-10-CM.codedx_code.onsetDateTimedx_dateObservation  LOINC21908-9 stage groupstage_group16112-5 ERer_status16113-3 PRpr_status48676-1 HER2her2_statusMedicationRequest  RxNorm.medicationCodeableConceptfirst_chemonot approved,not in study.csv10 of 24 columns drawndata-dictionary.jsoner_statusDESCRIPTIONEstrogen receptor status (LOINC16112-5)FHIRSOURCEObservation.valueCodeableConceptEXAMPLEPositiveall 24 columns are on theapprovalevery column has a fhirSource212 rows, the study cohortstudy.csv212 rows, 24 columnsdata-dictionary.json24 entries
From request to fileFour numbered steps. Step 1, request: the study request for study-0142 names the cohort (ICD-10-CM C50.x, diagnosed 2024, stage II or III) and 24 approved data elements, among them birth year, gender, diagnosis code and date, AJCC stage group, ER, PR and HER2 status, and first chemotherapy. Step 2, minimum necessary, highlighted: Clinical Extract kicks off GET /Group/{cohort-id}/$export with _type=Patient,Condition,Observation,MedicationRequest and Prefer: respond-async, so the NDJSON export holds 212 patients and 4 resource types. A cut line then separates the exported NDJSON elements from the study.csv columns. Patient.name, Patient.address, Patient.telecom and Patient.identifier (MRN) are not approved and stop at the line. Patient.id becomes patient_ref, Patient.birthDate birth_year, Patient.gender gender, Condition.code dx_code, Condition.onsetDateTime dx_date, Observation LOINC 21908-9 stage_group, 16112-5 er_status, 16113-3 pr_status, 48676-1 her2_status, and MedicationRequest.medicationCodeableConcept (RxNorm) first_chemo. Step 3, honest broker: the broker reads data-dictionary.json, for example the er_status entry with its description Estrogen receptor status (LOINC 16112-5), fhirSource Observation.valueCodeableConcept and example Positive, and checks that all 24 columns are on the approval, that every column has a fhirSource, and that the 212 rows are the study cohort. Step 4, analyst: study.csv, 212 rows and 24 columns, with data-dictionary.json, 24 entries.1REQUESTPROTOCOLstudy-0142IRB approvedCOHORTICD-10-CM C50.xdiagnosed 2024stage II or IIIDATA ELEMENTS24 approved, among thembirth year, genderdiagnosis code and dateAJCC stage groupER, PR, HER2 statusfirst chemotherapy2COHORT AND VARIABLESGET /Group/{cohort-id}/$export  ?_type=Patient,Condition,    Observation,MedicationRequestPrefer: respond-asyncNDJSON for 212 patients, 4 resource typesMINIMUM NECESSARYEXPORTED NDJSONSTUDY.CSVPatient.name.address.telecom.identifier (MRN).idpatient_ref.birthDatebirth_year.gendergenderCondition  ICD-10-CM.codedx_code.onsetDateTimedx_dateObservation  LOINC21908-9 stage groupstage_group16112-5 ERer_status16113-3 PRpr_status48676-1 HER2her2_statusMedicationRequest  RxNorm.medicationCodeableConceptfirst_chemonot approved,not in the CSV10 of 24 columns drawn3HONEST BROKERdata-dictionary.jsoner_statusDESCRIPTIONEstrogen receptor status (LOINC 16112-5)FHIRSOURCEObservation.valueCodeableConceptEXAMPLEPositiveall 24 columns are on the approvalevery column has a fhirSource212 rows, the study cohort4ANALYSTstudy.csv212 rows, 24 columnsdata-dictionary.json24 entries
Figure 1. From request to file. The Group export narrows the export to the study cohort and its resource types, and the minimum-necessary cut keeps only the approved variables, so identifiers the protocol did not approve never reach study.csv. The honest broker reviews the same data dictionary the analyst receives. Study-0142 and its counts are illustrative; 10 of its 24 columns are drawn.

2 The queue

Study pulls and the data warehouse.

A warehouse request waits in a shared queue for an analyst, a query, a chart check and a de-identification pass. A study pull goes from the approved variable list to study.csv in days, most of them the honest broker's review, and the export itself runs in hours. The warehouse still handles enterprise reporting.

Warehouse queue vs study pullTwo panels with illustrative durations. Panel a, two timelines on one axis of days from request, 0 to 8 weeks, each stage drawn as work or as waiting. The warehouse queue: the request waits in the intake queue for 21 days, a scoping meeting with a warehouse analyst takes 3 days, SQL is written and run against the warehouse in 12 days, results are checked against charts in 5 days, the request waits 4 days for the honest broker, and de-identification and delivery take 4 days; request to file, 49 days. The study pull with Clinical Extract: request to file, 3 days, highlighted. Panel b, the study pull on an axis of hours from request, 0 to 72, with the first 3 hours drawn enlarged: the cohort is built and the approved variables picked in 1 hour; the Group $export runs from kickoff to the last NDJSON file in a couple of hours for the study cohort; study.csv and data-dictionary.json are written in minutes, 3 hours from request; the files wait 61 hours for the honest broker, who reviews data-dictionary.json in 8 hours.aRequest to file, two timelinesdays from requestworkwaitingWarehouse queueanalyst-built query21 days  waits in the intake queue3 days  scoping with an analyst12 days  SQL written and run5 days  results checked against charts4 days  waits for the honest broker4 days  de-identified and delivered49 daysStudy pullwith Clinical Extract3 days  hour by hour in b01 wk2 wk3 wk4 wk5 wk6 wk7 wk8 wkbThe study pull, hour by hourhours from request, first 3 h enlargedStudy pullwith Clinical Extract1 h  cohort built, approved variables pickedabout 2 h  Group $export, kickoff to last NDJSON fileminutes  study.csv and data-dictionary.json written61 h  waits for the honest broker8 h  honest broker reviews data-dictionary.json72 h01 h2 h3 h12 h24 h36 h48 h60 h72 h
Warehouse queue vs study pullTwo panels with illustrative durations. Panel a, two timelines on one axis of days from request, 0 to 8 weeks, each stage drawn as work or as waiting. The warehouse queue: the request waits in the intake queue for 21 days, a scoping meeting with a warehouse analyst takes 3 days, SQL is written and run against the warehouse in 12 days, results are checked against charts in 5 days, the request waits 4 days for the honest broker, and de-identification and delivery take 4 days; request to file, 49 days. The study pull with Clinical Extract: request to file, 3 days, highlighted. Panel b, the study pull on an axis of hours from request, 0 to 72, with the first 3 hours drawn enlarged: the cohort is built and the approved variables picked in 1 hour; the Group $export runs from kickoff to the last NDJSON file in a couple of hours for the study cohort; study.csv and data-dictionary.json are written in minutes, 3 hours from request; the files wait 61 hours for the honest broker, who reviews data-dictionary.json in 8 hours.aRequest to filedays from requestworkwaitingWarehouse queue  analyst-built49 days11  waits in the intake queue21 days22  scoping meeting3 days33  SQL written and run12 days44  checked against charts5 days55  waits for the honest broker4 days66  de-identified, delivered4 daysStudy pull  with Clinical Extract3 dayshour by hour in b02 wk4 wk6 wk8 wkbThe study pullhours from request, first 3 h enlargedStudy pull  with Clinical Extract72 h11  cohort and variables picked1 h22  Group $export to last fileabout 2 h33  CSV and dictionary writtenminutes44  waits for the honest broker61 h55  broker reviews dictionary8 h03 h24 h48 h72 h
Figure 2 as a table
StageFrom requestDuration
Warehouse: waits in the intake queueday 0 to 2121 days, waiting
Warehouse: scoping meeting with a warehouse analystday 21 to 243 days
Warehouse: SQL written and run against the warehouseday 24 to 3612 days
Warehouse: results checked against chartsday 36 to 415 days
Warehouse: waits for the honest brokerday 41 to 454 days, waiting
Warehouse: de-identified and deliveredday 45 to 494 days
Warehouse: request to fileday 0 to 4949 days
Study pull: cohort built, approved variables pickedhour 0 to 11 hour
Study pull: Group $export, kickoff to last NDJSON filehour 1 to 2.8about 2 hours, illustrative for a study cohort
Study pull: study.csv and data-dictionary.json writtenhour 2.8 to 3minutes
Study pull: waits for the honest brokerhour 3 to 6461 hours, waiting
Study pull: honest broker reviews data-dictionary.jsonhour 64 to 728 hours
Study pull: request to filehour 0 to 723 days
Figure 2. Warehouse queue versus study pull. (a) A warehouse request waits for an analyst, a query, a chart check and a de-identification pass. A study pull goes from the approved variable list to study.csv in days. (b) Most of those days are the honest broker’s review; the export itself runs in hours. The first 3 hours of (b) are drawn enlarged. Every duration in (b) is illustrative for a study cohort.

3 The honest broker

How the honest broker reviews and releases an extract.

Honest broker research centers on one document: the list of variables to be released. Clinical Extract writes that list before the export runs. data-dictionary.json names every column, its description, the FHIR element it comes from and an example value. The broker approves those variables, keeps the re-identification key, and releases the coded extract, and the analyst receives the same dictionary. One dictionary entry, explained shows what the broker reads.

Research informatics teams keep their governance as written: the export runs inside your environment, the token carries read scopes only, and no patient data reaches us. Minimum necessary in the export covers the scopes and the data path.

4 The IRB

What the IRB application quotes.

Which patients: the cohort criteria and the count the cohort builder returned. Which variables: the dictionary. Where each one comes from: the fhirSource field. Preparatory-to-research counts come from building the cohort from coded criteria before any record moves, so the application carries a number and its definition instead of an estimate.

5 Questions

Questions from research informatics

What is an honest broker in research?

The person or office that stands between the investigator and the identified record: it holds the re-identification key, approves what leaves, and releases only the coded extract. Clinical Extract gives the broker a document to approve before anything moves: data-dictionary.json, one entry per column with its FHIR source.

What is the honest broker protocol?

The written procedure for that role: who requests, who approves the variable list, how identifiers are handled, where the key is kept, how the release is logged. Clinical Extract fits it as written: the approved list is the export’s variable list, the extract is study-scoped, and the export itself runs inside your environment.

What is clinical research informatics?

The discipline that turns clinical data into research data: the systems, the governance and the people between the EHR and the investigator. Research informatics teams run the data warehouse, the honest broker service and the study data requests, and Clinical Extract adds study pulls that answer a request in days.

How does a study pull differ from a warehouse request?

A warehouse request waits for an analyst, a query, a chart check and a de-identification pass. A study pull starts from one protocol: the approved variable list becomes the export, the cohort’s Group scopes it, and the CSV lands in days, most of them the broker’s review.

Where do the files land?

In a study folder inside your environment, on the server or in the cloud tenant you run Clinical Extract on: the NDJSON, study.csv and data-dictionary.json. Patient data moves only between your EHR and that folder.

Next

See Clinical Extract run on one of your studies.

Tell us which EHR you run and what the study or registry needs. We reply within one business day to set a meeting time.

Request a demo