Skip to main content

Obtaining U of A approval of EHR data requests

Regulatory requirements for accessing EHR data

In order to guide researchers, the Center for Biomedical Informatics and Biostatistics (CB2) has reviewed regulatory requirements by U of A and other regulators, and describes them in the sections below. However, please see more information about this page below, including a disclaimer, as CB2 nor U of A may not be aware of all the details.

Institutional Review Board (IRB)

Regulations require an IRB to determine if your study protocol needs their approval to ensure the safety and privacy of any possible human subjects. Below are instructions for how to get this determination and possible approval from the U of A IRB, including links to IRB forms and instructions, and information on additional approvals which may be needed before submitting to the IRB.

Efficiently and effectively using EHR data

Efficient approval of your study, and the effectiveness of your study, is partly dependent on the clarity and completeness of your phenotype definition(s) of the patient cohort(s) and electronic health record (EHR) data needed for your study. However, EHR data can be inaccurate and/or incomplete. Thus: 

Review EHR FAQs

Read answers to EHR frequently asked questions (FAQs) below about how to determine and specify patients and EHR details needed, including best practices, when drafting and completing your study protocol and other forms.

Initiate U of A approval to access EHR data

Initiate U of A IRB approval
Complete required steps:
  1. Draft a protocol using the appropriate IRB form, specifying:
    1. criteria for patient cohort(s):
      1. inclusions
      2. and/or exclusions
    2. "minimum necessary"* EHR data elements needed
    3. any planned AI: see next (AI) card section for more info 
    4. all other required and relevant info
  2. Create study in U of A eIRB:
    1. complete all required and relevant fields
    2. upload your protocol and any other necessary docs
    3. use the Printer Version button to get PDF of study for next step; but, do NOT submit, yet

*Read answers to EHR FAQs to determine criteria and data needed. Contact CB2 for more help.

Image
blurred person using finger to check box in row of checkboxes

Get IRB forms and instructions

IRBs and healthcare organizations (HCOs) need to know, and possibly approve, use of AI for your study.

Specify in your IRB protocol
  • What: 
    • is the intended use of the AI
    • data will be used for the AI
  • Where AI and data will be: 
    • handled and used
    • stored and shared
  • How the AI will be used

HCOs might need to approve if using AI
  • impacting an HCO’s clinical care 
  • using an HCO's data

Contact CB2 to help you specify necessary AI details and determine if you might need approval.

If an Investigator-Initiated Study or Trial (IIS/IIT)
If conducting a hypothesis-driven IIS/IIT at COM-T

(U of A College of Medicine in Tucson)

Before you submit a RAP form, U of A assists you by performing a Study Readiness Assessment (SRA):

Complete the IIS/IIT SRA form, inc. upload draft IRB protocol and other docs on form; and click Submit


IIS/IIT SRA

Clinical Research Administration Clinical Trials & Contracting (CRC) takes 1-2 weeks to conduct an SRA to reduce delays and work for you, by helping you:

  • ensure your IRB protocol is clear and has all necessary information 
  • complete any other necessary docs; for ex., trial schedule 

    Image
    wooden blocks in a row with a person draw about to walk across a gap, with a hand drawing a line to make bridge across the gap

Contact CRC with any SRA questions.  

Get budget help with your IIT

Color legend for card sections above: 

  • blue = required by U of A
  • gray = optional: only if applicable
Check your @arizona.edu email regularly

U of A only emails @arizona.edu addresses throughout this process

Next Steps:

Submit RAP form

Draft and gather the following IRB materials to complete the appropriate RAP form.  To request Banner Health data via the RAP, read obtaining Banner access step 3.

Submit to IRB

Submit your protocol to the U of A eIRB after you have confirmed feasibility with the owner of the EHR data (i.e. Banner), which you will receive after you submit the RAP form

Study team complete COI and training 

All study personnel must complete the following. Otherwise, your IRB application will not be approved.

  1. Conflict of interest (COI)) via the U of A eIRB: should pop up in each person’s IRB dashboard after protocol submitted
  2. CITI training (valid for 2 years): “Biomedical Research Investigators” course
  3. If appropriate, “GCP for Clinical Trials with Investigational Drugs and Medical Devices”.

Questions?

Read answers to frequently asked questions (FAQs) below from researchers about EHR data, accurately defining patient cohort(s), getting the data you actually need and approval(s).

If needed, see how to access Banner EHR data

FAQs:

Please see list of U of A Research Admin. IRB links to determine if you need IRB approval and which form(a) to complete. In general, if you need:

  1. individual patient level EHR data, you need IRB approval, as these types of studies are usually determined by IRBs to be human subjects research 
  2. only summary data (numbers of patients and other statistics), then an IRB exemption form should be completed and submitted

Start by reviewing the IRB process and forms

Faculty (not staff) can be the PI (principal investigator) for studies. Residents and students can be a PI only if they complete the PI eligibility form with a faculty member as Co-Investigator.

Most often simply using International Classification of Diseases (ICD) diagnosis codes is not accurate: as EHR data was built for billing, EHR data is not always diagnostically accurate. To accurately define a cohort of patients, you must specify as much detail as possible, including any codes, date ranges, encounter and clinician types, etc. to use from the EHR to find the patients you need. Thus, please:

  1. Reuse existing phenotype algorithms (or components thereof):
    1. Search EHR phenotype libraries published in the Wiley Online LIbrary to find existing computable phenotype algorithms
    2. Determine if algorithms are "fit," using the methods published in JAMIA, for your study
  2. Review design patterns published in JBI for extracting phenotypes from the EHR for best practices in defining inclusion and exclusion criteria to define the most accurate cohort(s) of patients

Most often, yes. Review design patterns published in JBI for extracting phenotypes from the EHR for best practices in defining inclusion and exclusion criteria to define the most accurate cohort of patients. Exclusion criteria most often excludes patients:
  1. with similar phenotypes: For ex.: Exclude type 1 diabetics from a cohort of type 2 diabetics, as sometimes patients are misdiagnosed, and certain billing codes are easily mis-typed
  2. and/or who have not been seen/assessed recently: For ex.: Exclude patients without a recent outpatient visit when searching for “healthy” controls, as patients not seen recently may have moved or passed away; thus, we can no longer know their current state of health
  3. and/or not recently tested for your phenotype: For ex.: Exclude patients without a recent glucose or HbA1c test, when searching for patients to be used as a control for a diabetes study

  1. The HIPAA Privacy Rule requires researchers request only "minimum necessary" Protected Health Information (PHI), per Health and Human Services (HHS), for their research. This does not mean researchers can request all available data, as long all data is de-identified using the Safe Harbor method from HHS. There is ample research proving even if all of those identifiers are removed, more data means patients are easier to identify from the data. Thus, it is recommended to specify both:
    1. minimum PHI, only if any is needed
    2. minimum necessary EHR data overall
  2. More data: 
    1. often results in overfitting statistical models (too many data points)
    2. does not mean you have more knowledge: as EHR data was built for billing, EHR data is not always diagnostically accurate 
    3. = more time for you to transform the data, such as:
      1. filing in missing data
      2. assessing and removing outliers
      3. formatting data for analysis
    4. = more time for honest (data) brokers to extract the data
  3. Most EHR data is not already de-identified; or, if it is, only summary data is available. For instance:
    1. Banner’s Clinical Research Data Warehouse (CRDW) contains PHI
    2. TriNetX only has summary data/statistics

  1. as with inclusion criteria, be as specific as possible and please review design patterns published in JBI for best practices in selecting EHR data for analysis
  2. create a "data abstraction sheet" to make things easier and faster: read answer to next FAQ for how to create this sheet

If requesting more than a few fields/columns of data, it is helpful to you and to Banner; thus faster, for you to create a "data abstraction sheet" or data dictionary: create a list, such as in MS (Microsoft) Excel, which lists the data variables/fields/columns of data needed from the EHR, and defines those variables.  Even if you are not doing a RCR, it is helpful to create this file for the Banner HBs so that they know exactly what data you need and in what format to export the data to you, so that you can easily begin to analyze the data once you have received it.  Include enough information so honest brokers can find and extract the data: for each variable specify as applicable: 

  • expected values and any applicable units of measure such as:
    • numeric minimum and maximum, especially if only a certain range is biologically plausible (for example, blood pressure in mmHg units from 10-220)
    • specific text values (for example: “yes” or “no”, “high” or “low”)
  • date or time ranges
  • data source such as from an encounter or procedure, or only administered medications instead of all ordered medications
  • any other information which would help honest brokers find the data

Yes, honest (data) brokers who extract data from the EHR can de-identify dates while still allowing you to compute the time between various events and diagnoses in the EHR.  
  1. Unless need exact dates to identify seasonal events, etc., request de-identified dates as either:
    1. ages at events
    2. shifted dates which preserve time between events:  If you do not need to know exact dates, but only time between events, then honest data brokers can typically shift dates by a different random integer for each patient, such that you will see dates for each patient, but they will not be the actual dates; rather, they will be dates with the same timeframes between them as the actual dates.  
  2. If you do need actual dates for your study, such as for determining seasonal events (for ex., summer vs. winter), you will need to request an "LDS" (limited data set), which is usually a data set with dates, and maybe zip codes, as the only PHI; but, usually takes longer to approve

Consider the “Context of Evidence" (derived from design patterns published in JBI):
  1. Who captured the evidence, such as: Can only a specialist accurately diagnose the condition?
  2. When it was captured, such as: age of patient, recent, near important events, etc.
  3. Where …: the setting, such as in- or out-patient
  4. Why ...: for billing / insurance claims, or for diagnosis, etc.
  5. How …: objective or definitive vs. subjective or possible

Adjust for erroneous, missing or duplicate data, esp. as can bias results (list below partly from design patterns published in JBI): 
  1. Absence of Evidence ≠ Evidence of Absence [for condition, lab, etc.]
    1. Confirm variable(s) checked, such as controls had lab test done, etc.
    2. Confirm patient and/or condition truly established: require >1 diagnosis date, recent visits, etc.
    3. Rx order ≠ patient took the med (or as Rx): use additional data as needed
    4. Deceased status: Use Social Security or other death data (most deaths occur outside HCOs)
  2. Account for negation and other qualifiers (ex. visit note: “patient was ‘not’ coughing”)
  3. Remove outliers (especially biologically implausible values such as weight of 2,000 pounds) and/or use summary statistics
  4. Delete records missing important data, or impute missing values
  5. Take time into account:
    1. Use distinct date/time intervals
    2. Add time constraints, esp. for transient conditions

If you are getting zero patients, or considerably less patients than you expect with a diagnosis, procedure, lab, etc., then try using more general search terms.  Also make sure if you are searching using diagnosis, procedure or other codes, include decimal points if the vocabulary uses those, such as ICD codes do, or try searching by a higher level code, i.e. just the part of the code before the decimal point. If you are still getting unexpected results, Contact CB2 so they can check your queries, and with Banner if needed.  Sometimes data is inadvertently missing and can be found by Banner even if you cannot find the data in query tools such as TriNetX.

Contact CB2 if you still have questions

About the information on this page

Updated: Spring/Summer 2026

Thank you to the following for the above information and answers to FAQs: the U of A IIT Review Committee; and A.G., data scientist from U of A COM-T.

Disclaimer

The information on this page was written by CB2 to guide researchers. However, CB2 may not be aware of all of the details, especially as processes are subject to change by the owners of the EHR data and the various regulatory bodies within and outside of U of A.  Please Contact CB2 if any instructions on this page do not work for you so CB2 can direct you to the appropriate person(s) with proper authority and verify the information.

Design and graphics by Manuel Snyder