Finding, Testing and Treating High-Risk Probationers and Parolees with HIV (Urban Health Studies: UHS II), California, 2012-2013 (ICPSR 39801)
The Seek, Test, Treat and Retain (STTR) Collaboration Project involved over twenty studies in the fields of HIV and drug abuse. These studies were independently developed, but were chosen for the collaboration because they focused on one or more steps of the HIV treatment cascade: Seek, Test, Treat and Retain. These studies were grouped into Criminal Justice-related studies and Vulnerable Population-related studies. The data collected by these studies included twelve common domains (e.g. demographic characteristics, mental health) in each of which a shared questionnaire or instrument was taken up by the studies and adapted to fit the study. This repository contains the collected data and documentation from the STTR collaboration.
This study in particular is part of the Urban Health Studies project, specifically assessing treatment outcomes of high-risk probationers and parolees with HIV in California from 2012 to 2013.
Proyecto PACTo: Enhanced HIV Care Access and Retention for Drug Users in San Juan, Puerto Rico, 2013-2014 (ICPSR 39791)
Model for Improving Patient Engagement and Data Integration with National Patient-Centered Clinical Research Network (PCORnet) Patient-Powered Research Networks and Payer Stakeholders [Methods Study], United States, 2015-2020 (ICPSR 39639)
Data from healthcare systems, patients and communities, and health plans can support health research. Two types of data sources are
- Patient-powered research networks, or PPRNs. In PPRNs, patients, families, caregivers, and community members share health data with the network. They work closely with researchers to plan and conduct research.
- Health plan research networks, or HPRNs. In HPRNs, networks of health plans have access to health claims data from members for research.
By linking patient records across PPRNs and HPRNs, researchers may be able to do more robust research. To link records, researchers use computer programs to connect the records of people in a PPRN with their claims data in an HPRN. Current methods to link records require use of personal information, such as names and dates of birth. But patients may not want to share this information.
In this project, the research team developed methods for linking data from PPRNs and HPRNs without using patients' personal information.
Statistical Methods for Phenotype Estimation and Analysis Using Electronic Health Records [Methods Study], 2016-2021 (ICPSR 39724)
Researchers can use data from electronic health records, or EHRs, in studies that compare two or more treatments. In these studies, researchers need to identify all patients with the same phenotype. Phenotypes are a person's known traits, like height and weight, or known health problems, like diabetes. However, in EHR data, some data on patient traits or health problems may be missing for some patients.
Missing data in EHRs make it hard to correctly identify all patients with the same phenotype. It's even harder when data are missing due to a patient's health status. For example, patients with uncontrolled diabetes may need more lab tests than patients with controlled diabetes. As a result, researchers who are looking at lab tests may not identify patients with controlled diabetes as having diabetes.
In this project, the research team developed and tested a new statistical method that accounts for missing EHR data to estimate patient phenotypes.
To access the methods and software, please visit the bias_correction GitHub repository.
Validating and Generalizing Personalized Treatment Rules by Leveraging Different Data Sources [Methods Study], United States, 2019-2022 (ICPSR 39735)
Researchers can use data on patient traits such as age, health problems, and treatment preferences, to create personalized treatment rules, or PTRs. PTRs provide doctors with guidance on how to treat patients' health problems based on their traits. But PTRs based on a single data source may not apply to all patients. For example, if researchers create a PTR using data from older people with heart failure, it may not apply to younger people with heart failure.
To avoid this problem, researchers can create PTRs by combining data from many sources. PTRs based on many data sources can help guide treatment for patients with different traits.
In this study, the research team created and tested a new method for creating PTRs using data from multiple sources.
Natural Language Processing (NLP) for Medication Adherence: Complex Semantics and Negation [Methods Study], United States, 2015-2022 (ICPSR 39736)
Clinical notes in electronic health records, or EHRs, can help researchers study treatments. For example, EHR notes may contain information about whether patients take their medicines as directed. But it takes researchers a lot of time to find this information.
Natural language processing, or NLP, methods can help researchers find information in EHR notes. With NLP, computer programs read and identify written language to make it easier to sort and study. But current NLP methods don't work well to find and label text about medicine use.
In this study, the research team created and tested a new NLP method to find and label EHR notes on patients' medicine use.
Using Topic Segmentation to Enhance Concept Parsing and Identification of Negations [Methods Study], Massachusetts, 2019-2023 (ICPSR 39740)
Clinical notes in electronic health records, or EHRs, may contain information that can help researchers study and compare treatments. But it takes researchers a lot of time to find information in EHR notes.
Natural language processing, or NLP, methods can help researchers find information in EHR notes. With NLP, computer programs read and identify written language to make it easier to sort and study. But in EHR notes, some sentences may contain more than one topic. Also, EHR notes may discuss a single topic over many sentences. In these cases, current NLP methods don't work well to find complete and accurate information about a specific topic.
In this study, the research team developed and tested new NLP methods to identify topics from EHR notes.
Randomize Everyone: Creating Valid Instrumental Variables for Learning Health Care Systems [Methods Study], New Hampshire, 2016-2022 (ICPSR 39717)
Comparative effectiveness research, or CER, compares two or more treatments. In some CER studies, researchers use patient data from electronic health records, or EHRs, to compare treatments. But patient traits like age may affect doctors' and patients' choice of treatments, which can bias results. Using EHR systems to identify eligible patients and assign them to treatments by chance could improve results of CER studies that use EHR data.
In this study, the research team explored the views of patients, clinic staff, and clinicians, such as doctors or nurses, on doing CER studies in clinics. The team also tested software with a widely used EHR system. The software finds patients who qualify for a study. During a clinic visit, the software prompts doctors to invite patients to take part in the study. If patients agree, the software assigns patients by chance to a treatment.
Improving Clinical Effectiveness Research (CER)/Patient-Centered Outcomes Research (PCOR) Methods for Analyzing Linked Data Sources in the Absence of Unique Identifiers [Methods Study], United States, 2011-2022 (ICPSR 39731)
Researchers often combine data from different sources, such as insurance claims and health records, to get a better picture of patients' health and use of health care. Researchers use unique identifiers, like Social Security numbers, to connect patient records and make them more complete. But sometimes this approach doesn't work well, especially when records don't have much personal information. Having limited personal data can lead to errors when linking records.
In this study, the research team created new methods to link data sets with limited personal information. Then they compared the new methods with existing ones. They also applied the new methods with real patient data.
Unlocking Clinical Text in Electronic Medical Records (EMR) by Query Refinement Using Both Knowledge Bases and Word Embedding [Methods Study], Ohio, 2006-2022 (ICPSR 39734)
Electronic health records, or EHRs, have information about a patient's health such as test results, diagnoses, and treatments. EHRs also have clinical notes that doctors and patients can use to track goals and decisions.
Clinical notes may be useful for research or to help improve care. But it's hard to get information from these notes across large groups of patients. The notes may use different ways to describe the same thing. For example, high blood pressure may be called hypertension. Also, the notes may use abbreviations or have spelling mistakes.
In this project, the research team designed and built a search engine to make EHR notes easier to search and use for patient care and research.
Statistical Methods and Designs for Addressing Correlated Errors in Outcomes and Covariates in Studies Using Electronic Health Records Data [Methods Study], Tennessee, 2016-2021 (ICPSR 39726)
Electronic health records, or EHRs, have data on patient traits, health problems, and treatments. Researchers can use EHR data to study how treatments work or which patient traits affect health outcomes. But EHR data can have errors.
The best way to get accurate EHR data is to closely review patients' original records. But reviewing all patient records isn't possible when many patients are in a study. In such cases, researchers can review and correct records for a few patients and use the revised records to adjust data for all patients. But existing methods for using revised records don't address some kinds of errors, such as errors that are related. For example, errors in a treatment starting date can lead to mistakes in the data on length of treatment.
In this project, the research team created and tested new methods to improve the accuracy of EHR data. The new methods corrected records from some patients. Then the team used the corrections to address related errors for all patients.
To access the methods and software, please visit the MeasurementErrorMethods GitHub repository.
Building Data Registries with Privacy and Confidentiality for Patient-Centered Outcomes Research (PCOR) [Methods Study], 2020 (ICPSR 39579)
Researchers can use patient health data to compare treatments. But these data may include information, like names or social security numbers, that could identify patients. Researchers use different methods to remove such information and protect patients' privacy. Some methods work well to protect privacy but may make data less useful for research. Other methods don't protect privacy well enough.
Current methods for protecting privacy don't work well when:
- The number of patients in the data set is smaller than the number of data fields, such as patient traits or health conditions, and data are updated many times
- Patients' health and treatments are measured at more than one point in time
- Data are displayed as a graph to better capture some types of content
In this study, the research team created three new methods. The team wanted to see if the new methods better protect patient privacy but also make sure data remain useful for research.
To access the methods and software, please visit the AIMS Group at Emory University.
Causal Analyses of Electronic Health Record Data for Assessing the Comparative Effectiveness of Treatment Regimens [Methods Study], United States, 2014-2019 (ICPSR 39581)
Patients with chronic health problems, such as diabetes, often need to change treatment plans over time to improve their health. To help with this process, doctors can monitor patients' health through follow-up clinic visits and lab tests. Doctors may also suggest changing a treatment plan in response to visits or lab test results. When a treatment plan changes in this way, it's called a dynamic treatment plan. In this study, the research team developed and tested new statistical methods to learn how dynamic treatment plans and choices about follow-up care affect patients' health. These methods use electronic health records, or EHRs. Using EHRs is helpful because they have data on
- What treatments patients have received over time
- How treatments have affected patients' health
- Follow-up information such as lab test results
But the data may differ for patients based on when and why they go to the doctor. These differences make it hard for researchers to accurately know the effect of dynamic treatment plans across many patients.
To access the methods and software, please visit the simcasual R Package.
Statistical Methods for Missing Data in Large Observational Studies [Methods Study], Georgia, 2013-2018 (ICPSR 39526)
Health registries record data about patients with a specific health problem. These data may include age, weight, blood pressure, health problems, medical test results, and treatments received. But data in some patient records may be missing. For example, some patients may not report their weight or all of their health problems.
Research studies can use data from health registries to learn how well treatments work. But missing data can lead to incorrect results. To address the problem, researchers often exclude patient records with missing data from their studies. But doing this can also lead to incorrect results. The fewer records that researchers use, the greater the chance for incorrect results.
Missing data also lead to another problem: it is harder for researchers to find patient traits that could affect diagnosis and treatment. For example, patients who are overweight may get heart disease. But if data are missing, it is hard for researchers to be sure that trait could affect diagnosis and treatment of heart disease.
In this study, the research team developed new statistical methods to fill in missing data in large studies. The team also developed methods to use when data are missing to help find patient traits that could affect diagnosis and treatment.
To access the methods, software, and R package, please visit the Long Research Group website.
Building Patient-Centered Outcomes Research Value and Integrity with Data Quality and Transparency Standards [Methods Study], United States, 2013 - 2018 (ICPSR 39529)
Many healthcare systems use electronic health records. Researchers use data from these records in their studies. Some records have missing or incorrect data. When this happens, people might not be able to trust a study's results. The research team wanted to:
- Create guidance to judge whether data that a study used were high quality
- Find new ways to display the quality of data
- Learn why researchers don't always report the quality of data that they used in studies
To access the methods and software, please visit the DQCODE-A-Thon GitHub.
Measuring and Talking to Patients About the Accuracy of Data Used in Patient-Centered Outcomes Research [Methods Study], North Carolina and Arkansas, 2013-2018 (ICPSR 39515)
For research studies, researchers can use data about patients' health and treatments from electronic health records, or EHRs. They may also collect self-reported data directly from patients. But a patient's EHR and self-reported data may not always agree. For example, differences may exist between the medicines that patients report taking and the medicines listed in their EHRs. Researchers don't know which of these two data sources is the most accurate.
In this project, the research team looked at EHR and self-reported data to learn which data source was more accurate.
Development of a Causal Inference Toolkit for Patient-Centered Outcomes Research [Methods Study], 2013-2018 (ICPSR 39533)
Comparative effectiveness research compares two or more treatments to see which one works better for which patients. One type of research study is a randomized controlled trial, or an RCT. In an RCT, the research team assigns patients to a treatment by chance.
Other types of studies use information from health records and registries. Registries store data about patients with a specific health problem. They often include information on how each patient responds to a treatment. Because researchers don't assign treatments by chance in such studies, differences in how patients respond to a treatment may be from the treatment or something else, such as a patient's age or the severity of their illness. In studies using registries and health records, researchers apply statistical approaches, called causal inference methods, to estimate how treatments work. At the same time, they look at other things that could affect results, like a patient's age.
Researchers can choose among many different causal inference methods. But they may have a hard time knowing which methods to use or how to use complex methods correctly. In this study, the research team made an interactive online guide for researchers. The guide, called CERBOT, helps researchers design studies and select these methods.
Methods for Analysis and Interpretation of Data Subject to Informative Visit Times [Methods Study], 2013-2018 (ICPSR 39474)
Comparative effectiveness research compares two or more treatments to see which one works better for certain patients. Researchers often use data from patients' electronic health records to compare different treatments. This study addresses some problems that can arise from this practice. In some long-term research studies, researchers use data collected when patients in the studies see their doctors. Regularly scheduled doctor visits, called well visits, include yearly checkups or periodic blood pressure checks. Other doctor visits, called sick visits, occur when a patient feels sick or needs special care.
Well and sick visits can produce different types of health record data. In addition, test results at sick visits may be different from results at well visits. Using data from sick visits may inappropriately influence, or bias, a study's results. Also, patients may go to the doctor more often when they have symptoms or chronic health problems. Researchers may then collect more data from these patients than they collect from the healthier patients. Unequal amounts of data per patient make it harder to compare treatment results.
For this study, the research team created three tests to find if data from sick visits lead to bias in a study's findings. The team also compared standard and newer statistical methods for analyzing data that include sick visits. Researchers designed the newer methods to reduce bias from data obtained at sick visits. With less biased results, doctors can be more certain about which treatment worked better for certain patients.
Cuyahoga County, Ohio, Heroin and Crime Initiative: Informing the Investigation and Prosecution of Heroin-Related Overdose, 2012-2021 (ICPSR 38295)
In 2013, the Cuyahoga County (Ohio) Medical Examiner's Office (CCMEO) and the Regional Forensic Science Laboratory developed the Heroin Involved Death Investigation (HIDI) alert system and protocol in response to a substantial increase in opioid-related overdose fatalities. The HIDI protocol is designed to support a safe, coordinated, and rapid response to an active, suspected opioid-overdose death scene, or suspected opioid-overdose deaths occurring at hospitals that are not considered active scenes, by alerting investigators to potential dangers and facilitating the timely protection of scene integrity and evidence collection in order to successfully investigate and prosecute drug traffickers.
The primary goals of the project were to:
- Complete extended coding of local medical examiner decedent data--investigative reports and toxicology to identify demographic or geographic trends or patterns of overdose deaths, as well as paraphernalia and evidence present at death scenes that may be useful to prosecutions;
- Examine the efficiency of how cases flow through the investigative and prosecutorial stages and how these could be improved;
- Identify key variables that may contribute to the successful indictment of traffickers connected to fatal and non-fatal overdose cases; and
- Evaluate the implementation and perceived effectiveness of the Cuyahoga County HIDI protocol.
This multi-method project involved three phases of data collection and analysis. First, a forensic epidemiologist coded and analyzed existing CCMEO records for decedent toxicology and death scene characteristics, focusing on drug-related fatalities. Second, county and federal cases prosecuted for drug trafficking, especially those linked to deaths, were systematically reviewed to determine what evidence was deemed important for successful indictment. Third, interviews and focus groups were conducted with key stakeholders from local and federal law enforcement, intelligence analysts, public health officials, and local and federal prosecutors to learn about the HIDI protocol.
Data and documentation for interviews and focus groups will be made available in a future update.
Functional Independence in Children at a Pediatric Clinic in Guanajuato, Mexico, 2004-2013 (ICPSR 37068)
This study sought to evaluate the functional independence in children at a Centers for Pediatric Rehabilitation Teleton (CRIT) facility in Guanajuato, Mexico through the use of the WeeFIM Instrument (0-3 Module). The dataset in this collection was generated in May 2013 from electronic health records for secondary analysis of de-identified data. The goal of CRIT, that this research sought to evaluate, was to improve social integration for children with disabilities in Mexico through comprehensive rehabilitation services, including physical therapy, occupational therapy, neurotherapy, speech therapy, physical and rehabilitation medicine, psychology, social integration, and school for parents.
The collection includes one dataset (35 variables, 5,993 cases). Demographic variables included in the collection: Age, gender, and city of residence.
Aging of Veterans of the Union Army: Surgeons' Certificates, United States, 1862-1940 (ICPSR 2877)
This data collection, Aging of Veterans of the Union Army: Surgeons' Certificates, United States, 1862-1940, constitutes a portion of the historical data collected by the project "Early Indicators of Later Work Levels, Disease, and Death." With the goal of constructing datasets suitable for longitudinal analyses of factors affecting the aging process, the project collects military, medical, and socioeconomic data on a sample of white males mustered into the Union Army during the Civil War. The surgeons' certificates contain information from examining physicians to determine eligibility for pension benefits. Also included are questions regarding the age, occupation, residence, and military experience of the veterans. These data can be linked to "Aging of Veterans of the Union Army: Military, Pension, and Medical Records, 1820-1940" (ICPSR 6837) and "Aging of Veterans of the Union Army: United States Federal Census Records, 1850, 1860, 1900, 1910" (ICPSR 6836) using the variable "recidnum."
Enhanced Data to Accelerate Complex Patient Comparative Effectiveness Research, 2006-2009 [United States] (ICPSR 34639)
Purpose: Develop an easy-to-use data product to facilitate comparative effectiveness research involving complex patients.
Scope: Claims data can be difficult to use, requiring experience to most appropriately aggregate to the patient level and to create meaningful variables such as treatments, covariates, and endpoints. Easy to use data products will accelerate meaningful comparative effectiveness research (CER).
Methods: This project used data from the Medicare Chronic Condition Data Warehouse for patients hospitalized with acute myocardial infarction (AMI) or stroke in 2007 with two-year follow-up and one-year pre-admission baseline. The project joined over 100 raw data files per condition to create research-ready person- and service-level analytic files, code templates, and macros while at the same time adding uniformity in measures of comorbid conditions and other covariates. The data product was tested in a project on statin effectiveness in older patients with multiple comorbidities.
Results: A programmer/analyst with no administrative claims data experience was able to use the data product to create an analytic dataset with minimal support aside from the documentation provided. Analytic dataset creation used the conditions, procedures, and timeline macros provided. The data structure created for AMI adapted successfully for stroke. Complexity increased and statin treatment decreased with age. The two-year survival benefit of statins post-AMI increased with age.
Conclusion: Claims data can be made more user-friendly for CER research on complex conditions. The data product should be expanded by refreshing the cohort and increasing follow-up. Action is warranted to increase the rate of statin use among the oldest patients.
Data Access: These data are not available from ICPSR. The data cannot be made publicly available. Data are stored on University of Iowa College of Public Health secure servers, and may be used only for projects covered within the aims of the original research protocol and Centers for Medicare and Medicaid Services (CMS)-approved data use agreement. Data sharing is allowed only for research protocols approved under data re-use requests by the CMS privacy board. The CMS process for data re-use requests is described at Research Data Assistance Center (ResDac). Please note that as of May 2013, the DUA covering this work is set to expire February 1, 2014. Thereafter, per the terms of the DUA, datasets created for this project may not be available.
User guides are available from ICPSR for detailed descriptions of the data products, including a user guide for Acute Myocardial Infarction (AMI) Analytic Files and a user guide for Stroke and Transient Ischemic Attack (TIA) Analytic Files. Data dictionaries are available upon request. Please contact Nick Rudzianski ([email protected] or 319-335-9783) for more information.
Clinical Database to Support Comparative Effectiveness Studies of Complex Patients, 2005-2010 [United States] (ICPSR 34644)
Overview: The goal of the project was to develop a unique database linking chronic disease clinical data from an electronic medical record (EMR) of a large academic healthcare system to multi-payer claims data. The longitudinal relational database can be used to study clinical effectiveness of many diagnostic and treatment interventions. The population of patients used consisted of those patients who were attributed to the University of Michigan Health System (UMHS) as continuing care patients, who are also in adjudicated and validated chronic disease registries.
Data Access: These data are not available from ICPSR. The data are restricted to use by the principal investigator and cannot be shared.
North Carolina Integrated Data for Researchers (NCIDR): Merged Behavioral Health Data from Four Publicly-Funded Sources in North Carolina, July 2007-June 2011 (ICPSR 34542)
Overview
The North Carolina Integrated Data for Researchers (NCIDR, pronounced "Insider") was funded to develop a robust research data warehouse for storing merged data from four different publicly-funded sources in North Carolina. Community Care of North Carolina maintains this unique database on behalf of the North Carolina Department of Health and Human Services, and facilitates requests for access to integrated behavioral health services data for research purposes. This expanded data set has great value to researchers in North Carolina and elsewhere. The NCIDR warehouse is a unique resource for obtaining the most complete picture of the health services delivered to people with severe mental illness in North Carolina. Few examples of such an integrated warehouse exist anywhere else, and NCIDR makes it possible for researchers and epidemiologists to conduct comparative effectiveness research related to people with these conditions.
The merged data sources include:
- Medicaid claims and enrollment data for nearly 1 million individuals with MH, DD and SA diagnoses.
- IPRS (Integrated Payment and Reporting System) -- covers primarily outpatient mental health services for people that do not qualify for Medicaid (approximately 250,000 individuals).
- HEARTS (Healthcare Enterprise Accounts Receivable Tracking System) -- documents services delivered by inpatient State Mental Health facilities (approximately 25,000 individuals).
- Piedmont Behavioral Health (Medicaid waiver) -- behavioral health encounter data from Medicaid's capitated arrangement in five counties (approximately 25,000 individuals).
Data are available for four state fiscal years, 2008 through 2011 (7/1/2007--6/30/2011). Each year has three data sets (claims, client, provider) in addition to multiple lookup tables with definitions. Population includes any Medicaid client with a claim that contains any MH, DD or SA (290xx through 319xx) diagnosis at least once in the four year time period, plus all clients appearing in the other 3 data sources. Requests will need to specify required time periods and clearly define the population being studied. Note that individuals dually enrolled with Medicare during months in which they are dually enrolled are excluded.
The data available for future use will include a claims file, a client file, a provider file and multiple lookup files. The claims file contains approximately 83 columns including 30 columns for diagnosis codes. The client file is approximately 78 columns which displays 12 columns each (one for each month in the SFY) for eligibility, enrollment, assigned network, primary care physician and dual status indicator. There are 5 columns in the Provider file. The lookup file will contain tables for every code that requires a description. The data will be parsed into individual state fiscal years.
Data Access
These data are not available from ICPSR. The process for requesting access to the integrated data is detailed on the NCIDR Web site, specifically the Request Process Overview page. Researchers interested in requesting access are strongly encouraged to contact the Director of Evaluation at [email protected] to discuss his/her intent to submit a Request Form. Some may also need to complete the Data Use Agreement if requesting data that are not completely de-identified.
Although IRB approval must be documented prior to release of data, NCIDR will accept applications with conditional IRB approval and researchers may discuss projects with the Director of Evaluation at any stage of development. A Research Oversight Committee (ROC) that includes stake holders from the NC Department of Health and Human Services (DHHS), the NC Division of Medical Assistance (DMA), the NC Division of State Operated Healthcare Facilities (DSOHF), the NC Division of Mental Health, Developmental Disabilities and Substance Abuse Services (DMHDDSAS), the NC Office of Rural Health and Community Care (ORHCC), the Community Care of North Carolina (CCNC) and other community partners will review research requests and grant approval when applicable. Once approved, please note that CCNC must charge a nominal fee of $3,000 to cover costs related to the preparation and transmission of files to the researcher (additional charges may apply depending on the specific programming needs).
Enhancing Analytic Abilities to Identify Complex Patients in 225 Practice Partner Research Network (PPRNet) Practices in 42 states: July 2010-July 2012 (ICPSR 34554)
Overview
Through electronic data collection and improving the efficiency of existing data processes to allow both more complete and specific identification of chronic illness, the study objectives included:
- Greatly enhance the scope of existing algorithms to permit comprehensive identification of the 20 chronic conditions key to primary care.
- Improve the specificity of the existing algorithms to permit more precise automated identification of chronic conditions, limiting the amount of human review required.
- Revise the algorithms to permit identification of more than one condition in a text string.
The investigators developed advanced SAS text string search algorithms and developed a modified parsing table that included inclusion and exclusion patterns and resultant diagnoses. The automation searches through each input text string for the inclusion pattern that is not equivalent to the exclusion pattern and maps the string to the corresponding resultant diagnosis. This technique allows the search functions to be easily modified to include additional search criteria and scaled to encompass additional conditions.
Data Dictionary
A data dictionary for 24 chronic conditions was developed. The dictionary assigns ICD-9 diagnosis codes to problem list text in electronic health record data. The dictionary contains 78,458 records and exists in two forms, a Microsoft Access database and a SAS 9.2 dataset. The Microsoft Access database contains 24 tables, one for each condition. The SAS 9.2 dataset contains four fields. The 24 chronic conditions for which problem list text data were examined and assigned to ICD-9 codes. Conditions include Alcohol Use Disorder, Asthma and Allergic Rhinitis, Atherosclerosis, Atrial Fibrillation, Cerebrovascular Disease, Chronic Liver Disease, COPD, Chronic Renal Disease, Coronary Disease, Dementia, Depression, Diabetes Mellitus, Epilepsy, GERD, Heart Failure, Hyperlipidemia, Hypertension, Migraine Headache, Obesity, Osteoarthritis, Osteopenia/Osteoporosis, Parkinson's Disease, Peptic Ulcer Disease and Rheumatoid Arthritis.
Data Access
The data dictionary is not available from ICPSR. For use arrangements, please contact Ruth G. Jenkins, PhD ([email protected]) or Steven M. Ornstein, MD ([email protected]) at the Practice Partner Research Network (PPRNet), Medical University of South Carolina.
Expansion Research Capability to Study Comparative Effectiveness in Complex Patients, 2007-2010 [Tampa, St. Petersburg, and Clearwater, Florida] (ICPSR 34544)
Overview
The Florida Department of Health and the Florida Cancer Data System (FCDS) collaborated with a hospital network composed of nine clinical facilities, to capture electronic medical records (EMR) data of patients who were diagnosed with or treated for invasive breast cancer from 2007 to 2010. Certain hospital data elements were available throughout 2006. An additional year of 2011 follow-up data was also available for a subset of patients receiving medication treatment. The purpose of the data capture was to advance patient-centered outcomes research to reduce the morbidity and mortality of cancer and other comorbidities.
A breast cancer pilot study was also conducted from a subset of all transmitted EMR records, consisting of admission records with a principal and/or secondary ICD-9-CM diagnosis between 174.0 and 174.9. The subset dataset was then linked to the central cancer registry using patient social security number, first and last name, and date of birth. Using a deterministic matching algorithm a total of 11,506 unique patients were matched to a patient in the FCDS database, resulting in 12,804 primary tumors and 53,940 unique hospital admission records. While the hospital EMR defined the patient dataset, all registry records for that patient were included in the final breast cancer pilot database, regardless of the reporting hospital or the date of diagnosis. This was to ensure capture of the entire diagnostic and treatment profile for each breast cancer patient.
Data Access
These data are not available from ICPSR. The data contain confidential information that can directly identify a patient. There are also reporting facility data. Therefore, to obtain these data, researchers will need to follow the Florida Cancer Data System data-sharing agreement process, as outlined on the FCDS data sharing request.