Hispanic Established Populations for the Epidemiologic Study of the Elderly (HEPESE) Wave 10, 2020-2021 [Arizona, California, Colorado, New Mexico, and Texas] (ICPSR 39219)
The Hispanic EPESE provides data on risk factors for mortality and morbidity in older Mexican Americans in order to contrast how these factors operate differently than in non-Hispanic Whites, African Americans, and other major ethnic groups.
The Wave 10 dataset comprises the ninth follow-up of the baseline Hispanic Established Populations for the Epidemiologic Studies of the Elderly, 1993-1994: [Arizona, California, Colorado, New Mexico, and Texas] (ICPSR 2851). The baseline Hispanic EPESE collected data on a representative sample of community-dwelling Mexican Americans, aged 65 years and older, residing in the five Southwestern states of Arizona, California, Colorado, New Mexico, and Texas.
The public-use data covers demographic characteristics (age, sex, type of Hispanic ethnicity, income, education, marital status, number of children, employment, and religion), height, weight, social and physical functioning, chronic conditions, related health problems, health behaviors, self-reported use of dental, hospital, and nursing home services, and depression. Subsequent follow-ups allow examination of the predictors of mortality, changes in health outcomes, institutionalization, changes in living arrangements, as well as changes in life situations and quality of life.
During this 10th Wave, 131 re-interviews were conducted either in person or by proxy, with 77 of the original respondents interviewed in 1993-1994. This Wave also includes 54 re-interviews from the 902 new respondents added at Wave 5 in 2004-2005. All respondents were aged 90 and over at Wave 10.
The wave 10, was conducted over 2020 and 2021 and consisted of two components, a pre-COVID in-person component and a post-COVID telephone component to the informant only. The pre-COVID in-person interviews were conducted from January 1, 2020 to March 17, 2020 (N=131 respondents; N=122 informants). In March 2020, the in-person interviews were suspended due to the COVID-19 pandemic. From April 1, 2021 to July 1, 2021, telephone interviews were conducted only with informants (n = 101). The study team collected information on health, function, social situation, finances, and general well-being of the older Hispanic EPESE respondents. Information was also collected on the informant's health, function, and caregiver responsibilities and burden. In Wave 10, during the telephone interviews conducted with the informant, the study team collected information related to their experiences during the first year of the COVID-19 pandemic and their contemporary experiences around the time of widespread vaccine availability in the United States.
Standardization of Uveitis Nomenclature ("SUN"), Global, 2004-2021 (ICPSR 38665)
The uveitides are a collection of >30 diseases characterized by intraocular inflammation. Collectively, they are the 5th or 6th leading cause of blindness in the United States, and the cost of treating them has been estimated be comparable to the cost of treating diabetic retinopathy. These diseases may be due to an intraocular or systemic infection, associated with a systemic rheumatic or other inflammatory disease or eye-limited and immune-mediated. They often are grouped by the primary site of inflammation as anterior, intermediate, posterior, or panuveitides, with the primary site of clinically detected inflammation in the anterior chamber, vitreous, retina and/or choroid, or entire eye, respectively. Clinical and translational research in the field of Uveitis has been hampered by the lack of gold standards for diagnosis and a lack of consistency in the diagnosis of these diseases. Agreement among uveitis experts on diagnosis has been modest at best with some pairs of experts having agreement no better than chance alone. Research in other branches of medicine has been greatly facilitated by the development of classification criteria. Classification criteria are a type of diagnostic criteria for research purposes. Classification criteria differ from clinical diagnostic criteria in that, if a trade-off is needed, classification criteria emphasize specificity, i.e. the classification of a group of patients definitely thought to have the disease. The Standardization of Uveitis Nomenclature (SUN) Working Group is an international group of 99 investigators from 64 centers in 22 countries, with expertise in uveitis, informatics, consensus techniques, database management, ophthalmic image interpretation, and machine learning.
The SUN Working Group's project "Developing Classification Criteria for the Uveitides" goal was to develop classification criteria for 25 of the most common uveitides. The project proceeded in 4 phases: 1) informatics, 2) case collection, 3) case selection, and 4) machine learning.
The informatics phase resulted in a standardized language to describe the uveitides and a successful mapping of terms and phrases to individual diseases. The informatics phase led to the creation of a menu-driven, hierarchical, data collection tool for the case collection phase. The case collection phase consisted of the SUN Working Group entering retrospective and de-identified data on 5766 cases (total) into a preliminary database. The goal of case collection was 100-250 cases of each of the 25 diseases. Because of the lack of gold standards for diagnosis and the modest agreement among experts on diagnosis, collected cases were reviewed, and only cases with a supermajority (>75%) agreement that they were the disease were selected for the final database. Case selection consisted of committees of 9 uveitis experts reviewing the cases and voting on whether or not they were the disease. This process used formal consensus techniques, including nominal group techniques. Committees were geographically and school-of-thought dispersed. Cases achieving a supermajority agreement that they represented the disease were included in the final database. Cases with a supermajority agreement that they were not the disease were excluded, and cases without a supermajority agreement were tabled. Only 1% of cases were tabled. The final database consisted of 4046 cases (70% of those collected). The consensus diagnosis was used as the accepted diagnosis in the machine learning phase.
Following case selection, the final database was subjected to machine learning as to features that distinguished the diseases. For machine learning the case data were split into a training set and a validation set. Cases were analyzed within anatomic class, with cases from those diseases with protean presentations used in more than one class. Multiple different machine learning approaches were used, including classification and regression trees, random forests, support vector machines and multinomial logistic regression, all of which tended to have a high degree of agreement on the distinguishing features and relatively similar accuracies. The method chosen for reporting was multinomial logistic regression. Boruta analyses were used to determine a parsimonious set of criteria, and the Quine-McCluskey algorithm to create a logical set of Boolean expressions that correctly classified the diseases. Because different tests or clinical features (e.g. hilar adenopathy in patients with sarcoid can be seen on chest radiography or on chest computed tomography) might be able to indicate the disease, feature engineering was used during machine learning. The set of Boolean expressions from the machine learning were then translated into English phrases ("final rules") for clinical use. As a back check on the translation, a random set of 10% of cases was subjected to classification by an observer masked as to the consensus diagnosis. These performance of these criteria (>90% accuracy within class for machine learning on the validation set and >95% accuracy of the "final rules" by the masked observer) suggest that they can be used in clinical and translational research.
Following the machine learning phase, a meeting of the SUN Working Group was held in December 2019 to review the work and the proposed criteria. The result of this meeting was an approval of the criteria. Twenty-six manuscripts were prepared, one dealing with the methods used, and 25 disease-specific manuscripts with the criteria for each disease. The individual diseases addressed in this project included: cytomegalovirus anterior uveitis, Fuchs uveitis syndrome, herpes simplex anterior uveitis, juvenile idiopathic arthritis-associated anterior uveitis, spondyloarthritis/HLA-B27-associated anterior uveitis, tubulointerstitial nephritis with uveitis, varicella zoster anterior uveitis, pars planitis, intermediate uveitis non-pars planitis type, multiple sclerosis-associated intermediate uveitis, acute posterior multifocal placoid pigment epitheliopathy, birdshot chorioretinitis, multiple evanescent white syndrome, multifocal choroiditis with panuveitis, punctate inner choroiditis, serpiginous choroiditis, Behçet disease uveitis, sympathetic ophthalmia, Vogt-Koyanagi-Harada disease, sarcoidosis-associated uveitis, acute retinal necrosis syndrome, cytomegalovirus retinitis, syphilitic uveitis, toxoplasmic retinitis, and tubercular uveitis. The goal is for these criteria to be used as the underpinning for future clinical and translational research in the field of Uveitis.
Cuyahoga County, Ohio, Heroin and Crime Initiative: Informing the Investigation and Prosecution of Heroin-Related Overdose, 2012-2021 (ICPSR 38295)
In 2013, the Cuyahoga County (Ohio) Medical Examiner's Office (CCMEO) and the Regional Forensic Science Laboratory developed the Heroin Involved Death Investigation (HIDI) alert system and protocol in response to a substantial increase in opioid-related overdose fatalities. The HIDI protocol is designed to support a safe, coordinated, and rapid response to an active, suspected opioid-overdose death scene, or suspected opioid-overdose deaths occurring at hospitals that are not considered active scenes, by alerting investigators to potential dangers and facilitating the timely protection of scene integrity and evidence collection in order to successfully investigate and prosecute drug traffickers.
The primary goals of the project were to:
- Complete extended coding of local medical examiner decedent data--investigative reports and toxicology to identify demographic or geographic trends or patterns of overdose deaths, as well as paraphernalia and evidence present at death scenes that may be useful to prosecutions;
- Examine the efficiency of how cases flow through the investigative and prosecutorial stages and how these could be improved;
- Identify key variables that may contribute to the successful indictment of traffickers connected to fatal and non-fatal overdose cases; and
- Evaluate the implementation and perceived effectiveness of the Cuyahoga County HIDI protocol.
This multi-method project involved three phases of data collection and analysis. First, a forensic epidemiologist coded and analyzed existing CCMEO records for decedent toxicology and death scene characteristics, focusing on drug-related fatalities. Second, county and federal cases prosecuted for drug trafficking, especially those linked to deaths, were systematically reviewed to determine what evidence was deemed important for successful indictment. Third, interviews and focus groups were conducted with key stakeholders from local and federal law enforcement, intelligence analysts, public health officials, and local and federal prosecutors to learn about the HIDI protocol.
Data and documentation for interviews and focus groups will be made available in a future update.
SFGR cross-sectional evaluation South Carolina (ICPSR 193704)
Care pathways of individuals with tuberculosis before and during the COVID-19 pandemic in Bandung, Indonesia (ICPSR 192709)
2010 United States Census Tract Community Type Classification and Neighborhood Social and Economic Environment Score for 2000 and 2010, from the Diabetes Location, Environmental Attributes, and Disparities (LEAD) Network (ICPSR 38645)
Substance Use Among American Indian Youth: Epidemiology and Etiology, [United States], 2015-2020 (ICPSR 37997)
This study is a continuation of an ongoing 40+ year surveillance effort assessing the levels and patterns of substance use among American Indian (AI) adolescents attending schools on or near reservations. The current set of data is from the most recent funding cycle, 2015-2020. During this funding cycle, annual samples across various geographic regions in which reservation-based AI residents reside were obtained and school-based surveys were completed. In addition to the annual epidemiology of substance use, data pertaining to risk and protective factors, including cultural-ethnic identity, perceived discrimination, family factors, and individual risk and protective factors were obtained. It should be noted that two major changes were made during this funding cycle:
1) The wording of substance use variables was altered to mirror wording from Monitoring the Future to allow for direct comparisons between the two studies.
2) All data during this funding cycle were obtained online using Qualtrics.
National Academy of Sciences-National Research Council Twin Registry (NAS-NRC Twin Registry), 1958-2013 [RESTRICTED] (ICPSR 36234)
In 1958, the Medical Follow-up Agency (MFUA) of the Institute of Medicine began a project to identify twins who had jointly entered military service during World War II. In the end, MFUA identified nearly 16,000 White male twin pairs born 1917-1927 in which both members had served in the military. These twins comprise the National Academy of Sciences-National Research Council World War II Twin Registry (NAS-NRC Twin Registry). This collection represents data from service records, a mailed questionnaire assessing zygosity, and repeating health surveys, including information on education, employment history, and earnings.
There are nine datasets associated with this restricted-use collection:
1) The Administrative dataset includes demographic, zygosity, service history, mortality, and questionnaire participation data;
2) The Service and Other Records dataset contains information collected from service records, physical exam data, cognitive test data, and dental records;
3) The Questionnaire 2 dataset consists of data collected in the first mailed questionnaire sent in 1965 about pain, illnesses, smoking habits, alcohol consumption, and employment;
4) The Questionnaire 3 dataset includes data from the baseline epidemiological questionnaire sent in 1974 about number and sex of children, religious attendance, education, income, and occupation;
5) Questionnaire 7, mailed in 1985, contains similar topics as in Questionnaire 2, and includes data about health conditions such as diabetes, as well as feelings about work and retirement;
6) Questionnaire 8 was mailed in 1998 was the third epidemiologic questionnaire. This dataset is comprised of overlapping topics with Q2 and Q7, and has additional data about feelings, prescription medications, activity levels, the Geriatric Depression Scale, and parental death status;
7) The NEO Personality Inventory dataset includes responses to the NEO Five-Factor Personality Inventory mailed in 2005-2006;
8) The Service and Death Records dataset (VDE access only) contains information about date and place of birth, state at induction, disciplinary measures during service, decorations received during service, indicator for those known to have been POWs, reason for separation from the military, age at death if died over age 90, and cause of death. Some of this information was obtained from the re-read of service records and is thus available only for a subset of 6357 men;
9) Diagnoses dataset (VDE access only) contains data about medical conditions diagnosed between 1935 and 1985 that were abstracted from a variety of medical records over the course of the study. The diagnoses were coded using the International Classification of Disease system (WHO, 2015).
Epidemiologic Catchment Area Program Sites 1-4, 1979-1983 with National Death Index Data through 2007 (ICPSR 36621)
The Epidemiologic Catchment Area (ECA) program of research was initiated in response to the 1977 report of the President's Commission on Mental Health. The purpose was to collect data on the prevalence and incidence of mental disorders and on the use of and need for services by the mentally ill. Independent research teams at five universities (Yale University, Johns Hopkins University, Washington University, Duke University, and University of California at Los Angeles), in collaboration with the National Institute for Mental Health, conducted the studies with a core of common questions and sample characteristics. The sites were areas that had previously been designated as Community Mental Health Center catchment areas: New Haven, Connecticut, Baltimore, Maryland, St. Louis, Missouri, Durham, North Carolina, and Los Angeles, California. Each site sampled over 3,000 community residents and 500 residents of institutions, yielding 20,861 respondents overall. The longitudinal ECA design incorporated two waves of personal interviews administered one year apart and a brief telephone interview in between (for the household sample). The diagnostic interview used in the ECA was the NIMH Diagnostic Interview Schedule (DIS), Version III (with the exception of the Yale Wave I survey, which used Version II). Diagnoses were categorized according to the DIAGNOSTIC AND STATISTICAL MANUAL OF MENTAL DISORDERS, 3rd Edition (DSM-III). Diagnoses derived from the DIS include manic episode, dysthymia, bipolar disorder, single episode major depression, recurrent major depression, atypical bipolar disorder, alcohol abuse or dependence, drug abuse or dependence, schizophrenia, schizophreniform, obsessive compulsive disorder, phobia, somatization, panic, antisocial personality, and anorexia nervosa. The DIS uses the Mini-Mental State Examination (MMSE), which measures cognitive functioning, as an indirect measure of the DSM-III Organic Mental Disorders. In the ECA survey, this diagnosis is called cognitive impairment.
This collection features data from 17,327 participants across 2,005 variables. Data from the Los Angeles, California, Catchment (UCLA) are not included. Baseline data (Wave 1) and Wave 2 data were linked to the National Death Index through 2007, which includes primary and contributing causes of death, International Classification of Disease (ICD) codes, and nature of injury variables.
The Community Vulnerability and Responses to Drug-User-Related HIV/AIDS, 1990-2013 [96 Metropolitan Statistical Areas, United States] (ICPSR 36575)
The Community Vulnerability and Responses to Drug-User-Related HIV/AIDS, 1990-2013 [96 Metropolitan Statistical Areas, United States] study (CVAR) was a research study of why large United States Metropolitan Statistical Areas (MSAs) vary over time in their vulnerability to HIV/AIDS among drug users and in MSA responses to HIV/AIDS. This collection contains estimates of HIV prevalence among people who injected drugs (PWID) and among sub-populations of PWID. This collection is comprised of ten datasets with differing amounts of variables and provides trend data that describe the following:
- Epidemiologic outcomes including population prevalence of PWIDs and Non-injecting drug users (NIDUs), and particularly their prevalence among youth; and, among PWIDs, HIV prevalence, late-diagnosis HIV cases, and AIDS incidence and mortality.
- Implementation of evidence-based drug-related interventions including drug abuse treatment, syringe exchange, HIV counseling and testing.
- Implementation of non-evidence-based drug-related interventions including incarceration and arrests of drug users.
The collection contains data on the MSA sub-populations including Black, Hispanic, White and "other" race categories. In addition, some statistics are presented in age range categories such as ages 15-29, 30-64 and 15-64.
Drug Use Among Young American Indians: Epidemiology and Prediction, 1993-2006 and 2009-2013 (ICPSR 35062)
The Drug Use Among Young Indians: Epidemiology and Prediction study is an annual surveillance effort assessing the levels and patterns of substance use among American Indian (AI) adolescents attending schools on or near reservations. In addition to annual epidemiology of substance use, data pertaining to the normative environment for adolescent substance use were also obtained. For this data collection data comes from annual in-school surveys completed between the years 1993 to 2006, and 2009 to 2013. Students completed the surveys at school during a specified class period. The dataset contains 534 variables for 26,451 students in grades 7 to 12.