CBS News/60 Minutes/Vanity Fair National Survey, November #2, 2012 (ICPSR 34689)
Chicago Male Drug Use and Health Survey (MSM Supplement), 2002-2003 (ICPSR 34303)
Civil Union Study 2000-2002, United States (ICPSR 31241)
Combined Generations Wave 1 and TransPop surveys, United States, 2016-2018 (ICPSR 38421)
This collection includes a combined dataset of the Generations study wave 1 (baseline) survey and the TransPop study transgender survey. The two studies have many overlapping variables, and they examined topics such as respondents' health outcomes and behaviors, experiences with discrimination, identity, and transition-related experiences. Data from these studies were merged to allow for analysis of the combined LGBT populations. This dataset has also been reweighted to be representative of these populations.
The complete Generations study data (baseline, wave 2, and wave 3 survey data) can be found under study number 37166, and the complete TransPop study data (transgender and cisgender survey data) can be found under study number 37938. For detailed information on the Generations and TransPop studies, including related publications, please refer to their respective DSDR/ICPSR study pages.
Community Health Center: Core Data Project, 2001-2002 [United States] (ICPSR 21520)
Culture-based Prediction of Adolescent HIV Risk (ICPSR 35922)
Eurobarometer 71.2: European Employment and Social Policy, Discrimination, Development Aid, and Air Transport Services, May-June 2009 (ICPSR 28183)
Family Life and Sexual Learning, 1976 (ICPSR 7755)
General Social Survey, 1972-2010 [Cumulative File] (ICPSR 31521)
General Social Survey, 1972-2012 [Cumulative File] (ICPSR 34802)
General Social Survey, 1972-2014 [Cumulative File] (ICPSR 36319)
General Social Survey, 1972-2016 [Cumulative File] (ICPSR 36797)
Generations: A Study of the Life and Health of LGB People in a Changing Society, United States, 2016-2019 (ICPSR 37166)
The Generations study is a five-year study designed to examine health and well-being across three generations of lesbians, gay men, and bisexuals (LGB). The study explored identity, stress, health outcomes, and health care and services utilization among LGBs in three generations of adults who came of age during different historical contexts. This collection includes baseline, wave 1, and wave 2 data collected as part of the Generations study.
The study aimed to assess whether younger cohorts of LGBs differed from older cohorts in how they viewed their LGB identity and experienced stress related to prejudice and everyday forms of discrimination, as well as whether patterns of resilience differed between different LGB cohorts. Additionally, the study sought to examine how differences in stress experience affected mental health and well-being, including depressive and anxiety symptoms, substance and alcohol use, suicide ideation and behavior, and how younger LGBs utilized LGB-oriented social and health services, relative to older cohorts.
In wave 2, respondents were re-interviewed approximately one year after completion of the baseline (wave 1) survey. Only respondents who participated in the original sample of participants were surveyed at wave 2 (i.e., the enhancement oversample was not included in the longitudinal design of this study).
In wave 3, respondents were re-interviewed approximately one year after the completion of the wave 2 survey.
Demographic variables collected as part of this study include questions related to age, education, race, ethnicity, sexual identity, gender identity, income, employment, and religiosity.
HIV Transmission Network Metastudy Project: An Archive of Data From Eight Network Studies, 1988--2001 (ICPSR 22140)
The purpose of this project was to establish a collection of datasets that could be used (1) to analyze the influence of partnership networks on the transmission of sexually transmitted and blood-borne infections, and (2) to examine the influence of study design on estimation of network properties and impacts. Eight studies contributed datasets to the collection.
They include:
- Colorado Springs Project 90, 1988-1992
- Bushwick [Brooklyn, NY] Social Factors and HIV Risk (SFHR) Study, 1991-1993
- Atlanta Urban Networks Project, 1996-1999
- Flagstaff Rural Network Study, 1996-1998
- Atlanta Antiretroviral Adherence Study, 1998-2001
- Houston Risk Networks Study, 1997-1998
- Baltimore SHIELD (Self-Help in Eliminating Life-Threatening Diseases), 1997-1999
- Manitoba Chlamydia Study, 1997-1998
Each study contains information on sexual, needle sharing, and/or social networks. Each dataset was harmonized to permit comparative analysis. Almost all of the studies were research projects funded by federal agency sources (e.g., United States Centers for Disease Control and the National Institutes of Health); one was funded by Canadian sources. These studies, all closed for further enrollment, provide a range of designs and study types as well as a range of transmitted diseases. This allows researchers to investigate the relative effect of personal behavior and network connections on the dynamics of disease transmission, and to explore the impact of sampling design on estimation of network properties. Respondents were asked questions about different test results such as HIV, chlamydia, syphilis and hepatitis. Demographic variables include race, ethnicity, marital status, age, and gender.
How Couples Meet and Stay Together (HCMST), Wave 1 2009, Wave 2 2010, Wave 3 2011, Wave 4 2013, Wave 5 2015, United States (ICPSR 30103)
How Couples Meet and Stay Together (HCMST) surveyed how Americans met their spouses and romantic partners, and compared traditional to non-traditional couples. This collection covers data that was gathered over five waves. During the first wave, respondents were asked about their relationship status, including the gender, ethnicity, and race of their current partner, as well as the level of education of their parents. They were also asked about their living arrangements with their partner, the country, state, and city the respondent and/or the respondent's partner resided in most from birth to age 16, and whether the couple attended the same high school/college/university, or grew up in the same town. Information was collected on the legal status of the relationship, the city/state where the partnership was legalized, and how many times the respondent had previously been married. Additionally, respondents were asked about how often they visited with relatives, which gender they were most attracted to, their earned income in 2008, and the length of their current relationship. Finally, respondents were asked to recall how, when, and where they met their partner, how their parents felt about their partner, and to describe the perceived quality of their relationship. The second wave followed up with respondents one year after Wave 1. Information was collected on respondents' changes, if any, in marital status, relationship status, living arrangements, and reasons for separation where applicable. The third wave followed up with respondents one year after the second wave, and collected information on respondents' relationships reported in the first two waves, again including any changes in the status of the relationship and reasons for separation. The fourth wave followed up with respondents two years after Wave 3. In addition to information on relationship status and reasons for separation, Wave 4 includes the subjective level of attractiveness for the respondent and their partner. Wave 5 collected updated data on respondents' changes, if any, in marital status, relationship status, and reasons for separation where applicable. Information about respondents' sexual orientations, sex frequencies, and attitudes towards sexual monogamy were also collected. Demographic information includes age, race/ethnicity, gender, level of education, household composition, religion, political party affiliation, and household income.
The data is being released in two parts: part one is available for public use and part two is available for restricted use. The public use data contains Waves 1-5, including the addition of nine variables collecting information such as race, household income, whether the respondent was born outside of the United States, zip code relative to rural area, and respondents' living arrangements between birth and 16 years of age. The restricted use data contains Waves 1-3, and differs from the public use data by including FIPS codes for state of marriage and state of residence, town or city where the respondent was raised, and qualitative variables revised by the Principal Investigator (Waves 1-5), consisting of respondent's answers to how they first met their partner, the quality of their relationship in their own words, why they broke up if applicable and if they have an open relationship.
Japanese General Social Survey (JGSS), 2003 (ICPSR 4242)
Japanese General Social Surveys (JGSS) Cumulative Data, 2000-2003 (ICPSR 4472)
National Transgender Discrimination Survey, [United States], 2008-2009 (ICPSR 37888)
This study brings to light what is both patently obvious and far too often dismissed from the human rights agenda. Transgender and gender non-conforming people face injustice at every turn: in childhood homes, in school systems that promise to shelter and educate, in harsh and exclusionary workplaces, at the grocery store, the hotel front desk, in doctors' offices and emergency rooms, before judges and at the hands of landlords, police officers, health care workers and other service providers.
The National Gay and Lesbian Task Force and the National Center for Transgender Equality are grateful to each of the 6,450 transgender and gender non-conforming study participants who took the time and energy to answer questions about the depth and breadth of injustice in their lives. A diverse set of people, from all 50 states, the District of Columbia, Puerto Rico, Guam and the U.S. Virgin Islands, completed online or paper surveys. This tremendous gift has created the first 360-degree picture of discrimination against transgender and gender non-conforming people in the U.S. and provides critical data points for policymakers, community activists and legal advocates to confront the appalling realities documented here and press the case for equity and justice.
These data provide information on discrimination in every major area of life, including housing, employment, health and health care, education, public accommodation, family life, criminal justice and government identity documents, and demographic information such as citizenship, race, ethnicity, employment, and income. In virtually every setting, the data underscores the urgent need for policymakers and community leaders to change their business-as-usual approach and confront the devastating consequences of anti-transgender bias.
Demographic information includes race, ethnicity, gender identity, sexual orientation, education, income, U.S citizenship, household size, and relationship status.
The public-use dataset was created in an earlier version of Stata that truncated write-in responses after 244 characters. The non-truncated write-in responses, plus Q10 zip codes and the essay responses to Q70, are included in the restricted-use dataset.
Population Assessment of Tobacco and Health (PATH) Study [United States] Public-Use Files (ICPSR 36498)
The Population Assessment of Tobacco and Health (PATH) Study began originally surveying 45,971 adult and youth respondents. The PATH Study was launched in 2011 to inform Food and Drug Administration's regulatory activities under the Family Smoking Prevention and Tobacco Control Act (TCA). The PATH Study is a collaboration between the National Institute on Drug Abuse (NIDA), National Institutes of Health (NIH), and the Center for Tobacco Products (CTP), Food and Drug Administration (FDA). The study sampled over 150,000 mailing addresses across the United States to create a national sample of people who use or do not use tobacco.
45,971 adults and youth constitute the first (baseline) wave of data collected by this longitudinal cohort study. These 45,971 adults and youth along with 7,207 "shadow youth" (youth ages 9 to 11 sampled at Wave 1) make up the 53,178 participants that constitute the Wave 1 Cohort. Respondents are asked to complete an interview at each follow-up wave. Youth who turn 18 by the current wave of data collection are considered "aged-up adults" and are invited to complete the Adult Interview. Additionally, "shadow youth" are considered "aged-up youth" upon turning 12 years old, when they are asked to complete an interview after parental consent.
At Wave 4, a probability sample of 14,098 adults, youth, and shadow youth ages 10 to 11 was selected from the civilian, noninstitutionalized population at the time of Wave 4. This sample was recruited from residential addresses not selected for Wave 1 in the same sampled Primary Sampling Units (PSUs) and segments using similar within-household sampling procedures. This "replenishment sample" was combined for estimation and analysis purposes with Wave 4 adult and youth respondents from the Wave 1 Cohort who were in the civilian, noninstitutionalized population at the time of Wave 4. This combined set of Wave 4 participants, 52,731 participants in total, forms the Wave 4 Cohort.
Dataset 0001 (DS0001) contains the data from the Master Linkage file. This file contains 14 variables and 67,276 cases. The file provides a master list of every person's unique identification number and what type of respondent they were for each wave.
At Wave 7, a probability sample of 14,863 adults, youth, and shadow youth ages 9 to 11 was selected from the civilian, noninstitutionalized population at the time of Wave 7. This sample was recruited from residential addresses not selected for Wave 1 or Wave 4 in the same sampled PSUs and segments using similar within-household sampling procedures. This second replenishment sample was combined for estimation and analysis purposes with Wave 7 adult and youth respondents from the Wave 4 Cohort who were at least age 15 and in the civilian, noninstitutionalized population at the time of Wave 7. This combined set of Wave 7 participants, 46,169 participants in total, forms the Wave 7 Cohort.
Please refer to the Public-Use Files User Guide that provides further details about children designated as "shadow youth" and the formation of the Wave 1, Wave 4, and Wave 7 Cohorts.
Dataset 1001 (DS1001) contains the data from the Wave 1 Adult Questionnaire. This data file contains 1,732 variables and 32,320 cases. Each of the cases represents a single, completed interview.
Dataset 1002 (DS1002) contains the data from the Youth and Parent Questionnaire. This file contains 1,228 variables and 13,651 cases.
Dataset 2001 (DS2001) contains the data from the Wave 2 Adult Questionnaire. This data file contains 2,197 variables and 28,362 cases. Of these cases, 26,447 also completed a Wave 1 Adult Questionnaire. The other 1,915 cases are "aged-up adults" having previously completed a Wave 1 Youth Questionnaire.
Dataset 2002 (DS2002) contains the data from the Wave 2 Youth and Parent Questionnaire. This data file contains 1,389 variables and 12,172 cases. Of these cases, 10,081 also completed a Wave 1 Youth Questionnaire. The other 2,091 cases are "aged-up youth" having previously been sampled as "shadow youth."
Dataset 3001 (DS3001) contains the data from the Wave 3 Adult Questionnaire. This data file contains 2,139 variables and 28,148 cases. Of these cases, 26,241 are continuing adults having completed a prior Adult Questionnaire. The other 1,907 cases are "aged-up adults" having previously completed a Youth Questionnaire.
Dataset 3002 (DS3002) contains the data from the Wave 3 Youth and Parent Questionnaire. This data file contains 1,309 variables and 11,814 cases. Of these cases, 9,769 are continuing youth having completed a prior Youth Interview. The other 2,045 cases are "aged-up youth" having previously been sampled as "shadow youth."
Datasets 3101, 3102, 3201 and 3202 (DS3101, DS3102, DS3201, and DS3202) are data files comprising the weight variables for Wave 3. The weight variables for Wave 1 and Wave 2 are included in the main data files. However, in Wave 3, the weight variables have been separated into individual data files for Adult and Youth Questionnaires. The "all-waves" weight files contain weights for those respondents who have completed an interview during all three waves of data collection. The "single-wave" weight files contain weights for all respondents in Wave 3 regardless of their participation in previous waves.
Dataset 3503 (DS3503) contains data derived from responses to Wave 1-3 questionnaires indicating if participants had ever/never used various tobacco products as of the Wave 3 study period. This data file contains 25 variables for all 53,178 study participants as of Wave 3. This file is provided for reference only to simplify the definitions of tobacco use variables in the Adult and Youth data files for subsequent waves.
Dataset 4001 (DS4001) contains the data from the Wave 4 Adult Questionnaire. This data file contains 2,182 variables and 33,822 cases. Of these cases, 25,857 are continuing adults having completed a prior Adult Questionnaire, 1,900 are "aged-up adults" having previously completed a Youth Questionnaire, and 6,065 are "replenishment sample adults" (also known as "new cohort adults" in the annotated instrument).
Dataset 4002 (DS4002) contains the data from the Wave 4 Youth and Parent Questionnaire. This data file contains 1,389 variables and 14,798 cases. Of these cases, 9,365 are continuing youth having completed a prior Youth Interview, 1,694 cases are "aged-up youth" having previously been sampled as "shadow youth," and 3,739 are "replenishment sample youth" (also known as "new cohort youth" in the annotated instrument).
Datasets 4111, 4112, 4211, 4212, 4321, and 4322 (DS4111, DS4112, DS4211, DS4212, DS4321, and DS4322) are data files comprising the weight variables for Wave 4. In Wave 4, the weight variables have been separated into individual data files corresponding to the Wave 1 and Wave 4 Cohorts and different weight types. The "all-waves" weight files contain weights for those Wave 1 Cohort respondents who completed an interview for all waves in which they were old enough or verified their information for waves in which they were not old enough to be interviewed. The "single-wave" weight files contain weights for Wave 1 Cohort respondents at Wave 4 who completed an interview at Wave 1, regardless of their participation in previous waves. The "cross-sectional" weight files contain weights for all respondents in the Wave 4 Cohort.
Dataset 4503 (DS4503) contains data derived from responses to Wave 1-4 questionnaires indicating if participants had ever/never used various tobacco products as of the Wave 4 data collection period. This data file contains 27 variables for all 67,276 study participants as of the Wave 4 data collection. This file is provided for reference only to simplify the definitions of tobacco use variables in the Adult and Youth data files for subsequent waves.
Dataset 5001 (DS5001) contains the data from the Wave 5 Adult Questionnaire. This data file contains 2,315 variables and 34,309 cases. Of these cases, 29,876 are continuing adults having completed a prior Adult Questionnaire, 4,433 are "aged-up adults" having previously completed a Youth Questionnaire.
Dataset 5002 (DS5002) contains the data from the Wave 5 Youth and Parent Questionnaire. This data file contains 1,530 variables and 12,098 cases. Of these cases, 10,446 are continuing youth having completed a prior Youth Interview, 1,652 cases are "aged-up youth" having previously been sampled as "shadow youth."
Datasets 5111, 5112, 5211, 5212, 5221, 5222, 5711, 5712, 5721, and 5722 (DS5111, DS5112, DS5211, DS5212, DS5221, DS5222, DS5711, DS5712, DS5721, and DS5722) are data files comprising the weight variables for Wave 5. In Wave 5, the weight variables are in individual data files corresponding to the Wave 1 and Wave 4 Cohorts and different weight types. The "all-waves" weight files contain weights for those Wave 1 Cohort participants who completed a Wave 5 interview and completed interviews (if old enough to do so) or verified their information (if not old enough to be interviewed) in Waves 1, 2, 3, and 4.
Dataset 5503 (DS5503) contains data derived from responses to Wave 1-5 (including Wave 4.5) questionnaires indicating if participants had ever/never used various tobacco products as of the Wave 5 data collection period. This data file contains 26 variables for all 67,276 study participants as of the Wave 5 data collection. This file is provided for reference only to simplify the definitions of tobacco use variables in the Adult and Youth data files for subsequent waves.
There are two separate sets of files with "single wave" weights: one for the Wave 1 Cohort and one for the Wave 4 Cohort. The "single-wave" weight files for the Wave 1 Cohort contain weights for participants who completed an interview in Wave 1 and in Wave 5, regardless of their participation in the intervening waves. The "single-wave" weight files for the Wave 4 Cohort contain weights for all Wave 5 interview respondents in the Wave 4 Cohort.
There are also two separate sets of files with "special collection all-waves" weights: one for the Wave 1 Cohort and one for the Wave 4 Cohort. The "special collection all-waves" weight files for the Wave 1 Cohort contains weights for participants who completed a Wave 5 interview and completed interviews (if old enough to do so) or verified their information (if not old enough to be interviewed) in Waves 1, 2, 3, 4, and the special collection in Wave 4.5. The "special collection all-waves" weight files for the Wave 4 Cohort contain weights for participants who completed a Wave 5 interview and completed interviews (if old enough to do so) or verified their information (if not old enough to be interviewed) in Wave 4 and the special collection in Wave 4.5.
Dataset 6001 (DS6001) contains the data from the Wave 6 Adult Questionnaire. This data file contains 2,589 variables and 30,516 cases. Of these cases, 28,852 are continuing adults having completed a prior Adult Questionnaire and 1,664 are "aged-up adults" having previously completed a Youth Questionnaire.
Dataset 6002 (DS6002) contains the data from the Wave 6 Youth and Parent Questionnaire. This data file contains 1,822 variables and 5,652 cases. Of these cases, 5,622 are continuing youth having completed a prior Youth interview and 30 cases are "aged-up youth" having previously been sampled as "shadow youth."
Datasets 6111, 6112, 6121, 6122, 6211, 6212, 6221, 6222, 6711, 6712, 6721, and 6722 (DS6111, DS6112, DS6121, DS6122, DS6211, DS6212, DS6221, DS6222, DS6711, DS6712, DS6721, and DS6722) are data files comprising the weight variables for Wave 6. In Wave 6, the weight variables are in individual data files corresponding to the Wave 1 and Wave 4 Cohorts and different weight types. There are two separate sets of files with "all-waves" weights: one for the Wave 1 Cohort and one for the Wave 4 Cohort. The "all-waves" weight files for the Wave 1 Cohort contain weights for participants who completed a Wave 6 interview and completed interviews (if old enough to do so) or verified their information (if not old enough to be interviewed) in Waves 1, 2, 3, 4, and 5. The "all-waves" weight files for the Wave 4 Cohort contain weights for participants who completed a Wave 6 interview and completed interviews (if old enough to do so) or verified their information (if not old enough to be interviewed) in Waves 4 and 5.
There are two separate sets of files with "single-wave" weights: one for the Wave 1 Cohort and one for the Wave 4 Cohort. The "single-wave" weight files for the Wave 1 Cohort contain weights for participants who completed an interview in Wave 1 and in Wave 6, regardless of their participation in the intervening waves. The "single-wave" weight files for the Wave 4 Cohort contain weights for participants who completed an interview in Wave 4 and in Wave 6, regardless of their participation in the intervening waves.
There are also two separate sets of files with "special collection all-waves" weights: one for the Wave 1 Cohort and one for the Wave 4 Cohort. The "special collection all-waves" weight files for the Wave 1 Cohort contain weights for participants who completed a Wave 6 interview and completed interviews (if old enough to do so) or verified their information (if not old enough to be interviewed) in Waves 1, 2, 3, 4, 5, and the special collections in Wave 4.5, and Wave 5.5 or PATH-ATS. The "special collection all-waves" weight files for the Wave 4 Cohort contain weights for participants who completed a Wave 6 interview and completed interviews (if old enough to do so) or verified their information (if not old enough to be interviewed) in Waves 4 and 5, and the special collections in Wave 4.5, and Wave 5.5 or PATH-ATS.
Dataset 6503 (DS6503) contains data derived from responses to Wave 1-6 (including Wave 4.5, Wave 5.5, and PATH-ATS) questionnaires indicating if participants had ever/never used various tobacco products as of the Wave 6 data collection period. This data file contains 24 variables for all 67,276 study participants as of the Wave 6 data collection. This file is provided for reference only to simplify the definitions of tobacco use variables in the Adult and Youth data files for subsequent waves.
Dataset 7001 (DS7001) contains the data from the Wave 7 Adult Questionnaire. This data file contains 2,813 variables and 30,801 cases. Of these cases, 27,258 are continuing adults having completed a prior Adult Questionnaire, 1,740 are "aged-up adults" having previously completed a Youth Questionnaire, and 1,803 are "replenishment sample adults" (also known as "new cohort adults" in the annotated instrument).
Dataset 7002 (DS7002) contains the data from the Wave 7 Youth and Parent Questionnaire. This data file contains 1,897 variables and 10,834 cases. Of these cases, 3,512 are continuing youth having completed a prior Youth Interview, 1 case is an "aged-up youth" having previously been sampled as "shadow youth," and 7,321 are "replenishment sample youth" (also known as "new cohort youth" in the annotated instrument).
Datasets 7111, 7112, 7121, 7122, 7211, 7212, 7221, 7222, 7331, 7332, 7711, 7712, 7721, and 7722 (DS DS7111, DS7112, DS7121, DS7122, DS7211, DS7212, DS7221, DS7222, DS7331, DS7332, DS7711, DS7712, DS7721, and DS7722) are data files comprising the weight variables for Wave 7. In Wave 7, the weight variables are in individual data files corresponding to the Wave 1, Wave 4, and Wave 7 Cohorts and different weight types.
There are two separate sets of files with "all-waves" weights: one for the Wave 1 Cohort and one for the Wave 4 Cohort. The "all-waves" weight files for the Wave 1 Cohort contain weights for participants who completed a Wave 7 interview and completed interviews (if old enough to do so) or verified their information (if not old enough to be interviewed) in Waves 1, 2, 3, 4, 5, and 6. The "all-waves" weight files for the Wave 4 Cohort contain weights for participants who completed a Wave 7 interview and completed interviews (if old enough to do so) or verified their information (if not old enough to be interviewed) in Waves 4, 5, and 6.
There are two separate sets of files with "single-wave" weights: one for the Wave 1 Cohort and one for the Wave 4 Cohort. The "single-wave" weight files for the Wave 1 Cohort contain weights for participants who completed an interview in Wave 1 and in Wave 7, regardless of their participation in the intervening waves. The "single-wave" weight files for the Wave 4 Cohort contain weights for participants who completed an interview in Wave 4 and in Wave 7, regardless of their participation in the intervening waves.
There are also two separate sets of files with "special collection all-waves" weights: one for the Wave 1 Cohort and one for the Wave 4 Cohort. The "special collection all-waves" weight files for the Wave 1 Cohort contain weights for participants who completed a Wave 7 interview and completed interviews (if old enough to do so) or verified their information (if not old enough to be interviewed) in Waves 1, 2, 3, 4, 5, 6, and the special collections in Wave 4.5, and Wave 5.5 or PATH-ATS. The "special collection all-waves" weight files for the Wave 4 Cohort contain weights for participants who completed a Wave 7 interview and completed interviews (if old enough to do so) or verified their information (if not old enough to be interviewed) in Waves 4, 5, 6, and the special collections in Wave 4.5, and Wave 5.5 or PATH-ATS.
The "cross-sectional" weight files contain weights for all respondents in the Wave 7 Cohort.
Dataset 8001 (DS8001) contains data from the Wave 8 Adult Questionnaire. This data file contains 3,467 variables and 31,477 cases. Of these cases, 30,021 are continuing adults having completed a prior Adult Questionnaire and 1,456 are "aged-up adults" having previously completed a Youth Questionnaire.
Dataset 8002 (DS8002) contains data from the Wave 8 Youth and Parent Questionnaire. This data file contains 2,393 variables and 8,002 cases. Of these cases, 7,046 are continuing youth having completed a prior Youth Interview and 956 are "aged-up youth" having previously been sampled as "shadow youth."
Datasets 8111, 8121, 8122, 8211, 8221, 8231, 8232, 8711, 8721, 8722, 8731, and 8732 (DS8111, DS8121, DS8122, DS8211, DS8221, DS8231, DS8232, DS8711, 8DS721, DS8722, DS8731, and DS8732) are data files comprising the weight variables for Wave 8. In Wave 8, the weight variables are in individual data files corresponding to the Wave 1, Wave 4, and Wave 7 Cohorts and different weight types.
There are two separate sets of files with "all-waves" weights: one for the Wave 1 Cohort and one for the Wave 4 Cohort. The "all-waves" weight files for the Wave 1 Cohort contain weights for participants who completed a Wave 8 interview and completed interviews (if old enough to do so) or verified their information (if not old enough to be interviewed) in Waves 1, 2, 3, 4, 5, 6, and 7. Note that only adults have "all-waves" weights for the Wave 1 Cohort; youth from the Wave 1 Cohort aged-up to adults by the time of Wave 8. The "all-waves" weight files for the Wave 4 Cohort contain weights for participants who completed a Wave 8 interview and completed interviews (if old enough to do so) or verified their information (if not old enough to be interviewed) in Waves 4, 5, 6, and 7.
There are three separate sets of files with "single-wave" weights: one for the Wave 1 Cohort, one for the Wave 4 Cohort, and one for the Wave 7 Cohort. The "single-wave" weight files for the Wave 1 Cohort contain weights for participants who completed an interview in Wave 1 and in Wave 8, regardless of their participation in the intervening waves. The "single-wave" weight files for the Wave 4 Cohort contain weights for participants who completed an interview in Wave 4 and in Wave 8, regardless of their participation in the intervening waves. Note that only adults have "single-wave" weights for the Wave 1 and Wave 4 Cohorts; youth from the Wave 1 Cohort aged-up to adults by the time of Wave 8 and youth from the Wave 4 Cohort were selected as shadow youth so they do not have any interview data from Wave 4. The "single wave" weights files for the Wave 7 Cohort contain weights for participants who completed an interview in Wave 7 and in Wave 8.
There are also three separate sets of files with "special collection all-waves" weights: one for the Wave 1 Cohort, one for the Wave 4 Cohort, and one for the Wave 7 Cohort. The "special collection all-waves" weight files for the Wave 1 Cohort contain weights for participants who completed a Wave 8 interview and completed interviews (if old enough to do so) or verified their information (if not old enough to be interviewed) in Waves 1, 2, 3, 4, 5, 6, 7 and the special collections in Wave 4.5, Wave 5.5, and Wave 7.5. Note that only adults have "special collection all-waves" weights for the Wave 1 Cohort; youth from the Wave 1 Cohort aged-up to adults by the time of Wave 8. The "special collection all-waves" weight files for the Wave 4 Cohort contain weights for participants who completed a Wave 8 interview and completed interviews (if old enough to do so) or verified their information (if not old enough to be interviewed) in Waves 4, 5, 6, 7, and the special collections in Wave 4.5, Wave 5.5, and Wave 7.5. The "special collection all-waves" weight files for the Wave 7 Cohort contain weights for participants who completed a Wave 8 interview and completed interviews (if old enough to do so) or verified their information (if not old enough to be interviewed) in Wave 7 and the special collection in Wave 7.5.
Each case in an Adult data file represents a single, completed interview. Each case in a Youth data file represents one youth and his or her parent's responses about that youth. Parents who provided permission for their child to participate in a Youth Interview were asked to complete a brief interview about their child. In all waves of data collection, less than 0.5 percent of the parents did not complete an interview. Most questions are asked about the child.
When multiple youth from the same household were selected to be in the study, the parent(s) completed separate interviews about each youth. If one parent completed two or more interviews, that parent only answered questions about himself/herself once. Those questions were then skipped in the subsequent interview(s) for the other child(ren) and the responses duplicated in that child(ren)'s data file(s).
Population Assessment of Tobacco and Health (PATH) Study [United States] Restricted-Use Files (ICPSR 36231)
The PATH Study was launched in 2011 to inform the Food and Drug Administration's regulatory activities under the Family Smoking Prevention and Tobacco Control Act (TCA). The PATH Study is a collaboration between the National Institute on Drug Abuse (NIDA), National Institutes of Health (NIH), and the Center for Tobacco Products (CTP), Food and Drug Administration (FDA). The study sampled over 150,000 mailing addresses across the United States to create a national sample of people who use or do not use tobacco.
45,971 adults and youth constitute the first (baseline) wave, Wave 1, of data collected by this longitudinal cohort study. These 45,971 adults and youth along with 7,207 "shadow youth" (youth ages 9 to 11 sampled at Wave 1) make up the 53,178 participants that constitute the Wave 1 Cohort. Respondents are asked to complete an interview at each follow-up wave. Youth who turn 18 by the current wave of data collection are considered "aged-up adults" and are invited to complete the Adult Interview. Additionally, "shadow youth" are considered "aged-up youth" upon turning 12 years old, when they are asked to complete an interview after parental consent.
At Wave 4, a probability sample of 14,098 adults, youth, and shadow youth ages 10 to 11 was selected from the civilian, noninstitutionalized population (CNP) at the time of Wave 4. This sample was recruited from residential addresses not selected for Wave 1 in the same sampled Primary Sampling Units (PSUs) and segments using similar within-household sampling procedures. This "replenishment sample" was combined for estimation and analysis purposes with Wave 4 adult and youth respondents from the Wave 1 Cohort who were in the CNP at the time of Wave 4. This combined set of Wave 4 participants, 52,731 participants in total, forms the Wave 4 Cohort.
At Wave 7, a probability sample of 14,863 adults, youth, and shadow youth ages 9 to 11 was selected from the CNP at the time of Wave 7. This sample was recruited from residential addresses not selected for Wave 1 or Wave 4 in the same sampled PSUs and segments using similar within-household sampling procedures. This "second replenishment sample" was combined for estimation and analysis purposes with the Wave 7 adult and youth respondents from the Wave 4 Cohort who were at least age 15 and in the CNP at the time of Wave 7. This combined set of Wave 7 participants, 46,169 participants in total, forms the Wave 7 Cohort.
Please refer to the Restricted-Use Files User Guide that provides further details about children designated as "shadow youth" and the formation of the Wave 1, Wave 4, and Wave 7 Cohorts.
Dataset 0002 (DS0002) contains the data from the State Design Data. This file contains 7 variables and 82,139 cases. The state identifier in the State Design file reflects the participant's state of residence at the time of selection and recruitment for the PATH Study.
Dataset 1011 (DS1011) contains the data from the Wave 1 Adult Questionnaire. This data file contains 2,021 variables and 32,320 cases. Each of the cases represents a single, completed interview.
Dataset 1012 (DS1012) contains the data from the Wave 1 Youth and Parent Questionnaire. This file contains 1,431 variables and 13,651 cases.
Dataset 1411 (DS1411) contains the Wave 1 State Identifier data for Adults and has 5 variables and 32,320 cases. Dataset 1412 (DS1412) contains the Wave 1 State Identifier data for Youth (and Parents) and has 5 variables and 13,651 cases. The same 5 variables are in each State Identifier dataset, including PERSONID for linking the State Identifier to the questionnaire and biomarker data and 3 variables designating the state (state Federal Information Processing System (FIPS), state abbreviation, and full name of the state). The State Identifier values in these datasets represent participants' state of residence at the time of Wave 1, which is also their state of residence at the time of recruitment.
Dataset 1611 (DS1611) contains the Tobacco Universal Product Code (UPC) data from Wave 1. This data file contains 32 variables and 8,601 cases. This file contains UPC values on the packages of tobacco products used or in the possession of adult respondents at the time of Wave 1. The UPC values can be used to identify and validate the specific products used by respondents and augment the analyses of the characteristics of tobacco products used by these respondents at the time of Wave 1.
Dataset 1801 (DS1801) contains Location Characteristics for Wave 1 Adults. This data file contains 4 variables and 32,320 cases.
Dataset 1802 (DS1802) contains Location Characteristics for Wave 1 Youth. This data file contains 4 variables and 13,651 cases.
Dataset 1901 (DS1901) contains Study Research Derived Variables for Wave 1 Adults created by PATH Study analysts. This data file contains 104 variables and 32,320 cases.
Dataset 1902 (DS1902) contains Study Research Derived Variables for Wave 1 Youth created by PATH Study analysts. This data file contains 89 variables and 13,651 cases.
Dataset 2011 (DS2011) contains the data from the Wave 2 Adult Questionnaire. This data file contains 2,421 variables and 28,362 cases. Of these cases, 26,447 also completed a Wave 1 Adult Questionnaire. The other 1,915 cases are "aged-up adults" having previously completed a Wave 1 Youth Questionnaire.
Dataset 2012 (DS2012) contains the data from the Wave 2 Youth and Parent Questionnaire. This data file contains 1,596 variables and 12,172 cases. Of these cases, 10,081 also completed a Wave 1 Youth Questionnaire. The other 2,091 cases are "aged-up youth" having previously been sampled as "shadow youth."
Dataset 2411 (DS2411) contains the Wave 2 State Identifier data for Adults and has 5 variables and 28,362 cases. Dataset 2412 (DS2412) contains the Wave 2 State Identifier data for Youth and Parents and has 5 variables and 12,172 cases. The same 5 variables are in each State Identifier dataset, including PERSONID for linking the State Identifier to the questionnaire and biomarker data and 3 variables designating the state (state FIPS, state abbreviation, and full name of the state). The State Identifier values in these datasets represent participants' state of residence at the time of Wave 2.
Dataset 2611 (DS2611) contains the Tobacco Universal Product Code (UPC) data from Wave 2. This data file contains 32 variables and 7,295 cases. This file contains UPC values on the packages of tobacco products used or in the possession of adult respondents at the time of Wave 2. The UPC values can be used to identify and validate the specific products used by respondents and augment the analyses of the characteristics of tobacco products used by these respondents at the time of Wave 2.
Dataset 2801 (DS2801) contains Location Characteristics for Wave 2 Adults. This data file contains 4 variables and 28,362 cases.
Dataset 2802 (DS2802) contains Location Characteristics for Wave 2 Youth. This data file contains 4 variables and 12,172 cases.
Dataset 2901 (DS2901) contains Study Research Derived Variables for Wave 2 Adults created by PATH Study analysts. This data file contains 178 variables and 28,362 cases.
Dataset 2902 (DS2902) contains Study Research Derived Variables for Wave 2 Youth created by PATH Study analysts. This data file contains 123 variables and 12,172 cases.
Dataset 3011 (DS3011) contains the data from the Wave 3 Adult Questionnaire. This data file contains 2,359 variables and 28,148 cases. Of these cases, 26,241 are continuing adults having completed a prior Adult Questionnaire. The other 1,907 cases are "aged-up adults" having previously completed a Youth Questionnaire.
Dataset 3012 (DS3012) contains the data from the Wave 3 Youth and Parent Questionnaire. This data file contains 1,492 variables and 11,814 cases. Of these cases, 9,769 are continuing youth having completed a prior Youth Interview. The other 2,045 cases are "aged-up youth" having previously been sampled as "shadow youth."
Datasets 3111, 3211, 3112, and 3212 (DS3111, DS3211, DS3112, and DS3212) are data files comprising the weight variables for Wave 3. The weight variables for Wave 1 and Wave 2 are included in the main data files. However, starting with Wave 3, the weight variables have been separated into individual data files. The "all-waves" weight files contain weights for respondents who completed an interview for all waves in which they were old enough to do so or verified their information with the study for waves in which they were not old enough to be interviewed. The "single-wave" weight files contain weights for all respondents in Wave 3 regardless of their participation in previous waves.
Dataset 3503 (DS3503) contains data derived from responses to Wave 1-3 questionnaires indicating if participants had ever/never used various tobacco products as of the Wave 3 study period. This data file contains 25 variables for all 53,178 study participants as of Wave 3. This file is provided for reference only to simplify the definitions of tobacco use variables in the Adult and Youth data files for subsequent waves.
Dataset 3411 (DS3411) contains the Wave 3 State Identifier data for Adults and has 5 variables and 28,148 cases. Dataset 3412 (DS3412) contains the Wave 3 State Identifier data for Youth and Parents and has 5 variables and 11,814 cases. The same 5 variables are in each State Identifier dataset, including PERSONID for linking the State Identifier to the questionnaire and biomarker data and 3 variables designating the state (state FIPS, state abbreviation, and full name of the state). The State Identifier values in these datasets represent participants' state of residence at the time of Wave 3.
Dataset 3611 (DS3611) contains the Tobacco Universal Product Code (UPC) data from Wave 3. This data file contains 32 variables and 6,768 cases. This file contains UPC values on the packages of tobacco products used or in the possession of adult respondents at the time of Wave 3. The UPC values can be used to identify and validate the specific products used by respondents and augment the analyses of the characteristics of tobacco products used by these respondents at the time of Wave 3.
Dataset 3801 (DS3801) contains Location Characteristics for Wave 3 Adults. This data file contains 4 variables and 28,148 cases.
Dataset 3802 (DS3802) contains Location Characteristics for Wave 3 Youth. This data file contains 4 variables and 11,814 cases.
Dataset 3901 (DS3901) contains Study Research Derived Variables for Wave 3 Adults created by PATH Study analysts. This data file contains 107 variables and 28,148 cases.
Dataset 3902 (DS3902) contains Study Research Derived Variables for Wave 3 Youth created by PATH Study analysts. This data file contains 88 variables and 11,814 cases.
Dataset 4001 (DS4001) contains the data from the Wave 4 Adult Questionnaire. This data file contains 2,504 variables and 33,822 cases. Of these cases, 25,857 are continuing adults having completed a prior Adult Questionnaire, 1,900 are "aged-up adults" having previously completed a Youth Questionnaire, and 6,065 are "replenishment sample adults" (also known as "new cohort adults" in the annotated instrument).
Dataset 4002 (DS4002) contains the data from the Wave 4 Youth and Parent Questionnaire. This data file contains 1,600 variables and 14,798 cases. Of these cases, 9,365 are continuing youth having completed a prior Youth Interview, 1,694 cases are "aged-up youth" having previously been sampled as "shadow youth," and 3,739 are "replenishment sample youth" (also known as "new cohort youth" in the annotated instrument).
Datasets 4111, 4211, 4321, 4112, 4212, and 4322 (DS4111, DS4211, DS4321, DS4112, DS4212, and DS4322) are data files comprising the weight variables for Wave 4. In Wave 4, the weight variables have been separated into individual data files corresponding to the Wave 1 and Wave 4 Cohorts and different weight types. The "all-waves" weight files contain weights for those Wave 1 Cohort respondents who completed an interview for all waves in which they were old enough or verified their information for waves in which they were not old enough to be interviewed. The "single-wave" weight files contain weights for Wave 1 Cohort respondents at Wave 4 who completed an interview at Wave 1, regardless of their participation in previous waves. The "cross-sectional" weight files contain weights for all respondents in the Wave 4 Cohort.
Dataset 4401 (DS4401) contains the Wave 4 State Identifier data for Adults and has 5 variables and 33,822 cases. Dataset 4402 (DS4402) contains the Wave 4 State Identifier data for Youth and Parents and has 5 variables and 14,798 cases. The same 5 variables are in each State Identifier dataset, including PERSONID for linking the State Identifier to the questionnaire and biomarker data and 3 variables designating the state (state FIPS, state abbreviation, and full name of the state). The State Identifier values in these datasets represent participants' state of residence at the time of Wave 4. For adults and youth from the replenishment sample, the values also represent state of residence at the time of recruitment.
Dataset 4503 (DS4503) contains data derived from responses to Wave 1-4 questionnaires, indicating if participants had ever/never used various tobacco products as of the Wave 4 data collection period. This data file contains 27 variables for all 67,276 study participants as of the Wave 4 data collection. This file is provided for reference only to simplify the definitions of tobacco use variables in the Adult and Youth data files for subsequent waves.
Dataset 4601 (DS4601) contains the Tobacco Universal Product Code (UPC) data from Wave 4. This data file contains 32 variables and 7,684 cases. This file contains UPC values on the packages of tobacco products used or in the possession of adult respondents at the time of Wave 4. The UPC values can be used to identify and validate the specific products used by respondents and augment the analyses of the characteristics of tobacco products used by these respondents at the time of Wave 4.
Dataset 4801 (DS4801) contains Location Characteristics for Wave 4 Adults. This data file contains 4 variables and 33,822 cases.
Dataset 4802 (DS4802) contains Location Characteristics for Wave 4 Youth. This data file contains 4 variables and 14,798 cases.
Dataset 5001 (DS5001) contains the data from the Wave 5 Adult Questionnaire. This data file contains 2,606 variables and 34,309 cases. Of these cases, 29,876 are continuing adults having completed a prior Adult Questionnaire and 4,433 are "aged-up adults" having previously completed a Youth Questionnaire.
Dataset 5002 (DS5002) contains the data from the Wave 5 Youth and Parent Questionnaire. This data file contains 1,776 variables and 12,098 cases. Of these cases, 10,446 are continuing youth having completed a prior Youth Interview and 1,652 cases are "aged-up youth" having previously been sampled as "shadow youth."
Datasets 5111, 5112, 5211, 5212, 5221, 5222, 5711, 5712, 5721, and 5722 (DS5111, DS5112, DS5211, DS5212, DS5221, DS5222, DS5711, DS5712, DS5721, and DS5722) are data files comprising the weight variables for Wave 5. In Wave 5, the weight variables are in individual data files corresponding to the Wave 1 and Wave 4 Cohorts and different weight types. The "all-waves" weight files contain weights for those Wave 1 Cohort participants who completed a Wave 5 interview and completed interviews (if old enough to do so) or verified their information (if not old enough to be interviewed) in Waves 1, 2, 3, and 4.
There are two separate sets of files with "single wave" weights: one for the Wave 1 Cohort and one for the Wave 4 Cohort. The "single-wave" weight files for the Wave 1 Cohort contain weights for participants who completed an interview in Wave 1 and in Wave 5, regardless of their participation in the intervening waves. The "single-wave" weight files for the Wave 4 Cohort contain weights for all Wave 5 interview respondents in the Wave 4 Cohort.
There are also two separate sets of files with "special collection all-waves" weights: one for the Wave 1 Cohort and one for the Wave 4 Cohort. The "special collection all-waves" weight files for the Wave 1 Cohort contain weights for participants who completed a Wave 5 interview and completed interviews (if old enough to do so) or verified their information (if not old enough to be interviewed) in Waves 1, 2, 3, 4, and the special collection in Wave 4.5. The "special collection all-waves" weight files for the Wave 4 Cohort contain weights for participants who completed a Wave 5 interview and completed interviews (if old enough to do so) or verified their information (if not old enough to be interviewed) in Wave 4 and the special collection in Wave 4.5.
Dataset 5401 (DS5401) contains the Wave 5 State Identifier data for Adults and has 5 variables and 34,309 cases. Dataset 5402 (DS5402) contains the Wave 5 State Identifier data for Youth and Parents and has 5 variables and 12,098 cases. The same 5 variables are in each State Identifier dataset, including PERSONID for linking the State Identifier to the questionnaire and biomarker data and 3 variables designating the state (state FIPS, state abbreviation, and full name of the state). The State Identifier values in these datasets represent participants' state of residence at the time of Wave 5.
Dataset 5503 (DS5503) contains data derived from responses to Wave 1-5 (including Wave 4.5) questionnaires indicating if participants had ever/never used various tobacco products as of the Wave 5 data collection period. This data file contains 26 variables for all 67,276 study participants as of the Wave 5 data collection. This file is provided for reference only to simplify the definitions of tobacco use variables in the Adult and Youth data files for subsequent waves.
Dataset 5601 (DS5601) contains the Tobacco Universal Product Code (UPC) data from Wave 5. This data file contains 33 variables and 6,678 cases. This file contains UPC values on the packages of tobacco products used or in the possession of adult respondents at the time of Wave 5. The UPC values can be used to identify and validate the specific products used by respondents and augment the analyses of the characteristics of tobacco products used by these respondents at the time of Wave 5.
Dataset 5801 (DS5801) contains Location Characteristics for Wave 5 Adults. This data file contains 4 variables and 34,309 cases.
Dataset 5802 (DS5802) contains Location Characteristics for Wave 5 Youth. This data file contains 4 variables and 12,098 cases.
Dataset 6001 (DS6001) contains the data from the Wave 6 Adult Questionnaire. This data file contains 2,935 variables and 30,516 cases
Of these cases, 28,852 are continuing adults having completed a prior Adult Questionnaire and 1,664 are "aged-up adults" having previously completed a Youth Questionnaire.
Dataset 6002 (DS6002) contains the data from the Wave 6 Youth and Parent Questionnaire. This data file contains 2,080 variables and 5,652 cases. Of these cases, 5,622 are continuing youth having completed a prior Youth Interview and 60 cases are "aged-up youth" having previously been sampled as "shadow youth."
Datasets 6111, 6112, 6121, 6122, 6211, 6212, 6221, 6222, 6711, 6712, 6721, and 6722 (DS6111, DS6112, DS6121, DS6122, DS6211, DS6212, DS62221, DS6222, DS6711, DS6712, DS6721, and DS6722) are data files comprising the weight variables for Wave 6. In Wave 6, the weight variables are in individual data files corresponding to the Wave 1 and Wave 4 Cohorts and different weight types. There are two separate sets of files with "all-waves" weights: one for the Wave 1 Cohort and one for the Wave 4 Cohort. The "all-waves" weight files for the Wave 1 Cohort contain weights for participants who completed a Wave 6 interview and completed interviews (if old enough to do so) or verified their information (if not old enough to be interviewed) in Waves 1, 2, 3, 4, and 5. The "all-waves" weight files for the Wave 4 Cohort contain weights for participants who completed a Wave 6 interview and completed interviews (if old enough to do so) or verified their information (if not old enough to be interviewed) in Waves 4 and 5.
There are two separate sets of files with "single-wave" weights: one for the Wave 1 Cohort and one for the Wave 4 Cohort. The "single-wave" weight files for the Wave 1 Cohort contain weights for participants who completed an interview in Wave 1 and in Wave 6, regardless of their participation in the intervening waves. The "single-wave" weight files for the Wave 4 Cohort contain weights for participants who completed an interview in Wave 4 and in Wave 6, regardless of their participation in the intervening waves.
There are also two separate sets of files with "special collection all-waves" weights: one for the Wave 1 Cohort and one for the Wave 4 Cohort. The "special collection all-waves" weight files for the Wave 1 Cohort contain weights for participants who completed a Wave 6 interview and completed interviews (if old enough to do so) or verified their information (if not old enough to be interviewed) in Waves 1, 2, 3, 4, 5, and the special collections in Wave 4.5, and Wave 5.5 or PATH-ATS. The "special collection all-waves" weight files for the Wave 4 Cohort contain weights for participants who completed a Wave 6 interview and completed interviews (if old enough to do so) or verified their information (if not old enough to be interviewed) in Waves 4 and 5, and the special collections in Wave 4.5, and Wave 5.5 or PATH-ATS.
Dataset 6401 (DS6401) contains the Wave 6 State Identifier data for Adults and has 5 variables and 30,516 cases. Dataset 6402 (DS6402) contains the Wave 6 State Identifier data for Youth and Parents and has 5 variables and 5,652 cases. The same 5 variables are in each State Identifier dataset, including PERSONID for linking the State Identifier to the questionnaire and biomarker data and 3 variables designating the state (state FIPS, state abbreviation, and full name of the state). The State Identifier values in these datasets represent participants' state of residence at the time of Wave 6.
Dataset 6503 (DS6503) contains data derived from responses to questionnaires in Waves 1-6 (including the special collections in Wave 4.5, Wave 5.5, and PATH-ATS) indicating if participants had ever/never used various tobacco products as of the Wave 6 data collection period. This data file contains 24 variables for all 67,276 study participants as of the Wave 6 data collection. This file is provided for reference only to simplify the definitions of tobacco use variables in the Adult and Youth data files for subsequent waves.
Dataset 6601 (DS6601) contains the Tobacco Universal Product Code (UPC) data from Wave 6. This data file contains 53 variables and 5,408 cases. This file contains UPC values on the packages of tobacco products used or in the possession of adult respondents at the time of Wave 6. The UPC values can be used to identify and validate the specific products used by respondents and augment the analyses of the characteristics of tobacco products used by these respondents at the time of Wave 6.
Dataset 6801 (DS6801) contains Location Characteristics for Wave 6 Adults. This data file contains 4 variables and 30,516 cases.
Dataset 6802 (DS6802) contains Location Characteristics for Wave 6 Youth. This data file contains 4 variables and 5,652 cases.
Dataset 7001 (DS7001) contains the data from the Wave 7 Adult Questionnaire. This data file contains 3,221 variables and 30,801 cases. Of these cases, 27,258 are continuing adults having completed a prior Adult Questionnaire, 1,740 are "aged-up adults" having previously completed a Youth Questionnaire, and 1,803 are "replenishment sample adults" (also known as "new cohort adults" in the annotated instrument).
Dataset 7002 (DS7002) contains the data from the Wave 7 Youth and Parent Questionnaire. This data file contains 2,171 variables and 10,834 cases. Of these cases, 3,512 are continuing youth having completed a prior Youth Interview, 1 case is an "aged-up youth" having previously been sampled as "shadow youth," and 7,321 are "replenishment sample youth" (also known as "new cohort youth" in the annotated instrument).
Datasets 7111, 7112, 7121, 7122, 7211, 7212, 7221, 7222, 7331, 7332, 7711, 7712, 7721, and 7722 (DS DS7111, DS7112, DS7121, DS7122, DS7211, DS7212, DS7221, DS7222, DS7331, DS7332, DS7711, DS7712, DS7721, and DS7722) are data files comprising the weight variables for Wave 7. In Wave 7, the weight variables are in individual data files corresponding to the Wave 1, Wave 4, and Wave 7 Cohorts and different weight types.
There are two separate sets of files with "all-waves" weights: one for the Wave 1 Cohort and one for the Wave 4 Cohort. The "all-waves" weight files for the Wave 1 Cohort contain weights for participants who completed a Wave 7 interview and completed interviews (if old enough to do so) or verified their information (if not old enough to be interviewed) in Waves 1, 2, 3, 4, 5, and 6. The "all-waves" weight files for the Wave 4 Cohort contain weights for participants who completed a Wave 7 interview and completed interviews (if old enough to do so) or verified their information (if not old enough to be interviewed) in Waves 4, 5, and 6.
There are two separate sets of files with "single-wave" weights: one for the Wave 1 Cohort and one for the Wave 4 Cohort. The "single-wave" weight files for the Wave 1 Cohort contain weights for participants who completed an interview in Wave 1 and in Wave 7, regardless of their participation in the intervening waves. The "single-wave" weight files for the Wave 4 Cohort contain weights for participants who completed an interview in Wave 4 and in Wave 7, regardless of their participation in the intervening waves.
There are also two separate sets of files with "special collection all-waves" weights: one for the Wave 1 Cohort and one for the Wave 4 Cohort. The "special collection all-waves" weight files for the Wave 1 Cohort contain weights for participants who completed a Wave 7 interview and completed interviews (if old enough to do so) or verified their information (if not old enough to be interviewed) in Waves 1, 2, 3, 4, 5, 6, and the special collections in Wave 4.5, and Wave 5.5 or PATH-ATS. The "special collection all-waves" weight files for the Wave 4 Cohort contain weights for participants who completed a Wave 7 interview and completed interviews (if old enough to do so) or verified their information (if not old enough to be interviewed) in Waves 4, 5, 6, and the special collections in Wave 4.5, and Wave 5.5 or PATH-ATS.
The "cross-sectional" weight files contain weights for all respondents in the Wave 7 Cohort.
Dataset 7401 (DS7401) contains the Wave 7 State Identifier data for Adults and has 5 variables and 30,801 cases. Dataset 7402 (DS7402) contains the Wave 7 State Identifier data for Youth and Parents and has 5 variables and 10,834 cases. The same 5 variables are in each State Identifier dataset, including PERSONID for linking the State Identifier to the questionnaire and biomarker data and 3 variables designating the state (state FIPS, state abbreviation, and full name of the state). The State Identifier values in these datasets represent participants' state of residence at the time of Wave 7.
Dataset 7503 (DS7503) contains data derived from responses to questionnaires in Waves 1-7 (including the special collections in Wave 4.5, Wave 5.5, and PATH-ATS) indicating if participants had ever/never used various tobacco products as of the Wave 7 data collection period. This data file contains 26 variables for all 82,139 study participants as of the Wave 7 data collection. This file is provided for reference only to simplify the definitions of tobacco use variables in the Adult and Youth data files for subsequent waves.
Dataset 7601 (DS7601) contains the Tobacco Universal Product Code (UPC) data from Wave 7. This data file contains 53 variables and 4,533 cases. This file contains UPC values on the packages of tobacco products used or in the possession of adult respondents at the time of Wave 7. The UPC values can be used to identify and validate the specific products used by respondents and augment the analyses of the characteristics of tobacco products used by these respondents at the time of Wave 7.
Dataset 7801 (DS7801) contains Location Characteristics for Wave 7 Adults. This data file contains 4 variables and 30,801 cases.
Dataset 7802 (DS7802) contains Location Characteristics for Wave 7 Youth. This data file contains 4 variables and 10,834 cases.
Dataset 8001 (DS8001) contains the data from the Wave 8 Adult Questionnaire. This data file contains 3,467 variables and 31,477 cases. Of these cases, 30,021 are continuing adults having completed a prior Adult Questionnaire and 1,456 are "aged-up adults" having previously completed a Youth Questionnaire.
Dataset 8002 (DS8002) contains the data from the Wave 8 Youth and Parent Questionnaire. This data file contains 2,393 variables and 8,002 cases. Of these cases, 7,046 are continuing youth having completed a prior Youth Interview and 956 are "aged-up youth" having previously been sampled as "shadow youth."
Datasets 8111, 8121, 8122, 8211, 8221, 8231, 8232, 8711, 8721, 8722, 8731, and 8732 (DS8111, DS8121, DS8122, DS8211, DS8221, DS8231, DS8232, DS8711, 8DS721, DS8722, DS8731, and DS8732) are data files comprising the weight variables for Wave 8. In Wave 8, the weight variables are in individual data files corresponding to the Wave 1, Wave 4, and Wave 7 Cohorts and different weight types.
There are two separate sets of files with "all-waves" weights: one for the Wave 1 Cohort and one for the Wave 4 Cohort. The "all-waves" weight files for the Wave 1 Cohort contain weights for participants who completed a Wave 8 interview and completed interviews (if old enough to do so) or verified their information (if not old enough to be interviewed) in Waves 1, 2, 3, 4, 5, 6, and 7. Note that only adults have "all-waves" weights for the Wave 1 Cohort; youth from the Wave 1 Cohort aged-up to adults by the time of Wave 8. The "all-waves" weight files for the Wave 4 Cohort contain weights for participants who completed a Wave 8 interview and completed interviews (if old enough to do so) or verified their information (if not old enough to be interviewed) in Waves 4, 5, 6, and 7.
There are three separate sets of files with "single-wave" weights: one for the Wave 1 Cohort, one for the Wave 4 Cohort, and one for the Wave 7 Cohort. The "single-wave" weight files for the Wave 1 Cohort contain weights for participants who completed an interview in Wave 1 and in Wave 8, regardless of their participation in the intervening waves. The "single-wave" weight files for the Wave 4 Cohort contain weights for participants who completed an interview in Wave 4 and in Wave 8, regardless of their participation in the intervening waves. Note that only adults have "single-wave" weights for the Wave 1 and Wave 4 Cohorts; youth from the Wave 1 Cohort aged-up to adults by the time of Wave 8 and youth from the Wave 4 Cohort were selected as shadow youth so they do not have any interview data from Wave 4. The "single wave" weights files for the Wave 7 Cohort contain weights for participants who completed an interview in Wave 7 and in Wave 8.
There are also three separate sets of files with "special collection all-waves" weights: one for the Wave 1 Cohort, one for the Wave 4 Cohort, and one for the Wave 7 Cohort. The "special collection all-waves" weight files for the Wave 1 Cohort contain weights for participants who completed a Wave 8 interview and completed interviews (if old enough to do so) or verified their information (if not old enough to be interviewed) in Waves 1, 2, 3, 4, 5, 6, 7 and the special collections in Wave 4.5, Wave 5.5, and Wave 7.5. Note that only adults have "special collection all-waves" weights for the Wave 1 Cohort; youth from the Wave 1 Cohort aged-up to adults by the time of Wave 8. The "special collection all-waves" weight files for the Wave 4 Cohort contain weights for participants who completed a Wave 8 interview and completed interviews (if old enough to do so) or verified their information (if not old enough to be interviewed) in Waves 4, 5, 6, 7, and the special collections in Wave 4.5, Wave 5.5, and Wave 7.5. The "special collection all-waves" weight files for the Wave 7 Cohort contain weights for participants who completed a Wave 8 interview and completed interviews (if old enough to do so) or verified their information (if not old enough to be interviewed) in Wave 7 and the special collection in Wave 7.5.
Dataset 8401 (DS8401) contains the Wave 8 State Identifier data for Adults and has 5 variables and 31,477 cases. Dataset 8402 (DS8402) contains the Wave 8 State Identifier data for Youth and Parents and has 5 variables and 8,002 cases. The same 5 variables are in each State Identifier dataset, including PERSONID for linking the State Identifier to the questionnaire and biomarker data and 3 variables designating the state (state FIPS, state abbreviation, and full name of the state). The State Identifier values in these datasets represent participants' state of residence at the time of Wave 8.
Dataset 8801 (DS8801) contains Location Characteristics for Wave 8 Adults. This data file contains 4 variables and 31,477 cases.
Dataset 8802 (DS8802) contains Location Characteristics for Wave 8 Youth. This data file contains 4 variables and 8,002 cases.
Each case in an Adult data file represents a single, completed interview. Each case in a Youth data file represents one youth and his or her parent's responses about that youth. Parents who provided permission for their child to participate in a Youth Interview were asked to complete a brief interview about their child. In all waves of data collection, less than 0.5 percent of the parents did not complete an interview. Most questions are asked about the child.
When multiple youth from the same household were selected to be in the study, the parent(s) completed separate interviews about each youth. If one parent completed two or more interviews, that parent only answered questions about himself/herself once. Those questions were then skipped in the subsequent interview(s) for the other child(ren) and the responses duplicated in that child(ren)'s data file(s).
Population Assessment of Tobacco and Health (PATH) Study [United States] Special Collection Public-Use Files (ICPSR 37786)
The PATH Study was launched in 2011 to inform the Food and Drug Administration's regulatory activities under the Family Smoking Prevention and Tobacco Control Act (TCA). The PATH Study is a collaboration between the National Institute on Drug Abuse (NIDA), National Institutes of Health (NIH), and the Center for Tobacco Products (CTP), Food and Drug Administration (FDA). The study sampled over 150,000 mailing addresses across the United States to create a national sample of people who do and do not use tobacco.
45,971 adults and youth constitute the first (baseline) wave, Wave 1, of data collected by this longitudinal cohort study. These 45,971 adults and youth along with 7,207 "shadow youth" (youth ages 9 to 11 sampled at Wave 1) make up the 53,178 participants that constitute the Wave 1 Cohort. Respondents are asked to complete an interview at each follow-up wave. Youth who turn 18 by the current wave of data collection are considered "aged-up adults" and are invited to complete the Adult Interview. Additionally, "shadow youth" are considered "aged-up youth" upon turning 12 years old, when they are asked to complete an interview after parental consent.
At Wave 4, a probability sample of 14,098 adults, youth, and shadow youth ages 10 to 11 was selected from the civilian, noninstitutionalized population (CNP) at the time of Wave 4. This sample was recruited from residential addresses not selected for Wave 1 in the same sampled Primary Sampling Units (PSUs) and segments using similar within-household sampling procedures. This "replenishment sample" was combined for estimation and analysis purposes with Wave 4 adult and youth respondents from the Wave 1 Cohort who were in the CNP at the time of Wave 4. This combined set of Wave 4 participants, 52,731 participants in total, forms the Wave 4 Cohort.
At Wave 7, a probability sample of 14,863 adults, youth, and shadow youth ages 9 to 11 was selected from the CNP at the time of Wave 7. This sample was recruited from residential addresses not selected for Wave 1 or Wave 4 in the same sampled PSUs and segments using similar within-household sampling procedures. This "second replenishment sample" was combined for estimation and analysis purposes with the Wave 7 adult and youth respondents from the Wave 4 Cohorts who were at least age 15 and in the CNP at the time of Wave 7. This combined set of Wave 7 participants, 46,169 participants in total, forms the Wave 7 Cohort.
Please refer to the Public-Use Files User Guide that provides further details about children designated as "shadow youth" and the formation of the Wave 1, Wave 4, and Wave 7 Cohorts.
Wave 4.5 was a special data collection for youth only who were aged 12 to 17 at the time of the Wave 4.5 interview. Wave 4.5 was the fourth annual follow-up wave for those who were members of the Wave 1 Cohort. For those who were sampled at Wave 4, Wave 4.5 was the first annual follow-up wave.
Wave 5.5, conducted in 2020, was a special data collection for Wave 4 Cohort youth and young adults ages 13 to 19 at the time of the Wave 5.5 interview. Also in 2020, a subsample of Wave 4 Cohort adults ages 20 and older were interviewed via the PATH Study Adult Telephone Survey (PATH-ATS).
Wave 7.5 was a special collection for Wave 4 and Wave 7 Cohort youth and young adults ages 12 to 22 at the time of the Wave 7.5 interview. For those who were sampled at Wave 7, Wave 7.5 was the first annual follow-up wave.
Dataset 1002 (DS1002) contains the data from the Wave 4.5 Youth and Parent Questionnaire. This file contains 1,395 variables and 13,131 cases. Of these cases, 11,378 are continuing youth having completed a prior Youth Interview. The other 1,753 cases are "aged-up youth" having previously been sampled as "shadow youth."
Datasets 1112, 1212, and 1222, (DS1112, DS1212, and DS1222) are data files comprising the weight variables for Wave 4.5. The "all-waves" weight file contains weights for participants in the Wave 1 Cohort who completed a Wave 4.5 Youth Interview and completed interviews (if old enough to do so) or verified their information with the study (if not old enough to be interviewed) in Waves 1, 2, 3, and 4.
There are two separate files with "single wave" weights: one for the Wave 1 Cohort and one for the Wave 4 Cohort. The "single-wave" weight file for the Wave 1 Cohort contains weights for youth who completed an interview in Wave 1 and in Wave 4.5, regardless of their participation in the intervening waves. The "single-wave" weight file for the Wave 4 Cohort contains weights for all Wave 4.5 Youth Interview respondents in the Wave 4 Cohort.
Dataset 1503 (DS1503) contains data derived from responses to questionnaires in Wave 1, Wave 2, Wave 3, Wave 4, and Wave 4.5 indicating if participants had ever/never used various tobacco products as of the Wave 4.5 data collection period. This data file contains 26 variables for all 67,276 study participants as of the Wave 4.5 data collection. This file is provided for reference only to simplify the definitions of tobacco use variables in the Adult and Youth data files for subsequent waves.
Dataset 2001 (DS2001) contains the data from the Wave 5.5 Adult Questionnaire. This file contains 2,323 variables and 3,628 cases. Of these cases, 1,014 are continuing adults having completed a prior Adult Questionnaire. The other 2,614 cases are "aged-up adults" having previously completed a Youth Questionnaire.
Dataset 2002 (DS2002) contains the data from the Wave 5.5 Youth and Parent Questionnaire. This file contains 1,625 variables and 7,129 cases. Of these cases, 7,076 are continuing youth having completed a prior Youth Interview. The other 53 cases are "aged-up youth" having previously been sampled as "shadow youth."
Datasets 2111, 2112, 2121, 2122, 2221, and 2222 (DS2111, DS2112, DS2121, DS2122, DS2221, and DS2222) are data files comprising the weight variables for Wave 5.5. In Wave 5.5, the weight variables are in individual data files corresponding to the Wave 1 and Wave 4 Cohorts and different weight types.
There are two separate sets of files with "all-waves" weights: one for the Wave 1 Cohort and one for the Wave 4 Cohort. The "all-waves" weight file for the Wave 1 Cohort contains weights for participants who completed a Wave 5.5 interview and completed interviews (if old enough to do so) or verified their information (if not old enough to be interviewed) in Waves 1, 2, 3, 4, 4.5, and 5. The "all-waves" weight file for the Wave 4 Cohort contains weights for participants who completed a Wave 5.5 interview and completed interviews (if old enough to do so) or verified their information (if not old enough to be interviewed) in Waves 4, 4.5, and 5.
The "single-wave" weight file for the Wave 4 Cohort contains weights for all Wave 5.5 interview respondents.
Dataset 3001 (DS3001) contains the data from PATH-ATS. This file contains 908 variables and 8,874 cases, all of which are continuing adults having completed a prior Adult Questionnaire, with their most recent interview in Wave 5.
Datasets 3111 and 3121 (DS3111 and DS3121) are data files comprising weights for PATH-ATS. In PATH-ATS, weight variables are in individual files corresponding to the Wave 1 and Wave 4 Cohorts.
The "all-waves" weight file for the Wave 1 Cohort contains weights for participants who completed an interview in PATH-ATS and completed interviews in Waves 1, 2, 3, 4, and 5. The "all-waves" weight file for the Wave 4 Cohort contains weights for participants who completed an interview in PATH-ATS; all PATH-ATS respondents completed interviews in Wave 4 and Wave 5.
Dataset 2503 (DS2503) contains data derived from responses to questionnaires in Wave 1, Wave 2, Wave 3, Wave 4, Wave 4.5, Wave 5, Wave 5.5, and PATH-ATS, indicating if participants had ever/never used various tobacco products as of the Wave 5.5/PATH-ATS data collection period. This data file contains 26 variables for all 67,276 study participants as of the Wave 5.5/PATH-ATS data collection. This file is provided for reference only to simplify the definitions of tobacco use variables in the Adult and Youth data files for subsequent waves.
Dataset 4001 (DS4001) contains the data from the Wave 7.5 Adult Questionnaire. This file contains 2,760 variables and 7,961 cases. Of these cases, 5,952 are continuing adults having completed a prior Adult Questionnaire. The other 2,009 cases are "aged-up adults" having previously completed a Youth Questionnaire.
Dataset 4002 (DS4002) contains the data from the Wave 7.5 Youth and Parent Questionnaire. This file contains 1,889 variables and 8,949 cases. Of these cases, 7,064 are continuing youth having completed a prior Youth Interview. The other 1,885 cases are "aged-up youth" having previously been sampled as "shadow youth."
Datasets 4111, 4112, 4121, 4122, 4221, 4222, 4231, and 4232 (DS4111, DS4112, DS4121, DS4122, DS4221, DS4222, DS4231, and DS4232) are data files comprising the weight variables for Wave 7.5. In Wave 7.5, the weight variables are in individual data files corresponding to the Wave 1, Wave 4, and Wave 7 Cohorts and different weight types.
There are two separate sets of files with "all-waves" weights: one for the Wave 1 Cohort and one for the Wave 4 Cohort. The "all-waves" weight file for the Wave 1 Cohort contains weights for participants who completed a Wave 7.5 interview and completed interviews (if old enough to do so) or verified their information (if not old enough to be interviewed) in Waves 1, 2, 3, 4, 4.5, 5, 5.5, 6, and 7. The "all-waves" weight file for the Wave 4 Cohort contains weights for participants who completed a Wave 7.5 interview and completed interviews (if old enough to do so) or verified their information (if not old enough to be interviewed) in Waves 4, 4.5, 5, 5.5, 6, and 7.
There are two separate sets of files with "single-waves" weights: one for the Wave 4 Cohort and one for the Wave 7 Cohort. The "single-wave" weight file for the Wave 4 Cohort contains weights for Wave 7.5 interview respondents in the Wave 4 Cohort, regardless of their response status at Waves 4.5, 5, 5.5, 6, or 7. The "single-wave" weight file for the Wave 7 Cohort contains weights for all Wave 7.5 interview respondents in the Wave 7 Cohort.
Dataset 4503 (DS4503) contains data derived from responses to questionnaires in Wave 1, Wave 2, Wave 3, Wave 4, Wave 4.5, Wave 5, Wave 5.5, PATH-ATS, Wave 6, Wave 7, and Wave 7.5, indicating if participants had ever/never used various tobacco products as of the Wave 7.5 data collection period. This data file contains 25 variables for all 82,139 study participants as of the Wave 7.5 data collection. This file is provided for reference only to simplify the definitions of tobacco use variables in the Adult and Youth data files for subsequent waves.
Population Assessment of Tobacco and Health (PATH) Study [United States] Special Collection Restricted-Use Files (ICPSR 37519)
The PATH Study was launched in 2011 to inform the Food and Drug Administration's regulatory activities under the Family Smoking Prevention and Tobacco Control Act (TCA). The PATH Study is a collaboration between the National Institute on Drug Abuse (NIDA), National Institutes of Health (NIH), and the Center for Tobacco Products (CTP), Food and Drug Administration (FDA). The study sampled over 150,000 mailing addresses across the United States to create a national sample of people who use or do not use tobacco.
45,971 adults and youth constitute the first (baseline) wave, Wave 1, of data collected by this longitudinal cohort study. These 45,971 adults and youth along with 7,207 "shadow youth" (youth ages 9 to 11 sampled at Wave 1) make up the 53,178 participants that constitute the Wave 1 Cohort. Respondents are asked to complete an interview at each follow-up wave. Youth who turn 18 by the current wave of data collection are considered "aged-up adults" and are invited to complete the Adult Interview. Additionally, "shadow youth" are considered "aged-up youth" upon turning 12 years old, when they are asked to complete an interview after parental consent.
At Wave 4, a probability sample of 14,098 adults, youth, and shadow youth ages 10 to 11 was selected from the civilian, noninstitutionalized population (CNP) at the time of Wave 4. This sample was recruited from residential addresses not selected for Wave 1 in the same sampled Primary Sampling Units (PSUs) and segments using similar within-household sampling procedures. This "replenishment sample" was combined for estimation and analysis purposes with Wave 4 adult and youth respondents from the Wave 1 Cohort who were in the CNP at the time of Wave 4. This combined set of Wave 4 participants, 52,731 participants in total, forms the Wave 4 Cohort.
At Wave 7, a probability sample of 14,863 adults, youth, and shadow youth ages 9 to 11 was selected from the CNP at the time of Wave 7. This sample was recruited from residential addresses not selected for Wave 1 or Wave 4 in the same sampled PSUs and segments using similar within-household sampling procedures. This "second replenishment sample" was combined for estimation and analysis purposes with the Wave 7 adult and youth respondents from the Wave 4 Cohorts who were at least age 15 and in the CNP at the time of Wave 7. This combined set of Wave 7 participants, 46,169 participants in total, forms the Wave 7 Cohort.
Please refer to the Restricted-Use Files User Guide that provides further details about children designated as "shadow youth" and the formation of the Wave 1, Wave 4, and Wave 7 Cohorts.
Wave 4.5 was a special data collection for youth only who were aged 12 to 17 at the time of the Wave 4.5 interview. Wave 4.5 was the fourth annual follow-up wave for those who were members of the Wave 1 Cohort. For those who were sampled at Wave 4, Wave 4.5 was the first annual follow-up wave.
Wave 5.5, conducted in 2020, was a special data collection for Wave 4 Cohort youth and young adults ages 13 to 19 at the time of the Wave 5.5 interview. Also in 2020, a subsample of Wave 4 Cohort adults ages 20 and older were interviewed via the PATH Study Adult Telephone Survey (PATH-ATS).
Wave 7.5 was a special collection for Wave 4 and Wave 7 Cohort youth and young adults ages 12 to 22 at the time of the Wave 7.5 interview. For those who were sampled at Wave 7, Wave 7.5 was the first annual follow-up wave.
Dataset 1002 (DS1002) contains the data from the Wave 4.5 Youth and Parent Questionnaire. This file contains 1,617 variables and 13,131 cases. Of these cases, 11,378 are continuing youth having completed a prior Youth Interview. The other 1,753 cases are "aged-up youth" having previously been sampled as "shadow youth"
Datasets 1112, 1212, and 1222, (DS1112, DS1212, and DS1222) are data files comprising the weight variables for Wave 4.5. The "all-waves" weight file contains weights for participants in the Wave 1 Cohort who completed a Wave 4.5 Youth Interview and completed interviews (if old enough to do so) or verified their information with the study (if not old enough to be interviewed) in Waves 1, 2, 3, and 4.
There are two separate files with "single wave" weights: one for the Wave 1 Cohort and one for the Wave 4 Cohort. The "single-wave" weight file for the Wave 1 Cohort contains weights for youth who completed an interview in Wave 1 and in Wave 4.5, regardless of their participation in the intervening waves. The "single-wave" weight file for the Wave 4 Cohort contains weights for all Wave 4.5 Youth Interview respondents in the Wave 4 Cohort.
Dataset 1402 (DS1402) contains the Wave 4.5 State Identifier data for Youth and Parents and has 5 variables and 13,131 cases. The State Identifier dataset includes PERSONID for linking the State Identifier to the questionnaire data and 3 variables designating the state (state Federal Information Processing System (FIPS), state abbreviation, and full name of the state). The State Identifier values in this dataset represent participants' state of residence at the time of Wave 4.5.
Dataset 1503 (DS1503) contains data derived from responses to questionnaires in Wave 1, Wave 2, Wave 3, Wave 4, and Wave 4.5 indicating if participants had ever/never used various tobacco products as of the Wave 4.5 data collection period. This data file contains 26 variables for all 67,276 study participants as of the Wave 4.5 data collection. This file is provided for reference only to simplify the definitions of tobacco use variables in the Adult and Youth data files for subsequent waves.
Dataset 2001 (DS2001) contains the data from the Wave 5.5 Adult Questionnaire. This file contains 2,619 variables and 3,628 cases. Of these cases, 1,014 are continuing adults having completed a prior Adult Questionnaire. The other 2,614 cases are "aged-up adults" having previously completed a Youth Questionnaire.
Dataset 2002 (DS2002) contains the data from the Wave 5.5 Youth and Parent Questionnaire. This file contains 1,871 variables and 7,129 cases. Of these cases, 7,076 are continuing youth having completed a prior Youth Interview. The other 53 cases are "aged-up youth" having previously been sampled as "shadow youth."
Datasets 2111, 2112, 2121, 2122, 2221, and 2222 (DS2111, DS2112, DS2121, DS2122, DS2221, and DS2222) are data files comprising the weight variables for Wave 5.5. In Wave 5.5, the weight variables are in individual data files corresponding to the Wave 1 and Wave 4 Cohorts and different weight types.
There are two separate sets of files with "all-waves" weights: one for the Wave 1 Cohort and one for the Wave 4 Cohort. The "all-waves" weight file for the Wave 1 Cohort contains weights for participants who completed a Wave 5.5 interview and completed interviews (if old enough to do so) or verified their information (if not old enough to be interviewed) in Waves 1, 2, 3, 4, 4.5, and 5. The "all-waves" weight file for the Wave 4 Cohort contains weights for participants who completed a Wave 5.5 interview and completed interviews (if old enough to do so) or verified their information (if not old enough to be interviewed) in Waves 4, 4.5 and 5.
The "single-wave" weight file for the Wave 4 Cohort contains weights for all Wave 5.5 interview respondents.
Dataset 2401 (DS2401) contains the Wave 5.5 State Identifier data for Adults and has 5 variables and 3,628 cases. Dataset 2402 (DS2402) contains the Wave 5.5 State Identifier data for Youth and Parents and has 5 variables and 7,129 cases. The same 5.5 variables are in each State Identifier dataset, including PERSONID for linking the State Identifier to the questionnaire and biomarker data and 3 variables designating the state (state FIPS, state abbreviation, and full name of the state). The State Identifier values in these datasets represent participants' state of residence at the time of Wave 5.5.
Dataset 2503 (DS2503) contains data derived from responses to questionnaires in Wave 1, Wave 2, Wave 3, Wave 4, Wave 4.5, Wave 5, and Wave 5.5 indicating if participants had ever/never used various tobacco products as of the Wave 5.5 data collection period. This data file contains 26 variables for all 67,276 study participants as of the Wave 5.5 data collection. This file is provided for reference only to simplify the definitions of tobacco use variables in the Adult and Youth data files for subsequent waves.
Dataset 3001 (DS3001) contains the data from PATH-ATS. This file contains 977 variables and 8,874 cases, all of which are continuing adults having completed a prior Adult Questionnaire, with their most recent interview in Wave 5.
Datasets 3111 and 3121 (DS3111 and DS3121) are data files comprising weights for PATH-ATS. In PATH-ATS, weight variables are in individual files corresponding to the Wave 1 and Wave 4 Cohorts.
The "all-waves" weight file for the Wave 1 Cohort contains weights for participants who completed an interview in PATH_-ATS and completed interviews in Waves 1, 2, 3, 4, and 5. The "all-waves" weight file for the Wave 4 Cohort contains weights for participants who completed an interview in PATH-ATS; all PATH-ATS respondents completed interviews in Wave 4 and Wave 5.
Dataset 3401 (DS3401) contains the PATH-ATS State Identifier data and has 5 variables and 8,874 cases. The State Identifier dataset includes PERSONID for linking the State Identifier to the questionnaire and biomarker data and 3 variables designating the state (state FIPS, state abbreviation, and full name of the state). The State Identifier values in this dataset represents participants' state of residence at the time of PATH-ATS.
Dataset 4001 (DS4001) contains the data from the Wave 7.5 Adult Questionnaire. This file contains 3,142 variables and 7,961 cases. Of these cases, 5,952 are continuing adults having completed a prior Adult Questionnaire. The other 2,009 cases are "aged-up adults" having previously completed a Youth Questionnaire.
Dataset 4002 (DS4002) contains the data from the Wave 7.5 Youth and Parent Questionnaire. This file contains 2,169 variables and 8,949 cases. Of these cases, 7,064 are continuing youth having completed a prior Youth Interview. The other 1,885 cases are "aged-up youth" having previously been sampled as "shadow youth."
Datasets 4111, 4112, 4121, 4122, 4221, 4222, 4231, and 4232 (DS4111, DS4112, DS4121, DS4122, DS4221, DS4222, DS4231, and DS4232) are data files comprising the weight variables for Wave 7.5. In Wave 7.5, the weight variables are in individual data files corresponding to the Wave 1, Wave 4, and Wave 7 Cohorts and different weight types.
There are two separate sets of files with "all-waves" weights: one for the Wave 1 Cohort and one for the Wave 4 Cohort. The "all-waves" weight file for the Wave 1 Cohort contains weights for participants who completed a Wave 7.5 interview and completed interviews (if old enough to do so) or verified their information (if not old enough to be interviewed) in Waves 1, 2, 3, 4, 4.5, 5, 5.5, 6, and 7. The "all-waves" weight file for the Wave 4 Cohort contains weights for participants who completed a Wave 7.5 interview and completed interviews (if old enough to do so) or verified their information (if not old enough to be interviewed) in Waves 4, 4.5, 5, 5.5, 6, and 7.
There are two separate sets of files with "single-waves" weights: one for the Wave 4 Cohort and one for the Wave 7 Cohort. The "single-wave" weight file for the Wave 4 Cohort contains weights for Wave 7.5 interview respondents in the Wave 4 Cohort, regardless of their response status at Waves 4.5, 5, 5.5, 6, or 7. The "single-wave" weight file for the Wave 7 Cohort contains weights for all Wave 7.5 interview respondents in the Wave 7 Cohort.
Dataset 4401 (DS4401) contains the Wave 7.5 State Identifier data for Adults and has 5 variables and 7,961 cases. Dataset 4402 (DS4402) contains the Wave 7.5 State Identifier data for Youth and Parents and has 5 variables and 8,949 cases. The same 7.5 variables are in each State Identifier dataset, including PERSONID for linking the State Identifier to the questionnaire and biomarker data and 3 variables designating the state (state FIPS, state abbreviation, and full name of the state). The State Identifier values in these datasets represent participants' state of residence at the time of Wave 7.5.
Dataset 4503 (DS4503) contains data derived from responses to questionnaires in Wave 1, Wave 2, Wave 3, Wave 4, Wave 4.5, Wave 5, Wave 5.5, PATH-ATS, Wave 6, Wave 7, and Wave 7.5 indicating if participants had ever/never used various tobacco products as of the Wave 7.5 data collection period. This data file contains 25 variables for all 82,139 study participants as of the Wave 7.5 data collection. This file is provided for reference only to simplify the definitions of tobacco use variables in the Adult and Youth data files for subsequent waves.
Dataset 4601 (DS4601) contains the Tobacco Universal Product Code (UPC) data from Wave 7.5. This data file contains 53 variables and 157 cases. This file contains UPC values on the packages of tobacco products used or in the possession of adult respondents at the time of Wave 7.5. The UPC values can be used to identify and validate the specific products used by respondents and augment the analyses of the characteristics of tobacco products used by these respondents at the time of Wave 7.5.
Project STRIDE: Stress, Identity, and Mental Health, New York City, 2004-2005 (ICPSR 35525)
Project STRIDE is a three-year research project that examines the effect of stress and minority identity related to sexual orientation, race/ethnicity and gender on mental health. The research describes social stressors that affect minority populations, explores the coping and social support resources that they utilize as they confront these social stressors, and assesses the associations of stress and coping with mental health outcomes including mental disorders and wellbeing. The study also explores the impact of various identity characteristics, such as whether an identity is viewed positively or negatively, or whether it is prominent or not to the relationship of stress and mental health outcomes.
The study, using extensive quantitative and some qualitative measures, is a longitudinal survey of 525 men and women between the ages 18 and 59 who are residents of New York City. Socio-demographic information collected about respondents included age, education, race and Hispanic ethnicity, adopting the measures developed and used by the United States Census Bureau in the United States population survey of 2000. In addition to these items, racial/ethnic identity was also assessed with the question "What is the country of origin related to your or your family's ethnic or national background, if any?" Respondents were allowed to select up to two nations from a comprehensive listing. For the purposes of the study, the instrument also assessed whether or not participants were natives of New York City or migrated as adults. Additional demographic variables include employment status, religion, relationship status, and sexual orientation.
The Science of BDSM Data, Phoenix, Arizona, 2014 (ICPSR 37395)
The goals of this study were to test whether participants who engaged in an extreme ritual in a naturalistic setting would evidence signs of altered states of consciousness, to examine other physiological and affective effects of the ritual, and to determine whether these effects varied based on the role the individual performed within the ritual. A multi-method approach was used that utilized various psychological self-report measures, a measure of cognitive functioning, and a measure of physiological stress. The data collection took place at the "Dance of Souls," a ritual conducted on the last day of the annual Southwest Leather Conference in Phoenix, Arizona, in which participants received temporary piercings with hooks or weights attached to the piercings and danced to music provided by drummers.
The associated publication, Altered States of Consciousness during an Extreme Ritual, was used to accompany the data in this collection. Users are encouraged to consult the publication for additional information. The data collection includes one de-identified dataset with 164 variables for 83 cases. Demographic variables include sex, gender, pierced vs. non-pierced, and the role the participant played in the ceremony.
Seek, Test, Treat and Retain Strategies Leveraging Mobile Health Technologies (Connect4Care), San Francisco, California, 2013-2015 (ICPSR 39783)
This study is part of the Seek, Test, Treat and Retain (STTR) Collaboration Project that involved over twenty studies in the fields of HIV and drug abuse. All studies were independently developed, but were chosen for the collaboration because they focused on one or more steps of the HIV treatment cascade: Seek, Test, Treat and Retain. As part of STTR Collaboration Project, the studies were grouped into Criminal Justice-related studies and Vulnerable Population-related studies. The data collected by these studies included twelve common domains (e.g., Demographic characteristics, Mental Health) in each of which a shared questionnaire or instrument was taken up by the studies and adapted to fit the study.
Connect4Care (C4C) was a single site, randomized year-long study of Short Message Service (SMS) primary care appointment reminders vs. SMS primary care appointment reminders plus thrice-weekly supportive, informational, and motivational SMS messages. Eligible consenting patients were allocated 1:1 to the two arms within strata defined by HIV diagnosis within the past 12 months (i.e. "newly diagnosed") vs. earlier.
Social Justice Sexuality Project: 2010 National Survey, including Puerto Rico (ICPSR 34363)
The Social Justice Sexuality Project (SJS) is one of the largest national surveys of Black, Latina/o, Asian and Pacific Islander, and multiracial lesbian, gay, bisexual, and transgender (LGBT) people. With over 5,000 respondents, the final sample includes respondents from all 50 states; Washington, DC, and Puerto Rico; in rural and suburban areas, in addition to large urban areas; and from a variety of ages, racial/ethnic identities, sexual orientations, and gender identities. The purpose of the SJS Project is to document and celebrate the experiences of lesbian, gay, bisexual and transgender (LGBT) people of color. All too often, when we think about LGBT people of color, it's from a perspective of pathology. In contrast, the SJS Project is designed and dedicated to describing a more dynamic experience. It's a knowledge-based study that investigates the sociopolitical experiences of this population around five themes: racial and sexual identity; spirituality and religion; mental and physical health; family formations and dynamics; civic and community engagement. Demographic variables include: race/ethnicity, sexual orientation, gender identity, age, education, religion, household, income, height, weight, location, birthplace, and political affiliation.
Additional information about the SJS Project can be found on the Social Justice Sexuality Project Web site.
Trends in Undiagnosed Chlamydia Prevalence in Baltimore, 1997-1998 and 2006-2009 (ICPSR 35064)
Tsogolo La Thanzi (TLT): Seventh Wave, Malawi, 2011 [Healthy Futures] (ICPSR 37831)
Tsogolo la Thanzi (TLT) is a longitudinal study in Balaka, Malawi designed to examine how young people navigate reproduction in an AIDS epidemic. Tsogolo la Thanzi (TLT) means "Healthy Futures" in Chichewa, Malawi's most widely spoken language. The TLT research team is collecting new data to develop better understandings of the reproductive goals and behavior of young adults in Malawi -- the first cohort to never have experienced life without AIDS. To understand these patterns of family formation in a rapidly changing setting, TLT used a unique approach: an intensive longitudinal design where respondents are interviewed every fourth months at TLT's centralized research center. Data collection began in May of 2009 and was completed in June of 2012. To assess changes on a longer time-horizon, a follow-up survey we refer to as TLT-2 was fielded between June and August of 2016.
This study contains data collected from the seventh wave of the multi-wave study.
Each of waves 1-8 are comprised of three data files. The Women dataset (dataset 1) is a random sample of women aged 15-25 in 2009 (N=1,505 at wave 1), drawn from a census of the area. Likewise, the Random Men dataset (dataset 3) is a random-sample of men aged 15-25 in 2009 (N=574 at wave 1) drawn from a census of the area. The Male Partners dataset (dataset 2) contains survey data from sexual and romantic partners who were referred into the study by respondents in the women's file; this is a non-random sample of male partners, so analysts should be especially cautious with inferences.
Topics covered across all waves include relationships, religion, HIV/AIDS, politics, family composition, mental health, sex and protection, pregnancy, marriage, sexually transmitted diseases, future expectations, school enrollment status, goods purchased/received, and diet.
Modules specific to wave 7 include: best friend characteristics, literacy, treatment optimism, travel, and health services with an expanded education section (interrupted education).
Additional demographic variables in each dataset include age and education.
Understanding Online Hate Speech as a Motivator and Predictor of Hate Crime, Los Angeles, California, 2017-2018 (ICPSR 37470)
In the United States, a number of challenges prevent an accurate assessment of the prevalence of hate crimes in different areas of the country. These challenges create huge gaps in knowledge about hate crime--who is targeted, how, and in what areas--which in turn hinder appropriate policy efforts and allocation of resources to the prevention of hate crime. In the absence of high-quality hate crime data, online platforms may provide information that can contribute to a more accurate estimate of the risk of hate crimes in certain places and against certain groups of people. Data on social media posts that use hate speech or internet search terms related to hate against specific groups has the potential to enhance and facilitate timely understanding of what is happening offline, outside of traditional monitoring (e.g., police crime reports). This study assessed the utility of Twitter data to illuminate the prevalence of hate crimes in the United States with the goals of (i) addressing the lack of reliable knowledge about hate crime prevalence in the U.S. by (ii) identifying and analyzing online hate speech and (iii) examining the links between the online hate speech and offline hate crimes.
The project drew on four types of data: recorded hate crime data, social media data, census data, and data on hate crime risk factors. An ecological framework and Poisson regression models were adopted to study the explicit link between hate speech online and hate crimes offline. Risk terrain modeling (RTM) was used to further assess the ability to identify places at higher risk of hate crimes offline.