National Neighborhood Data Archive (NaNDA): Arts, Entertainment, and Leisure Establishments by Census Tract and ZCTA, United States, 1990-2022 (ICPSR 209163)
This dataset contains measures of the count and density of arts, entertainment, and leisure establishments per United States Census Tract or ZIP Code Tabulation Area (ZCTA) from 1990 through 2022. Business establishment data were drawn from the National Establishment Time Series (NETS) database and geocoded to 2010 and 2020 Census tract and ZCTA boundaries. The dataset includes four files — Census Tract 2010, Census Tract 2020, ZCTA 2010, and ZCTA 2020 — each containing one observation per geographic unit per year across ten establishment categories including museums, theaters, amusement parks, movie theaters, zoos and gardens, gambling facilities, bowling alleys, hotels, casino hotels, and an aggregate arts and entertainment total.
National Neighborhood Data Archive (NaNDA): Eating and Drinking Places by Census Tract and ZCTA, United States, 1990-2022 (ICPSR 208751)
This dataset provides annual measures of the number and density of eating and drinking places — including bars and night clubs, retail bakeries, coffee shops, fast food restaurants, delis, pizza restaurants, and sit-down restaurants — per census tract and ZIP Code Tabulation Area (ZCTA) across the United States from 1990 through 2022. Data are derived from the National Establishment Time Series (NETS) database and are available for four geographies: Census Tract 2010, Census Tract 2020, ZCTA 2010, and ZCTA 2020.
National Neighborhood Data Archive (NaNDA): Parks and Proximity to Polluting Sites by Census Tract and ZIP Code Tabulation Area (ZCTA), United States, 2024 (ICPSR 305511)
This dataset measures the number and area of parks in each U.S. census tract and ZIP Code Tabulation Area (ZCTA), as well as the spatial proximity of parks to two types of EPA-designated polluting sites: Toxics Release Inventory (TRI) facilities and Superfund sites. Park measures are derived from the 2024 ParkServe database (Trust for Public Land); polluting site measures use 2023 TRI data and 2024 Superfund Site data, with proximity calculated within park boundaries and at 0.5-, 1-, and 2-mile buffers. Geographic boundaries are drawn from the U.S. Census Bureau's 2020 TIGER/Line shapefiles.
National Neighborhood Data Archive (NaNDA): Retail Establishments by Census Tract and ZCTA, United States, 1990-2022 (ICPSR 208682)
This dataset contains measures of the number and density of retail establishments per United States Census Tract or ZIP Code Tabulation Area (ZCTA) from 1990 through 2022. Retail establishments are classified into eight categories based on Standard Industrial Classification (SIC) codes: clothing and shoe stores, furniture and appliance stores, music stores, hardware and garden stores, department/variety/general merchandise stores, used merchandise stores, pet stores and pet supplies, and shoe repair shops. The dataset is derived from the National Establishment Time Series (NETS) database and is available in four geographic versions: Census Tract 2010, Census Tract 2020, ZCTA 2010, and ZCTA 2020.
National Neighborhood Data Archive (NaNDA): Liquor, Tobacco, Cannabis, Vape, and Convenience Stores by Census Tract and ZCTA, United States, 1990-2022 (ICPSR 208907)
This dataset provides annual measures of the number and density of liquor, tobacco, cannabis, vape, and convenience stores per census tract and ZIP Code Tabulation Area (ZCTA) across the United States from 1990 through 2022. Data are derived from the National Establishment Time Series (NETS) database and are available for four geographies: Census Tract 2010, Census Tract 2020, ZCTA 2010, and ZCTA 2020.
National Neighborhood Data Archive (NaNDA): Recreational Establishments by Census Tract and ZCTA, United States, 1990-2022 (ICPSR 209164)
This dataset provides annual measures of the number and density of recreational services — including fitness centers, golf courses, skating rinks and pools, membership sports clubs, and specialized recreational establishments — per census tract and ZIP Code Tabulation Area (ZCTA) across the United States from 1990 through 2022. Data are derived from the National Establishment Time Series (NETS) database and are available for four geographies: Census Tract 2010, Census Tract 2020, ZCTA 2010, and ZCTA 2020.
National Neighborhood Data Archive (NaNDA): Personal Care Services and Laundry by Census Tract and ZCTA, United States, 1990-2022 (ICPSR 208906)
This dataset provides annual measures of the number and density of personal care services and laundry establishments — including barber shops, beauty shops, coin-operated laundromats, and laundry and dry cleaning services — per census tract and ZIP Code Tabulation Area (ZCTA) across the United States from 1990 through 2022. Data are derived from the National Establishment Time Series (NETS) database and are available for four geographies: Census Tract 2010, Census Tract 2020, ZCTA 2010, and ZCTA 2020.
National Neighborhood Data Archive (NaNDA): Grocery and Food Stores by Census Tract and ZCTA, United States, 1990-2022 (ICPSR 209313)
This dataset provides annual measures of the number and density of grocery and food stores — including grocery stores, supermarkets, meat and fish markets, fruit and vegetable markets, warehouse clubs selling food, and total food stores — per census tract and ZIP Code Tabulation Area (ZCTA) across the United States from 1990 through 2022. Data are derived from the National Establishment Time Series (NETS) database and are available for four geographies: Census Tract 2010, Census Tract 2020, ZCTA 2010, and ZCTA 2020.
National Neighborhood Data Archive (NaNDA): Law Enforcement by Census Tract and ZCTA, United States, 1990-2022 (ICPSR 208684)
This dataset measures the number and density of law enforcement organizations—including police departments, fire departments, courts, correctional facilities, and legal counsel and prosecution offices—across United States census tracts and ZIP Code Tabulation Areas (ZCTAs) from 1990 through 2022. Data are derived from the National Establishment Time Series (NETS) database and geocoded to 2010 and 2020 TIGER/Line shapefiles from the US Census Bureau.
National Neighborhood Data Archive (NaNDA): Civic, Social, and Religious Organizations by Census Tract and ZCTA, United States, 1990-2022 (ICPSR 207966)
This dataset contains measures of the number and density of select types of civic, social, and religious organizations per United States Census Tract or ZIP Code Tabulation Area (ZCTA) from 1990 through 2022.
National ZIP Code Crosswalk, [United States], 1990-2020 (ICPSR 39431)
ZIP Codes are administrative codes generated by the United States Postal Service (USPS) that refer to the geographic area covered by a specific set of mail delivery routes. The U.S. Census Bureau calculates and distributes aggregated social, economic, and demographic information for the population associated with "ZIP Code Tabulation Areas" (ZCTAs), which are roughly analogous to ZIP Codes and serve as identifiers for specific neighborhoods and communities. These aggregated census data, however, are unable to account for changes in ZIP Code boundaries that occur between decennial censuses, leading to measurement error and missing data problems for scholars who attempt to use the aggregated ZCTA data. The purpose of this crosswalk file is to allow researchers to overcome this limitation, enabling them to appropriately link spatial reference information (ZIP Codes) with characteristics of the populations to which they refer.
Most ZIP Codes do not change boundaries in a decade, but a large enough percentage do as to create a problem with missing or mis-specified data. Boundary changes typically involve one or more of the following three processes, although a small number of cases do not conform to these typologies: (1) two or more existing ZIP Codes are combined to create a single surviving ZIP Code, (2) an existing ZIP Code is divided into multiple resulting ZIP Codes, and (3) boundaries between two or more existing ZIP Codes are altered.
Each of these types of changes alters the geographic area that a ZIP Code refers to, and as such, the spatial unit identified by the ZIP Code includes a different population, with a different array of characteristics. By linking the spatial units associated with ZIP Codes as these boundary changes are enacted, the research team can both prevent the loss of observations due to missing data, and more accurately measure social, demographic, and economic characteristics associated with each ZIP Code.
This data set identifies changes in ZIP Code boundaries between 1990 and 2020, and provides numeric codes that cluster the ZIP Codes into the smallest geographic unit, or group of ZIP Codes, that are consistent across a decade: 1990 - 2000, 2000 - 2010, and 2010 - 2020. This "crosswalk" covers the contiguous United States, Alaska, Hawaii, and the District of Columbia. Since much administrative data is available with ZIP Code as the smallest identifiable geography, ZIP Codes are often used to embed observations from administrative data (patients, businesses, survey respondents, etc.) within their social, demographic, and economic contexts. However, ZIP Code boundaries change over time, resulting in measurement error (matching observations to the wrong contextual unit) or missing data (due to an observation reporting a ZIP Code that did not exist at the beginning of the observational period). These data were collected, and the crosswalk created, in an attempt to resolve these data quality issues.
National Neighborhood Data Archive (NaNDA): Hospitals by Census Tract and ZIP Code Tabulation Area, United States, 2023 (ICPSR 39378)
This dataset contains measures of the number and density of hospitals per United States Census Tract or ZIP Code Tabulation Area (ZCTA) in 2023. The dataset includes four separate files for four different geographic areas (GIS shapefiles from the United States Census Bureau). The four geographies include:
- Census Tract 2010
- Census Tract 2020
- ZIP Code Tabulation Area (ZCTA) 2010
- ZIP Code Tabulation Area (ZCTA) 2020
National ZIP Code Crosswalk (1990-2020) (ICPSR 194404)
Strategic Prevention Framework State Incentive Grant (SPF SIG) National Cross-Site Evaluation [Restricted Use] (ICPSR 28921)
Practice Patterns of Young Physicians, 1987: [United States] (ICPSR 9277)
This study investigated the factors that influenced the career decisions of young physicians and the characteristics of their practices. The collection has five datasets: Public-Use Version of the Young Physicians Survey (Dataset 1), Socioeconomic Monitoring System Study (Dataset 2), ZIP Code Data (Dataset 3), Verbatim Responses to the Open-Ended Questions (Dataset 4), and Restricted-Use Version of the Young Physicians Survey (Dataset 5).
The Public-Use Version of the Young Physicians Survey comprises responses from the Young Physicians Survey (YPS), plus merged data from the American Medical Association (AMA) Masterfile and the Association of American Medical Colleges' Student and Applicant Information Management System (SAIMS) database. The YPS interviewed physicians below 40 years of age who recently completed graduate medical training and were in their early years of practice. These physicians were queried about their graduate medical training, perceptions of the medical profession, current practice arrangements, career decisions, family background, patient care activities, and current income and expenses. To obtain information on current practice arrangements, respondents were questioned about the practices they worked in, including who owned the practices, the number of physicians in each practice, specialties or subspecialties practiced, usual fees for selected services, percentages of revenues from HMOs, PPOs, and IPAs, and percentages of patients who were Medicare patients, had no health insurance coverage, or were poor, Black, Hispanic, severely physically disabled, or chronically mentally ill. Questions on career decisions asked respondents about factors that influenced their career choices, such as reasons for working in multiple practices, reasons for leaving past practices, and reasons for deciding in favor of or against self-employment. Information on family background elicited by the survey includes the respondent's race, marital status, and educational debt, parents' income class and education, number of children living in the respondent's home, and whether the respondent's spouse or parents were physicians. Questions on patient care activities included questions on the number of hours spent providing uncompensated health care to the poor, and the number of hours spent with patients in a variety of settings, such as the office, emergency rooms, hospital outpatient clinics, and operating rooms. Information from the AMA Masterfile and the SAIMS database includes board certification status, AMA membership, school and year of graduation, Medical College Admission Test scores, primary undergraduate institution, most recent grade point averages, place of birth, number of acceptances to United States medical schools, parents' occupations, preferred medical specialty, and preferred practice setting.
Dataset 2 comprises responses from the AMA's Socioeconomic Monitoring System (SMS), a semiannual survey of nonfederal physicians that collected data on topics similar to those in the YPS, such as practice ownership, hours spent seeing patients in various settings, income, expenses, and opinions on practice procedures. The SMS data can be used for comparative analyses of young, prime, and senior physicians.
The ZIP Code Data contain estimates for the composition of the population residing in the ZIP code areas of the YPS respondents' main practices. This includes estimates of the size of each ZIP code area population, as well as its components with respect to gender, age, race, Hispanic ethnicity, and income. Also included are estimates of the number of physicians and their composition with respect to age, sex, practice type, and specialty.
Dataset 4 contains verbatim responses to open-ended questions asked in the YPS.
The Restricted-Use Version of the Young Physicians Survey is the same as the Public-Use Version of the Young Physicians Survey, except for some variables that were restricted from general dissemination for reasons of confidentiality. The restricted-use version includes the restricted variables, but the public-use version does not.