Net Migration of the Population by Age, Sex, and Race, 1950-1970 (ICPSR 8493)
Net Migration of the Population of the United States by Age, Race and Sex, 1970-1980 (ICPSR 8697)
County-Specific Net Migration Estimates, 1980-1990 [United States] (ICPSR 26761)
This data collection represents a set of United States county net migration estimates by age and sex for the 1980-1990 decade, and is part of a series of estimates done for each decade since 1950 (1950-1970: see NET MIGRATION OF THE POPULATION BY AGE, SEX, AND RACE, 1950-1970 [ICPSR 8493]; 1970-1980: see NET MIGRATION OF THE POPULATION OF THE UNITED STATES BY AGE, RACE, AND SEX, 1970-1980 [ICPSR 8697]; 1990-2000: see COUNTY-SPECIFIC NET MIGRATION BY FIVE-YEAR AGE GROUPS, HISPANIC ORIGIN, RACE, AND SEX, 1990-2000 [ICPSR 4171]).
Net migration, the difference between the number of people moving into an area and the number moving out over a period, is measured here, and in all the other sets of estimates in the series, by the residual method. That is, net migration is equal to the population change over the period minus the natural increase (births -- deaths). Full details on how natural increase is estimated for each county, as well as other details of the data collection, are described in the codebook.
County-Specific Net Migration by Five-Year Age Groups, Hispanic Origin, Race and Sex: 2000-2010 (ICPSR 34638)
These data include county-level, net migration estimates by five-year age cohorts and sex, and by race and Hispanic origin, for the intercensal period from 2000 to 2010. The estimates were prepared using a vital statistics version of the forward cohort residual method. These estimates (and the net migration rates derivable from them) extend the set of decennial estimates of net migration that have been produced following each decennial census beginning with 1960 (net migration for the 1950s: Bowles and Tarver, 1965; 1960s: Bowles, Beale and Lee, 1975; 1970s: White, Mueser and Tierney, 1987; 1980s: Fuguitt, Beale, and Voss 2010; and 1990s: Voss, McNiven, Hammer, Johnson and Fuguitt, 2004).
Further information about this project is available on the Net Migration Patterns for US Counties Web site.
County-Specific Net Migration by Five-Year Age Groups, Hispanic Origin, Race, and Sex, 2010-2020: [United States] (ICPSR 39582)
County-Specific Net Migration by Five-Year Age Groups, Hispanic Origin, Race, and Sex, 1990-2000: [United States] (ICPSR 4171)
Population Redistribution and Economic Growth in the United States: Population Data, 1870-1960 (ICPSR 7753)
Population Estimates for States and Counties with Components of Change, 1981-1987 (ICPSR 9261)
Population Estimates by County with Components of Change, 1981-1985 (Provisional) (ICPSR 8613)
Migration Spillover: Spatial Panel Analysis in Europe (ICPSR 219921)
Federal-State Cooperative Program: 1975-1976 Population Estimates (ICPSR 7841)
Federal-State Cooperative Program: 1976-1977 Population Estimates (ICPSR 7842)
Federal-State Cooperative Program: 1977-1978 Population Estimates (ICPSR 7843)
International Data Base, World Population: 1983 Extract (ICPSR 8320)
Population Estimates of Counties in the United States, 1971-1974 (ICPSR 7500)
Population Estimates of Counties in the United States, 1973-1975 (ICPSR 7578)
Real Interest Rates and Population Growth across Generations (ICPSR 193943)
Projections of the Population of States by Age, Sex, and Race [United States]: 1988 to 2010 (ICPSR 9270)
Demographic Characteristics of the Population of the United States, 1930-1950: County-Level (ICPSR 20)
Puerto Rico Population Estimates (2020-2023) (ICPSR 248670)
Urban and Regional Migration Estimates (ICPSR 201260)
Data and Code for: The Likelihood of Persistently Low Global Fertility (ICPSR 239496)
Data and Code for: Labor Mobility and Unemployment over the Business Cycle (ICPSR 185901)
Replication Package for "Climate and Migration in the United States" (ICPSR 232122)
Great Plains Population and Environment Data: Social and Demographic Data, 1870-2000 [United States] (ICPSR 4296)
The social and demographic data included in this collection consist of a single data file for each decennial year between 1870 and 2000, covering 10 of the 12 Great Plains states. Information on a variety of social and demographic topics was gathered to historically characterize populations living in counties within the United States Great Plains, in terms of: (1) urban, rural, and total population, (2) vital statistics, (3) net migration, (4) age and sex, (5) nativity and ancestry, (6) education and literacy, (7) religion, (8) industry, and (9) housing and other characteristics. These data include selected material compiled as part of the United States population census. The United States Census of Population and Housing has been conducted since 1790 on a regular schedule that is decennial. The county-level social and demographic data produced by the United States government as a result constitute a consistent series of measures capturing changes in the United States population's size, composition, and other characteristics. A subset of the variables available from the short and long-form survey questionnaires of the United States Census of Population and Housing (as compiled for counties) were extracted from previously existing digital files. Besides the decennial census of the population, county-level data were drawn from an assortment of existing digital files as well as sources that were manually digitized. Other data include compilations of county-level information gathered from various federal agencies and private organizations as well as the agriculture and economic censuses. Supplementing these compilations are manually digitized consumer market data, religious data, and vital statistics, including information about births, deaths, marriage, and divorce.
Spatial Analysis of Crime in Appalachia [United States], 1977-1996 (ICPSR 3260)
Drug Offending in Cleveland, Ohio Neighborhoods, 1990-1997 and 1999-2001 (ICPSR 3929)
Baumol’s migrants:Productive and unproductive entrepreneurship and between-MSA migration (ICPSR 237784)
Replication data for: The Long-Run Effects of Labor Migration on Human Capital Formation in Communities of Origin (ICPSR 113660)
U.S. Inter-State Migration by Age (Radaris Data Sample, Anonymized) (ICPSR 251436)
- person_id — A random surrogate ID. Not derived from any real identifier and not reversible. A row key only — not a feature.
- age_group — Age band from year of birth: <25, 25-39, 40-54, 55-69, 70+.
- first_state — The person's earliest recorded state of residence (origin), as a 2-letter code.
- last_state — The person's current state of residence (destination), as a 2-letter code.
- Sampling. A uniform random sample of ~500,000 records was drawn from the full source database, so the sample's distributions reflect the source population.
- Endpoint extraction. Each source record carried a residential history. We reduced each history to its two endpoints — the earliest state and the current state — and dropped everything in between. (The source stored histories most-recent-first, so the origin is taken from the end of the sequence and the current location from the current-residence field.)
- Cleaning. Military postal codes (AA, AE, AP, used by APO/FPO/DPO overseas addresses rather than real states) were removed before extracting endpoints, so they never contaminate origin or destination.
- De-identification. All direct identifiers — names, source IDs, cities, and full address histories — were removed. Year of birth was generalized into five age bands. The original ID was replaced with a random surrogate.
- Re-identification control. The file enforces k-anonymity with k = 5 over the combination {age_group, first_state, last_state}: every published combination is shared by at least five people. The rare combinations that fell below this threshold (~0.6% of rows) were removed prior to release.
- geographic composition — how residents are distributed across states;
- age composition — the share of the population in each age band;
- interstate-mobility rate — the overall fraction of people who have crossed state lines, and how that fraction varies by age.
- The rare-flow tail is intentionally thinned. The k-anonymity step removed the least-common origin→destination pairs. So while common flows and overall rates are representative, the rarest corridors are under-represented by design. Do not treat tail frequencies as population estimates.
- It only speaks at the state level. Cities, neighborhoods, intermediate stops, and the timing of moves are not in this file. The dataset is representative of the source's state-level structure and says nothing below that resolution.
- Model whether a person has moved (moved = first_state != last_state) from age_group. This is a deliberately low-dimensional, interpretable problem — a good teaching or baseline example rather than a high-capacity modeling task.
- Build and analyze an origin→destination transition matrix: cluster states by their inflow/outflow profiles, rank net-gain vs net-loss states, visualize corridors as a flow map or chord diagram.
- Practice categorical/tabular workflows: contingency tables, chi-square tests of independence between age and mobility, proportion estimation with confidence intervals.
- Estimate the mover-vs-stayer rate by age band and test whether interstate mobility differs significantly across the life course.
- Quantify net migration per state (inflow − outflow), gross flows, and how these shift by age group.
- Validate against external sources — U.S. Census ACS migration tables and IRS county-to-county migration data — to benchmark or enrich the flows seen here.
- Is the classic finding that mobility declines with age visible here — is the <25/25-39 mover rate higher than the 55-69/70+ rate?
- Which states are the largest net receivers and net senders within each age band, and do retirement-age flows differ from early-career flows?
- Are there age-specific corridors — routes that dominate for the young but not the old, or vice versa (for example, Sun Belt destinations concentrated in older bands)?
- Do younger age groups show more geographic dispersion in their origins and destinations than older ones?
- State-level only; nothing below states is recoverable by design.
- Endpoints only; intermediate states, number of moves, and return migration are not represented.
- No dates; this is a cross-sectional snapshot, not a time series.
- "Stayer" means no interstate move was recorded, not necessarily no move at all.
- Sampling + suppression slightly thin the rarest flows.
- Download / mirrors: Kaggle, Hugging Face, Zenodo
- DOI: 10.5281/zenodo.21321002
- License: CC-BY-4.0
- Cite as: Zara Mann, 2026, U.S. Inter-State Migration by Age, v1.0, 10.5281/zenodo.21321002
Comparative Socio-Economic, Public Policy, and Political Data,1900-1960 (ICPSR 34)
Data and Code for: "The Welfare Magnet Hypothesis: Evidence from an Immigrant Welfare Scheme in Denmark" (ICPSR 118585)
Replication data for: Taxation and International Migration of Superstars: Evidence from the European Football Market (ICPSR 112661)
Malawi Longitudinal Study of Families and Health (MLSFH), 1998-2021 (ICPSR 20840)
The Malawi Longitudinal Study of Families and Health (MLSFH) is one of very few long-standing longitudinal cohort studies in a poor Sub-Saharan African (SSA) context. It provides a record of more than 25 years of demographic, socioeconomic, and health conditions in one of the world's poorest countries. Initial data collection began in 1998 under the Malawi Diffusion and Ideational Change Project (MDICP) to examine social networks and fertility decisions among married women and their husbands. While this initial study population is still followed, the scope of the project and population expanded to a broader focus on social and contextual determinants of health across the lifecourse in Malawi.
This collection includes Rounds 1 through 9 of the MLSFH, as well as supplemental data collections from Sexual Diaries, Migration Follow-Ups (MHM), a Biomarker Survey, Adverse Childhood Experiences (ACE), and a Benefits of Knowledge Intervention Survey. The MLSFH Data web page contains additional information and cohort profiles for all MLSFH data collections, including those not made available through ICPSR-DSDR.
COVID-19 High Frequency Phone Survey of Households, Ethiopia, 2020-2021 (ICPSR 38419)
The potential impacts of the COVID-19 pandemic in Ethiopia are expected to be severe on Ethiopian households' welfare. To monitor these impacts on households, the team selected a subsample of households that had been interviewed for the Living Standards Measurement Study (LSMS) in 2019, covering urban and rural areas in all regions of Ethiopia. The 15-minute questionnaire covers a series of topics, such as knowledge of COVID and mitigation measures, access to routine healthcare as public health systems are increasingly under stress, access to educational activities during school closures, employment dynamics, household income and livelihood, income loss and coping strategies, and external assistance.
The survey is implemented using Computer Assisted Telephone Interviewing, using a modular approach, which allows for modules to be dropped and/or added in different waves of the survey. Survey data collection started at the end of April 2020 and households are called back every three to four weeks for a total of seven survey rounds to track the impact of the pandemic as it unfolds and inform government action. This provides data to the government and development partners in near real-time, supporting an evidence-based response to the crisis.
The sample of households was drawn from the sample of households interviewed in the 2018/2019 round of the Ethiopia Socioeconomic Survey (ESS). The extensive information collected in the ESS, less than one year prior to the pandemic, provides a rich set of background information on the COVID-19 High Frequency Phone Survey of households which can be leveraged to assess the differential impacts of the pandemic in the country.