Search results

Search tips
Showing 1 – 50 of 3,056 results.
Curated
Simple Crosstabs

Aggregate Data, Regions of Russia (RoR), 1990-2010 (ICPSR 35355)

Released/updated on: 2014-10-14
Geographic coverage: Global, Russia
Time period: 1990-01-01--2010-12-31
The "Aggregate Data, Regions of Russia (RoR), 1990-2010" study is a collection of aggregate statistical data for the Russian regions, made available in English. It includes a large range of variables that characterize a wide scope of economic and social factors for the period from 1990 to 2010. This collection comprises data from 82 regions of Russia on topics including trade, production, demography, labor, investment, climate, crime, education, health care, culture, banks, insurance, services, communication, and many industries.
Curated

Aggregate Data Bank and Indices of Brazil: 1940-1960 (ICPSR 58)

Released/updated on: 1992-02-16
Geographic coverage: South America, Brazil, Global, Latin America
Time period: 1940-01-01--1960-12-31
This study contains data on the social, economic, and population characteristics of 22 states of Brazil in 1940, 1950, and 1960. For each of the three time periods, data are provided on the total population in urban and rural areas, industrial and commercial employment, and rural employment. Information is also provided on the literate population, eligible electorate, and actual voting electorate. The data ascertain the numbers of industrial and commercial establishments as well as membership in various unions, in art and literary associations, in sports organizations, in charitable organizations, and in Roman Catholic organizations.
Self-published

Lyme disease public use aggregated data with geography, 2022-2023 (ICPSR 242737)

Released/updated on: 2026-01-08
Overview: Public health surveillance data are collected and reported voluntarily to CDC by U.S. states and territories through the National Notifiable Diseases Surveillance System (NNDSS) (https://www.cdc.gov/nndss/index.html). Data include demographic, clinical, and geographic information; data do not include direct identifiers. Two types of datasets of human Lyme disease case data collected through public health surveillance are available: one includes annual case count aggregated by county of residence according to specific demographic variables and one is line-listed with patient demographic factors, month of illness onset, and clinical presentation information but without corresponding geographic information. These privacy-protected datasets were implemented in accordance with methodology described in Lee et al. Protecting Privacy and Transforming COVID-19 Case Surveillance Datasets for Public Use. Public Health Rep. 2021 Sep-Oct;136(5):554-561. doi: 10.1177/00333549211026817. Lyme disease became nationally notifiable in 1991. Different surveillance case definitions have been in effect over time; details are available here: https://ndc.services.cdc.gov/conditions/lyme-disease/. In 2008, a probable case definition was included in public health surveillance for the first time. In 2022, states with a high incidence of Lyme disease started reporting cases based on laboratory evidence alone without requirement for a clinical investigation, precluding comparison with historical data (for more information: https://www.cdc.gov/mmwr/volumes/73/wr/mm7306a1.htm?s_cid=mm7306a1_w). As such, Lyme disease surveillance data are grouped into separate datasets based on when these major changes occurred; data are provided for download separately for 1992–2007, 2008–2021, and 2022 to current. Data will be updated annually upon final verification of Lyme disease surveillance data by health departments. Data Limitations: Surveillance data have significant limitations that must be considered in the analysis, interpretation, and reporting of results. 1. Under-reporting and misclassification are features common to all surveillance systems. Not every case of Lyme disease is reported to CDC, and some cases that are reported may be reflect illness due to another cause. 2. Please note that before the 2022 surveillance case definition went into effect, several states with high Lyme disease incidence had initiated alternative methods of surveillance and those data were not reportable to CDC. 3. Final case data are subject to each state’s abilities to capture and classify cases, which is dependent upon budget and personnel. This can vary not only between states, but also from year to year within a given state. Consequently, a sudden or marked change in reported cases does not necessarily represent a true change in disease incidence. Every effort should be made to construct analyses to limit overinterpretation of this variation (see the following reference for more context: Kugeler KJ, Eisen RJ. Challenges in Predicting Lyme Disease Risk. JAMA Netw Open. 2020 Mar 2;3(3):e200328. doi: 10.1001/jamanetworkopen.2020.0328.) Read less
Self-published

Lyme disease public use aggregated data with geography, 2008-2021 (ICPSR 242738)

Released/updated on: 2026-01-08
Overview: Public health surveillance data are collected and reported voluntarily to CDC by U.S. states and territories through the National Notifiable Diseases Surveillance System (NNDSS) (https://www.cdc.gov/nndss/index.html). Data include demographic, clinical, and geographic information; data do not include direct identifiers. Two types of datasets of human Lyme disease case data collected through public health surveillance are available: one includes annual case count aggregated by county of residence according to specific demographic variables and one is line-listed with patient demographic factors, month of illness onset, and clinical presentation information but without corresponding geographic information. These privacy-protected datasets were implemented in accordance with methodology described in Lee et al. Protecting Privacy and Transforming COVID-19 Case Surveillance Datasets for Public Use. Public Health Rep. 2021 Sep-Oct;136(5):554-561. doi: 10.1177/00333549211026817. Lyme disease became nationally notifiable in 1991. Different surveillance case definitions have been in effect over time; details are available here: https://ndc.services.cdc.gov/conditions/lyme-disease/. In 2008, a probable case definition was included in public health surveillance for the first time. In 2022, states with a high incidence of Lyme disease started reporting cases based on laboratory evidence alone without requirement for a clinical investigation, precluding comparison with historical data (for more information: https://www.cdc.gov/mmwr/volumes/73/wr/mm7306a1.htm?s_cid=mm7306a1_w). As such, Lyme disease surveillance data are grouped into separate datasets based on when these major changes occurred; data are provided for download separately for 1992–2007, 2008–2021, and 2022 to current. Data will be updated annually upon final verification of Lyme disease surveillance data by health departments. Data Limitations: Surveillance data have significant limitations that must be considered in the analysis, interpretation, and reporting of results. 1. Under-reporting and misclassification are features common to all surveillance systems. Not every case of Lyme disease is reported to CDC, and some cases that are reported may be reflect illness due to another cause. 2. Please note that before the 2022 surveillance case definition went into effect, several states with high Lyme disease incidence had initiated alternative methods of surveillance and those data were not reportable to CDC. 3. Final case data are subject to each state’s abilities to capture and classify cases, which is dependent upon budget and personnel. This can vary not only between states, but also from year to year within a given state. Consequently, a sudden or marked change in reported cases does not necessarily represent a true change in disease incidence. Every effort should be made to construct analyses to limit overinterpretation of this variation (see the following reference for more context: Kugeler KJ, Eisen RJ. Challenges in Predicting Lyme Disease Risk. JAMA Netw Open. 2020 Mar 2;3(3):e200328. doi: 10.1001/jamanetworkopen.2020.0328.) Read less
Self-published

Lyme disease public use aggregated data with geography, 1992-2007 (ICPSR 242739)

Released/updated on: 2026-01-08
Overview: Public health surveillance data are collected and reported voluntarily to CDC by U.S. states and territories through the National Notifiable Diseases Surveillance System (NNDSS) (https://www.cdc.gov/nndss/index.html). Data include demographic, clinical, and geographic information; data do not include direct identifiers. Two types of datasets of human Lyme disease case data collected through public health surveillance are available: one includes annual case count aggregated by county of residence according to specific demographic variables and one is line-listed with patient demographic factors, month of illness onset, and clinical presentation information but without corresponding geographic information. These privacy-protected datasets were implemented in accordance with methodology described in Lee et al. Protecting Privacy and Transforming COVID-19 Case Surveillance Datasets for Public Use. Public Health Rep. 2021 Sep-Oct;136(5):554-561. doi: 10.1177/00333549211026817. Lyme disease became nationally notifiable in 1991. Different surveillance case definitions have been in effect over time; details are available here: https://ndc.services.cdc.gov/conditions/lyme-disease/. In 2008, a probable case definition was included in public health surveillance for the first time. In 2022, states with a high incidence of Lyme disease started reporting cases based on laboratory evidence alone without requirement for a clinical investigation, precluding comparison with historical data (for more information: https://www.cdc.gov/mmwr/volumes/73/wr/mm7306a1.htm?s_cid=mm7306a1_w). As such, Lyme disease surveillance data are grouped into separate datasets based on when these major changes occurred; data are provided for download separately for 1992–2007, 2008–2021, and 2022 to current. Data will be updated annually upon final verification of Lyme disease surveillance data by health departments. Data Limitations: Surveillance data have significant limitations that must be considered in the analysis, interpretation, and reporting of results. 1. Under-reporting and misclassification are features common to all surveillance systems. Not every case of Lyme disease is reported to CDC, and some cases that are reported may be reflect illness due to another cause. 2. Please note that before the 2022 surveillance case definition went into effect, several states with high Lyme disease incidence had initiated alternative methods of surveillance and those data were not reportable to CDC. 3. Final case data are subject to each state’s abilities to capture and classify cases, which is dependent upon budget and personnel. This can vary not only between states, but also from year to year within a given state. Consequently, a sudden or marked change in reported cases does not necessarily represent a true change in disease incidence. Every effort should be made to construct analyses to limit overinterpretation of this variation (see the following reference for more context: Kugeler KJ, Eisen RJ. Challenges in Predicting Lyme Disease Risk. JAMA Netw Open. 2020 Mar 2;3(3):e200328. doi: 10.1001/jamanetworkopen.2020.0328.) Read less
Curated

World Handbook of Political and Social Indicators II: Cross-National Aggregate Data, 1950-1965 (ICPSR 5027)

Released/updated on: 1992-02-16
Geographic coverage: Benin, Papua New Guinea, Angola, Cambodia, Sudan, Paraguay, Portugal, Syria, North Korea, Greece, Mongolia, Morocco, Iran, Mali, Panama, Guatemala, Guyana, Iraq, Chile, Laos, Nepal, Argentina, Tanzania, Zambia, Ghana, Belize, India, Canada, Turkey, Belgium, Taiwan, Finland, South Africa, Trinidad and Tobago, Central African Republic, Jamaica, Peru, Germany, Yemen, Vietnam (Socialist Republic), Puerto Rico, United States, Guinea, China (Peoples Republic), Chad, Somalia, Madagascar, Ivory Coast, Thailand, Libya, Costa Rica, Sweden, Malawi, Poland, Kuwait, Jordan, Nigeria, Bulgaria, Tunisia, Uruguay, Sri Lanka, Kenya, Switzerland, Spain, Lebanon, Liberia, Cuba, Venezuela, Czech Republic, Burkina Faso, Mauritania, Israel, Australia, Soviet Union, Myanmar, Cameroon, Cyprus, Malaysia, Iceland, Global, Gabon, South Korea, Great Britain, Austria, Yugoslavia, Mozambique, El Salvador, Luxembourg, Brazil, Algeria, Ecuador, Colombia, Hungary, Japan, Mauritius, Albania, New Zealand, Senegal, Italy, Honduras, Ethiopia, Haiti, Afghanistan, Burundi, Singapore, Egypt, Sierra Leone, Bolivia, Malta, Saudi Arabia, Netherlands, Pakistan, Ireland, France, Romania, Togo, Niger, Philippines, Rwanda, Nicaragua, Barbados, Norway, Democratic Republic of Congo, Botswana, Denmark, Dominican Republic, Mexico, Uganda, Zimbabwe, Indonesia
Time period: 1950-01-01--1965-12-31
This data collection consists of aggregate political, economic, and educational data for 136 countries in the period 1950-1965. Included are indicators of population size and growth, communications, education, culture, economics, and politics for the four base years: 1950, 1960, 1966, and 1965. Data are provided for the percentage of population living in cities of 100,000 or more and 20,000 or more, the total economically active male population engaged in agricultural occupations, and the total economically active male population as a percentage of the total male population. Information is also provided for the number of telephones, radios, televisions, and newspapers per 1,000 population, cinema attendance per capita, literacy rates, and school enrollment ratio. Other variables provide information for steel consumption, energy consumption per capita growth rates, gross national product (GNP) per capita, total trade as a percentage of the GNP, total number of current scientific and technical serials published, percentage of contribution to the total world scientific authors, percentage of gross domestic product (GDP) originating in agriculture, industry, transportation, and communications, and gross domestic fixed capital formation as a percentage of the GNP. Additional information is provided on sectorial income inequality, land inequality, total number of physicians, and number of physicians per one million population. Other items include total military manpower, defense, education, and health expenditure in million United States dollars, total United States economic and military aid and Soviet aid, number of memberships in United Nations organizations and in other international organizations, diplomatic representation, electoral irregularity score, press freedom index, total internal security forces, the beginning and ending year of modernization, the date of independence, and the date of founding of the present constitution.
Curated

Aggregate Economic Data, United States, 1947-1989 (ICPSR 1093)

Released/updated on: 1996-01-03
Geographic coverage: United States
These data and/or computer programs are part of ICPSR's Publication-Related Archive and are distributed exactly as they arrived from the data depositor. ICPSR has not checked or processed this material. Users should consult the INVESTIGATOR(S) if further information is desired.
Self-published

Data and Code for: Using aggregate relational data to feasibly identify network structure without network data (ICPSR 110841)

Released/updated on: 2022-01-06
Social network data is often prohibitively expensive to collect, limiting empirical network research. We propose an inexpensive and feasible strategy for network elicitation using Aggregated Relational Data (ARD) - responses to questions of the form "how many of your links have trait $k$?" Our method uses ARD to recover parameters of a network formation model, which permits sampling from a distribution over node- or graph-level statistics. We replicate the results of two field experiments that used network data and draw similar conclusions with ARD alone.
Self-published

How are SNAP Benefits Spent? Aggregated Replication Data from a Retail Panel (ICPSR 121689)

Released/updated on: 2020-09-15
This archive contains aggregated retail panel data based on the data described in Hastings and Shapiro (2018). The archive also contains code to produce plots analogous to those in figures 4 and 5 in Hastingsand Shapiro (2018).Hastings, Justine and Jesse M. Shapiro. 2018. How are SNAP benefits spent? Evidence from a retail panel. American Economic Review 108(12): 3493-3540. 
Self-published

Replication data for: Aggregate Recruiting Intensity (ICPSR 113158)

Released/updated on: 2019-10-12
We develop an equilibrium model of firm dynamics with random search in the labor market where hiring firms exert recruiting effort by spending resources to fill vacancies faster. Consistent with microevidence, fast-growing firms invest more in recruiting activities and achieve higher job-filling rates. These hiring decisions of firms aggregate into an index of economy-wide recruiting intensity. We study how aggregate shocks transmit to recruiting intensity, and whether this channel can account for the dynamics of aggregate matching efficiency during the Great Recession. Productivity and financial shocks lead to sizable procyclical fluctuations in matching efficiency through recruiting effort. Quantitatively, the main mechanism is that firms attain their employment targets by adjusting their recruiting effort in response to movements in labor market slackness.
Self-published

Replication data for: Aggregation and the Gravity Equation (ICPSR 116453)

Released/updated on: 2019-12-07
One of the most successful empirical relationships in international trade is the gravity equation. A key decision for researchers in estimating this relationship is the level of aggregation, since the gravity equation is log linear, whereas aggregation involves summing the level rather than the log of trade. In this paper, we derive a Jensen's inequality correction term for nested constant elasticity of substitution preferences, such that a log-linear gravity equation holds exactly for each nest. We provide evidence that sectoral composition is quantitatively relevant for the aggregate effect of distance on international trade, particularly for more disaggregated definitions of sectors.
Self-published

The Effect of SNAP on the Composition of Purchase Foods: Aggregated Replication Data from a Retail Panel (ICPSR 121688)

Released/updated on: 2020-09-15
This archive contains aggregated data based on the data described in Hastings et al. (Forthcoming). The archive also contains code to produce analogues of select plots and estimates reported in Hastings et al. (Forthcoming).
Hastings, Justine, Ryan Kessler, and Jesse M. Shapiro. Forthcoming. The effect of SNAP on the composition of purchased foods: Evidence and implications. American Economic Journal: Economic Policy
Self-published

NACP Regional: National greenhouse gas inventories and aggregated gridded model data (ICPSR 250100)

Released/updated on: 2026-06-21
Time period: 2000-01-01--2007-12-01
This data set provides two products that were derived from the recently published North American Carbon Program (NACP) Regional Synthesis 1-degree terrestrial biosphere model (TBM) and inverse model (IM) outputs (Gridded 1-deg Observation Data and Biosphere and Inverse Model Outputs, Wei et al., 2013). The first product is the aggregation of the standardized gridded 1-degree TBM and IM outputs to the Greenhouse Gas (GHG) inventory zones as defined for North America (United States, Canada, and Mexico). Depending on the data availability, the monthly/yearly Net Ecosystem Exchange (NEE), Net Primary Production (NPP), Total Vegetation Carbon (VegC), Heterotrophic Respiration (Rh), and Fire Emissions (FE) outputs from the 22 TBM and 7 IM models were aggregated from the 1-degree resolution gridded format to the inventory zones and then, further divided into Forest Lands, Crop Lands, and Other Lands sectors within each inventory zone based on the 1-kilometer (km) resolution GLC2000 land cover map (GLC2000, 2003). The second product is the North American national GHG inventories on the scale of inventory zones which contain estimated land-atmosphere exchange of CO2 (NEE) in forest lands, crop lands, and other lands sectors. NEE estimates were synthesized from inventory-based data on productivity, ecosystem carbon stock change, and harvested product stock change, and additional information from national-level GHG inventories of the United States, Canada, and Mexico including EPA (2011) and Environment Canada (2011). An additional summary file of annual mean NEE (2000-2006) is provided for both land sectors and reporting zones in North America and was created by combining the aggregated model output and the national GHG database and is provided. The aggregated monthly and yearly model output data and the national GHG inventories data are available in comma separated value (*.csv) format files. Also provided are detailed inventory zone spatial data as an ESRI Shapefile. Included are zone names, boundaries, and zone and land cover type area attributes. For mapping convenience, the inventory zones shapefile was merged with 1-km forest, crop, and other lands masks to create a 1-km resolution reference data file that was converted to GeoTIFF format. The GeoTIFF defines to which inventory zone and land cover type each 1-km grid cell belongs.
Self-published

Replication data for: Markups, Aggregation, and Inventory Adjustment (ICPSR 116033)

Released/updated on: 2019-12-06
In this paper I suggest a unified explanation for two puzzles in the inventory literature: first, estimates of inventory speeds of adjustment in aggregate data are very small relative to the apparent rapid reaction of stocks to unanticipated variations in sales. Second, estimates of inventory speeds of adjustment in firm-level data are significantly higher than in aggregate data. The paper develops a multisector model where inventories are held to avoid stockouts, and price markups vary along the business cycle. The omission of countercyclical markup variations from inventory targets introduces a downward bias in estimates of adjustment speeds obtained from partial adjustment models. When the cyclicality of markups differs across sectors, this downward bias is shown to be more severe with aggregate rather than firm-level data. Similar results apply not only to inventories, but also to labor and prices. Montercarlo simulations of a calibrated version of the model suggest that these biases are quantitatively significant.
Curated

Uniform Crime Reports, 1966-1976: Data Aggregated by Standard Metropolitan Statistical Areas (ICPSR 7743)

Released/updated on: 1992-02-16
Geographic coverage: United States
Time period: 1966-01-01--1976-12-31
This data collection contains a revised SMSA (Standard Metropolitan Statistical Area) aggregate version of the FBI's Uniform Crime Reports (UCR) statistics gathered from 1966-1976, in which original UCR agency records are combined to produce several types of crime rates, by SMSA, for eight crimes. The data were prepared by the Hoover Institution for Economic Studies of the Criminal Justice System, at Stanford University. The data in the file are an aggregation of all relevant law enforcement reporting agencies into 291 SMSAs, and corresponding approximate aggregations of crime rates and dispositions. Each record contains crime rates for one SMSA in one specific year, with data including annual statistics of eight index crimes, i.e., murder, manslaughter, rape, robbery, assault, burglary, larceny, and motor vehicle theft. Calculations include offense-based clearance rates (the number of clearances of juvenile clearances per reported offense), clearance-based rates (the number of persons charged per offense cleared by arrest), and charge-based rates (the number of persons whose cases were disposed in a particular manner per person charged). A related study is UNIFORM CRIME REPORTS, 1966-1976 (ICPSR 7676).
Self-published

Replication data for: Aggregation and the PPP Puzzle in a Sticky-Price Model (ICPSR 112459)

Released/updated on: 2019-10-11
We study the purchasing power parity (PPP) puzzle in a multisector, two-country, sticky-price model. Sectors differ in the extent of price stickiness, leading to heterogeneous sectoral real exchange rate dynamics. Deviations from PPP are more volatile and persistent than in an otherwise identical one-sector world economy with the same average frequency of price changes. Under the empirical distribution of price stickiness of the US economy, the model produces PPP deviations with a half-life of 39 months. We provide a structural interpretation of the approaches found in the empirical literature on aggregation and PPP, and reconcile its apparently conflicting findings. (JEL F31, G31)
Self-published

Replication data for: Aggregate and Idiosyncratic Risk in a Frictional Labor Market (ICPSR 112469)

Released/updated on: 2019-10-11
This paper develops a tractable extension of a Mortensen-Pissarides style matching model that allows for risk averse workers with limited ability to smooth consumption. I show that this leads to a form of equilibrium wage rigidity, as the inability of workers to smooth their consumption across unemployment and employment spells changes how unemployed workers value wage offers, and hence also the offers that employers find profitable to make. In the model risk-averse entrepreneurs use optimal long-term contracts to attract risk averse workers facing limited access to asset markets. A simple analytic representation for the equilibrium is derived. (JEL D81, E21, E24, E32, J31, J41, J64)
Self-published

Code and Public Data for "Aggregate Nominal Wage Adjustments: New Evidence from Administrative Payroll Data" (ICPSR 120647)

Released/updated on: 2021-01-28
Time period: 2008-05-01--2016-12-31
Using administrative payroll data from the largest U.S. payroll processing company, we measure the extent of nominal wage rigidity in the United States.   The data allow us to define a worker's per-period base contract wage separately from other forms of compensation such as overtime premiums and bonuses.  We provide evidence that firms use base wages to cyclically adjust the marginal cost of their workers.  Nominal base wage declines are much rarer than previously thought with only 2% of job-stayers receiving a nominal base wage cut during a given year. Approximately 35% of workers receive no base wage change year over year. We document strong evidence of both time and state dependence in nominal base wage adjustments.   In addition, we provide evidence that the flexibility of new hire base wages is similar to that of existing workers. Collectively, our results can be used to discipline models of nominal wage rigidity.
Self-published

Forest Inventory and Analysis (FIA) invasive plant data aggregated by U.S. county, 2005-2018 (ICPSR 248919)

Released/updated on: 2026-06-01
Time period: 2005-01-01--2018-12-31
Nonnative invasive plant species cause long-term detrimental effects on forest ecosystems, including declines in biological diversity, alterations to forest succession, and changes in nutrient, carbon, and water cycles. The damage caused by these exotic species, and the efforts to control them, are costly, even before accounting for the impacts to nonmarket economic services such as recreation and landscape aesthetics. The Forest Inventory and Analysis (FIA) program collects invasive plant data based on expert-derived lists of problematic invasive plant species defined as those of any growth form likely to cause economic or environmental harm. Using each state's most recent evaluation period between 2005 and 2018, we determined the number and percent of FIA plots invaded by non-native plant species for each U.S. county, as well as the mean number of invasive species and percent cover of invasive species on the plots inventoried for invasive species in each county. These county-level data are provided as both a shapefile and Geopackage.
Self-published

Replication data for: Measured Aggregate Gains from International Trade (ICPSR 114051)

Released/updated on: 2019-10-12
We examine the implications of workhorse trade models for how aggregate productivity, real GDP and real consumption, as measured by statistical agencies, respond to changes in trade costs. In a range of models, changes in measured productivity are equal to the inverse of an export-share weighted average of changes in variable trade costs incurred domestically. Under certain conditions, despite the multiple biases in the CPI, measured real consumption captures the first-order effects of changes in variable trade costs on welfare. Through the lens of these results, we interpret some of the empirical work on measured gains from trade. (JEL E21, E23, F11, F43)
Self-published

Replication data for: Aggregate Implications of Lumpy Investment: New Evidence and a DSGE Model (ICPSR 114283)

Released/updated on: 2019-10-12
The sensitivity of US aggregate investment to shocks is procyclical. The response upon impact increases by approximately 50 percent from the trough to the peak of the business cycle. This feature of the data follows naturally from a DSGE model with lumpy microeconomic capital adjustment. Beyond explaining this specific time variation, our model and evidence provide a counterexample to the claim that microeconomic investment lumpiness is inconsequential for macroeconomic analysis.
Self-published

FHFA Data: Uniform Appraisal Dataset Aggregate Statistics (ICPSR 219961)

Released/updated on: 2025-02-24
Time period: 2013-01-01--2024-12-31
The Uniform Appraisal Dataset (UAD) Aggregate Statistics Data File and Dashboards are the nation’s first publicly available datasets of aggregate statistics on appraisal records, giving the public new access to a broad set of data points and trends found in appraisal reports. The UAD Aggregate Statistics for Enterprise Single-Family, Enterprise Condominium, and Federal Housing Administration (FHA) Single-Family appraisals may be grouped by neighborhood characteristics, property characteristics and different geographic levels.DocumentationOverview (10/28/2024)Data Dictionary (10/28/2024)Data File Version History and Suppression Rates (12/18/2024)Dashboard Guide (2/3/2025)UAD Aggregate Statistics DashboardsThe UAD Aggregate Statistics Dashboards are the visual front end of the UAD Aggregate Statistics Data File.  The Dashboards are designed to provide easy access to customized maps and charts for all levels of users. Access the UAD Aggregate Statistics Dashboards here.UAD Aggregate Statistics DatasetsNotes:
  1. Some of the data files are relatively large in size and will not open correctly in certain software packages, such as Microsoft Excel. All the files can be opened and used in data analytics software such as SAS, Python, or R.
  2. All CSV files are zipped.
Self-published

Replication data for: Composition and Aggregate Real Wage Growth (ICPSR 113514)

Released/updated on: 2019-10-12
Time period: 1981-01-01--2016-12-31
Aggregate real wages exhibit less procyclicality than most macroeconomic models predict. We use 35 years of Current Population Survey data to confirm that the puzzling behavior of wages largely owes to changes in the composition of the employed over the business cycle. This composition effect relates to changes in both the number and the relative wage levels of those entering and exiting. The changing gap in wages of entrants and exiters is especially important for the unemployed. A large part of this wage gap is due to differences in average Mincer residuals between entrants and exiters.
Self-published

County-level Aggregate Expenditure and Risk Score Data on Assignable Beneficiaries (ICPSR 247049)

Released/updated on: 2026-03-21
Time period: 2016-01-01--2024-12-31
The Shared Savings Program County-level Aggregate Expenditure and Risk Score Data on Assignable Beneficiaries Public Use File (PUF) for the Medicare Shared Savings Program (Shared Savings Program) provides aggregate data consisting of per capita Parts A and B FFS expenditures, average CMS-HCC prospective risk scores and total person-years for Shared Savings Program assignable beneficiaries by Medicare enrollment type (End Stage Renal Disease (ESRD), disabled, aged/dual eligible, aged/non-dual eligible).
Self-published

Data and Code: Aggregation Bias in the Measurement of U.S. Global Value Chains (ICPSR 201621)

Released/updated on: 2024-05-13
Time period: 2002-01-01--2012-12-31
This paper analyzes aggregation bias in measuring global value chain (GVC) activity -- the bias arising when an entire industry is essentially treated as a single establishment -- by employing novel U.S. Census microdata. We find negative aggregation bias that has worsened between 2002 and 2012. Unlike with industry-level measures, we see little slowdown in GVC integration by U.S. manufacturers during this period. A decomposition reveals most of the increase in the bias is from within establishments with high export and import intensities. Such granular-level measures of GVC will provide further understanding into how economies adjust to increasingly prevalent shocks
Self-published

Replication data for: Optimal Contracts, Aggregate Risk, and the Financial Accelerator (ICPSR 116393)

Released/updated on: 2019-12-07
This paper derives the optimal lending contract in the financial accelerator model of Bernanke, Gertler, and Gilchrist (1999), henceforth, BGG. The optimal contract includes indexation to the aggregate return on capital, household consumption, and the return to internal funds. This triple indexation results in a dampening of fluctuations in leverage and the risk premium. Hence, compared with the contract originally imposed by BGG, the privately optimal contract implies essentially no financial accelerator. (JEL D11, D81, D86, D92, E13, G31, L26)
Self-published

Replication data for: Menu Costs, Aggregate Fluctuations, and Large Shocks (ICPSR 116411)

Released/updated on: 2019-12-07
We document that the aggregate price level responds flexibly and asymmetrically to large positive and negative value-added tax changes. We present a price-setting model with menu costs, trend inflation, and fat-tailed product-level shocks that is consistent with these observations. The model predicts a flexible price-level response to standard monetary policy shocks because it anticipates a large number of firms on the verge of price adjustment and far from their optimal prices when the shock hits.
Self-published

Replication data for: Heterogeneity and Aggregation: Implications for Labor-Market Fluctuations (ICPSR 116297)

Released/updated on: 2019-12-07
We demonstrate that aggregate employment and consumption can increase without a corresponding movement in productivity in a model with heterogeneous agents where the only aggregate disturbance is a productivity shock. The interaction between incomplete capital markets and indivisible labor results in a low employment-productivity correlation and creates a time-varying wedge between the marginal rate of substitution (for commodity consumption and hours) and productivity. Our results caution against viewing the measured wedge as an inefficiency due to a failure of labor-market clearing or as a fundamental driving force behind business cycles. (JEL D31, E32, J22, J24, J31)
Self-published

Replication data for: Heterogeneity and Aggregation: Implications for Labor-Market Fluctuations: Comment (ICPSR 112762)

Released/updated on: 2019-10-11
Chang and Kim (2007) develop an incomplete asset markets model incorporating discrete labor supply and idiosyncratic labor productivity. Their results resolve long-standing puzzles for business cycle models. Specifically, they produce a low correlation between aggregate hours worked and labor productivity (0.23) and a labor wedge with 76 percent the volatility of output. I show that these results arise from errors in their computational method. I resolve their model using a corrected method and find a strong, positive correlation between hours and productivity (0.80). Fluctuations in the labor wedge decrease to 24 percent of those in output.
Self-published

Replication data for: Heterogeneity and Aggregation: Implications for Labor-Market Fluctuations: Reply (ICPSR 112763)

Released/updated on: 2019-10-11
Takahashi (2014) has uncovered coding errors in our paper, Chang and Kim (2007)-henceforth, CK. We acknowledge and are embarrassed by these mistakes. We are grateful to Takahashi for uncovering them. While the correction decreases the volatility of the labor market wedge, we find that the main message of CK remains valid: the measured labor market wedge arises endogenously in an economy with incomplete capital markets and indivisible labor supply. For example, our model accounts for 18 percent of the volatility in the labor market wedge in the data; it was 43 percent in CK.
Self-published

Data and Code for: Voter Turnout and Preference Aggregation (ICPSR 111041)

Released/updated on: 2021-10-21
We study how voter turnout affects the aggregation of preferences in
elections. Under voluntary voting, election outcomes disproportionately
aggregate the preferences of voters with low voting cost and high preference
intensity. We show identification of the correlation structure among
preferences, costs, and perceptions of voting efficacy, and explore how the
correlation affects preference aggregation. Using 2004 U.S. presidential
election data, we find that young, low-income, less-educated, and minority
voters are underrepresented. All of these groups tend to prefer Democrats,
except for the less-educated. Democrats would have won the majority of the
electoral votes if all eligible voters had turned out.
Self-published

Replication data for: Establishment Size Dynamics in the Aggregate Economy (ICPSR 113221)

Released/updated on: 2019-10-12
This paper presents a theory of establishment size dynamics based on the accumulation of industry-specific human capital that simultaneously rationalizes the economy- wide facts on establishment growth rates, exit rates, and size distributions. The theory predicts that establishment growth and net exit rates should decline faster with size, and that the establishment size distribution should have thinner tails, in sectors that use specific human capital less intensively. We establish that there is substantial cross-sector heterogeneity in US establishment size dynamics and distributions, which is well explained by relative factor intensities. (JEL L11 , L16, L25).
Self-published

Replication data for: The Aggregate Impact of Household Saving and Borrowing Constraints: Designing a Field Experiment in Uganda (ICPSR 116121)

Released/updated on: 2019-12-06
We develop a model of households with multiple needs (smoothing shocks, financing investment) and constraints (limited credit, self-control issues) in order to examine the nature of household's financing constraints in a developing country, and the impact of relaxing them. We show that increased access to credit has very different implications for the aggregate model economy depending on its form: asset-financed or cash. We then illustrate how a short-term increase in access to loans leads to very distinct behavior in the short run. The relevance of the model can be evaluated using a field experiment, which we are currently implementing in Uganda.
Self-published

Data and Code for: Unemployment Insurance Generosity and Aggregate Employment (ICPSR 118525)

Released/updated on: 2021-04-07
Time period: 2007-01-01--2014-12-31
This paper examines the impact of unemployment insurance (UI) on aggregate employment by exploiting cross-state variation in the maximum benefit duration during the Great Recession. Comparing adjacent counties located in neighboring states, there is no statistically significant impact of increasing UI generosity on aggregate employment. Point estimates are uniformly small in magnitude, and the most precise estimates rule out employment-to-population ratio reductions in excess of 0.35 percentage points from the UI extension. The results contrast with the negative effects implied by most micro-level labor supply studies and are consistent with both job rationing and aggregate demand channels.
Self-published

Data and Code for: Sectoral Media Focus and Aggregate Fluctuations (ICPSR 145301)

Released/updated on: 2021-11-19
Time period: 1987-01-01--2018-12-01
We formalize the editorial role of news media in a multi-sector economy and
show that media can be an independent source of business cycle fluctuations,
even when they report accurate information. Public reporting about a subset
of sectoral developments that are newsworthy but unrepresentative causes
firms across all sectors to hire too much or too little labor. We construct
historical measures of US sectoral news coverage and use them to calibrate
our model. Time-varying media focus generates demand-like fluctuations that
are orthogonal to productivity, even in the absence of non-TFP shocks.
Presented with historical sectoral productivity, the model reproduces the
2009 Great Recession.
Self-published

Data and Code for: "Aggregating Distributional Treatment Effects: A Bayesian Hierarchical Analysis of the Microcredit Literature" (ICPSR 155821)

Released/updated on: 2022-05-25
Time period: 2009-01-01--2015-12-31
This repository contains all R scripts needed to generate the main results of the paper entitled “Aggregating Distributional Treatment Effects: A Bayesian Hierarchical Analysis of the Microcredit Literature.” In its default state the masterfile.R generates the paper's tables and figures from saved MCMC output which takes about 15 minutes. If rStan is installed, then the masterfile can be toggled to run the MCMC scripts from the raw data, which takes about 72 hours on a high performance computing server.
Self-published

Supplementary data for “Wealth Inequality, Aggregate Consumption, and Macroeconomic Trends under Incomplete Markets” (ICPSR 197181)

Released/updated on: 2024-06-13
Geographic coverage: United States
Time period: 1983-01-01--2018-12-31
I construct an incomplete market model featuring a closed-form expression for optimal consumption. In the model, individual consumption is an isoelastic function of wealth, inclusive of income, yielding partial consumption smoothing based on borrowing and lending in response to income shocks. I show that the model replicates several empirical characteristics of inequality in consumption, income, and wealth and their dynamics at the individual level. Using the model, I show that the rising wealth inequality since the 1980s, induced by an increase in idiosyncratic income risk, has substantially contributed to trend-level changes in real interest rates, capital-to-income ratios, and consumption-to-wealth ratios
Self-published

Replication data for: Entry, Exit, Firm Dynamics, and Aggregate Fluctuations (ICPSR 116400)

Released/updated on: 2019-12-07
Firm entry and exit amplify and propagate the effects of aggregate shocks, leading to greater persistence and unconditional variation of aggregate quantities. Following a positive aggregate shock, entry rises. As in the data, entrants are small and their initial impact on aggregate dynamics is negligible. However, as the common productivity component reverts to its unconditional mean, the youngsters that survive grow larger, generating a wider and longer expansion than in a scenario without entry or exit. The model also identifies a causal link between the drop in establishments at the outset of the Great Recession and the subsequent slow recovery.
Self-published

Data and Code for: The Elasticity of Aggregate Output with Respect to Capital and Labor (ICPSR 194858)

Released/updated on: 2024-08-16
Time period: 1948-01-01--2018-12-31
It is often assumed that the elasticity of GDP with respect to capital is one-third, but this assumes zero markups and an aggregate production function. I estimate the elasticity allowing markups to vary by industry and with a rich input-output structure. Assumptions about capital costs provide bounds on the elasticity. In the U.S. from 1948-1995 the capital elasticity was in the range 0.19- 0.32 and this shifted to 0.24-0.37 by 1996-2018. Excluding housing or de-capitalizing intellectual property lowers those bounds to as low as 0.11-0.26. Based on these elasticities, common estimates of total factor productivity growth represent a lower bound.
Self-published

Replication data for: Labor Market Heterogeneity and the Aggregate Matching Function (ICPSR 114096)

Released/updated on: 2019-10-12
We estimate an aggregate matching function and find that the regression residual, which captures movements in matching efficiency, displays procyclical fluctuations and a dramatic decline after 2007. Using a matching function framework that explicitly takes into account worker heterogeneity as well as market segmentation, we show that matching efficiency movements can be the result of variations in the degree of heterogeneity in the labor market. Matching efficiency declines substantially when, as in the Great Recession, the average characteristics of the unemployed deteriorate substantially, or when dispersion in labor market conditions—the extent to which some labor markets fare worse than others—increases markedly. (JEL E24, E32, J41, J42)
Self-published

Replication data for: How Transparency Kills Information Aggregation: Theory and Experiment (ICPSR 114353)

Released/updated on: 2019-10-12
Time period: 2013-05-08--2013-05-30
We investigate the potential of transparency to influence committee decision-making. We present a model in which career concerned committee members receive private information of different type-dependent accuracy, deliberate, and vote. We study three levels of transparency under which career concerns are predicted to affect behavior differently and test the model's key predictions in a laboratory experiment. The model's predictions are largely borne out—transparency negatively affects information aggregation at the deliberation and voting stages, leading to sharply different committee error rates than under secrecy. This occurs despite subjects revealing more information under transparency than theory predicts.
Self-published

Replication data for: Evaluating Microfoundations for Aggregate Price Rigidities: Evidence from Matched Firm-Level Data on Product Prices and Unit Labor Cost (ICPSR 112531)

Released/updated on: 2019-10-11
Time period: 1990-01-01--2002-12-31
Using matched data on product-level prices and the producing firm's unit labor cost, we find a moderate pass-through of current idiosyncratic marginal-cost changes. Also, the response does not vary across firms facing very different idiosyncratic shock variances, but identical aggregate conditions. These results do not fit the predictions of Mackowiak and Wiederholt (2009). Neither do firms react strongly to predictable marginal-cost changes, as expected from Mankiw and Reis (2002). We find that firms consider both current and expected future marginal cost when setting prices. This points toward impediments to continuous price adjustments as a key driver of monetary non-neutrality.
Self-published

Data and Code for: Firm Entry and Exit and Aggregate Growth (ICPSR 145861)

Released/updated on: 2022-12-09
Time period: 1992-01-01--2014-12-31
Data and programs required for replicating analysis, tables, and figures in “Firm Entry and Exit and Aggregate Growth” by Jose Asturias, Sewon Hur, Timothy J. Kehoe, and Kim J. Ruhl.
Self-published

Replication data for: Knowledge Capital and Aggregate Income Differences: Development Accounting for US States (ICPSR 114144)

Released/updated on: 2019-10-12
Improvement in human capital is often presumed to be important for state economic development, but little research links better education to state incomes. We develop detailed measures of worker skills in each state that incorporate cognitive skills from state- and country-of-origin achievement tests. These new measures of knowledge capital permit development accounting analyses calibrated with standard production parameters. Differences in knowledge capital account for 20-30 percent of the state variation in per capita GDP, with roughly even contributions by school attainment and cognitive skills. Similar results emerge from growth accounting analyses. These estimates support school improvement as a strategy for state economic development.
Self-published

Data and Code for: Confidence, Self-Selection and Bias in the Aggregate (ICPSR 185741)

Released/updated on: 2023-06-19
The influence of behavioral biases on aggregate outcomes depends in part on self-selection: whether rational people opt more strongly into aggregate interactions than biased individuals. In betting market, auction and committee experiments, we document that some errors are strongly reduced through self-selection, while others are not affected at all or even amplified. A large part of this variation is explained by differences in the relationship between confidence and performance. In some tasks, they are positively correlated, such that self-selection attenuates errors. In other tasks, rational and biased people are equally confident, such that self-selection has no effects on aggregate quantities.
Self-published

Data and Code for "The Aggregate-Demand Doom Loop: Precautionary Motives and the Welfare Costs of Sovereign Risk" (ICPSR 197281)

Released/updated on: 2025-06-06
Time period: 2000-01-01--2018-12-31
The data and code in this replication package produces all 26 figures and 8 tables for the paper "The Aggregate-Demand Doom Loop: Precautionary Motives and the Welfare Costs of Sovereign Risk," by Francisco Roldán.
The paper examines the role of households' precautionary savings motive in amplifying and propagating movements in sovereign spreads. It studies this mechanism in a model where the government of a small open economy borrows from foreigners but the debt is then partially held by heterogeneous domestic savers. In a calibration to Spain in the 2000s, it finds that default risk accounts for about half of the output contraction. More generally, sovereign risk exacerbates volatility in consumption over time and across agents, creating large and unequal welfare costs even if default does not materialize.
Self-published

Replication data for: The Relative Importance of Aggregate and Sectoral Shocks and the Changing Nature of Economic Fluctuations (ICPSR 116397)

Released/updated on: 2019-12-07
A principal components decomposition of sectoral IP data reveals that the contribution of aggregate shocks to the variance of aggregate output declined from about 70 percent in the period 1967–1983 to about 30 percent after 1983. We develop an "islands" model with two sectors and costly labor reallocation to investigate how this change in the relative importance of shocks alters business cycle moments. A version of the model with relatively more important sectoral shocks results in a sizeable decline in the cyclicality of labor productivity and is consistent with changes in several other business cycle moments observed in the data.
Back to top