How Extractive Was Russian Serfdom? Income Inequality in Moscow Province in the Early Nineteenth Century (ICPSR 301023)
This replication package contains the complete dataset used in "How Extractive Was Russian Serfdom? Income Inequality in Moscow Province in the Early Nineteenth Century" (Korchmina and Malinowski, 2026).
The package includes: (1) individual-level income records for 7,399 asset-holding households in Moscow Province in 1811, including all registered aristocrats and merchants; (2) estimated incomes for 21 additional social groups derived from government and private business financial records; (3) the constructed social table with pre- and post-tax income measures; (4) calculated inequality measures (Gini coefficients, Extraction Ratio, top income shares) for 1811 and 1904; (5) two maps showing geographic distribution of income and social groups; and (6) detailed documentation of data sources, variable definitions, and calculation methods.
All data are provided in Excel format with accompanying Word documents containing methodological notes, and data source descriptions.
Mapping Language Literacy At Scale: A Case Study on Facebook (ICPSR 300445)
Literacy is one of the most fundamental skills for people to access and navigate today’s digital environment. This work systematically studies the language literacy skills of online populations for more than 160 countries and regions across the world, including many low-resourced countries where official literacy data are particularly sparse. Leveraging public data on Facebook, we develop a population-level literacy estimate for the online population that is based on aggregated and de-identified public posts written by adult Facebook users globally, significantly improving both the coverage and resolution of existing literacy tracking data. We found that, on Facebook, women collectively show higher language literacy than men in many countries, but substantial gaps remain in Africa and Asia. Further, our analysis reveals a considerable regional gap within a country that is associated with multiple socio-technical inequalities, suggesting an “inequality paradox” – where the online language skill disparity interacts with offline socioeconomic inequalities in complex ways. These findings have implications for global women’s empowerment and socioeconomic inequalities.
The Return to College, Marriage, and Intergenerational Mobility (ICPSR 302706)
This dataset is a custom extract from the public-use Panel Study of Income Dynamics (PSID), produced through the PSID Data Center (Job ID J320456). The extract contains 775 family-level variables drawn from multiple survey waves spanning 1968 through 2019. Variables were selected to construct demographic, geographic, and educational measures for the household reference person (head) and spouse/partner.
The extract includes variables on age, sex, race, education, parental education, place of birth, state of residence, state grew up, state at age 25, and related interview year identifiers. The data are organized in wide format across waves in the original extract and are subsequently processed using the accompanying replication code to create harmonized cross-sectional measures based on the most recent non-missing report for each individual. The replication files also reshape the data into an individual-level long format separating head and spouse observations.