Macroeconomic Stars (ICPSR 227362)
- U-star: long-run level of unemployment rate
- R-star: long-run real rate of interest
- Pi-star: long-run level of price inflation
- P-star: long-run level of productivity growth
- W-star: long-run level of nominal wage inflation
- G-star: growth rate of potential output
- Output Gap: cyclical assessment of the US economy
- Persistence in price inflation (gap)
- Persistence in nominal wage inflation (gap)
- Slope of the price Phillips Curve
- Slope of the wage Phillips Curve
- Short-run passthrough from prices to wages
- Wedge: between W-star and (P-star + Pi-star)
- D: the catch all component in R-star equation
- Stochastic volatility price inflation gap
- Stochastic volatility nominal wage inflation gap
- Stochastic volatility labor productivity gap
- Stochastic volatility interest rate gap
- Stochastic volatility output gap
- Stochastic volatility UR gap
Statistical Methods for Development, Validation, and Implementation of Absolute Risk Models [Methods Study], 2016-2022 (ICPSR 39730)
Factors, such as personal traits, behaviors, or the environment, can affect a person's risk of getting an illness. Doctors can use risk models, which account for these factors, to predict a person's chance of getting an illness. The risk models group patients into different levels for certain illnesses, such as high risk or low risk.
Most risk models look at only a small number of factors, which affects how well the models can separate patients into different levels. Combining factors from different studies into a single risk model may improve how well the model works. Researchers can use statistical methods to combine data from different studies. But current methods don't work when the studies look at different traits or other factors.
In this study, the research team developed a new method for combining data from studies that have information on different risk factors. The new method is called Generalized Meta-Analysis, or GENMETA.
To access the R package, please visit the Implements Generalized Meta-Analysis Using Iterated Reweighted Least Square Algorithm CRAN webpage.
Strengthening Work Requirements? Forecasting Impacts of Reforming Cash Assistance Rules (ICPSR 212481)
Forecasting bilateral asylum seeker flows with high-dimensional data and machine learning techniques (ICPSR 198322)
Using Machine Learning to Identify High-Risk Domestic Violence Offenders in New York City, New York, 2006-2017 (ICPSR 38540)
To address the relative difficulty in predicting domestic violence incidents and effectively targeting resources, the University of Chicago Crime Lab and the New York Police Department (NYPD) collaborated to develop and test a machine learning-based statistical model to predict the risk of domestic violence victimization in New York City.
Phase 1 of the project was to develop a statistical model using machine learning techniques. NYPD administrative records dated between January 2006 and January 2017 were used as input data to build and refine the tool. Due to the lack of unique identifiers for victims in the records, the research team also used data from the Chicago Police Department to create a probabilistic record linkage toolkit (Name Match) to identify which records belonged to the same person within and across data sources.
In Phase 2, the researchers aimed to field test the tool's capability to identify individuals at risk of repeated domestic violence through a large-scale randomized control trial. Measuring the effects of regular home visits of high-priority individuals thought to be at risk of serious domestic assault, the test intended to compare the selections of individuals made by officers versus those predicted by the tool.
This collection contains only the machine learning code files (R and Python) created during secondary analysis, which have been released as a zipped package. Please refer to the Data Roadmap for instructions on how to obtain the original NYPD data. To access the Name Change algorithm and documentation, please visit the Github repository.
Comparing the Growth and Predictive Performance of a Traditional Oral Reading Fluency Measure to an Experimental Novel Measure (ICPSR 156501)
National Center for Early Development and Learning Multistate Study of Pre-Kindergarten, 2001-2003 (ICPSR 4283)
The National Center for Early Development and Learning (NCEDL) Multi-State Study of Pre-Kindergarten examined the pre-kindergarten programs of six states: California, Illinois, New York, Ohio, Kentucky, and Georgia. For this study, pre-kindergarten (pre-k) included center-based programs for four-year-olds that are fully or partially funded by state education agencies and that are operated in schools or under the direction of state and local education agencies.
The study had two primary purposes:
To describe the variations of experiences for children in pre-kindergarten and kindergarten programs in school-related settings (public schools and state-funded pre-k classrooms in community-based settings), and
To examine the relationships between variations in pre-kindergarten/kindergarten experiences and children's outcomes in early elementary school.
The study addressed six primary groups of research questions:
What is the nature and distribution of education and experience of teachers and teacher assistants in pre-k public school programs?
What is the nature and distribution of global quality and specific practices in key areas such as literacy, math, and teacher-child relationships in a diverse sample of pre-k public school programs for four-year-olds as well as in a similarly diverse sample of kindergarten classes?
How do quality and practices vary as a result of child and teacher characteristics (e.g., child gender, race, home language, family income, and teacher's years of education) and classroom, program, community, and state structural variables (e.g., teacher-child ratio, funding base of the program, teacher salary, and degree of state regulation) for children with different demographic characteristics (e.g., race, gender, home language, and family income)?
Do quality and practice vary in relation to combinations of these variables? For example, are quality and practice a function of family poverty and teacher pay or education?
Can children's outcomes at the end of their pre-kindergarten year be predicted by the children's experiences in pre-k programs? Are the various dimensions of quality and/or practice differentially related to outcomes? Are these relationships constant across a population of children with different characteristics (e.g., race, gender, home language, and family income)?
Do pre-kindergarten program quality and practices predict children's transitions to kindergarten and children's skills at the end of the kindergarten year? Are these transitions moderated by children's characteristics, like race, gender, and family income?
The six states in the study were selected based on the significant amount of resources they have committed to pre-k initiatives. States were also selected to maximize the diversity in geography, program settings (public school or community), program intensity (full day versus part day), and educational requirements for teachers. Within each state, a random sample of 40 centers/schools was selected. One classroom in each center/school was selected at random for observation, and four children in each classroom were selected for individual assessment. The children were followed from the beginning of pre-k through the end of kindergarten. In five of the six states, families were also visited in their homes.
Classroom Services and Specific Instructional Practices
Within the 40 classrooms in each participating state, carefully trained data collectors conducted classroom observations twice each year, while additional surveys were used to gather information from administrators/principals, teachers, and parents. Data were gathered on program services, (e.g., healthcare, meals, and transportation), program curriculum, teacher training and education, teachers' opinions of child development, and their instructional practices on subjects such as language, literacy, mathematics concepts, and social-emotional competencies. Data were also collected as to what types of steps were taken to aid children in their transitions from pre-k to kindergarten.
Children
Within each participating pre-k classroom, four randomly selected children were assessed using a battery of individual instruments to measure language, literacy, mathematics, and related concept development, as well as social competence. A panel of expert reviewers aided the researchers in selecting a variety of standardized and nonstandardized assessments. The pre-k child assessments were conducted in the fall and spring of 2001-2002. The same children were followed into kindergarten and assessed in the fall and spring of 2002-2003 to examine whether specific practices employed by pre-k teachers made a difference in their transitions to kindergarten.
Families
In individual home-based interviews, information on socio-economic, socio-cultural, and familial contexts were obtained through open-ended questions, structured ratings, and videotaped parent-child interactions. Specifically, parents were asked about (1) family life as it relates to socio-economic status and socio-cultural environment, (2) family educational practices and beliefs about the comparative roles of school and family in educating children, (3) the nature and quality of the home-school relationship, and (4) their own ratings of their children's psychological development and social competence.
Demographic information collected includes race, gender, family income, and mother's education level.
The above information pertains to the Main Child Level Public-Use Version and the Main Child Level Restricted-Use Version. From these main datasets, subsets were created at the classroom level for Pre-Kindergarten (Pre-K Classroom Level Public-Use Version and Pre-K Classroom Level Restricted-Use Version) and for Kindergarten (Kindergarten Classroom Level Public-Use Version and Kindergarten Classroom Level Restricted-Use Version).
Crime Hot Spot Forecasting with Data from the Pittsburgh [Pennsylvania] Bureau of Police, 1990-1998 (ICPSR 3469)
This study used crime count data from the Pittsburgh, Pennsylvania, Bureau of Police offense reports and 911 computer-aided dispatch (CAD) calls to determine the best univariate forecast method for crime and to evaluate the value of leading indicator crime forecast models.
The researchers used the rolling-horizon experimental design, a design that maximizes the number of forecasts for a given time series at different times and under different conditions. Under this design, several forecast models are used to make alternative forecasts in parallel. For each forecast model included in an experiment, the researchers estimated models on training data, forecasted one month ahead to new data not previously seen by the model, and calculated and saved the forecast error. Then they added the observed value of the previously forecasted data point to the next month's training data, dropped the oldest historical data point, and forecasted the following month's data point. This process continued over a number of months.
A total of 15 statistical datasets and 3 geographic information systems (GIS) shapefiles resulted from this study.
The statistical datasets consist of
- Univariate Forecast Data by Police Precinct (Dataset 1) with 3,240 cases
- Output Data from the Univariate Forecasting Program: Sectors and Forecast Errors (Dataset 2) with 17,892 cases
- Multivariate, Leading Indicator Forecast Data by Grid Cell (Dataset 3) with 5,940 cases
- Output Data from the 911 Drug Calls Forecast Program (Dataset 4) with 5,112 cases
- Output Data from the Part One Property Crimes Forecast Program (Dataset 5) with 5,112 cases
- Output Data from the Part One Violent Crimes Forecast Program (Dataset 6) with 5,112 cases
- Input Data for the Regression Forecast Program for 911 Drug Calls (Dataset 7) with 10,011 cases
- Input Data for the Regression Forecast Program for Part One Property Crimes (Dataset 8) with 10,011 cases
- Input Data for the Regression Forecast Program for Part One Violent Crimes (Dataset 9) with 10,011 cases
- Output Data from Regression Forecast Program for 911 Drug Calls: Estimated Coefficients for Leading Indicator Models (Dataset 10) with 36 cases
- Output Data from Regression Forecast Program for Part One Property Crimes: Estimated Coefficients for Leading Indicator Models (Dataset 11) with 36 cases
- Output Data from Regression Forecast Program for Part One Violent Crimes: Estimated Coefficients for Leading Indicator Models (Dataset 12) with 36 cases
- Output Data from Regression Forecast Program for 911 Drug Calls: Forecast Errors (Dataset 13) with 4,936 cases
- Output Data from Regression Forecast Program for Part One Property Crimes: Forecast Errors (Dataset 14) with 4,936 cases
- Output Data from Regression Forecast Program for Part One Violent Crimes: Forecast Errors (Dataset 15) with 4,936 cases.
- The GIS Shapefiles (Dataset 16) are provided with the study in a single zip file: Included are polygon data for the 4,000 foot, square, uniform grid system used for much of the Pittsburgh crime data (grid400); polygon data for the 6 police precincts, alternatively called districts or zones, of Pittsburgh(policedist); and polygon data for the 3 major rivers in Pittsburgh the Allegheny, Monongahela, and Ohio (rivers).
Validation of Risk Assessment Tools for Predicting Re-offending at Different Developmental Periods, 1951-2010 (ICPSR 32761)
Pre-Kindergarten in Eleven States: NCEDL's Multi-State Study of Pre-Kindergarten and Study of State-Wide Early Education Programs (SWEEP) (ICPSR 34877)
The National Center for Early Development and Learning (NCEDL) combined the data of two major studies in order to understand variations among state-funded pre-kindergarten (pre-k) programs and in turn, how these variations relate to child outcomes at the end of pre-k and in kindergarten. The Multi-State Study of Pre-Kindergarten and the State-Wide Early Education Programs (SWEEP) Study provide detailed information on pre-kindergarten teachers, children, and classrooms in 11 states. By combining data from both studies, information is available from 721 classrooms and 2,982 pre-kindergarten children in these 11 states.
Pre-kindergarten data collection for the Multi-State Study of Pre-Kindergarten took place during the 2001-2002 school year in six states: California, Georgia, Illinois, Kentucky, New York, and Ohio. These states were selected from among states that had committed significant resources to pre-k initiatives. States were selected to maximize diversity with regard to geography, program settings (public school or community setting), program intensity (full-day vs. part-day), and educational requirements for teachers. In each state, a stratified random sample of 40 centers/schools was selected from the list of all the school/centers or programs (both contractors and subcontractors) provided to the researchers by each state's department of education.
In total, 238 sites participated in the fall and two additional sites joined the study in the spring. Participating teachers helped the data collectors recruit children into the study by sending recruitment packets home with all children enrolled in the classroom. On the first day of data collection, the data collectors determined which of the children were eligible to participate. Eligible children were those who (1) would be old enough for kindergarten in the fall of 2002, (2) did not have an Individualized Education Plan, according to the teacher, and (3) spoke English or Spanish well enough to understand simple instructions, according to the teacher.
Pre-kindergarten data collection for the SWEEP Study took place during the 2003-2004 school year in five states: Massachusetts, New Jersey, Texas, Washington, and Wisconsin. These states were selected to complement the states already in the Multi-State Study of Pre-K by including programs with significantly different funding models or modes of service delivery. In each of the five states, 100 randomly selected state-funded pre-kindergarten sites were recruited for participation in the study from a list of all sites provided by the state.
In total, 465 sites participated in the fall. Two sites declined to continue participation in the spring, resulting in 463 sites participating in the spring. Participating teachers helped the data collectors recruit children into the study by sending recruitment packets home with all children enrolled in the classroom. On the first day of data collection, the data collectors determined which of the children were eligible to participate. Eligible children were those who (1) would be old enough for kindergarten in the fall of 2004, (2) did not have an Individualized Education Plan, according to the teacher, and (3) spoke English or Spanish well enough to understand simple instructions, according to the teacher.
Demographic information collected across both studies includes race, teacher gender, child gender, family income, mother's education level, and teacher education level.
The researchers also created a variable for both the child-level data and the class-level data which allows secondary users to subset cases according to either the Multi-State or SWEEP study.