Search results

Search tips
Showing 1 – 39 of 39 results.
Self-published

False Positives in AI Writing Detection: A Small-Scale Empirical Study Using Authentic Filipino Student Essays (ICPSR 251455)

Released/updated on: 2026-07-29
Geographic coverage: Philippines
Writer(s)
1. Mhel Cedric D. Bendo
Abstract
This research note reports a small-scale exploratory study into how two AI writing detectors, ZeroGPT and Copyleaks, classify authentic student essays. A total of ten anonymised college-student essays were analysed to observe misclassification patterns, particularly false positives, where human-written content has been incorrectly flagged as AI-generated. The assessment was conducted focusing on essays produced by Filipino undergraduates whose nonnative English writing may have features that could lead to misclassification by the detectors. Results show that both detectors inconsistently classified the same set of essays: five essays were labelled as “AI-generated,” while the other five were labelled as “Human,” when in fact all the texts were authentically written by students. These results point to the potential misclassification risks when AI detection tools are used within educational contexts, where it is commonplace for teachers to make important decisions about academic integrity based on the outputs of detectors. The present study underlines the need to validate the outputs of AI detectors with human judgment and advises educators against the use of these tools as sole evidence of misconduct. Implications for ICT-supported assessment practices and policies of academic honesty are discussed, together with recommendations for a more responsible integration of AI detection tools in educational contexts. The results described above have implications for the practice of academic integrity in all countries, but also specifically for those countries with multilingual student bodies whose writing styles may differ from the writing style used to train the detectors.https://doi.org/10.64233/VYVI9613
Self-published

Intelligence Theory 2026: Supplementary Research Data and Documentation (ICPSR 251126)

Released/updated on: 2026-07-07
Geographic coverage: Earth
Time period: 2024-01-01--2026-12-31
Intelligence Theory 2026: Supplementary Research Data and Documentation is a supporting research collection accompanying the publication Intelligence Theory 2026. The deposit contains supplementary documentation, conceptual frameworks, analytical models, figures, reference materials, and related research documents developed during the preparation of the study. These materials are intended to improve research transparency, facilitate scholarly reuse where appropriate, and provide additional context for topics including artificial intelligence, open-source intelligence (OSINT), intelligence analysis, strategic decision-making, and contemporary intelligence theory. The collection does not contain personally identifiable information or confidential research data.
Self-published

A systematic discourse analysis of how U.S. political leaders frame disability: Implications for students with disabilities (ICPSR 307656)

Released/updated on: 2026-07-06
Geographic coverage: United States
Time period: 2025-01-01--2025-12-31

This study examined how U.S. political leaders publicly framed disability during calendar year 2025, using a systematic discourse analysis of public statements and federal legislation. It also examined how this discourse was reflected in federal education policy, including statements by the Secretary of Education and education-related bills.

The study analyzed 121 public statements about people with disabilities made by White House and Cabinet officials, and 32 federal bills introduced in 2025 that could affect the rights, services, or educational opportunities of individuals with disabilities. Statements were identified through a two-stage process that combined a custom Python-based web-scraping tool, which extracted verbatim, attributed quotations from news articles, press releases, interview transcripts, official speeches, and social media posts using the OpenAI GPT-4 API, with manual verification searches conducted in ChatGPT Plus and Perplexity AI Pro. Federal bills were identified through Congress.gov. All statements were reviewed by the authors to confirm accuracy, attribution, and date. Each statement and bill was independently scored by two human coders and by GPT-4 using an author-developed four-point rubric grounded in the social and human rights models of disability, ranging from 1 (dehumanizing) to 4 (affirming).

The data contain one record per statement, including the speaker's name and title, date, verbatim quotation, source, context, an analysis of the framing, a score from 1 (dehumanizing) to 4 (affirming), and the rationale for the score. The bills data are available on the project website and are not included in this deposit.

The study was approved by the university Institutional Review Board (March 2025) and was preregistered on the Open Science Framework (https://osf.io/hbfe5/).

Self-published

AI-Supported Inquiry, AI Literacy, and Authentic Performance Among Preservice Teachers in Central China, 2024 (ICPSR 306435)

Released/updated on: 2026-06-03
Geographic coverage: Jiangxi, China
Time period: 2024-09-01--2024-12-31

This study examined the effects of QUEST+AI, an AI-supported inquiry model, on AI literacy and authentic performance among preservice teachers in Central China. The study used a nonequivalent-groups, quasi-experimental pretest–posttest design with two intact sections of an undergraduate Educational Research Methods course at a public university. Ninety-five preservice teachers participated, including 52 students in the experimental group and 43 students in the comparison group. Both groups received the same face-to-face course instruction over a 10-week period. The experimental group completed two QUEST+AI inquiry cycles with coached use of generative artificial intelligence, while the comparison group completed conventional homework assignments.

The data include participant demographic variables, pretest and posttest responses to a multidimensional AI literacy questionnaire, AI literacy total and subscale scores, final research proposal scores, and group assignment indicators. AI literacy measures cover use and application of AI, knowledge and understanding of AI, AI detection, AI ethics, AI creation, AI-supported problem solving, AI persuasion literacy, and AI emotion regulation. Authentic performance is represented by scores on a final educational research proposal, evaluated with a common rubric by two independent raters. The dataset also includes variables used in the study’s comparative analyses, including group condition, pretest scores, gender, grade level, age, and major.

Curated

Using Topic Segmentation to Enhance Concept Parsing and Identification of Negations [Methods Study], Massachusetts, 2019-2023 (ICPSR 39740)

Released/updated on: 2026-03-23
Geographic coverage: United States, Massachusetts
Time period: 2019-01-01--2023-12-31

Clinical notes in electronic health records, or EHRs, may contain information that can help researchers study and compare treatments. But it takes researchers a lot of time to find information in EHR notes.

Natural language processing, or NLP, methods can help researchers find information in EHR notes. With NLP, computer programs read and identify written language to make it easier to sort and study. But in EHR notes, some sentences may contain more than one topic. Also, EHR notes may discuss a single topic over many sentences. In these cases, current NLP methods don't work well to find complete and accurate information about a specific topic.

In this study, the research team developed and tested new NLP methods to identify topics from EHR notes.

Curated

Development of Computational Methods for Evaluating Doctor-Patient Communication [Methods Study], United States, 2016-2021 (ICPSR 39720)

Released/updated on: 2026-03-18
Geographic coverage: United States
Time period: 2016-01-01--2021-12-31

The way doctors communicate with patients during office visits can affect the quality of care. Studying conversations between doctors and patients can help doctors improve their communication skills.

To study conversations, researchers rely on written records, or transcripts, of office visits. They read the transcripts and give each conversation topic a label. For example, topics may include smoking or pain. But labeling topics in this way may take a lot of time.

In this project, the research team created and tested a new method to make this work easier using natural language processing, or NLP. With NLP, computer programs interpret written language. NLP methods use a process called machine learning, where computer programs use data to learn how to perform different tasks with little or no human input.

Curated

Unlocking Clinical Text in Electronic Medical Records (EMR) by Query Refinement Using Both Knowledge Bases and Word Embedding [Methods Study], Ohio, 2006-2022 (ICPSR 39734)

Released/updated on: 2026-03-16
Geographic coverage: United States, Ohio
Time period: 2006-01-01--2022-12-31

Electronic health records, or EHRs, have information about a patient's health such as test results, diagnoses, and treatments. EHRs also have clinical notes that doctors and patients can use to track goals and decisions.

Clinical notes may be useful for research or to help improve care. But it's hard to get information from these notes across large groups of patients. The notes may use different ways to describe the same thing. For example, high blood pressure may be called hypertension. Also, the notes may use abbreviations or have spelling mistakes.

In this project, the research team designed and built a search engine to make EHR notes easier to search and use for patient care and research.

Curated
Partially restricted
Simple Crosstabs

Detroit Metro Area Communities Study (DMACS) Wave 22, Michigan, 2025 (ICPSR 39692)

Released/updated on: 2026-02-26
Geographic coverage: Detroit, United States, Michigan
Time period: 2025-08-06--2025-10-01

The Detroit Metro Area Communities Study (DMACS) is a panel survey of Detroit residents aged 18 and older. The original panel of respondents was drawn from an address-based probability sample of all occupied Detroit households in 2016 and has since been refreshed through additional address-based sampling annually. Between August 6, 2025 and October 1, 2025, 3,170 previously enrolled panelists were invited to participate in a self-administered online or interviewer-administered telephone survey.

Topics included: household composition; housing status; perceptions of neighborhood; social connection and loneliness; election; mayoral priorities; crime and safety; violence reduction; artificial intelligence; flood management; mental health; employment.

Self-published

ChatGPT and Willingness to share data (ICPSR 244947)

Released/updated on: 2026-02-04
Time period: 2025-01-15--2025-04-15
This study explores how the desire for unique consumer products (DUCP) affects users' willingness to share personal data (WTS) within personalized interactions with ChatGPT
Curated
Simple Crosstabs

National Assessment of Demand Reduction Efforts, Part II: New Developments in the Primary Prevention of Sex Trafficking, [United States], 2021 (ICPSR 38928)

Released/updated on: 2026-01-14
Geographic coverage: United States
Time period: 2021-01-01--2021-12-31

To combat prostitution and sex trafficking, criminal justice strategies and collaborative programs have emerged that focus on reducing consumer-level demand. From 2008 to 2012, the National Institute of Justice (NIJ) sponsored a study entitled "A National Overview of Prostitution and Sex Trafficking Demand Reduction Efforts" (referred to as Part I) that featured the systematic collection of information to determine the types and distribution of demand reduction tactics implemented throughout the United States. These efforts gave rise to a typology of law enforcement and community-based tactics identifying 12 different methods for deterring people (mostly men) from buying sex or which sanction those individuals who solicit sex acts. The essential product of that study was the Demand Forum website, launched in January 2013 by Abt Associates. In the years that followed, Demand Forum provided information about demand reduction interventions in the United States, and its content was updated and expanded through daily web searches and supplemented by periodic literature reviews or direct contact with a network of practitioners and other experts. During the website's first seven years of operation, it was viewed by more than 262,000 individuals from 179 countries and was used to shape policy and practice within the United States. However, innovations in the field, primarily new tactics using information technology (IT) to deter buyers and develop evidence to apprehend those actively seeking to purchase sex, have emerged since Demand Forum's launch.

The current study (referred to as Part II) builds upon the methodology and knowledge base of the initial study to keep the field informed of innovations and evolving responses to buyer behaviors and to continue to provide support for practice and policy. Beginning January 2021, the National Center on Sexual Exploitation (NCOSE), which now maintains Demand Forum, conducted a systematic assessment of current demand reduction tactics and created an expanded tactic typology to reflect recent innovations intended to reduce the demand that drives sex trafficking markets. The project also aimed to provide updated information and resources that could be used by practitioners. The methodology for identifying new information about existing tactics and their implementation in U.S. cities and counties featured a web-based survey distributed to more than 3,200 law enforcement agencies, more than 50 interviews with expert practitioners and survivors, searches of thousands of open source reports, reviews of the research and practice literature, and reviews of prostitution laws within all 50 states.

NOTE: Data collected from the survey and interviews during this project were intended to verify information about demand reduction tactics and were not meant for analysis. This collection is limited to the online survey data.

Self-published

A pplication of Machine Learning Approaches to D evelop P redictive M odels for Diabetes and Hypertension among Bangladesh Adults (ICPSR 241544)

Released/updated on: 2026-01-12
AbstractIntroduction:With rapidurbanization, lifestyle changes, and an aging population, non-communicablediseases (NCDs),includinghypertension and diabetes,pose significant public health challenges inBangladesh andmanyotherlow-and middle-incomecountries. This studyusedmachine learning (ML)approachesto develop predictive models forhypertension and diabetesamong adultsin this country.Methods:BangladeshDemographic andHealthSurvey2022datawere analyzed. This isa nationallyrepresentative cross-sectional survey.Participantre classified as hypertensivewhen theirsystolic bloodpressurewas≥140mmHg, diastolicblood pressure was≥90mmHg, orif they usedantihypertensivemedication.They were classified asdiabetic if theirfasting plasma glucosewas≥7.0mmol/L orthey usedglucose-lowering drugs. Potential predictors included age, gender, education, wealth quintile,overweight/obesity,rural-urbanresidence, and divisionof residence.Descriptiveanalysis was conducted,andsix ML modelswere applied: artificialneuralnetwork (ANN),randomforest,adaptive boosting(AdaBoost),gradientboosting, XGBoost, andsupportvectormachine (SVM). Models’performancewasevaluated via accuracy,area under the curve (AUC), sensitivity, specificity, and F1-score. Featureimportance was assessed to rankrisk factors.Results:The study included13,847 adults, 55% of whom were females.Sensitivity was high across models(up to 0.96 for diabetes and 0.90 for hypertension). However, the overall specificity was low, particularlyfor diabetes (as low as 0.13 in XGBoost).Diabetes and hypertension had prevalence of 16.3% and 20.5%,respectively.The prevalence of both conditionsincreasedwith age, and the highest prevalencewas24.4%for diabetes and 43.3% for hypertensionamongindividuals aged 65 and older. Wealthier and urban residentsexperienced higher rates (diabetes: 24.9% among the richest compared to 9.9% among the poorest;hypertension: 23.3% in urban versus 19.2% in rural areas). Additionally, overweight/obesity was a strongpredictor for both conditions.For diabetes, AdaBoosthadthe highest AUC (0.699) and SVMhadthehighest accuracy (0.836); for hypertension, AdaBoosthad the greatestAUC (0.775) and accuracy (0.799).Hypertension topped diabetes predictors, while overweight/obesitywas the top predictorfor hypertension,followed by age and diabetes. Wealth and gender were moderately influential, with education andgeographic factors less so. Low specificity across models indicated challenges in identifying non-cases.Conclusion:This ML-driven analysisidentifiedthe bidirectional relationship ofhypertensionanddiabetesalong with several other predictors, includingoverweight/obesity,older age, and richer household wealthquintiles.Ourfindings underscore the need for integrated screening and lifestyle interventions targetinghigh-risk groups to mitigatefutureNCD burden.
Curated
Simple Crosstabs

Work in America Survey, [United States], 2022-2024 (ICPSR 39280)

Released/updated on: 2026-01-05
Geographic coverage: United States
Time period: 2022-01-01--2024-12-31
The Work in America survey was commissioned by the American Psychological Association and conducted online in the United States by The Harris Poll. The survey questions measure psychological safety, positive experiences, negative workplace outcomes, self-measures of performance and productivity, and workplace practices, policies, and programs.
Self-published

ChatGPT in education: A discourse analysis of worries and concerns on social media (ICPSR 300429)

Released/updated on: 2025-12-18
Time period: 2022-12-01--2023-03-31

These data are restricted and require an application. To apply, see SOMAR’s Application Portal and Application Guide.

The rapid advancements in generative AI models present new opportunities in the education sector. However, it is imperative to acknowledge and address the potential risks and concerns that may arise with their use. We collected Twitter data to identify key concerns related to the use of ChatGPT in education. This dataset is used to support the study "ChatGPT in education: A discourse analysis of worries and concerns on social media."

In this study, we particularly explored two research questions. RQ1 (Concerns): What are the key concerns that Twitter users perceive with using ChatGPT in education? RQ2 (Accounts): Which accounts are implicated in the discussion of these concerns? In summary, our study underscores the importance of responsible and ethical use of AI in education and highlights the need for collaboration among stakeholders to regulate AI policy.

Self-published

X (Twitter) IDs and LLM-Generated Analyses for Economic Narratives: Datasets for Pre-pandemic (2007-2020) and Post-LLM training cutoff (2021-2023) (ICPSR 300498)

Released/updated on: 2025-12-17
Time period: 2007-01-01--2023-07-31

This research comprises two distinct collections of economy-related posts from the X (formerly Twitter) platform – one spanning 2007-2020 (pre-pandemic) and the other 2021-2023 (post LLM training cutoff) – alongside corresponding LLM-generated analyses of the 2021-2023 posts. These collections, curated using targeted keywords, along with the LLM analyses, are provided to facilitate investigations into the potential of economic narratives and their influence. For more information about the data collection methodology, please refer to the associated paper.

The data provided here are the post (tweet) IDs for the pre-pandemic dataset and the LLM-generated analyses for the pre-pandemic data collection. The post-LLM training cutoff data collection could not be shared due to platform data sharing restrictions.

Self-published

Casual Conversations v2 Dataset (ICPSR 300465)

Released/updated on: 2025-12-15
Geographic coverage: Vietnam, United States, Philippines, Brazil, Mexico, India, Indonesia
Time period: 2023-01-01--2023-12-31

Casual Conversations v2 is composed of over 5,567 participants (26,467 videos) and intended mainly to be used for assessing the performance of already trained models in computer vision and audio applications for the purposes permitted in our data license agreement. The videos feature paid individuals who agreed to participate in the project and explicitly provided Age, Gender, Language/Dialect, Geo-location, Disability, Physical adornments, Physical attributes labels themselves. The videos were recorded in Brazil, India, Indonesia, Mexico, Philippines, United States, and Vietnam with a diverse set of adults in various categories. A group of trained annotators labeled the participants’ apparent skin tone using the Fitzpatrick scale and Monk Scale, in addition to annotations of Voice timbre, Activity and Recording setups. Spoken words in all videos are either scripted (a sample paragraph from The Idiot by Fyodor Dostoevsky provided with the dataset) or nonscripted (answering one of five predetermined questions).

Curated
Restricted

Using Physician Behavioral Big Data for High Precision Fraud Prediction and Detection, United States, 2000-2019 (ICPSR 38811)

Released/updated on: 2025-12-02
Geographic coverage: United States
Time period: 2000-01-01--2020-12-31
This project used big data from non-clinical physician behavior. These include traffic violations, substance abuse, property ownership, stressors (e.g., bankruptcy and divorce), social media data, and other life events data. These variables, all based on public records, were used to construct a predictive model of Medicare fraud using machine learning techniques.
Curated

Methods for Heterogeneity of Treatment Effects: Random Forest Counterfactual Machines [Methods Study], Cleveland, Ohio, 2014-2019 (ICPSR 39559)

Released/updated on: 2025-11-24
Geographic coverage: United States, Ohio, Cleveland
Time period: 2014-01-01--2019-12-31

Patients may respond differently to the same treatment due to individual traits such as age or gender. Knowing how different traits can affect a patient's response to treatment can help doctors and patients make better treatment decisions. For example, this information can help doctors know what types of cancer medicines work better for certain patients. This project focuses on improving the methods that researchers use to compare how treatments work for different patients.

In this project, the research team developed and tested a statistical method called random forests, or RF. RF is a way to analyze data using a technique called machine learning. In machine learning, computers use data to learn how to perform different tasks with little or no human input. Many types of RF methods exist. The team compared multiple RF methods to learn how well the methods would work to find out how patients with different traits respond to the same treatment.

To access the R package, please visit the randomForestSRC CRAN webpage.

Self-published

Análisis cualitativo asistido por LLMs: Una metodología híbrida para el estudio territorial de la participación ciudadana (ICPSR 239202)

Released/updated on: 2025-10-26
Geographic coverage: Colombia
Time period: 2024-01-01--2024-12-31
Prompts para el artículo: "Análisis cualitativo asistido por LLMs: Una metodología híbrida para el estudio territorial de la participación ciudadana"Esta investigación desarrolla una metodología que integra modelos de lenguaje en análisis cualitativo, argumentando que es posible superar limitaciones de escalabilidad sin sacrificar rigor interpretativo. Aplicada al estudio de participación ciudadana en Colombia, combinó transcripción automática, análisis asistido por IA y validación humana. Los resultados mostraron alta eficiencia (99 entrevistas analizadas en una semana) manteniendo profundidad analítica para identificar patrones territoriales, confirmando el potencial de este enfoque híbrido para la investigación en ciencias sociales.
Curated

Semiparametric Causal Inference Methods for Adaptive Statistical Learning in Trauma Patient-Centered Outcomes Research [Methods Study], 2013-2018 (ICPSR 39471)

Released/updated on: 2025-08-26
Geographic coverage: United States
Time period: 2013-01-01--2018-12-31

Electronic health records store a lot of data about a patient. These data often include age, health problems, current medicines, and lab results. Looking at these data may help doctors treating patients after a trauma predict how likely it is that they will respond well to a treatment and survive. This information can help doctors make better treatment decisions. But first, researchers need to figure out how to combine and analyze data to make accurate predictions. In this study, the research team created new statistical methods to combine data from patient records. They used these methods to predict patient health outcomes. Then the team used health record data collected from patients in hospital trauma centers to test their predictions.

To access the methods and software, please visit the following GitHubs:

  • origami
  • varimpact
  • opttx
Curated
Simple Crosstabs

Ithaka S+R Instructor Survey, United States, 2024 (ICPSR 39221)

Released/updated on: 2025-08-12
Geographic coverage: United States
Time period: 2024-01-01--2024-12-31
The first cycle of the 2024 US Instructor Survey queried a random sample of faculty members and instructors in the United States to gain a better understanding of their attitudes, perceptions, and practices regarding teaching, learning, and instructional support at their respective campuses. This survey is a renewed adaptation of Ithaka S+R's triennial US Faculty Survey, fielded since 2000, with a special focus on instruction, diverse teaching, and learning modalities.
Self-published

Identification and Appraisal of AI-Generated vs. Human-Created Artworks (2024 Survey Dataset) (ICPSR 228723)

Released/updated on: 2025-07-01
Geographic coverage: United States
Time period: 2024-06-01--2024-06-10
This dataset supports the doctoral dissertation The AI of the Beholder: A Quantitative Study on Human Perception and Appraisal of AI-Generated Images by Joshua Cunningham (Robert Morris University, 2025). The study investigates how individuals perceive and appraise artwork generated by artificial intelligence (AI) in comparison to human-created pieces. Specifically, it examines: (1) whether participants can accurately distinguish AI-generated from human-created artwork, (2) how age and exposure to AI art influence this ability and related appraisals, and (3) how digital versus traditional visual styles of AI art are perceived. The dataset includes anonymized survey responses collected from a diverse group of adult participants. Respondents were asked to evaluate a series of visual artworks—some created by humans, others by AI—across a range of styles, including both digital and traditional aesthetics. Additional demographic information such as age and prior exposure to AI tools was collected to assess moderating effects. The data were analyzed using SPSS to evaluate participant accuracy, preferences, and perceptions. This dataset can support further research into the psychological, aesthetic, and cultural dynamics of AI-generated content, as well as human-machine interaction in the creative arts.You may find the images used in this study, the original survey instrument, as well as a legend detailing each of the variables here: https://drive.google.com/drive/folders/125oaW82HpJUjI7EbaQWk_51puz0DgWKc?usp=sharing
Self-published

Public Perceptions of AI in Healthcare Based on Reddit Posts (2020–2025) (ICPSR 231282)

Released/updated on: 2025-05-29
Geographic coverage: Earth
Time period: 2020-03-01--2025-03-31
This dataset contains 18,754 Reddit posts and comments related to artificial intelligence (AI) in healthcare, collected from March 2020 to March 2025. The dataset was used to analyze public sentiment and topic trends using BERTopic and sentiment classification models. The data includes original texts, sentiment labels, topic assignments, and c-TF-IDF keyword weights.
Curated
Simple Crosstabs

Broadening the Reach, Impact, and Delivery of Genetic Services (BRIDGE) Chatbot or Standard of Care Trial for Genetic Cancer Counseling, New York and Utah, 2020-2023 (ICPSR 39256)

Released/updated on: 2025-02-20
Geographic coverage: United States, New York (state), Utah
Time period: 2020-01-01--2023-12-31

The Broadening the Reach, Impact, and Delivery of Genetic Services (BRIDGE) randomized controlled trial included 3,073 eligible patients between 2020-2023. The trial examined whether chatbot and standard of care approaches are equivalent in completion of pre-test cancer genetic services and genetic testing.

Self-published

AI Libraries Outreach (ICPSR 209346)

Released/updated on: 2024-09-27
The data includes the response for prompts from Claude 2 and ChatGPT 4 on library outreach ideas for specific populations.
Curated
Simple Crosstabs

Applying Artificial Intelligence to Person-Based Policing Practices, 2019-2023 (ICPSR 39074)

Released/updated on: 2024-09-26
Time period: 2019-01-01--2023-12-31
In this project, the research team developed and evaluated an artificial intelligence (AI) tool using agent-based modeling methods for crime analysis and risk evaluation (CARE): CAREsim. The purpose of this tool was to improve the effectiveness of person-based patrol strategies, where police take preemptive actions upon selected high-risk individuals (determined based on factors known to police such as violent crime history) when predicted risks of committing crimes are high. CARESim was developed and tested with a simulated randomized controlled experiment within the jurisdiction of Hampton, Virginia. 240 high-risk individuals (120 in each group) were followed for a 12-month period, with the simulation lasting 23 months. The treatment group received additional crime analyses using the AI tool and more focused patrols, while the control group received analyses as usual and random patrols in the simulated environment. The tool was evaluated on a series of outcomes (e.g., number of crimes and arrests) comparing the control and treatment groups. This collection contains the simulated high-risk individual data (DS1) and the simulated crimes data (DS2) used for the experiment.
Self-published

Bridging the Gap for ALICE: Charitable Organizations Acting Amid Rising Inflation and Further Solutions Using AI (ICPSR 208441)

Released/updated on: 2024-08-11
Geographic coverage: United States
Time period: 2013-01-01--2021-12-31
Abstract — Nationwide, nearly 37.9 million U.S. households fall into the Asset Limited, Income Constrained, Employed (ALICE) category. These households earn above the Federal Poverty Line but below the ALICE Household Survival Budget, making them ineligible for many public assistance programs. Nonprofit organizations like United Way and Feeding America have developed tailored assistance programs to support ALICE households. This study utilizes data from ALICE reports, U.S. Census data, and tools such as the U.S. Bureau of Labor Statistics Data Retrieval Tools and the Federal Reserve Bank of Atlanta’s Policy Rules Database to analyze the challenges faced by ALICE households. Additionally, data from the Integrated Public Use Microdata Series - Current Population Survey (IPUMS CPS) is examined using Microsoft Excel’s Pivottable to quantify the impact of nonprofit programs through the Household Rasch Food Security Score (FSRASCH) metric. The analysis highlights a significant gap in the number of ALICE households receiving assistance, attributed to inadequate access to information about available programs. To address this, AskALICE, an AI chatbot, is proposed. Developed using Yellow.ai’s Orchestrator LLM, AskALICE provides an accessible, centralized solution to inform ALICE households about eligibility and available assistance programs. By leveraging various data tables, assistance program information, and machine learning capabilities, AskALICE aims to bridge the information gap and enhance economic stability for millions of ALICE households. This study underscores the importance of accessible information in improving support for the ALICE population and proposes a practical solution with significant implications for economic stability.
Self-published

Forecasting bilateral asylum seeker flows with high-dimensional data and machine learning techniques (ICPSR 198322)

Released/updated on: 2024-08-05
We develop monthly asylum seeker flow forecasting models for 157 origin countries to the EU27, using machine learning and high-dimensional data, including digital trace data from Google Trends. Comparing different models and forecasting horizons and validating out-of-sample, we find that an ensemble forecast combining Random Forest and Extreme Gradient Boosting algorithms outperforms the random walk over horizons between 3 and 12 months. For large corridors, this holds in a parsimonious model exclusively based on Google Trends variables, which has the advantage of near real-time availability. We provide practical recommendations how our approach can enable ahead-of-period asylum seeker flow forecasting applications.
Self-published

ECIN Replication Package for "Specialization Trends in Economics Research: A Large-Scale Study Using Natural Language Processing and Citation Analysis" (ICPSR 198921)

Released/updated on: 2024-08-02
Time period: 1970-01-01--2016-12-31
We conduct a comprehensive analysis of specialization trends within and across fields of economics research. We collect data on 24,273 articles published between 1970 and 2016 in general research economics outlets and employ machine learning techniques to enrich the collected data. Results indicate that theory and econometric methods papers are becoming increasingly specialized, with a narrowing scope and steady or declining citations from outside economics and from other fields of economics research. Conversely, applied papers are covering a broader range of topics, receiving more extramural citations from fields like medicine, and psychology. Trends in applied theory articles are unclear.
Self-published

Face Sketches and Personality (ICPSR 207764)

Released/updated on: 2024-07-08
Geographic coverage: China
Time period: 2022-01-01--2024-12-31
Applying Machine Learning, we are looking for a number of elusive lincks between appearance and behavior. The dataset obtained after the moderation of images and validation of questionnaires. The test dataset consisted of twenty personality traits of 220 male and 160 female sketches self-reported and predicted by ML algorithms from face sketches.
External data

AI Enabled Community Supervision for Criminal Justice Services, 2020-2023 (ICPSR 38996)

Released/updated on: 2023-12-20

This project aimed to revolutionize the reentry process for justice-involved individuals (JII) by harnessing the power of artificial intelligence (AI) and advanced technologies. The centerpiece of the endeavor is the AI-based Support and Monitoring System, or AI-SMS, a cutting-edge platform designed to assist JII and their dedicated caseworkers in their journey to reintegrate seamlessly into the community. While the primary focus is on JII, the researchers recognize the critical role played by caseworkers-clinically trained individuals who facilitate the reentry process from a community perspective.

AI-SMS was conceived to be a multifaceted tool that provides case workers with early warning indicators of risky behavior and equips JII with the means and strategies to mitigate these risks, aligning with best practices in hybrid supervision. At its core, the system is committed to delivering personalized resources and opportunities to JII, complementing the support offered by caseworkers.

Self-published

Attitudes Toward Artificial Intelligence–Enabled Mental Health Tools Among Prospective Psychotherapists (ICPSR 195822)

Released/updated on: 2023-12-14
Geographic coverage: Canada, United States, United Kingdom, Germany
Time period: 2022-11-02--2023-03-12
Background: Recent efforts to make artificial intelligence (AI) applications in clinical care more user-friendly have faced adoption barriers. There is a lack of research on the application of AI systems in mental health care specifically.
Objective: In an attempt to fill this research gap, this study focuses on factors influencing the likelihood of psychology students and early practitioners adopting two specific AI-enabled mental health tools. These tools have been evaluated with reference to the Unified Theory of Acceptance and Use of Technology.
Methods: A cross-sectional study with a sample size of 206 psychology students and trainee psychotherapists was undertaken. The participants' openness to using two AI tools was evaluated. The first tool provides feedback to therapists on their motivational interviewing techniques and the second one uses patient voice samples to provide mood scores for better treatment decisions. Participants were presented with visual explanations of each tool's functions before their responses were measured based on the Unified Theory of Acceptance and Use of Technology. Two structural equation models were created to predict tool use intentions.
Results: Both perceived usefulness and social influence had a positive impact on participants' willingness to use both tools. However, users' trust in the technology did not appear to influence their choice to use it. Interestingly, perceived ease of use was found to be unrelated or even negatively related to use intentions. Additionally, a positive correlation was found between readiness to adopt technology and the intention to use the feedback tool while AI anxiety had a negative correlation with the use intention for both tools.
Conclusions: The data sheds light on both generic and tool-specific drivers of AI technology adoption in the mental health care field. Further study is encouraged to better understand the technological and user group dynamics impacting the adoption of AI in this area.
Curated

Eurobarometer 87.1: Two Years Until the 2019 European Elections, Attitudes of Europeans Towards Tobacco and Electronic Cigarettes, Climate Change, Attitudes Towards the Impact of Digitization and Automation on Daily Life, and Coach Services, March 2017 (ICPSR 38334)

Released/updated on: 2022-06-13
Geographic coverage: Cyprus, Portugal, Malta, Greece, Netherlands, Sweden, Great Britain, Austria, Latvia, Luxembourg, Ireland, Poland, Slovenia, Slovakia, France, Bulgaria, Lithuania, Croatia, Romania, Hungary, Northern Ireland, Spain, Czech Republic, Belgium, Finland, Denmark, Italy, Germany, Estonia
Time period: 2017-01-01--2017-12-31

The Eurobarometer series is a unique cross-national and cross-temporal survey program conducted on behalf of the European Commission. These surveys regularly monitor public opinion in the European Union (EU) member countries and consist of standard modules and special topic modules. The standard modules address attitudes towards European unification, institutions and policies, measurements for general socio-political orientations, as well as respondent and household demographics. The special topic modules address such topics as agriculture, education, natural environment and resources, public health, public safety and crime, and science and technology.

Eurobarometer 87.1 covered the following special topics: two years until the 2019 European elections, tobacco and electronic cigarettes, climate change, the impact of digitization and automation on daily life, and coach services. Questions regarding the European elections in 2019 included information on and the role of the European Parliament (EP), the knowledge about European institutions and the EP, the present and future of the EP, European values and policies, European identity, and media use. Further questions were asked regarding smoking habits and various tobacco/nicotine products. Respondents were queried about their smoking habits, their efforts to quit smoking, passive smoking inside, and banning advertisements for tobacco products. Respondent's opinions were collected on which world issues they believed were the most serious problems, how serious the issue of climate change was and if the EU should be responsible for addressing it, and what actions they have personally taken to fight climate change. Respondents were also asked their awareness of, usage of, and attitude towards autonomous systems including robots, artificial intelligence, driverless cars, civil drones, cyber security, online social networks, and health services. Lastly, several questions were asked regarding the usage frequency of coach services, the reasons for using coaches, and the appraisal of services.

Demographic and other background information collected includes left or right self-placement on political scale, age, gender, nationality, marital status, occupation, age when stopped full-time education, household composition, ownership of a fixed or mobile telephone and other goods, difficulties in paying bills, self-assessed social class, internet use, life satisfaction, political discussion frequency, and opinions on whether their voice counts in their country/EU. Country-specific data includes type and size of locality, region of residence, and language of interview (select countries).

Self-published

Student emotions, bad professors, and course ratings. A machine learning evaluation of one million student reviews. (ICPSR 163141)

Released/updated on: 2022-02-22
Geographic coverage: United States
Time period: 1999-01-01--2020-12-31
For the first time in the literature, Natural Language Processing models based on deep neural networks are used to identify student emotions and extract bad performance by professors from the course reviews of students. I study how emotions and harmful performance are associated with high and low faculty ratings, the perceived level of difficulty of the course, and the university rating. I use a random sample of nearly one million student reviews from the ratemyprofessors.com website and the performance indicators of US universities from the Times Higher Education ranking. The linear and ordinal logistic regression models reveal strong relationships, and the estimated parameters have the expected signs.   
Description of variables in the Results replication file, in Rdata format. 975,860 observations (rows). Raw data scraped from ratemyprofessors.com
- rev - textual student review
- rat - course rating (1-5)
- dif - course perceived difficulty (1-5)
- uname - University name
- emo - emotion with the highest score
- action_zs - professor bad performance with the highest probability calculated by zero-shot classification (DistilBERT model available at https://huggingface.co/typeform/distilbert-base-uncased-mnli)
- sadness_sc - sadness emotion score calculated by DistilBERT model trained on Twitter corpus (available at https://huggingface.co/bhadresh-savani/distilbert-base-uncased-emotion)
- similar for the other five basic emotions
- nwords - number of words in the review
- m, y - month and year when the review was written
Self-published

Promises and pitfalls of using computer vision to make inferences about landscape preferences: Evidence from an urban-proximate park system (data and code) (ICPSR 139681)

Released/updated on: 2021-10-31
Geographic coverage: Boulder, Colorado, United States
Time period: 2004-01-01--2018-12-31
We compare preferences for landscape features derived through a computer vision algorithm (Google Cloud Vision) used to analyze social media photographs with preferences derived through a traditional on-site intercept survey. We surveyed visitors in Boulder Open Space and Mountain Parks lands in Colorado (USA) in May and June, 2018. We downloaded all Flickr photographs within Boulder Open Space and Mountain Parks lands from 2004 - 2018, and ran the photographs through Google Cloud Vision to get up to 10 labels for each image. We compare the content in Flickr photographs to the features that visitors say positively impacted their experience on surveys.This paper is currently under review.Contents of this repository:GoogleVision: Contains raw data exported from Google Vision, as well as a codebook for how we coded each label to match the landscape categories we asked about in the survey. Also contains a full database connecting the Google Vision labels and presence/absence of each feature to the Flickr data.Code:Contains one R script that makes the maps and runs spatial cluster analysis, and one R script that does all the data cleaning and analysis for the Flickr and survey data. You will need to download the contents in the GoogleVision, shapefiles, and survey folders to run this code. This folder also contains a Python script that we used to download Flickr data within Boulder through the Flickr API, and another R script used only to generate table E.1 in the supplementary material.Shapefiles: Contains all the spatial data needed to reproduce maps and run R code. This includes OSMP trails, trailheads, lands, landscape character areas, survey locations, and coordinates of Flickr points.Survey:Contains the data from a visitor survey in Boulder OSMP lands from May and June 2018 (in a CSV), as well as a codebook to interpret the data, and the survey instrument.
Self-published

Public Understanding of Artificial Intelligence through Entertainment Media (ICPSR 152542)

Released/updated on: 2021-10-15
Our data was collected for a project on how media representations shape public perceptions of AI and then use what we learn to explore how we might better represent everyday interactions with AI to the public. We began by developing lists, taxonomies, and case studies of popular representations of AI and hypotheses about how the general public and elite groups discriminate between “good” and “bad” AI, and about how media representations help shape these perceptions. Our research questions include:
  • What drives popular negativity about AI? Does the public have its reasons that the experts know not of? Or have the public adopted views of AI borne of misrepresentations?
  • Do overtly dystopian representations of AI feed, or perhaps temper, public outrage about insidious issues with AI and machine learning such as the biases of search algorithms?
  • Can we produce narratives on AI in different media modalities that are more nuanced and complex than just the false dichotomy of good and bad? Could a more accurately critical (and yet still exciting) model of AI themed entertainment be developed, once we’ve gained an understanding of how the public has been encountering AI?
Self-published

Health and the built environment in U.S. cities: Measuring associations using Google Street View-derived indicators of the built environment (ICPSR 115264)

Released/updated on: 2019-11-01
Geographic coverage: United States
Background: The built environment is a structural determinant of health and has been shown to influence health expenditures, behaviors, and outcomes. Traditional methods of assessing built environment characteristics are time-consuming and difficult to combine or compare. Google Street View (GSV) images represent a large, publicly available data source that can be used to create indicators of characteristics of the physical environment with machine learning techniques. 
Methods: We used computer vision techniques to derive built environment indicators from approximately 31 million GSV images at 7.8 million intersections. Associations between derived indicators and health behaviors and outcomes on the census-tract level were assessed using multivariate regression models, controlling for demographic factors and socioeconomic position. 
Results: Street greenness was associated with decreased prevalence of physical and mental distress, as well as decreased binge drinking, but with increased obesity. Single lane roads were associated with increased diabetes and obesity, while non-single-family home buildings were associated with decreased obesity, diabetes and inactivity. 
Conclusions: Structural determinants of health such as the built environment can influence population health. Our study suggests that higher levels of urban development have mixed effects on health and adds further evidence that socioeconomic distress has adverse impacts on multiple physical and mental health outcomes.
Self-published

SISA/SISAL Dataset (ICPSR 115165)

Released/updated on: 2019-10-25
Geographic coverage: Ecuador
Time period: 2013-11-20--2017-09-13
This dataset includes the symptoms and basic demographic information for subjects who had a diagnosis of suspected arboviral illness, including dengue, chikungunya, or Zika virus infection. They were recruited in Machala, Ecuador from 2013-2017. The collection details are available in: Stewart-Ibarra AM, Ryan SJ, Kenneson A, King CA, Abbott M, Barbachano-Guerrero A, et al. The Burden of Dengue Fever and Chikungunya in Southern Coastal Ecuador: Epidemiology, Clinical Presentation, and Phylogenetics from the First Two Years of a Prospective Study. Am J Trop Med Hyg. 2018;98: 1444–1459. doi:10.4269/ajtmh.17-0762There are two datasets available: the full set of subjects (SISA) and a subset of subjects with available laboratory data (SISAL).
Self-published

Face Shape and Personality (ICPSR 109082)

Released/updated on: 2019-03-27
Geographic coverage: Russia
Time period: 2016-01-01--2018-12-31
Applying Machine Learning, we are looking for a number of elusive lincks between appearance and behavior. The dataset obtained after the moderation of images and validation of questionnaires. The test (holdout) dataset included self-reported and predicted from face shape by a ML algorithm the ​​Big Five personality traits for  505 males  and 740 females. 
Back to top