False Positives in AI Writing Detection: A Small-Scale Empirical Study Using Authentic Filipino Student Essays (ICPSR 251455)
Intelligence Theory 2026: Supplementary Research Data and Documentation (ICPSR 251126)
A systematic discourse analysis of how U.S. political leaders frame disability: Implications for students with disabilities (ICPSR 307656)
This study examined how U.S. political leaders publicly framed disability during calendar year 2025, using a systematic discourse analysis of public statements and federal legislation. It also examined how this discourse was reflected in federal education policy, including statements by the Secretary of Education and education-related bills.
The study analyzed 121 public statements about people with disabilities made by White House and Cabinet officials, and 32 federal bills introduced in 2025 that could affect the rights, services, or educational opportunities of individuals with disabilities. Statements were identified through a two-stage process that combined a custom Python-based web-scraping tool, which extracted verbatim, attributed quotations from news articles, press releases, interview transcripts, official speeches, and social media posts using the OpenAI GPT-4 API, with manual verification searches conducted in ChatGPT Plus and Perplexity AI Pro. Federal bills were identified through Congress.gov. All statements were reviewed by the authors to confirm accuracy, attribution, and date. Each statement and bill was independently scored by two human coders and by GPT-4 using an author-developed four-point rubric grounded in the social and human rights models of disability, ranging from 1 (dehumanizing) to 4 (affirming).
The data contain one record per statement, including the speaker's name and title, date, verbatim quotation, source, context, an analysis of the framing, a score from 1 (dehumanizing) to 4 (affirming), and the rationale for the score. The bills data are available on the project website and are not included in this deposit.
The study was approved by the university Institutional Review Board (March 2025) and was preregistered on the Open Science Framework (https://osf.io/hbfe5/).
AI-Supported Inquiry, AI Literacy, and Authentic Performance Among Preservice Teachers in Central China, 2024 (ICPSR 306435)
This study examined the effects of QUEST+AI, an AI-supported inquiry model, on AI literacy and authentic performance among preservice teachers in Central China. The study used a nonequivalent-groups, quasi-experimental pretest–posttest design with two intact sections of an undergraduate Educational Research Methods course at a public university. Ninety-five preservice teachers participated, including 52 students in the experimental group and 43 students in the comparison group. Both groups received the same face-to-face course instruction over a 10-week period. The experimental group completed two QUEST+AI inquiry cycles with coached use of generative artificial intelligence, while the comparison group completed conventional homework assignments.
The data include participant demographic variables, pretest and posttest responses to a multidimensional AI literacy questionnaire, AI literacy total and subscale scores, final research proposal scores, and group assignment indicators. AI literacy measures cover use and application of AI, knowledge and understanding of AI, AI detection, AI ethics, AI creation, AI-supported problem solving, AI persuasion literacy, and AI emotion regulation. Authentic performance is represented by scores on a final educational research proposal, evaluated with a common rubric by two independent raters. The dataset also includes variables used in the study’s comparative analyses, including group condition, pretest scores, gender, grade level, age, and major.
Using Topic Segmentation to Enhance Concept Parsing and Identification of Negations [Methods Study], Massachusetts, 2019-2023 (ICPSR 39740)
Clinical notes in electronic health records, or EHRs, may contain information that can help researchers study and compare treatments. But it takes researchers a lot of time to find information in EHR notes.
Natural language processing, or NLP, methods can help researchers find information in EHR notes. With NLP, computer programs read and identify written language to make it easier to sort and study. But in EHR notes, some sentences may contain more than one topic. Also, EHR notes may discuss a single topic over many sentences. In these cases, current NLP methods don't work well to find complete and accurate information about a specific topic.
In this study, the research team developed and tested new NLP methods to identify topics from EHR notes.
Development of Computational Methods for Evaluating Doctor-Patient Communication [Methods Study], United States, 2016-2021 (ICPSR 39720)
The way doctors communicate with patients during office visits can affect the quality of care. Studying conversations between doctors and patients can help doctors improve their communication skills.
To study conversations, researchers rely on written records, or transcripts, of office visits. They read the transcripts and give each conversation topic a label. For example, topics may include smoking or pain. But labeling topics in this way may take a lot of time.
In this project, the research team created and tested a new method to make this work easier using natural language processing, or NLP. With NLP, computer programs interpret written language. NLP methods use a process called machine learning, where computer programs use data to learn how to perform different tasks with little or no human input.
Unlocking Clinical Text in Electronic Medical Records (EMR) by Query Refinement Using Both Knowledge Bases and Word Embedding [Methods Study], Ohio, 2006-2022 (ICPSR 39734)
Electronic health records, or EHRs, have information about a patient's health such as test results, diagnoses, and treatments. EHRs also have clinical notes that doctors and patients can use to track goals and decisions.
Clinical notes may be useful for research or to help improve care. But it's hard to get information from these notes across large groups of patients. The notes may use different ways to describe the same thing. For example, high blood pressure may be called hypertension. Also, the notes may use abbreviations or have spelling mistakes.
In this project, the research team designed and built a search engine to make EHR notes easier to search and use for patient care and research.
Detroit Metro Area Communities Study (DMACS) Wave 22, Michigan, 2025 (ICPSR 39692)
The Detroit Metro Area Communities Study (DMACS) is a panel survey of Detroit residents aged 18 and older. The original panel of respondents was drawn from an address-based probability sample of all occupied Detroit households in 2016 and has since been refreshed through additional address-based sampling annually. Between August 6, 2025 and October 1, 2025, 3,170 previously enrolled panelists were invited to participate in a self-administered online or interviewer-administered telephone survey.
Topics included: household composition; housing status; perceptions of neighborhood; social connection and loneliness; election; mayoral priorities; crime and safety; violence reduction; artificial intelligence; flood management; mental health; employment.
ECIN Replication Package for "On the environmental monitoring of firms: Does the centralization of inspections matter?" (ICPSR 237487)
ChatGPT and Willingness to share data (ICPSR 244947)
National Assessment of Demand Reduction Efforts, Part II: New Developments in the Primary Prevention of Sex Trafficking, [United States], 2021 (ICPSR 38928)
To combat prostitution and sex trafficking, criminal justice strategies and collaborative programs have emerged that focus on reducing consumer-level demand. From 2008 to 2012, the National Institute of Justice (NIJ) sponsored a study entitled "A National Overview of Prostitution and Sex Trafficking Demand Reduction Efforts" (referred to as Part I) that featured the systematic collection of information to determine the types and distribution of demand reduction tactics implemented throughout the United States. These efforts gave rise to a typology of law enforcement and community-based tactics identifying 12 different methods for deterring people (mostly men) from buying sex or which sanction those individuals who solicit sex acts. The essential product of that study was the Demand Forum website, launched in January 2013 by Abt Associates. In the years that followed, Demand Forum provided information about demand reduction interventions in the United States, and its content was updated and expanded through daily web searches and supplemented by periodic literature reviews or direct contact with a network of practitioners and other experts. During the website's first seven years of operation, it was viewed by more than 262,000 individuals from 179 countries and was used to shape policy and practice within the United States. However, innovations in the field, primarily new tactics using information technology (IT) to deter buyers and develop evidence to apprehend those actively seeking to purchase sex, have emerged since Demand Forum's launch.
The current study (referred to as Part II) builds upon the methodology and knowledge base of the initial study to keep the field informed of innovations and evolving responses to buyer behaviors and to continue to provide support for practice and policy. Beginning January 2021, the National Center on Sexual Exploitation (NCOSE), which now maintains Demand Forum, conducted a systematic assessment of current demand reduction tactics and created an expanded tactic typology to reflect recent innovations intended to reduce the demand that drives sex trafficking markets. The project also aimed to provide updated information and resources that could be used by practitioners. The methodology for identifying new information about existing tactics and their implementation in U.S. cities and counties featured a web-based survey distributed to more than 3,200 law enforcement agencies, more than 50 interviews with expert practitioners and survivors, searches of thousands of open source reports, reviews of the research and practice literature, and reviews of prostitution laws within all 50 states.
NOTE: Data collected from the survey and interviews during this project were intended to verify information about demand reduction tactics and were not meant for analysis. This collection is limited to the online survey data.
A pplication of Machine Learning Approaches to D evelop P redictive M odels for Diabetes and Hypertension among Bangladesh Adults (ICPSR 241544)
Work in America Survey, [United States], 2022-2024 (ICPSR 39280)
ChatGPT in education: A discourse analysis of worries and concerns on social media (ICPSR 300429)
These data are restricted and require an application. To apply, see SOMAR’s Application Portal and Application Guide.
The rapid advancements in generative AI models present new opportunities in the education sector. However, it is imperative to acknowledge and address the potential risks and concerns that may arise with their use. We collected Twitter data to identify key concerns related to the use of ChatGPT in education. This dataset is used to support the study "ChatGPT in education: A discourse analysis of worries and concerns on social media."
In this study, we particularly explored two research questions. RQ1 (Concerns): What are the key concerns that Twitter users perceive with using ChatGPT in education? RQ2 (Accounts): Which accounts are implicated in the discussion of these concerns? In summary, our study underscores the importance of responsible and ethical use of AI in education and highlights the need for collaboration among stakeholders to regulate AI policy.
X (Twitter) IDs and LLM-Generated Analyses for Economic Narratives: Datasets for Pre-pandemic (2007-2020) and Post-LLM training cutoff (2021-2023) (ICPSR 300498)
This research comprises two distinct collections of economy-related posts from the X (formerly Twitter) platform – one spanning 2007-2020 (pre-pandemic) and the other 2021-2023 (post LLM training cutoff) – alongside corresponding LLM-generated analyses of the 2021-2023 posts. These collections, curated using targeted keywords, along with the LLM analyses, are provided to facilitate investigations into the potential of economic narratives and their influence. For more information about the data collection methodology, please refer to the associated paper.
The data provided here are the post (tweet) IDs for the pre-pandemic dataset and the LLM-generated analyses for the pre-pandemic data collection. The post-LLM training cutoff data collection could not be shared due to platform data sharing restrictions.
Casual Conversations v2 Dataset (ICPSR 300465)
Casual Conversations v2 is composed of over 5,567 participants (26,467 videos) and intended mainly to be used for assessing the performance of already trained models in computer vision and audio applications for the purposes permitted in our data license agreement. The videos feature paid individuals who agreed to participate in the project and explicitly provided Age, Gender, Language/Dialect, Geo-location, Disability, Physical adornments, Physical attributes labels themselves. The videos were recorded in Brazil, India, Indonesia, Mexico, Philippines, United States, and Vietnam with a diverse set of adults in various categories. A group of trained annotators labeled the participants’ apparent skin tone using the Fitzpatrick scale and Monk Scale, in addition to annotations of Voice timbre, Activity and Recording setups. Spoken words in all videos are either scripted (a sample paragraph from The Idiot by Fyodor Dostoevsky provided with the dataset) or nonscripted (answering one of five predetermined questions).
Using Physician Behavioral Big Data for High Precision Fraud Prediction and Detection, United States, 2000-2019 (ICPSR 38811)
Methods for Heterogeneity of Treatment Effects: Random Forest Counterfactual Machines [Methods Study], Cleveland, Ohio, 2014-2019 (ICPSR 39559)
Patients may respond differently to the same treatment due to individual traits such as age or gender. Knowing how different traits can affect a patient's response to treatment can help doctors and patients make better treatment decisions. For example, this information can help doctors know what types of cancer medicines work better for certain patients. This project focuses on improving the methods that researchers use to compare how treatments work for different patients.
In this project, the research team developed and tested a statistical method called random forests, or RF. RF is a way to analyze data using a technique called machine learning. In machine learning, computers use data to learn how to perform different tasks with little or no human input. Many types of RF methods exist. The team compared multiple RF methods to learn how well the methods would work to find out how patients with different traits respond to the same treatment.
To access the R package, please visit the randomForestSRC CRAN webpage.
Análisis cualitativo asistido por LLMs: Una metodología híbrida para el estudio territorial de la participación ciudadana (ICPSR 239202)
Semiparametric Causal Inference Methods for Adaptive Statistical Learning in Trauma Patient-Centered Outcomes Research [Methods Study], 2013-2018 (ICPSR 39471)
Electronic health records store a lot of data about a patient. These data often include age, health problems, current medicines, and lab results. Looking at these data may help doctors treating patients after a trauma predict how likely it is that they will respond well to a treatment and survive. This information can help doctors make better treatment decisions. But first, researchers need to figure out how to combine and analyze data to make accurate predictions. In this study, the research team created new statistical methods to combine data from patient records. They used these methods to predict patient health outcomes. Then the team used health record data collected from patients in hospital trauma centers to test their predictions.
To access the methods and software, please visit the following GitHubs:
- origami
- varimpact
- opttx
Ithaka S+R Instructor Survey, United States, 2024 (ICPSR 39221)
Identification and Appraisal of AI-Generated vs. Human-Created Artworks (2024 Survey Dataset) (ICPSR 228723)
Public Perceptions of AI in Healthcare Based on Reddit Posts (2020–2025) (ICPSR 231282)
Broadening the Reach, Impact, and Delivery of Genetic Services (BRIDGE) Chatbot or Standard of Care Trial for Genetic Cancer Counseling, New York and Utah, 2020-2023 (ICPSR 39256)
The Broadening the Reach, Impact, and Delivery of Genetic Services (BRIDGE) randomized controlled trial included 3,073 eligible patients between 2020-2023. The trial examined whether chatbot and standard of care approaches are equivalent in completion of pre-test cancer genetic services and genetic testing.
AI Libraries Outreach (ICPSR 209346)
Applying Artificial Intelligence to Person-Based Policing Practices, 2019-2023 (ICPSR 39074)
Bridging the Gap for ALICE: Charitable Organizations Acting Amid Rising Inflation and Further Solutions Using AI (ICPSR 208441)
Forecasting bilateral asylum seeker flows with high-dimensional data and machine learning techniques (ICPSR 198322)
ECIN Replication Package for "Specialization Trends in Economics Research: A Large-Scale Study Using Natural Language Processing and Citation Analysis" (ICPSR 198921)
Face Sketches and Personality (ICPSR 207764)
AI Enabled Community Supervision for Criminal Justice Services, 2020-2023 (ICPSR 38996)
This project aimed to revolutionize the reentry process for justice-involved individuals (JII) by harnessing the power of artificial intelligence (AI) and advanced technologies. The centerpiece of the endeavor is the AI-based Support and Monitoring System, or AI-SMS, a cutting-edge platform designed to assist JII and their dedicated caseworkers in their journey to reintegrate seamlessly into the community. While the primary focus is on JII, the researchers recognize the critical role played by caseworkers-clinically trained individuals who facilitate the reentry process from a community perspective.
AI-SMS was conceived to be a multifaceted tool that provides case workers with early warning indicators of risky behavior and equips JII with the means and strategies to mitigate these risks, aligning with best practices in hybrid supervision. At its core, the system is committed to delivering personalized resources and opportunities to JII, complementing the support offered by caseworkers.
Attitudes Toward Artificial Intelligence–Enabled Mental Health Tools Among Prospective Psychotherapists (ICPSR 195822)
Eurobarometer 87.1: Two Years Until the 2019 European Elections, Attitudes of Europeans Towards Tobacco and Electronic Cigarettes, Climate Change, Attitudes Towards the Impact of Digitization and Automation on Daily Life, and Coach Services, March 2017 (ICPSR 38334)
The Eurobarometer series is a unique cross-national and cross-temporal survey program conducted on behalf of the European Commission. These surveys regularly monitor public opinion in the European Union (EU) member countries and consist of standard modules and special topic modules. The standard modules address attitudes towards European unification, institutions and policies, measurements for general socio-political orientations, as well as respondent and household demographics. The special topic modules address such topics as agriculture, education, natural environment and resources, public health, public safety and crime, and science and technology.
Eurobarometer 87.1 covered the following special topics: two years until the 2019 European elections, tobacco and electronic cigarettes, climate change, the impact of digitization and automation on daily life, and coach services. Questions regarding the European elections in 2019 included information on and the role of the European Parliament (EP), the knowledge about European institutions and the EP, the present and future of the EP, European values and policies, European identity, and media use. Further questions were asked regarding smoking habits and various tobacco/nicotine products. Respondents were queried about their smoking habits, their efforts to quit smoking, passive smoking inside, and banning advertisements for tobacco products. Respondent's opinions were collected on which world issues they believed were the most serious problems, how serious the issue of climate change was and if the EU should be responsible for addressing it, and what actions they have personally taken to fight climate change. Respondents were also asked their awareness of, usage of, and attitude towards autonomous systems including robots, artificial intelligence, driverless cars, civil drones, cyber security, online social networks, and health services. Lastly, several questions were asked regarding the usage frequency of coach services, the reasons for using coaches, and the appraisal of services.
Demographic and other background information collected includes left or right self-placement on political scale, age, gender, nationality, marital status, occupation, age when stopped full-time education, household composition, ownership of a fixed or mobile telephone and other goods, difficulties in paying bills, self-assessed social class, internet use, life satisfaction, political discussion frequency, and opinions on whether their voice counts in their country/EU. Country-specific data includes type and size of locality, region of residence, and language of interview (select countries).
Student emotions, bad professors, and course ratings. A machine learning evaluation of one million student reviews. (ICPSR 163141)
Promises and pitfalls of using computer vision to make inferences about landscape preferences: Evidence from an urban-proximate park system (data and code) (ICPSR 139681)
Public Understanding of Artificial Intelligence through Entertainment Media (ICPSR 152542)
- What drives popular negativity about AI? Does the public have its reasons that the experts know not of? Or have the public adopted views of AI borne of misrepresentations?
- Do overtly dystopian representations of AI feed, or perhaps temper, public outrage about insidious issues with AI and machine learning such as the biases of search algorithms?
- Can we produce narratives on AI in different media modalities that are more nuanced and complex than just the false dichotomy of good and bad? Could a more accurately critical (and yet still exciting) model of AI themed entertainment be developed, once we’ve gained an understanding of how the public has been encountering AI?