Adaptive Alignment: Designing AI for a Changing World - Frauke Kreuter

Frauke Kreuter

International Conference on Machine Learning 2025 · Invited Talk

Overview

In an insightful talk at ICML 2025, social scientist Frauke Kreuter presented a compelling argument for a more nuanced and data-centric approach to AI alignment, emphasizing the critical importance of understanding what models are being aligned with. While the machine learning community often focuses on the technical mechanisms of alignment—how to make models behave in desired ways—Kreuter zoomed out to explore the underlying societal values, norms, and preferences that should inform these technical efforts. Her presentation highlighted the inherent difficulties in capturing these dynamic and diverse human elements, particularly across different subgroups and over time.

Watch on SlidesLive

Visual summary for Adaptive Alignment: Designing AI for a Changing World - Frauke Kreuter by Frauke Kreuter
Visual summary for Adaptive Alignment: Designing AI for a Changing World - Frauke Kreuter by Frauke Kreuter

Key moments

  1. 0:00 Introduction: Aligning AI with changing human values
  2. 2:00 Why alignment is difficult: Three research tasks
  3. 4:00 The challenge of description: Getting the denominator right
  4. 4:30 Role of statistical agencies in measuring society
  5. 6:00 Increasing difficulty and cost of data collection

Adaptive Alignment: Designing AI for a Changing World

Speakers: Frauke Kreuter

Conference: ICML 2025

YouTube: https://slideslive.com/39043347

Overview

In an insightful talk at ICML 2025, social scientist Frauke Kreuter presented a compelling argument for a more nuanced and data-centric approach to AI alignment, emphasizing the critical importance of understanding what models are being aligned with. While the machine learning community often focuses on the technical mechanisms of alignment—how to make models behave in desired ways—Kreuter zoomed out to explore the underlying societal values, norms, and preferences that should inform these technical efforts. Her presentation highlighted the inherent difficulties in capturing these dynamic and diverse human elements, particularly across different subgroups and over time.

Kreuter's core message underscored that human preferences are not static or monolithic, as evidenced by shifting societal views on topics like same-sex relationships or working mothers. Drawing on over 80 years of experience in survey research and survey methodology, she shared invaluable lessons on measuring human attitudes, opinions, and behaviors, emphasizing the meticulous effort required to collect high-quality, representative data. This talk served as a powerful call for interdisciplinary collaboration, urging the ML community to leverage existing social science expertise and data archives to build more robust, ethically aligned, and adaptively responsive AI systems that can navigate the complexities of a changing world.

The talk is particularly relevant for the AI/ML community as large language models (LLMs) are increasingly deployed in applications that require interaction with human values, from content moderation to generating synthetic data for policy decisions. Kreuter illuminated how current data collection practices in ML, such as annotation tasks for Reinforcement Learning with Human Feedback (RLHF), often overlook fundamental principles of good measurement and representation. By bridging the gap between social science and machine learning, Kreuter envisioned a future where AI systems are not only technically aligned but also deeply informed by a comprehensive, representative, and continuously updated understanding of human society.

Background

▶ Watch: Introduction: Aligning AI with changing human values (0:00)

The problem of AI alignment, particularly for large language models, has predominantly been framed as a technical challenge within the machine learning and systems space. Researchers often focus on developing algorithms and frameworks to ensure models adhere to predefined ethical guidelines, safety protocols, or user preferences. However, Frauke Kreuter argued that this technical focus often bypasses a more fundamental question: what precisely are we trying to align these models with? Human values, norms, and attitudes are neither universal nor static; they vary significantly across cultures, demographics, and evolve over time. For instance, questions about the morality of same-sex relationships or the societal role of working mothers have garnered vastly different answers historically and geographically, even within a single country like the US. Capturing this intricate and shifting landscape of human preferences is a profoundly difficult task.

This challenge is exacerbated by the increasing pressure on traditional statistical agencies, such as the US Census Bureau or the Bureau of Labor Statistics, to produce more timely and cost-effective data. These agencies historically invest immense resources—years of design, extensive fieldwork, and rigorous sampling—to produce high-quality, representative population statistics, like inflation numbers or rates of depression. However, the cost and difficulty of traditional data collection have risen dramatically. Kreuter presented data showing that survey response rates have plummeted, with some well-known telephone surveys now seeing single-digit response rates, compared to 70% or more a decade ago for official statistics. This decline threatens the precision and representativeness of crucial public data.

In response to these pressures, the ML community has been increasingly called upon to assist, offering solutions like predictive models for survey non-response, automated classification, imputation methods, and even AI-driven interviewers. More controversially, there's a growing interest in generating entirely synthetic data using large language models, raising questions about the validity and representativeness of such approaches. Kreuter highlighted that while ML excels at prediction and causation tasks—which often only require instances of phenomena and relevant covariates—it struggles with description tasks. Descriptive statistics, such as determining the prevalence of depression in a population, are uniquely difficult because they require accurately knowing the sampling probability and getting the denominator right, ensuring every individual has a known likelihood of appearing in the dataset. This distinction is often overlooked in ML-driven data generation, potentially leading to biased or unrepresentative outputs that misinform policy and public understanding.

Key Findings

▶ Watch: Why alignment is difficult: Three research tasks (2:00)

Frauke Kreuter's talk highlighted several critical findings derived from both social science research and early explorations into AI-driven data generation:

  1. Challenges of Synthetic Data Generation: While promising, current approaches to generating synthetic survey data using LLMs like GPT-3 (e.g., Lisa Argyle et al. from Duke) show significant limitations. These models can reliably answer closed-ended questions in ways that mirror human respondents for certain contexts (like US elections with binary choices). However, they often exhibit less variance than human responses and struggle to accurately capture minority or extreme views, particularly in multi-party systems (as shown by Sarah Ball and Simon Armendinger in Germany, where left and extreme right parties were poorly represented). Even "aligned" models, while performing better than unaligned ones, are not yet sufficiently accurate, suggesting a fundamental disconnect with the nuances of real human populations.
  1. Relevance of Survey Methodology for ML: Decades of research in survey methodology offer a robust framework for improving data quality and mitigating bias in ML-related data collection, especially for Reinforcement Learning with Human Feedback (RLHF) and annotation tasks. Kreuter introduced a framework that dissects potential sources of error into representation (from population to sample to respondents) and measurement (from concept to measurement to response). Each step presents a risk of bias or misalignment.
  1. Impact of Annotation Task Design: Subtle design choices in annotation tasks—analogous to web surveys—can significantly impact results. Kreuter presented experimental evidence where the display order of questions (e.g., showing two yes-no questions below a tweet vs. one by one) and the order of items (tweets) influenced how annotators classified offensive language and hate speech. Rating questions all at once, for instance, led to statistically fewer answers. This demonstrates that the "instrument design" for human feedback is as crucial in ML as it is in traditional surveys.
  1. Significance of Annotator Demographics: The characteristics of annotators matter. A Pew Research Center study revealed that online panel workers, often used for annotation, tend to be overrepresented by individuals receiving unemployment compensation, with lower incomes, in single-adult households, and without children. If these demographic characteristics correlate with the task at hand, the resulting labels and model training can be severely biased, missing crucial perspectives from other population segments.
  1. Underutilized High-Quality Data Archives: A vast repository of meticulously curated, high-quality social science data exists in archives like ICPSR, IPUMS, the Health and Retirement Survey (HRS), and the Panel Study of Income Dynamics (PSID). These datasets, which include longitudinal studies spanning decades, biomarkers, physical measurements, and detailed attitudinal information, are largely untapped by the ML community due to format (e.g., Stata, SPSS files), access restrictions, and siloed research practices. Kreuter highlighted their immense potential for training robust and adaptively aligned models.
  1. Dynamic Nature of Societal Values: Crucial societal attitudes are not static. Kreuter illustrated this with data from the US showing significant shifts in opinions on same-sex relationships (from 90% endorsement of "always wrong" in certain subgroups in the 1970s to 50% recently) and the relationship between working mothers and their children (showing increased divergence across political parties). These shifts underscore the need for "adaptive alignment" in AI, where models can learn and evolve with changing human values, rather than being aligned to a fixed, potentially outdated, or unrepresentative snapshot of preferences.

Technical Deep Dive

▶ Watch: The challenge of description: Getting the denominator right (4:00)

Kreuter's technical deep dive began by categorizing research tasks into prediction, causation, and description. While machine learning models often excel at prediction (e.g., identifying hateful tweets) and causation (e.g., isolating mechanisms in a randomized controlled trial), these tasks are "easy" in a specific statistical sense. They primarily require instances of the phenomena and relevant covariates. The "hard" task, surprisingly, is description—accurately quantifying a phenomenon across a population (e.g., the total number of people suffering from depression). This difficulty stems from the unique requirement of knowing the sampling probability for every case in the dataset, ensuring a correct denominator and representative sample. Without this, statistics are biased.

Traditional statistical agencies dedicate immense resources to achieving this representativeness. Kreuter illustrated this with the Bureau of Labor Statistics' (BLS) effort to calculate inflation numbers, which involves two data streams: scanner data for prices and laborious, hour-long interviews with sampled households to accurately capture the "basket of goods." This process includes developing a sampling frame (e.g., walking through streets to list house numbers, using postal files) and conducting costly interviews to ensure access and participation.

However, the efficacy of these traditional methods is declining. Kreuter presented stark evidence of falling response rates in surveys: data for the Consumer Price Index (CPI) saw rates drop from around 70% to 40% over a decade, while prominent telephone surveys like those by the Pew Research Center now face single-digit response rates. This necessitates increased effort to maintain statistical precision, leading to a growing interest in leveraging ML models for efficiency.

The ML community's response has included using models to predict survey respondents, automate data classification and imputation, and even creating synthetic data. Kreuter discussed the concept of personas in market research, where language models are given demographic profiles (e.g., "I'm a 20-year-old female, college degree...") and then asked survey questions. While some studies (e.g., Lisa Argyle et al. using GPT-3 for US election prediction) found that models could reliably mirror human answers for closed-ended questions, others (Ball and Armendinger in Germany) revealed significant limitations: synthetic data exhibited less variance than human data, and models struggled to capture minority or extreme political views.

A core contribution from survey methodology is the Total Survey Error framework, which systematically identifies sources of error. Kreuter implicitly applied this by detailing two main categories:

  1. Representation Error: Occurs when the observed data does not accurately reflect the target population. This includes sampling error (due to selecting only a subset), coverage error (when the sampling frame doesn't include all population members), and non-response error (when those who respond differ systematically from those who don't).
  2. Measurement Error: Occurs when the recorded response differs from the true value. This involves the complex cognitive steps a respondent takes: comprehending the question, recalling relevant information, making a judgment, and formulating a response. Factors like social desirability bias can distort answers.

Kreuter emphasized that these errors are highly relevant to ML annotation tasks, which she likened to web surveys. Drawing on decades of survey research, she outlined best practices for question design, such as Gallup's five-step approach from the 1940s: using filter questions to gauge respondent knowledge (e.g., about "filibuster"), measuring intensity, allowing for interpretation, and simplifying language.

She presented experimental results demonstrating how annotation interface design impacts outcomes. For instance, when labeling tweets for offensive language or hate speech, displaying multiple questions simultaneously versus sequentially, or the order in which tweets appeared, significantly affected annotator responses and subsequent ML classification models. These effects were statistically significant, highlighting the sensitivity of annotation quality to seemingly minor design choices.

Furthermore, the demographics of annotators themselves introduce bias. Studies show that workers on online platforms (e.g., crowd-sourcing platforms) often represent specific socioeconomic groups, which can lead to unrepresentative labels if their characteristics correlate with the task. To counter this, traditional survey research employs extensive efforts, including machine learning models to predict non-response and targeted follow-up, to ensure diverse participation. This process allows for statistical adjustment for non-response bias, a practice largely absent in many ML data collection efforts.

Finally, Kreuter pointed to the vast, underutilized resource of high-quality research data archives. Institutions like the Inter-university Consortium for Political and Social Research (ICPSR) and Integrated Public Use Microdata Series (IPUMS) house hundreds of thousands of meticulously curated datasets. Examples include the Health and Retirement Survey (HRS), which tracks health, economic, and social factors internationally, and the Panel Study of Income Dynamics (PSID), following 18,000 individuals in 5,000 families since the 1960s. These datasets contain rich, longitudinal information, including biomarkers and physical measurements, alongside attitudes and opinions. Their current formats (e.g., Stata, SPSS) and data use agreements, however, restrict easy access and direct use for training large language models.

Experimental Setup & Results

▶ Watch: Role of statistical agencies in measuring society (4:30)

Kreuter presented several key pieces of empirical evidence to underscore her points:

  1. Declining Survey Response Rates:
  • Consumer Price Index (CPI) Related Data: A graph showed a significant decline in response rates for data contributing to the CPI. Approximately 10 years prior to the talk, response rates were around 70%. By the time of the talk, these had fallen to 40%. This indicates a substantial increase in the effort required to gather the same amount of reliable data.
  • Pew Research Center Telephone Surveys: For well-known telephone surveys conducted by the Pew Research Center, response rates had plummeted to single digits. This drastic reduction raises serious concerns about the representativeness and statistical validity of such samples, even if they start with a probability-based sampling frame.
  1. Synthetic Data Performance:
  • GPT-3 for Election Prediction (Lisa Argyle et al., Duke): This study explored using GPT-3 to answer closed-ended survey questions for election prediction. The finding was that GPT-3 "reliably answers closed-ended survey questions in a way that closely mirrors answers given by human respondents." This was particularly effective in contexts like the US, with a two-party system where many votes are driven by geographic location or simple covariates.
  • GPT-3 for German Multi-Party System (Sarah Ball and Simon Armendinger, Munich): In contrast, applying similar methods to Germany's multi-party system yielded less accurate results. Specifically, the predictions for the left party and the extreme right were "not at all captured correctly." Furthermore, model predictions consistently showed less variance than human predictions within covariate sets, and while "aligned models" performed better than "unaligned models," their overall accuracy was still not "great."
  1. Impact of Annotation Task Design:
  • Hate Speech/Offensive Language Annotation: Kreuter described an experiment where hundreds of tweets were randomized to thousands of human raters to classify them for offensive language or hate speech. The experimental setup varied the display of questions and the order of tweets.
  • Results: The study found that "seeing order" (the order in which tweets appeared) and whether "rating questions at once" (multiple questions per tweet on one screen) or sequentially significantly mattered. Specifically, rating questions all at once led to fewer answers overall (statistically significant). Moreover, tweets appearing later in a segment were "not seen as hateful, not seen as offensive." These findings demonstrate that seemingly minor interface design choices in annotation tasks can introduce significant biases that propagate to trained ML classification models.
  1. Annotator Demographics (Pew Research Center):
  • A study by the Pew Research Center compared the demographics of online panel workers (often used for annotation tasks) to benchmark data.
  • Results: Online panel workers were found to be "much more likely to receive unemployment compensation, much more likely to be of low in low income categories, much more likely to be single family homes or a single adult households and not to have children in the household." This highlights a potential for significant demographic bias in crowd-sourced annotation efforts, which can impact the representativeness of the resulting training data.
  1. Longitudinal Shifts in Societal Attitudes:
  • Attitudes on Same-Sex Relationships (US): Using data from the US General Social Survey (GSS), Kreuter showed how attitudes on whether "it is wrong for same-sex adults to have sexual relations" have shifted over decades. In the mid-1970s, certain subgroups (e.g., those with a high school degree or less) showed a 90% endorsement rate for the statement that it is "always wrong." This dropped to around 50% in recent years, though a slight uptick was observed in the last year.
  • Attitudes on Working Mothers (US): Similarly, the question of whether "a working mother can have as good of a relationship with her child as a non-working mother" showed a dramatic shift. Initially, there was broad agreement across political parties. Over time, however, there has been "quite a bit of divergence in the answers to these question across the political parties."

These examples powerfully illustrate that "the ground is shifting" on fundamental societal values, emphasizing the need for AI alignment strategies that are dynamic and adaptive, rather than static.

Practical Implications

▶ Watch: Increasing difficulty and cost of data collection (6:00)

Frauke Kreuter's talk carries profound practical implications for various stakeholders in the AI/ML ecosystem, from model builders to policy makers.

For practitioners and model builders, the most critical takeaway is the need for heightened awareness regarding data quality and representation, especially when dealing with human preferences, attitudes, and behaviors. The common practice of "hunting for the highest accuracy" on readily available datasets often overlooks the "hard problem" of descriptive statistics and population representation. This means that models trained on unrepresentative data, even if technically proficient, may perpetuate or amplify existing societal biases, misinterpret diverse viewpoints, or fail to generalize effectively across different demographic groups. When designing RLHF or other human annotation tasks, principles from survey methodology—such as careful question wording, structured response options, understanding cognitive load, and considering display effects—are not merely "good to have" but are essential for eliciting accurate and unbiased feedback. Ignoring annotator demographics, for instance, risks building models aligned with a narrow, unrepresentative slice of society.

Infrastructure teams and those responsible for data pipelines should recognize the immense value of high-quality, curated social science data archives. Instead of relying solely on easily scraped internet data, efforts should be made to integrate datasets from ICPSR, IPUMS, HRS, or PSID. This integration, however, is not trivial. It requires:

  1. Data Translation: Converting data from traditional statistical software formats (Stata, SPSS) into ML-friendly formats (e.g., Python packages).
  2. Navigating Data Use Agreements: Working with data archives to understand and adapt consent statements and data use agreements that currently restrict easy scraping or direct use in training large models. This often involves exploring different levels of data aggregation that are acceptable for privacy.
  3. Public-Private Partnerships: Establishing collaborations between academic archives and industry, as exemplified by the ICPSR-Meta partnership, to share and leverage research data responsibly.

For model deployers and policy makers, the talk highlights the inherent tradeoffs between speed and accuracy. In situations where "getting an okay number fast might be better and more important than getting the correct number 10 years later" (e.g., for rapidly evolving issues), synthetic data might offer a temporary solution. However, this must be balanced against the risks of misinformed policy due to unrepresentative or biased data. The dynamic nature of societal values means that AI alignment cannot be a one-time fix. Models need mechanisms for adaptive alignment, continuously learning and adjusting to shifting norms and subgroup differences. This necessitates the creation of "living benchmarks"—regular, high-quality population surveys designed specifically to monitor model drift and provide updated ground truth from real human beings.

Finally, Kreuter's call for interdisciplinary collaboration is a practical directive. The ML community needs the expertise of social scientists to ask the right questions about what to align with, design robust data collection instruments, and interpret the nuances of human behavior. Conversely, social scientists can benefit from ML tools for imputation, classification, and potentially for making their vast data archives more accessible and useful. This partnership is crucial for addressing the transparency needs in government statistics (as highlighted by CNSTAT) and updating criteria for transparency initiatives in public opinion research (e.g., AAPOR), ensuring that AI development is grounded in a comprehensive understanding of human society and its complexities.

Key Takeaways

  • Adaptive Alignment is Crucial: AI alignment must move beyond static, technical definitions to consider the dynamic, diverse, and subgroup-specific nature of human values, norms, and attitudes, which continuously shift over time.
  • High-Quality Data is Foundational: Robust AI models, especially for descriptive tasks (population statistics), depend on meticulously collected, representative data. This is a "hard problem" that requires significant effort and attention to sampling probability and measurement accuracy.
  • Survey Methodology Offers Invaluable Lessons: Decades of research in survey methodology provide a rich framework for designing effective data collection, annotation, and human feedback processes (e.g., RLHF) to mitigate bias, improve measurement, and ensure representativeness.
  • Vast Data Archives Remain Untapped: Hundreds of thousands of high-quality, curated social science datasets (e.g., ICPSR, IPUMS, HRS, PSID) exist but are underutilized by the ML community due to format, access restrictions, and disciplinary silos. These represent a massive opportunity for training more aligned and robust AI.
  • Design Choices Matter Immensely: Subtle decisions in annotation task design, question wording, and annotator selection can introduce significant biases into training data, leading to models that misrepresent population views or fail to capture minority perspectives.
  • Interdisciplinary Collaboration is Essential: Bridging the gap between machine learning and social sciences through partnerships, data translation, and shared expertise is vital for building AI systems that are not only technically proficient but also ethically aligned and responsive to the complexities of a changing human world.

About the Speaker(s)

Frauke Kreuter is a distinguished social scientist with a keen interest in data quality and data-centricness. Her work focuses on the critical intersection of social science methodology and the rapidly evolving field of machine learning, particularly concerning AI alignment. She brings a wealth of knowledge from survey research and survey methodology, fields that have spent over 80 years grappling with the challenges of accurately measuring human preferences, behaviors, attitudes, and opinions. Kreuter advocates for leveraging this extensive social science expertise to inform and improve how AI models are trained and aligned, ensuring they reflect a more accurate and nuanced understanding of society. Her perspective emphasizes the importance of moving beyond purely technical alignment to address the fundamental question of what human values and norms AI systems should be aligned with, and how these dynamic elements can be effectively captured and integrated.

Reviews

Maya Iyer (Theoretical ML Researcher) — SOLID

Kreuter delivers a well-informed interdisciplinary argument that the ML community's alignment efforts are methodologically underspecified — not in the formal sense, but in the measurement sense. The core contribution is a transfer of the Total Survey Error framework into the RLHF/annotation context, illustrated with empirical evidence on response rate decay, annotator demographic skew, and interface-driven label noise. The talk is honest about its scope: this is a call for methodological borrowing, not a new theorem or algorithm. It earns its place at ICML as a corrective voice from a neighboring discipline, but it does not change how a theoretician would formulate a problem. Solid…

Chen Zhao (Applied ML Researcher & Empiricist) — SOLID

Kreuter delivers a well-grounded interdisciplinary talk that makes a genuine contribution by importing survey methodology's Total Survey Error framework into ML alignment discourse. The core argument — that RLHF and annotation pipelines have representation and measurement problems that social scientists solved decades ago — is correct, practically actionable, and underappreciated in the ML community. The empirical anchors (annotation order effects, annotator demographic skew, declining response rates, LLM synthetic data failures in multi-party systems) are real and well-chosen. But this is a position/advocacy talk, not a research paper with new experimental results, and it should be…

→ Top-rated talks at International Conference on Machine Learning 2025

All talks from International Conference on Machine Learning 2025