The Value of Prediction in Identifying the Worst-Off

Unai Fischer Abaigar (PhD student · University of Munich), Christoph Kern, Juan Perdomo

International Conference on Machine Learning 2025 · Oral

Overview

This article delves into a crucial dilemma faced by public agencies: how to effectively allocate scarce resources to individuals most in need, especially when leveraging machine learning (ML) for decision-making. Presented by Unai Fischer Abaigar from the University of Munich, alongside co-authors Christoph Kern and Juan Perdomo, the talk "The Value of Prediction in Identifying the Worst-Off" investigates the complex tradeoffs between improving predictive model accuracy and expanding institutional capacity. It challenges the intuitive assumption that better predictions always lead to better outcomes, particularly in public sector contexts focused on welfare.

Watch on SlidesLive

Visual summary for The Value of Prediction in Identifying the Worst-Off by Unai Fischer Abaigar, Christoph Kern, Juan Perdomo
Visual summary for The Value of Prediction in Identifying the Worst-Off by Unai Fischer Abaigar, Christoph Kern, Juan Perdomo

Key moments

  1. 0:00 Introduction: Identifying the worst-off with algorithmic systems
  2. 2:00 Core question: When are predictive capability improvements worth it?
  3. 2:18 Main findings: Value of prediction vs. capacity improvements
  4. 3:25 Methodology: Theoretical models and real-world data
  5. 4:00 Introducing the prediction access ratio concept
  6. 5:00 Formalizing the screening problem: features and outcomes
  7. 6:00 How institutions use predictions for optimal screening

The Value of Prediction in Identifying the Worst-Off

Speakers: Unai Fischer Abaigar, PhD Student, University of Munich; Christoph Kern, Junior Professor, University of Munich; Juan Perdomo, Incoming Faculty, NYU

Conference: ICML 2025

YouTube: https://slideslive.com/39044107

Overview

This article delves into a crucial dilemma faced by public agencies: how to effectively allocate scarce resources to individuals most in need, especially when leveraging machine learning (ML) for decision-making. Presented by Unai Fischer Abaigar from the University of Munich, alongside co-authors Christoph Kern and Juan Perdomo, the talk "The Value of Prediction in Identifying the Worst-Off" investigates the complex tradeoffs between improving predictive model accuracy and expanding institutional capacity. It challenges the intuitive assumption that better predictions always lead to better outcomes, particularly in public sector contexts focused on welfare.

The core of the research explores when investments in predictive capabilities are truly "worth it" compared to other policy decisions, such as simply increasing the number of people an institution can serve. This is particularly salient for public agencies like employment offices, which use ML to identify individuals at high risk of adverse events, such as long-term unemployment, to prioritize support. The findings offer critical guidance for social planners and policymakers on where to direct limited resources for maximum welfare impact, moving beyond average performance metrics to focus on the most vulnerable.

The talk introduces and rigorously analyzes the prediction access ratio, a formal notion that quantifies the marginal value of improving bureaucratic capacity versus enhancing predictive quality. Through a combination of theoretical modeling and empirical validation using a large-scale administrative dataset of German job seekers, the authors provide actionable insights into the conditions under which one investment strategy outweighs the other. This work is essential for anyone designing or deploying algorithmic systems in resource-constrained public service environments, urging a holistic view of system design that considers both technological advancements and operational realities.

Background

▶ Watch: Introduction: Identifying the worst-off with algorithmic systems (0:00)

Public agencies globally are grappling with the dual pressures of modernizing operations and functioning within severe resource constraints. In response, many have increasingly turned to algorithmic systems and machine learning tools to streamline early-stage decision-making. These tools are often employed to direct institutional attention, manage screening processes, and solve complex resource allocation problems—identifying which individual cases should receive faster follow-up, quicker screening, or prioritized support. A pervasive goal in the public sector, particularly in social welfare contexts, is to identify and assist individuals who are most at risk of experiencing adverse events, often referred to as the "worst-off."

Consider, for instance, an employment office tasked with estimating the likely duration of an unemployment episode for a new job seeker. The aim is to prioritize support for those predicted to become long-term unemployed. Such prediction-based allocation systems are fundamentally shaped by two factors: the quality of predictions (how accurately can we estimate proxies of individual need) and the administrative capacity (how many people can the institution actually process or prioritize). Improving either of these factors incurs costs and binds institutional resources. For example, enhancing data collection standards to improve downstream prediction quality is a significant investment. This highlights a critical, often overlooked, tradeoff: from a social planner's perspective, investments in prediction are not isolated but must be weighed against other potential uses of limited resources, such as expanding the human or operational capacity of the system.

The "prediction access ratio," a formal concept initially introduced by Perdomo (2024), provides a framework for understanding this tradeoff. It quantifies the relative marginal value gained from small improvements in bureaucratic capacity versus small improvements in predictive quality. This ratio allows policymakers to ask: how many more at-risk job seekers can be successfully screened if capacity is raised by a few percentage points, compared to investing in improving predictive accuracy? The underlying problem is that it's not immediately obvious which investment strategy yields greater returns in terms of welfare for the at-risk population, and this optimal balance may shift depending on the specific operational regime an institution finds itself in. This research aims to illuminate these tradeoffs, providing a principled understanding of when improvements in predictive capability genuinely justify their cost in the context of public service delivery.

Key Findings

▶ Watch: Main findings: Value of prediction vs. capacity improvements (2:18)

The research yields several pivotal findings that challenge conventional wisdom regarding investments in predictive accuracy within resource-constrained public service settings. The primary discovery is that the value of improving predictions is highly contextual and often non-linear, exhibiting peak utility at the extremes of predictive capability. Specifically, improvements in prediction are most valuable under two distinct conditions:

  1. When baseline predictions explain a very limited amount of the outcome variance (R-squared close to 0): In this "first-mile effort" scenario, even small gains in predictability can establish a crucial baseline signal, moving the system from near-random allocation towards some level of informed decision-making.
  2. When the allocation system is nearly perfect (R-squared close to 1): Here, prediction improvements serve to "squeeze out the last percentage points," refining an already highly accurate system to near-optimal performance.

However, the talk highlights that in practical settings, institutions often operate in an "in-between" regime: they possess a decent predictive model that explains some, but not all, of the outcome variance (e.g., an R-squared of approximately 0.15 in the empirical study), yet the model is far from perfect. In this common scenario, especially when institutional capacity is significantly limited, the relative value of prediction improvements diminishes. The research robustly demonstrates that capacity improvements—such as increasing the number of individuals an institution can process or prioritize—have a much larger and more immediate impact on the welfare of those at risk than marginal gains in predictive accuracy.

These conclusions are drawn from a two-pronged methodological approach. First, a theoretical investigation employs simple statistical linear models with Gaussian assumptions to derive fundamental insights into the prediction access ratio. Second, these theoretical insights are rigorously tested and validated using a real-world, nationwide administrative dataset on German job seekers. Crucially, the empirical study confirms that the core intuitions gleaned from the simpler theoretical models hold up remarkably well even in a complex, messy social prediction setting with non-Gaussian outcomes and temporal distribution shifts. This underscores the robustness and practical applicability of the findings for guiding resource allocation decisions in public agencies.

Technical Deep Dive

▶ Watch: Methodology: Theoretical models and real-world data (3:25)

The core of the technical investigation revolves around formalizing the screening problem faced by public agencies and defining a welfare metric that captures the impact on the "worst-off." The authors consider a setting with features X and real-valued individual outcomes Y, where Y encodes how pressing a case is (e.g., the true number of months a job seeker will remain unemployed). A social planner's normative goal is to identify individuals whose outcomes Y fall above a specific threshold, denoted as beta. This beta parameter defines the segment of the population considered "at risk."

Since institutions typically lack oracle access to these true outcomes Y, they rely on predictive models to estimate them. These models are then used to formulate optimal screening policies, which involve prioritizing and screening in the top fraction of predictions, constrained by the available institutional capacity, denoted as alpha. It's critical to note that alpha (actual capacity) and beta (normative goal for identifying at-risk individuals) do not necessarily align. For instance, an institution might aim to identify the top 5% most at-risk individuals (beta=0.05) but only possess the resources to serve 2% (alpha=0.02).

The welfare metric chosen for this analysis is the True Positive Rate (TPR). This is defined as the conditional probability that an individual who is genuinely at risk (i.e., Y > beta) will actually be successfully screened in by the allocation policy. Graphically, this corresponds to the ratio of individuals in the upper-right quadrant (correctly identified at-risk and screened in) to the entire "at-risk" population (the purple band in the presenter's scatter plot). The objective is to understand how small improvements in capacity (increasing alpha) trade off against improvements in the accuracy of the allocation (making predictions more accurate).

To quantify this tradeoff, the authors introduce the Prediction Access Ratio (PAR). In their simple statistical linear Gaussian model, predictive accuracy is captured by the R-squared value, which represents the proportion of outcome variance explained by the model. The PAR is formally expressed as the ratio of the marginal value gained from a small improvement in capacity (differentiating the TPR with respect to alpha) to the marginal value gained from a small improvement in predictive accuracy (differentiating the TPR with respect to R-squared).

The theoretical analysis, based on linear models with Gaussian assumptions, yields key insights into the behavior of the PAR:

  • Extremes of Predictive Signal: The marginal value of better prediction is highest when the baseline predictions have almost no predictive signal (R-squared approaches 0) or when the system is already close to an optimal allocation (R-squared approaches 1), assuming sufficient capacity to screen the entirety of the at-risk population. Formally, the marginal value of prediction diverges as R-squared goes to 0 or 1. This suggests that prediction improvements are crucial for establishing an initial signal or for fine-tuning a near-perfect system.
  • Intermediate Regimes and Limited Capacity: In the more common scenario where a predictor explains a positive but imperfect share of the outcome variance (0 < R-squared < 1), and especially when institutional capacities are very limited (small alpha), the relative value of prediction is smaller. The PAR is lower-bounded by alpha^(-1) / (1 - R-squared). This formula reveals that as alpha (capacity) becomes smaller, this lower bound explodes, indicating that capacity improvements become disproportionately valuable. Essentially, when resources are very scarce, expanding the pool of people that can be served (increasing alpha) has a much greater impact on welfare than striving for perfectly accurate predictions for a small, fixed pool. This formalization provides a powerful tool for guiding resource allocation decisions based on the current state of predictive models and institutional capacity.

Experimental Setup & Results

▶ Watch: Formalizing the screening problem: features and outcomes (5:00)

To validate the theoretical insights derived from simplified Gaussian models, the researchers conducted an empirical study using a rich, real-world administrative dataset.

Dataset:

The study utilized a nationwide administrative dataset on German job seekers. This unique dataset was provided by the German Institute for Employment Research and comprised a 2% sample of all unemployment spells in Germany over the past 50 years. The data is highly confidential and represents the exact type of information that employment offices across Germany collect and would use to build their own statistical prediction systems. This makes the experimental setting exceptionally realistic and relevant to the public sector challenges discussed. The focus of the empirical investigation was specifically on job seekers who were unemployed for over one year, which aligns with the legal definition of long-term unemployment in Germany, representing a critical "worst-off" population segment.

Models & Predictive Performance:

For the predictive task—estimating unemployment duration—the researchers primarily employed gradient boosting style models. These models are widely used in practice for their strong performance on tabular data and ability to handle complex relationships. After training, these models explained roughly 15% of the outcome variance (R-squared ≈ 0.15). This R-squared value is significant because it places the empirical setting squarely in the "in-between" regime identified by the theoretical analysis—a scenario where there is some predictive signal, but the model is far from perfectly accurate. This mirrors the reality often observed in social prediction settings, which are inherently complex and noisy.

Key Empirical Results:

The empirical analysis consistently confirmed the intuitions gained from the theoretical models. When examining the prediction access ratio (PAR) across different screening capacities (varying alpha) and with the observed R-squared of 0.15:

  • The PAR was generally large, especially as institutional capacity became more limited. This directly supports the theoretical finding that in intermediate predictive regimes, when capacity is scarce, investments in expanding capacity yield a much higher marginal return on welfare than investments in marginal prediction accuracy improvements. In other words, if an institution can only serve a very small fraction of the at-risk population, it's more impactful to expand that fraction than to make the predictions for that small fraction slightly more accurate.
  • The study also explored how the PAR shifts across different R-squared regimes. If the system was closer to a randomized allocation (i.e., an R-squared near 0, indicating almost no predictive signal), the PAR shifted, indicating that investments in prediction would indeed become more valuable again. This is because establishing even a basic predictive signal from a state of near-randomness has a high initial value.

The authors explicitly acknowledge that many assumptions from their theoretical investigation (e.g., IID data, Gaussian outcomes, lack of temporal distribution shifts) do not hold in the real-world administrative dataset. However, the remarkable consistency between the theoretical predictions and the empirical observations underscores the robustness of their core message. The empirical results provide strong evidence that the derived principles for balancing investments in prediction versus capacity are highly applicable to practical public sector allocation problems.

Practical Implications

▶ Watch: How institutions use predictions for optimal screening (6:00)

The findings of this research carry significant practical implications for a wide range of stakeholders involved in the design, deployment, and management of AI/ML systems in the public sector. For practitioners and model builders, the primary takeaway is a crucial re-evaluation of the default assumption that "better predictions always lead to better outcomes." Instead, the value of prediction improvements must be rigorously contextualized within the broader operational and resource landscape.

  • Strategic Investment Decisions: Public agencies and social planners must adopt a more nuanced approach to resource allocation. Rather than solely focusing on pushing model accuracy metrics, they need to explicitly consider the tradeoff between investing in improved data collection, more sophisticated models, or advanced training techniques versus expanding the raw capacity of the institution. This could mean hiring more caseworkers, increasing the number of available screening slots, or streamlining non-ML-related administrative processes. The prediction access ratio serves as a vital analytical tool for guiding these strategic investment decisions, indicating when one type of investment will yield greater returns in terms of welfare for the "worst-off."
  • Context-Dependent Value of Prediction: The research clearly demonstrates that the utility of prediction improvements is not uniform. If an agency's current models are extremely weak (near-random) or exceptionally strong (near-perfect), then investing in prediction accuracy can be highly valuable. However, in the common "intermediate" scenario, where models offer some but imperfect signal, and institutional capacity is constrained, expanding capacity often becomes the more impactful lever for improving welfare outcomes. This means model builders should not pursue marginal accuracy gains blindly; they should understand the operational context their models will inhabit.
  • Holistic System Design: The talk implicitly advocates for a holistic perspective on algorithmic system design in the public sector. ML models are only one component of a larger system that includes human caseworkers, specific intervention programs, and overall administrative capacity. The effectiveness of a predictive model is intrinsically linked to the institution's ability to act on its predictions. As Unai Fischer Abaigar points out, the research opens avenues for further study into how prediction systems interact with downstream interventions or work in tandem with caseworkers who operate these systems in practice.
  • Limitations and Further Considerations: The study, while robust, has inherent limitations. It primarily focuses on the identification of the "worst-off" and their successful screening (true positive rate) as the welfare metric. It does not delve into the efficacy or potential harms of the subsequent interventions themselves. The Q&A session highlighted this, with a participant questioning how the research ties into methodologies that account for demographic disparities or cases that "fall through the cracks." The speaker acknowledged that the paper focuses on the welfare of the "at-risk population" rather than an average welfare notion, and future work could investigate settings where interventions might be harmful to certain subgroups, complicating allocation decisions further. This underscores that while the PAR provides a powerful framework for resource allocation between prediction and capacity, it should be integrated into a broader ethical and operational framework that considers equity, fairness, and the full lifecycle of interventions.

In essence, the research serves as a critical reminder that technological solutions, however advanced, must be evaluated not in isolation but in concert with the human and organizational infrastructure that supports them, especially when the goal is to improve the welfare of vulnerable populations.

Key Takeaways

  • Prediction Value is Context-Dependent: Improvements in predictive accuracy are most valuable at the extremes of predictive capability: when models are either very weak (establishing a baseline signal) or nearly perfect (fine-tuning an optimal system).
  • Capacity Over Prediction in Intermediate Regimes: In the common scenario where predictive models offer some signal but are imperfect, and institutional capacity is limited, expanding capacity often yields greater welfare improvements for the "worst-off" than marginal gains in prediction accuracy.
  • The Prediction Access Ratio (PAR) as a Decision Tool: The PAR formally quantifies the tradeoff between investing in predictive accuracy versus institutional capacity, providing a principled framework for social planners to guide resource allocation decisions.
  • Real-World Validation: Theoretical insights, derived from simplified models, are robustly confirmed by empirical analysis using a nationwide administrative dataset of German job seekers, highlighting the practical applicability of the findings despite complex real-world conditions.
  • Beyond Accuracy Metrics: For ML systems in public sector allocation, a holistic view is crucial. Investments in predictive capabilities must be contextualized against other system components, such as institutional processing capacity, and evaluated based on their impact on the welfare of vulnerable populations, rather than just average performance or model accuracy.
  • Strategic Policy Implications: Policymakers should weigh the costs and benefits of improving data collection and model sophistication against expanding human resources and operational capacity, using tools like the PAR to optimize for genuine welfare outcomes.

About the Speaker(s)

The talk was presented by Unai Fischer Abaigar, a PhD student from the University of Munich. His work represents a collaboration with his colleagues and advisors. Juan Perdomo is an incoming faculty member at NYU, actively seeking students for his research group. Christoph Kern is one of Unai Fischer Abaigar's PhD advisors and holds a position as a junior professor at the University of Munich. Together, their interdisciplinary expertise in machine learning, economics, and public policy underpins this impactful research.

Reviews

Maya Iyer (Theoretical ML Researcher) — SOLID

A careful, policy-relevant theoretical paper that formalizes the tradeoff between predictive accuracy and institutional capacity in public-sector screening problems. The central object — the Prediction Access Ratio — is clean and the core message is directionally correct and underappreciated in applied ML circles. The theoretical contribution is real but modest: linear-Gaussian assumptions do most of the heavy lifting, and the empirical validation, while honest, confirms rather than stress-tests the framework. A solid contribution for researchers working at the ML-policy interface; less essential for theorists.

Chen Zhao (Applied ML Researcher & Empiricist) — SOLID

A theoretically grounded and empirically validated investigation into when improving predictive accuracy matters for welfare-focused allocation problems versus expanding institutional capacity. The prediction access ratio is a clean formal contribution, and the empirical confirmation on a nationally representative German employment dataset is credible. However, the result is bounded in significance — the finding that capacity dominates in intermediate R-squared regimes under limited alpha is intuitive once formalized, the empirical evaluation is relatively thin (one dataset, one model class, one welfare metric), and the paper's practical uplift depends heavily on policymakers being able to…

→ Top-rated talks at International Conference on Machine Learning 2025

All talks from International Conference on Machine Learning 2025