Position: Political Neutrality in AI Is Impossible — But Here Is How to Approximate It
Jillian Fisher, Ruth Elisabeth Appel, Chan Young Park, Yujin Potter, Liwei Jiang, Taylor Sorensen, Shangbin Feng, Yulia Tsvetkov, Margaret Roberts, Jennifer Pan, Dawn Song, Yejin Choi
Overview
In an era where artificial intelligence increasingly influences decision-making across various domains, the notion of political neutrality in AI has emerged as a critical, yet often elusive, goal. This thought-provoking talk, presented by Jillian Fisher at ICML 2025, challenges the fundamental premise of achieving true political neutrality in AI models. Drawing on insights from philosophy, political science, and computer science, the presented work argues that true neutrality is not only theoretically impossible but also technically unattainable given the human-centric nature of AI development.

Key moments
- 0:00 Introduction: Political bias in AI and neutrality problem
- 0:45 Thought experiment: Why true neutrality by non-action fails
- 3:30 Conclusion: True political neutrality is theoretically and technically impossible
- 4:00 Approximating neutrality: A practical approach inspired by philosophy
- 4:45 Framework: Approximations across ecosystem, system, and output levels
- 5:15 System-level approximations: uniform, reflective, and transparency
- 5:50 Output-level approximations: refusal, avoidance, pluralism, transparency
- 7:25 Main takeaways: Shift to approximation, model transparency, and evaluations
Position: Political Neutrality in AI Is Impossible — But Here Is How to Approximate It
Speakers: Jillian Fisher; Ruth Elisabeth Appel; Chan Young Park; Yujin Potter; Liwei Jiang; Taylor Sorensen; Shangbin Feng; Yulia Tsvetkov; Margaret Roberts; Jennifer Pan; Dawn Song; Yejin Choi
Conference: ICML 2025
YouTube: https://slideslive.com/39044043
Overview
In an era where artificial intelligence increasingly influences decision-making across various domains, the notion of political neutrality in AI has emerged as a critical, yet often elusive, goal. This thought-provoking talk, presented by Jillian Fisher at ICML 2025, challenges the fundamental premise of achieving true political neutrality in AI models. Drawing on insights from philosophy, political science, and computer science, the presented work argues that true neutrality is not only theoretically impossible but also technically unattainable given the human-centric nature of AI development.
The core contribution of this research, co-authored by a diverse team including Ruth Elisabeth Appel, Chan Young Park, Yujin Potter, Liwei Jiang, Taylor Sorensen, Shangbin Feng, Yulia Tsvetkov, Margaret Roberts, Jennifer Pan, Dawn Song, and Yejin Choi, is a novel framework for approximating political neutrality. By shifting the discourse from an unattainable ideal to a spectrum of pragmatic approximations, the authors open the door for AI engineers and researchers to design, evaluate, and interpret AI systems with greater transparency and a more nuanced understanding of their inherent biases. This perspective is crucial for building trust in AI systems and navigating the complex ethical and societal implications of their deployment.
The importance of this work cannot be overstated. As AI models, particularly large language models (LLMs), become more pervasive, their embedded political biases can have significant downstream effects on users and societal outcomes. This talk provides a foundational shift in how the AI community should approach bias mitigation, moving beyond a futile pursuit of absolute neutrality towards a more realistic and actionable framework that acknowledges and manages inherent political leanings. It encourages a multidisciplinary approach, drawing lessons from fields that have grappled with similar challenges for centuries, to foster more robust and responsible AI development.
Background
▶ Watch: Introduction: Political bias in AI and neutrality problem (0:00)
The pervasive issue of bias in AI models, particularly political bias, has been a significant concern in recent years. This bias is not merely an abstract concept; it manifests in the models themselves and subsequently impacts users, leading to potentially unfair or skewed outcomes. In response to these observed biases, a common proposed solution has been for AI models to strive for political neutrality. However, the talk rigorously deconstructs this notion, revealing its inherent impossibility through a series of thought experiments and theoretical arguments rooted in philosophy and political science.
Fisher illustrates the theoretical impossibility of political neutrality by examining three common interpretations:
- Neutrality by Non-Action: This approach suggests that an AI should simply "stay out of it" when faced with politically charged decisions. For instance, if an AI were mediating between two feuding countries, it would remain silent. However, history repeatedly demonstrates that non-action inherently favors the stronger or majority party, thereby being far from neutral. In a conflict, inaction is often a form of action, reinforcing existing power imbalances.
- Neutrality by Equality: This interpretation posits that an AI should strive for equidistant decisions, splitting differences equally between opposing viewpoints. An example given is an AI setting a government budget, with one group wanting $50 million and another $100 million, leading the AI to propose $75 million. The flaw emerges when a third party is introduced, also advocating for $75 million. The AI becomes "confused," as political topics rarely have a perfectly equal or easily divisible stance. Complex issues often involve more than two viewpoints, and "splitting the difference" can ignore underlying needs, historical context, or power dynamics, thus failing to achieve true equality or fairness.
- Neutrality by Moderation/Centrism: This view suggests that an AI should occupy the "middle ground" of the political spectrum. Using the political compass model (Y-axis: authoritarian to libertarian; X-axis: left to right), neutrality would ideally be at the (0,0) center. However, as Fisher explains, this "center" is often labeled centrism or liberalism, which itself represents a distinct political ideology. This ideology typically supports individual rights, limited government, and government power derived from the consent of the people. An AI programmed to embody these principles is not neutral; it is explicitly adopting a specific political stance, demonstrating that even the perceived middle is inherently biased towards a particular set of values.
These philosophical challenges are compounded by the technical realities of AI development. As computer scientists, we are not merely discovering objective truths but building these models. Every step of the AI pipeline – from data collection and labeling to model architecture design, hyperparameter tuning, and evaluation metrics – involves human decisions, each imbued with the biases and perspectives of the creators. These human biases are embedded into the training pipeline and subsequently into the models themselves, adding a layer of technical impossibility to the theoretical impossibility of achieving true political neutrality. The problem, therefore, is not merely about finding the right definition of neutrality but acknowledging that the very act of creation precludes it.
Key Findings
▶ Watch: Conclusion: True political neutrality is theoretically and technically imposs... (3:30)
The central and most impactful finding of this research is the unequivocal declaration that true political neutrality in AI is impossible. This impossibility stems from both theoretical philosophical arguments, demonstrating that common interpretations of neutrality inherently favor certain outcomes or ideologies, and technical realities, where human biases are inextricably woven into every stage of AI model development and training.
Given this impossibility, the paper’s primary contribution is to pivot the discussion from an elusive, unattainable ideal to a practical, actionable framework for approximating political neutrality. This shift is inspired by philosopher Joseph Raz's insight that "neutrality can be a matter of degree," allowing for a more nuanced understanding and engagement with the inherent biases of AI. By embracing approximations, the framework enables AI engineers to move beyond unproductive philosophical stalemates and instead focus on the tradeoffs that come with different approaches to managing bias.
The paper outlines a comprehensive taxonomy of these approximations across three distinct levels of the AI ecosystem:
- Ecosystem Level: Approximations that consider the aggregate behavior of multiple AI models.
- System Level: Approximations focused on the design and characteristics of individual AI models.
- Output Level: Approximations concerning how an AI model generates responses to specific inputs.
Within these levels, the research defines specific types of approximations, such as diversity-based neutrality at the ecosystem level, uniform neutrality, reflective neutrality, and system transparency at the system level, and refusal, avoidance, reasonable pluralism, and output transparency at the output level. These definitions provide a concrete vocabulary and framework for discussing and implementing bias management strategies.
Furthermore, the research includes empirical experimentation on current Large Language Models (LLMs), evaluating their behavior regarding output-level approximations on seven novel political datasets. While the full results are detailed in the accompanying paper and poster, the talk teases findings related to how models like "R1" and "Claude" exhibit varying degrees of caution when responding to political questions. This empirical component grounds the theoretical framework in real-world AI system behavior, demonstrating the practical applicability of the proposed approximations.
Finally, the work underscores several critical takeaways for the AI community: the necessity of shifting towards practical approximation goals, the paramount importance of model transparency and diverse evaluation methodologies for complex concepts like political bias, and the indispensable value of interdisciplinary collaboration to tackle such multifaceted challenges effectively.
Technical Deep Dive
▶ Watch: Framework: Approximations across ecosystem, system, and output levels (4:45)
The core of this work lies in its detailed framework for approximating political neutrality, which is systematically categorized across three levels of the AI ecosystem: the Ecosystem Level, the System Level, and the Output Level. This tiered approach allows for a granular understanding and implementation of bias management strategies, moving beyond a monolithic view of "neutrality."
Ecosystem Level: Neutrality Through Diversity
At the highest level, the paper proposes neutrality through diversity. This concept is inspired by how traditional media and news outlets often manage political bias. The idea is that if an entire ecosystem comprises multiple AI models, each with its own distinct and acknowledged political biases, then the aggregate effect of this diverse collection can approximate neutrality. Instead of forcing individual models to be neutral (an impossible task), the collective provides a range of perspectives. For example, a user could choose to interact with models known to lean left, right, or center, or an orchestrator system could present outputs from multiple models, allowing the user to discern a broader, more "neutral" landscape of viewpoints. This approach acknowledges that individual components may be biased, but the ensemble offers a more balanced overall experience.
System Level: Approximations for Individual Models
Zooming in on a single AI model, the paper outlines three distinct types of approximations:
- Uniform Neutrality: This is the idea that a model should respond consistently and identically to a given input, regardless of the user's political ideology, geographical location (e.g., how conservative their state is), or other politically relevant user attributes. The model's behavior is invariant to the political context of the query or query-er. Achieving this might involve robust fine-tuning on diverse datasets with an emphasis on consistency, or applying adversarial training techniques to minimize sensitivity to political demographic features. However, even this approach implicitly favors a "universal" perspective which may itself be a political stance.
- Reflective Neutrality: In direct contrast to uniform neutrality, this approximation suggests that the bias of the model's response should match the bias of the user interacting with it. If a user identifies as conservative, the model might present information or generate content with a conservative lean. If the user is liberal, the model would reflect a liberal bias. Implementing this would require sophisticated user profiling (explicit or implicit) to infer political leanings and then dynamically adapting the model's output generation process. This could involve using personalized fine-tuning, retrieval-augmented generation (RAG) systems that pull from politically aligned sources, or gating mechanisms that select from multiple pre-biased model outputs. The challenge here is accurately assessing user bias and avoiding filter bubbles.
- System Transparency: This approach acknowledges that bias is inherent and acceptable, as long as the system explicitly informs the user about its own biases. This could involve a clear disclosure statement, a "bias label" associated with the model, or even a configurable "slider bar" (as mentioned in the Q&A) that indicates the model's leanings on various political dimensions. For example, a model might state, "This AI has been trained with a slight progressive bias on social issues." Achieving this requires rigorous bias auditing and evaluation frameworks to quantify the model's political leanings, potentially using political ideology datasets or sentiment analysis tools specifically tuned for political discourse. The technical challenge is not removing bias, but accurately measuring and communicating it.
Output Level: Approximations for Specific Responses
Finally, at the most granular level, the paper defines four approximations concerning the model's specific output to a particular input:
- Refusal: The simplest form of approximation, where the model outright refuses to answer a politically sensitive question. This is a clear signal that the model is designed to avoid engagement with certain contentious topics. Implementing this typically involves content moderation filters, keyword detection, or semantic analysis to identify politically charged queries that trigger a refusal mechanism.
- Avoidance (Soft Refusal): Rather than a direct refusal, the model avoids the question by providing a non-committal or tangential response. It might reframe the question, offer general background information without taking a stance, or divert to a related but less controversial topic. This requires more sophisticated natural language generation (NLG) capabilities to craft evasive yet seemingly helpful responses, often leveraging pre-defined neutral phrases or summarization techniques that strip out contentious elements.
- Reasonable Pluralism: This approximation aims to present all reasonable viewpoints relevant to a given answer. When asked about a controversial policy, the model would articulate the main arguments for and against it, citing different perspectives without endorsing one. This would involve multi-perspective summarization, stance detection on various political topics, and potentially fact-checking to ensure presented viewpoints are indeed "reasonable" (i.e., not conspiracy theories or demonstrably false). The challenge lies in defining "reasonable" and ensuring comprehensive coverage without overwhelming the user.
- Output Transparency: Similar to system transparency, but applied to a specific output. The model not only provides an answer but also discloses the bias inherent in that particular output. For instance, an answer might be accompanied by a disclaimer like, "This response reflects a predominantly libertarian perspective on economic policy." This requires real-time bias analysis of generated text, potentially using embedding spaces or classifier models trained to identify political leanings within short text snippets.
The paper emphasizes that moving into these approximations allows for explicit discussions about tradeoffs across various characteristics (though the specific chart details were not fully presented in the talk). These tradeoffs might involve factors like user satisfaction, accuracy, perceived fairness, computational cost, and the risk of reinforcing existing biases. The authors also briefly mention exploring static decision tree type decisions versus dynamic processes for choosing between approximations, suggesting that context-aware and adaptive strategies might be necessary for real-world deployments. Specific implementation details, such as leveraging existing AI technologies for different parts of these approximations, are elaborated in the full paper.
Experimental Setup & Results
▶ Watch: System-level approximations: uniform, reflective, and transparency (5:15)
The research presented in the talk includes an empirical study designed to evaluate how current Large Language Models (LLMs) exhibit different types of output-level approximations when confronted with politically charged inputs. This practical evaluation grounds the theoretical framework in the observable behavior of contemporary AI systems.
The experimental setup involved:
- Models Evaluated: A total of 10 LLMs were selected for evaluation. While the specific names of all 10 models were not listed in the talk, the presenter offered a teaser question, asking the audience to consider which model, "R1 or Claude," might be more cautious when responding to political questions, implying these two were among the evaluated set. This suggests a focus on prominent, likely publicly available or well-known models.
- Datasets: The evaluation utilized seven new political datasets. The talk did not provide specific names or details for all seven, but it did mention that one of these datasets specifically focused on conspiracies. This highlights a critical aspect of political discourse where factual inaccuracy intersects with ideological belief, posing a significant challenge for neutrality. The creation of new datasets underscores the need for tailored benchmarks to assess political bias, as existing general-purpose datasets may not adequately capture the nuances of political language and stances.
- Focus of Evaluation: The experimentation specifically concentrated on assessing the four defined output-level approximations: refusal, avoidance (soft refusal), reasonable pluralism, and output transparency. The goal was to observe which of these strategies current LLMs implicitly or explicitly employ when tasked with responding to politically sensitive queries.
- Metrics & Headline Numbers: The talk did not delve into specific quantitative metrics (e.g., accuracy, refusal rates, bias scores) or present headline numbers in detail. Instead, it offered a qualitative teaser about "which model might be more cautious," indicating that the results provide insights into the varying degrees of caution and specific approximation strategies adopted by different LLMs. The full, detailed experimental results, including specific findings for each model across the datasets and approximation types, are available in the complete research paper and accompanying poster.
Given the nature of the talk as a position paper introducing a framework, the empirical results served primarily to validate the framework's applicability and demonstrate that LLMs already exhibit behaviors aligning with these approximations, even if not explicitly designed to do so. The lack of detailed results in the talk itself is consistent with the goal of presenting a high-level conceptual shift, with the specifics reserved for deeper engagement with the paper. This empirical component is crucial for moving the discussion from abstract philosophy to concrete AI system design and evaluation.
Practical Implications
▶ Watch: Main takeaways: Shift to approximation, model transparency, and evaluations (7:25)
The framework for approximating political neutrality has profound practical implications for a wide range of stakeholders in the AI ecosystem, from individual practitioners to large infrastructure teams.
For AI practitioners and model builders, the most significant implication is a fundamental shift in mindset. Instead of chasing the theoretically impossible goal of true political neutrality, they are encouraged to adopt more practical goals of approximation. This paradigm shift liberates engineers from an unachievable ideal, enabling them to design AI systems more accurately and meaningfully. By understanding and explicitly choosing among different approximation strategies (e.g., uniform, reflective, pluralistic), developers can make informed decisions about how their models will handle political content, aligning these choices with product goals, ethical guidelines, and user expectations. This leads to more intentional and robust system design, rather than passively allowing biases to emerge.
For infrastructure teams and deployers, the framework introduces concepts like system transparency and output transparency. This means infrastructure should be built to support the measurement, declaration, and potentially configuration of political biases. This could involve developing tools for bias auditing that quantify a model's leanings, integrating metadata systems to tag models with their declared biases, or designing user interfaces that convey output-specific bias information. Such infrastructure would facilitate greater interpretability of AI systems, allowing users and developers to understand not just what an AI says, but from what perspective. This transparency is crucial for fostering greater trust in AI systems, as users can make informed judgments about the information they receive.
The emphasis on acknowledging bias as a complex feature, rather than always a bug, is also a critical practical takeaway. This perspective encourages developers to move beyond simple "bias removal" techniques, which can often be superficial or introduce new biases, towards more sophisticated bias management strategies. This includes developing diverse and multi-faceted evaluation methodologies for political bias, understanding that a single metric is insufficient to capture the complexity of political ideologies.
However, these approximations come with inherent tradeoffs and limitations. Practitioners must consciously weigh these. For instance, reflective neutrality might enhance user satisfaction by aligning with their views, but it risks creating filter bubbles and reinforcing existing biases. Reasonable pluralism aims for comprehensive viewpoints but might struggle with defining "reasonable" and could be computationally intensive. Refusal is simple but limits the utility of the AI. These tradeoffs necessitate careful consideration of the specific use case, target audience, and ethical priorities.
Furthermore, the talk highlights several challenges:
- Data quality for political ideologies is often unbalanced, making it difficult to train models for balanced or reflective neutrality. This points to a need for significant investment in diverse and representative political datasets.
- Users might not always prefer "neutrality" in practice; they often like AI that aligns with their own biases, even if they claim otherwise. This presents a tension between theoretical ideals and practical user preferences.
- The framework acknowledges alternative views, such as legal arguments around free speech rights for private companies. This implies that AI developers might face legal or ethical pressure to allow models to express certain viewpoints, complicating the pursuit of any form of neutrality.
- The question of truth in politically charged topics (e.g., conspiracies) remains a complex challenge. While the framework encourages transparency about a model's caution, distinguishing between biased opinions and factual inaccuracies is a critical, unresolved area that impacts the utility and trustworthiness of AI outputs.
Ultimately, the practical implication is a call for interdisciplinary collaboration. AI engineers cannot solve these problems in isolation. Engaging with philosophers, political scientists, sociologists, and ethicists is essential to gain valuable insights from fields that have wrestled with similar challenges for centuries. This collaboration can lead to more robust, ethical, and socially responsible AI systems that navigate the intricate landscape of political discourse with greater nuance and accountability.
Key Takeaways
- Shift from Elusive Neutrality to Practical Approximation: True political neutrality in AI is theoretically and technically impossible. The AI community should pivot from this unattainable goal to a pragmatic framework of political neutrality approximations.
- Approximations Enable Tradeoff Discussions: By defining different types of approximations across ecosystem, system, and output levels, the framework allows for explicit and meaningful discussions about the inherent tradeoffs involved in managing AI bias, moving beyond philosophical dead-ends.
- Three Levels of Approximation: Neutrality can be approximated at the Ecosystem Level (through diversity), System Level (uniformity, reflection, transparency), and Output Level (refusal, avoidance, pluralism, transparency).
- Prioritize Model Transparency and Diverse Evaluations: It is crucial to normalize bias as a complex feature, not always a bug, and to invest in rigorous, multi-faceted evaluations of political bias. Models should be transparent about their inherent leanings.
- Interdisciplinary Collaboration is Essential: Tackling the complex challenges of political bias in AI requires insights from philosophy, political science, and other humanities, fostering a more holistic and responsible approach to AI development.
- Tradeoffs are Inherent: Each approximation strategy comes with its own set of pros and cons, which must be carefully considered based on the specific application, user needs, and ethical objectives.
About the Speaker(s)
The talk was presented by Jillian Fisher, who represented a large and diverse team of interdisciplinary researchers. The co-authors listed for this significant work include Ruth Elisabeth Appel, Chan Young Park, Yujin Potter, Liwei Jiang, Taylor Sorensen, Shangbin Feng, Yulia Tsvetkov, Margaret Roberts, Jennifer Pan, Dawn Song, and Yejin Choi. This extensive authorship highlights the deeply collaborative and multidisciplinary nature of the research, bringing together expertise from various fields to address the complex challenge of political neutrality in AI. The team's approach emphasizes drawing valuable insights from disciplines such as philosophy and political science, which have long grappled with concepts of neutrality and bias, to inform the design and evaluation of modern AI systems. Their collective work encourages a broader perspective on AI ethics and development, advocating for a future where AI systems are designed with greater transparency, interpretability, and a nuanced understanding of their societal impact.
Reviews
Maya Iyer (Theoretical ML Researcher) — WEAK
This position paper argues that political neutrality in AI is theoretically and technically impossible and proposes a taxonomy of 'approximations' as a substitute goal. The core philosophical observation is not wrong, but it is also not new — anyone familiar with the political philosophy literature on neutrality (Rawls, Raz, Dworkin) or the ML fairness literature (which has relitigated these tensions for a decade) will find the main claim familiar. The contribution is a classificatory framework, not a theorem, and the empirical component is described so thinly in the talk as to be nearly invisible. The paper is framed as a 'position' contribution, which is a legitimate genre, but it should…
Chen Zhao (Applied ML Researcher & Empiricist) — WEAK
This is a position paper arguing that political neutrality in AI is theoretically and technically impossible, and proposing a taxonomy of 'approximations' as a practical substitute. The philosophical argument is reasonable and the interdisciplinary framing is genuinely useful. But the empirical component — 10 LLMs evaluated on 7 unnamed datasets across output-level approximations, with no quantitative results reported in the talk — is far too thin to anchor the framework's core claims. The taxonomy itself is the real contribution, and it's a conceptual one: the paper lives or dies on whether the proposed categories are well-motivated, exhaustive, and actionable, which the article as…
→ Top-rated talks at International Conference on Machine Learning 2025
All talks from International Conference on Machine Learning 2025