Long-term risks from ideological fanaticism

Back to the Research Library

FIG Fellows: Will Aldred

Authors: David Althaus, Jamie Harris, Vanessa Sarre, Clare Diane, Will Aldred

Abstract:

  • History’s most destructive ideologies—like Nazism, totalitarian communism, and religious fundamentalism—exhibited remarkably similar characteristics:

    • epistemic and moral certainty

    • extreme tribalism dividing humanity into a sacred “us” and an evil “them”

    • a willingness to use whatever means necessary, including brutal violence.

  • Such ideological fanaticism was a major driver of eight of the ten greatest atrocities since 1800, including the Taiping Rebellion, World War II, and the regimes of Stalin, Mao, and Hitler.

  • We focus on ideological fanaticism over related concepts like totalitarianism partly because it better captures terminal preferences, which plausibly matter most as we approach superintelligent AI and technological maturity.

  • Ideological fanaticism is considerably less influential than in the past, controlling only a small fraction of world GDP. Yet at least hundreds of millions still hold fanatical views, many regimes exhibit concerning ideological tendencies, and the past two decades have seen widespread democratic backsliding.

  • The long-term influence of ideological fanaticism is uncertain. Fanaticism faces many disadvantages including a weak starting position, poor epistemics, and difficulty assembling broad coalitions. But it benefits from greater willingness to use extreme measures, fervent mass followings, and a historical tendency to survive and even thrive amid technological and societal upheaval. Beyond complete victory or defeat, multipolarity may persist indefinitely, with fanatics permanently controlling a non-trivial fraction of the universe, potentially using superintelligent AI to entrench their rule.

  • Ideological fanaticism increases existential risks and risks of astronomical suffering through multiple mutually-reinforcing pathways.

    • Ideological fanaticism exacerbates most common causes of war. Fanatics' sacred values and outgroup hostility often preclude compromise, while their irrational overconfidence and differential commitment credibility make bargaining failures more likely. Fanatics may even welcome conflict, rather than viewing it as a costly last resort.

    • Fanatical retributivism may lead to astronomical suffering. In our survey of 1,084 people, 11–14% in the US, UK, and Pakistan agreed that if hell didn't exist, we should create it to punish evil people with extreme suffering forever, and separately selected 'forever' when asked how long evil people should suffer unbearable pain, while also stating that at least 1% of humanity deserves this fate. Rates ranged from 19–25% in China, Saudi Arabia, and Turkey. Similar questions showed roughly comparable patterns. Advanced AI could enable fanatics to actually instantiate such preferences.

    • Certain of their righteousness, fanatics resist further reflection and seek to lock in their current values, which threatens long-reflection-style proposals that envision humanity carefully deliberating on how to achieve its potential. Viewing compromise and cooperation as betrayal, fanatics also seem more likely to oppose moral trade and use hostile bargaining tactics. Their intolerant ‘fussy’ preferences may regard almost all configurations of matter as immoral, including those containing vast flourishing, potentially resulting in astronomical waste.

    • AI intent alignment alone won't help if the human principal is fanatical or malevolent: an AI aligned with Stalin probably won't usher in utopia. Fanatics may reflectively endorse their existing values, even after preference idealization. The worst futures may therefore arise from misuse of intent-aligned AI by ideological fanatics, rather than from misaligned AI.

    • Ideological fanaticism also poses other risks, including extreme optimization and differential intellectual regress.

  • Most relevant interventions, while not novel, fall into two overlapping categories.

    • Political and societal interventions include strengthening and safeguarding liberal democracies, reducing political polarization, promoting anti-fanatical principles like classical liberalism, and fostering international cooperation.

    • AI-related interventions appear higher-leverage. Compute governance and information security can reduce the likelihood that transformative AI falls into the hands of fanatical and malevolent actors. Preventing AI-enabled coups could be particularly important given such actors' propensity for power grabs. Other promising interventions include proactively using AI to improve epistemics at scale, developing fanaticism-resistant post-AGI governance frameworks, and making transformative AIs themselves less fanatical—e.g., by guiding their character towards wisdom and benevolence.

Previous
Previous

A Progressive Global Corporate Tax For the Age of AI

Next
Next

Beyond Mimicry: Preference Coherence in LLMs