Towards Shutdownable Agents: Generalizing Stochastic Choice in RL Agents and LLMs

Back to the Research Library

FIG Fellows: Carissa Cullen, Christos Ziakas

Authors: Elliott Thornley, Carissa Cullen, Christos Ziakas, alexr, LAThomson, Harry Garland

Abstract:

Eleni Angelou, in Artificial Intelligence Safety as an Emerging Paradigm, frames AI safety not as a technological add-on, but as a Kuhnian paradigm shift that should be widely adopted, exploring how safety concerns – both short-term and existential – are reshaping regulatory norms and forcing developers to reimagine risk, uncertainty, and control.

Previous
Previous

Do as I Say, Not as I Do: Instruction-Induction Conflict in LLMs

Next
Next

How's it going? Reinforcement learning in language models recruits a functional welfare axis