Probing the Preferences of a Language Model: Integrating Verbal and Behavioral Tests of AI Welfare

Back to the Research Library

FIG Fellows: Eleni Angelou

Authors: Eleni Angelou

Abstract:

Eleni Angelou, in Artificial Intelligence Safety as an Emerging Paradigm, frames AI safety not as a technological add-on, but as a Kuhnian paradigm shift that should be widely adopted, exploring how safety concerns – both short-term and existential – are reshaping regulatory norms and forcing developers to reimagine risk, uncertainty, and control.

Previous
Previous

Lessons from Source Code Inspection Facilities for AI Verification

Next
Next

A timing problem for instrumental convergence