tl;dr: For increasingly capable AIs to gain widespread adoption, they will have to earn our trust by reliably doing what we want in the ways we want. Four promising areas for ensuring reliability and corrigibility in AI products are: refining AI behavior, interpreting models, evaluating models, and […] Read more
