SECURITY Signal 56
My recommended resources for AI safety, alignment, and existential risks
Illustration only Photo by Mick Haupt on Unsplash
elseif has not written about this yet · Lesswrong describes it this way
I have studied AI safety and existential risks for the last four years and, in this post, I share the resources I liked reading/watching the most. I have read/watched all of them. A few of these resources are in French. Not all the articles are peer-reviewed. I will update this post over time.Articles I liked the mostAI safety via debateSafe uses of AI OraclesThe Off-Switch GameFormalizing Two Problems of Realistic World-ModelsGoal Misgeneralization in Deep Reinforcement LearningConcrete Problems in AI SafetyThe Value Learning ProblemPenalizing side effects using stepwise relative reachabilityEmbedded AgencyAI Safety GridworldsConservative AgencyParametrically Retargetable Decision-Makers Tend To Seek PowerTiling Agents for Self-Modifying AI, and the Löbian ObstacleProgram Equilibrium in the Prisoner’s Dilemma via Löb's TheoremAGI Safety Literature ReviewFormalizing Convergent Instrumental GoalsEliciting Latent Knowledge: How to tell if your eyes deceive youCooperative Inverse Reinforcement LearningYouTube channels and videosRational AnimationsRobert Miles AI SafetyAI In ContextSciencePetrMonsieur PhiSiliconversationsThe video Writing Doom – Award-Winning Short Film on Superintelli
THE CLUSTER