Safeguards Enforcement Analyst, Violence & Extremism
Description
About Anthropic Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole. Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together to build beneficial AI systems. About the role As a Safeguards Enforcement Analyst focused on Violence & Extremism, you will be responsible for building and executing operational workflows to assess model behavior, drive enforcement decisions, and develop evals across a technically demanding range of policy areas. Your work spans detecting and mitigating attempts to misuse Anthropic's AI systems to facilitate real-world harm, including weapons and dangerous technology, critical infrastructure attacks, violent extremism, and threats of violence. Important context for this role: In this position you may be exposed to and engage with explicit content spanning a range of topics, including those of a violent, graphic, hateful, or psychologically disturbing nature. Key responsibilities - Design and architect automated enforcement systems and review workflows that scale effectively while maintaining high accuracy - Develop and maintain evals that measure model performance on these policy areas, surface regressions, and inform policy and model improvements - Partner with Engineering and Data Science to optimize detection and automated enforcement systems for potential policy violations - Review flagged content to drive enforcement decisions and surface policy gaps, with particular attention to novel or technically sophisticated misuse attempts + emerging extremist movements, ideologies, and mobilization tactics - Support the Safeguards policy design team by providing structured feedback on policy gaps and enforcement ambiguities based on real enforcement scenarios - Develop and maintain enforcement guidelines and reviewer documentation that enable accurate, consistent enforcement across a wide range of content - Keep up to date with emerging threats, terrorist and extremist movements, regulatory changes, and AI policy enforcement best practices, and apply these to inform our workflows and evals - Identify and escalate emerging misuse patterns, novel attack vectors, and signs of coordinated violent extremist activity Minimum qualifications - Experience in policy enforcement, threat intelligence, counterterrorism, government, or a closely related field, with direct exposure to harmful content, dangerous technology, violent extremism, or physical harm facilitation - Experience standing up and scaling policy enforcement or content review workflows - Proficiency in SQL and/or other data analysis tools to draw insights from large datasets and monitor enforcement workflow health - Experience identifying emerging risks and threat actors, and communicating findings to a diverse set of stakeholders, such as Product, Policy, Engineering, and Legal teams - Experience working with generative AI products, including writing effective prompts for content review and enforcement - Understanding of the challenges involved in implementing product policies at scale, including in the content moderation space Preferred qualifications - Subject ma