3,847 results
Research Engineer/Research Scientist, Pre-training
About Anthropic Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole. Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together to build beneficial AI systems. Anthropic is at the forefront of AI research, dedicated to developing safe, ethical, and powerful artificial intelligence. Our mission is to ensure that transformative AI systems are aligned with human interests. We are seeking a Research Engineer to join our Pre-training team, responsible for developing the next generation of large language models. In this role, you will work at the intersection of cutting-edge research and practical engineering, contributing to the development of safe, steerable, and trustworthy AI systems. Key Responsibilities: - Conduct research and implement solutions in areas such as model architecture, algorithms, data processing, and optimizer development - Independently lead small research projects while collaborating with team members on larger initiatives - Design, run, and analyze scientific experiments to advance our understanding of large language models - Optimize and scale our training infrastructure to improve efficiency and reliability - Develop and improve dev tooling to enhance team productivity - Contribute to the entire stack, from low-level optimizations to high-level model design Qualifications: - Advanced degree (MS or PhD) in Computer Science, Machine Learning, or a related field - Strong software engineering skills with a proven track record of building complex systems - Expertise in Python and experience with deep learning frameworks (PyTorch preferred) - Familiarity with large-scale machine learning, particularly in the context of language models - Ability to balance research goals with practical engineering constraints - Strong problem-solving skills and a results-oriented mindset - Excellent communication skills and ability to work in a collaborative environment - Care about the societal impacts of your work Preferred Experience: - Work on high-performance, large-scale ML systems - Familiarity with GPUs, Kubernetes, and OS internals - Experience with language modeling using transformer architectures - Knowledge of reinforcement learning techniques - Background in large-scale ETL processes You'll thrive in this role if you: - Have significant software engineering experience - Are results-oriented with a bias towards flexibility and impact - Willingly take on tasks outside your job description to support the team - Enjoy pair programming and collaborative work - Are eager to learn more about machine learning research - Are enthusiastic to work at an organization that functions as a single, cohesive team pursuing large-scale AI research projects - Are working to align state of the art models with human values and preferences, understand and interpret deep neural networks, or develop new models to support these areas of research - View research and engineering as
ML Research Engineer, ML Systems
Scale’s ML platform (RLXF) team builds our internal distributed framework for large language model training and inference. The platform has been powering MLEs, researchers, data scientists and operators for fast and automatic training and evaluation of LLM's, as well as evaluation of data quality. Scale is uniquely positioned at the heart of the field of AI as an indispensable provider of training and evaluation data and end-to-end solutions for the ML lifecycle. You will work closely across Scale’s ML teams and researchers to build the foundation platform that supports all our ML research and development. You will be building and optimizing the platform to enable our next generation of LLM training, inference and data curation. If you are excited about shaping the future AI via fundamental innovations, we would love to hear from you! You will: - Build, profile and optimize our training and inference framework - Collaborate with ML teams to accelerate their research and development and enable them to develop the next generation of models and data curation - Research and integrate state-of-the-art technologies to optimize our ML system Ideally you’d have: - Strong excitement about system optimization - Experience with multi-node LLM training and inference - Experience with developing large-scale distributed ML systems - Strong software engineering skills, proficient in frameworks and tools such as CUDA, Pytorch, transformers, flash attention, etc. - Strong written and verbal communication skills and the ability to operate in a cross functional team environment Nice to haves: - Demonstrated expertise in post-training methods &/or next generation use cases for large language models including instruction tuning, RLHF, tool use, reasoning, agents, and multimodal, etc. Compensation packages at Scale for eligible roles include base salary, equity, and benefits. The range displayed on each job posting reflects the minimum and maximum target for new hire salaries for the position and may be inclusive of several career levels at Scale; it will be determined during the interview process based on work location and additional factors, including job-related skills, experience, qualifications, interview performance, and relevant education or training. Scale employees in eligible roles are also granted equity based compensation, subject to Board of Director approval. Your recruiter can share more about the specific salary range for your preferred location during the hiring process, and confirm whether the hired role will be eligible for equity grant. You'll also receive benefits including, but not limited to: comprehensive health, dental and vision coverage, retirement benefits, a learning and development stipend, and generous PTO. Additionally, this role may be eligible for additional benefits such as a commuter stipend. Please reference the job posting's subtitle for where this position will be located. For pay transparency purposes, the base salary range for this full-time position in the locations of San Francisco, New York, Seattle is: $189,600 - $237,000 USD PLEASE NOTE: Our po
Researcher, Alignment Science
ABOUT THE TEAM The Alignment Science team at OpenAI studies the science of intent alignment: how to train models to understand what users are actually asking for, act faithfully on that intent while respecting safety constraints, verify what they did, and report their limitations honestly. Our work sits alongside broader value alignment efforts, but this team focuses on scalable methods for ensuring instruction-following, honesty, and robustness as models become more capable. We work on both sides of alignment research: producing externally publishable results and integrating promising techniques into the models OpenAI deploys. Recent team research on model confessions studies how models can be trained to honestly report shortcomings after their original answer, including failures involving hallucination, instruction following, scheming, and reward hacking. That work reflects a broader agenda: build scalable and general methods to ensure models follow human intent. The team uses a mix of training and evaluation methods, with a focus on reinforcement learning. We care about rigorous, quantitative research that can translate into safer model behavior. ABOUT THE ROLE As a Research Engineer / Research Scientist on the Alignment team, you will design and run experiments that help increasingly capable models follow user intent, remain calibrated about correctness and risk, and honestly surface their own mistakes. You will work on hands-on model training, evaluation design, and research infrastructure, while helping turn promising alignment methods into techniques that can be used in frontier model development. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. We are also open to exceptional remote candidates who can operate independently and collaborate closely with the team. IN THIS ROLE, YOU WILL: - Design and implement alignment experiments focused on intent following, honesty, calibration, and robustness. - Train and evaluate models using reinforcement learning, and other empirical ML methods. - Develop evaluations for failure modes such as hallucination, instruction-following failures, reward hacking, covert actions, and scheming. - Study methods that encourage models to verify their behavior and report shortcomings honestly, including confession-style training objectives. - Build monitoring and inference-time interventions that ensure compliant behavior or surface model issues to users or downstream systems. - Investigate how alignment methods scale with model capability, compute, data, context length, action length, and adversarial pressure. - Integrate successful techniques into model training and deployment workflows. - Produce externally publishable research when results advance the broader science of alignment. - Collaborate with researchers and engineers across post-training, RL, evaluations, safety, and product-facing teams. YOU MIGHT THRIVE IN THIS ROLE IF YOU: - Have strong hands-on experience training, evaluating, or debugging large ML models, especially LLMs. - Have excellent engineering skills in Python and modern ML frameworks such as PyTorch. - Bring mathematical rigor, quantitative taste, and comfort turning ambiguous research questions into measurable experiments. - Have experience with reinforcement learning, post-training, preference optimization, scalable oversight, model evaluation, or adjacent empirical ML research. - Can operate with high independence and do not need close day-to-day handholding. - Enjoy fast-paced, collaborative research environments where priorities shift as models and evidence change. - Have a strong record in technical problem solving, such as competitive programming, math contests, systems work, or similarly rigorous engineering and research projects. - Care about building AI systems that are trustworthy, honest, and reliable in high-stakes settings. - Are motivated by making concrete
Research Engineer, Codex
ABOUT THE TEAM The Codex Research team creates the frontier agents OpenAI ships to the world. We are training the models behind our agents in Codex, ChatGPT, the API, and other frontier products: persistent, proactive intelligence that can operate computers, collaborate with people and other agents, and expand what people and organizations can imagine, attempt, and achieve. We define what the next generation of agents should be able to do, build the training signal that teaches those abilities, and run the experiments that make them real. Our work spans coding, tool use, computer use, multi-agent coordination, long-horizon execution, factuality, instruction following, calibrated reasoning, and taste. Our team is where new model capabilities get made. We build the data, environments, graders, training methods, and feedback loops that shape what OpenAI's next agents can do, then carry those capabilities through major training runs and into the products people use. ABOUT THE ROLE As a member of the Codex Research team, you will improve the capabilities, reliability, and product fit of OpenAI's agentic models. You might own a research direction, build the infrastructure that makes large training runs faster and more trustworthy, create evals that reveal where models fail, or drive a capability from an idea through experimentation, integration, and launch. This role is intentionally broad. The strongest candidates are not defined by one method or subfield; they are people who can take an ambiguous capability problem and make progress across research, engineering, data, evals, and product. You should be excited to work on models that act in the world: writing and debugging code, using tools, calling functions, operating computers, collaborating with other agents, and completing valuable work on behalf of users. You will work with researchers, engineers, product teams, infrastructure teams, and safety/alignment partners to decide what should go into major model runs, measure whether it worked, and ship improvements into products used by real people. This is a high-agency role for people who want their work to land directly in frontier models. IN THIS ROLE, YOU MIGHT: - Design and run experiments that improve agentic model behavior across coding, tool use, function calling, computer use, multi-agent collaboration, long-horizon tasks, factuality, instruction following, and calibrated reasoning. - Own end-to-end improvements to the post-training stack, including RL, data pipelines, graders, reward signals, evals, diagnostics, and model-behavior analysis. - Build evals and environments that expose the next set of model failures, then turn those failures into training data, product fixes, or new research directions. - Partner with Codex, API/platform, ChatGPT, and general-agent product teams to understand what users need and translate product signal into model improvements. - Work on early-training and alignment interventions, including data mixtures, objectives, synthetic data, and eval loops that shape downstream agent behavior. - Help decide which integrations, capabilities, and fixes are ready for inclusion in major model runs. - Improve the machinery for large-scale training and launch: experiment velocity, reliability, observability, reproducibility, cost, latency, and production readiness. - Take on cross-functional projects that touch model training, product infrastructure, and the production agent harness, such as multi-agent systems or training directly against production-like environments. - Debug hard failures in shipped or near-shipped models and turn messy qualitative behavior into concrete hypotheses, experiments, and fixes. YOU MIGHT THRIVE IN THIS ROLE IF YOU: - Have strong technical fundamentals in machine learning, software engineering, systems, statistics, or a related field, and can learn quickly across the parts you have not worked in before. - Have hands-on experience with LLMs, RL, RLHF/RLAIF, post-training, evals,
Model Policy, Chemical & Biological Risk
About the Team Our Safety Systems https://openai.com/safety/safety-systems team is at the forefront of OpenAI's mission to build and deploy safe AGI, driving our commitment to AI safety and fostering a culture of trust and transparency. The Model Policy team aligns model behavior with desired human values and norms. We co-design policy with models and for models by driving rapid policy taxonomy iteration based on data and defining evaluation criteria for foundational models’ ability to reason about safety. Key focus areas include: catastrophic risk, mental health, teen safety and multimodal safety. About the Role Providing access to frontier AI systems raises complex questions around dual-use science and catastrophic risk. How should models respond to requests involving chemical synthesis, biological experimentation, or pathogen research? Where is the boundary between legitimate scientific inquiry and information that could enable misuse? How do we design policies that meaningfully reduce risk without unnecessarily restricting beneficial research? This is a senior role in which you’ll help shape policy creation and development at OpenAI for addressing biological and chemical risks. You will develop structured policy frameworks and taxonomies to guide safe model behavior. This role sits at the intersection of biosecurity expertise, AI safety research, and policy design. You will help ensure that frontier AI systems can support beneficial life sciences research, such as drug discovery, public health, and biosafety, while reducing the risk that these capabilities could be misused. Our relevant publications: - Preparedness framework https://openai.com/index/updating-our-preparedness-framework/ - Preparing for future AI capabilities in biology https://openai.com/index/preparing-for-future-ai-capabilities-in-biology/ - Safety evaluations hub https://openai.com/safety/evaluations-hub/ - OpenAI GPT5 System Card https://openai.com/index/gpt-5-system-card/ - Evaluating Fairness in ChatGPT https://openai.com/index/evaluating-fairness-in-chatgpt/ - Improving Model Safety Behavior with Rule-Based Rewards https://openai.com/index/improving-model-safety-behavior-with-rule-based-rewards/ - OpenAI Model Spec https://openai.com/index/introducing-the-model-spec/ Your Responsibilities: - Design and maintain model policies governing chemical and biological risk, defining how models should safely handle dual-use scenarios. - Develop structured taxonomies of chemical and biological risk that inform model training data, evaluation benchmarks, and safety monitoring systems. - Translate biosecurity and chemical security expertise into actionable model behavior, working closely with research and engineering teams to operationalize policy in training and evaluation pipelines. - Develop a broad range of subject matter expertise while maintaining agility across topics. - Identify emerging risk vectors where frontier AI capabilities could meaningfully lower barriers to harmful activity and develop mitigation strategies. - Engage with internal and external subject-matter experts in biosecurity, biodefense, and chemical safety to ensure policies reflect real-world risk landscapes. You might thrive in this role if you: - Have strong domain expertise in chemistry, biology, biosecurity, or related fields and are motivated to translate that expertise into principled, operational policies that scale to frontier AI systems. - Have experience researching or working with LLMs, machine learning, AI governance, technology policy, or related areas, and enjoy tackling structured reasoning and classification problems—such as defining boundaries between legitimate scientific inquiry and potentially harmful applications. - Have experience designing, refining, or enforcing policies or safeguards for complex systems, whether in AI/ML environments, scientific research governance, national security contexts, or other high-stakes technical domains. - Are comfortable navigating a
Research Engineer, Pretraining Scaling
About Anthropic Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole. Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together to build beneficial AI systems. About the Role: Anthropic's ML Performance and Scaling team trains our production pretrained models, work that directly shapes the company's future and our mission to build safe, beneficial AI systems. As a Research Engineer on this team, you'll ensure our frontier models train reliably, efficiently, and at scale. This is demanding, high-impact work that requires both deep technical expertise and a genuine passion for the craft of large-scale ML systems. This role lives at the boundary between research and engineering. You'll work across our entire production training stack: performance optimization, hardware debugging, experimental design, and launch coordination. During launches, the team works in tight lockstep, responding to production issues that can't wait for tomorrow. Responsibilities: - Own critical aspects of our production pretraining pipeline, including model operations, performance optimization, observability, and reliability - Debug and resolve complex issues across the full stack—from hardware errors and networking to training dynamics and evaluation infrastructure - Design and run experiments to improve training efficiency, reduce step time, increase uptime, and enhance model performance - Respond to on-call incidents during model launches, diagnosing problems quickly and coordinating solutions across teams - Build and maintain production logging, monitoring dashboards, and evaluation infrastructure - Add new capabilities to the training codebase, such as long context support or novel architectures - Collaborate closely with teammates across SF and London, as well as with Tokens, Architectures, and Systems teams - Contribute to the team's institutional knowledge by documenting systems, debugging approaches, and lessons learned You May Be a Good Fit If You: - Have hands-on experience training large language models, or deep expertise with JAX, TPU, PyTorch, or large-scale distributed systems - Genuinely enjoy both research and engineering work—you'd describe your ideal split as roughly 50/50 rather than heavily weighted toward one or the other - Are excited about being on-call for production systems, working long days during launches, and solving hard problems under pressure - Thrive when working on whatever is most impactful, even if that changes day-to-day based on what the production model needs - Excel at debugging complex, ambiguous problems across multiple layers of the stack - Communicate clearly and collaborate effectively, especially when coordinating across time zones or during high-stress incidents - Are passionate about the work itself and want to refine your craft as a research engineer - Care about the societal impacts of AI and responsible scaling Strong Candidates May Also Have: - Previous experience training LLM’s or working extensively with JAX/TPU, PyTorch, or other ML frameworks at scale - Contributed to open-source LLM frame
[Expression of Interest] Research Manager, Interpr...
About Anthropic Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole. Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together to build beneficial AI systems. Note: we don't have open Research Manager positions on the Interpretability team at this time. However, we're actively growing our team of Research Engineers and Research Scientists . If you're excited about interpretability research and open to an individual contributor role, we encourage you to apply. About the Interpretability team When you see what modern language models are capable of, do you wonder, "How do these things work? How can we trust them?" The Interpretability team’s mission is to reverse engineer how trained models work, and Interpretability research is one of Anthropic’s core research bets on AI safety. We believe that a mechanistic understanding is the most robust way to make advanced systems safe. People mean many different things by "interpretability". We're focused on mechanistic interpretability, which aims to discover how neural network parameters map to meaningful algorithms. Some useful analogies might be to think of us as trying to do "biology" or "neuroscience" of neural networks, or as treating neural networks as binary computer programs we're trying to "reverse engineer". We aim to create a solid scientific foundation for mechanistically understanding neural networks and making them safe (see our vision post ). We have focused on resolving the issue of "superposition" (see Toy Models of Superposition , Superposition, Memorization, and Double Descent , and our May 2023 update ), which causes the computational units of the models, like neurons and attention heads, to be individually uninterpretable, and on finding ways to decompose models into more interpretable components. Our subsequent work which found millions of features in Claude 3.0 Sonnet, one of our production language models, represents progress in this direction. In our most recent work , we developed methods that allow us to build circuits using features and use these circuits to understand the mechanisms associated with a model's computation and study specific examples of multi-hop reasoning, planning, and chain-of-thought faithfulness on Claude Haiku 3.5, one of our production models.” This is a stepping stone towards our overall goal of mechanistically understanding neural networks. A few places to learn more about our work and team are this introduction to Interpretability from our research lead, Chris Olah, Stanford CS25 lecture given by Josh Batson, and TWIML AI podcast with E
Research Engineer/Research Scientist, RL/Reasoning
About the Team The RL and Reasoning team drives the core reasoning paradigm and has created groundbreaking innovations such as o1 and o3. They focus on pushing the boundaries of reinforcement learning research, building next-generation generative models, and deploying them at scale. About the Role As a Research Engineer/Research Scientist at OpenAI, you will advance the frontier of AI alignment and capabilities through cutting-edge RL methods. Your work will sit at the heart of training intelligent, aligned, and general-purpose agents, including the systems that power various models. We’re looking for people who have a background in reinforcement learning research, are able to iterate quickly, and are proficient at coding. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. You might thrive in this role if: - You love being on the cutting edge of RL and language model research. - You’re a self-starter who takes initiative and ownership of ideas, driving them to completion. - You value principled approaches, simple experiments in tightly-controlled settings, and reaching trustworthy conclusions which stand the test of time. - You thrive in a fast-paced, dynamic, and technically complex environment where rapid iteration is key. - You’re comfortable diving into a large ML codebase to debug and improve it. - You have a deep understanding of machine learning and machine learning applications. About OpenAI OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. We push the boundaries of the capabilities of AI systems and seek to safely deploy them to the world through our products. AI is an extremely powerful tool that must be created with safety and human needs at its core, and to achieve our mission, we must encompass and value the many different perspectives, voices, and experiences that form the full spectrum of humanity. We are an equal opportunity employer, and we do not discriminate on the basis of race, religion, color, national origin, sex, sexual orientation, age, veteran status, disability, genetic information, or other applicable legally protected characteristic. For additional information, please see OpenAI’s Affirmative Action and Equal Employment Opportunity Policy Statement https://cdn.openai.com/policies/eeo-policy-statement.pdf. Background checks for applicants will be administered in accordance with applicable law, and qualified applicants with arrest or conviction records will be considered for employment consistent with those laws, including the San Francisco Fair Chance Ordinance, the Los Angeles County Fair Chance Ordinance for Employers, and the California Fair Chance Act, for US-based candidates. For unincorporated Los Angeles County workers: we reasonably believe that criminal history may have a direct, adverse and negative relationship with the following job duties, potentially resulting in the withdrawal of a conditional offer of employment: protect computer hardware entrusted to you from theft, loss or damage; return all computer hardware in your possession (including the data contained therein) upon termination of employment or end of assignment; and maintain the confidentiality of proprietary, confidential, and non-public information. In addition, job duties require access to secure and protected information technology systems and related data security obligations. To notify OpenAI that you believe this job posting is non-compliant, please submit a report through this form https://form.asana.com/?d=57018692298241&k=5MqR40fZd7jlxVUh5J-UeA. No response will be provided to inquiries unrelated to job posting compliance. We are committed to providing reasonable accommodations to applicants with disabilities, and requests can be made via this link https://form.asana.com/?k=bQ7w9h3iexRlicUdWRiwvg&d=57018692298241. OpenAI Global
Research Engineer, Pretraining Scaling - London
About Anthropic Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole. Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together to build beneficial AI systems. About the Role: Anthropic's ML Performance and Scaling team trains our production pretrained models, work that directly shapes the company's future and our mission to build safe, beneficial AI systems. As a Research Engineer on this team, you'll ensure our frontier models train reliably, efficiently, and at scale. This is demanding, high-impact work that requires both deep technical expertise and a genuine passion for the craft of large-scale ML systems. This role lives at the boundary between research and engineering. You'll work across our entire production training stack: performance optimization, hardware debugging, experimental design, and launch coordination. During launches, the team works in tight lockstep, responding to production issues that can't wait for tomorrow. Responsibilities: - Own critical aspects of our production pretraining pipeline, including model operations, performance optimization, observability, and reliability - Debug and resolve complex issues across the full stack—from hardware errors and networking to training dynamics and evaluation infrastructure - Design and run experiments to improve training efficiency, reduce step time, increase uptime, and enhance model performance - Respond to on-call incidents during model launches, diagnosing problems quickly and coordinating solutions across teams - Build and maintain production logging, monitoring dashboards, and evaluation infrastructure - Add new capabilities to the training codebase, such as long context support or novel architectures - Collaborate closely with teammates across SF and London, as well as with Tokens, Architectures, and Systems teams - Contribute to the team's institutional knowledge by documenting systems, debugging approaches, and lessons learned You May Be a Good Fit If You: - Have hands-on experience training large language models, or deep expertise with JAX, TPU, PyTorch, or large-scale distributed systems - Genuinely enjoy both research and engineering work—you'd describe your ideal split as roughly 50/50 rather than heavily weighted toward one or the other - Are excited about being on-call for production systems, working long days during launches, and solving hard problems under pressure - Thrive when working on whatever is most impactful, even if that changes day-to-day based on what the production model needs - Excel at debugging complex, ambiguous problems across multiple layers of the stack - Communicate clearly and collaborate effectively, especially when coordinating across time zones or during high-stress incidents - Are passionate about the work itself and want to refine your craft as a research engineer - Care about the societal impacts of AI and responsible scaling Strong Candidates May Also Have: - Previous experience training LLM’s or working extensively with JAX/TPU, PyTorch, or other ML frameworks at scale - Contributed to open-source LLM frame
Machine Learning Engineer, Integrity
About the Team The Integrity team at OpenAI is dedicated to ensuring that our cutting-edge technology is not only revolutionary, but also secure from a myriad of adversarial threats. We strive to maintain the integrity of our platforms as they scale. The Integrity team is at the front lines of defending against misuse in all its forms: content abuse, scaled attacks, and other actions that could undermine the user experience or harm our operational stability. About the Role As a Machine Learning Engineer in OpenAI's Integrity team, you will have the opportunity to work with some of the brightest minds in AI. You’ll work on state-of-the-art models and classifiers, experiment with new architecture and approaches, and push forward our abilities in content and user understanding. You’ll help turn research breakthroughs into tangible solutions that improve the trust and safety of our platform. If you're excited about training LLMs and building ML models, this role is your chance to make a significant mark. In this role, you will: - Innovate and Deploy: Design and deploy advanced machine learning models that solve real-world problems. Bring OpenAI's research from concept to implementation, creating AI-driven applications with a direct impact. - Collaborate with the Best: Work closely with researchers, software engineers, and product managers to understand complex business challenges and deliver AI-powered solutions. Be part of a dynamic team where ideas flow freely and creativity thrives. - Optimize and Scale: Implement scalable data pipelines, optimize models for performance and accuracy, and ensure they are production-ready. Contribute to projects that require cutting-edge technology and innovative approaches. - Learn and Lead: Stay ahead of the curve by engaging with the latest developments in machine learning and AI. Take part in code reviews, share knowledge, and lead by example to maintain high-quality engineering practices. - Make a Difference: Monitor and maintain deployed models to ensure they continue delivering value. Your work will directly influence how AI benefits individuals, businesses, and society at large. You might thrive in this role if you: - Master's/ PhD degree in Computer Science, Machine Learning, Data Science, or a related field. - Demonstrated experience in deep learning and transformers models - Experience with content understanding or abuse prevention with LLMs is a plus - Proficiency in frameworks like PyTorch or Tensorflow - Strong foundation in data structures, algorithms, and software engineering principles. - Are familiar with methods of training and fine-tuning large language models, such as distillation, supervised fine-tuning, and policy optimization - Excellent problem-solving and analytical skills, with a proactive approach to challenges. - Ability to work collaboratively with cross-functional teams. - Ability to move fast in an environment where things are sometimes loosely defined and may have competing priorities or deadlines - Enjoy owning the problems end-to-end, and are willing to pick up whatever knowledge you're missing to get the job done About OpenAI OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. We push the boundaries of the capabilities of AI systems and seek to safely deploy them to the world through our products. AI is an extremely powerful tool that must be created with safety and human needs at its core, and to achieve our mission, we must encompass and value the many different perspectives, voices, and experiences that form the full spectrum of humanity. We are an equal opportunity employer, and we do not discriminate on the basis of race, religion, color, national origin, sex, sexual orientation, age, veteran status, disability, genetic information, or other applicable legally protected characteristic. For additional information, please see OpenAI’s Affirmative Action and Equal Employm
Research Engineer, Machine Learning (Reinforcement...
About Anthropic Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole. Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together to build beneficial AI systems. About the teams Our Reinforcement Learning teams lead Anthropic's reinforcement learning research and development, playing a critical role in advancing our AI systems. We've contributed to all Claude models, with significant impacts on the autonomy and coding capabilities of Claude Sonnet 4.5 and Opus 4.5. Our work spans several key areas: - Developing systems that enable models to use computers effectively - Advancing code generation through reinforcement learning - Pioneering fundamental RL research for large language models - Building scalable RL infrastructure and training methodologies - Enhancing model reasoning capabilities We collaborate closely with Anthropic's alignment and frontier red teams to ensure our systems are both capable and safe. We partner with the applied production training team to bring research innovations into deployed models, and are dedicated to implement our research at scale. Our Reinforcement Learning teams sit at the intersection of cutting-edge research and engineering excellence, with a deep commitment to building high-quality, scalable systems that push the boundaries of what AI can accomplish. About the Role As a Research Engineer within Reinforcement Learning, you will collaborate with a diverse group of researchers and engineers to advance the capabilities and safety of large language models. This role blends research and engineering responsibilities, requiring you to both implement novel approaches and contribute to the research direction. You'll work on fundamental research in reinforcement learning, creating 'agentic' models via tool use for open-ended tasks such as computer use and autonomous software generation, improving reasoning abilities in areas such as mathematics, and developing prototypes for internal use, productivity, and evaluation. Representative projects: - Architect and optimize core reinforcement learning infrastructure, from clean training abstractions to distributed experiment management across GPU clusters. Help scale our systems to handle increasingly complex research workflows. - Design, implement, and test novel training environments, evaluations, and methodologies for reinforcement learning agents which push the state of the art for the next generation of models. - Drive performance improvements across our stack through profiling, optimization, and benchmarking. Implement efficient caching solutions and debug distributed systems to accelerate both training and evaluation workflows. - Collaborate across research and engineering teams to develop automated testing frameworks, design clean APIs, and build scalable infrastructure that accelerates AI research. You may be a good fit if you: - Are proficient in Python and async/concurrent programming with frameworks like Trio - Have experience with machine learning frameworks (PyTorch, TensorFlow, JAX) - Have industry experience in machine learning research - Can balance research exploration with engineering implementation<
Researcher, Frontier Cybersecurity Risks
ABOUT THE TEAM Preparedness is a critical Safety Research team at OpenAI, which is focused on mitigating AI threats to global security https://openai.com/index/updating-our-preparedness-framework/ that could scale to an extreme level of severity. Our work involves: 1. Measurement. Monitoring and predicting the evolving capabilities of frontier AI systems. 2. Mitigation. Keeping misuse safeguards, alignment tools, and security measures on track to adequately address extreme threats that might arise in the future. 3. Coordination. Setting mitigation targets by maintaining OpenAI’s preparedness framework https://openai.com/index/updating-our-preparedness-framework/, and partnering with other staff to achieve these targets. This is urgent, fast-paced work that has far-reaching implications for the company and for society. ABOUT THE ROLE Models are becoming increasingly capable—moving from tools that assist humans to agents that can plan, execute, and adapt in the real world. As we push toward AGI, cybersecurity becomes one of the most important and urgent frontiers: the same systems that can accelerate productivity can also accelerate exploitation. As a Researcher for cybersecurity risks, you will help design and implement an end-to-end mitigation stack to reduce severe cyber misuse across OpenAI’s products. This role requires strong technical depth and close cross-functional collaboration to ensure safeguards are enforceable, scalable, and effective. You’ll contribute directly to building protections that remain robust as products, model capabilities, and attacker behaviors evolve. IN THIS ROLE, YOU WILL: - Design and implement mitigation components for model-enabled cybersecurity misuse—spanning prevention, monitoring, detection, and enforcement—under the guidance of senior technical and risk leadership. - Integrate safeguards across product surfaces in partnership with product and engineering teams, helping ensure protections are consistent, low-latency, and scale with usage and new model capabilities. - Evaluate technical trade-offs within the cybersecurity risk domain (coverage, latency, model utility, and user privacy) and propose pragmatic, testable solutions. - Collaborate closely with risk and threat modeling partners to align mitigation design with anticipated attacker behaviors and high-impact misuse scenarios. - Execute rigorous testing and red-teaming workflows, helping stress-test the mitigation stack against evolving threats (e.g., novel exploits, tool-use chains, automated attack workflows) and across different product surfaces—then iterate based on findings. YOU MIGHT THRIVE IN THIS ROLE IF YOU: - Have a passion for AI safety and are motivated to make cutting-edge AI models safer for real-world use. - Bring demonstrated experience in deep learning and transformer models. - Are proficient with frameworks such as PyTorch or TensorFlow. - Possess a strong foundation in data structures, algorithms, and software engineering principles. - Are familiar with methods for training and fine-tuning large language models, including distillation, supervised fine-tuning, and policy optimization. - Excel at working collaboratively with cross-functional teams across research, security, policy, product, and engineering. - Have significant experience designing and deploying technical safeguards for abuse prevention, detection, and enforcement at scale. - (Nice to have) Bring background knowledge in cybersecurity or adjacent fields. About OpenAI OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. We push the boundaries of the capabilities of AI systems and seek to safely deploy them to the world through our products. AI is an extremely powerful tool that must be created with safety and human needs at its core, and to achieve our mission, we must encompass and value the many different perspectives, voices, and experiences that form the full spectr
Hardware / Software CoDesign Engineer - 3P
About the Team OpenAI’s Hardware organization develops silicon and system-level solutions designed for the unique demands of advanced AI workloads. The team is responsible for building the next generation of AI-native silicon while working closely with software and research partners to co-design hardware tightly integrated with AI models. In addition to delivering production-grade silicon for OpenAI’s supercomputing infrastructure, the team also creates custom design tools and methodologies that accelerate innovation and enable hardware optimized specifically for AI. About the Role As an Engineer on our hardware optimization and co-design team, you will co-design future hardware from different vendors for programmability and performance. You will work with our kernel, compiler and machine learning engineers to understand their unique needs related to ML techniques, algorithms, numerical approximations, programming expressivity, and compiler optimizations. You will evangelize these constraints with various vendors to develop and influence future hardware architectures towards efficient training and inference on our models. If you are excited about efficiently distributing a large language model across devices, dealing with and optimizing system-wide/rack-wide networking bottlenecks and eventually tailoring the compute pipe and memory hierarchy of the hardware platform, simulating workloads at different abstractions and working closely with our partners, this is the perfect opportunity! This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. Key Responsibilities - Co-design future hardware for programmability and performance with our hardware vendors - Assist hardware vendors in developing optimal kernels and add support for it in our compiler - Develop performance estimates for critical kernels for different hardware configurations and drive decisions on compute core and memory hierarchy features - Build system performance models at different abstraction levels and carry out analysis to drive decisions on scale up, scale out, front end networking - Work with machine learning engineers, kernel engineers and compiler developers to understand their vision and needs from high performance accelerators - Manage communication and coordination with internal and external partners - Influence the roadmap of hardware partners to optimize them for OpenAI’s workloads. - Evaluate potential partners’ accelerators and platforms. - As the scope of the role and team grows, understand and influence roadmaps for hardware partners for our datacenter networks, racks, and buildings. Qualifications - 4+ years of industry experience, including experience harnessing compute at scale and optimizing ML platform code to run efficiently on target hardware. - Strong experience in software/hardware co-design - Deep understanding of GPU and/or other AI accelerators - Experience with CUDA, Triton or a related accelerator programming language - Experience driving Machine Learning accuracy with low precision formats - Experience with system performance modeling and analysis to optimize ML model deployment - Strong coding skills in C/C++ and Python - Are familiar with the fundamentals of deep learning computing and chip architecture/microarchitecture. - Able to actively collaborate with ML engineers, kernel writers, compiler developers, system engineers, chip architects/microarchitects Preferred Skills - PhD in Computer Science and Engineering with a specialization in Computer Architecture, Parallel Computing. Compilers or other Systems - Strong understanding of LLMs and challenges related to their training and inference About OpenAI OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. We push the boundaries of the capabilities of AI systems and seek to safely deploy them to the world
Research Engineer, Pretraining
About Anthropic Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole. Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together to build beneficial AI systems. Anthropic is at the forefront of AI research, dedicated to developing safe, ethical, and powerful artificial intelligence. Our mission is to ensure that transformative AI systems are aligned with human interests. We are seeking a Research Engineer to join our Pretraining team, responsible for developing the next generation of large language models. In this role, you will work at the intersection of cutting-edge research and practical engineering, contributing to the development of safe, steerable, and trustworthy AI systems. Key Responsibilities: - Conduct research and implement solutions in areas such as model architecture, algorithms, data processing, and optimizer development - Independently lead small research projects while collaborating with team members on larger initiatives - Design, run, and analyze scientific experiments to advance our understanding of large language models - Optimize and scale our training infrastructure to improve efficiency and reliability - Develop and improve dev tooling to enhance team productivity - Contribute to the entire stack, from low-level optimizations to high-level model design Qualifications: - Advanced degree (MS or PhD) in Computer Science, Machine Learning, or a related field - Strong software engineering skills with a proven track record of building complex systems - Expertise in Python and experience with deep learning frameworks (PyTorch preferred) - Familiarity with large-scale machine learning, particularly in the context of language models - Ability to balance research goals with practical engineering constraints - Strong problem-solving skills and a results-oriented mindset - Excellent communication skills and ability to work in a collaborative environment - Care about the societal impacts of your work Preferred Experience: - Work on high-performance, large-scale ML systems - Familiarity with GPUs, Kubernetes, and OS internals - Experience with language modeling using transformer architectures - Knowledge of reinforcement learning techniques - Background in large-scale ETL processes You'll thrive in this role if you: - Have significant software engineering experience - Are results-oriented with a bias towards flexibility and impact - Willingly take on tasks outside your job description to support the team - Enjoy pair programming and collaborative work &l
Frontier Agents Engineer
About Scale AI Scale AI is the data foundation for AI, helping organizations build and deploy reliable production AI applications. We partner with leading enterprises and government organizations to accelerate their AI initiatives through our data annotation platform, generative AI solutions, and enterprise AI capabilities. Role Overview As a Frontier Agents Engineer on our Enterprise team, you'll be the technical bridge between Scale AI's cutting-edge AI capabilities and our most strategic customers. You'll work with enterprise clients to understand their unique challenges, architect custom AI solutions, and ensure successful deployment and adoption of AI systems in production environments. This is a hands-on technical role that combines deep engineering expertise with customer-facing problem solving. You'll work directly with customer engineering teams to integrate AI into their critical workflows. Key Responsibilities Customer Integration & Deployment - Partner directly with enterprise customers to understand their technical infrastructure, data pipelines, and business requirements - Design and implement custom integrations between Scale AI's platform and customer data environments (cloud platforms, data warehouses, internal APIs) - Build robust data connectors and ETL pipelines to ingest, process, and prepare customer data for AI workflows - Deploy and configure AI models and agents within customer security and compliance boundaries AI Agent Development - Develop production-grade AI agents tailored to customer use cases across domains like customer support, data analysis, content generation, and workflow automation - Architect multi-agent systems that orchestrate between different models, tools, and data sources - Implement evaluation frameworks to measure agent performance and iterate toward business objectives - Design human-in-the-loop workflows and feedback mechanisms for continuous agent improvement Prompt Engineering & Optimization - Create sophisticated prompt engineering strategies optimized for customer-specific domains and data - Build and maintain prompt libraries, templates, and best practices for customer use cases - Conduct systematic prompt experimentation and A/B testing to improve model outputs - Implement RAG (Retrieval Augmented Generation) systems and fine-tuning pipelines where appropriate Technical Leadership & Collaboration - Serve as the primary technical point of contact for strategic enterprise accounts - Collaborate with customer data scientists, ML engineers, and software developers to ensure smooth integration - Provide technical training and knowledge transfer to customer teams - Work closely with Scale's product and engineering teams to translate customer needs into product improvements - Document technical architectures, integration patterns, and best practices Problem Solving & Innovation - Debug complex technical issues across the entire stack, from data pipelines to model outputs - Rapidly prototype solutions to unblock customers and prove out new use cases - St
Research Engineer / Research Scientist, Pre-traini...
About Anthropic Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole. Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together to build beneficial AI systems. About the team We are seeking passionate Research Scientists and Engineers to join our growing Pre-training team in Zurich. We are involved in developing the next generation of large language models. The team primarily focuses on multimodal capabilities: giving LLMs the ability to understand and interact with modalities other than text. In this role, you will work at the intersection of cutting-edge research and practical engineering, contributing to the development of safe, steerable, and trustworthy AI systems. Responsibilities In this role you will interact with many parts of the engineering and research stacks. - Conduct research and implement solutions in areas such as model architecture, algorithms, data processing, and optimizer development - Independently lead small research projects while collaborating with team members on larger initiatives - Design, run, and analyze scientific experiments to advance our understanding of large language models - Optimize and scale our training infrastructure to improve efficiency and reliability - Develop and improve dev tooling to enhance team productivity - Contribute to the entire stack, from low-level optimizations to high-level model design Qualifications & Experience We encourage you to apply even if you do not believe you meet every single criterion. Because we focus on so many areas, the team is looking for both experienced engineers and strong researchers, and encourage anyone along the researcher/engineer spectrum to apply. - Degree (BA required, MS or PhD preferred) in Computer Science, Machine Learning, or a related field - Strong software engineering skills with a proven track record of building complex systems - Expertise in Python and deep learning frameworks - Have worked on high-performance, large-scale ML systems, particularly in the context of language modeling - Familiarity with ML Accelerators, Kubernetes, and large-scale data processing - Strong problem-solving skills and a results-oriented mindset - Excellent communication skills and ability to work in a collaborative environment You'll thrive in this role if you - Have significant software engineering experience - Are able to balance research goals with practical engineering constraints - Are happy to take on tasks outside your job description to support the team - Enjoy pair programming and collaborative work - Are eager to learn more about machine learning research &l
Research Scientist, Safety Post Training
Scale Labs, Research Scientist — Safety Post Training As the leading data and evaluation partner for frontier AI companies, Scale plays an integral role in understanding the capabilities and safeguarding AI models and systems. Building on this expertise, Scale Labs has launched a new team focused on policy research, to bridge the gap between AI research and global policymakers to make informed, scientific decisions about AI risks and capabilities. Our research tackles the hardest problems in agent robustness, AI control protocols, and AI risk evaluations to help governments, industry, and the public understand and mitigate AI risk while maximizing AI adoption. This team collaborates broadly across industry, the public sector, and academia and regularly publishes our findings. We are actively seeking talented researchers to join us in shaping this vision. As a Research Scientist working on Safety Post-Training you will develop and apply post-training methods and interpretability techniques to make frontier AI systems safer, and better understood by researchers and policymakers.. For example, you might: - Design and run post-training pipelines to study how training choices affect model safety, robustness, and alignment properties; - Develop interpretability-informed evaluations that reveal how and why models produce unsafe, deceptive, or otherwise undesirable behaviors, and use those insights to guide targeted mitigations; - Collaborate with policymakers, engineers, and other researchers to translate post-training and interpretability findings into actionable safety standards, evaluation benchmarks, and best practices. Ideally you’d have: - Commitment to our mission of promoting safe, secure, and trustworthy AI deployments in the industry as frontier AI capabilities continue to advance. - Experience with post-training and RL techniques such as RLHF, DPO, GRPO, and similar approaches. - A track record of published research in machine learning, particularly in generative AI. - At least three years of experience addressing sophisticated ML problems, whether in a research setting or in product development. - Strong written and verbal communication skills to operate in a cross-functional team. Nice to have: - Experience with mechanistic interpretability, probing, or other techniques for understanding model internals. - Familiarity with red-teaming or adversarial evaluation of post-trained models. - Experience studying failure modes introduced or masked by post-training, such as reward hacking, sycophancy, or alignment faking. Our research interviews are crafted to assess candidates' skills in practical ML prototyping and debugging, their grasp of research concepts, and their alignment with our organizational culture. We will not ask any LeetCode-style questions. If you’re excited about advancing AI safety and contributing to our mission, we encourage you to apply, even if your experience doesn’t perfectly align with every requirement. Compensation packages at Scale for eligible roles include base salary, equity, and benefits. The range displayed on each job posting reflects the minimum and maximum target for new hire salaries for the position and may be inclusive of several career levels at Scale; it will be determined during the interview process based on work location and additional facto
Data Science Manager, Integrity
ABOUT THE TEAM Integrity Data Science sits at the center of OpenAI’s mission to deploy powerful AI responsibly. We help ensure people can trust our products by building measurement systems, experimentation practices, and detection/mitigation strategies that protect OpenAI and our users from misuse, fraud, and evolving adversarial behaviors. As the scope and urgency of Integrity work expands across product surfaces and go-to-market motion, we’re hiring a dedicated Data Science Manager to scale the team, strengthen execution across multiple Integrity domains, and deepen partnership with Product, Engineering, Operations, and adjacent orgs (e.g., Growth, Ads). This role is based in our San Francisco HQ (in-office). ABOUT THE ROLE As Data Science Manager, Integrity, you will lead a team of data scientists working across trust & safety, fraud prevention, risk analysis, measurement, and modeling. You’ll be accountable for building a high-performing DS function that can keep pace with fast-moving threats—and for shaping the analytical strategy that informs how OpenAI detects, measures, and mitigates integrity risks at scale. This is a highly cross-functional leadership role. You’ll help set the roadmap with Integrity Product/Engineering leaders, evolve team structure and operating rhythms, raise the bar on technical rigor (experimentation, causal inference, modeling, metrics), and develop a culture of proactive, high-leverage impact. Many of the challenges in this space are emergent—new misuse patterns appear as the technology and ecosystem evolves—so this role requires strong judgment, comfort with ambiguity, and an ability to build systems that scale. IN THIS ROLE, YOU WILL: - Lead and scale a high-impact Integrity Data Science team—hiring, coaching, and developing DS ICs (and potentially future managers) while setting a strong technical and cultural bar. - Drive strategy across multiple Integrity domains (policy enforcement, bot detection, fraud prevention, IP theft, risk measurement, abuse prevention), balancing near-term response with durable systems. - Build and institutionalize analytical rigor: clear metric frameworks, experimentation standards, monitoring/alerting, and repeatable evaluation approaches for Integrity interventions. - Partner deeply with Product & Engineering to shape roadmaps, prioritize the right bets, and translate ambiguous risk signals into practical product and platform decisions. - Evolve team structure and operating model as the org scales—defining ownership boundaries, improving processes, and creating leverage through better tooling and AI-assisted workflows. - Enable cross-org outcomes, supporting partners outside Integrity (e.g., Growth, Ads, GTM) where integrity risks intersect with product and business goals. - Communicate clearly with senior leadership, synthesizing complex tradeoffs, surfacing risk, and driving alignment on priorities and success metrics. - Push the team toward an AI-leveraged operating mode, using modern tooling and model capabilities to accelerate detection, triage, analysis, and iteration. YOU MIGHT THRIVE IN THIS ROLE IF YOU: - Have deep experience leading and scaling Data Science teams, ideally in trust & safety, fraud/abuse, security, risk, or other adversarial problem spaces in fast-moving environments. - Bring strong technical grounding across modern DS techniques (experimentation, causal inference, anomaly detection, risk modeling, measurement design) and can coach others to execute with rigor. - Have a track record of building durable partnerships across DS, Engineering, Product, and Operations—able to influence without authority and create shared accountability. - Are excellent at hiring, mentoring, and developing technical talent, and can build a culture that is both high-bar and supportive. - Can translate messy, evolving threats into clear frameworks, metrics, and decisions—and keep the team focused on the highest-leverage work. - Are comfortable operating in ambigu
Anthropic Fellows Program
About Anthropic Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole. Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together to build beneficial AI systems. Apply using this link . Applications for the next cohort of Anthropic Fellows close at 11:59pm PT on July 26 . The cohort is expected to start November 2 . In some circumstances, we can accommodate fellows starting outside the usual cohort timelines — please note in your application if the November start date doesn't work for you. Anthropic Fellows Program overview The Anthropic Fellows Program is designed to foster AI research and engineering talent. We provide funding and mentorship to promising technical talent - regardless of previous experience. Fellows will primarily use external infrastructure (e.g. open-source models, public APIs) to work on an empirical project aligned with our research priorities, with the goal of producing a public output (e.g. a paper submission). In one of our earlier cohorts, over 80% of fellows produced papers. We run multiple cohorts of Fellows each year and review applications on a rolling basis. What to expect - 4 months of full-time research - Direct mentorship from Anthropic researchers - Access to a shared workspace (in either Berkeley, California or London, UK) - Connection to the broader AI safety and security research community - Weekly stipend of 3,850 USD / 2,310 GBP / 4,300 CAD + benefits (these vary by country) - Funding for compute (~$15k/month) and other research expenses Interview process The interview process will include an initial application & reference check, technical assessments & interviews, and a research discussion. We encourage you to apply even if you do not believe you meet every single qualification. Not all strong candidates will meet every single qualification as listed. Research shows that people who identify as being from underrepresented groups are more prone to experiencing imposter syndrome and doubting the strength of their candidacy, so we urge you not to exclude yourself prematurely and to submit an application if you're interested in this work. We think AI systems like the ones we're building have enormous social and ethical implications. We think this makes representation even more important, and we strive to include a range of diverse perspectives on our team. Compensation The expected base stipend for this role is 3,850 USD / 2,310 GBP / 4,300 CAD per week, with an expectation of 40 hours per week for 4 months (with possible extension). Fellows workstreams Due to the success of the Anthropic Fellows for AI Safety Research program, we are now expanding it across teams at Anthropic. We expect there to be significant overlap in the types of skills and responsibilities across the roles and will by default consider candidates for all the workstreams. Some of t
Research Scientist, Agent Robustness
Scale Labs, Research Scientist — Agent Robustness As the leading data and evaluation partner for frontier AI companies, Scale plays an integral role in understanding the capabilities and safeguarding AI models and systems. Building on this expertise, Scale Labs has launched a new team focused on policy research, to bridge the gap between AI research and global policymakers to make informed, scientific decisions about AI risks and capabilities. Our research tackles the hardest problems in agent robustness, AI control protocols, and AI risk evaluations to help governments, industry, and the public understand and mitigate AI risk while maximizing AI adoption. This team collaborates broadly across industry, the public sector, and academia and regularly publishes our findings. We are actively seeking talented researchers to join us in shaping this vision. As a Research Scientist working on Agent Robustness you will work on the fundamental challenges of building AI agents that are safe and aligned with humans. For example, you might: - Research the science of AI agent capabilities with a focus on how they relate to safety, risk factors, and methodologies for benchmarking them; - Design and build harnesses to test AI agents’ tendency to take harmful actions when pressured to do so by users or tricked into doing so by elements of their environment; - Design and build exploits and mitigations for new and unique failure modes that arise as AI agents gain affordances like coding, web browsing, and computer use; - Characterize and design mitigations for potential failure modes or broader risks of systems involving multiple interacting AI agents. Ideally you’d have: - Commitment to our mission of promoting safe, secure, and trustworthy AI deployments in the industry as frontier AI capabilities continue to advance. - Practical experience conducting technical research collaboratively. You should be comfortable building and leveraging agent scaffolding, designing evaluation harnesses, and quickly turning new ideas from the research literature into working prototypes. - Experience with post-training and RL techniques such as RLHF, DPO, GRPO, and similar approaches. - A track record of published research in machine learning, particularly in generative AI. - At least three years of experience addressing sophisticated ML problems, whether in a research setting or in product development. - Strong written and verbal communication skills to operate in a cross-functional team. Nice to have: - Hands-on experience with agent evaluation frameworks such as SWE-bench, WebArena, OSWorld, Inspect, or similar tools. - Experience with red-teaming, prompt injection, or adversarial testing of AI systems. Our research interviews are crafted to assess candidates' skills in practical ML prototyping and debugging, their grasp of research concepts, and their alignment with our organizational culture. We will not ask any LeetCode-style questions. If you’re excited about advancing AI safety and contributing to our mission, we encourage you to apply, even if your experience doesn’t perfectly align with every requirement. Compensation packages at Scale for eligible roles include base salary, equity, and benefits. The range displayed on each job posting reflects the minimum and maximum target for new hire salaries for the position and m
Solutions Engineer, Enterprise
Scale plays a vital role in the development of AI applications. Our customer base is growing exponentially, and you will be on the front lines, ensuring that the world's most innovative companies become passionate, lifelong Scale customers. Solutions Engineers partner closely with AEs, Product, and MLEs to lead prospective customers through pre-sales, delivering customized demos and pilots to secure the “technical win”. Solutions Engineers scope customer technical requirements and develop an actionable SOW. They will work closely with the delivery team to help with initial implementation. Solutions Engineers are relentlessly curious about customer needs and pain points. They employ their expert Scale product knowledge and GenAI knowledge to design solutions that best address these needs. Solutions Engineers are strong relationship builders, great project managers, and provide technical expertise. You will: - Partner with Scale AEs on the customer journey, delivering tailored demos and prototypes according to the customer's requirements. - Develop technical domain expertise in Generative AI / large language model applications for Enterprise use cases, including customers in financial services, insurance, SaaS, and similar enterprises. - Be accountable for securing the “technical win” by unblocking technical challenges - Interact with customers daily to understand their needs and design solutions to better serve them. - Design and develop “Scopes of Work” by breaking down customer challenges into a project plan - Work closely with forward-deployed Software and Machine learning Engineers to develop agents in the initial post-sales stage - Work with AEs and PMs to identify customer-specific feature requests. - Drive strategic initiatives to improve the efficiency and effectiveness of the Solution Engineering team. Ideally, you'd have: - Strong engineering background with prior experience working with clients in a pre or post-sales capacity to realize business goals. - Prior experience developing with Python, Java and/or other web development languages. - Experience working in enterprise SaaS, cloud tech, finance, fintech or similar industries in a technical capacity with end-customer engagement. - A track record as a self-starter, motivated to independently unblock technical issues in the field with the customer, away from the mothership. - Presentation skills with a high degree of technical credibility when speaking with executives and front-line engineers. - High level of comfort communicating effectively across internal and external organizations. - Intellectual curiosity, empathy, and ability to operate with high velocity. Nice to haves: - GenAI Experience - Forward deployed engineering experience - Machine Learning Experience Compensation packages at Scale for eligible roles include base salary, equity, and benefits. The range displayed on each job posting reflects the minimum and maximum target for new hire salaries for the position and may be inclusive of several career levels at Scale; it will be determined during the interview process based on work location and additional factors, including job-related skills, experience, qualifications, interview performance, and relevant education or training. Scale employees in eligible roles are also granted equ
Infrastructure Software Engineer, Enterprise GenAI
Scale GP (Scale Generative AI Platform) is an enterprise-grade AI platform that provides APIs for knowledge retrieval, inference, evaluation, and more. We are looking for a strong engineer to join our team and help us build and scale our core infrastructure in a fast-paced environment. The ideal candidate will have a strong understanding of software engineering principles and practices, as well as experience with large-scale distributed systems. You will implement solutions across multiple cloud providers (GCP, Azure, AWS) for customers in diverse, highly-regulated industries like healthcare, telecom, finance, and retail. What You’ll Do: - Architect multi-cloud systems and abstractions to allow the SGP platform to run on top of existing Cloud providers - Implement custom integrations between Scale AI's platform and customer data environments (cloud platforms, data warehouses, internal APIs) - Collaborate with platform, product teams and our customers directly to develop and implement innovative infrastructure that scales to meet evolving needs. - Deliver experiments at a high velocity and level of quality to engage our customers - Work across the entire product lifecycle from conceptualization through production - Be able, and willing, to multi-task and learn new technologies quickly What We’re Looking For: - 4+ years of full-time engineering experience, post-graduation - Experience scaling products at hyper growth startups - Experience tinkering with or productizing LLMs, vector databases, and the other latest AI technologies - Proficient in Python or Javascript/Typescript, and SQL - Experience with Kubernetes - Experience with major cloud providers (AWS, Azure, GCP) - Excellent communication skills with the ability to explain technical concepts to both technical and non-technical audiences Compensation packages at Scale for eligible roles include base salary, equity, and benefits. The range displayed on each job posting reflects the minimum and maximum target for new hire salaries for the position and may be inclusive of several career levels at Scale; it will be determined during the interview process based on work location and additional factors, including job-related skills, experience, qualifications, interview performance, and relevant education or training. Scale employees in eligible roles are also granted equity based compensation, subject to Board of Director approval. Your recruiter can share more about the specific salary range for your preferred location during the hiring process, and confirm whether the hired role will be eligible for equity grant. You'll also receive benefits including, but not limited to: comprehensive health, dental and vision coverage, retirement benefits, a learning and development stipend, and generous PTO. Additionally, this role may be eligible for additional benefits such as a commuter stipend. Please reference the job posting's subtitle for where this position will be located. For pay transparency purposes, the base salary range for this full-time position in the locations of San Francisco, New York, Seattle is: $179,400 - $224,250 USD PLEASE NOTE:&
Research Engineer, Knowledge Team
About Anthropic Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole. Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together to build beneficial AI systems. About the role: We are looking for Research Engineers to help us redesign how Claude interacts with external data sources. Many of the paradigms for how data and knowledge bases are organized assume human consumers and constraints. This is no longer true in a world of LLMs! Your job will be to design new architectures for how information is organized, and train language models to optimally use those architectures. Responsibilities: - Designing and implementing from scratch new information architecture strategies - Performing finetuning and reinforcement learning to teach language models how to interact with new information architectures - Building “hard” knowledge base eval sets to help identify failure modes of how language models work with external data - Designing and evaluating advanced agentic search capabilities. You may be a good fit if you: - Are a very experienced Python programmer who can quickly produce reliable, high quality code that your teammates love using - Have good machine learning research experience - Have experience developing software that utilizes Large Language Models such as Claude - Are results-oriented, with a bias towards flexibility and impact - Pick up slack, even if it goes outside your job description - Enjoy pair programming (we love to pair!) - Want to partner with world-class ML researchers to develop new LLM capabilities - Care about the societal impacts of your work - Have clear written and verbal communication Strong candidates will also have experience with: - Collaborating with product teams to quickly prototype and deliver innovative solutions - Building complex agentic systems that utilize LLMs - Developing scalable distributed information retrieval systems, such as search engines, knowledge graphs, RAG, indexing, ranking, query understanding, and distributed data processing The annual compensation range for this role is listed below. For sales roles, the range provided is the role’s On Target Earnings ("OTE") range, meaning that the range includes both the sales commissions/sales bonuses target and annual base salary for the role. Annual Salary: $350,000 - $850,000 USD Logistics Minimum education: Bachelor’s degree or an equivalent combination of education, training, and/or experience Required field of study: A field relevant to the role as demonstrated through coursework, training, or professional experience Minimum years of experience: Years of expe
Research Engineer, Interpretability
About Anthropic Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole. Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together to build beneficial AI systems. About the role: When you see what modern language models are capable of, do you wonder, "How do these things work? How can we trust them?" The Interpretability team at Anthropic is working to reverse-engineer how trained models work because we believe that a mechanistic understanding is the most robust way to make advanced systems safe. Think of us as doing "neuroscience" of neural networks using "microscopes" we build - or reverse-engineering neural networks like binary programs. More resources to learn about our work: - Our research blog - covering advances including Monosemantic Features and Circuits - An Introduction to Interpretability from our research lead, Chris Olah - The Urgency of Interpretability from CEO Dario Amodei - Engineering Challenges Scaling Interpretability - directly relevant to this role - 60 Minutes segment - Around 8:07, see a demo of tooling our team built - New Yorker article - what it's like to work on one of AI's hardest open problems Even if you haven’t worked on interpretability before, the infrastructure expertise is similar to what's needed across the lifecycle of a production language model: - Pretraining: Training dictionary learning models looks a lot like model pretraining - creating stable, performant training jobs for massively parameterized models across thousands of chips - Inference: Interp runs a customized inference stack. Day-to-day analysis requires services that allow editing a model's internal activations mid-forward-pass - for example, adding a "steering vector" - Performance: Like all LLM work, we push up against the limits of hardware and software. Rather than squeezing the last 0.1%, we are focused on finding bottlenecks, fixing them and moving ahead given rapidly evolving research and safety mission The science keeps scaling - and it's now applied directly in safety audits on frontier models, with real deadlines. As our research has matured, engineering and infrastructure have become a bottleneck. Your work will have a direct impact on one of the most important open problems in AI. Responsibilities: - Build and maintain the specialized inference and training infrastructure that powers interpretability research - including instrumented forward/backward passes, activation extraction, and steering vector a
ML/Research Engineer, Safeguards
About Anthropic Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole. Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together to build beneficial AI systems. About the role We are looking for ML Engineers and Research Engineers to help detect and mitigate misuse of our AI systems. As a member of the Safeguards ML team, you will build systems that identify harmful use—from individual policy violations to sophisticated, coordinated attacks—and develop defenses that keep our products safe as capabilities advance. You will also work on systems that protect user wellbeing and ensure our models behave appropriately across a wide range of contexts. This work feeds directly into Anthropic's Responsible Scaling Policy commitments. Responsibilities - Develop classifiers to detect misuse and anomalous behavior at scale. This includes developing synthetic data pipelines for training classifiers and methods to automatically source representative evaluations to iterate on - Build systems to monitor for harms that span multiple exchanges, such as coordinated cyber attacks and influence operations, and develop new methods for aggregating and analyzing signals across contexts - Evaluate and improve the safety of agentic products—developing both threat models and environments to test for agentic risks, and developing and deploying mitigations for prompt injection attacks - Conduct research on automated red-teaming, adversarial robustness, and other research that helps test for or find misuse You may be a good fit if you - Have 4+ years of experience in ML engineering, research engineering, or applied research, in academia or industry - Have proficiency in Python and experience building ML systems - Are comfortable working across the research-to-deployment pipeline, from exploratory experiments to production systems - Are worried about misuse risks of AI systems, and want to work to mitigate them - Have strong communication skills and ability to explain complex technical concepts to non-technical stakeholders Strong candidates may also have experience with - Language modeling and transformers - Building classifiers, anomaly detection systems, or behavioral ML - Adversarial machine learning or red-teaming - Interpretability or probes - Reinforcement learning - High-performance, large-scale ML systems The annual compensation range for this role is listed below. For sales roles, the range provided is the role’s On Target Earnings ("OTE") range, meaning that the range includes both the sales commissions/sales bonuses target and annual base salary for the role. Annual Salary: $350,000 - $500,000 USD Logistics Minimum education: Bac
Research Engineer, Performance RL (Reinforcement L...
About Anthropic Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole. Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together to build beneficial AI systems. About the RL Teams Our Reinforcement Learning teams lead Anthropic's reinforcement learning research and development, playing a critical role in advancing our AI systems. We've contributed to all Claude models, with significant impacts on the autonomy and coding capabilities of Claude Sonnet 4.6 and Opus 4.6. Our work spans several key areas: - Developing systems that enable models to use computers effectively - Advancing code generation through reinforcement learning - Pioneering fundamental RL research for large language models - Building scalable RL infrastructure and training methodologies - Enhancing model reasoning capabilities We collaborate closely with Anthropic's alignment and frontier red teams to ensure our systems are both capable and safe. We partner with the applied production training team to bring research innovations into deployed models, and are dedicated to implement our research at scale. Our Reinforcement Learning teams sit at the intersection of cutting-edge research and engineering excellence, with a deep commitment to building high-quality, scalable systems that push the boundaries of what AI can accomplish. About the Role We're hiring for the Code RL team within the RL organization. As a Research Engineer, you'll advance our models' ability to safely write correct, fast code for accelerators. You'll need to know accelerator performance well to turn it into tasks and signals models can learn from. Specifically, you will: - Invent, design and implement RL environments and evaluations. - Conduct experiments and shape our research roadmap. - Deliver your work into training runs. - Collaborate with other researchers, engineers, and performance engineering specialists across and outside Anthropic. You may be a good fit if you: - Have expertise with accelerators (CUDA, ROCm, Triton, Pallas), ML framework programming (JAX or PyTorch). - Have worked across the stack – kernels, model code, distributed systems. - Know how to balance research exploration with engineering implementation. - Are passionate about AI's potential and committed to developing safe and beneficial systems. Strong candidates may also have: - Experience with reinforcement learning. - Experience porting ML workloads between different types of accelerators. - Familiarity with LLM training methodologies. The annual compensation range for this role is listed below. For sales roles, the range provided is th
Research Engineer, Production Model Post-Training
About Anthropic Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole. Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together to build beneficial AI systems. About the role Anthropic's production models undergo sophisticated post-training processes to enhance their capabilities, alignment, and safety. As a Research Engineer on our Post-Training team, you'll train our base models through the complete post-training stack to deliver the production Claude models that users interact with. You'll work at the intersection of cutting-edge research and production engineering, implementing, scaling, and improving post-training techniques like Constitutional AI, RLHF, and other alignment methodologies. Your work will directly impact the quality, safety, and capabilities of our production models. Note: For this role, we conduct all interviews in Python. This role may require responding to incidents on short-notice, including on weekends. Responsibilities: - Implement and optimize post-training techniques at scale on frontier models - Conduct research to develop and optimize post-training recipes that directly improve production model quality - Design, build, and run robust, efficient pipelines for model fine-tuning and evaluation - Develop tools to measure and improve model performance across various dimensions - Collaborate with research teams to translate emerging techniques into production-ready implementations - Debug complex issues in training pipelines and model behavior - Help establish best practices for reliable, reproducible model post-training You may be a good fit if you: - Thrive in controlled chaos and are energised, rather than overwhelmed, when juggling multiple urgent priorities - Adapt quickly to changing priorities - Maintain clarity when debugging complex, time-sensitive issues - Have strong software engineering skills with experience building complex ML systems - Are comfortable working with large-scale distributed systems and high-performance computing - Have experience with training, fine-tuning, or evaluating large language models - Can balance research exploration with engineering rigor and operational reliability - Are adept at analyzing and debugging model training processes - Enjoy collaborating across research and engineering disciplines - Can navigate ambiguity and make progress in fast-moving research environments Strong candidates may also: - Have experience with LLMs - Have a keen interest in AI safety and responsible deployment We welcome candidates at various experience levels, with a preference for senior engineers who have hands-on experience with frontier AI systems. However, proficiency in Python, deep learning frameworks, and distributed computing is required for this role. The annual com
Research Scientist, Life Sciences (Computational)
About Anthropic Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole. Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together to build beneficial AI systems. About the team Anthropic's Life Sciences team is building a world-class research group focused on making fundamental biological discoveries. The team combines cutting-edge AI with hands-on biological research, positioning Anthropic at the forefront of AI-accelerated scientific discovery. About the role We're seeking an exceptional Research Scientist to join the team. This role combines deep computational biology expertise with frontier AI capabilities, positioning Anthropic at the forefront of AI-driven scientific discovery. As one of the first computational members of this Life Sciences research group, you'll work on a high-impact team that operates at the intersection of computational and experimental biology. You'll bring broad computational biology experience to bear across the team's projects, driving discoveries from large-scale computational analysis of biological data through to results our experimental scientists can test, and moving flexibly between problems as the science demands. You'll have substantial access to Claude and you'll help establish how computational biology operates at Anthropic. This role offers a unique opportunity to shape how AI transforms biological research. You'll work with some of the world's best AI researchers while tackling problems that matter deeply for scientific understanding and biomedicine. If you're excited about using your computational expertise to make fundamental biological discoveries and guide the development of transformative AI systems, we want to hear from you. Key responsibilities - Build, run, and maintain the analysis pipelines that back the team's experimental programs: sequence analysis at petabyte scale, structural bioinformatics, phylogenetic and comparative genomics, design and analysis of high-throughput functional screens, biological sequence modeling, etc. - Partner directly with experimental biologists to design experiments that produce high-quality data, and turn results around fast enough to immediately inform the next experiment - Draw on the literature and curated biological knowledge bases alongside primary data to generate and prioritize hypotheses for experimental follow-up - Stand up and maintain the team's computational infrastructure: data ingestion, workflow orchestration, internal databases, and the interfaces that make all of it accessible to both researchers and AI agents <li class="font-claude-response-body whitespace-normal b
AI Success Engineer, Government
ABOUT THE TEAM OpenAI’s AI Success Engineer team partners with the world’s most ambitious government & partner organizations to translate cutting edge AI into real business and mission impact for governments of all levels from Local, State, Federal, and International. We guide customers and users journey from the first time they try ChatGPT Enterprise, automate a workflow, develop and execute a new skill, and create their first agent to scaled enterprise adoption of ChatGPT, Codex, our API and other novel capabilities. Our work spans technical integration and enablement, workflow transformation, inspiring and upskilling AI literacy and confidence across the workforce, sustained program, product and new capability delivery. Most importantly, we help each member of our customer's workforce, their teams, programs and missions meet their total potential. Our government customers have vital missions, and we must meet them with game-changing technology. Every engagement is an opportunity to shape how AI changes work, productivity, and innovation. This role sits at the center of that mission. ABOUT THE ROLE Governments work at a scale that is truly exponential on missions that are of critical importance to people, communities and nations. The AI Success Engineer role is the primary post-sales relationship for OpenAI’s most important customers. You are responsible for the end-to-end account management of critical Government and Partner customers. You will be helping Government Leaders/Partners appropriately and effectively use AI for their mission, while simultaneously investing in ensuring their people are AI-enabled and ready to advance positive outcomes that their constituents depend on them for. You will drive: the impact of our tools on their mission, account health and adoption, ensuring technical readiness, creating and executing on the deployment strategy, enabling, educating and training their workforce, identifying new use cases and upsell opportunities, and delivering measurable value to our customers with OpenAI’s ambitiously growing capabilities. This role blends technical depth, program and account management, customer advisory, training and enablement and product influence. You will partner deeply with customer teams, map workflows, lead configuration, oversee deployment plans, and guide customers toward high impact use cases that showcase the ways OpenAI tools can make a difference to the mission.. You drive our customers’ success and journey in an AI age. You will work closely with Sales, Solutions Architecture, Product, and Research to ensure the customer experience is connected and successful across every touchpoint. Success in this role means accelerating adoption, increasing customer use and value from our tools, guiding strategic use cases that get to production, and helping customers demonstrate tangible business and mission impact. You will bring key product feedback and insights to our product teams to ensure our capabilities continue to advance our customers' mission. IN THIS ROLE, YOU WILL: - Lead the relationship for post-sale customers and act as their trusted advisor on technical deployment, adoption, and value realization, this includes setting up, configuring and running API instances of our products. - Own customer success: account strategy & health; breadth, depth, velocity of adoption that drives mission impact, enablement and education; and ongoing technical deployment and success across your portfolio. - Be an expert in all of OpenAI products across our API and agentic platform, Codex, ChatGPT Enterprise, and more and conduct technical enablement and configuration sessions across them. - Train, educate and enable ChatGPT users to drive adoption and value. - Create and show customers how to make custom GPT’s, Skills, Agents, Plugins, Connectors, Codex and use all of the features and capabilities of our tools. - Design and lead hands-on activities like workshops, hackathons, and training sessions acr
Research Engineer / Research Scientist- Personal A...
About the Team The Personal AGI team is responsible for training and improving pre-trained models to be deployed into ChatGPT, the API, and potential future products. The team partners closely with research and product teams across the company, and conducts research as a final step to prepare for real world deployment to millions of users, ensuring that our models are safe, efficient, and reliable. About the Role As a Research Engineer / Scientist, you will research and develop improvements to our models. Our team works in research areas combining reinforcement learning and products. We're looking for individuals with strong ML engineering skills and research experience, especially with novel and highly capable models. An ideal candidate is passionate about product-driven research. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: - Own and pursue a research agenda to improve model capability and performance. - Collaborate closely with the other research and product teams, allowing customers to optimize their own models. - Build robust evaluations for tracking modeling improvements. - Design, implement, test, and debug code across our research stack. You might thrive in this role if you: - Have a deep understanding of machine learning and machine learning applications. - Have a working knowledge of relevant models, and building evaluations for model capability improvement. - Are comfortable diving into a large ML codebase to debug. - Thrive in a dynamic and technically complex environment. About OpenAI OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. We push the boundaries of the capabilities of AI systems and seek to safely deploy them to the world through our products. AI is an extremely powerful tool that must be created with safety and human needs at its core, and to achieve our mission, we must encompass and value the many different perspectives, voices, and experiences that form the full spectrum of humanity. We are an equal opportunity employer, and we do not discriminate on the basis of race, religion, color, national origin, sex, sexual orientation, age, veteran status, disability, genetic information, or other applicable legally protected characteristic. For additional information, please see OpenAI’s Affirmative Action and Equal Employment Opportunity Policy Statement https://cdn.openai.com/policies/eeo-policy-statement.pdf. Background checks for applicants will be administered in accordance with applicable law, and qualified applicants with arrest or conviction records will be considered for employment consistent with those laws, including the San Francisco Fair Chance Ordinance, the Los Angeles County Fair Chance Ordinance for Employers, and the California Fair Chance Act, for US-based candidates. For unincorporated Los Angeles County workers: we reasonably believe that criminal history may have a direct, adverse and negative relationship with the following job duties, potentially resulting in the withdrawal of a conditional offer of employment: protect computer hardware entrusted to you from theft, loss or damage; return all computer hardware in your possession (including the data contained therein) upon termination of employment or end of assignment; and maintain the confidentiality of proprietary, confidential, and non-public information. In addition, job duties require access to secure and protected information technology systems and related data security obligations. To notify OpenAI that you believe this job posting is non-compliant, please submit a report through this form https://form.asana.com/?d=57018692298241&k=5MqR40fZd7jlxVUh5J-UeA. No response will be provided to inquiries unrelated to job posting compliance. We are committed to providing reasonable accommodations to applicants with disa
Researcher, Automated Red Teaming
ABOUT THE TEAM Preparedness is a critical Safety Research team at OpenAI, which is focused on mitigating AI threats to global security https://openai.com/index/updating-our-preparedness-framework/ that could scale to an extreme level of severity. Our work involves: 1. Measurement. Monitoring and predicting the evolving capabilities of frontier AI systems. 2. Mitigation. Keeping misuse safeguards, alignment tools, and security measures on track to adequately address extreme threats that might arise in the future. 3. Coordination. Setting mitigation targets by maintaining OpenAI’s preparedness framework https://openai.com/index/updating-our-preparedness-framework/, and partnering with other staff to achieve these targets. This is urgent, fast-paced work that has far-reaching implications for the company and for society. ABOUT THE ROLE This role leads the Automated Red Teaming (ART) effort: building scalable, research-driven systems that continuously uncover failure modes in our models and safeguards, and translate those findings into actionable, production-facing improvements. The goal is to reduce expected harm by finding the highest-leverage, least-covered weaknesses early and reliably. IN THIS ROLE, YOU'LL: - Own the research and technical direction for automated red teaming across catastrophic risk areas, with an initial emphasis on: - Automated classifier jailbreak discovery (cyber and bio). - Automated bio threat-development elicitation (worst-feasible planning uplift). - CoT monitoring evasion probing (and adjacent loss-of-control evaluations). - Partner closely with: - Vertical risk teams (Cyber, Bio, Loss of Control) to define threat models, prioritize targets, and land mitigations. - The Classifiers team to turn discovered attacks into training data, evals, and measurable robustness gains. - Product / Engineering / Safety stakeholders to ensure ART outputs are operationally useful. YOU MIGHT THRIVE IN THIS ROLE IF YOU: - Feel a strong pull toward AI safety, and you’re motivated by reducing real-world catastrophic risk (not just publishing cool results). - Love breaking systems (responsibly) — you get energy from finding weird, high-severity failure modes and turning them into concrete fixes. - Have strong applied research instincts, especially around evaluations: you’re good at designing experiments that are reproducible, interpretable, and hard to fool. - Bring hands-on experience with LLMs and agents, including multi-turn behaviors, tool use, and the ways models adapt to constraints. - Are comfortable building scalable automation, not just prototypes — you can turn red-teaming ideas into pipelines that run continuously and produce high-signal outputs. - Have solid software engineering fundamentals (data structures, algorithms, testing discipline) and you can work effectively in a production-adjacent environment. - Think in threat models and incentives, and you naturally ask “what would an attacker do next?” or “how would this fail under pressure?” - Can translate messy findings into action, communicating clearly with researchers, engineers, product, and policy — and driving alignment on what to fix first. - Care about efficiency and prioritization, and you’re happy to say “no” to low-leverage work to focus on what moves the risk needle. - Nice to have: - Experience in adversarial ML, security research / red teaming, abuse prevention systems, or large-scale eval infrastructure. About OpenAI OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. We push the boundaries of the capabilities of AI systems and seek to safely deploy them to the world through our products. AI is an extremely powerful tool that must be created with safety and human needs at its core, and to achieve our mission, we must encompass and value the many different perspectives, voices, and experiences that form the full spectrum of humanity. We are an equal opportun
Research Scientist, Life Sciences
About Anthropic Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole. Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together to build beneficial AI systems. We're seeking an exceptional Research Scientist to join our Life Sciences team at Anthropic. Our team is building a world-class research group focused on making Claude a superhuman life sciences research assistant. This role sits at the intersection of machine learning, software engineering, and biology — you'll directly improve model capabilities on scientific tasks through post-training, evaluation design, and RL environment development. As a core member of our Life Sciences team, you'll work in a high-impact team that translates deep biological domain knowledge into model training objectives, benchmarks, and agentic workflows. You'll help establish Anthropic as a leader in AI-accelerated biology while shaping how frontier models reason about and execute computational biology tasks. This role offers a unique opportunity to shape how frontier AI models learn to do biology. You'll work alongside some of the world's best AI researchers while tackling problems that matter for human health and scientific understanding. If you're excited about turning your computational biology expertise into model capabilities, we want to hear from you. Key Responsibilities - Build and ship agentic tools and integrations that let Claude execute real life science workflows — bioinformatics pipelines, database queries, analysis notebooks, literature review - Design and build evaluation benchmarks that measure model capabilities on biology tasks — figure interpretation, bioinformatics, protocol reasoning, literature synthesis - Work closely with product and design teams to scope, prototype, and ship features for life sciences users - Partner with external biotech, pharma, and academic users to understand their workflows and turn feedback into product improvements - Build and maintain the engineering infrastructure behind our biology product surface — tool scaffolding, data pipelines, eval harnesses - Translate biological domain knowledge into product requirements and evaluation criteria that guide model improvement Minimum Qualifications - Experience applying ML and software engineering to biological problems — computational biology, bioinformatics, protein ML, genomics, or similar - Experience working in drug discovery or development at a biotech or pharma company, or conducted fundamental research in an academic setting — with an understanding of what real scientific workflows look like and where they break down - Strong software engineering skills: comfortable building production-quality Python, working in large codebases, and owning infrastructure end-to-end - Hands-on experience training or fine-tuning ML models (LLMs, protein language models, or other deep learning architectures) - A track record of shipping computational tools or pipelines that biologists actually use - Comfortable navigating ambiguity and defining problems in a rapidly evolving research environme
Research Engineer / Research Scientist -Personal A...
About the Team The Proactivity Research team, within OpenAI’s broader Personal AGI team, is focused on making our models in ChatGPT and future potential products proactive in ways that are truly useful. We're laying the technical foundations for AI that can anticipate what users need in real time, adapt as their goals and preferences shift, and build a deeper, evolving understanding of the person it's helping. About the Role As a Research Engineer / Scientist, you will research and develop improvements to our models’ personalization and agentic capabilities. Our team works on reinforcement learning, dataset creation, evaluations, and other post-training methods. We partner closely with research and product teams across the company to realize the vision of a highly personalized, collaborative, and proactive assistant. We're looking for individuals with strong ML engineering skills and research experience, especially with novel and highly capable models. An ideal candidate is passionate about product-driven research. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: - Own and pursue a research agenda to improve the proactivity and ability of our models to further user goals. - Build robust evaluations for tracking modeling improvements. - Design, implement, test, and debug code across our research stack. - Collaborate closely with the other research and product teams to influence the shape of technical solutions in the product You might thrive in this role if you: - Have a deep understanding of machine learning and machine learning applications. - Have a working knowledge of LLM post-training and evaluation approaches - Are passionate about, or have experience thinking about, personalization and enabling users to achieve their goals - Are comfortable diving into a large ML codebase to debug. - Thrive in a dynamic and technically complex environment. About OpenAI OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. We push the boundaries of the capabilities of AI systems and seek to safely deploy them to the world through our products. AI is an extremely powerful tool that must be created with safety and human needs at its core, and to achieve our mission, we must encompass and value the many different perspectives, voices, and experiences that form the full spectrum of humanity. We are an equal opportunity employer, and we do not discriminate on the basis of race, religion, color, national origin, sex, sexual orientation, age, veteran status, disability, genetic information, or other applicable legally protected characteristic. For additional information, please see OpenAI’s Affirmative Action and Equal Employment Opportunity Policy Statement https://cdn.openai.com/policies/eeo-policy-statement.pdf. Background checks for applicants will be administered in accordance with applicable law, and qualified applicants with arrest or conviction records will be considered for employment consistent with those laws, including the San Francisco Fair Chance Ordinance, the Los Angeles County Fair Chance Ordinance for Employers, and the California Fair Chance Act, for US-based candidates. For unincorporated Los Angeles County workers: we reasonably believe that criminal history may have a direct, adverse and negative relationship with the following job duties, potentially resulting in the withdrawal of a conditional offer of employment: protect computer hardware entrusted to you from theft, loss or damage; return all computer hardware in your possession (including the data contained therein) upon termination of employment or end of assignment; and maintain the confidentiality of proprietary, confidential, and non-public information. In addition, job duties require access to secure and protected information technology systems and relat
Engineering Manager - Machine Learning
Your work will change lives. Including your own. The Impact You’ll Make You will lead a team working to build, scale, and optimize the machine learning infrastructure that powers Recursion's drug discovery platform. From model training pipelines to production deployment systems, to agent infrastructure and Large Language Models, you will ensure our ML models can operate at massive scale across our supercomputing infrastructure, both on prem and in the cloud. You will work cross-functionally across ML engineering, data science, and research teams to translate requirements into robust, scalable ML infrastructure solutions. In This Role You Will: - Enable AI/ML, LLM, and Agentic Systems teams for scale - The ML infrastructure team is responsible for building and operating platforms that allow data scientists and ML engineers to train, deploy, and monitor models across Recursion's massive datasets. With billions of compounds, 30+ petabytes of experimental data, and complex deep learning workloads, your team enables everything from automated compound screening models to clinical trial prediction systems. You will work closely with researchers and ML engineers to understand their infrastructure needs and build scalable solutions for model development, training, and deployment. - Act as a mentor, coach, and sponsor - You will share your technical, leadership and managerial skills in MLOps, distributed computing, and infrastructure engineering, delivering impact, learning, and growth across teams at Recursion. We believe that the best work comes from working across organizational boundaries and you will have opportunities to partner with ML research, platform engineering, and business teams. - Enable a model-driven culture - Machine learning is at the core of everything we do. You will work with stakeholders across the business to ensure our ML infrastructure supports rapid experimentation, reliable model deployment, and continuous improvement. Problems you will work on could range from optimizing GPU cluster utilization to implementing Agentic orchestration and establishing company-wide MLOps standards The Team You’ll Join: You'll be part of a group of technical leaders who work together on the craft of engineering leadership as well as debate ML system architecture, MLOps patterns, and infrastructure optimization strategies. We all work better when we have the support of those around us and are learning together to solve complex problems around model scalability, deployment reliability, and infrastructure efficiency across our teams. You will report to the Executive Director of Engineering who broadly oversees Cloud Infrastructure, High Performance Compute and Machine Learning Infrastructure space. The Experience You Will Need: - Experience in a hands-on technical role as a tech lead or a manager with a focus on infrastructure, MLOps and distributed systems. Excitement for deeply engaging in technical details with your team around machine learning, orchestration and agentic systems. - A people-first mindset. We deliver in a way that prioritizes supporting our coworkers in their growth and experience and understand how Conway's Law shapes our ML system outcomes. - Demonstrated past record of learning from and teaching peers in areas of ML infrastructure, model deploy
Research Engineer/Research Scientist - Personal AG...
About the Team The Personal AGI team seeks to empower all of humanity to benefit from frontier intelligence in whatever way they choose. We are responsible for training models to deploy to millions of users globally via ChatGPT, the API, and future products. We aim to evolve ChatGPT from a chatbot to an infinitely capable and personalized superassistant supporting human flourishing. We work on defining, measuring, and improving capabilities across the training stack. Our focus areas include but are not limited to model behavior, personalization, safety, factuality, instruction following, personality, interactivity, multilingual fluency, world interaction, and bringing agents to everyone. We chart the course for what to strive towards. We partner closely with research and product teams across the company ensuring that our models are safe, efficient, and reliable. About the Role You’ll work as a Research Engineer / Scientist on the North Stars team within the broader Personal AGI research org. You will work on bringing the next generation of AI-enabled experiences to all of humanity by closing the capability overhang between power users and the average consumer, including areas like tool-use, feature discovery, connectors, and instruction following. You will think deeply about the current bottlenecks in model behavior, translate these insights into robust evals, training data, reward signals, and model and harness improvements. We're looking for individuals with strong ML engineering skills and research experience passionate about creative, product-driven research. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: - Own and pursue a research agenda to improve model capability and performance. - Collaborate closely with the other research and product teams, allowing customers to optimize their own models. - Build robust evaluations for tracking modeling improvements. - Design, implement, test, and debug code across our research stack. You might thrive in this role if you: - Have a deep understanding of machine learning and machine learning applications. - Have a working knowledge of relevant models, and building evaluations for model capability improvement. - Are comfortable diving into a large ML codebase to debug. - Thrive in a dynamic and technically complex environment. About OpenAI OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. We push the boundaries of the capabilities of AI systems and seek to safely deploy them to the world through our products. AI is an extremely powerful tool that must be created with safety and human needs at its core, and to achieve our mission, we must encompass and value the many different perspectives, voices, and experiences that form the full spectrum of humanity. We are an equal opportunity employer, and we do not discriminate on the basis of race, religion, color, national origin, sex, sexual orientation, age, veteran status, disability, genetic information, or other applicable legally protected characteristic. For additional information, please see OpenAI’s Affirmative Action and Equal Employment Opportunity Policy Statement https://cdn.openai.com/policies/eeo-policy-statement.pdf. Background checks for applicants will be administered in accordance with applicable law, and qualified applicants with arrest or conviction records will be considered for employment consistent with those laws, including the San Francisco Fair Chance Ordinance, the Los Angeles County Fair Chance Ordinance for Employers, and the California Fair Chance Act, for US-based candidates. For unincorporated Los Angeles County workers: we reasonably believe that criminal history may have a direct, adverse and negative relationship with the following job duties, potentially resulting in the withdrawal of a condit
Engagement Manager
As an Engagement Manager on Scale's Generative AI team, you'll be the face of Scale to the world's leading AI labs and model builders — the labs training the foundational LLM and agent capabilities defining the field. You'll own these customer relationships end to end: translating what researchers need into the complex human-data programs our delivery teams execute, advocating for customer success internally, and earning Scale more work through precise, trustworthy follow-through. You'll work cross-functionally with Operations, ML, Engineering, and Finance to ensure the seamless delivery of Scale's GenAI products and services. You won't run data pipelines yourself — you own the customer relationship, shape requirements into delivery outcomes, and partner closely with the teams who execute them. The ideal candidate is organized, execution-focused, and thrives on managing complexity. Precision across tracking, documentation, and customer follow-through is what builds trust and wins repeat work. What you’ll do: - Own end-to-end account management for assigned frontier-lab customers, from opportunity scoping through project kickoff to completion. - Translate customer requirements into well-defined human-data programs run by Scale's internal delivery operations teams. - Manage customer communications, action items, and follow-ups across multiple concurrent projects with strong attention to detail. - Act as a consultative thought partner to Scale's customers, shaping their needs into clear, deliverable project outcomes. - Monitor delivery against customer throughput and quality standards, proactively flagging risks and escalating blockers with proposed solutions. - Partner cross-functionally with operations, engineering, and planning teams to improve delivery efficiency and drive better customer outcomes. What we’re looking for: - 3+ years of work experience in consulting, technical program management, or project management. - 2+ years of experience in B2B client-facing roles. - Experience in a high-growth environment, working cross-functionally and wearing multiple hats. - A technical background (education or professional experience with CS, Engineering, Economics, Mathematics, or another STEM field) - Enough understanding of the ML training lifecycle to discuss use cases meaningfully with technical customers — or a clear ability to learn it quickly. Compensation packages at Scale for eligible roles include base salary, equity, and benefits. The range displayed on each job posting reflects the minimum and maximum target for new hire salaries for the position and may be inclusive of several career levels at Scale; it will be determined during the interview process based on work location and additional factors, including job-related skills, experience, qualifications, interview performance, and relevant education or training. Scale employees in eligible roles are also granted equity based compensation, subject to Board of Director approval. Your recruiter can share more about the specific salary range for your preferred location during the hiring process, and confirm whether the hired role will be eligible for equity grant. You'll also receive benefits including, but not limited to: comprehensive health, dental and vision coverage, retirement benefits, a learning and development stipend, and generous PTO. Additionally, this role may be eligible for additional benefits such as a commuter stipend.</em&g
Machine Learning Solutions Engineer – Robotics
The next frontier for AI is the physical world. At Scale, we're pioneering this shift, moving artificial intelligence from digital spaces into robotics. Our Robotics team builds the critical infrastructure that empowers the most sophisticated robotic efforts. We are currently seeking a strategic MLSE to join our team and make a direct impact on our platform’s success. As a Machine Learning Solutions Engineer, you'll be a trusted technical partner to the world's most innovative Foundation Model builders and renowned robotics companies. You will partner closely with Product, Sales, and ML Engineers to guide prospective customers through the pre-sales process, delivering customized demos and Proof of Concepts that secure the "technical win" as well as being a critical role in the strategy shaping of our next generation of products covering the full E2E process from data collection to deployments. You’ll help customers bridge the gap from demos to real-world deployment in production by defining technical requirements for multi-modal real-world data pipelines (from data collection, curation, training to deploying models at customer sites and evaluating the performance). You'll develop actionable Statements of Work and collaborate with the delivery team on high-fidelity ground truth implementation. Your expert knowledge of Scale's products will allow you to design creative, impactful solutions. This is a critical role that directly influences multi-million dollar contracts and initiatives. You'll travel globally to conduct on-site technical workshops and scope new projects, while also leading demos and pilots for new prospects. You'll be part of a tight-knit, specialized team, influencing a rapidly growing business that is expanding into new product areas. In this role, you will: - Partner with Scale Account Executives and Engagement Managers to deliver new customer pilots and grow technical relationships with existing clients. - Work with Product Engineering and Product Management to influence our product roadmap based on your frontline insights and help implementing and developing PoC’s - Become a domain expert in next-generation Robotics and physical AI (e.g. VLMs, VLAs, World Models) - Develop technical domain expertise in areas of 2D and 3D imaging and annotation, multi-sensor fusion and calibration, GPS/INS navigation systems, computer vision and other autonomy-adjacent concepts - Be accountable for the technical customer experience and commercial growth, expanding relationships and use cases with existing customers. - Collaborate with highly technical engineers at our customer sites to ensure satisfaction with our data, software platforms, and workflows. - Design and develop playbooks, demos, and other tools to ensure efficient and successful pilots and customer expansions. - Pioneer the development of a global Robotics Data Marketplace, actively seeking out and engaging with key international partners to build a comprehensive data ecosystem. - Evangelize Scale by interacting with customers at major industry events and academic conferences. You have: - PhD in Robotics, or an M.S. in Robotics with a strong track record of deploying VLAs to production. - 3+ years of experience developing with Python, C++, Java, and/or other scripting languages. - Exceptional project management and interpersonal skills, strong attention to detail, and a strong sense of ownership. - The presentation skills and technical credibility to speak confidently with a variety of stakeholders, from execu
Full-Stack Software Engineer, Reinforcement Learni...
About Anthropic Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole. Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together to build beneficial AI systems. About the Role As a Full-Stack Software Engineer in RL, you'll build the platforms, tools, and interfaces that power environment creation, data collection, and training observability. The quality of Claude's next generation depends on the quality of the data we train it on — and the systems you build are what make that data possible. You'll own product surfaces end-to-end — from backend services and APIs to the web UIs that researchers, external vendors, and thousands of data labelers use every day. You don't need a background in ML research. What matters is that you can take an ambiguous, high-stakes problem and ship a polished, reliable product against it, fast. This team moves very quickly. Claude writes a lot of the code we commit, which means the bottleneck isn't typing — it's judgment, taste, and the ability to react to what researchers need next. You'll iterate on data collection strategies to distill the knowledge of thousands of human experts around the world into our models, and you'll do it in a loop that closes in hours and days, not quarters or months. Anthropic's Reinforcement Learning organization leads the research and development that trains Claude to be capable, reliable, and safe. We've contributed to every Claude model, with significant impact on the autonomy and coding capabilities of our most advanced models. Our work spans teaching models to use computers effectively, advancing code generation through RL, pioneering fundamental RL research for large language models, and building the scalable training methodologies behind our frontier production models. The RL org is organized around four goals: solving the science of long-horizon tasks and continual learning, scaling RL data and environments to be comprehensive and diverse, automating software engineering end-to-end, and training the frontier production model. Our engineering teams build the environments, evaluation systems, data pipelines, and tooling that make all of this possible — from realistic agentic training environments and scalable code data generation to human data collection platforms and production training operations. What You'll Do - Build and extend web platforms for RL environment creation, management, and quality review — including environment configuration, versioning, and validation workflows - Develop vendor-facing interfaces and tooling that let external partners create, submit, and iterate on training environments with minimal friction - Design and implement platforms for human data collection at scale, including labeling workflows, quality assurance systems, and feedback mechanisms that surface reward signal integrity issues early - Build evaluation dashboards and observability UIs that give researchers real-time insight into environment quality, training run health, and reward hacking - Create backend services and APIs that connect environment authoring tools, data collection systems, and RL training infrastructure - Build and expand scalable code data generation pipelines, producing diverse programming tasks with robust reward signals across languages and difficulty levels - Develop onboarding automation and documentation tooling so new vendors and internal users r
Engineering Manager, MLE
About the Team The Integrity team at OpenAI is dedicated to ensuring that our cutting-edge technology is not only revolutionary, but also secure from a myriad of adversarial threats. We strive to maintain the integrity of our platforms as they scale. The Integrity team is at the front lines of defending against misuse in all its forms: content abuse, scaled attacks, and other actions that could undermine the user experience or harm our operational stability. About the Role As a Machine Learning Engineer in OpenAI's Integrity team, you will have the opportunity to work with some of the brightest minds in AI. You’ll work on state-of-the-art models and classifiers, experiment with new architecture and approaches, and push forward our abilities in content and user understanding. You’ll help turn research breakthroughs into tangible solutions that improve the trust and safety of our platform. If you're excited about training LLMs and building ML models, this role is your chance to make a significant mark. In this role, you will: - Innovate and Deploy: Design and deploy advanced machine learning models that solve real-world problems. Bring OpenAI's research from concept to implementation, creating AI-driven applications with a direct impact. - Collaborate with the Best: Work closely with researchers, software engineers, and product managers to understand complex business challenges and deliver AI-powered solutions. Be part of a dynamic team where ideas flow freely and creativity thrives. - Optimize and Scale: Implement scalable data pipelines, optimize models for performance and accuracy, and ensure they are production-ready. Contribute to projects that require cutting-edge technology and innovative approaches. - Learn and Lead: Stay ahead of the curve by engaging with the latest developments in machine learning and AI. Take part in code reviews, share knowledge, and lead by example to maintain high-quality engineering practices. - Make a Difference: Monitor and maintain deployed models to ensure they continue delivering value. Your work will directly influence how AI benefits individuals, businesses, and society at large. You might thrive in this role if you: - Master's/ PhD degree in Computer Science, Machine Learning, Data Science, or a related field. - Demonstrated experience in deep learning and transformers models - Experience with content understanding or abuse prevention with LLMs is a plus - Proficiency in frameworks like PyTorch or Tensorflow - Strong foundation in data structures, algorithms, and software engineering principles. - Are familiar with methods of training and fine-tuning large language models, such as distillation, supervised fine-tuning, and policy optimization - Excellent problem-solving and analytical skills, with a proactive approach to challenges. - Ability to work collaboratively with cross-functional teams. - Ability to move fast in an environment where things are sometimes loosely defined and may have competing priorities or deadlines - Enjoy owning the problems end-to-end, and are willing to pick up whatever knowledge you're missing to get the job done About OpenAI OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. We push the boundaries of the capabilities of AI systems and seek to safely deploy them to the world through our products. AI is an extremely powerful tool that must be created with safety and human needs at its core, and to achieve our mission, we must encompass and value the many different perspectives, voices, and experiences that form the full spectrum of humanity. We are an equal opportunity employer, and we do not discriminate on the basis of race, religion, color, national origin, sex, sexual orientation, age, veteran status, disability, genetic information, or other applicable legally protected characteristic. For additional information, please see OpenAI’s Affirmative Action and Equal Employm
Agent Post-Training, Computer Use Research
ABOUT THE TEAM The Agent Post-Training team creates the frontier agents OpenAI ships to the world. We are training the models behind our agents in Codex, ChatGPT, the API, and other frontier products: persistent, proactive intelligence that can operate computers, collaborate with people and other agents, and expand what people and organizations can imagine, attempt, and achieve. We define what the next generation of agents should be able to do, build the training signal that teaches those abilities, and run the experiments that make them real. Our work spans coding, tool use, computer use, multi-agent coordination, long-horizon execution, factuality, instruction following, calibrated reasoning, and taste. Our team is where new model capabilities get made. We build the data, environments, graders, training methods, and feedback loops that shape what OpenAI's next agents can do, then carry those capabilities through major training runs and into the products people use. ABOUT THE ROLE As a member of Agent Post-Training, Computer Use, you will teach models to operate computers. You will help train models that can navigate browsers and desktops, use tools and applications, reason through complex workflows, collaborate with users and other agents, and complete long-horizon tasks with reliability and judgment. This work sits at the intersection of frontier model training, product behavior, evaluation, and systems engineering, and will directly shape the computer-use capabilities shipped in OpenAI’s next generation of agents. Currently, our models are the best in the world at this behavior! You will work with researchers, engineers, product teams, infrastructure teams, and safety/alignment partners to decide what should go into major model runs, measure whether it worked, and ship improvements into products used by real people. This is a high-agency role for people who want their work to land directly in frontier models. IN THIS ROLE, YOU MIGHT - Design and run experiments that improve agentic model behavior for complex computer use https://openai.com/index/codex-for-almost-everything/, including desktop and browser. - Own end-to-end improvements to the post-training stack, including RL, data pipelines, graders, reward signals, evals, diagnostics, and model-behavior analysis. - Build evals and environments that expose the next set of model failures, then turn those failures into training data, product fixes, or new research directions. - Partner with Codex and ChatGPT product teams to understand what users need and translate product signal into model improvements. - Work on early-training and alignment interventions, including data mixtures, objectives, synthetic data, and eval loops that shape downstream agent behavior. - Help decide which integrations, capabilities, and fixes are ready for inclusion in major model runs. - Improve the machinery for large-scale training and launch: experiment velocity, reliability, observability, reproducibility, cost, latency, and production readiness. - Take on cross-functional projects that touch model training, product infrastructure, and the production agent harness, such as multi-agent systems or training directly against production-like environments. - Debug hard failures in shipped or near-shipped models and turn messy qualitative behavior into concrete hypotheses, experiments, and fixes. YOU MIGHT THRIVE IN THIS ROLE IF YOU - Have strong technical fundamentals in machine learning, software engineering, systems, statistics, or a related field, and can learn quickly across the parts you have not worked in before. - Have hands-on experience with LLMs, RL, RLHF/RLAIF, post-training, evals, graders, synthetic data, model training, coding agents, tool-using agents, or production ML systems. - Are excited by open-ended problems where the path is unclear, the signal is noisy, and the right answer requires both research taste and engineering execution. - Care about product impact and model behavior, n
Agent Post-Training, Frontier Evals and Environmen...
ABOUT THE TEAM The Agent Post-Training team creates the frontier agents OpenAI ships to the world. We are training the models behind our agents in Codex, ChatGPT, the API, and other frontier products: persistent, proactive intelligence that can operate computers, collaborate with people and other agents, and expand what people and organizations can imagine, attempt, and achieve. We define what the next generation of agents should be able to do, build the training signal that teaches those abilities, and run the experiments that make them real. Our work spans coding, tool use, computer use, multi-agent coordination, long-horizon execution, factuality, instruction following, calibrated reasoning, and taste. Our team is where new model capabilities get made. We build the data, environments, graders, training methods, and feedback loops that shape what OpenAI's next agents can do, then carry those capabilities through major training runs and into the products people use. ABOUT THE ROLE As a researcher working on Frontier Evals & Environments, you will help build north star model environments to drive progress towards safe AGI/ASI. Your work will directly guide the research programs of the most ambitious training runs happening at OpenAI. Some prior open-sourced evaluations built by researchers in this role include GDPval https://openai.com/index/gdpval/, SWE-bench Verified https://openai.com/index/introducing-swe-bench-verified/, MLE-bench https://openai.com/index/mle-bench/, PaperBench https://openai.com/index/paperbench/, and SWE-Lancer https://openai.com/index/swe-lancer/. If you are interested in feeling firsthand the fast progress of our models, and steering them towards good outcomes, this is the role for you. You will work with researchers, engineers, product teams, infrastructure teams, and safety/alignment partners to decide what should go into major model runs, measure whether it worked, and ship improvements into products used by real people. This is a high-agency role for people who want their work to land directly in frontier models. IN THIS ROLE, YOU MIGHT - Create ambitious RL environments to push our models to their limits, and measure frontier model capabilities, skills, and behaviors - Develop new methodologies for automatically exploring the behavior of these models - Dive deep into the science of measurement, including understanding scalability, reliability, and variance of our evaluation methodology - Help steer training for our largest training runs, and see the future first - Design scalable systems and processes to support continuous evaluation - Build self-improvement loops to automate model understanding YOU MIGHT THRIVE IN THIS ROLE IF YOU - Have strong technical fundamentals in machine learning, software engineering, systems, statistics, or a related field, and can learn quickly across the parts you have not worked in before. - Have hands-on experience with LLMs, RL, RLHF/RLAIF, post-training, evals, graders, synthetic data, model training, coding agents, tool-using agents, or production ML systems. - Are excited by open-ended problems where the path is unclear, the signal is noisy, and the right answer requires both research taste and engineering execution. - Care about product impact and model behavior, not just benchmark movement. You have opinions about what makes an agent useful, reliable, honest, tasteful, and easy to work with. - Can move from a vague behavioral problem to a concrete experiment: define the hypothesis, build the pipeline, run the model, analyze the result, and decide what to do next. - Are comfortable working across research, product, infrastructure, data, evals, and safety boundaries, and can communicate clearly with each group. - Like building load-bearing systems and processes when that is what the team needs, even if the work is not glamorous. - Want to train and ship the models that make agents genuinely useful for developers, enterprises, researchers, and everyday users. Abo
Agent Post-Training, Personality
ABOUT THE TEAM The Agent Post-Training team creates the frontier agents OpenAI ships to the world. We are training the models behind our agents in Codex, ChatGPT, the API, and other frontier products: persistent, proactive intelligence that can operate computers, collaborate with people and other agents, and expand what people and organizations can imagine, attempt, and achieve. We define what the next generation of agents should be able to do, build the training signal that teaches those abilities, and run the experiments that make them real. Our work spans coding, tool use, computer use, multi-agent coordination, long-horizon execution, factuality, instruction following, calibrated reasoning, and taste. Our team builds the data, environments, graders, training methods, and feedback loops that shape what OpenAI’s next agents can do and what they are like to work with, then carries those improvements through major training runs and into products used by people every day. ABOUT THE ROLE As a member of the Agent Post-training Personality team, you will help make OpenAI’s agents exceptional collaborators. You will study what makes an agent thoughtful, clear, perceptive, appropriately proactive, and genuinely easy to work with, then translate those insights into evals, training data, reward signals, and model improvements. We use “personality” to mean much more than writing style or general likability. It includes whether an agent understands what the user is trying to accomplish, communicates with good judgment, adapts to context, asks useful questions, handles disagreement honestly and takes initiative at the right moments. The goal is to create a strong, tasteful default that can adapt to different people and situations. This work combines behavioral research, product thinking, research and communication taste. You will collaborate with product teams, human experts, and researchers across post-training and pretraining to ensure that improvements survive the full training stack and reach the models people use every day. IN THIS ROLE, YOU MIGHT - Develop a rigorous understanding of what makes an agent a great collaborator across professional, creative, technical, and everyday work. - Turn qualitative judgments about model behavior into concrete hypotheses, evals, graders, and training interventions. - Study explicit and implicit user signals to understand which behaviors create trust, satisfaction, continued use, and successful outcomes. - Work with human experts and trainers to produce high-quality, tasteful rollouts and preference data that capture excellent collaborative behavior. - Improve reward models and RL objectives for model behaviors. - Work with pretraining and early-training teams on data mixtures, objectives, synthetic data, and other upstream choices that shape downstream personality. - Build sustainable pipelines for updating older training data as our understanding of excellent model behavior evolves. - Partner closely with ChatGPT, Codex, and other product teams to turn consumer insight into model improvements and validate them in real workflows. - Own projects end to end, from observing a subtle behavioral failure through experimentation, training, evaluation, and launch. YOU MIGHT THRIVE IN THIS ROLE IF YOU - Think instinctively from the user’s perspective and care deeply about how models feel to work with, not only how they perform on benchmarks. - Can translate subjective-seeming product questions into falsifiable hypotheses and rigorous evaluations without losing the nuance that made the question important. - Care about preserving individuality, adaptability, and behavioral diversity rather than optimizing every model toward one narrow style. - Want to shape how frontier agents communicate, collaborate, and build trust with millions of people. - Have strong technical foundations in machine learning, software engineering, statistics, behavioral science, HCI, or a related field, and can quickly learn across u
Backend Software Engineer (Evals)
About the Team The Support Automation team at OpenAI scales the organization by applying cutting-edge AI models to real-world challenges, automating and enhancing work across the organization. From customer operations to engineering, we develop an ecosystem of automation products that empower our colleagues and drive impact. We're passionate about crafting products that serve those around us, blending rapid prototyping with a focus on long-term quality and reliability. By creating reusable solutions, we create patterns that can be applied across diverse domains within OpenAI. TLDR: this team leverages OpenAI technology to improve OpenAI, and you’ll have the opportunity to leverage the full extent of our tech (both public and pre-released) to accomplish this mission. About the Role We’re looking for a Backend Software Engineer with experience working in ML/LLM-heavy domains to help to design and build an evals infrastructure that measures the quality of OpenAI’s support automation. This is a deeply technical and highly cross-functional role where you’ll build robust systems and backend services that serve as the foundation for how knowledge is created, accessed, and applied across OpenAI. The role will especially focus on working closely with Data Science and Research partners to design and build evals at scale. In this role, you will: - Design eval pipelines that are reliable, reproducible, and extendable - Build the infrastructure for continuous eval monitoring frameworks (regression/drift monitoring, building robust golden datasets) along with feedback loops that ultimately strengthen support automation - Design, build, and maintain backend services and APIs to support intelligent automation and knowledge systems - Integrate and structure data across internal platforms, transforming it into formats optimized for use by downstream systems and AI workflows. - Collaborate closely with data, research, and engineering teams to integrate OpenAI models into high-leverage workflows - Own the full development lifecycle of new backend systems and internal platform capabilities - Build with scale and maintainability in mind, while rapidly iterating on new ideas You might be a great fit if you have: - 4+ years of backend engineering experience at product-driven companies (excluding internships) - Proficiency in backend technologies. Our tech stack includes Python, FastAPI, and Postgres - Experience designing and scaling distributed systems, APIs, or data processing pipelines - Have experience building AI agents or applications, including designing evals and improving performance through prompting or scaffolding - Are familiar with evaluation methods for LLMs and have worked with patterns like multi-agent workflows, tool use, or long context. - Experience creating production evals and/or measuring performance of ML/LLM models at scale - A pragmatic mindset. You’re comfortable shipping iteratively while building toward a long-term vision About OpenAI OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. We push the boundaries of the capabilities of AI systems and seek to safely deploy them to the world through our products. AI is an extremely powerful tool that must be created with safety and human needs at its core, and to achieve our mission, we must encompass and value the many different perspectives, voices, and experiences that form the full spectrum of humanity. We are an equal opportunity employer, and we do not discriminate on the basis of race, religion, color, national origin, sex, sexual orientation, age, veteran status, disability, genetic information, or other applicable legally protected characteristic. For additional information, please see OpenAI’s Affirmative Action and Equal Employment Opportunity Policy Statement https://cdn.openai.com/policies/eeo-policy-statement.pdf. Background checks for applicants will be administered in accordan
Machine Learning Engineer, Distributed Data System...
About the Team The OpenAI Robotics team is focused on unlocking general-purpose robotics and pushing towards AGI-level intelligence in dynamic, real-world settings. Working across the entire model stack, we integrate cutting-edge hardware and software to explore a broad range of robotic form factors. We strive to seamlessly blend high-level AI capabilities with the constraints of physical systems to improve peoples’ lives. About the Role As a Research Engineer, Distributed Data Systems, you will design and scale the infrastructure that powers large-scale multimodal training and evaluation at OpenAI. You’ll manage distributed data pipelines, collaborate closely with researchers to translate requirements into robust systems, and harden pipelines that serve as the backbone for OpenAI's rapid iteration cycles. We’re looking for engineers who are detail-oriented, have strong experience with distributed systems, and excel at building reliable infrastructure in high-stakes environments. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: - Design, build, and maintain data infrastructure systems such as distributed compute, data orchestration, distributed storage, streaming infrastructure, machine learning infrastructure while ensuring scalability, reliability, and security. - Ensure our data platform can scale by orders of magnitude while remaining reliable and efficient. - Partner with researchers to deeply understand requirements and translate them into production-ready systems. - Harden, optimize, and maintain critical data infrastructure systems that power multimodal training and evaluation. You might thrive in this role if you: - Have strong experience with distributed systems and large-scale infrastructure with a strong interest in data. - Are detail-oriented and bring rigor to building and maintaining reliable systems. - Demonstrate excellent software engineering fundamentals and organizational skills. - Are comfortable with ambiguity and rapid change. About OpenAI OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. We push the boundaries of the capabilities of AI systems and seek to safely deploy them to the world through our products. AI is an extremely powerful tool that must be created with safety and human needs at its core, and to achieve our mission, we must encompass and value the many different perspectives, voices, and experiences that form the full spectrum of humanity. We are an equal opportunity employer, and we do not discriminate on the basis of race, religion, color, national origin, sex, sexual orientation, age, veteran status, disability, genetic information, or other applicable legally protected characteristic. For additional information, please see OpenAI’s Affirmative Action and Equal Employment Opportunity Policy Statement https://cdn.openai.com/policies/eeo-policy-statement.pdf. Background checks for applicants will be administered in accordance with applicable law, and qualified applicants with arrest or conviction records will be considered for employment consistent with those laws, including the San Francisco Fair Chance Ordinance, the Los Angeles County Fair Chance Ordinance for Employers, and the California Fair Chance Act, for US-based candidates. For unincorporated Los Angeles County workers: we reasonably believe that criminal history may have a direct, adverse and negative relationship with the following job duties, potentially resulting in the withdrawal of a conditional offer of employment: protect computer hardware entrusted to you from theft, loss or damage; return all computer hardware in your possession (including the data contained therein) upon termination of employment or end of assignment; and maintain the confidentiality of proprietary, confidential, and non-public information
ML Research Engineer - Hardware Codesign
ABOUT THE TEAM OpenAI’s Hardware organization develops AI-native silicon and system-level solutions for the unique demands of advanced AI workloads. Building on efforts like Jalapeño, the team is developing future generations of AI-native silicon and tightly integrated systems to power the next generation of frontier models. By co-designing chips, systems, tools, and methodologies, the team helps deliver faster, more efficient, and production-ready hardware for OpenAI’s supercomputing platform. ABOUT THE ROLE We’re seeking a Research-Hardware Codesign Engineer to operate at the boundary between model research and silicon/system architecture. You’ll help shape the numerics, architecture, and technology bets of future OpenAI silicon in collaboration with both Research and Hardware. Your work will include debugging gaps between rooflines and reality, writing quantization kernels, derisking numerics via model evals, quantifying system architecture tradeoffs, and implementing novel numeric RTL. This is a hands-on role for people who go looking for hard problems, get to ground truth, and drive it to production. Strong prioritization and clear, honest communication are essential. Location: San Francisco, CA (Hybrid: 3 days/week onsite) Relocation assistance available. IN THIS ROLE YOU WILL: - Build on our roofline simulator to track evolving workloads, and deliver analyses that quantify the impact of system architecture decisions and support technology pathfinding. - Debug gaps between performance simulation and real measurements; clearly communicate root cause, bottlenecks, and invalid assumptions. - Write emulation kernels for low-precision numerics and lossy compression schemes, and get Research the information they need to trade efficiency with model quality. - Prototype numerics modules by pushing RTL through synthesis; hand off novel numerics cleanly, or occasionally own an RTL module end-to-end. - Proactively pull in new ML workloads, prototype them with rooflines and/or functional simulation, and drive initial evaluation of new opportunities or risks. - Understand the whole picture from ML science to hardware optimization, and slice this end-to-end objective into near-term deliverables. - Build ad-hoc collaborations across teams with very different goals and areas of expertise, and keep progress unblocked. - Communicate design tradeoffs clearly with explicit assumptions and confidence levels; produce a trail of evidence that enables confident execution. YOU WILL THRIVE IN THIS ROLE IF: - An exceptional track record of high-quality technical output, and a bias for shipping a prototype now and iterating later in the absence of clear requirements. - Strong Python, and C++ or Rust, with a cautious attitude toward correctness and an intuition for clean extensibility. - Experience writing Triton, CUDA, or similar, and an understanding of the resulting mapping of tensor ops to functional units. - Working knowledge of PyTorch or JAX; experience in large ML codebases is a plus. - Practical understanding of floating point numerics, the ML tradeoffs of reduced precision, and the current state of the art in model quantization. - Deep understanding of transformer models, and strong intuition for transformer rooflines and the tradeoffs of sharded training and inference in large-scale ML systems. - Experience writing RTL (especially for floating point logic) and understanding of PPA tradeoffs is a plus. - Strong cross-functional communication (e.g. across ML researchers and hardware engineers); ability to slice ambiguous early-incubation ideas into concrete arenas in which progress can be made. To comply with U.S. export control laws and regulations, candidates for this role may need to meet certain legal status requirements as provided in those laws and regulations. About OpenAI OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. We push the boundaries
Recruiter, AI/ML Research
About the Team OpenAI’s mission is to build safe artificial general intelligence (AGI) that benefits all of humanity. Achieving this requires bringing the world’s most exceptional talent under one roof to push the boundaries of what’s possible. Our Research Recruiting team plays a critical role in this effort. We are an embedded part of the research organization, working side by side with our research staff to deeply understand evolving priorities, build trust, and strategically shape the future of OpenAI’s talent. About the Role You will own and execute long-term talent strategies to identify, engage, and recruit many of the world’s leading and emerging AI researchers, research engineers, and technical scientists working at the frontier of machine learning. This is not a traditional execution-focused recruiting role. You will operate as a strategic partner to OpenAI’s research staff, helping define hiring priorities, shape search strategy, influence candidate evaluation, and guide hiring decisions that directly impact the direction and quality of our frontier-model research and fulfillment of our mission. In this role, you will: - Partner directly with research and technical staff to define hiring priorities, shape search strategies, and anticipate future talent needs as technical roadmaps evolve. - Proactively identify and cultivate exceptional AI/ML research talent across industry, academia, and emerging labs, often before formal hiring needs exist. - Use market insights and candidate signals to influence hiring decisions, leveling, and compensation strategy for highly specialized research roles. - Serve as a trusted advisor throughout candidate evaluation and closing — helping leaders calibrate for research excellence, long-term potential, and organizational fit. - Collaborate closely with your sourcing partner to execute complex, high-impact searches in ambiguous or rapidly evolving technical domains. You might thrive in this role if you: - Significant experience recruiting within highly technical or specialized environments. - Deep interest in AI research and a desire to engage directly with global research communities. - Experience recruiting within highly technical or specialized environments such as ML/AI, distributed systems, infrastructure, scientific computing, or quantitative research. - Track record of leading complex, ambiguous technical searches from early talent mapping through close. - Experience navigating high-stakes negotiations with senior technical or research candidates. - Comfort operating in fast-moving environments where hiring priorities and role definitions may evolve over time. Workplace & Location This role is based in our San Francisco office and we aren’t considering remote applications at this time. We use a hybrid work model of 3 days in the office with optional work from home on Thursdays and Fridays. We also offer relocation assistance to new employees. Our open-plan offices have height-adjustable desks, conference rooms, phone booths, well-stocked kitchens full of snacks and drinks, three in-house prepared meals daily, outdoor space for working and socializing, wellness rooms, private bike storage, and more. About OpenAI OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. We push the boundaries of the capabilities of AI systems and seek to safely deploy them to the world through our products. AI is an extremely powerful tool that must be created with safety and human needs at its core, and to achieve our mission, we must encompass and value the many different perspectives, voices, and experiences that form the full spectrum of humanity. We are an equal opportunity employer, and we do not discriminate on the basis of race, religion, color, national origin, sex, sexual orientation, age, veteran status, disability, genetic information, or other applicable legally protected characteristic. For additional i
Recruiter, AI/ML Research EMEA
About the Team OpenAI’s mission is to build safe artificial general intelligence (AGI) that benefits all of humanity. Achieving this requires bringing the world’s most exceptional talent under one roof to push the boundaries of what’s possible. Our Research Recruiting team plays a critical role in this effort. We are an embedded part of the research organization, working side by side with our research staff to deeply understand evolving priorities, build trust, and strategically shape the future of OpenAI’s talent. About the Role You will own and execute long-term talent strategies to identify, engage, and recruit many of the world’s leading and emerging AI researchers, research engineers, and technical scientists working at the frontier of machine learning. This is not a traditional execution-focused recruiting role. You will operate as a strategic partner to OpenAI’s research staff, helping define hiring priorities, shape search strategy, influence candidate evaluation, and guide hiring decisions that directly impact the direction and quality of our frontier-model research and fulfillment of our mission. In this role, you will: - Partner directly with research and technical staff to define hiring priorities, shape search strategies, and anticipate future talent needs as technical roadmaps evolve. - Proactively identify and cultivate exceptional AI/ML research talent across industry, academia, and emerging labs, often before formal hiring needs exist. - Use market insights and candidate signals to influence hiring decisions, leveling, and compensation strategy for highly specialized research roles. - Serve as a trusted advisor throughout candidate evaluation and closing — helping leaders calibrate for research excellence, long-term potential, and organizational fit. - Collaborate closely with your sourcing partner to execute complex, high-impact searches in ambiguous or rapidly evolving technical domains. You might thrive in this role if you: - Significant experience recruiting within highly technical or specialized environments. - Deep interest in AI research and a desire to engage directly with global research communities. - Experience recruiting within highly technical or specialized environments such as ML/AI, distributed systems, infrastructure, scientific computing, or quantitative research. - Track record of leading complex, ambiguous technical searches from early talent mapping through close. - Experience navigating high-stakes negotiations with senior technical or research candidates. - Comfort operating in fast-moving environments where hiring priorities and role definitions may evolve over time. Workplace & Location This role is based in our London office and we aren’t considering remote applications at this time. We use a hybrid work model of 3 days in the office with optional work from home on Thursdays and Fridays. We also offer relocation assistance to new employees. Our open-plan offices have height-adjustable desks, conference rooms, phone booths, well-stocked kitchens full of snacks and drinks, three in-house prepared meals daily, outdoor space for working and socializing, wellness rooms, private bike storage, and more. About OpenAI OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. We push the boundaries of the capabilities of AI systems and seek to safely deploy them to the world through our products. AI is an extremely powerful tool that must be created with safety and human needs at its core, and to achieve our mission, we must encompass and value the many different perspectives, voices, and experiences that form the full spectrum of humanity. We are an equal opportunity employer, and we do not discriminate on the basis of race, religion, color, national origin, sex, sexual orientation, age, veteran status, disability, genetic information, or other applicable legally protected characteristic. For additional informat
Data Operations Manager, Human Data
About Anthropic Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole. Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together to build beneficial AI systems. About the Role As Data Operations Manager, you'll build and scale data operations across research teams working on frontier AI capabilities. You'll partner with researchers to design and execute data strategies, manage vendor relationships, and own the entire data pipeline from requirements to production. This role requires operational excellence combined with technical depth to understand what makes high-quality training data, but your focus will be on strategy and execution. About the Impact The data operations you build will directly determine how well our models perform on critical capabilities—tool use accuracy, prompt injection robustness, long-horizon reasoning, and safety alignment. You'll work with world-class researchers advancing the frontier while building the operational infrastructure to scale these efforts. We're looking for someone who gets excited about the challenge of scaling quality across diverse research areas—someone who can understand nuanced technical requirements, build the right partnerships, and execute flawlessly. If you thrive at the intersection of operational excellence and cutting-edge AI research, we'd love to hear from you. Responsibilities: - Own and execute data strategy for research teams advancing frontier AI capabilities across RLHF, safety, tool use, and agentic workflows - Drive strategic vendor partnerships and build scalable frameworks for technical data collection at scale - Design and implement operational systems that translate research requirements into high-quality data pipelines - Build evaluation frameworks and quality standards that ensure data meets the bar for training state-of-the-art AI systems - Lead cross-functional initiatives to optimize research velocity while maintaining rigorous quality standards - Proactively identify risks, bottlenecks, and opportunities to improve efficiency and effectiveness across data operations - Partner with senior research leaders to align data operations with model development roadmaps and strategic priorities You may be a good fit if you: - Have 3+ years in operations, consulting, product management, or program management roles - Have exceptional project management skills with ability to handle multiple complex projects simultaneously - Have strong communication skills and can engage effectively with technical and non-technical stakeholders - Are familiar with how LLMs work or have strong interest in understanding AI training methodologies - Are highly organized and can navigate ambiguity effectively - Have experience with data analysis tools (SQL, Python, Tableau, spreadsheets, or similar) - Thrive in fast-paced research environments with shifting priorities - Are passionate about AI safety an
Get new jobs by email
A weekly edit of the newest roles. No spam.