Researcher, Agent Safety, Training and Evaluations
About the Team The Agent Safety team works to ensure that increasingly capable AI agents act safely, exercise sound judgment, and remain aligned with user intent. Our mission is to reduce the probability of severe unintended outcomes from increasingly capable AI agents while preserving their ability to act effectively and autonomously. Our work spans three areas: - Training: Create training methods, environments and data that teach agents to make better decisions in...