From Evaluation to Mitigation: Building Safer Agents

As AI agents become more capable, measuring and mitigating risk is more important than ever. In this episode of Human in the Loop, our guest Mehrnoosh Sameki discusses the shift from evaluation to mitigation, the new open-source measurement framework ASSERT, and the role of agent control systems in building safer and more trustworthy agents.

What You Will Learn:

  • Understand the unique risks posed by autonomous AI agents.

  • Learn why AI agent testing requires more than standard safety checks.

  • Discover how ASSERT helps evaluate AI systems against custom policies and regulations.

  • Explore how context-aware testing improves AI reliability and safety.

  • Learn why sandboxing is essential for safe adversarial testing.

  • See how Agent Control Spec (ACS) enables scalable AI governance and controls.

  • Understand the journey from AI evaluation to mitigation and continuous monitoring.

Guest bio

Mehrnoosh leads the Developer Experiences and Tools product group within Microsoft CoreAI’s Trustworthy AI organization. The team builds tools for generative AI and agentic evaluations across many dimensions, with a particular focus on safety, AI red teaming agents, AI governance tooling through the Foundry Control Plane, and agent control specifications and runtimes.
Beyond this role at Microsoft, she serves as a curriculum developer and lecturer with Break Through Tech, helping underrepresented women pursue meaningful careers in data science.

Enjoy

Chris Huntingford πŸ‘‰ LinkedIn | YouTube

Ioana Tanase πŸ‘‰ LinkedIn

Mehrnoosh Sameki πŸ‘‰ LinkedIn

Next
Next

MCP, Copilot Studio & Responsible AI