Skip to main content
The Quantum Dispatch
Back to Home
Cover illustration for run-assert-eval: How Microsoft Tests AI Agent Fixes

run-assert-eval: How Microsoft Tests AI Agent Fixes

Microsoft’s run-assert-eval links AI agent risk discovery, repeatable tests and runtime policies, helping teams check whether a proposed fix works.

Kai Aegis
Kai Aegis★Sep 26, 2026★3 min read

Microsoft's run-assert-eval connects an AI agent's safety requirements with tests of what that agent actually does. Introduced on September 24, the workflow links risk discovery, evaluation and proposed runtime controls. Its useful promise is a repeatable way to investigate a failure and check a proposed correction.

  • Clarity helps identify risks specific to the application.
  • ASSERT turns specified behaviors into executable evaluations.
  • Agent Control Specification policies can govern runtime actions.
  • Generated controls require human review before the governed test run.

What does run-assert-eval test?

Microsoft's walkthrough starts with the application's risks, develops test cases, runs the agent and examines the resulting evidence. It then proposes a policy and repeats the evaluation under that control. The guidance preserves the test cases and judging setup so the comparison remains meaningful.

The distinction between an unsafe action and an unnecessary refusal is central. A support assistant should protect one customer's information while still helping an authorized customer. Blocking every request would satisfy neither the product's purpose nor a useful safety standard.

This extends the practical testing conversation in our earlier coverage of Microsoft's Clarity safety tooling. The September development is the connected evaluation-and-policy workflow.

How does ASSERT record the evidence?

ASSERT's project documentation describes an evaluation pipeline driven by natural-language requirements. It generates single-turn or multi-turn cases and uses a model-based judge to assess behavior. Trace integration can supply tool calls and other intermediate activity as evidence, going beyond the final answer shown to a user.

Runs produce local artifacts that teams can inspect, compare and retain. The project also includes a viewer for examining results side by side. These are useful building blocks for AI security work, where an explanation of the failure is often as important as its score.

What should a team review before accepting a fix?

Our practical reading is to treat the output as reviewable evidence. Start with the failed case, inspect the attempted action and check whether the proposed control addresses that action. Then examine legitimate requests that should continue to succeed.

An automated judge is an evaluation component, not an independent guarantee. A strong review also asks whether the test set covers the actual application and whether the policy is connected at the point where actions occur. Microsoft's examples demonstrate a method; they do not establish that every agent using the tools is secure. The constructive step forward is making those questions easier to test and revisit.

Sources: Microsoft run-assert-eval introduction — September 24, 2026; ASSERT project documentation — accessed September 26, 2026. These are primary project sources, not an independent security audit.

More Ai Security Stories