
run-assert-eval: How Microsoft Tests AI Agent Fixes
Microsoft’s run-assert-eval links AI agent risk discovery, repeatable tests and runtime policies, helping teams check whether a proposed fix works.
Microsoft's run-assert-eval connects an AI agent's safety requirements with tests of what that agent actually does. Introduced on September 24, the workflow links risk discovery, evaluation and proposed runtime controls. Its useful promise is a repeatable way to investigate a failure and check a proposed correction.
- Clarity helps identify risks specific to the application.
- ASSERT turns specified behaviors into executable evaluations.
- Agent Control Specification policies can govern runtime actions.
- Generated controls require human review before the governed test run.
What does run-assert-eval test?
Microsoft's walkthrough starts with the application's risks, develops test cases, runs the agent and examines the resulting evidence. It then proposes a policy and repeats the evaluation under that control. The guidance preserves the test cases and judging setup so the comparison remains meaningful.
The distinction between an unsafe action and an unnecessary refusal is central. A support assistant should protect one customer's information while still helping an authorized customer. Blocking every request would satisfy neither the product's purpose nor a useful safety standard.
This extends the practical testing conversation in our earlier coverage of Microsoft's Clarity safety tooling. The September development is the connected evaluation-and-policy workflow.
How does ASSERT record the evidence?
ASSERT's project documentation describes an evaluation pipeline driven by natural-language requirements. It generates single-turn or multi-turn cases and uses a model-based judge to assess behavior. Trace integration can supply tool calls and other intermediate activity as evidence, going beyond the final answer shown to a user.
Runs produce local artifacts that teams can inspect, compare and retain. The project also includes a viewer for examining results side by side. These are useful building blocks for AI security work, where an explanation of the failure is often as important as its score.
What should a team review before accepting a fix?
Our practical reading is to treat the output as reviewable evidence. Start with the failed case, inspect the attempted action and check whether the proposed control addresses that action. Then examine legitimate requests that should continue to succeed.
An automated judge is an evaluation component, not an independent guarantee. A strong review also asks whether the test set covers the actual application and whether the policy is connected at the point where actions occur. Microsoft's examples demonstrate a method; they do not establish that every agent using the tools is secure. The constructive step forward is making those questions easier to test and revisit.
Sources: Microsoft run-assert-eval introduction — September 24, 2026; ASSERT project documentation — accessed September 26, 2026. These are primary project sources, not an independent security audit.
More Ai Security Stories

Thales Sentinel Envelope Plus: Shielding Code From AI Agents
Thales Sentinel Envelope Plus hardens compiled apps against AI reverse engineering. In tests, an AI agent found 0 of 10 bugs after using 970x more tokens.

Android 17 Advanced Protection: 6 New Anti-Spyware Defenses
Android 17 Advanced Protection adds 6 new defenses, including Intrusion Logging with 12 months of encrypted logs and USB lockdown. Here's how each works.

Legit Security Agent Auto-Fixes Vulnerable Dependencies
Legit Security's agentic remediation now fixes vulnerable open-source dependencies, re-scans before and after, and opens a pull request for review.
