The AX Practitioner Playbook is our method for evaluating what AI coding agents do with your technology, diagnosing why they get it wrong, and fixing it at the source. It puts everything we’ve learned from hundreds of agent sessions in one place. Download it at aka.ms/ax-playbook.
Ask a coding agent to build something with the technology you work on, and watch what happens. It generates code that looks reasonable and might even compile. Look closer and you’ll find the wrong SDK version, a deprecated authentication pattern, or a setup nobody on the product team would recommend. The agent didn’t make a random mistake. It did exactly what its training data told it to do.
Developers increasingly let agents pick the SDK, the version, and the pattern, and then judge your technology by the result. Waiting for models to get better isn’t a strategy: a knowledge cutoff tells you little about what a model knows about your product. What you can change are the sources agents rely on: docs, MCP tools, skills, plugins, instructions, CLIs, and APIs. The hard part is knowing which change actually helps.
From a series of articles to one method
Since fall 2025, we in DevRel at Microsoft have measured how agents perform with Azure, Cosmos DB, SharePoint Framework (SPFx), and Microsoft 365 Copilot extensions, using the same kinds of prompts real developers use. We’ve shared what we learned piece by piece in the Agent Experience series. Each article covers one part of the method. The playbook puts the parts together, end to end, so you can run it on your own technology.
What you get from it
The playbook follows an evaluation from setup to shipped fix.
Results you can defend
An evaluation can lie to you convincingly. We’ve seen perfect scores for code that never compiled, and a “used Platform X” check that passed whether or not the agent used Platform X. The playbook’s evaluation model sets out what a result needs before you can trust it, from criteria that judge meaning to gates that prove the code runs.
Criteria that measure what matters
Writing criteria is where your domain expertise becomes measurable, and it’s the hardest step. The playbook shows how to write criteria a judge rules on the same way every time, calibrate them before you trust them, and version them when your product changes. It also covers the traps, like letting a model write your criteria for you.
A shorter path from symptom to cause
A readout tells you what failed while the trajectory tells you why. You’ll learn to tell apart an extension that never loaded, one that loaded but was never called, and one that was called but applied wrong, because each needs a different fix. Nine failure patterns that repeat across every technology we’ve evaluated each point to the surface to inspect first.
Fixes that ship
An evaluation has no impact until someone changes the source of the behavior. The playbook covers how to test each proposed change as a hypothesis before you ship it, and how to bring evidence to the owning team, whether that’s your own team, a team you work with directly, or an open source project you contribute to.
Examples from real evaluations
Every step in the playbook is grounded in scenarios we ran on Cosmos DB, SPFx, and other technologies. Those evaluations led to dozens of shipped fixes to docs and agent extensions, including 46 improvements to the Azure Cosmos DB Agent Kit. The SPFx project upgrade runs end to end in the playbook, from the scenario through the fixes that shipped. When you’re ready to run your own evaluation, a quality checklist helps you review it first.
Who it’s for
The playbook is for anyone who builds technology or extensions that AI coding agents use. If you build an SDK, API, service, CLI, MCP server, skill, plugin, or instruction set, or write the docs behind them, you control what agents discover and apply. If you advocate for a technology rather than own it, you bring the scenario knowledge to evaluate it and the evidence to help the owning team decide what to change.
The method doesn’t depend on a specific evaluation system. You can use one your team built or an existing one, and the playbook lists the capabilities to check for.
Learn it interactively with the AX Practitioner skill
Along with the PDF, we’re releasing the AX Practitioner skill, so you can learn the method while you apply it. Install it in your coding agent and ask it questions as you work on your own evaluation, like how to word a criterion or why the agent ignored your extension. It can also review your scenarios and criteria against the playbook’s standards.
The skill answers only from the playbook. When a question goes beyond it, the skill says so and asks before it uses any other source. Across 330 questions, its answers scored 95% on average against what the playbook says.
Start evaluating
Download the AX Practitioner Playbook and run your first evaluation. To learn as you go, install the AX Practitioner skill and keep it at hand while you work.
You already know what correct looks like for your technology. The playbook shows you how to find out whether agents know it too.
0 comments
Be the first to start the discussion.