October 6th, 2026
0 reactions

What is Agent Experience (AX)?

Principal Developer Advocate

Agent Experience is the experience AI agents have when discovering, choosing, and using your technology. Practicing AX means systematically improving that experience. Measurement is how you know whether you did.

Developers now ask AI agents to build with your SDK, API, or CLI, and they judge your technology by what the agent produces. If the agent picks a competing service, calls an API you deprecated, or reports success on a broken build, your technology looks broken, even when your docs are excellent for people. Mathias Biilmann of Netlify coined the term in January 2025 and described it as “the holistic experience AI agents will have as the user of a product or platform.” It applies to any AI agent, from a coding agent in your editor to ChatGPT calling your MCP server.

What Agent Experience measures

AX comes down to two questions. The first is propensity: does the agent find your technology and choose it? Give the agent a task that doesn’t name your product, like adding email sending to an app, and see what it reaches for. The second is efficacy: once the agent uses your technology, does it use it correctly? Name your product in the task, and see whether the agent calls your current API or one you deprecated. Your eval tool might use other names, like discovery and selection for propensity, or quality for efficacy.

For both, you score what each run produced and what it cost. Cost is easy to skip and easy to get wrong. For example, Claude Sonnet 5 launched with 33% lower per-token pricing than Sonnet 4.6. Yet when upgrading SharePoint Framework (SPFx) projects, Microsoft’s framework for building Microsoft 365 Copilot, SharePoint and Teams extensions, it cost 3.7x more per run in GitHub Copilot Chat, across 3 scenarios and 15 runs per model (Not all model upgrades are upgrades).

So to recap: propensity tells you whether agents reach your technology. Efficacy tells you what happens when they do. Quality and cost tell you whether the outcome is actually better.

Where you practice Agent Experience

You can’t change the model, and you can only influence the harness, the agent app that runs the model, through its vendor. Everything the agent reads and calls is yours, and the AX stack maps out where that line runs.

  • Docs. When you update your docs, a model that learned them in training won’t know until the next model ships. An agent that reads them while it works sees your update on its very next run.
  • Agent extensions. Skills, MCP servers, instruction files, and custom agents you ship to steer the agent.
  • Your interfaces. APIs, SDKs, CLIs, error messages, and scaffolders. Agents use them the way developers do, and they trust what comes back. In one case, an unpinned scaffolder created a project on a version from July 2020 instead of the current one, and the agent took it as a success.

Best practices are hypotheses until you measure them

Lists of AX best practices, like the principles of AX and best practices for agent-friendly products, are a good place to start. But a change that sounds agent-friendly can still make agents worse with your technology, or more expensive. And the obvious fix isn’t always the best one.

Take SPFx upgrades again. When we ran that task five times in GitHub Copilot Chat with Claude Sonnet 4.6 on Windows, the agent passed only 30 of 80 configuration checks. Telling it to use CLI for Microsoft 365, an open-source tool that upgrades SPFx projects, raised that to 75 of 80. But an instruction like that only works for developers who either know of the CLI or who install the skill that carries it.

The agent’s step-by-step trail showed us something more useful. It fetched the SPFx release notes while it worked, so rather than shipping a skill, we fixed the release notes instead. After that, every run used the CLI without being told, for every developer, with no extra skill installed (Behind SPFx Dev Skills). What we took away from this evaluation is that you must watch how agents use your technology, change what you control, and measure whether it got better. These activities combined are the practice of Agent Experience.

Here are a few more things we measured.

Sounds agent-friendly What we measured
Give your CLI a JSON input mode for agents With regular arguments, every agent profile got all 5 runs right. In JSON mode, Claude Haiku 4.5 got 2 of 5 deployments right, and every model cost 4x to 11x more per task (Don’t rewrite your CLI for agents)
Add a tip in the docs that points to the right tool With a more direct tip, 1 of 5 runs used the tool. A warning that named the failing approach got 5 of 5 (Your agent already has a plan)
Give the agent another documentation source On SPFx upgrades, the context7 MCP server added no meaningful lift. Its tools didn’t load in 3 of 5 runs and weren’t called in the other 2 (Behind SPFx Dev Skills)
Switch to the model with cheaper tokens Claude Sonnet 5 cost 3.7x more per run than Sonnet 4.6 on SPFx code upgrades, despite 33% lower per-token pricing (Not all model upgrades are upgrades)

Each result holds for the scenarios and agent profile we measured. Yours may differ, and that’s exactly why you measure.

Key Agent Experience terms

AX has its own vocabulary for talking about how agents interact with your technology. Here are some key terms that we use regularly.

  • Agent extensions: everything you put in front of the agent to shape its behavior, like skills, MCP servers, instruction files, and custom agents.
  • Bare baseline: the same tasks run with no extensions, just the harness and the model. It’s what you compare every change against.
  • Lift and drag: against the bare baseline, a change either improves outcomes (lift) or makes them the same or worse (drag).
  • Task cost: what it costs to complete one task with a given setup, counting every token at its price, not just the token count. You report it next to quality, every time.
  • Agent profile: the setup you ran with, meaning the operating system, harness, model and its settings, and the extensions loaded. A result belongs to the profile it was measured with. (It’s not the same as the agent profile files some tools use for custom agents.)
  • Propensity and efficacy: whether the agent chooses your technology, and whether it uses it correctly once it has.

How Agent Experience relates to DX, GEO, and llms.txt

Developer experience (DX). AX builds on DX, including one of its oldest lessons. Usability conventions help, but you still watch real users to know whether your product works. Agents are users now too, and they’re easier to watch. You can run the same task five times with and without a change, which you’d never ask of a room full of developers.

Generative engine optimization (GEO). GEO is about getting AI search to mention your product. It overlaps with propensity, but AX goes further and asks whether the agent uses your technology correctly once it has picked it.

llms.txt, MCP, and other formats. These are surfaces and formats you can offer agents, which makes them part of AX. Like any best practice, they’re hypotheses until you’ve measured them with your technology and your developers’ tasks.

How to start practicing Agent Experience

Here’s a pragmatic approach to start evaluating Agent Experience for your technology.

  1. Pick your technology and write a few tasks your developers actually give agents, in their words. Include some that don’t name your product, to test propensity, and some that do, to test efficacy.
  2. Run each task a few times on a bare baseline, using the same agent profile. Look at where the agent goes wrong, and what it read along the way.
  3. Fix what caused it, starting with the docs or interfaces agents already use. Reach for a new extension last, since it only helps developers who install it. Then run the same tasks again with your change.
  4. Compare quality and task cost, and decide whether your change is lift, drag, or lift that’s too expensive to keep.
  5. Keep your conclusion scoped to that agent profile, and run again when it changes.

Already shipped a skill or an MCP server? Start at step 2, and compare runs with and without it.

How to measure AI agent extension effectiveness walks through the method, including the mistakes that produce false signals. Any evaluation beats none, because it gets you out of the “trust me, bro” zone, where the only evidence is a demo that looked fine.

Further reading

Agent Experience is an evolving field, and there are many resources to help you understand and improve it. Here are a few resources we recommend reading.

Category

Author

Waldek Mastykarz
Principal Developer Advocate

Waldek is a Principal Developer Advocate at Microsoft focusing on AI Coding Agents. He researches AI Coding Agents, and evaluates and improves Agent Experience for Microsoft's products and services.

0 comments