Showing category results for AI

Sep 29, 2026
Post comments count0
Post likes count0

Build with Azure Canvases: a shared workspace for agents

Matt Hernandez

See how Azure Canvases bring plans, code, previews, and deployment feedback together in a collaborative workspace for agent-assisted development.

AIMicrosoft for DevelopersGitHub Copilot
Sep 21, 2026
Post comments count2
Post likes count1

Knowledge cutoff is a poor proxy for model capability

Waldek Mastykarz

A model can fail on features released before its knowledge cutoff, then succeed on ones released after it. We tested hundreds of product changes and found that the date tells you far less than the work does.

AI
Sep 16, 2026
Post comments count2
Post likes count1

Your AI coding agent evaluation is only as good as its sandbox

Waldek Mastykarz

Your AI coding agent passed the eval. But did the model know the answer, or did it find it somewhere on your machine? A correct answer can still invalidate your measurement.

AI
Sep 15, 2026
Post comments count1
Post likes count1

Build an interview coach app with the GitHub Copilot SDK

Justin Yoo

An interview coach has to do more than ask questions. It needs to read a resume, follow up on an incomplete answer, and save enough context to give useful feedback at the end. Some of that work is conversation. Some of it requires calling an application service. The GitHub Copilot SDK lets you use the runtime behind Copilot CLI for that work insid...

GitHub CopilotAI
Sep 9, 2026
Post comments count2
Post likes count1

Your work might not need the smartest model

Waldek Mastykarz

The smartest model can cost five times more and deliver the same result, or even a worse one. See how evaluating your own work helps you get more value from your agent budget.

AI
Jul 21, 2026
Post comments count0
Post likes count1

How to test agent experience changes without shipping them

Waldek,
Garry

Most changes you think will improve AI agent behavior won't. We tested a dozen hypotheses on a real project upgrade scenario and the majority failed. Learn how to emulate documentation, API, and MCP server changes locally so you can validate what works before shipping anything to production.

AI
Jul 17, 2026
Post comments count0
Post likes count1

How to test agent skills without hitting real APIs

Waldek Mastykarz

Your agent skill calls an API. The moment you start evaluating it, every run either costs money or mutates production data. Learn how to mock APIs transparently so you can run evals without changing your skill or hitting real endpoints.

AI
Jul 15, 2026
Post comments count0
Post likes count0

Building AX evals that actually work

Waldek Mastykarz

This is the eighth and final article in a series about Agent Experience (AX): the practice of making AI coding agents work correctly with your technology. The series covers what you can and can't control in the agent stack, how to measure whether your extensions are helping or hurting, and how to iterate toward better outcomes. You've read seven a...

AI