Copilot Studio · Power Platform

🚀 Adventures with Copilot Studio: Choosing the Right Model and Harness

I wanted to put together a cheat sheet for the models available in Copilot Studio. Not just a list of names, but something that helps answer a more useful question: which one would I use for the work I need the agent to do?

There are really two decisions here. Which harness fits the job? And which model fits the reasoning?

Picking the newest model does not answer both questions.

🧭 First, which harness do I need?

The model gets most of the attention. But the harness matters just as much. It is the runtime around the model, and it affects how the agent works with tools, files, and the process you want it to complete.

📊 Harness selection at a glance
🧭 Harness💡 Choose it when…📖 Illustrative business scenario🤖 Model-selection reference
Standard harnessYou want authored topics, controlled branching, and repeatable processes.A plant employee requests a replacement device; the agent collects required fields and follows an approved routing path.Standard harness models and prompt-builder models.
GitHub Copilot harnessWork requires adaptive planning across tools, systems, and documents.An accounts-payable agent investigates an invoice mismatch, gathers supporting records, prepares an exception package, and requests approval.GitHub Copilot harness models.
Copilot chat harnessThe main goal is employee access to enterprise knowledge inside Microsoft Copilot Chat.A new hire asks where to find onboarding requirements and relevant SharePoint guidance.The harness overview says it uses current chat models; it does not provide a selectable, versioned model inventory. Do not assume the other harnesses' pickers apply.

Scenarios and selection recommendations are illustrative guidance, not deployed customer examples.

One thing worth calling out: The GitHub Copilot harness in Copilot Studio is not the GitHub Copilot service. Similar name, different context.

⚙️ Which models can I use with the standard harness?

For a predictable service conversation, I would start by evaluating the default model against the requests the agent actually needs to handle. Then compare alternatives using the same test cases.

Selection surface: Agent Overview → Model. These choices concern orchestration; they are not automatically the same choices as prompt builder.

📊 Standard harness — United States model availability
🤖 Model🏷️ Category🌎 US availability💡 Suggested business use / action
GPT-5.5 ChatGeneralDefaultStarting candidate for grounded employee help, customer-service conversations, and routine tool use.
GPT-5 ChatGeneralGACompare against the default using an existing agent's regression test set.
GPT-4.1GeneralGARetain or evaluate where existing prompts and workflows have already been validated.
Claude Sonnet 4.6GeneralGAAlternative candidate for grounded service conversations and drafting; validate against the same cases.
Claude Opus 4.6DeepGACandidate for complex policy interpretation or multi-document exception analysis.
Claude Opus 4.7DeepGAAnother deep-model candidate for difficult cases; measure rather than assume superiority.
GPT-5 ReasoningDeepPreviewNonproduction evaluation of nuanced troubleshooting or policy analysis.
GPT-5 AutoAutoPreviewNonproduction evaluation of a help desk with both simple and complex requests.
GPT-5.3 ChatGeneralExperimental; early-access environmentControlled evaluation only.
GPT-5.4 ReasoningDeepExperimental; early-access environmentControlled evaluation of complex reasoning only.
GPT-5.5 ReasoningDeepExperimental; early-access environmentControlled evaluation of complex reasoning only.
Grok 4.1 Fast (Non-reasoning)GeneralExperimental; early-access environmentEvaluation only; review Microsoft's additional safety warning before testing.
Mistral Medium 3.5GeneralExperimental; cross-geoEvaluation only; review data movement and provider requirements.
GPT-4oGeneralRetired in commercial regionsNot a new commercial deployment choice. See government note below.
Claude Sonnet 4.5GeneralRetiredPlan migration, not a new deployment.

Suggested uses are category-based starting points, not model-specific benchmark claims.

Government exception: The same reference lists GPT-4o as the default for GCC, GCC High, and DoD. Do not carry the commercial-cloud inventory into government environments.

🤖 What about the GitHub Copilot harness?

If the job involves investigating a problem, coordinating several tools, or producing files, this is the harness I would evaluate. That still does not mean every request needs a Deep model.

📊 GitHub Copilot harness — United States model availability
🤖 Model🏷️ Category🌎 US availability💡 Suggested business use / action
GPT-5 ChatGeneralGACandidate for straightforward document drafting and routine tool-supported work.
GPT-5.5 ChatGeneralGAGeneral-workload candidate; validate quality and consumption on representative tasks.
GPT-6 AstraDeepGACandidate for cross-system investigation, exception handling, and evidence synthesis.
Claude Sonnet 4.6GeneralGAAlternative candidate for drafting and routine knowledge-plus-action tasks.
Claude Sonnet 5GeneralGA; early-access environmentGeneral-workload candidate where the early-access environment exposes it.
Claude Fable 5DeepGA; early-access environmentComplex-workflow candidate where available.
Claude Fable 5.1DeepGA; early-access environmentComplex-workflow candidate where available; compare on your own evaluation set.
Claude Opus 4.8DeepGACandidate for document-heavy analysis and multi-step investigations.
Claude Opus 5DeepGAAnother deep-workload candidate; select on measured business outcomes.
GPT-5.6 ReasoningDeepExperimental; early-access environmentControlled evaluation only.
Mistral Medium 3.5GeneralExperimental; cross-geoControlled evaluation only.

The availability table does not label a default model. Early-access availability is an environment condition, not synonymous with experimental status.

A quick regional gotcha: The reference shows no availability for Sonnet 5, Fable 5, or Fable 5.1 in Australia or Saudi Arabia. Many other entries require cross-geo processing outside the US. Check the region matrix before recommending a model.

✍️ Wait, does prompt builder use the same model?

This is where it is easy to mix things up. The primary agent model and the model selected for a prompt are separate choices. A prompt might just extract an asset ID or classify a ticket. It does not necessarily need the same model doing the orchestration.

📊 Prompt builder — United States release status and published rate tiers
🤖 Prompt model💳 Published rate tier🌎 US release status💡 Suggested task
GPT-4.1 miniBasic; defaultGAExtract purchase-order references, classify tickets, or summarize a short handoff.
GPT-4.1StandardGAProduce structured summaries from more demanding source material.
GPT-5 chatStandardGADraft a customer response from approved facts.
GPT-5.3 chatStandardExperimental — evaluation onlyEvaluation only: compare general-purpose prompt outputs. Not a production recommendation.
GPT-5 reasoningPremiumGAAnalyze an invoice exception with multiple constraints.
GPT-5.2 reasoningPremiumExperimental — evaluation onlyEvaluation only: test complex comparisons and reasoning. Not a production recommendation.
Claude Sonnet 4.6StandardExperimental — evaluation onlyEvaluation only: test general-purpose prompt tasks. Verify this surface’s status before production use.
Claude Opus 4.6PremiumExperimental — evaluation onlyEvaluation only: test reasoning prompt tasks. Verify this surface’s status before production use.
Grok 4.1 Fast (Non-reasoning)StandardExperimental — evaluation onlyEvaluation only; additional safety review required. Not a production recommendation.

A billing tier does not establish production readiness. Rate tiers are not dollar quotes or guarantees of cost per business transaction.

⚠️ Here is a gotcha: Check release status for the specific selection surface—not just the model name. Claude Sonnet 4.6 and Claude Opus 4.6 are experimental in prompt builder, even though agent model tables list them as GA. Treat these prompt choices as evaluation only. The tables above should not be read as a shared release-status list across surfaces.

The primary-agent documentation also identifies separate deep-reasoning and generative-response settings. Their inventories are not established by the tables above.

🏭 What does this look like in a real business scenario?

A list of model names is helpful. But what would I actually do with them? Here are a few places I would start. These are examples to evaluate, not a claim that one model wins every time.

📊 Business scenarios and suggested starting points
💼 Business request🧭 Harness🤖 Starting model approach🛡️ Why / controls
“Where is the approved supplier onboarding policy?”Copilot chatCurrent chat models supplied by that experienceKnowledge access, not a long-running transaction. Keep source permissions and citations intact.
“Submit a standard IT equipment request.”StandardGPT-5.5 Chat plus authored stepsA known process benefits from required fields and controlled routing.
“Explain my order status and open a case if it is overdue.”StandardGPT-5.5 Chat; evaluate other GA General modelsGround answers in actual order records; define the condition for creating a case.
“Classify these maintenance tickets and extract asset IDs.”Standard process with a promptGPT-4.1 mini as the initial prompt candidateA bounded transformation can be tested against an expected output schema. Use rules for identifier validation.
“Investigate this invoice/PO mismatch and prepare an approval package.”GitHub CopilotEvaluate GPT-6 Astra, Claude Opus 4.8, or Claude Opus 5Multi-system evidence and documents; require human approval before payment or ledger changes.
“Compare supplier agreements and flag conflicting obligations.”GitHub CopilotEvaluate GA Deep candidatesDocument-heavy synthesis; require clause citations and legal review rather than autonomous legal decisions.
“Build a weekly manufacturing operations briefing from ERP records and approved reports.”GitHub CopilotGeneral for straightforward assembly; evaluate Deep for exception analysisFile creation and synthesis; validate every metric against its source. Schedule only after testing.
“Diagnose a recurring quality issue across inspection, maintenance, and supplier records.”GitHub CopilotEvaluate GA Deep candidatesSeveral evidence sources and hypotheses; label uncertainty and keep safety-critical actions with authorized personnel.
“Test adaptive handling of mixed-complexity help-desk questions.”Standard, nonproductionGPT-5 Auto, PreviewAn experiment—not the production default. Keep a validated GA alternative.

💳 What does this cost?

There is one more decision to make before picking a model: how much will it cost to get the job done?

The model matters. So does the harness. And a billing tier is not the same thing as the total cost of an agent interaction.

✍️ Prompt builder: Basic, Standard, or Premium?

The published token-based rates for text and generative AI tools are:

📊 Published prompt-tool token rates — checked October 8, 2026
🏷️ Tier💳 Copilot Credits per 1,000 tokens🤖 Example models
Basic0.1GPT-4.1 mini
Standard1.5GPT-4.1, GPT-5 chat
Premium10GPT-5 reasoning

That gives me a starting point for comparing a small extraction task with a more demanding reasoning task. It does not give me a dollar price for the whole business process.

Experimental does not mean free. Release status and billing tier are separate checks.

⚙️ Standard harness: features plus reasoning

Standard-harness billing has feature charges. The published rates include 2 Copilot Credits for a generative answer and 5 for an agent action. One interaction can use several billed features.

If the agent uses a reasoning model, premium token usage is charged in addition to the applicable feature charge. Do not treat a reasoning answer as just a 2-credit answer.

These are standard-harness rates, not a universal price list for every harness. Bring-your-own-model configurations, including Azure Foundry models, are billed separately.

🤖 GitHub Copilot harness: budget for building, too

With this harness, consumption covers model tokens, tools, and the runtime. Building, previewing, testing, evaluating, and running agents can consume Copilot Credits.

The important bit: this usage is not included in a Microsoft 365 Copilot user license. A prototype can consume credits before it ever reaches production.

I would not assign a fixed per-model price from the references used here. Instead, run representative tasks and review the actual consumption. Longer investigations, repeated tool calls, and additional model calls all belong in that evaluation.

🎫 Does a Microsoft 365 Copilot license cover it?

Eligible standard-harness employee scenarios can be included for authenticated Microsoft 365 Copilot–licensed users in Microsoft 365 channels, subject to the published conditions and fair-use limits. Agent-flow inclusion applies to the “When an agent calls the flow” trigger; other triggers are billed. Computer use is excluded.

The Copilot chat harness can be consumption-billed or covered by an applicable user subscription. Do not extend either inclusion to GitHub Copilot harness usage.

💡 What would I measure?

  • Cost per completed task: not just the cost of a single answer.
  • Quality and retries: does the cheaper candidate actually finish the job reliably?
  • Development consumption: include your evaluation and testing budget.
  • Controls: assign an owner and set agent limits and notifications. Azure budget alerts alone do not stop consumption.

My take: use the least expensive option that reliably completes the business task. A deeper model can be worth it—but let the results, not the name, make that decision.

Rates are a documentation snapshot, not a billing quote. Dollar cost depends on your credit purchase plan and applicable agreement. Check the current licensing guidance before budgeting.

💡 So, where would I start?

  • A question about company knowledge: start with the knowledge experience, not an autonomous workflow.
  • A fixed, repeatable business path: start with the standard harness and explicit process controls.
  • A goal requiring investigation, tool coordination, or a document deliverable: evaluate the GitHub Copilot harness.
  • A small classification or extraction step: start with a mini prompt candidate and validate its output.
  • Several interacting constraints: evaluate GA Deep models against representative cases.
  • Mixed workloads: test whether one GA General model is sufficient before adding complexity.
  • Production: select a GA model that meets region, policy, quality, latency, and consumption requirements. Do not select Preview or Experimental models for production.

These are implementation recommendations, not model-specific performance guarantees.

✅ What should I check before putting this into production?

  • Confirm harness, selection surface, cloud, environment region, and release cycle.
  • Check the actual model picker; documentation availability does not establish tenant enablement.
  • For external models, check both Power Platform environment/environment-group controls and provider access in Microsoft 365 admin center.
  • Review cross-region data movement and preview/experimental controls separately.
  • Test representative successful cases, missing data, ambiguous requests, denied permissions, and failed tools.
  • Measure task completion, factual grounding, response time, and credits consumed; do not choose solely by model name.
  • Keep human approval for consequential financial, legal, production, or safety decisions.
  • Re-test when models, defaults, prompts, tools, or policies change.

My take: Start with the business problem, not the newest model name. Use the simplest harness and GA model that can do the job reliably. If deeper reasoning improves the outcome in your tests, great. If it does not, keep it simple.

📚 References

1. Harnesses in Copilot Studio — “GitHub Copilot harness,” “Standard harness,” “Copilot chat harness,” and “Compare harnesses.” Updated October 1, 2026. https://learn.microsoft.com/en-us/microsoft-copilot-studio/harnesses-overview

2. Select a primary AI model for your agent — “Standard harness availability,” “US Government availability,” model categories, selection, and administrative controls. Updated September 18, 2026. https://learn.microsoft.com/en-us/microsoft-copilot-studio/authoring-select-agent-model

3. Model availability and controls — GitHub Copilot harness inventory, regional availability, release types, and administrative controls. Updated October 2, 2026. https://learn.microsoft.com/en-us/microsoft-copilot-studio/agents-experience/authoring-agent-model-availability

4. Change the model version and settings — Prompt builder inventory, rate tiers, and release-stage caveats. Updated August 4, 2026. https://learn.microsoft.com/en-us/microsoft-copilot-studio/prompt-model-settings

5. Prompt model availability by region and updates — “Public availability,” United States release status. Updated August 3, 2026. https://learn.microsoft.com/en-us/microsoft-copilot-studio/prompt-model-availability

6. Billing rates and management — Standard-harness feature charges, prompt-tool token rates, reasoning charges, and license inclusion conditions. https://learn.microsoft.com/en-us/microsoft-copilot-studio/requirements-messages-management

7. Overview of usage-based billing — GitHub Copilot harness billing scope and development consumption. https://learn.microsoft.com/en-us/microsoft-copilot-studio/agents-experience/billing-credit-overview

8. Manage costs for agents powered by the GitHub Copilot harness — License coverage, agent limits, monitoring, and cost controls. https://learn.microsoft.com/en-us/power-platform/admin/manage-usage-github-copilot-harness

Availability changes frequently. Recheck these references and your environment before sharing this as deployment guidance.