Tag: AI Governance

  • Copilot Cowork First Month: How Admins Should Check July Consumption

    Copilot Cowork First Month: How Admins Should Check July Consumption

    The first month of Copilot Cowork is a baseline, not a verdict.

    For Microsoft’s qualifying Frontier grace-period cohort—and for the tenant in this guide—July was the first full billed calendar month of Copilot Cowork consumption.

    But opening one dashboard, copying the big number, and calling it “usage” will give you an incomplete story.

    There are really two questions:

    1. What did Cowork consume?
    2. Are people adopting it in a healthy, valuable way?

    Microsoft answers those questions in two different areas of the Microsoft 365 admin center. The Cost Management report shows Copilot Credit consumption. The Cowork Usage report shows adoption, tasks, activity, and retention.

    They use different definitions, date controls, and refresh cycles.

    This guide walks through both, shows what I found in a real first-month tenant, and gives admins a practical checklist for configuring and reviewing month one.

    Important July context: Copilot Cowork became generally available on June 16, 2026. Tenants with at least one user who used Cowork during the Frontier period from March 30 through June 16 received a billing grace period through June 30, with billing beginning July 1. That makes July the first full billed month for that cohort—not automatically for every tenant. Confirm your own activation and billing dates before labeling July “month one.”

    In this guide

    The two dashboards answer different questions

    Admin viewBest forDate behaviorRefresh behavior
    Copilot > Cost ManagementCredits, limits, policies, users, groups, agents/services, prepaid versus PayGoSupports an exact month such as July 2026Overview refreshes every 4 hours; Consumption refreshes every 2 hours
    Copilot > Cowork > UsageActive users, tasks, scheduled versus user-initiated work, retention, adoption trendsRolling 7, 28, 90, or 180 days; default is 28 daysCheck the report’s visible Last updated timestamp

    That distinction matters. The Cowork Usage report I reviewed on August 5 offered a 28-day range of July 5 through August 2. It did not offer a calendar-month July option. I could use it to understand adoption near the end of July, but I could not honestly present it as an exact July report.

    Cost Management Consumption page with summary cards for users, groups, and agents and services
    Cost Management can report the exact July 2026 consumption period.

    How to check July Copilot Cowork consumption

    1. Start in Cost Management, not the generic Copilot credits report

    In the Microsoft 365 admin center, go to:

    Copilot > Cost Management > Consumption

    For Cowork, this is the useful consumption surface. It lets you move between Users, Groups, and Agents and services.

    Do not confuse it with Reports > Usage > Microsoft 365 Copilot > Credits. That report has a different scope and is not the report I would use to isolate Cowork consumption.

    2. Select July 2026

    Open Time period and select July 2026.

    July month selector
    Use the calendar-month selector in Cost Management when you need an exact July baseline.

    Here is the small trap I saw in the August 5 tenant UI: changing between Users, Groups, and Agents and services reset the time period to Month to date. Re-select July on every view and confirm the label before recording a number or exporting data.

    3. Use Agents and services to isolate Cowork

    Start with Agents and services. This is the cleanest way to separate Cowork from Work IQ API or another consumptive service covered by the same policy.

    Cowork service consumption for July
    The Agents and services view is the best starting point for the Cowork total.

    Record at least:

    • Cowork Credits used
    • Active users
    • Sessions
    • The report period
    • The time and date you captured the report

    That last item is not busywork. Microsoft says the Consumption tab refreshes every two hours, while exported reports are point-in-time snapshots. A dashboard and an export taken at different moments can temporarily disagree.

    4. Review users for concentration and limit pressure

    Move to Users, re-select July, then inspect:

    • Credits used
    • Monthly credit limit
    • Percentage of the limit consumed
    • Sessions
    • Last activity date
    • Assigned spending policy
    • Daily credit usage
    User consumption details with spending policy and monthly limit
    The user drawer connects total usage to the assigned policy, service access, and individual limit.
    Daily user credit usage
    Daily usage shows when concentration occurred. Pair it with Cowork task data or Purview audit evidence before deciding why it occurred.

    Do not just chase the largest user. Ask whether the work was intentional, repeatable, and valuable. A heavy user may be delivering the best ROI in the tenant. The useful signal is high consumption without an understood scenario, owner, or outcome.

    5. Review groups, then re-check the period

    The Groups view helps you compare the audiences targeted by spending policies. It is useful for spotting:

    • One group consuming most of the credits
    • A group with low consumption after you compare the Consumption view with its enabled scope under Configuration
    • A pilot team whose limits are too low for its intended workload
    • Overlapping memberships that make policy behavior harder to explain
    July consumption by group
    Group totals help admins compare the audiences governed by different spending policies.

    User and group totals do not always match perfectly. Users can belong to multiple groups and policies, and policy changes during a billing period can affect what appears at user level. Treat these views as different analytical cuts, then reconcile material differences against exports and policy history.

    6. Check the July spending trend separately

    Return to Cost Management > Overview and set the Spending trend card to July.

    July cumulative spending trend
    The Spending trend card has its own period control. Do not assume the page-level cards and chart are showing the same month.

    Look for:

    • Large daily spikes
    • A steady upward ramp after enablement
    • Weekend spikes to compare with scheduled-task history
    • A shift from prepaid capacity into PayGo
    • A sudden drop that coincides with a limit or access change

    In the tenant reviewed for this guide, July was spiky rather than smooth. The largest single day used more than 12,000 credits, while several days were below 1,000. That is a prompt to investigate which scenarios or scheduled tasks ran on the high days—not proof that the spikes were waste.

    What the first full month looked like in this tenant

    The following numbers are an anonymized, point-in-time snapshot captured on August 5, 2026.

    Exact July consumption

    • 136,956 Cowork Credits in the Agents and services view
    • 19 active Cowork users in that service view
    • 260 sessions across the user rows
    • 18 users with non-zero recorded credits at capture time
    • 6,150 median credits across the 19 user rows
    • 7,208 average credits across the 19 user rows
    • 526.7 credits per session across the user rows
    • The largest user represented 22.0% of user-level credits
    • The top five users represented 59.8% of user-level credits
    • 6 users had consumed at least half of their individual limit
    • 0 users had consumed 80% or more of their individual limit

    The service, group, user, and trend views differed by a few hundred credits at capture time: roughly two-tenths of one percent. That is small enough to be consistent with refresh timing and aggregation differences, but it should still be documented. Use one defined source for the headline number—here, the Cowork row under Agents and services—and preserve exports when you close the month.

    Do not turn credits directly into an invoice unless you know the billing path

    Microsoft lists PayGo at $0.01 USD per Copilot Credit. At pure retail PayGo, 136,956 credits would be a $1,369.56 USD usage equivalent.

    That is not automatically the tenant’s invoice. This tenant’s policies used prepaid capacity. When capacity packs and a subscription with P3/pay-as-you-go are both selected, Microsoft documents the order as capacity packs first, then P3 prepaid credits, then PayGo. Discounts, prepaid commitments, shared capacity, and billing arrangements can all change the actual cost. Report credits and billing source together.

    Now check adoption—but label the date range honestly

    Go to:

    Copilot > Cowork > Usage

    Cowork Usage rolling date selector
    Cowork Usage uses rolling ranges. The 28-day view here covers July 5 through August 2, not calendar July.

    The report’s four headline metrics are:

    • Active Cowork users
    • Total Cowork tasks
    • Average tasks per active user
    • Retained Cowork users
    Cowork adoption cards
    Always capture the date range and Last updated timestamp with the headline metrics.

    Microsoft defines a task as one Cowork conversation thread, including user-initiated conversations and scheduled task executions. Follow-up questions in the same thread remain part of the same task.

    The same report defines a retained user as someone active in both the previous seven-day period and the most recent seven-day period. It is a useful continuity signal, but it is not a conventional month-over-month cohort retention rate.

    Microsoft's task definition in the report
    A task is not the same thing as a prompt, run, or Cost Management session.

    The report was viewed on August 5 and displayed Last updated: August 3. For the rolling 28 days from July 5 through August 2, this tenant showed:

    • 18 active users
    • 332 total tasks
    • 18.44 average tasks per active user
    • 18 retained users
    • 103 user-initiated tasks
    • 229 scheduled tasks

    That means about 69% of tasks were scheduled. It is a valuable operational signal: Cowork was being used as automation, not only as a conversational tool.

    Daily active users and task mix
    Read daily active users beside the task-type mix to understand how engagement is occurring.
    Scheduled and user-initiated task trend
    The late-July scheduled-task increase deserves a scenario-level review before changing limits.

    But the average hid a highly concentrated pattern: one user accounted for 206 of the 332 tasks, or about 62%. The median active user had only 6 tasks. Fifteen users were active on two days or fewer.

    This does not mean the rollout failed. It means the first-month story is a scheduled-heavy tenant and a highly concentrated user-level pattern, plus a low-frequency long tail. Month two should validate the dominant user’s scenarios while improving repeat, intentional use across the rest of the pilot.

    The first-month admin playbook

    Days 1–3: Prove the plumbing

    • Confirm each intended user has a Microsoft 365 Copilot license.
    • Confirm usage-based billing is enabled and tied to the correct subscription or prepaid capacity.
    • Confirm Cowork is discoverable to the intended audience.
    • Confirm the intended security groups are assigned to an active spending policy.
    • Run a controlled task and verify it appears after each report updates. Use the visible Last updated timestamp for Cowork Usage and allow for Cost Management’s documented refresh window.
    • Capture the initial policy, billing method, and access configuration before changing anything.

    Discoverability and access are separate controls. The discovery setting controls tenant-wide visibility, while Cost Management controls scoped access and spend. Microsoft says usage-based billing setup overrides the discovery setting for users covered by that setup, so validate both the tenant-wide experience and the intended scoped audience.

    Week 1: Establish safe limits and alerts

    Start with group-based policies that match the rollout audiences. A practical pilot configuration should include:

    • A policy-level monthly cap
    • A per-user monthly cap
    • Alerts well before the hard cap
    • Named operational and financial owners for those alerts
    • Explicit agents and services included in the policy
    • A documented billing method

    Also make an explicit decision about Allow new services and agents as they become available. Microsoft’s current policy setup selects that option by default. That can be convenient, but a controlled pilot should not inherit future service access without an owner and review process.

    Spending policies impose access and consumption limits; they do not reserve or allocate a slice of the underlying Copilot Credit capacity to a group. Treat policy limits as guardrails, not as capacity reservations.

    Policy precedence warning and configuration
    When multiple policies apply, Microsoft enforces the most permissive applicable policy; settings are not combined.

    Microsoft’s current precedence order for overlapping policies is:

    1. Highest per-user limit
    2. If tied, highest overall policy limit
    3. If still tied, newest policy

    That is easy to get wrong. Adding a “special” policy can accidentally make someone less restricted than intended.

    Monthly hard-cap behavior
    A hard cap blocks covered services for the rest of the month and resets on the first day of the next month.
    Policy alert threshold
    Set an early threshold so an alert creates time to investigate instead of merely announcing that the budget is gone.
    Per-user monthly cap
    Per-user caps protect a shared policy from one runaway or misunderstood workload.

    When the run rate is unknown, I start the first alert at 50%. Once you understand the run rate, use multiple operational checkpoints outside the platform—such as weekly reviews at 50%, 70%, and 85%—even if the product sends a single configured alert threshold.

    Use least privilege for the operating team. AI Administrators and License Administrators can create spending policies, manage limits and alerts, and view Cost Management. Selecting or overriding the billing method requires a Global Administrator or Billing Administrator. After a spending policy is created, its billing method cannot be changed; Microsoft requires deleting and recreating that policy.

    Week 2: Review adoption, not just activation

    Ask five questions:

    1. How many targeted users became active through a direct prompt or a scheduled execution?
    2. How many came back in the next seven-day period?
    3. Is work primarily user-initiated or scheduled?
    4. Is usage concentrated in one person, team, or scenario?
    5. Which enabled users have no visible activity and need a use case, training, or removal from the pilot?

    The Usage table lists people with Cowork activity dating back to April 1, 2026, while its summary cards respond to the selected rolling period. Do not use the total table row count as your active-user count. Also remember that an active user can be counted because they prompted Cowork or because a scheduled task executed on their behalf. Active-user coverage is not identical to direct human adoption.

    Week 3: Connect consumption to scenarios

    For the top users and largest days, document:

    • Business scenario
    • Owner
    • Frequency
    • User-initiated or scheduled
    • Inputs and data sources
    • Outputs and downstream actions
    • Credits consumed
    • Time saved or risk reduced
    • Whether a lower-cost model or tighter scope can produce an acceptable result

    Consumption alone is not value. A 5,000-credit task that replaces two days of expert work may be excellent. A 500-credit task running every hour with nobody reading the output may not be.

    Week 4: Freeze the baseline and decide month two

    • Re-select the exact July period in every Cost Management view.
    • Export users, groups, and agents/services after the reporting refresh.
    • Record the Cowork Usage rolling range and Last updated timestamp.
    • Preserve policy settings and group membership as of month-end.
    • Explain material differences between views.
    • Review users near limits and compare the configured audience with users showing no repeat activity.
    • Classify scheduled tasks as keep, tune, pause, or investigate.
    • Set month-two targets before raising budgets.

    The metrics I would put on an admin scorecard

    MetricFormulaWhat it tells you
    Active coverageActive users ÷ enabled target usersHow much of the intended audience generated direct or scheduled activity
    Retained-to-active ratioRetained users ÷ active usersA simple continuity signal, not a conventional cohort-retention rate; retained users compare the two most recent seven-day periods while active users cover the selected range
    Scheduled-task shareScheduled tasks ÷ total tasksHow much usage is recurring automation versus direct interaction
    Credits per active userCowork credits ÷ active usersOverall consumption intensity when the periods align
    Credits per sessionCowork credits ÷ Cost Management sessionsRelative session intensity within the consumption report
    Limit utilizationCredits used ÷ applicable policy limitHow close a user or policy is to its cap; the limit does not reserve underlying capacity
    Top-five concentrationCredits from top five users ÷ total user creditsWhether the rollout depends on a few users
    Average-to-median gapAverage usage compared with median usageWhether a small number of heavy users are distorting the average

    Only combine figures whose date ranges and definitions match. A Cowork task is not a Cost Management session. In this review, the adoption report covered July 5 through August 2 while the consumption report covered July 1 through July 31, so I did not calculate credits per task.

    Month-one red flags worth investigating

    • A view silently returned to Month to date after you changed tabs.
    • The Overview headline cards and Spending trend chart are showing different periods.
    • A user or group is above 80% of its cap with time left in the month.
    • Scheduled work dominates consumption but has no named owner or reviewed output.
    • Average usage is high while median usage is low.
    • A service other than Cowork is consuming credits under the same policy.
    • A group has access but almost no activity.
    • A spending policy includes new services automatically without a review process.
    • Several groups overlap, producing a more permissive policy than expected.
    • Finance is treating credits multiplied by retail PayGo as the invoice even though prepaid capacity or P3 applies.
    • Security assumes all Microsoft Purview controls apply identically to Cowork. As of this writing, Microsoft documents DLP for general Cowork AI interactions as unsupported, even though auditing, sensitivity labels, eDiscovery, data lifecycle, communication compliance, and other controls have documented support. Cowork’s local Edge browser tasks are narrower: Microsoft says they inherit existing browser DLP, Conditional Access, and site policies. Review both control paths rather than inheriting assumptions from another Copilot surface.

    What I would change for month two

    For this tenant, I would not start by cutting the heavy user or scheduled workload. I would first confirm whether that automation is producing a real business outcome. If it is, preserve it and establish an owner, expected run rate, and value measure.

    Then I would:

    1. Create a repeat-use goal for the low-activity users.
    2. Review the five largest credit consumers and the largest daily spikes.
    3. Audit scheduled tasks for frequency, duplicate runs, and unused output.
    4. Separate Cowork and Work IQ API policy reporting wherever the rollout design allows it.
    5. Keep alerts early enough to react before a hard cap blocks work.
    6. Raise limits only for a documented scenario with evidence of value.
    7. Review plugins, model availability, browser use, audit coverage, and Purview support as part of the same operating rhythm—not as a one-time launch checklist.

    Final thought

    The goal of month one is not to prove that every user became a power user. It is to leave with a trustworthy baseline: who used Cowork, what they ran, what it consumed, which work repeated, and which controls need to change next.

    That is the admin story July can tell—if you use both dashboards and keep their definitions straight.

    Sources

  • GPT-5.6 Is Missing in Copilot Cowork? Turn On This Admin Setting

    GPT-5.6 Is Missing in Copilot Cowork? Turn On This Admin Setting

    GPT-5.6 is available in Copilot Cowork.

    But if you open the model picker and only see GPT 5.5, the problem may not be your Cowork license or the rollout.

    There is one Microsoft 365 admin setting you need to check first.

    OpenAI must be enabled as a Microsoft subprocessor before users can access the new OpenAI-operated models.

    Copilot Cowork model picker before GPT-5.6 models are enabled
    Before OpenAI is enabled as a Microsoft subprocessor, the Copilot Cowork model picker may show GPT 5.5 but not GPT 5.6 Sol or Terra.

    The hidden dependency behind GPT-5.6 in Cowork

    Microsoft now supports OpenAI-operated models in Microsoft 365 Copilot experiences. That is different from OpenAI models operated by Microsoft through Azure OpenAI.

    For GPT-5.6, Microsoft’s documentation is direct: if you want to use GPT-5.6, you need to enable OpenAI-operated models.

    That means the model picker is not the first place to troubleshoot. Start in the Microsoft 365 admin center.

    Where to enable OpenAI as a subprocessor

    You need to be a Global Administrator or AI Administrator to change this setting.

    Go to the Microsoft 365 admin center and follow these steps:

    1. Open Copilot.
    2. Select Settings.
    3. Select View all.
    4. Open AI providers operating as Microsoft subprocessors.
    Microsoft 365 admin center Copilot settings showing AI providers operating as Microsoft subprocessors
    In the Microsoft 365 admin center, go to Copilot > Settings > View all, then open AI providers operating as Microsoft subprocessors.

    Choose who gets access

    Expand OpenAI. From there, choose which users can access OpenAI-operated models:

    • All users
    • No users
    • Specific users and groups

    Then select Save.

    OpenAI subprocessor access settings in the Microsoft 365 admin center
    OpenAI is disabled by default in the current admin experience. Admins can enable it for everyone or scope access to specific users and security groups.

    For a pilot, I would start with Specific users and groups.

    This setting is applied at the provider level across supported Microsoft 365 Copilot and Copilot Studio experiences. Scoping it to a security group gives you a clean way to test the new models with a small group before expanding access.

    What users see after the setting is enabled

    After the setting is saved, go back to Copilot Cowork and open the model picker again.

    In my tenant, the new choices appeared as:

    • GPT 5.6 Sol — the most capable GPT-5.6 model for complex work.
    • GPT 5.6 Terra — the balanced GPT-5.6 model for high-volume work.
    Copilot Cowork model picker showing GPT 5.6 Sol and GPT 5.6 Terra
    After enabling OpenAI as a Microsoft subprocessor, GPT 5.6 Sol and GPT 5.6 Terra appear in the Copilot Cowork model picker.

    That is the before-and-after test.

    If GPT-5.6 is missing, check the provider setting before you spend time troubleshooting the user, the browser, or Cowork itself.

    The governance detail admins should not skip

    Enabling this setting is not just a feature toggle. It is a data-processing decision.

    Microsoft says OpenAI operates these models as a Microsoft subprocessor under contractual safeguards and appropriate technical and organizational measures. Microsoft Product Terms, the Microsoft Data Protection Addendum, and Enterprise Data Protection apply, subject to Microsoft’s documented exclusions.

    There are also two details worth reviewing with your governance team:

    • OpenAI-operated models are currently excluded from in-country processing commitments where those commitments apply.
    • They are not currently available in GCC, GCC High, DoD, or sovereign clouds.

    Do not turn this on just because the new model names look interesting. Decide who needs access, document the data-processing change, and start with a controlled group.

    One rollout date to pay attention to

    Microsoft says OpenAI-operated models are currently disabled for customers. Starting July 24, 2026, they will be enabled for all users in eligible commercial tenants unless an admin explicitly selects No users.

    That makes this a setting every Microsoft 365 Copilot admin should review, even if the organization is not ready to use GPT-5.6 yet.

    The question is no longer just, “How do I turn GPT-5.6 on?”

    It is also, “Who should get it, and what should our policy be before the default changes?”

    The practical rollout pattern

    • Confirm the Global Administrator or AI Administrator who owns the setting.
    • Review the OpenAI subprocessor terms and data-processing notes.
    • Create a security group for the first Cowork users.
    • Enable OpenAI for that group.
    • Confirm GPT 5.6 Sol and Terra appear in the Cowork model picker.
    • Test real work with a small group before expanding access.
    • Review the setting again before the July 24 default change.

    The technical step takes a minute.

    The admin decision behind it deserves more attention.

    If GPT-5.6 is missing in Copilot Cowork, start with the OpenAI subprocessor setting.

    Sources

  • Build Your Agent Factory: 10 Moves That Ship Fast (and Scale)

    Build Your Agent Factory: 10 Moves That Ship Fast (and Scale)

    Build Your Agent Factory: 10 Moves That Ship Fast (and Scale)

    Agents at scale. Not POCs.

    Here’s the playbook I’d hand any exec or builder who wants working agents in production—without turning the org into a science fair.

    1) Stand up an AI Agents Workforce

    What it is: A small cross-functional crew with authority to hunt repetitive work and ship agents.

    Who’s in:

    • 1 product owner
    • 1 engineer (Copilot Studio/Power Automate)
    • 1 data person
    • 1 security/governance lead
    • 1 domain SME.

    Ship this week: Write a one-page charter with scope, decision rights, and a 30-day roadmap (first 5 agents + metrics).

    2) Win with horizontals first, then go vertical

    Horizontals (1-hour wins): drafting, summarizing, policy Q&A, meeting notes to actions, form-fill helpers.

    Verticals (outsized ROI): pick 1–2 per business unit where there’s money, risk, or SLA pain.

    Guardrail: don’t start with the hardest workflow; start where you can close the loop and measure value inside two weeks.

    3) Make an Agents Directory the front door

    Why: Ideas die in email. A directory turns “we should build X” into spec and governance.

    Minimum intake fields:

    • use case name
    • goal
    • users
    • decision rights
    • data sources + who owns it
    • tools
    • PII/sensitivity
    • KPIs
    • business owner
    • risk level
    • rollout plan.

    Outcome: Every request auto-generates a lightweight PRD (goal, inputs, outputs, metrics, guardrails) and a yes/no gate.

    4) Create the 1-Hour Agent template

    Template anatomy:

    Goal + success criteria Input schema (what the user provides) Tools (actions/connectors) and permissions Knowledge sources (files, sites, indexes) Safety rules (allowed/blocked actions, escalation) Evaluation set (10–20 test prompts with expected outcomes) Deploy script (Dev → Test → Prod)

    Rule: If a use case can’t fit this page, it’s not a 1-hour agent—park it for later.

    5) Tie every agent to a visible scorecard

    Metrics to publish: time saved, cost avoided, error rate, CO₂/efficiency (where relevant), user satisfaction.

    Simple formula: monthly users × average minutes saved × loaded cost = value.

    Make it public internally: green/red status, owner, last review, next improvement.

    6) Run on a secure, managed agent runtime

    Non-negotiables: identity passthrough, content safety, audit logs, tool call restrictions, data boundary controls, environment isolation.

    Practical tip: standardize a “sensitive sources” policy and block tools by default; allow case-by-case.

    7) Split the stack to move fast without breaking things

    Experience layer: Copilot Studio for UX, channels, and connectors.

    Agent runtime/orchestration: managed agent service for threads, tool calls, safety, and evaluations.

    Why it works: builders ship quickly at the edge; platform team keeps shared guardrails, monitoring, and upgrades stable.

    8) Mix knowledge + action (or you’ll stall)

    Knowledge: structured grounding (SharePoint/Fabric/Search), doc versioning, citations-on by default.

    Action: flows/Logic Apps, Graph, line-of-business APIs; always ship with a dry-run mode first.

    Design pattern: Answer → show sources → propose actions → execute on approval. When confidence is high and stakes are low, allow auto-execute.

    9) Keep humans in the loop—by design

    HITL patterns that work:

    Shadow mode (observe only) → suggest mode → execute with approval → auto-execute.

    Confidence thresholds where low confidence routes to a human. Escalation logic when guardrails trip or data is missing.

    UX rule: one click to approve, one click to undo.

    10) Plan to scale on day one

    Pipelines: Dev → Test → Prod with approvals and rollback.

    Evals: pre-ship test set per agent; weekly drift checks; quarterly red-team.

    Ops: central logging, cost dashboards, incident playbook.

    Program ritual: a quarterly “Agent Backlog Day” to harvest new ideas and retire underperformers.

    Starter Architecture (fast and boring)

    Experience: Copilot Studio (web, Teams, M365, chat, plugins)

    Actions: Power Automate/Logic Apps + custom APIs

    Knowledge: SharePoint/Fabric/AI Search with retrieval policies

    Runtime: managed agent service for tool orchestration, identity, safety

    Observability: evaluations, telemetry, and a simple agent scorecard per app

    Security: Entra ID RBAC, private endpoints, DLP, approval gates

    Prompts and policies that save you pain

    Prompt contract (keep it in the repo): role, goals, inputs, allowed tools, forbidden actions, decision rights, escalation, output format, citation rules.

    Data contract: what sources are permitted, freshness expectations, sensitivity tags.

    Failure modes: what the agent must do when unsure (ask for clarification, route to human, or stop).

    Anti-patterns I keep seeing

    • Starting with an “AI strategy deck” instead of shipping 3 agents.
    • Agents that answer but can’t act—users stop coming back.
    • No owner, no scorecard, no sunset date.
    • Canary-testing in production without a rollback plan.
    • Letting one giant use case block 20 small wins.

    Your first week mapped

    Day 1: Form the team and publish the charter.

    Day 2: Launch the Agents Directory (intake + PRD autogeneration).

    Day 3–4: Build two 1-hour agents (drafting + policy Q&A) with eval sets.

    Day 5: Ship to a pilot group with scorecards visible. Book the first backlog day.