Tag: Copilot Credits

  • Copilot Cowork Model Benchmark: Five Models, One Prompt, One Winner

    Copilot Cowork Model Benchmark: Five Models, One Prompt, One Winner

    I put five Copilot Cowork models through the same executive decision stress test: one locked prompt, the same six-file evidence pack, and a fresh session every time.

    The job was not to summarize the packet. Each model had to reconcile conflicting numbers, stale executive claims, security blockers, contract gates, and future retest dates—then choose the correct launch path and turn it into a memo, PowerPoint, and Excel action tracker.

    All five reached the expected decision. The real separation came after that: evidence discipline, unsupported assumptions, file quality, and a ninefold spread in Cowork credits. Claude Opus 5 delivered the strongest overall result while using the fewest credits.

    Decision accuracy came first. Polish only counted after the answer was right.

    The result in 30 seconds

    Every model chose the expected Option B: approve a conditional seven-site phased launch. Opus was the only run to pair a perfect quality score with the top design score and the lowest credit total.

    ModelQualityDesignCredits
    Claude Opus 5100 / 1009.5 / 10294
    Claude Sonnet 598 / 1008.9 / 10324
    GPT-5.6 Sol99 / 1008.4 / 10771
    Claude Fable 599 / 1009.4 / 102,658
    GPT-5.6 Terra97 / 1007.5 / 10360
    All five selected Option B. Quality used the locked 100-point rubric. Design was scored separately. Credits are the task totals shown by Cowork’s /cost response; Elapsed time was not scored.

    In this benchmark

    The executive decision stress test

    I built one fictional but realistic executive decision scenario, created a hidden answer key, and ran the exact same job in five fresh Copilot Cowork sessions. The selected model was the variable.

    • Same input: one prompt and the same six files for every model.
    • Fresh sessions: no previous conversation or model output could influence the next run.
    • No rescue prompt: no web access, extra explanation, or follow-up designed to repair a weak answer.
    • Outcome first: decision accuracy was checked before design and presentation quality.

    The fictional company, Kavora Industries, was preparing a steering committee decision for a Hartland Energy software launch. The evidence pack deliberately mixed current data with stale summaries, future retest dates, security blockers, contractual gates, and conflicting executive claims.

    Every session received:

    • An executive context document
    • An Excel launch-readiness workbook
    • A security and contract brief
    • An executive email thread with conflicting claims
    • A fictional company brand guide
    • A board presentation template

    The model had to reconcile the evidence, calculate the real readiness position, choose one of three launch paths, and create three usable business files: a two-page decision memo, a steering committee deck of six slides or fewer, and an Excel action tracker.

    The answer key

    I locked the answer key before reviewing any model output.

    Expected decision: approve Option B—a conditional seven-site phased launch on August 17. Both Priority 1 defects had to pass retest with recorded evidence, the final seven-site list had to be confirmed, and the Hartland sponsor had to provide written approval naming the phase and facilities. If any gate failed, launch activity moved to August 31.

    Two fictional P1 defects controlled the decision:

    • SEC-014: a contractor supervisor could access another company’s attachment through a predictable file URL.
    • AUTH-022: a legacy SSO fallback could bypass the enforced MFA path during a retry condition.
    Open Defect Register showing SEC-014 and AUTH-022 as Priority 1 defects with plain-language descriptions
    Source evidence: The readiness workbook classified SEC-014 and AUTH-022 as Priority 1 defects.

    Both defects were still in progress. The listed fix dates were targets, the retests had not happened, and neither issue met the documented closure rule.

    Defect register showing affected sites, in-progress status, target fix dates, retest dates, workarounds, and launch blocker flags
    Source evidence: Future fix and retest dates did not mean the P1 defects were already closed.

    The contract allowed an initial phase of at least seven facilities, but only after every acceptance gate was met. The security gate required zero open P1 defects and successful retest evidence.

    Security and contract brief showing the four acceptance gates for a seven-site production phase
    Source evidence: Section 6.1 allowed a seven-site phase and made security evidence and sponsor approval hard gates.

    Option A—the immediate ten-site launch—was ineligible. Sites 8 through 10 were not ready, the two P1 defects remained open, and the Hartland sponsor had rejected a ten-site approval on the current evidence. Option C, delaying everything to August 31, was the fallback if Option B’s gates failed.

    Security release policy showing the Priority 1 closure standard, security controls, decision timeline, and fallback rule
    Source evidence: A P1 stayed open until its fix was deployed, the prescribed retest passed, and the evidence was recorded.

    The six traps

    The packet was built to punish shallow summarization. A model could extract plenty of correct facts and still answer the wrong question.

    1. 91.9% looked like a failed UAT gate. The contract actually measured three core flows at 535 of 555 tests, or 96.4%. The model had to report both numbers and explain which one governed.
    2. The packet contained two defect counts. An older PMO summary said 12 defects were open. The newer readiness workbook said 19: 2 P1, 6 P2, and 11 P3.
    3. An executive claimed all ten sites were approved. A later message from Hartland’s sponsor said the opposite. Authority and timestamp made the sponsor’s conditional seven-site position the governing evidence.
    4. A target fix date did not close a security defect. Closure required deployment, a passed retest, and recorded evidence—not a workaround, target date, or code-complete status.
    5. The reported CAD $730,000 delay cost mixed two different things. It combined CAD $250,000 in direct cost with a CAD $480,000 milestone invoice. The invoice was delayed cash timing, not lost revenue.
    6. The decision date was easy to move accidentally. The committee was deciding on August 6. Retests on August 12 and 13 controlled execution on August 17; they were not missing inputs for the August 6 recommendation.

    My Prompt

    I used the same prompt on all the models, the prompt:

    You are the engagement lead at Kavora Industries. The executive steering committee for the Hartland Energy Kavora Smart Operations program meets in 45 minutes. Review all six attached files and prepare an evidence-based launch recommendation.
    Your job is to determine whether Kavora should: (A) launch all 10 sites on August 17, 2026; (B) launch a conditional seven-site phase on August 17; or (C) delay all sites to August 31.
    Working rules:
    1. Use only the attached files. Do not browse the web.
    2. Do not ask clarifying questions. Make conservative assumptions and flag unknowns.
    3. Where sources conflict, use the most recent dated evidence and state the conflict.
    4. Cite the supporting file name and sheet, section, or message for every material conclusion.
    5. Recalculate key percentages and financial totals. Distinguish direct cost, cash timing, and contingency.
    6. Do not invent facts, approvals, dates, owners, or commitments.
    7. Do not send email, publish, post, or take any external action. Draft only.
    8. Follow the attached Kavora brand guide and use the attached board template as the visual basis for the deck.
    Create these three downloadable files with these exact names:
    - Kavora_Hartland_Launch_Decision_Memo.docx — maximum two pages. Lead with the recommendation and fallback; include evidence, financial treatment, conditions, risks, and the exact steering decision requested.
    - Kavora_Hartland_Steering_Committee.pptx — maximum six slides. Make it executive-ready, visual, concise, and branded.
    - Kavora_Hartland_Launch_Action_Tracker.xlsx — include action, owner, due date, gate, status, dependency, evidence location, and escalation path; add useful formulas or validation where appropriate.
    When the files are complete, reply in chat with a compact draft executive email of no more than 180 words. It must state the recommendation, hard conditions, fallback, and decision required. Do not send it.​‌
    [01_COWORK_INPUT_FILES]

    I have also included the input files I used for this session

    How each model performed

    Claude Opus 5: the winner

    Quality 100/100 · Design 9.5/10 · Credits 294

    Opus was the cleanest run. It reached the expected decision, reconciled every material conflict, used precise citations,

    separated direct cost from cash timing, and built the strongest action tracker. It also used the fewest credits.

    Opus steering deck control-plan slide with conditions, owners, dates, and status
    Deck receipt: A compact control plan with the conditions, owners, dates, and status visible.
    Opus action tracker gate-status sheet showing thresholds, owners, evidence, status, and linked actions
    Tracker receipt: The gate sheet makes the thresholds, owners, evidence, and linked actions easy to audit.
    Cowork cost receipt showing 294 credits for Claude Opus 5
    Cost receipt: 294 credits used for the task.

    Claude Sonnet 5: close and efficient

    Quality 98/100 · Design 8.9/10 · Credits 324 so far

    Sonnet’s decision and math were right, and the files were polished. I deducted two points because it proposed delegating the no-go call to the Program Director and Security Lead without evidence that they already held that authority.

    Sonnet steering deck control-plan slide with four gates and an August 14 checkpoint
    Deck receipt: Polished, clear, and executive-friendly.
    Sonnet memo excerpt delegating the no-go call to the Program Director and Security Lead
    Deduction receipt: The memo assigns decision authority the source pack never established.
    Cowork cost receipt showing 324 credits so far for Claude Sonnet 5
    Cost receipt: 324 credits “so far” when captured.

    GPT-5.6 Sol: accurate, but more expensive

    Quality 99/100 · Design 8.4/10 · Credits 771

    Sol showed strong evidence discipline and avoided inventing facts. It lost one point because the go/no-go checkpoint owner remained unresolved in the tracker. The larger difference was cost: 771 credits, more than 2.6 times the Opus run.

    Sol steering deck control-plan slide listing four dated gate controls
    Deck receipt: A clear six-slide story that uses the supplied template well.
    Sol action tracker excerpt showing the go or no-go checkpoint owner as not specified in source
    Deduction receipt: Sol flags the missing owner, but the operational tracker leaves the accountability unresolved.
    Cowork cost receipt showing 771 credits for GPT-5.6 Sol
    Cost receipt: 771 credits used for the task.

    Claude Fable 5: polished, but costly

    Quality 99/100 · Design 9.4/10 · Credits 2,658

    Fable reached the expected answer and produced some of the sharpest-looking files. I deducted one point because it inserted my name as the engagement lead even though the prompt and source pack never provided it. The run used just over nine times the Opus total without improving the decision.

    Fable steering deck control-plan slide with clear owners, dates, and status
    Deck receipt: One of the sharpest-looking decks in the set.
    Fable memo header showing Josh Cook as engagement lead without source support
    Deduction receipt: The memo invents an author identity that was absent from the prompt and packet.
    Cowork cost receipt showing 2,658 credits for Claude Fable 5
    Cost receipt: 2,658 credits used for the task.

    GPT-5.6 Terra: right answer, visible QA misses

    Quality 97/100 · Design 7.5/10 · Credits 360

    Terra handled the scenario date correctly, reconciled the stale claims, calculated the readiness and financial figures, and created all three files. Its PowerPoint still shipped with a literal \n on slide 1, text collisions on slide 6, and a low-contrast source line.

    Terra steering deck slide 1 showing a literal backslash-n in the decision card
    Slide 1 receipt: The decision card exposes a literal \n.
    Terra steering deck slide 6 showing overlapping text and low-contrast source text
    Slide 6 receipt: Both cards contain text collisions, and the source line is difficult to read.
    Cowork cost receipt showing 360 credits for GPT-5.6 Terra
    Cost receipt: 360 credits used for the task.

    Design review

    I kept design separate from decision accuracy. These are my visual scores across hierarchy, readability, template fidelity, spacing, consistency, QA, and day-to-day usability.

    ModelMemoDeckTrackerOverall
    Claude Opus 59.49.59.69.5
    Claude Fable 59.39.59.49.4
    Claude Sonnet 59.09.28.58.9
    GPT-5.6 Sol8.29.08.08.4
    GPT-5.6 Terra8.26.57.87.5
    A 10 means I would present the file without visual cleanup. An 8 means it is solid with a few practical issues. A 6 means visible mistakes survived into the deliverable. Overall is the simple average of the three artifact scores.

    Opus had the best balance. Fable was nearly as polished; its unsupported author name affected factual quality, not its design score. Sonnet delivered a strong memo and deck. Sol’s files were accurate but more functional than finished. Terra’s memo and tracker were usable, but the deck needed another QA pass.

    Where Terra lost design points

    • Decision memo: clean two-page render, but noisy raw-filename citations and inconsistent typography.
    • Steering deck: strong template match and clear story, but the literal \n, overlapping text, and low-contrast source line were visible QA failures.
    • Action tracker: complete and usable, but extremely wide, awkward to navigate, and unfinished for printing.
    Two-page Terra decision memo rendered side by side for design review
    Memo check: Both pages render cleanly. The remaining issues are hierarchy, citation noise, and typography consistency.
    Wide Terra action tracker with owner, due date, gate, status, dependency, evidence, escalation, and date-offset columns
    Tracker check: The content is strong. Its width and frozen-pane choice make horizontal navigation harder than it needs to be.

    Methodology and limits

    This was a controlled practical benchmark, not a permanent model leaderboard.

    • One final submission per model was evaluated.
    • Every run used the same prompt, six files, and output requirements.
    • No corrective follow-up prompt was used.
    • Decision correctness was checked before presentation quality.
    • Credits were recorded separately from quality and design.
    • Elapsed time was excluded.

    A different task could produce a different order. Coding, research, financial analysis, image work, and browser automation stress different capabilities. The series will keep the controls stable while changing the work.

    Final take

    Claude Opus 5 won this benchmark.

    All five models found the expected launch path. Opus separated itself by producing the best supporting work, the strongest design score, and the lowest credit total. Sonnet stayed efficient. Sol and Fable produced high-quality work at much higher task costs. Terra was accurate but needed a better presentation QA pass.

    Give the models real work. Lock the answer key. Keep the receipts.

  • Copilot Cowork First Month: How Admins Should Check July Consumption

    Copilot Cowork First Month: How Admins Should Check July Consumption

    The first month of Copilot Cowork is a baseline, not a verdict.

    For Microsoft’s qualifying Frontier grace-period cohort—and for the tenant in this guide—July was the first full billed calendar month of Copilot Cowork consumption.

    But opening one dashboard, copying the big number, and calling it “usage” will give you an incomplete story.

    There are really two questions:

    1. What did Cowork consume?
    2. Are people adopting it in a healthy, valuable way?

    Microsoft answers those questions in two different areas of the Microsoft 365 admin center. The Cost Management report shows Copilot Credit consumption. The Cowork Usage report shows adoption, tasks, activity, and retention.

    They use different definitions, date controls, and refresh cycles.

    This guide walks through both, shows what I found in a real first-month tenant, and gives admins a practical checklist for configuring and reviewing month one.

    Important July context: Copilot Cowork became generally available on June 16, 2026. Tenants with at least one user who used Cowork during the Frontier period from March 30 through June 16 received a billing grace period through June 30, with billing beginning July 1. That makes July the first full billed month for that cohort—not automatically for every tenant. Confirm your own activation and billing dates before labeling July “month one.”

    In this guide

    The two dashboards answer different questions

    Admin viewBest forDate behaviorRefresh behavior
    Copilot > Cost ManagementCredits, limits, policies, users, groups, agents/services, prepaid versus PayGoSupports an exact month such as July 2026Overview refreshes every 4 hours; Consumption refreshes every 2 hours
    Copilot > Cowork > UsageActive users, tasks, scheduled versus user-initiated work, retention, adoption trendsRolling 7, 28, 90, or 180 days; default is 28 daysCheck the report’s visible Last updated timestamp

    That distinction matters. The Cowork Usage report I reviewed on August 5 offered a 28-day range of July 5 through August 2. It did not offer a calendar-month July option. I could use it to understand adoption near the end of July, but I could not honestly present it as an exact July report.

    Cost Management Consumption page with summary cards for users, groups, and agents and services
    Cost Management can report the exact July 2026 consumption period.

    How to check July Copilot Cowork consumption

    1. Start in Cost Management, not the generic Copilot credits report

    In the Microsoft 365 admin center, go to:

    Copilot > Cost Management > Consumption

    For Cowork, this is the useful consumption surface. It lets you move between Users, Groups, and Agents and services.

    Do not confuse it with Reports > Usage > Microsoft 365 Copilot > Credits. That report has a different scope and is not the report I would use to isolate Cowork consumption.

    2. Select July 2026

    Open Time period and select July 2026.

    July month selector
    Use the calendar-month selector in Cost Management when you need an exact July baseline.

    Here is the small trap I saw in the August 5 tenant UI: changing between Users, Groups, and Agents and services reset the time period to Month to date. Re-select July on every view and confirm the label before recording a number or exporting data.

    3. Use Agents and services to isolate Cowork

    Start with Agents and services. This is the cleanest way to separate Cowork from Work IQ API or another consumptive service covered by the same policy.

    Cowork service consumption for July
    The Agents and services view is the best starting point for the Cowork total.

    Record at least:

    • Cowork Credits used
    • Active users
    • Sessions
    • The report period
    • The time and date you captured the report

    That last item is not busywork. Microsoft says the Consumption tab refreshes every two hours, while exported reports are point-in-time snapshots. A dashboard and an export taken at different moments can temporarily disagree.

    4. Review users for concentration and limit pressure

    Move to Users, re-select July, then inspect:

    • Credits used
    • Monthly credit limit
    • Percentage of the limit consumed
    • Sessions
    • Last activity date
    • Assigned spending policy
    • Daily credit usage
    User consumption details with spending policy and monthly limit
    The user drawer connects total usage to the assigned policy, service access, and individual limit.
    Daily user credit usage
    Daily usage shows when concentration occurred. Pair it with Cowork task data or Purview audit evidence before deciding why it occurred.

    Do not just chase the largest user. Ask whether the work was intentional, repeatable, and valuable. A heavy user may be delivering the best ROI in the tenant. The useful signal is high consumption without an understood scenario, owner, or outcome.

    5. Review groups, then re-check the period

    The Groups view helps you compare the audiences targeted by spending policies. It is useful for spotting:

    • One group consuming most of the credits
    • A group with low consumption after you compare the Consumption view with its enabled scope under Configuration
    • A pilot team whose limits are too low for its intended workload
    • Overlapping memberships that make policy behavior harder to explain
    July consumption by group
    Group totals help admins compare the audiences governed by different spending policies.

    User and group totals do not always match perfectly. Users can belong to multiple groups and policies, and policy changes during a billing period can affect what appears at user level. Treat these views as different analytical cuts, then reconcile material differences against exports and policy history.

    6. Check the July spending trend separately

    Return to Cost Management > Overview and set the Spending trend card to July.

    July cumulative spending trend
    The Spending trend card has its own period control. Do not assume the page-level cards and chart are showing the same month.

    Look for:

    • Large daily spikes
    • A steady upward ramp after enablement
    • Weekend spikes to compare with scheduled-task history
    • A shift from prepaid capacity into PayGo
    • A sudden drop that coincides with a limit or access change

    In the tenant reviewed for this guide, July was spiky rather than smooth. The largest single day used more than 12,000 credits, while several days were below 1,000. That is a prompt to investigate which scenarios or scheduled tasks ran on the high days—not proof that the spikes were waste.

    What the first full month looked like in this tenant

    The following numbers are an anonymized, point-in-time snapshot captured on August 5, 2026.

    Exact July consumption

    • 136,956 Cowork Credits in the Agents and services view
    • 19 active Cowork users in that service view
    • 260 sessions across the user rows
    • 18 users with non-zero recorded credits at capture time
    • 6,150 median credits across the 19 user rows
    • 7,208 average credits across the 19 user rows
    • 526.7 credits per session across the user rows
    • The largest user represented 22.0% of user-level credits
    • The top five users represented 59.8% of user-level credits
    • 6 users had consumed at least half of their individual limit
    • 0 users had consumed 80% or more of their individual limit

    The service, group, user, and trend views differed by a few hundred credits at capture time: roughly two-tenths of one percent. That is small enough to be consistent with refresh timing and aggregation differences, but it should still be documented. Use one defined source for the headline number—here, the Cowork row under Agents and services—and preserve exports when you close the month.

    Do not turn credits directly into an invoice unless you know the billing path

    Microsoft lists PayGo at $0.01 USD per Copilot Credit. At pure retail PayGo, 136,956 credits would be a $1,369.56 USD usage equivalent.

    That is not automatically the tenant’s invoice. This tenant’s policies used prepaid capacity. When capacity packs and a subscription with P3/pay-as-you-go are both selected, Microsoft documents the order as capacity packs first, then P3 prepaid credits, then PayGo. Discounts, prepaid commitments, shared capacity, and billing arrangements can all change the actual cost. Report credits and billing source together.

    Now check adoption—but label the date range honestly

    Go to:

    Copilot > Cowork > Usage

    Cowork Usage rolling date selector
    Cowork Usage uses rolling ranges. The 28-day view here covers July 5 through August 2, not calendar July.

    The report’s four headline metrics are:

    • Active Cowork users
    • Total Cowork tasks
    • Average tasks per active user
    • Retained Cowork users
    Cowork adoption cards
    Always capture the date range and Last updated timestamp with the headline metrics.

    Microsoft defines a task as one Cowork conversation thread, including user-initiated conversations and scheduled task executions. Follow-up questions in the same thread remain part of the same task.

    The same report defines a retained user as someone active in both the previous seven-day period and the most recent seven-day period. It is a useful continuity signal, but it is not a conventional month-over-month cohort retention rate.

    Microsoft's task definition in the report
    A task is not the same thing as a prompt, run, or Cost Management session.

    The report was viewed on August 5 and displayed Last updated: August 3. For the rolling 28 days from July 5 through August 2, this tenant showed:

    • 18 active users
    • 332 total tasks
    • 18.44 average tasks per active user
    • 18 retained users
    • 103 user-initiated tasks
    • 229 scheduled tasks

    That means about 69% of tasks were scheduled. It is a valuable operational signal: Cowork was being used as automation, not only as a conversational tool.

    Daily active users and task mix
    Read daily active users beside the task-type mix to understand how engagement is occurring.
    Scheduled and user-initiated task trend
    The late-July scheduled-task increase deserves a scenario-level review before changing limits.

    But the average hid a highly concentrated pattern: one user accounted for 206 of the 332 tasks, or about 62%. The median active user had only 6 tasks. Fifteen users were active on two days or fewer.

    This does not mean the rollout failed. It means the first-month story is a scheduled-heavy tenant and a highly concentrated user-level pattern, plus a low-frequency long tail. Month two should validate the dominant user’s scenarios while improving repeat, intentional use across the rest of the pilot.

    The first-month admin playbook

    Days 1–3: Prove the plumbing

    • Confirm each intended user has a Microsoft 365 Copilot license.
    • Confirm usage-based billing is enabled and tied to the correct subscription or prepaid capacity.
    • Confirm Cowork is discoverable to the intended audience.
    • Confirm the intended security groups are assigned to an active spending policy.
    • Run a controlled task and verify it appears after each report updates. Use the visible Last updated timestamp for Cowork Usage and allow for Cost Management’s documented refresh window.
    • Capture the initial policy, billing method, and access configuration before changing anything.

    Discoverability and access are separate controls. The discovery setting controls tenant-wide visibility, while Cost Management controls scoped access and spend. Microsoft says usage-based billing setup overrides the discovery setting for users covered by that setup, so validate both the tenant-wide experience and the intended scoped audience.

    Week 1: Establish safe limits and alerts

    Start with group-based policies that match the rollout audiences. A practical pilot configuration should include:

    • A policy-level monthly cap
    • A per-user monthly cap
    • Alerts well before the hard cap
    • Named operational and financial owners for those alerts
    • Explicit agents and services included in the policy
    • A documented billing method

    Also make an explicit decision about Allow new services and agents as they become available. Microsoft’s current policy setup selects that option by default. That can be convenient, but a controlled pilot should not inherit future service access without an owner and review process.

    Spending policies impose access and consumption limits; they do not reserve or allocate a slice of the underlying Copilot Credit capacity to a group. Treat policy limits as guardrails, not as capacity reservations.

    Policy precedence warning and configuration
    When multiple policies apply, Microsoft enforces the most permissive applicable policy; settings are not combined.

    Microsoft’s current precedence order for overlapping policies is:

    1. Highest per-user limit
    2. If tied, highest overall policy limit
    3. If still tied, newest policy

    That is easy to get wrong. Adding a “special” policy can accidentally make someone less restricted than intended.

    Monthly hard-cap behavior
    A hard cap blocks covered services for the rest of the month and resets on the first day of the next month.
    Policy alert threshold
    Set an early threshold so an alert creates time to investigate instead of merely announcing that the budget is gone.
    Per-user monthly cap
    Per-user caps protect a shared policy from one runaway or misunderstood workload.

    When the run rate is unknown, I start the first alert at 50%. Once you understand the run rate, use multiple operational checkpoints outside the platform—such as weekly reviews at 50%, 70%, and 85%—even if the product sends a single configured alert threshold.

    Use least privilege for the operating team. AI Administrators and License Administrators can create spending policies, manage limits and alerts, and view Cost Management. Selecting or overriding the billing method requires a Global Administrator or Billing Administrator. After a spending policy is created, its billing method cannot be changed; Microsoft requires deleting and recreating that policy.

    Week 2: Review adoption, not just activation

    Ask five questions:

    1. How many targeted users became active through a direct prompt or a scheduled execution?
    2. How many came back in the next seven-day period?
    3. Is work primarily user-initiated or scheduled?
    4. Is usage concentrated in one person, team, or scenario?
    5. Which enabled users have no visible activity and need a use case, training, or removal from the pilot?

    The Usage table lists people with Cowork activity dating back to April 1, 2026, while its summary cards respond to the selected rolling period. Do not use the total table row count as your active-user count. Also remember that an active user can be counted because they prompted Cowork or because a scheduled task executed on their behalf. Active-user coverage is not identical to direct human adoption.

    Week 3: Connect consumption to scenarios

    For the top users and largest days, document:

    • Business scenario
    • Owner
    • Frequency
    • User-initiated or scheduled
    • Inputs and data sources
    • Outputs and downstream actions
    • Credits consumed
    • Time saved or risk reduced
    • Whether a lower-cost model or tighter scope can produce an acceptable result

    Consumption alone is not value. A 5,000-credit task that replaces two days of expert work may be excellent. A 500-credit task running every hour with nobody reading the output may not be.

    Week 4: Freeze the baseline and decide month two

    • Re-select the exact July period in every Cost Management view.
    • Export users, groups, and agents/services after the reporting refresh.
    • Record the Cowork Usage rolling range and Last updated timestamp.
    • Preserve policy settings and group membership as of month-end.
    • Explain material differences between views.
    • Review users near limits and compare the configured audience with users showing no repeat activity.
    • Classify scheduled tasks as keep, tune, pause, or investigate.
    • Set month-two targets before raising budgets.

    The metrics I would put on an admin scorecard

    MetricFormulaWhat it tells you
    Active coverageActive users ÷ enabled target usersHow much of the intended audience generated direct or scheduled activity
    Retained-to-active ratioRetained users ÷ active usersA simple continuity signal, not a conventional cohort-retention rate; retained users compare the two most recent seven-day periods while active users cover the selected range
    Scheduled-task shareScheduled tasks ÷ total tasksHow much usage is recurring automation versus direct interaction
    Credits per active userCowork credits ÷ active usersOverall consumption intensity when the periods align
    Credits per sessionCowork credits ÷ Cost Management sessionsRelative session intensity within the consumption report
    Limit utilizationCredits used ÷ applicable policy limitHow close a user or policy is to its cap; the limit does not reserve underlying capacity
    Top-five concentrationCredits from top five users ÷ total user creditsWhether the rollout depends on a few users
    Average-to-median gapAverage usage compared with median usageWhether a small number of heavy users are distorting the average

    Only combine figures whose date ranges and definitions match. A Cowork task is not a Cost Management session. In this review, the adoption report covered July 5 through August 2 while the consumption report covered July 1 through July 31, so I did not calculate credits per task.

    Month-one red flags worth investigating

    • A view silently returned to Month to date after you changed tabs.
    • The Overview headline cards and Spending trend chart are showing different periods.
    • A user or group is above 80% of its cap with time left in the month.
    • Scheduled work dominates consumption but has no named owner or reviewed output.
    • Average usage is high while median usage is low.
    • A service other than Cowork is consuming credits under the same policy.
    • A group has access but almost no activity.
    • A spending policy includes new services automatically without a review process.
    • Several groups overlap, producing a more permissive policy than expected.
    • Finance is treating credits multiplied by retail PayGo as the invoice even though prepaid capacity or P3 applies.
    • Security assumes all Microsoft Purview controls apply identically to Cowork. As of this writing, Microsoft documents DLP for general Cowork AI interactions as unsupported, even though auditing, sensitivity labels, eDiscovery, data lifecycle, communication compliance, and other controls have documented support. Cowork’s local Edge browser tasks are narrower: Microsoft says they inherit existing browser DLP, Conditional Access, and site policies. Review both control paths rather than inheriting assumptions from another Copilot surface.

    What I would change for month two

    For this tenant, I would not start by cutting the heavy user or scheduled workload. I would first confirm whether that automation is producing a real business outcome. If it is, preserve it and establish an owner, expected run rate, and value measure.

    Then I would:

    1. Create a repeat-use goal for the low-activity users.
    2. Review the five largest credit consumers and the largest daily spikes.
    3. Audit scheduled tasks for frequency, duplicate runs, and unused output.
    4. Separate Cowork and Work IQ API policy reporting wherever the rollout design allows it.
    5. Keep alerts early enough to react before a hard cap blocks work.
    6. Raise limits only for a documented scenario with evidence of value.
    7. Review plugins, model availability, browser use, audit coverage, and Purview support as part of the same operating rhythm—not as a one-time launch checklist.

    Final thought

    The goal of month one is not to prove that every user became a power user. It is to leave with a trustworthy baseline: who used Cowork, what they ran, what it consumed, which work repeated, and which controls need to change next.

    That is the admin story July can tell—if you use both dashboards and keep their definitions straight.

    Sources

  • Copilot Cowork Billing Setup: Turn On GA Without Opening the Credit Floodgates

    Copilot Cowork Billing Setup: Turn On GA Without Opening the Credit Floodgates

    Copilot Cowork is generally available.

    That is the headline.

    But before you run into the tenant, turn everything on, and let users start throwing goals and long-running tasks at it, you need to understand the billing side.

    Copilot Cowork runs on Copilot Credits. That means the admin work is not just, “Can the user access it?” It is also, “Who can spend credits, how much can they spend, which billing method gets charged, and who gets warned before usage goes sideways?”

    Turning Copilot Cowork on is easy. Turning it on responsibly is where the real admin work starts.

    In this walkthrough, I am setting up Copilot Cowork billing from the Microsoft 365 admin center and showing the choices I would pay attention to before giving users access.

    What changed with Copilot Cowork GA

    Microsoft announced Copilot Cowork general availability on June 16, 2026. Cowork requires a Microsoft 365 Copilot user subscription license, and Cowork usage is billed separately on a usage basis using Copilot Credits.

    That matters because Copilot Cowork is not just another chat surface. It is designed for complex, long-running, multi-tool work. It can retrieve context, call tools, use models, create artifacts, and keep working through a task. All of that value has a meter behind it.

    Microsoft’s current Cost Management experience for usage-based billing applies to Copilot Cowork and Work IQ API right now, with more agents and services expected to come into that experience over time.

    Microsoft 365 admin center Cost management page showing Copilot Cowork and Work IQ API usage-based billing
    Cost Management in the Microsoft 365 admin center is where Copilot Cowork and Work IQ API usage-based billing is configured.

    Start with the right mental model

    Access and consumption are two different things.

    Access lets a user get into the experience. Consumption happens when the experience starts doing work. If Cowork runs a goal, uses a model, retrieves context, calls tools, or performs longer-running work, that activity can consume Copilot Credits.

    That is why billing policies matter. Without them, you are basically handing out an AI gas card and hoping the bill looks reasonable at the end of the month.

    That is not a strategy.

    The better approach is billing plus governance. Set the billing method, define the spending policy, decide who is in scope, add limits, configure alerts, and then expand once you understand real usage.

    Check your admin roles before you start

    If you do not see the option in the admin center, check your role first.

    Billing setup and policy governance may not be handled by the same person in a real organization. Your Microsoft 365 admin, AI admin, Power Platform admin, billing admin, and finance owner might all be different people.

    That matters because the person configuring the billing method needs the right permissions, and the person managing policy limits and alerts also needs the right permissions. Do not discover that halfway through the rollout call. Get the right people in the room before you begin.

    Choose the billing method intentionally

    In the setup flow, you may see more than one billing method. In my tenant example, I had Capacity Packs available and also had the option to use a pay-as-you-go subscription.

    Capacity Packs let you bill against prepaid Copilot Credit capacity. Pay-as-you-go can keep services running once capacity pack credits run out, but it also means the meter can keep running against the connected subscription.

    Neither option is automatically good or bad. The point is to make the choice on purpose.

    Billing method options for Copilot Credits showing Capacity Packs and pay-as-you-go subscription
    Select the billing method intentionally. Capacity Packs and pay-as-you-go behave differently when credits run out.

    For a personal tenant or a controlled pilot, I would usually start with the most bounded option. For a production rollout where interruption would be a problem, you may choose to include pay-as-you-go, but that decision should involve the billing owner.

    What if you have no Copilot Credits?

    If your tenant does not already have Copilot Credits available, do not let that push you into an unlimited rollout. Start with pay-as-you-go, but treat it like a metered pilot, not an open tab.

    Microsoft lists PayGo for Copilot Cowork at $0.01 per Copilot Credit. That means 25,000 credits would be about $250 if you let usage land entirely on pay-as-you-go.

    That is the point where the economics should make you stop and re-check the licensing path. Microsoft lists Copilot Studio credit packs at 25,000 Copilot Credits for $200 per pack per month. So if you expect usage to get anywhere near 25,000 credits in a month, PAYG is no longer just a convenient starter option. It may be more expensive than buying a capacity pack.

    The practical setup I would use is this: enable PAYG to get started, set the monthly policy limit at 25,000 credits, turn alerts on well before that point, and review usage before the policy hits the cap.

    • At 10,000 credits, check whether the pilot group is using Cowork the way you expected.
    • At 20,000 credits, start the capacity pack conversation because the $200 pack economics are already close.
    • At 25,000 credits, do not just increase the PAYG limit without a decision. Buy a P3 or Copilot Credit capacity pack if monthly usage is becoming predictable.

    PAYG is still useful. It is the fastest way to start when you have no credits, and it is a good safety net for overages. But once usage becomes steady, capacity packs are the cleaner budgeting conversation.

    Do not assume the visible credit pool is unused

    This is the part admins need to slow down on.

    In my example, the Cost Management setup showed 50,000 credits available. But in the Power Platform admin center, 10,000 of those Copilot Studio messages were already allocated to a specific environment.

    Power Platform admin center Capacity add-ons page showing Microsoft Copilot Studio messages allocated from a 50,000 credit pool
    Before assigning a Cowork policy, check whether Copilot Credits are already supporting other Power Platform or Copilot Studio workloads.

    The lesson is simple: do not treat the number in one setup panel as the whole story.

    Copilot Credits can support more than one AI workload. If your organization is already using Copilot Studio, Dynamics 365, Power Platform, or other metered AI capabilities, make sure you understand what that credit pool is already expected to cover.

    The last thing you want is for a Cowork pilot to burn through credits that another production agent was relying on.

    Do not start unlimited

    The setup flow gives you the choice to limit monthly spending or leave monthly spending unlimited.

    My recommendation for most organizations starting with Copilot Cowork is simple: do not start unlimited.

    Launch it like a controlled pilot. Give it a monthly policy budget, watch usage, learn what normal looks like, then increase the limit later if the usage is healthy.

    In my example, I set the policy budget to 40,000 credits because I wanted to keep 10,000 credits available for other Copilot Studio usage in the tenant. Your number will be different. The principle is the same.

    Set a per-user monthly limit

    The per-user monthly spending limit is optional, but I would seriously consider enabling it.

    Without a user limit, one person can potentially consume a large chunk of the available credits. They may not be doing anything malicious. They might just be experimenting, running long tasks, testing browser automation, asking for big research outputs, or sending Cowork work that should have stayed in regular Copilot Chat.

    That kind of experimentation is a sign adoption is happening. But adoption without guardrails becomes waste.

    Set a policy-level monthly budget, consider a per-user monthly limit, and configure alerts before activating the policy.

    In the demo tenant, I used a 40,000 credit policy limit and a 20,000 credit per-user limit because the pilot only had two users. That is an example, not a universal recommendation.

    The right number depends on how many users are in scope, what tasks they will run, how much existing Copilot Credit capacity is already committed, and how much risk you are willing to tolerate during the first rollout.

    Turn alerts on before usage surprises you

    Do not wait until the end of the month to learn usage went sideways.

    Cost Management lets you define alerts so the right people get notified when usage reaches a threshold. That could be the Microsoft 365 admin, billing owner, finance stakeholder, platform owner, or whoever is accountable for the pilot.

    In my example, with a 40,000 credit monthly policy limit, I used a 30,000 credit alert threshold. That gives the owner time to investigate before the policy hits the cap.

    This is how you move from reactive admin to proactive admin.

    Do not activate for everyone by default

    This is the part of the setup that is easy to rush.

    If you activate broadly, you may be enabling access across the whole organization depending on how the policy is configured. For some organizations, that might be fine. For most first rollouts, I would not start there.

    Customize the setup configuration and scope the policy to a security group.

    Microsoft 365 admin center security group creation screen with the name Copilot Cowork Access
    Create a clear security group for the pilot, such as Copilot Cowork Access, instead of starting with the entire tenant.

    Use a clear name like Copilot Cowork Access or Copilot Cowork Pilot Users. Add a small group of trusted users first: admins, builders, business champions, finance, operations, or the people who will give you useful feedback.

    Think of this like a pilot program. You want people who will use it seriously, report what worked, report what wasted credits, and help you understand whether the budget is realistic.

    Cost Management access group configuration scoped to specific security groups
    Scope the policy to specific groups so you know exactly who can consume Copilot Credits through Cowork.

    Review the policy before you activate it

    Before clicking activate, review the setup like a change request:

    • Which services are enabled by the policy?
    • Which billing method will be charged?
    • Is the monthly spending limit set?
    • Is the per-user limit set?
    • Who receives alerts?
    • What threshold triggers those alerts?
    • Which users or groups are in scope?
    • Are other workloads already using the same credit pool?

    In my setup, the final shape was:

    • Billing method: Capacity Packs
    • Policy monthly limit: 40,000 credits
    • Per-user monthly limit: 20,000 credits
    • Alert threshold: 30,000 credits
    • Access scope: a specific security group for Copilot Cowork access

    Again, those are demo values. Do not copy them blindly. Copy the pattern: limit, alert, scope, observe, then scale.

    The rollout pattern I would use

    If I were rolling out Copilot Cowork in a tenant, I would not start with the whole organization. I would use this pattern:

    1. Confirm the admin roles needed for billing and policy management.
    2. Review existing Copilot Credit usage in Microsoft 365 and Power Platform.
    3. Choose the billing method with the billing owner involved.
    4. Create a bounded monthly spending policy.
    5. Add a per-user limit for the pilot group.
    6. Set alerts before the policy limit is reached.
    7. Create a security group for pilot access.
    8. Add a small number of serious users.
    9. Monitor usage and identify what tasks burn credits.
    10. Increase limits or expand access only after you understand normal usage.

    This gives you control. You know who has access, what they can consume, which budget they are under, and when someone needs to pay attention.

    That is how you test Copilot Cowork without turning the tenant into a free-for-all.

    Final thought

    Copilot Cowork is powerful because it can take on real work. But the more real the work gets, the more important the admin controls become.

    Cost Management is not the boring part of Copilot Cowork. It is the part that lets you adopt it without surprising finance, burning shared credits, or giving every user an unlimited meter on day one.

    Start controlled. Learn the usage. Then scale with confidence.

    Sources

  • Copilot Cowork /goal Prompt Vault:

    Copilot Cowork /goal Prompt Vault:

    I’ve been using Copilot Cowork since March, and the biggest lesson is simple: stop treating it like a chat box.

    Give it a real goal.

    The /goal skill is where Cowork gets serious. When you point it at the right files, context, assets, and instructions, it can move from “give me ideas” to “go figure this out, build the plan, show me the sources, and wait for approval.”

    Each prompt will include the use case, the exact prompt, the expected result, and the follow-up prompts you can use to push Cowork further.

    Prompt #1 — Create an Organization Branding Kit

    Use case

    Use this when you want Copilot Cowork to pull together a practical branding kit for an organization using available internal context, public website branding, existing files, logos, icons, and PowerPoint decks.

    This is useful when the organization already has branding scattered across folders, decks, documents, images, and websites, but there is no clean source of truth.

    The goal is simple: get Cowork to discover what already exists first, then build the branding kit from real assets instead of guessing.

    The /goal prompt

    /goal Create a complete branding kit for [Organization Name].
    Use all available context you can access, including internal files, shared documents, PowerPoint decks, Word documents, PDFs, images, logos, icons, and the organization’s public domain: [Website URL].
    Your first job is discovery.
    Look for:
    - Existing logos and icon files
    - PNG and JPEG image files containing official logos, icons, or brand graphics
    - PowerPoint decks that appear to be templates or close to templates
    - Sales decks, pitch decks, marketing decks, one-pagers, proposals, product sheets, and internal documents
    - Existing colors, fonts, layouts, screenshots, visual patterns, and repeated design styles
    - Website branding, tone, product language, and positioning from the public domain
    Asset handling guardrails:
    - Do not recreate, redraw, regenerate, restyle, or reimagine any existing logo or icon.
    - Use the actual existing image assets when they are available.
    - Prefer the original PNG or JPEG files for logos, icons, and other brand images.
    - If multiple versions exist, identify the best available source file and note the others.
    - Preserve the original appearance of logos and icons, including proportions, colors, spacing, and transparency.
    - Do not generate substitute logos or substitute icons.
    - If an official asset cannot be found, clearly mark it as missing and recommend that it be provided.
    - If an asset is unclear, low quality, duplicated, or inconsistent, flag it instead of attempting to remake it.
    Do not invent brand rules. If something is not clearly available, mark it as “recommended” instead of “confirmed.”
    Create a practical branding kit that includes:
    1. Brand overview
    Summarize the organization’s visual identity, positioning, tone, and audience.
    2. Logo guidance
    Identify the available logo and icon assets, especially the PNG and JPEG files, and recommend how they should be used across slides, documents, social posts, and web assets.
    3. Color palette
    Extract or infer the main brand colors from existing assets. Include hex codes when possible. Separate confirmed colors from recommended supporting colors.
    4. Typography guidance
    Identify any fonts used in existing materials if possible. If the fonts are unclear, recommend a clean Microsoft-friendly font stack that matches the brand.
    5. Presentation style
    Review existing PowerPoint decks and identify the best deck or slides to use as the closest template. Explain what makes it the best starting point.
    6. Slide design rules
    Create clear rules for title slides, section dividers, content slides, screenshots, diagrams, callout slides, and closing slides.
    7. Icon and image style
    Define the icon style, image style, screenshot style, and visual treatment the organization should use consistently, based only on existing assets and patterns you find.
    8. Voice and tone
    Define how the organization should sound in external content, internal content, sales material, and social posts.
    9. Reusable asset recommendations
    List the assets that should be created next, including PowerPoint template, Word template, social post template, proposal template, one-page product sheet, and icon set.
    10. Gaps and questions
    List anything missing, unclear, inconsistent, or worth confirming before the branding kit becomes official.
    Before creating the final branding kit, show me:
    - The files and sources you found
    - Which logo and icon image files you found, especially PNG and JPEG assets
    - Which PowerPoint deck is the best starting point
    - Any assumptions you are making
    - Any missing or questionable assets
    - The proposed structure for the final branding kit
    Wait for my approval before creating the final output.

    Expected result

    Cowork should search the available context, identify the real brand assets, review existing decks and documents, and build a branding kit grounded in what already exists.

    The important part is that it should separate confirmed brand rules from recommended brand rules.

    • Confirmed means Cowork found evidence in the files, website, images, decks, or documents.
    • Recommended means Cowork is filling a gap based on the closest available context.

    That distinction matters because branding work can go sideways fast when the AI starts inventing things. The prompt forces Cowork to find the source material first, flag gaps, and wait for approval before creating the final output.

    The guardrails around logos and icons are also important. Cowork should use the actual image files where possible. It should not recreate or redesign official assets.

    How to extend / prompt more

    Once Cowork creates the first version of the branding kit, keep driving it with follow-up prompts like these:

    Add this <logo.png> as the main logo for the branding kit. Use the actual image file. Do not recreate, redraw, or modify the logo.
    Use this <icon.png> as the primary app or product icon. Keep the original proportions, colors, and transparency.
    Using this brand kit, generate a new Word template I can use for general letter headings.
    Using this brand kit, create a Word proposal template with a cover page, section headings, body styles, callout sections, and a closing page.
    Using this brand kit, create a PowerPoint template structure with a title slide, section divider, agenda slide, content slide, screenshot slide, quote or callout slide, and closing slide.
    Use this existing <presentation.pptx> as the closest visual reference and create a cleaner template based on its style.
    Create a one-page brand cheat sheet for employees that shows the logo usage, colors, fonts, tone, and common do and don’t rules.
    Create a social post template guide using this branding kit for LinkedIn and X.
    Create a reusable prompt I can paste into Cowork any time I want it to generate on-brand content for this organization.
    Create a list of missing brand assets I should upload next, including logo formats, icons, PowerPoint templates, Word templates, screenshots, and product images.
    Create a lightweight version of this branding kit for sales, delivery, and support teams.
    Review this generated template against the branding kit and tell me what is off-brand before I use it.

    Final note

    This is where Copilot Cowork starts to become more than a chat experience.

    You are giving it a goal, pointing it at real organizational context, adding guardrails, and forcing it to work from actual assets.

    That is how you get better output.


    Prompt #2: Find Copilot-Only Automation Use Cases from Work IQ

    This prompt is designed for Copilot Cowork when you want to look across your Microsoft 365 work patterns and find practical automation opportunities that can be handled with Copilot first.

    The key constraint is important:

    The first three use cases must be possible using only Copilot.

    No Power Automate. No Azure. No custom code. No developer work.

    Then the prompt asks for one bonus use case where Power Automate is allowed, but only if the workflow truly needs a trigger, system action, approval, notification, or reliable automation beyond what a prompt can do.

    This is a strong Frontier grace period test because it helps you answer a very practical question:

    What recurring work can we improve with Copilot alone before we start building anything heavier?

    What this /goal does

    • Looks for recurring work patterns across Microsoft 365
    • Identifies repeated meeting prep, follow-ups, updates, summaries, and content work
    • Shortlists possible Copilot-only automation use cases
    • Selects the top 3 practical use cases
    • Builds setup blueprints for each one
    • Drafts scheduled prompts, Cowork skill drafts, /goal workflows, or Copilot Chat prompts
    • Adds one bonus Power Automate use case where automation is actually justified

    The /goal prompt

    /goal
    Analyze my Work IQ and Microsoft 365 work patterns to identify the top Copilot-only automation use cases.
    Purpose:
    I want you to look deeply at my available Microsoft 365 work context and Work IQ signals to find the best recurring work patterns that can be improved using only Microsoft 365 Copilot and Copilot Cowork.
    This should not be a generic brainstorm.
    I want practical use cases that can be implemented with Copilot capabilities only, such as:
    - Microsoft 365 Copilot scheduled prompts
    - Copilot Cowork skills
    - Copilot Cowork /goal workflows
    - Microsoft 365 Copilot Chat prompts
    - Copilot agents where available
    For the first 3 use cases, do not require Power Automate, Azure, custom code, external connectors, or developer work.
    After the top 3 Copilot-only use cases, include a 4th bonus use case where Power Automate is allowed.
    Primary goal:
    Find the top 3 recurring work patterns that can be improved using only Copilot, then create a practical setup plan for each one.
    Then identify one additional use case where Power Automate would be the better option.
    Important guardrails:
    - Do not force Power Automate into the first 3 use cases
    - Do not recommend custom code
    - Do not recommend Azure Functions
    - Do not recommend complex integrations
    - Do not recommend external tools
    - Do not create anything unless I approve
    - Do not schedule anything unless I approve
    - Do not send emails
    - Do not post in Teams
    - Do not modify files
    - Do not create automations
    - Clearly separate Copilot-only use cases from the Power Automate use case
    - Base recommendations on real work patterns where possible
    - If Work IQ signals are limited, use available Microsoft 365 context such as meetings, emails, chats, files, recurring collaboration patterns, and repeated tasks
    - Be practical
    - Recommend the lightest approach that solves the problem
    Step 1: Create a work plan
    Before starting, create a short work plan.
    Explain:
    - What Microsoft 365 work context you will analyze
    - What Work IQ signals you will look for
    - How you will identify repeated work patterns
    - How you will decide whether something can be handled with Copilot only
    - How you will decide whether something needs Power Automate
    - What outputs you will create
    Then proceed.
    Step 2: Analyze Work IQ and Microsoft 365 work patterns
    Review available work signals and look for recurring patterns such as:
    - Repeated meeting prep
    - Repeated meeting follow-up
    - Weekly status updates
    - Recurring stakeholder updates
    - Repeated executive summaries
    - Repeated project updates
    - Repeated customer updates
    - Repeated inbox triage
    - Repeated action item tracking
    - Repeated document drafting
    - Repeated research tasks
    - Repeated content creation
    - Repeated review workflows
    - Recurring decisions that need summaries
    - Recurring risks that need monitoring
    - Work that happens before meetings
    - Work that happens after meetings
    - Work that requires pulling context from multiple Microsoft 365 sources
    - Work that could run on a schedule
    - Work that could be standardized with a skill
    - Work that could be delegated to Cowork
    - Work that should stay lightweight in Copilot Chat
    Create a short summary of the main patterns you found.
    Step 3: Create a shortlist of Copilot-only candidates
    Create a shortlist of at least 10 possible Copilot-only use cases.
    For each candidate, include:
    - Use case name
    - Work pattern it is based on
    - Why it matters
    - Who benefits
    - Frequency
    - Business value
    - Time savings potential
    - Complexity
    - Data sensitivity
    - Best Copilot capability
    - Why this can be done with Copilot only
    - Why it does not need Power Automate
    The best Copilot capability should be one of:
    - Microsoft 365 Copilot scheduled prompt
    - Copilot Cowork skill
    - Copilot Cowork /goal workflow
    - Microsoft 365 Copilot Chat prompt
    - Copilot agent, where available
    - Combination of Copilot capabilities only
    Step 4: Select the top 3 Copilot-only use cases
    Rank the candidates and select the top 3 use cases that can be implemented with Copilot only.
    Score each candidate from 1 to 5 for:
    - Business value
    - Frequency
    - Time savings
    - Ease of setup
    - User adoption likelihood
    - Scheduled prompt fit
    - Cowork skill fit
    - Cowork /goal fit
    - Governance risk
    - Data sensitivity
    - Repeatability
    For each top 3 use case, explain:
    - Why it made the top 3
    - Why it is more valuable than the other candidates
    - Why it can be handled with Copilot only
    - Why Power Automate is not required
    - Whether the best approach is a scheduled prompt, Cowork skill, /goal workflow, Copilot Chat prompt, or combination
    Step 5: Build the setup blueprint for each Copilot-only use case
    For each of the top 3 Copilot-only use cases, create a complete setup blueprint.
    Include:
    - Use case name
    - Business problem
    - Work pattern evidence
    - Main users
    - Trigger or schedule
    - Required context
    - Microsoft 365 sources needed
    - Recommended Copilot capability
    - Setup steps
    - Prompt text
    - Expected output
    - Human review points
    - Governance notes
    - What success looks like
    - What could go wrong
    - What should not be automated
    - How to test it
    - How to improve it after testing
    Step 6: Draft the actual Copilot assets
    For each top 3 use case, draft the actual Copilot setup assets.
    If it should be a scheduled prompt, include:
    - Scheduled prompt name
    - Recommended schedule
    - Why that schedule makes sense
    - Full scheduled prompt text
    - Expected output
    - Who should review it
    - What the user should do with the output
    - Guardrails
    - What the scheduled prompt should not do
    If it should be a Copilot Cowork skill, include:
    - Skill name
    - Skill purpose
    - When users should use it
    - What the skill should ask the user
    - What context it needs
    - What steps it should follow
    - What outputs it should create
    - What it should never do
    - How it should handle uncertainty
    - Approval points
    - Validation checklist
    If it should be a Copilot Cowork /goal workflow, include:
    - Goal name
    - Goal purpose
    - Full /goal prompt
    - Expected outputs
    - Why Cowork is the right fit
    - Approval points
    - How to avoid unnecessary credit usage
    If it should stay in Microsoft 365 Copilot Chat, include:
    - Prompt name
    - Full prompt text
    - When to use it
    - When not to use it
    - Expected output
    - How to reuse it
    Step 7: Add one Power Automate bonus use case
    After the top 3 Copilot-only use cases, identify one additional use case where Power Automate is actually justified.
    This should be a workflow where:
    - A real trigger is needed
    - A system action needs to happen automatically
    - A record needs to be created or updated
    - An approval needs to be routed
    - A notification needs to be sent
    - A process needs reliability beyond a prompt
    - Copilot alone is not enough
    For the Power Automate use case, include:
    - Use case name
    - Why Copilot alone is not enough
    - Why Power Automate is justified
    - Trigger
    - Core flow steps
    - Inputs
    - Outputs
    - Approvals
    - Error handling
    - Where Copilot fits
    - Where Cowork fits
    - Governance notes
    - What should stay manual
    - What should be tested first
    Step 8: Compare the top 4
    Create a comparison table with:
    - Use case
    - Copilot-only or Power Automate
    - Best tool
    - Why this tool fits
    - Setup effort
    - Business value
    - Risk
    - Data sensitivity
    - User adoption likelihood
    - First test to run
    Step 9: Final output
    At the end, give me:
    - Executive summary
    - Work IQ pattern summary
    - Top 10 Copilot-only candidate list
    - Top 3 Copilot-only use cases
    - Why those 3 won
    - Full setup blueprint for each top 3 use case
    - Scheduled prompt drafts
    - Cowork skill drafts
    - Cowork /goal drafts
    - Copilot Chat prompt drafts
    - One Power Automate bonus use case
    - Comparison table
    - Rollout recommendations
    - Governance concerns
    - What should be tested first
    - What should not be automated yet
    - Assumptions
    - Data gaps
    - Confidence level
    Final instruction:
    Be strict.
    The first 3 use cases must be possible with Copilot only.
    Do not sneak Power Automate into the first 3.
    Only the 4th bonus use case can use Power Automate.
    I want the strongest practical Copilot-only use cases, not flashy demos.

    Why this prompt works?

    This prompt forces Cowork to stay practical.

    The first three use cases must be possible with Copilot only, so the output should focus on things like scheduled prompts, Cowork skills, /goal workflows, Copilot Chat prompts, and agents where available.

    Then the fourth use case gives you a clean escalation path for Power Automate when the process actually needs a trigger, approval, notification, record update, or reliable system action.

    That is the real value: start with Copilot, then only move to automation when the work pattern justifies it.