Tag: claude-code

  • Copilot Cowork Model Benchmark: Five Models, One Prompt, One Winner

    Copilot Cowork Model Benchmark: Five Models, One Prompt, One Winner

    I put five Copilot Cowork models through the same executive decision stress test: one locked prompt, the same six-file evidence pack, and a fresh session every time.

    The job was not to summarize the packet. Each model had to reconcile conflicting numbers, stale executive claims, security blockers, contract gates, and future retest dates—then choose the correct launch path and turn it into a memo, PowerPoint, and Excel action tracker.

    All five reached the expected decision. The real separation came after that: evidence discipline, unsupported assumptions, file quality, and a ninefold spread in Cowork credits. Claude Opus 5 delivered the strongest overall result while using the fewest credits.

    Decision accuracy came first. Polish only counted after the answer was right.

    The result in 30 seconds

    Every model chose the expected Option B: approve a conditional seven-site phased launch. Opus was the only run to pair a perfect quality score with the top design score and the lowest credit total.

    ModelQualityDesignCredits
    Claude Opus 5100 / 1009.5 / 10294
    Claude Sonnet 598 / 1008.9 / 10324
    GPT-5.6 Sol99 / 1008.4 / 10771
    Claude Fable 599 / 1009.4 / 102,658
    GPT-5.6 Terra97 / 1007.5 / 10360
    All five selected Option B. Quality used the locked 100-point rubric. Design was scored separately. Credits are the task totals shown by Cowork’s /cost response; Elapsed time was not scored.

    In this benchmark

    The executive decision stress test

    I built one fictional but realistic executive decision scenario, created a hidden answer key, and ran the exact same job in five fresh Copilot Cowork sessions. The selected model was the variable.

    • Same input: one prompt and the same six files for every model.
    • Fresh sessions: no previous conversation or model output could influence the next run.
    • No rescue prompt: no web access, extra explanation, or follow-up designed to repair a weak answer.
    • Outcome first: decision accuracy was checked before design and presentation quality.

    The fictional company, Kavora Industries, was preparing a steering committee decision for a Hartland Energy software launch. The evidence pack deliberately mixed current data with stale summaries, future retest dates, security blockers, contractual gates, and conflicting executive claims.

    Every session received:

    • An executive context document
    • An Excel launch-readiness workbook
    • A security and contract brief
    • An executive email thread with conflicting claims
    • A fictional company brand guide
    • A board presentation template

    The model had to reconcile the evidence, calculate the real readiness position, choose one of three launch paths, and create three usable business files: a two-page decision memo, a steering committee deck of six slides or fewer, and an Excel action tracker.

    The answer key

    I locked the answer key before reviewing any model output.

    Expected decision: approve Option B—a conditional seven-site phased launch on August 17. Both Priority 1 defects had to pass retest with recorded evidence, the final seven-site list had to be confirmed, and the Hartland sponsor had to provide written approval naming the phase and facilities. If any gate failed, launch activity moved to August 31.

    Two fictional P1 defects controlled the decision:

    • SEC-014: a contractor supervisor could access another company’s attachment through a predictable file URL.
    • AUTH-022: a legacy SSO fallback could bypass the enforced MFA path during a retry condition.
    Open Defect Register showing SEC-014 and AUTH-022 as Priority 1 defects with plain-language descriptions
    Source evidence: The readiness workbook classified SEC-014 and AUTH-022 as Priority 1 defects.

    Both defects were still in progress. The listed fix dates were targets, the retests had not happened, and neither issue met the documented closure rule.

    Defect register showing affected sites, in-progress status, target fix dates, retest dates, workarounds, and launch blocker flags
    Source evidence: Future fix and retest dates did not mean the P1 defects were already closed.

    The contract allowed an initial phase of at least seven facilities, but only after every acceptance gate was met. The security gate required zero open P1 defects and successful retest evidence.

    Security and contract brief showing the four acceptance gates for a seven-site production phase
    Source evidence: Section 6.1 allowed a seven-site phase and made security evidence and sponsor approval hard gates.

    Option A—the immediate ten-site launch—was ineligible. Sites 8 through 10 were not ready, the two P1 defects remained open, and the Hartland sponsor had rejected a ten-site approval on the current evidence. Option C, delaying everything to August 31, was the fallback if Option B’s gates failed.

    Security release policy showing the Priority 1 closure standard, security controls, decision timeline, and fallback rule
    Source evidence: A P1 stayed open until its fix was deployed, the prescribed retest passed, and the evidence was recorded.

    The six traps

    The packet was built to punish shallow summarization. A model could extract plenty of correct facts and still answer the wrong question.

    1. 91.9% looked like a failed UAT gate. The contract actually measured three core flows at 535 of 555 tests, or 96.4%. The model had to report both numbers and explain which one governed.
    2. The packet contained two defect counts. An older PMO summary said 12 defects were open. The newer readiness workbook said 19: 2 P1, 6 P2, and 11 P3.
    3. An executive claimed all ten sites were approved. A later message from Hartland’s sponsor said the opposite. Authority and timestamp made the sponsor’s conditional seven-site position the governing evidence.
    4. A target fix date did not close a security defect. Closure required deployment, a passed retest, and recorded evidence—not a workaround, target date, or code-complete status.
    5. The reported CAD $730,000 delay cost mixed two different things. It combined CAD $250,000 in direct cost with a CAD $480,000 milestone invoice. The invoice was delayed cash timing, not lost revenue.
    6. The decision date was easy to move accidentally. The committee was deciding on August 6. Retests on August 12 and 13 controlled execution on August 17; they were not missing inputs for the August 6 recommendation.

    My Prompt

    I used the same prompt on all the models, the prompt:

    You are the engagement lead at Kavora Industries. The executive steering committee for the Hartland Energy Kavora Smart Operations program meets in 45 minutes. Review all six attached files and prepare an evidence-based launch recommendation.
    Your job is to determine whether Kavora should: (A) launch all 10 sites on August 17, 2026; (B) launch a conditional seven-site phase on August 17; or (C) delay all sites to August 31.
    Working rules:
    1. Use only the attached files. Do not browse the web.
    2. Do not ask clarifying questions. Make conservative assumptions and flag unknowns.
    3. Where sources conflict, use the most recent dated evidence and state the conflict.
    4. Cite the supporting file name and sheet, section, or message for every material conclusion.
    5. Recalculate key percentages and financial totals. Distinguish direct cost, cash timing, and contingency.
    6. Do not invent facts, approvals, dates, owners, or commitments.
    7. Do not send email, publish, post, or take any external action. Draft only.
    8. Follow the attached Kavora brand guide and use the attached board template as the visual basis for the deck.
    Create these three downloadable files with these exact names:
    - Kavora_Hartland_Launch_Decision_Memo.docx — maximum two pages. Lead with the recommendation and fallback; include evidence, financial treatment, conditions, risks, and the exact steering decision requested.
    - Kavora_Hartland_Steering_Committee.pptx — maximum six slides. Make it executive-ready, visual, concise, and branded.
    - Kavora_Hartland_Launch_Action_Tracker.xlsx — include action, owner, due date, gate, status, dependency, evidence location, and escalation path; add useful formulas or validation where appropriate.
    When the files are complete, reply in chat with a compact draft executive email of no more than 180 words. It must state the recommendation, hard conditions, fallback, and decision required. Do not send it.​‌
    [01_COWORK_INPUT_FILES]

    I have also included the input files I used for this session

    How each model performed

    Claude Opus 5: the winner

    Quality 100/100 · Design 9.5/10 · Credits 294

    Opus was the cleanest run. It reached the expected decision, reconciled every material conflict, used precise citations,

    separated direct cost from cash timing, and built the strongest action tracker. It also used the fewest credits.

    Opus steering deck control-plan slide with conditions, owners, dates, and status
    Deck receipt: A compact control plan with the conditions, owners, dates, and status visible.
    Opus action tracker gate-status sheet showing thresholds, owners, evidence, status, and linked actions
    Tracker receipt: The gate sheet makes the thresholds, owners, evidence, and linked actions easy to audit.
    Cowork cost receipt showing 294 credits for Claude Opus 5
    Cost receipt: 294 credits used for the task.

    Claude Sonnet 5: close and efficient

    Quality 98/100 · Design 8.9/10 · Credits 324 so far

    Sonnet’s decision and math were right, and the files were polished. I deducted two points because it proposed delegating the no-go call to the Program Director and Security Lead without evidence that they already held that authority.

    Sonnet steering deck control-plan slide with four gates and an August 14 checkpoint
    Deck receipt: Polished, clear, and executive-friendly.
    Sonnet memo excerpt delegating the no-go call to the Program Director and Security Lead
    Deduction receipt: The memo assigns decision authority the source pack never established.
    Cowork cost receipt showing 324 credits so far for Claude Sonnet 5
    Cost receipt: 324 credits “so far” when captured.

    GPT-5.6 Sol: accurate, but more expensive

    Quality 99/100 · Design 8.4/10 · Credits 771

    Sol showed strong evidence discipline and avoided inventing facts. It lost one point because the go/no-go checkpoint owner remained unresolved in the tracker. The larger difference was cost: 771 credits, more than 2.6 times the Opus run.

    Sol steering deck control-plan slide listing four dated gate controls
    Deck receipt: A clear six-slide story that uses the supplied template well.
    Sol action tracker excerpt showing the go or no-go checkpoint owner as not specified in source
    Deduction receipt: Sol flags the missing owner, but the operational tracker leaves the accountability unresolved.
    Cowork cost receipt showing 771 credits for GPT-5.6 Sol
    Cost receipt: 771 credits used for the task.

    Claude Fable 5: polished, but costly

    Quality 99/100 · Design 9.4/10 · Credits 2,658

    Fable reached the expected answer and produced some of the sharpest-looking files. I deducted one point because it inserted my name as the engagement lead even though the prompt and source pack never provided it. The run used just over nine times the Opus total without improving the decision.

    Fable steering deck control-plan slide with clear owners, dates, and status
    Deck receipt: One of the sharpest-looking decks in the set.
    Fable memo header showing Josh Cook as engagement lead without source support
    Deduction receipt: The memo invents an author identity that was absent from the prompt and packet.
    Cowork cost receipt showing 2,658 credits for Claude Fable 5
    Cost receipt: 2,658 credits used for the task.

    GPT-5.6 Terra: right answer, visible QA misses

    Quality 97/100 · Design 7.5/10 · Credits 360

    Terra handled the scenario date correctly, reconciled the stale claims, calculated the readiness and financial figures, and created all three files. Its PowerPoint still shipped with a literal \n on slide 1, text collisions on slide 6, and a low-contrast source line.

    Terra steering deck slide 1 showing a literal backslash-n in the decision card
    Slide 1 receipt: The decision card exposes a literal \n.
    Terra steering deck slide 6 showing overlapping text and low-contrast source text
    Slide 6 receipt: Both cards contain text collisions, and the source line is difficult to read.
    Cowork cost receipt showing 360 credits for GPT-5.6 Terra
    Cost receipt: 360 credits used for the task.

    Design review

    I kept design separate from decision accuracy. These are my visual scores across hierarchy, readability, template fidelity, spacing, consistency, QA, and day-to-day usability.

    ModelMemoDeckTrackerOverall
    Claude Opus 59.49.59.69.5
    Claude Fable 59.39.59.49.4
    Claude Sonnet 59.09.28.58.9
    GPT-5.6 Sol8.29.08.08.4
    GPT-5.6 Terra8.26.57.87.5
    A 10 means I would present the file without visual cleanup. An 8 means it is solid with a few practical issues. A 6 means visible mistakes survived into the deliverable. Overall is the simple average of the three artifact scores.

    Opus had the best balance. Fable was nearly as polished; its unsupported author name affected factual quality, not its design score. Sonnet delivered a strong memo and deck. Sol’s files were accurate but more functional than finished. Terra’s memo and tracker were usable, but the deck needed another QA pass.

    Where Terra lost design points

    • Decision memo: clean two-page render, but noisy raw-filename citations and inconsistent typography.
    • Steering deck: strong template match and clear story, but the literal \n, overlapping text, and low-contrast source line were visible QA failures.
    • Action tracker: complete and usable, but extremely wide, awkward to navigate, and unfinished for printing.
    Two-page Terra decision memo rendered side by side for design review
    Memo check: Both pages render cleanly. The remaining issues are hierarchy, citation noise, and typography consistency.
    Wide Terra action tracker with owner, due date, gate, status, dependency, evidence, escalation, and date-offset columns
    Tracker check: The content is strong. Its width and frozen-pane choice make horizontal navigation harder than it needs to be.

    Methodology and limits

    This was a controlled practical benchmark, not a permanent model leaderboard.

    • One final submission per model was evaluated.
    • Every run used the same prompt, six files, and output requirements.
    • No corrective follow-up prompt was used.
    • Decision correctness was checked before presentation quality.
    • Credits were recorded separately from quality and design.
    • Elapsed time was excluded.

    A different task could produce a different order. Coding, research, financial analysis, image work, and browser automation stress different capabilities. The series will keep the controls stable while changing the work.

    Final take

    Claude Opus 5 won this benchmark.

    All five models found the expected launch path. Opus separated itself by producing the best supporting work, the strongest design score, and the lowest credit total. Sonnet stayed efficient. Sol and Fable produced high-quality work at much higher task costs. Terra was accurate but needed a better presentation QA pass.

    Give the models real work. Lock the answer key. Keep the receipts.