The skill that decides whether AI pays back in 2026 is not prompting. It is delegation. Four products shipped the same idea inside eight months. ChatGPT Work, Claude Cowork, Gemini Spark and Microsoft Scout all take an objective, work across your files and apps for a while, and hand back something finished. The prompt stopped being the unit of work. The task took over, and that changes what you should be teaching your team.
By Toni Dos Santos, Co-Founder, Spicy Advisory. Published 10 August 2026. Written for UK SMB and mid-market leaders deciding what to train their people on next.
Key Takeaways
- The unit of AI work has changed twice. 2023 was the prompt. 2024 to 2025 was the conversation. 2026 is the task. Any training programme still organised around prompt engineering is teaching a unit of work that has been retired.
- Four products, one design decision. ChatGPT Work (9 July 2026), Claude Cowork (January 2026), Gemini Spark (Google I/O 2026) and Microsoft Scout (Frontier preview) all accept a described outcome and return a finished artefact. None of them is selling a better prompt box.
- The replacement skill is delegation, which your managers already have. What they lack is the habit of writing it down, because they have spent their careers delegating verbally to people who could ask a follow-up question in the corridor.
- Anthropic already published the framework. The 4Ds of AI fluency (Delegation, Description, Discernment, Diligence) map onto the fields of a written brief almost line for line.
- Teach a Delegation Brief, not a prompt library. Six fields: objective, context, resources, constraints, expected output, verification. One page, reusable, tool-agnostic.
- The two fields nobody writes are the two that matter. Resources decides what counts as truth when your CRM and your spreadsheet disagree. Verification decides how you know the output is right and who signs it off.
- Draft is not send. Customer-facing and financial outputs need a named human reviewer. Under UK GDPR the accountability stays with you regardless of which vendor's agent produced the work.
Not sure which of your workflows are worth delegating to an agent?
Run the free AI diagnostic →Book an audit call10 minutes, no signup wall. A maturity score plus the three highest-value workflows in your business.
The unit of AI work has changed twice since 2023
Every wave of AI enablement has had a unit: the smallest thing you hand to the machine and get something useful back from. Get the unit wrong and your training is well delivered and useless.
In 2023 the unit was the prompt. One instruction in, one block of text out. Prompt libraries made sense, because the phrasing genuinely was the variable.
By 2024 the unit had become the conversation. Context windows grew, uploads arrived, custom instructions and memory landed, and the skill shifted from writing one good line to steering a thread over twenty minutes without losing the plot.
In 2026 the unit is the task. You hand over an outcome, the agent plans, works, browses, opens your files, asks you a question when it is stuck, and comes back with a spreadsheet you can actually send. That is not a bigger prompt. It is a different act.
| Era | Unit of work | What you supplied | What came back | Skill we trained |
|---|---|---|---|---|
| 2023 | The prompt | A phrasing | A block of text | Prompt engineering |
| 2024–25 | The conversation | Context and follow-ups | A steered draft | Context management |
| 2026 | The task | A brief and access | A finished deliverable | Delegation |
This is why so many UK teams tell us their AI training “went well” and changed nothing. The course was competent. It was aimed at 2023.
Four products, one design decision
The clearest evidence that the unit has moved is that four competitors independently arrived at the same product shape within eight months of each other.
| Product | Shipped | Where it runs | What it hands back |
|---|---|---|---|
| ChatGPT Work (OpenAI) | 9 July 2026 | Desktop app, web and mobile on eligible paid plans | Documents, spreadsheets, presentations, reports and Sites, plus scheduled and monitoring tasks |
| Claude Cowork (Anthropic) | January 2026 | Desktop, working on local files with permission | Multi-step deliverables from your own folders and connected tools |
| Gemini Spark (Google) | Google I/O 2026 | Always-on agent on Google Cloud VMs, Workspace and Chrome | Work that continues when your laptop is shut, briefed from Gmail, Docs and Calendar |
| Microsoft Scout | Frontier preview, 2026 | Windows 11 and macOS desktop, Microsoft 365 and local shell | Coordination work: meeting prep, follow-ups, recurring reports |
OpenAI has pushed ChatGPT Work steadily further from chat and towards longer-running work: research and analysis, actions across connected apps and files, finished documents, spreadsheets, presentations, reports and Sites, and tasks that run on a schedule or watch for a change. On 6 August 2026 OpenAI ran a small-business session on Work, framed explicitly around everyday business tasks and finished deliverables rather than model capability.
Anthropic got there first with Claude Cowork in January. Google's Gemini Spark pushed the same idea onto always-on cloud VMs, so the work continues when your devices are off. Microsoft Scout put it on the desktop with its own governed identity, and Copilot Cowork extends it into the Microsoft 365 estate.
Four vendors. Four go-to-market stories. One design decision: describe the outcome, hand over the resources, approve the risky steps, receive a finished thing.
Which is why prompt engineering training has stopped paying back
A prompt library is a phrasebook. It is genuinely useful when the unit of work is one sentence and one answer. It falls apart the moment the unit is a task, for three reasons we see repeatedly in UK companies.
Libraries encode phrasing, not judgement. Your best account director is not good because of the words she uses. She is good because she knows which three of the eleven numbers in the report matter to that client this quarter. A prompt template cannot carry that. A brief can.
Long prompts got worse, not better. The mega-prompt arms race produced instructions so overloaded that models started dropping half of them. We wrote about why mega-prompts now make your outputs worse when this first became obvious in our own delivery.
The tool moved the bottleneck. When an agent can spend forty minutes reading your Drive, the constraint is no longer how you phrase the request. It is whether you told it which folder is authoritative, what the output must look like, and how anyone will know if it is wrong.
We hit this ourselves. Our own production work moved from prompts to reusable methods, which is the story behind building fifty Claude Skills instead of a prompt library, and it is why our tool-agnostic sessions now start from workflows rather than tools. The prompt-literacy layer still matters for individual contributors, and we still teach the small set of prompt skills non-technical managers actually need. It is now the first hour of a programme, not the whole programme.
Anthropic already wrote the framework, and it aged unusually well
Anthropic's AI fluency work defines four competencies, published as the 4D framework in its free AI Fluency: Framework & Foundations course. It was written when chat was still the unit. It reads better now than it did then.
| The D | What it asks of a person | What breaks when it is missing |
|---|---|---|
| Delegation | Deciding whether, when and how to engage AI on this piece of work at all | People automate the wrong thing, usually the part they enjoyed |
| Description | Describing the goal well enough to produce useful behaviour and output | Generic deliverables, endless rewriting, blame aimed at the model |
| Discernment | Judging whether the output and the process behind it are any good | Confident fiction ships, because polish gets mistaken for accuracy |
| Diligence | Owning what you do with it: sources, disclosure, consequences | No audit trail, no named owner, a compliance problem waiting to be found |
Under agents, each D got heavier. Delegation now includes deciding what an agent may touch, not just what it may draft. Discernment means reviewing a twenty-page artefact rather than a paragraph, which is a genuinely harder job. Diligence stopped being a policy sentence and became an audit trail your ICO-facing paperwork depends on.
“Every UK team we work with can already delegate. They do it every day, to juniors, to agencies, to freelancers. What they have never had to do is write the brief down, because a human could always find them in the kitchen and ask.”
The Delegation Brief: six fields, one page
This is what we teach instead of a prompting framework. It is deliberately boring, because the point is that it survives contact with a busy Tuesday.
| Field | The question it answers | A weak answer looks like |
|---|---|---|
| Objective | What decision does this work have to serve? | “A report on last month” |
| Context | Who reads it, and what do they already know? | Nothing written, because “everyone knows” |
| Resources | Which files and systems count as truth, and which wins in a conflict? | “Use my Drive” |
| Constraints | Tone, length, house rules, and what it must not do | No mention of what is off limits |
| Expected output | Which file type, which sections, in what order? | “Something I can send” |
| Verification | How do we know it is right, and who signs it off? | Left blank, every time |
Here is one from a Manchester operations lead, unedited apart from the client name. It runs every Friday in ChatGPT Work.
Objective. Give the ops lead and the MD a decision-ready list of stalled work before Monday's stand-up, so we can chase or write off by Monday lunchtime.
Context. Audience is two people who already know the accounts. They do not need background, they need exceptions and a recommendation. Last week they complained the list was too long to act on.
Resources. The CRM export in /Ops/Weekly and the #ops-digest Slack channel for the last seven days. If the CRM and Slack disagree on stage, the CRM wins and you flag the disagreement. Nothing else is a source.
Constraints. British English. Maximum twenty rows. No commentary on people, only on work items. Do not contact anyone.
Expected output. A spreadsheet with columns: item, owner, days stalled, last movement, recommended action, confidence. Plus a memo of under 200 words naming the three items that matter most and why.
Verification. Every row must cite the source record. If a deal has no stage history, mark it “unknown” rather than inferring. If you are unsure about scope, period or audience, ask me one question before you start. I review and sign before anything is shared.
Two lines in that brief do most of the work, and they are the two people leave out.
The first is the conflict rule: if the CRM and Slack disagree, the CRM wins and you flag it. Without it, an agent quietly averages two versions of reality and produces a number that exists nowhere.
The second is the question rule: ask me one question before you start. That single sentence converts an agent from something that guesses into something that checks, and it costs you thirty seconds on a Friday morning.
Store the brief where the work lives, not in a chat. Put it in the ChatGPT Project, the Claude Project, or the shared workspace instruction field, so it survives when the thread scrolls away and when the person who wrote it is on holiday. A brief that lives in someone's message history is not a process, it is a habit that leaves with them.
Want your team writing briefs like this by the end of the month?
Book a 30-minute audit call →See our UK SMB programmesWe work with UK companies from 10 to 500 people. Tool-agnostic, workflow-first, measured on hours to an approved deliverable.
Ten delegations UK companies are running in ChatGPT Work right now
These are patterns we see working in UK SMBs and mid-market firms. Each one is written as a delegation, not a prompt, and each one names the human gate. Compressed here for space, the real briefs use all six fields.
1. The Monday client performance pack
Who. Independent agencies of six to thirty people, where an account lead loses most of Monday to formatting.
Delegation. Objective: give the account lead a client-ready pack that supports next week's spend decision. Resources: last week's metrics export, the brand kit, and the previous deck as the format of record. Output: an editable presentation with wins, misses and three recommendations, plus a flag on any metric with no source cell. Verification: every figure traces to a cell in the export, and unsourced claims are marked rather than smoothed over.
Human gate. The account lead reviews before it reaches the client. The agent never emails anyone.
What changes. The reformatting hours disappear. The judgement hours stay, which is the point.
2. Month-end management accounts commentary
Who. Finance leads in mid-market firms who write the same narrative twelve times a year.
Delegation. Objective: explain the variance to a board that will ask about margin. Resources: the trial balance export and the prior three months, in sterling, financial year ending 31 March. Constraints: no invented figures, flag anything not in source. Output: a one-page commentary and a variance table over 5%.
Human gate. The finance director owns every number quoted externally. See our finance team workflows for the fuller version.
What changes. The first draft arrives before the meeting rather than during it.
3. PQQ and tender first drafts
Who. UK professional services firms bidding for public sector frameworks, where the bid team is two people and a deadline.
Delegation. Objective: produce a first-pass response mapped question by question to the tender document. Resources: the ITT pack, the last three winning submissions, the accreditations folder. Constraints: never claim an accreditation not in the folder. Output: a response document with a coverage table showing which questions have evidence and which have gaps.
Human gate. A bid lead rewrites every claim about capability and price. Nothing about certifications ships unchecked.
What changes. The gap list arrives on day one instead of day nine. More on this pattern in how UK professional services firms are using AI to win more work.
4. The proposal from a discovery call
Who. B2B services founders and small sales teams.
Delegation. Objective: turn call notes into a priced proposal in the house template. Resources: the call transcript, the rate card, the last signed proposal as the shape of record. Constraints: scope, assumptions and exclusions must each be explicit. Output: the template, filled, with anything uncertain left blank and listed.
Human gate. The founder writes pricing and legal language personally. Always.
What changes. Proposals go out the same day, which is usually worth more than the hours saved.
5. Competitor and price monitoring, on a schedule
Who. UK retail, DTC and B2B SaaS teams who currently do this in a burst every six months.
Delegation. Objective: tell the commercial lead what changed this week that affects our pricing or positioning. Resources: a named list of competitor URLs and the pricing sheet. Constraints: report changes only, no speculation about intent. Output: a short digest, plus a spreadsheet row appended per change.
Human gate. Nobody changes a price off the digest alone.
What changes. This is where ChatGPT Work's scheduled and monitoring tasks earn their keep, because the work is small, recurring and easy to forget.
6. The supplier invoice exception report
Who. Operations and finance in businesses processing a few hundred invoices a month.
Delegation. Objective: surface invoices that do not match a PO or a delivery note before payment run. Resources: the invoice export, the PO ledger. Constraints: never approve, never pay, never message a supplier. Output: an exception spreadsheet with a reason code per row.
Human gate. Payment approval stays exactly where it is today.
What changes. The exceptions get found before the money leaves. See operations workflows for the wider set.
7. The policy refresh when the rules change
Who. HR and compliance leads in regulated or data-heavy UK businesses.
Delegation. Objective: identify which clauses of our policies are affected by a specific regulatory change and draft replacement wording. Resources: the current policy set, the text of the change. Constraints: quote the source clause for every proposed edit. Output: a redline table of clause, issue, proposed wording, source.
Human gate. Legal signs. The agent produces the shortlist, not the decision.
What changes. The scan takes an afternoon instead of a fortnight. Useful when you are working through the Data (Use and Access) Act or, for anyone trading into the EU, Article 4 AI literacy obligations.
8. Support macros and knowledge base refresh
Who. Customer support teams of four to forty.
Delegation. Objective: find the ten questions we answered most last quarter that have no knowledge base article, and draft them. Resources: the ticket export and the current help centre. Constraints: house tone, no promises about SLAs or refunds. Output: ten draft articles plus a list of macros that contradict current policy.
Human gate. Support lead approves before publication. More patterns in AI workflows for customer support.
What changes. The backlog that never gets prioritised finally gets a draft.
9. The board pre-read
Who. Mid-market companies of 100 to 500 people with a board or investor cadence.
Delegation. Objective: give non-executives what they need to arrive with questions rather than requests for data. Resources: the KPI sheet, last board minutes, the department updates. Constraints: five pages maximum, no adjectives on performance without a number attached. Output: the pack, plus a page listing what materially changed since last time.
Human gate. The MD reads and edits. Nothing goes to a board unread.
What changes. The pack goes out five working days early because producing it stopped being a two-day job.
10. The campaign kit from a brief
Who. In-house marketing teams of two to fifteen.
Delegation. Objective: take an approved campaign brief and produce everything needed to launch. Resources: brief, brand kit, last campaign's performance export, ASA guidance on disclosure. Constraints: no claims that are not evidenced in the brief. Output: landing copy, five emails, fifteen social variants, and a measurement plan naming one primary metric.
Human gate. Marketing lead approves claims and disclosure before anything is scheduled.
What changes. The kit lands in hours. The strategy conversation, which is the valuable part, gets the week.
| Start here if you are… | Delegation | Time to first useful output | Risk if unsupervised |
|---|---|---|---|
| An agency | Monday performance pack | One week | Medium: client-facing |
| A mid-market finance team | Month-end commentary | One cycle | High: numbers |
| A services firm that bids | Tender first draft | One tender | High: claims and accreditations |
| An ops-heavy business | Invoice exceptions | Two weeks | Low: internal only |
| A marketing team | Campaign kit | One campaign | Medium: advertising rules |
Five ways delegation goes wrong
1. The verification field is blank. This is the most common failure and the most expensive. A polished twenty-page artefact with one invented figure is worse than a rough draft, because nobody checks the polished one. If you cannot describe how you would catch an error, you are not ready to delegate that task.
2. Every connector switched on in week one. Each integration is blast radius. Connect the two systems this month's delegation actually needs, review permissions quarterly, and offboard people from the agent and the drive on the same day. Our governance framework for mid-market companies covers the controls that matter without a six-month policy project.
3. Delegating the judgement instead of the work. The failure mode is subtle: people hand over the interesting 20% and keep the formatting. Reverse it. The agent assembles, the human decides.
4. No gate on outbound actions. Draft is not send. Draft is not publish. Draft is not pay. Write that into the project instructions in plain words, because a general instruction to “be careful” is not a control.
5. Using an agent where a trigger would do. If the requirement is “every time a form is submitted, create the record and notify the channel”, that is not a delegation, it is automation. Agents improvise. For deterministic routing, improvisation is a defect.
These are the same failure patterns behind most stalled programmes, which we unpacked in why AI adoption fails in companies.
How we have always approached this at Spicy Advisory
We have never run a “ChatGPT course”. Not because we are precious about it, but because the tool changes every quarter and the workflow does not. Every programme we deliver starts with a map of how work actually moves through your business, and only then asks which model, licence or agent belongs at which step.
That is why this shift did not require us to rewrite anything. When your unit of analysis was always the task, an industry moving from prompts to tasks is a tailwind rather than a rebuild. What changed is that the thing we were already teaching now has product support behind it.
| Week | What happens | What you leave with |
|---|---|---|
| 1 | Workflow mapping with the people who do the work, not just the leadership team | A ranked list of candidate delegations, scored on hours, risk and repeatability |
| 2 | Brief writing clinic: your managers write real briefs for their own real work | Three to five delegation briefs your team wrote themselves |
| 3 | Supervised runs in your own tools, with the failures used as teaching material | Working deliverables plus a documented gate for each one |
| 4 | Governance, measurement and handover to a named internal owner | A baseline, a target, and a review cadence that survives us leaving |
We measure one thing above all others: hours from trigger to an approved deliverable. Not licences bought, not prompts written, not attendance. If that number does not move, the programme did not work, and we would rather find that out in week three. The full method behind that sits in AI training that sticks and how to measure AI training ROI in a UK business.
Why UK SMBs and mid-market teams work with us
We are a UK and EU AI adoption partner built for companies between 10 and 500 people: too small for a six-figure transformation programme, too big to leave everyone experimenting on personal accounts. We are tool-agnostic across ChatGPT, Claude, Gemini and Copilot, we start from your workflows rather than a vendor's roadmap, and we hand over an internal owner rather than a dependency on us.
Explore our AI adoption programmes for UK SMBs →Prefer to start with evidence? Take the free AI maturity audit → or book a 30-minute audit call →
Which agent for which delegation
The brief is portable. The tool is not the strategy. That said, the four are not interchangeable, and the buying decision for a UK company usually comes down to where your source of truth already lives.
| Choose | When | Check before you buy |
|---|---|---|
| ChatGPT Work | The delegation spans many non-Google, non-Microsoft tools and ends in a finished document, sheet or deck | Which plan unlocks Work on which surface, and how usage is counted on long runs |
| Claude Cowork | The work is desktop and file-heavy, and you want approval gates built into the flow | Commercial plan terms, so your work stays out of model training |
| Gemini Spark | Your source of truth is Gmail, Docs, Sheets and Calendar, and you want work continuing off-device | Agent data scopes per role, and what “acting on your behalf” is allowed to include |
| Microsoft Scout | You are a Microsoft 365 shop and the pain is coordination rather than creation | Frontier enrolment, Intune-managed devices, and Copilot licensing |
If you are still choosing platforms rather than delegations, start with the platform comparison for companies, then come back to the brief. And if two or three of these are already in use across your business without anyone deciding, that is worth knowing before you buy a fourth: AI spend visibility tends to be the fastest cost saving we find.
What to do in the next fortnight
Five steps, in order, and none of them require a budget approval.
One. Pick one recurring deliverable that someone dreads. Weekly, monthly, always late, always reformatted.
Two. Write the six-field brief for it. On one page. The person who does the work writes it, not the person who commissions it.
Three. Run it in whichever agent you already pay for. Do not buy anything yet.
Four. Time it. Trigger to approved deliverable, before and after, honestly.
Five. If it worked three times running, move the brief into a shared project with a named owner and a schedule. Then pick the next one.
Teams that do this find their AI budget stops being a licence line and starts being a workflow line. Which is the only version of this that survives a finance review. If you want the ranked list of which workflows to start with, that is exactly what our free AI diagnostic produces, and what we work through on an audit call.
Move your team from prompting to delegating, in four weeks
Book a 30-minute audit call →See UK training optionsTool-agnostic, workflow-first, delivered across the UK and remotely.
Frequently Asked Questions
Is prompt engineering dead in 2026?
Not dead, but demoted. Prompt skills still matter for short, individual tasks and for writing the description part of a brief well. What has ended is prompt engineering as the organising principle of corporate AI training. When the unit of work is a task rather than a sentence, the decisive skills are choosing what to delegate, specifying it in writing, judging the output, and owning the consequences. Prompting is now roughly the first hour of a programme rather than the whole syllabus.
What is a delegation brief for AI?
A delegation brief is a one-page written specification you hand to an AI agent instead of a prompt. It has six fields: objective (the decision the work serves), context (who reads it and what they know), resources (which files and systems count as truth, and which wins in a conflict), constraints (tone, length, and what it must not do), expected output (file type and sections), and verification (how anyone knows it is right, and who signs it off). It is tool-agnostic, so the same brief runs in ChatGPT Work, Claude Cowork, Gemini Spark or Microsoft Scout.
What is Anthropic's 4D framework?
The 4Ds are the four competencies in Anthropic's AI fluency work: Delegation (deciding whether, when and how to engage AI), Description (describing goals well enough to get useful behaviour), Discernment (accurately assessing outputs and the process behind them), and Diligence (taking responsibility for what you do with AI and how). Anthropic teaches them in its free AI Fluency: Framework & Foundations course. They map almost line for line onto the fields of a written delegation brief, which is why the framework has aged well into the agent era.
What can ChatGPT Work actually do for a small business?
It takes longer-running jobs across your connected apps and files and returns finished artefacts: documents, spreadsheets, presentations, reports and Sites, and it can run tasks on a schedule or monitor for changes. For UK SMBs the strongest early uses are recurring deliverables that currently eat a person's day: the Monday client pack, month-end commentary, tender first drafts, proposals from call notes, competitor monitoring, invoice exception reports, policy redlines, support content refreshes, board pre-reads and campaign kits. Every one of them needs a named human reviewing before anything leaves the business.
How is ChatGPT Work different from Claude Cowork, Gemini Spark and Microsoft Scout?
They share the same shape and differ on where they live. ChatGPT Work is strongest when the job spans many tools and ends in a finished document, sheet or deck. Claude Cowork is desktop and file-centric with approval gates in the flow. Gemini Spark is an always-on agent that keeps working when your devices are off and is deepest in Gmail, Docs and Calendar. Microsoft Scout is a desktop agent for Microsoft 365 estates, aimed at coordination work, and it is still an experimental Frontier preview requiring Intune-managed devices and Copilot Business or Enterprise licensing.
Do we still need Zapier, Make or n8n if we have an agent?
Yes, for anything deterministic. Agents plan and improvise, which is what you want for judgement work and exactly what you do not want for “every time a form is submitted, create the record and notify the channel”. Use automation platforms for reliable triggers and routing, and use agents for the assembly and analysis that sits between those triggers. The two are complements, and replacing a working automation with an agent is one of the more expensive mistakes we see.
What are the UK compliance risks of delegating work to an AI agent?
Three worth naming. First, standing access: an agent connected to email and file storage widens your breach surface, and under UK GDPR accountability stays with your organisation regardless of vendor. Second, verification: a confident, well-formatted output containing an invented figure is a real risk in regulated reporting, so the verification field is a control rather than a nicety. Third, outbound actions: sending, publishing and paying need explicit human gates and an audit trail. Scope connectors tightly, name a reviewer per workflow, and review permissions quarterly.
How long does it take to move a team from prompting to delegating?
About four weeks for the first workflows, in our experience with UK SMB and mid-market teams. Week one maps workflows and ranks candidates by hours, risk and repeatability. Week two is a brief-writing clinic on real work. Week three runs the delegations supervised in your own tools. Week four covers governance, measurement and handover to a named internal owner. The metric that matters is hours from trigger to approved deliverable, measured before and after. If that number has not moved by week four, something in the design was wrong.