Back

Prompt Improver for Teams: How to Cut AI Token Spend

Most budget holders think the long, detailed prompt is the expensive one. Per finished task it is the cheap one - here is where team token spend actually goes, and how a prompt improver brings it down.

Date

Reading time

9

min

Amelia Miller

Co-founder and CEO

A prompt improver is a tool that rewrites a prompt before it reaches the AI model, adding the context, constraints and output format the model needs to answer properly first time. For a team, it is also the cheapest lever on the AI bill that most budget holders have not pulled yet.

To reduce AI token costs across a team, cut the back-and-forth, not the prompt. Under-specified prompts trigger repair turns, and every repair turn resends the whole conversation. Give the model full context up front, send routine work to a lighter model, and share the prompts that already work, so nobody pays to rediscover them.

That runs against instinct. Most people assume the long, detailed prompt is the expensive one, and per message they are right. This guide covers where team token spend actually goes, how much model choice moves it, and what a prompt improver like the ivee AI Learning Companion changes when it sits inside ChatGPT, Claude, Gemini and Copilot for a whole team.

Why do under-specified prompts cost more?

Under-specified prompts cost more because they rarely land first time, and the fix is paid for in extra turns. Each follow-up message carries the entire conversation back to the model as input, so a four-turn repair job bills the opening prompt four times over, plus every answer you did not want.

ivee's August 2026 research put a number on it: a thin, under-specified prompt costs 14 times more on average than a well-loaded one, because of the back-and-forth needed to repair it. Asked to estimate that multiplier themselves, decision makers gave answers ranging from 2x to 100x. That spread is the problem. The people signing off the budget are guessing, and the 14x cost of an under-specified prompt is the figure that replaces the guess.

The same group of 500 UK AI decision makers was asked live which costs more money, a thin prompt or a loaded, context-rich prompt. They split 80/20, and the 80% picked the loaded prompt, which is the wrong answer. The 80% are right about tokens and wrong about the unit of account. Per call, a loaded prompt costs more, because every input token is billed. Per finished piece of work, the thin prompt costs more, because it does not land.

Output makes it worse. Anthropic's published API pricing bills output tokens at five times the input rate on its current models: Claude Sonnet 5.5 is $2 per million input tokens and $10 per million output. A vague prompt invites a long, generic answer, and the long, generic answer is the expensive part.

How much does model choice change the AI bill?

Model choice changes the bill by up to two orders of magnitude, which dwarfs anything you can save by trimming words. On Anthropic's published rates, Claude Fable 5.1 costs $10 per million input tokens and $50 per million output. Claude Haiku 5.5 costs $0.10 and $0.50 for prompts under 100,000 tokens. That is a hundredfold gap for the same number of tokens.

Model

Input, per million tokens

Output, per million tokens

Suited to

Claude Haiku 5.5

$0.10

$0.50

Summaries, extraction, quick answers

Claude Sonnet 5.5

$2

$10

Writing, analysis, everyday work

Claude Opus 5.5

$4

$20

Deep research and complex reasoning

Claude Fable 5.1

$10

$50

Long, complex projects with few check-ins

Prices are Anthropic API list prices in US dollars, checked October 2026. Haiku 5.5 rises to $0.50 and $2.50 above 100,000 tokens.

Seat-based plans hide this, but they do not remove it. Anthropic's own guide to choosing a Claude model says each model draws on your usage limit at a different rate, and that using Opus or Fable on a task Sonnet or Haiku could handle may use more of your limit unnecessarily. On a team plan, that shows up as people hitting their cap on a Tuesday afternoon and asking for an upgrade.

The difficulty is that nobody wants to look up a pricing table before every prompt. One AI decision maker in ivee's research put it plainly: "How easy is it to get used to which models to use for which tasks?" For most people the answer is never, so the default model does everything.

What does a prompt improver actually do?

A prompt improver takes the prompt you were about to send, rewrites it with the missing context and structure, and hands it back before anything reaches the model. The good ones explain the changes, so the person typing learns what a strong prompt looks like instead of depending on the tool forever.

That matches what the model makers recommend. Anthropic's prompting best practices describe Claude as "a brilliant but new employee who lacks context on your norms and workflows", and add that the more precisely you explain what you want, the better the result. A prompt improver does that explaining on the user's behalf, every time.

The ivee AI Learning Companion is a Mac app that runs alongside the AI tools a team already uses. It does four jobs:

  • Rewrites the prompt in one shortcut. Type as normal in ChatGPT, Claude, Gemini or Copilot, press ⇧+⌘+Space, and ivee sharpens the prompt and shows why it changed.

  • Suggests the right model. ivee flags when a cheaper or faster model would handle the job, which is where the hundredfold price gap gets closed.

  • Follows the whole conversation. ivee keeps the chat on track as it goes, so fewer threads drift into five turns of repair.

  • Spots what you could automate. ivee learns from repeated requests and suggests automation builds, with the steps to get there.

A global context setting sits underneath all four. You state your role, your organisation and how you like answers once, and ivee carries that into every prompt instead of relying on people to retype it.

How do you roll a prompt improver out across a team?

Roll a prompt improver out by starting with the people who prompt least, sharing what works, and setting team context once. The savings sit unevenly across a team, and the rollout should follow them.

  1. Start with the beginners. New starters and occasional users write the thinnest prompts and pay the 14x most often. Most of them will not read a guide to writing prompts that actually work, but they will press a shortcut. Seeing their own prompt rewritten, with the reasons, builds the habit faster than a workshop slide.

  2. Give power users their time back. Heavy users already prompt well. Their gain is speed: dictating a rough thought and getting a structured prompt back, and a nudge off Opus or Fable when the task does not need it.

  3. Share the prompts that work. On ivee's Team plan, prompt libraries are shared. When the head of operations writes the client update prompt that works, it becomes everyone's starting point, which stops ten people paying to reinvent it.

  4. Set global team context. Team plans carry one shared context, so every prompt starts with the organisation's tone, clients and constraints already in place.

Typical effort is small. The app installs per person, the shortcut works inside the tools already licensed, and there is no new interface to train anyone on. ivee is on the Mac now, with a Windows waitlist open.

How do you measure whether token spend is coming down?

Measure it per finished piece of work, not per message. Track how many turns a task takes, which models people reach for, and whether usage caps are being hit less often. A falling turn count with a stable or rising volume of work is the signal that spend is coming down for the right reason.

Most organisations cannot do this today. In the same survey, 4% of AI decision makers could prove the ROI of their organisation's AI spend, tracking it and showing the numbers, according to ivee's research with 500 UK AI decision makers. The sample was drawn from organisations already building with AI, so the wider picture is unlikely to be better.

ivee's Team plan adds an admin dashboard and individual usage reporting, so a budget holder can see who is using AI, how well, and where tokens and time are being saved. For the bigger number, the free AI training ROI calculator turns hours saved per person per week, team size and industry into an annual figure you can put in front of a finance director.

What happens to your prompts?

Your prompts pass through ivee to be rewritten, so the data terms matter before a team rollout. ivee runs zero data retention with Anthropic, which means the model provider does not keep your prompts after processing them. Team plans add on-device redaction of sensitive information and team-wide admin controls.

The free Basic plan works differently: prompts on Basic may be used to improve ivee. If the prompts involve client names, financials or personal data, that is the reason to put a team on a paid plan rather than letting people install the free tier one by one. The full plan-by-plan breakdown sits on ivee's pricing page.

Frequently asked questions

Is a token the same as a Claude usage credit?

No. A token is a chunk of text, roughly four characters in English, and it is what API usage is billed on. Seat-based chat plans convert token use into a usage limit, so you do not see a per-token bill, but heavier models and longer conversations still use the limit up faster.

Is token cost on top of the per-user subscription?

On seat-based chat plans, usually not: the subscription covers usage up to a limit. Token costs appear separately when a team builds on the API directly, for automations or internal tools. Both are worth tracking, because the habits that burn through a seat limit are the same ones that inflate an API bill.

Is it cheaper to use smaller prompts on a lighter model than one big prompt on a heavier model?

For routine work, a well-loaded prompt on a lighter model is usually cheapest. A fuller prompt adds input tokens but cuts repair turns, and dropping from Claude Fable 5.1 to Haiku 5.5 cuts the per-token price a hundredfold. Several thin prompts on any model tend to cost more than one good one.

How do you figure out if the lighter model is enough?

Try the task on the lighter model first and only move up if the answer falls short. Quick answers, summaries and extraction rarely need more than Haiku; writing and analysis suit Sonnet. A prompt improver with model suggestions makes that call on every prompt, so nobody has to memorise the table.

The ivee AI Learning Companion rewrites prompts, picks the right model and shares what works across your team, inside the AI tools you already pay for. See how the ivee AI Learning Companion works, then join the waitlist.

Don't know what you don't know? Book a call.

Book a call and tell us where you're at. We'll show you how other teams are tackling AI, and, crucially, what's actually paying off.

Don't know what you don't know? Book a call.

Book a call and tell us where you're at. We'll show you how other teams are tackling AI, and, crucially, what's actually paying off.

Don't know what you don't know? Book a call.

Book a call and tell us where you're at. We'll show you how other teams are tackling AI, and, crucially, what's actually paying off.