In short: Choose by how often the work repeats and what a wrong answer costs, not by price - because the cheap tier hides its cost in limits nobody publishes, the bought tier in a bill that grows with your success, and the built tier in maintenance nobody scoped.
Every business owner we speak to has already made this decision, usually without noticing. Someone put ChatGPT on the company card. Someone else bought an AI add-on because the vendor demoed it well. A third person asked a developer to “build something with the API”.
Eighteen months later there are three AI spends on the books, none of them measured, and the question “should we build or buy?” arrives far too late to be useful.
This decision goes wrong because owners make it on price, and price is the one dimension where all three options mislead you. The cheap tier hides its cost in limits it will not publish. The bought tool hides it in a bill that grows with your success. The custom build hides it in maintenance nobody scoped.
Here is the framework we use, with prices checked on official pages as of 31 July 2026, and the trade-offs that appear in no comparison table.
1. The three tiers, honestly described
Forget the vendor language. There are only three things you can do. Renting a capability is a sound choice, right up to the point the vendor withdraws it, as the Sora API shutdown shows.
Use a chat tool. A person opens ChatGPT or Claude, types, reads, decides. The intelligence is rented; the judgment and the effort stay human. ChatGPT Business is $25 per user per month billed monthly or $20 annually, with a two-seat minimum. Claude’s Team plan is $20 per seat annually, $25 monthly. Microsoft 365 Copilot is $18 per user per month on an annual commitment, promotional against a $21 regular price, and it needs a Microsoft 365 business plan underneath.
Buy a product that does one job. A support AI, a document AI, a sales AI. Somebody else built the workflow, the integrations and the guardrails. You configure and feed it.
Build something shaped like your business. Your process, your data, your systems of record, your rules about what the AI may and may not do alone. Whichever way you go, the same operating habits decide the result - see our 27 AI rules for business owners.
Most owners treat these as three price points on one ladder. They are not. They are three answers to a different question - who owns the workflow - and that is the question that should decide it.
2. What each option actually costs (beyond the sticker)
The chat tier hides its cost in limits nobody will publish
This surprised us most while checking prices. Neither OpenAI nor Anthropic publishes numeric usage caps for its business plans. OpenAI’s business documentation describes usage as unlimited subject to abuse guardrails, and notes usage may be restricted during high demand. Anthropic states that Pro and Max limits are shared across Claude and Claude Code, without giving numbers. OpenAI is explicit that seats include “baseline access” and that you buy workspace credits to exceed the included rate limits.
So the seat price is a floor, not a ceiling.
The trade-off nobody lists
You cannot capacity-plan a business process against an undisclosed, vendor-adjustable limit. That is fine for a person drafting emails. It is not fine for anything a customer is waiting on, or anything that must run on the last day of the month alongside everything else. If a process has to complete, the chat tier is the wrong home for it - whatever it costs.
There is a second, quieter problem with living at this tier. Thoughtworks put “AI-accelerated shadow IT” in the caution ring of its April 2026 Technology Radar, describing what happens when “the spreadsheet that quietly runs the business evolves into customized agentic workflows that lack governance”. That is the chat tier’s real failure mode: not cost, but critical processes quietly taking up residence in tools nobody governs. We wrote about the staff-side of that in shadow AI.
Two practical notes as well: self-serve ChatGPT Business is card-only, with no invoicing, purchase orders or net terms, which is genuine procurement friction if you need a PO. And on Claude’s Team plan, Claude Code sits on the premium seat at $100 per seat per month against $20 for standard - a five-fold fork that catches buyers out.
The buy tier charges for outcomes now, which cuts both ways
The pricing model here genuinely changed. Instead of paying per seat, you increasingly pay per result. Intercom’s Fin charges $0.99 per resolution, minimum 50 outcomes a month, one outcome per conversation however many actions it takes. Zendesk charges $1.50 per automated resolution, defining resolved as a 72-hour quiet period with no reopening. Sierra prices on outcomes too but publishes no rate.
This is more honest than per-seat pricing and we generally like it. But be clear what you signed: your AI bill now scales with your volume, not your headcount. Good if you are seasonal or shrinking. A different conversation if you are growing, or if one campaign triples your ticket volume. Check the definition too - what counts as a resolution is set by the vendor, not by you.
The other buy-tier risk is older and duller: Zylo’s research puts 53% of software licences unused or under-used. AI tools do not escape that. They are just the most confidently purchased line on the list.
The build tier hides its cost in a clock you did not know was running
Model retirement is a contractual reality, and the notice periods differ more than people expect. Anthropic commits to at least 60 days’ notice before retiring a publicly released model, and requests to retired models fail. OpenAI commits to at least six months for generally available models. Google’s Gemini documentation publishes no minimum notice period at all, listing only earliest possible shutdown dates.
Anthropic’s recent practice has sat near its 60-day floor: Sonnet 4 and Opus 4 launched in May 2025, were deprecated in April 2026 and retired on 15 June 2026 - roughly a 13-month production life on 62 days’ notice.
Migration is not a find-and-replace either. On Claude Opus 4.7 and later, the temperature, top_p and top_k parameters are deprecated and return a 400 error if set to a non-default value, with prompting recommended instead. An upgrade can mean re-tuning behavior, not renaming a string.
And the tooling has a clock too, which is the part almost nobody prices in. OpenAI is retiring its Assistants API in August 2026, and announced in June 2026 that its Evals platform, reusable prompts and Agent Builder shut down on 30 November 2026. If you build on the vendor’s agent-building tools, those tools are on the same treadmill as the models.
The two facts that belong together
LangChain’s State of Agent Engineering survey (fielded November to December 2025, 1,340 respondents) found only 52.4% of teams run offline evaluations and 37.3% run online evaluations, even though observability is near-universal. Set that beside a 60-day retirement notice and you get the real switching cost: roughly half of teams cannot measure whether a new model is worse than the one they are being forced off. They can neither leave easily nor verify the move they have to make.
None of this argues against building. It argues for budgeting maintenance from day one, keeping evaluations in place so a swap is measurable, and holding the model behind an interface you control - which is the practical case for building on the Model Context Protocol rather than hard-wiring one vendor’s SDK through your codebase.
One honest note on maintenance. There is a widely repeated convention that annual maintenance runs 15% to 30% of build cost, and it is a reasonable band to budget against. Treat it as a rule of thumb rather than a finding: every source we could trace for it is an agency or vendor page with no methodology, no sample size and no disclosed dataset. Useful for planning. Not evidence, and anybody presenting it as a measured figure is overstating what exists. We quote maintenance against the actual scope rather than a flat percentage.

3. What the evidence says, including the part usually left out
The clearest signal is that the market already decided. Menlo Ventures surveyed 495 US enterprise decision-makers in November 2025 and found 76% of AI use cases were purchased rather than built internally, against 47% built and 53% purchased just a year earlier. That is a decisive one-year reversal.
a16z’s survey of 100 CIOs (June 2025) gives the reason in buyers’ own words: internally developed tools proved “difficult to maintain and frequently don’t give them a business advantage”. Notably, their top purchasing criteria had become security first and cost second, with accuracy third - because, as one leader put it, for most tasks the models now perform well enough that pricing dominates.
You will also see MIT’s State of AI in Business 2025 quoted here, reporting that “external partnerships with learning-capable, customized tools reached deployment ~67% of the time, compared to ~33% for internally built tools”. Treat that one carefully. The same report warns in the same section that the difference “may reflect organizational capabilities rather than implementation approach alone” and that correlation “does not necessarily prove causation”. The sample was 52 organizations and 153 leaders, self-reported, and the document is labeled preliminary findings.
Our reading: buying is not better than building. Buying is more forgiving of a business that has never shipped software before. The gap is mostly measuring organizational muscle, not architecture - and you can assess your own muscle honestly in about ten minutes. If you do decide to build, the gap that actually sinks projects is the distance between a working prototype and production software.
The strongest case against building
“We are not a software company, and every hour on this is an hour off our actual business.” This is right more often than the AI industry admits. If the work is genuinely generic - answering common questions, summarizing documents, drafting first passes - somebody has already built it better than you will, and maintains it for a living. Build when the workflow is the business: your pricing logic, your compliance rules, your operational sequence. Nobody sells that, because nobody else has it.
Should I build or buy AI for my business?
Buy when the workflow is standard across your industry and a vendor maintains it for you. Build when the workflow is specific to your business, repeats constantly, and a wrong answer costs money. Use a chat tool when the work is genuinely one-off human thinking. Most businesses need two of the three at once.
Two questions settle it faster than any feature comparison.
How often does this exact work repeat? A few times a week, and a person with a chat tool is correct; anything else is over-engineering. Dozens of times a day, every day, and the human cost of the chat tier will quietly exceed both other options inside a year.
Would a wrong answer cost money or only time? If only time, buy the cheapest thing that works. If it costs money, reputation or a regulator’s attention, you need control over what the AI may do unsupervised - which means a vendor who exposes those controls, or a system you own. We set out why that boundary matters in what AI hallucinations really cost a business.
The decision table
| If this is true | Choose | Watch out for |
|---|---|---|
| Work is varied, occasional, human-judged | Chat tool, business tier | Undisclosed limits; keep it off anything time-critical |
| Work is standard for your industry | Buy a product | Outcome pricing scales with growth; shelfware risk |
| Work repeats daily and follows your own rules | Build | Maintenance, and retirement notice as short as 60 days |
| A wrong answer moves money | Build, or buy with hard controls | Anything that cannot enforce a cap or an approval step |
| You have never shipped software before | Buy first, build later | Treating a pilot as proof you can operate it |
| The workflow is your competitive advantage | Build | Buying here quietly hands your differentiator to a vendor |
Myth vs Facts
Myth: “Just use ChatGPT, it is the cheap option.”
Fact: At $20 to $25 a seat it is cheap per person and expensive per task, because every task still costs a human’s time. It is also bounded by limits neither major vendor publishes numerically, so it cannot safely carry a process that must complete.
Myth: “Building is cheaper than paying per seat forever.”
Fact: Only if you budget maintenance. Retirement notice can be 60 days, and API parameters get deprecated between versions - temperature and top_p now return an error on newer Claude models rather than being ignored. The vendor’s agent tooling gets retired too.
Myth: “MIT proved buying beats building two to one.”
Fact: The report published that gap and then warned it may reflect organizational capability rather than approach, from a self-reported sample of 52 organizations labeled preliminary. For the market shift itself, Menlo’s 76% figure is the better-sourced number.
Myth: “The model is where the money goes, so we should optimize tokens.”
Fact: In a 2025 survey of 372 organizations, the top sources of unexpected AI spend were data platforms and network access costs. LLMs ranked fifth. Four in five organizations missed their AI cost forecasts by 25% or more.
What this means if you are running AI in your business
Businesses that get this right did not choose correctly on day one. They chose separately for each piece of work, and then re-checked.
A realistic end state for a mid-sized business is all three at once: business-tier chat accounts for everyone, one bought product for a standard high-volume function, and one built system around the workflow that is genuinely yours. What makes that a strategy rather than sprawl is knowing which is which, and having a number attached to each.
- List every AI spend you have, including the ones on personal cards
- For each, write down how often the work repeats and what a wrong answer costs
- Move anything time-critical off the chat tier, whatever it is costing you
- On any outcome-priced tool, model the bill at double your current volume before signing
- On anything built, name who handles the next model retirement, and check you could measure whether the replacement is worse
- Set one review date a quarter, because vendor terms move faster than your handbook
If you are earlier than that, the honest first step is to buy or rent, prove value on one workflow, and build only once you know which part of it is actually yours. That sequencing is also what separates deployments that reach production from those that stall, which we unpack in why agentic AI projects fail.
Are you in the wrong tier?
Tick each one that applies to you.
- A customer-facing or deadline-bound process depends on a chat-tier AI account
- We have an outcome-priced AI tool and have never modelled the bill at higher volume
- Someone built us an AI feature and nobody owns it now
- We pay for an AI product that fewer than half the licensed staff open weekly
- Our AI logic is written directly against one vendor’s SDK with no layer in between
- We could not tell whether a new model version made our output worse
- We chose our current approach because of the monthly price
If this were your decision, here is our first move
We would not start with vendors. We would take your three highest-volume repeating tasks and put numbers against each: how many times a week, how long a human spends, and what happens when it is wrong. Those numbers pick the tier on their own, and they usually pick differently for each task.
The businesses that waste money on AI are rarely the ones that chose the wrong tool. They are the ones that chose one tool for everything and then defended the choice. Every price in this article will be stale within months. The two questions - how often does it repeat, and what does wrong cost - will not be.
If you want a straight read on which of your workflows belongs in which tier, talk to us. We will tell you where buying beats building, including when that means not hiring us to build it.