No single model, and that's the honest answer. The buyers who manage AI spend well sort their tasks by what each one requires, choose by price within that tier, and keep switching costs at zero so every decision stays reversible.
Copy link to headingWhy is there no one answer?
Every page ranking AI models answers with a leaderboard, and leaderboards go stale within weeks. The August 2026 AI Gateway Production Index describes what production buyers do instead. They treat buying as tiered rather than as a single-price race. The task sets the tier, and price competition happens inside it. The Index notes that the cheapest and most expensive models were never competing for the same work.
The spend numbers reflect that sorting. In July, Anthropic collected 65.1% of gateway spend on 30% of volume. In the Index's terms, a large share of buyers pay a premium for the models they place in the top tier for their tasks. The figure describes how buyers sort their work. The question for your own work is which tier each task belongs to.
Copy link to headingWhat does the July 2026 data show?
Average price per token across the gateway fell 13.6% in July, according to the same Index edition. List prices didn't drive the fall. Holding June's mix of models constant, the average would have held essentially flat. Teams paid less because they changed which models they used.
The average hides a split. Among teams that ran more than 10 million tokens in both June and July, a quarter cut their cost per token by more than 30%, while another quarter paid at least 20% more. Same market, same list prices, opposite outcomes. Model-mix decisions made the difference.
Copy link to headingHow to decide what to pay for
The framework needs no particular product. Three steps:
Tier your tasks by what they require. Some tasks punish a wrong or mediocre answer (contract review, customer-facing writing, hard debugging), while others reward volume over brilliance (summaries, first drafts, reformatting, routine questions). Nothing about any model enters this step.
Price inside the tier. Comparing a frontier model's price to a budget model's price tells you nothing, because you'd never hand them the same task. Compare only models that can handle the given tier, and choose the cheapest one that performs well on your tasks. The cheap tier carries real production load. Open-weight models ran over a third of July's gateway volume at about a seventh of frontier rates, and if that tier tempts you but data handling holds you back, there's a guide to trying open-weight models safely.
Keep the decision reversible. Whatever you pick, make sure moving off it costs nothing: no rewrites, no migrations, no renegotiated contracts. The July spread shows why. The teams that paid less were the ones that changed their mix, so revisit the choice monthly; the data that justified it ages that fast.
Copy link to headingWhat does this look like in practice?
The Vercel AI Gateway is one way to run this framework with zero switching costs. It puts 300+ models behind a single account, and any of them work with any of the 9 supported coding agents. Switching is a one-line change in the agent's model picker, so nothing locks you in, and re-pricing a tier next month means changing that line again. Budgets and a single spend dashboard come with it; the hub guide to using any coding agent with 300+ models covers both. Using the gateway requires a Vercel account and the Vercel CLI.
Copy link to headingWhat this framework is not
It's not a ranked list, and it names no winner. Prices and rankings move monthly, and a specific recommendation printed today would be out of date within a month. The framework survives that churn; a model name wouldn't.
It's not a claim that cheap models cover every task. They don't. Some work justifies premium models, and the July spend data shows buyers act on it. When the task needs frontier output, the tier framing says pay for it.
Copy link to headingFrequently asked questions
Copy link to headingWhich AI model is the cheapest?
Price per token only makes sense within a task tier. The August 2026 AI Gateway Production Index found that the lowest-priced and highest-priced models never competed for the same work, and rankings move monthly. Sort your tasks into tiers first, then compare prices only among the models suited to each tier. The Vercel AI Gateway model catalog lists current per-token prices.
Copy link to headingDo cheap models work for every task?
No. Some work justifies a premium model, and the spend data bears that out. A large share of gateway spend goes to the models buyers reserve for top-tier tasks. Use cheap models where the task allows, and pay premium rates where the output warrants it.
Copy link to headingHow often should I revisit which model I pay for?
Monthly is a sound cadence. The AI Gateway Production Index publishes monthly, and its August 2026 edition showed that cost changes came from teams changing their model mix rather than from list prices moving. A model choice that was right two months ago isn't guaranteed to be right now.
Copy link to headingAm I locked in when I pick a model?
No. Through the Vercel AI Gateway, switching models is a one-line change in your coding agent's model picker, so the decision stays reversible. Treat each choice as provisional, and if a cheaper model in the same tier can handle your work, move it.
Copy link to headingHow do teams decide which model to use?
Teams that manage AI spend well group their tasks by how much a wrong answer would cost, then pick the model that best fits each group by price. In the August 2026 AI Gateway Production Index, the teams cutting costs in July didn't do so by lowering list prices; they changed which models handled which tasks.
Build a software factory with Vercel Connect
Ship a software factory with GitHub and Linear connectors provisioned for you from the first deployment.
Deploy now