These are not the models you're looking for
I read the LLM benchmark charts, or I try to. The truth is I do not fully understand what they are telling me, and I do not think I am the only one. They get posted as if the verdict should be obvious. Really what I use them for is checking whether a model is even still on the board, or has it fallen off the map. I do not need to understand the score for that; I just need to see if the name shows up. I started out running models locally on a cheap, cobbled-together rig. That was fine for the simple stuff, but when the work got more complex, the hardware demands shot up fast. So now I buy inference from someone else. The benchmark answers a question I never asked.
The charts stopped working
The rankings have stopped being useful, and it is not a hot take to say so. The tests they use to rank these models are either maxed out, with everyone scoring near the top so the spread is noise, or they measure something that has nothing to do with the work you need done. A model can ace every test on the board and still be useless for your task. The people who study this admit the surviving numbers are partial and that a ranking is a guess, not a verdict. I have no argument with the testing; I just skip to the part where I open a price list.
Data terms
What happens to my data comes first. Both things a provider can do with my data are offensive to me: keeping it after the inference is done, and using it to train their models. I want a US-hosted provider with a published retention term and a published statement about training, and here is what that actually buys: “zero retention” is a policy promise, not a measurement. My prompt still leaves my machine in plain text, arrives at their server readable, and sits in memory while it runs. What the term gives me is a contract not to keep it and not to train on it, which is worth having, but it is not something I can verify. I take that on faith, not on evidence. As one person buying one seat, the published page is the only handle I get.
I assumed open-weight models were the safe choice on data and proprietary models were the risk, because the small user with no contract gets the default terms and the default terms are usually “we keep it.” The retention tables do not split along that line. Some open-weight models are hosted by providers listed as may-train, and some closed models are served under no-train terms. The line that matters is whether the terms are published and readable, and whether I can walk away. DeepInfra, which I’ll get to, mostly meets that bar: US-hosted, published no-train terms, data held in memory only during inference.
Billing
I prefer paying by the token over the subscription models. The subscription providers cap your usage, and their advice for controlling it is unnatural once you are in the zone on a project. You hit the limit and they slam the door, forcing you to wait. That breaks your flow in a way that costs more than the tokens would. By-the-token billing has no cap; it just costs what it costs, and you watch the meter move. The risk is that “what it costs” can get away from you fast, and the big companies proved it. Microsoft told its own engineers to stop “tokenmaxxing” this summer and set division-level AI budgets, because even the company hosting the models could not afford the usage-based bills. Uber exhausted its entire 2026 AI coding budget by April. These were not subscription problems; they were by-the-token problems on expensive models. The reason I can make by-the-token work is that I opt for open models, where the cost is lower and the output is good enough for most things. The bill adds up, but slowly enough that I do not have to watch it like a hawk. The per-token price on the page is the least useful number there; what I actually pay depends on how much of the conversation the provider can reuse, and I get there by feel rather than by spreadsheet.
Reliability
Reliability is where the job lives, and it is where my last provider lost me. I had been a customer of the same cloud provider for thirteen years, with servers running there, and I wanted it to work. The model I was using was fine for chat, but broke in weird and interesting ways the moment I asked it to do something structured. I was coding with it in OpenCode, an agent harness that lets the model take actions on your behalf, not just chat. When the model is supposed to send a request to my harness in a format it can read, and it comes over garbled instead, the whole thing falls over. I also ran into rate-limiting with several models; I guess I wasn’t the only person out there looking for low-cost inference. The final straw was a change to their billing policy: a prepaid inference pool applying against my hosted server that had always been billed in arrears. A prepaid pool means you’re loaning the vendor money upfront, and it is also the vendor protecting itself from users who run up crazy big AI bills they cannot pay. I did not leave over a benchmark score; I left over broken results and a billing model that flipped the script on me.
I moved to DeepInfra because the prices were the best I found, the reliability was good, and the model list was wide. I won’t pretend it was a rigorous evaluation. The data terms, the billing shape, the reliability, and the model breadth were the lens, and this one passed.
Model breadth
Model breadth matters because I pick a model per task. A vendor carrying one family is a vendor I outgrow, and a plan that pins me to one family is lock-in with a discount painted on it. There is a sharper version of this: some providers route the model you named to a different version without telling you in the UI. In September, OpenAI was caught serving GPT-5.6 Luna when people selected GPT-6 Astra on their Codex endpoint, confirmed by response headers that named the cheaper model. The users paid the top-tier rate and got the cheaper model, and the only trace was buried in metadata most people would never check. If you keep your own notes on what works, that kind of swap quietly ruins them.
The method
The method is what actually replaces the chart. Use the cheapest model that clears the task and bump up when it fails. The signal that matters is the kind of work, not the model tier. My former go-to cheap model was fine for single-file scripts and normal work but produced absolute garbage for web apps; that is a cliff no chart showed me. A mid-tier open-weight model is what I use for apps now, not because it wins a ranking but because it cleared my task.
This is called a model cascade, and the research shows it can match the most expensive model at a fraction of the cost, at least for the kind of work where you can tell a good answer from a bad one. The hard part is not the idea; it is knowing where to draw the line, which is the thing I had been doing by eyeball. The economist Dean Baker made the same point from the other direction: companies do not need frontier models, the previous versions are plenty good enough, and the question is who is going to pay for the expensive ones. If the cheap model clears the task, the expensive one is a luxury, not a requirement.
Anyone writing about inference should show the bill. I have spent about $85 in the past month, which is more than the twenty-dollar subscription everyone points at, and I know it. What I get for the difference is terms I can read, models I can change, a service that behaves inside my agent, and no contract I am too small to sign. Whether that trade is worth it depends on your work.
The charts can wait.