By now, you’ve probably heard that open models have caught up with proprietary ones. I think that claim needs a closer look.
Look at Xiaomi’s MiMo-V2.6-Pro. The open-weight model scored 46 on Artificial Analysis’ Intelligence Index, just two points behind OpenAI’s GPT-6 Sol at max effort.
Yet the measured cost per benchmark task was $0.13 through Xiaomi’s API, versus $1.06 through OpenAI’s. That’s close overall performance at roughly one-eighth the API cost.
So why spend engineering time running open models yourself when you could use a hosted API?
Well, the shift toward open-source models hasn’t got much to do with them being dramatically smarter. It’s just that they’re perfectly capable of doing all that work.
Add better cost control, data residency, and freedom from vendor lock-in, and the case for open source becomes pretty compelling.
Where open models fit in an agent
Frontier models still take the gold medal for the hardest reasoning. But your agent doesn’t need that level of reasoning on every call.
Say a customer sends an agent an invoice and asks why a payment hasn’t appeared. The agent may use OCR to read the scan, an embedding model to search payment policies, a reranker to choose the right passage, and an extractor to pull the invoice number. Only then does it need to answer the customer’s question.
Teams can put high-volume, everyday work on open models and reserve frontier-model calls for cases that need them. Retrieval, extraction, classification, and routine generation are candidates, provided the model passes tests on your data.
What pushes teams to make the switch
These are the four reasons why teams usually switch to open-source models:
Cost is often the top reason.
With a closed API, you pay for every token, whereas a GPU in a self-hosted setup is a fixed cost whether it’s busy or not (excluding electricity, a nearby sea, etc.)
However, keep in mind that self-hosting isn’t automatically cheaper.
It becomes more affordable when your volume is high and steady enough to keep those GPUs busy doing the math. In practice, utilization is the main factor when deciding whether to go fully self-hosted.
I’d count successful tasks, not just tokens or GPU hours. If a cheaper extraction model gets a field wrong and you have to run the document through a frontier model anyway, the first call didn’t save you much.
Add the time spent operating the system, too. Someone will have to deal with slow requests, failed deployments, and the occasional model that won’t fit where you expected.
Data control is sometimes a hard requirement.
With closed APIs, your prompts and documents are processed by another company under its terms.
Major providers do offer controls. OpenAI, for example, says API data isn’t used to train its models unless you opt in. For plenty of applications, the provider’s controls and a contract are enough.
Still, if you handle regulated data or run in an air-gapped network, keeping inference in your own cloud may be a requirement. You also get to decide where requests are logged and how long those logs stay around.
Of course, that means your team has to set those rules up and make sure they work.
Model control is another reason I find interesting.
You can pin a model version while you test its replacement. That’s useful when an agent depends on a particular output format or when a small change in retrieval quality sends it the wrong documents.
But swapping models is rarely as easy as changing a name. Change an embedding model, and you’ll usually need to rebuild your search index.
Change a generation model, and you need to check its prompts, tool calls, and structured output again. You have more choices, but you’re also the one testing them.
When This Move Isn’t for You
Open models are not universal solutions. A closed API is still a better choice when:
A task needs the top reasoning tier
Your demand is small, new, or spiky
You need broad multimodal work
You want the newest model without deploying it yourself
Your team has no resources to run and maintain the infrastructure
Measure twice, cut once. And I’m not just being metaphorical here. Measure for real, then let those measurements drive your decisions. If it turns out you really need frontier models, use them without giving it a second thought.
For retrieval, I usually check whether the right document appears in the results, then whether the reranker puts the useful passage near the top.
For extraction, compare the returned fields with the source file. Use your own documents, especially the messy ones.
A benchmark score won’t tell you whether a model understands your product names or reads your PDFs correctly.
Hybrid Pattern
You don’t need to abandon closed APIs entirely. You could split the workload depending on the task at hand.
Keep high-volume, low-difficulty work on open models you control.
Pay for frontier-model tokens only when the value is obvious.
A good place to start experimenting with local models is the retrieval and extraction layer. These steps are easier to check than an agent’s final answer, so you can see fairly quickly whether the model is good enough.
Imagine a support agent reading an email with an attached invoice. An open model could extract the invoice number; another could find the relevant account notes.
The frontier model can handle the actual reply if the case needs judgment. If the invoice number is missing or the search results look weak, send that case down a different path.
I recommend you track how often that happens. If most requests end up at the frontier model anyway, you’ve added work without cutting much of the bill.
Superlinked’s SIE vs. OpenAI breakdown and its walkthrough of five agents show what this kind of split can look like.
If You Decide to Run Open Models Yourself
Deciding to make this move is one thing. Running the models, however, is a completely different beast. Encoders, rerankers, extractors, OCR, and small generation models all have their own serving quirks and their own spots on a GPU.
That’s the market SIE (Superlinked Inference Engine) is built for.
https://github.com/superlinked/sie
It’s an open-source inference engine that gives you one endpoint for a catalog of open models. It covers embeddings, reranking, extraction, OCR, and generation.
You can run it on your own hardware or in the Superlinked cloud.
The main reason why it should appeal to you is its practicality. An agent using several small models shouldn’t need a separate serving system for every task. SIE can load models as needed, though you should test what that does to response times with your own traffic. It also won’t decide which model is accurate enough for your application. You still have to do that part.
Take it for a test drive, and see which models it serves. I’d begin with one retrieval or extraction task, compare it with your current API over a normal week, and look at quality and the full operating cost before moving anything else.
Frequently Asked Questions
Where should I try an open model first?
Start with a task you can check, such as retrieving a document, reranking search results, or extracting an invoice number. Test it on your own data, including the messy cases.
Will running open models myself save money?
Only if the work keeps your hardware busy enough. Compare your current API bill with GPU time, operations, and the cost of sending failed requests to a stronger model.
Do I need to replace my closed API?
No. Keep it for requests that need stronger reasoning. An open model can handle the routine steps, then pass the relevant information to a frontier model when needed.
What do I need to run open models in production?
An inference layer that serves your encoders, rerankers, extractors, and generators and shares GPUs across them. You can build that layer yourself or use something like SIE.
Hi there! Thanks for making it to the end of this post! If you enjoyed this content and would like to support my work, consider becoming a paid subscriber. Your support means a lot!







Great article and you make a lot of good points about the choice points for local vs frontier. I've had great success with local models in my own work but still rely on Claude for a lot of heavy lifting. Hopefully as these models advance that will change.