How to Use Open Source AI Models (Stop Wasting Money)
Open source AI models now handle a surprising amount of real work at a fraction of frontier pricing, and most people are still paying Opus rates to rename files.
Watch Me Test Open Source AI Models Against My Own Landing Page
This episode of AI Rabbit Holes is a conversation with Aaron Makelky, an AI consultant who has been running open models for years.
Partway through, I pointed the three top-ranked open models at my real opt-in page and asked each one to redesign it.
The results are further down, and they are not what the benchmark charts would lead you to expect.
If you want the Claude Code setup behind tests like this one, my Claude Code Skills Stack ships the same skills and systems I run daily.
Open Source vs Closed Source AI: What You Are Actually Choosing
Most comparisons frame this as a loyalty test. Pick a side, defend it.
That framing is wrong, and neither of us runs a business that way. We both pay for frontier models and route the cheap work elsewhere.
Aaron uses a right to repair analogy. Farmers bought John Deere equipment, then had to haul it back to the dealership for any repair.
Closed models work the same way. To use OpenAI's desktop app, you use OpenAI's models.
Worth getting the term right, because most roundups do not. Nearly everything sold as open source is really open weight, meaning the weights are published while the training data and code usually are not.
That distinction changes the ownership question, not the quality question. You hold the blueprint, so the hosting and the bill are yours to decide.
Frontier models still win on hard reasoning, layout work, and anything where a subtle mistake gets expensive. Open source AI models win on volume, price, and privacy.
The skill is knowing which bucket a task belongs in before you spend anything.

The local hosting question, answered honestly
Aaron's estimate is that 99% of people are not going to self-host a full-size model. He puts the hardware for running something like Kimi or GLM 5.3 at home around eight to ten thousand dollars in compute.
Do that math against a $100 monthly subscription and the hardware rarely wins. Aaron and I ran the subscription side of that comparison in our episode on whether a $100 plan is worth it.
My own Mac Mini and MacBook Pro are not enough for serious local work at scale either.
Small local models are a different story. Aaron runs a 3GB Qwen vision model on a Mac desktop that renames his screenshots, offline.
Twenty or thirty screenshots a day, each saved as a meaningless date string, become things like "Google Chrome Wells Fargo account September 14."
That task is the whole argument in miniature. Bank statements should not travel to a frontier API just to get a filename, and that filename is not worth frontier rates.
Where to Run Open Source AI Models Without Buying Hardware
You have spent money on a model you cannot open anywhere. This is the gap nobody explains.
OpenRouter is the middleman I use, and the pitch is simple. One account, one API key, hundreds of models.
That key plugs into Claude Code, Codex, Cursor, or any IDE you already run. The alternative is a separate key from Z.ai, DeepSeek, and Moonshot, all managed in environment files.

Look at the input column on the day I captured this. GPT-5.6 Luna, OpenAI's cheaper tier, was listed at $0.20 per million input tokens.
DeepSeek V4 Flash 0731 read $0.03 and GLM 5.3 Flash $0.075, though that one carries a visible discount badge, so treat it as a promo rate and not a standing price. Nemotron reads $0, because NVIDIA publishes a rate-limited free endpoint alongside the paid one.
Three traps before you budget off any of this. Model names carry dated snapshots, so V4 Flash 0731 and the newer V4.1 Flash are different models at different prices.
Open weights also get served by dozens of competing providers at different rates, so the headline number on a listing page is one provider's price rather than the model's price.
And these figures move week to week. Every number here is a snapshot, not a quote, so read the live rate for whichever provider you actually select.
At one prompt none of this matters. At a few million tokens a month of background work, it stops not mattering fast.
Start with fifteen or twenty dollars of credits. Open models are cheap enough that a small balance lasts through a lot of testing.
Why I skip the discount portals
Portals like Kie.ai advertise steep discounts against the official rates on the same models. I avoid them, and this one is a hunch rather than a tested finding.
Heavy discounting usually means distillation.
A distilled model is a smaller, cheaper imitation of the original. It answers faster and costs less, and quality drops in ways that are hard to see until the work matters.
Aaron reaches these models through Hermes instead, which holds his ChatGPT subscription, a Z.ai coding plan, and a portal of open models in one interface.
His setup does one thing I want to steal. He fires a single prompt at DeepSeek, Kimi, and a frontier model at the same time, then reads all three answers side by side.
Last I checked, Anthropic's own app will not do that with two Anthropic models. You open a second tab and paste the prompt again.
Think about what that does to copywriting. Four cheap drafts of a subject line beats one expensive draft, every time.
The Best Open Source AI Models Right Now (And Who Ranks Them)
Picking a model from memory is how you end up using something that was current six months ago. Let the ranking decide instead.

The Arena LLM leaderboard ranks models by blind human preference across math, coding, and creative writing. Filter it to open source and you get a live answer instead of a guess.
When we recorded, that filtered list read kimi-k3-max first, then glm-5.3-max, glm-5.3-flash, glm-5.2-max, and mimo-v2.5-pro.
Now read the second column, because it is the most honest number on this page. The top open source model sits at overall rank 17 once closed models are back in the running.
Hold onto that gap. These models are genuinely good and genuinely not the frontier, and both facts are visible in one screenshot.
My own rotation is narrow. Kimi K3 and the GLM models get most of my open source work, with DeepSeek V4 for workhorse tasks.
Plenty sit outside that rotation untested. MiniMax, most of the Qwen family, Tencent's Hunyuan, Mistral.
Aaron has actually settled on one. GLM Flash has been his default for at least nine months, for background tasks rather than creative or strategic ones.
His examples are deliberately boring. Pull this file up, run a search, draft this.
Those tasks are hard to get wrong, and GLM Flash does them for almost nothing.
One correction worth making, because we got this slightly wrong on the recording. Vision lives in the Flash line, not the main GLM 5.3 model, which is still text only.
So if you need to hand a model a screenshot, check the modality before you pick the bigger number.
I Let the Leaderboard Pick Three Open Source AI Models and Tested Them
Benchmarks are votes on chat responses. They do not tell you whether a model can rebuild a page you actually use.
So I did not pick the models. I handed over the leaderboard URL and told it to fetch the current top three, then test those.
That is the part worth copying. Run this next quarter and the leaderboard hands you a different three, and the method still works.
The target was my opt-in page for the free AI marketing guide, a deliberately simple page built to convert.

One detail surfaced immediately. Those max and flash suffixes do not map onto OpenRouter model IDs, so ranks two and three both resolve back to Z.ai GLM.
Kimi K3 produced the most designed result. Editorial headline, floating proof cards, sticky navigation, the most visual personality of the three.
It also mangled my branding at the top and rendered a crooked frame around my headshot. Aaron called the frame wonky, and he was right.
Close, but not quite shippable.
GLM 5.3 Flash came back plain. Black and white, almost no branding, nothing above the fold that grabs you.
Which is roughly what a flash model should produce. Aaron noted he would rather ship a clean black and white page than a busy one chasing gradients, and for many opt-in pages that is the correct instinct.
GLM 5.3 landed closest to usable. It underlined one phrase in the headline to pull the eye toward the offer, a real copywriting decision nobody asked it to make.
Spacing problems remained. Usable with a few fixes is the fair verdict.
These outputs sit roughly where the frontier models were six months to a year ago. That is the calibration to keep.
Aaron put a sharper point on it, saying the GLM result looked like Opus 4.8 work, a model less than a year old.
If you want the deeper version of this, I put six models including Qwen and MiniMax against a Claude-built homepage and logged what each run cost. That breakdown lives in my open source model bake off for web design.
A trick for making cheap models smarter
If you already pay for a frontier plan, you are sitting on an asset you are probably not using.
Ask Fable 5.1 or GPT-6 Astra, in a high thinking mode, to write a skill markdown file capturing how it approaches a specific task. Then point a cheap open model at that file.
The open model inherits some of the frontier reasoning without the frontier bill. It is not a one to one transfer and I would not claim it is, but it noticeably lifts output on repetitive work.
This pairs with knowing your own model tiers, which I broke down in my guide on which Claude model to use.
When Open Source AI Models Are the Wrong Answer
Here is where I would not touch an open model.
High-stakes creative and strategic work still belongs on frontier models. Marketing strategy, complex visual design, and anything a client sees first are not places to save three cents.
Free models deserve particular caution. Aaron's warning is blunt, that they are free for a reason, they train on your data, and you cannot opt out.
The same goes for code-named mystery models on these portals. A model listed under a codename with no stated lab is a lab testing compute and collecting data.
Use those for side projects and throwaway agents, never for client data.
Hosted open models carry a quieter version of the same issue. Aaron assumes his GLM prompts are being read somewhere, and he uses the model accordingly.
Open weights do not mean private inference unless the weights run on your own machine.
The offline option nobody talks about
For genuinely sensitive work, there is a free path that skips the hardware question.
Google AI Edge Gallery runs Gemma models directly on a phone, and it is now on the App Store and Google Play rather than a sideloaded file.
The current build features the Gemma 4 family, with the smaller variants landing a couple of gigabytes. Google describes the app as in active development, so treat it as a preview rather than a finished product.
Once a model is downloaded, turn off wifi and it still answers. It handles images too, so you can hand it a photo and ask about it.
Aaron uses it on airplanes and for medical or banking questions. Nothing leaves the device, which is a different privacy guarantee than any subscription offers.
Nobody would call it a great model. It is a decent one that works when nothing else does.
Route the Boring Work Somewhere Cheaper This Week
Open source AI models are rate limit insurance as much as a cost play. When a provider retires a plan or throttles you mid-project, a second lane matters.
At real volume the math only points one way. Teams in China and elsewhere ship upgraded open models continuously, and more of them land free.
So keep your frontier subscription for the work that earns it. Put fifteen dollars on OpenRouter and move your background tasks over this week.
Spend the savings on whatever actually moves your business.
Open Source AI Models FAQs
Are open source AI models as good as ChatGPT or Claude?
Not at the top end. On the Arena leaderboard, the best open source model ranks 17th overall once closed models are included. For mid and low-level knowledge work the gap rarely matters. In my landing page test, the open models produced results comparable to frontier models from roughly six months to a year ago.
What is the difference between open source and closed source AI?
Closed source means you rent access and the provider controls where it runs, so your data goes to their servers and your bill follows their pricing. Open weight means you can run the model anywhere, including on hardware you own. The difference is control and cost rather than automatic quality, and the top closed models still lead on benchmarks.
How much does it cost to run open source AI models?
Less than frontier pricing, though the number moves faster than most guides admit. The same open model is served by dozens of competing providers at different rates, and promotional discounts expire, so the headline figure on a listing page is one provider's price on one day. Budget from the live rate of the provider you actually select.
Can I run open source AI models on my own computer?
Small ones, yes. A 3GB vision model runs comfortably on a Mac desktop, and Google AI Edge Gallery runs Gemma models on a phone offline once you download one. Full-size open models need serious hardware that rarely pays for itself against a monthly subscription.
Are free open source AI models safe to use?
Treat anything free as public. Free tiers generally train on your inputs with no opt out, and free endpoints are usually rate limited versions of a paid model rather than a different product. Run side projects on them and keep client data, financial records and anything personal off them entirely.
