Software Ownership

How to Run AI Models Locally in 2026

Open-weight MoE models now rival frontier APIs, and unified-memory hardware can hold them. Here's how to run one, and the usage threshold that decides if it's worth it.

At a glance
  1. 01AI application spend doubled in 2026, pushing teams to reconsider local models over costly metered APIs.
  2. 02New desktop hardware with unified memory runs 320-billion-parameter models natively without server racks.
  3. 03Local deployment breaks even with proprietary APIs once daily volume reaches 2 to 3 million tokens.
  4. 04Shadow AI creates $670,000 in excess breach costs, making local enterprise models a critical governance fix.
Cutaway illustration of a unified-memory computer block with a mixture-of-experts module ring above it, showing only a few modules actively connected to the shared memory slab, alongside a threshold gauge bar.
Illustration generated by Remy for this story.

Running an AI model locally in 2026 means downloading an open-weight model like GLM-5.3-Flash, quantizing it to fit your machine's memory, and serving it through a tool like Ollama, LM Studio, or vLLM. No API key, no metered bill, no data leaving your building. The catch: it only pays off once your usage crosses a specific threshold, and finding that threshold is the actual work.

The API bill that made local AI worth a second look

The trigger for most teams isn't ideology. It's the invoice.

AI-native application spend rose 108% year over year across organizations tracked in Zylo's 2026 SaaS Management Index, and 393% at companies with more than 10,000 employees.1 That index covers more than 40 million licenses and $75 billion in spend, and it puts median SaaS spend per employee at $9,455 a year.1 Seventy-eight percent of IT leaders reported unexpected charges tied to consumption-based or AI pricing in the past year.1 Consumption pricing is efficient for the vendor and unpredictable for the buyer. That's the setup that makes owning the model, not renting the endpoint, worth a second look.

Figure 1
The invoice that's driving the local-AI conversation
108%
AI-native app spend growth YoY, all organizations
393%
Spend growth at 10,000+ employee organizations
$9,455
Median SaaS spend per employee, per year
Source: Zylo

What changed in 2026: models good enough to own

For years, the honest answer to "should I run this locally" was no. Open models lagged the frontier by too much to bother. That gap has mostly closed.

GLM-5.3-Flash, released by Z.ai under an MIT license with full weights posted on Hugging Face, is a mixture-of-experts model with 320 billion total parameters but only 18 billion active per token.23 It approaches Claude Opus 4.8 on coding and agentic benchmarks while running at roughly one-tenth the price of its predecessor.2 Z.ai's own hosted API prices it at about $0.15 per million input tokens and $0.50 per million output tokens, already remarkably cheap.4 That's worth sitting with: even the rented version of this model is inexpensive. The case for running it yourself isn't about beating that price per token. It's about ownership, privacy, and not being one pricing-page update away from a different bill.

Figure 2
GLM-5.3-Flash: total parameters vs. active parameters per token
6%Active per token
Active per token6%
Routed but inactive per token94%
Mixture-of-experts design: memory must be sized for all 320B parameters, but compute cost per token tracks only the 18B active.
Source: Z.ai

What changed in 2026: hardware that can hold them

The other half of the shift is memory, not compute.

Apple's Mac Studio with M5 Ultra, announced in August 2026, scales to 512GB of unified memory with 1.2TB/s of bandwidth, and Apple markets it specifically for running local models with hundreds of billions of parameters on a desktop.5 That matters because of how mixture-of-experts models work. Total parameters determine how much memory you need to load the model. Active parameters determine how much compute each token actually costs.46 GLM-5.3-Flash needs memory sized for a 320-billion-parameter model, but at inference time it behaves like an 18-billion-parameter model: fast and cheap to run per token.4

At 4-bit quantization, the full model needs roughly 180GB of memory.4 Unsloth's 2-bit and 3-bit dynamic quants push that closer to 100GB.4 A 512GB Mac Studio holds either comfortably, with room left over for the model's context window. That's the hardware story in one sentence: unified memory turned a server-rack problem into a desktop one.

Figure 3
Memory footprint of GLM-5.3-Flash by quantization
memory required (GB)
1002-bit / 3-bit dynamic quant1804-bit quant
Source: MindStudio

How do you actually run GLM-5.3-Flash locally?

The mechanics are more approachable than the parameter count suggests.

  1. Pick your quantization. A 4-bit quant gets you close to full quality at about 180GB. Dynamic 2-bit and 3-bit quants trade a little accuracy for a smaller footprint, near 100GB, which opens the door to more machines.4
  2. Match hardware to that footprint. A 512GB Mac Studio with M5 Ultra handles either quant with headroom.5 Multi-GPU Linux boxes and DGX Spark-class machines work too, at higher power draw and cost. If you're starting smaller, a CPU-only build with 128GB of RAM can run the model, just slower.
  3. Choose a serving stack. Ollama is the fastest path to a working local API for one developer. vLLM and SGLang are built for production throughput and multi-user load, and both officially support GLM-5.3-Flash.36 LM Studio gives you a polished desktop GUI if you'd rather not touch a terminal. Jan is the fully open-source option if extensibility matters more than polish.6
  4. Pull the weights and serve. Download the quantized model from Hugging Face, point your serving tool at it, and expose an OpenAI-compatible endpoint. Most existing tooling that talks to a cloud API can be repointed at localhost with a one-line config change.
  5. Test latency and context before you commit. Local inference typically returns a first token in 50 to 200 milliseconds, against 200 to 800 milliseconds for a round trip to a cloud API.76 Confirm your quant handles the context length you actually need before you build workflows around it.
Figure 4
Choosing a serving stack for GLM-5.3-Flash
Choosing a serving stack for GLM-5.3-Flash
Ease of setupProduction throughputDesktop GUIFully open-source
RecommendedOllamaa single developer's local APIHighLowNoNo
vLLM / SGLangproduction throughput and multi-user loadLowHighNoYes
LM Studioa polished GUI without touching a terminalHighLowYesNo
Janopen-source extensibilityMediumLowYesYes
Ratings are relative across these four options, not absolute measurements.
Source: Remy analysis

For a walkthrough that goes deeper on the CPU-only path specifically, the step-by-step commands for running GLM-5.3-Flash on a 128GB RAM machine are worth reading before you buy anything.

What does local AI actually cost versus API calls?

This is where the ownership argument either holds up or falls apart, and the answer depends entirely on volume.

Independent total-cost-of-ownership modeling breaks it into tiers. At light usage, around 500,000 tokens a day, a consumer local build runs about $6,457 over 12 months, an effective $35.37 per million tokens.7 Hosted APIs at that volume cost $360 to $1,800 over the same period.7 Renting wins, clearly, at light usage. At heavy usage, 50 million tokens a day on enterprise-grade hardware running vLLM, the effective cost drops to $1.69 to $7.15 per million tokens over 36 months, beating OpenAI's roughly $6.90 and Anthropic's roughly $9.86 per million.7 The crossover against proprietary APIs like GPT-4.1 sits around 2 to 3 million tokens a day at a 12-month horizon. Against cheaper hosted open-weight APIs, local doesn't break even until 15 to 20 million tokens a day.7

Figure 5
Effective cost per million tokens: local vs. hosted APIs
Local, light usage (consumer/Ollama)$35.37Anthropic, heavy-usage comparison$9.86OpenAI, light-usage comparison$6.90OpenAI, heavy-usage comparison$6.90Local, heavy usage (enterprise/vLLM)$1.69
Light-tier local build loses to renting; heavy-tier local build (50M tokens/day, 36-month TCO) beats both proprietary APIs on a per-token basis.
Source: SitePoint

A simpler model tells the same story from a different angle. An RTX 4090 running an 8B model at Q4 costs about $59 a month once you include amortized hardware and electricity, and it never beats a $6-a-month budget API tier. It reliably beats flagship-tier pricing around $150 a month, with a break-even point near 20 months.8 For the full arithmetic behind that number, the break-even math on local models versus API pricing walks through it line by line. The pattern across both models is consistent: local ownership rewards sustained, predictable, heavy usage. It punishes light or spiky workloads.

Figure 6
Consumer GPU economics: RTX 4090 running an 8B model at Q4
$59/mo
Amortized monthly cost (hardware + electricity)
$150/mo
Flagship API tier it reliably beats
20 months
Break-even point

When does renting still win?

Local isn't a universal replacement, and treating it as one leads to bad purchasing decisions.

  • Low or unpredictable volume. If your usage is under a few million tokens a day and irregular, a metered API is cheaper and simpler, full stop.78
  • Frontier-only quality needs. If your workload genuinely requires the top proprietary model and no open-weight substitute clears the bar, you're renting regardless of the math.
  • No one to run it. Local deployment has a labor cost the TCO models bake in but that's easy to underestimate in practice. Someone has to manage quantization, updates, and uptime.

The decision isn't "local versus cloud" as a permanent stance. It's a spreadsheet question you should rerun every time your volume shifts meaningfully.

Local models as a governance answer, not just a cost one

There's a second reason this matters beyond the invoice, and it's arguably more urgent.

Seventy-eight percent of AI users already bring their own AI tools to work, sanctioned or not.9 Organizations with high levels of that kind of shadow AI incur an average of $670,000 more per data breach, and 63% of breached organizations had no AI governance policy in place at all.9 That's not a cost problem. It's a control problem, and it's the same gap that pushes employees toward unsanctioned tools in the first place, a pattern covered in depth in the enterprise playbook for running AI models locally. Giving a team a model they can point at real, sanctioned infrastructure, rather than leaving them to find their own way to get work done, is a governance fix disguised as an infrastructure decision. A platform like Remy exists for exactly that gap: turning the ad hoc tools people build for themselves into something the organization can actually see and stand behind, instead of pretending the shadow usage isn't happening.

Figure 7
Shadow AI is a governance problem, not just a cost one
78%
AI users who bring their own AI tools to work
$670,000
Extra breach cost at high-shadow-AI organizations
63%
Breached organizations with no AI governance policy

Quick-start checklist for 2026

If you're deciding whether to run AI locally this year, work through it in this order:

  1. Measure your actual daily token volume for at least a month before buying anything. Spiky, low volume favors renting.
  2. Pick a model sized to your hardware budget. GLM-5.3-Flash at 2-bit to 4-bit quantization fits a 100 to 180GB memory footprint.4
  3. Match hardware to that footprint, not the other way around. A 512GB M5 Ultra Mac Studio, a multi-GPU Linux box, or a high-RAM CPU build all work at different price points.5
  4. Choose your serving stack based on who's using it. Ollama for a solo developer, vLLM or SGLang for production load, LM Studio for a GUI, Jan for open-source flexibility.6
  5. Rerun the break-even math at your real volume before committing capital, and rerun it again if usage changes.

The threshold that decides this isn't philosophical. It's a number on a spreadsheet, and in 2026, for the first time, the hardware and the models exist to make owning that number worth the trouble.

Frequently asked
Questions readers ask
What hardware do I need to run GLM-5.3-Flash locally?

At 4-bit quantization the model needs roughly 180GB of memory, dropping toward 100GB with dynamic 2-bit and 3-bit quants. An Apple Mac Studio with M5 Ultra, which scales to 512GB of unified memory at 1.2TB/s bandwidth, handles either comfortably, and multi-GPU Linux servers or high-RAM CPU builds also work at different price and speed points.

Is it cheaper to run AI models locally or pay for an API in 2026?

It depends on volume. At light usage, around 500,000 tokens a day, local hardware costs roughly $35 per million tokens versus $6.90 to $9.86 for hosted APIs, so renting wins. At heavy sustained usage, around 50 million tokens a day, local drops to $1.69 to $7.15 per million tokens over three years, beating major API providers. The break-even point against proprietary APIs is roughly 2 to 3 million tokens a day.

What is the best tool to serve a local model like GLM-5.3-Flash?

Ollama is the fastest path to a working local API for a solo developer. vLLM and SGLang are built for production-grade multi-user throughput and both officially support GLM-5.3-Flash. LM Studio offers a polished desktop GUI, and Jan is the fully open-source option for teams that want more extensibility.

Why does GLM-5.3-Flash run well on consumer-adjacent hardware despite having 320 billion parameters?

It's a mixture-of-experts model, so only 18 billion of its 320 billion parameters activate per token. Total parameters determine the memory needed to load the model, but active parameters determine compute cost per token, so it behaves like a much smaller, faster model at inference time even though it needs memory sized for a dense model many times larger.

Does running AI locally actually solve shadow AI problems?

It addresses the governance side, not just the cost side. Seventy-eight percent of AI users already bring their own AI tools to work, and organizations with high shadow AI levels incur an average of $670,000 more per data breach, often with no formal AI governance policy in place. Giving teams a sanctioned, owned model to use closes some of that gap.

Sources
  1. 1Zylo's 2026 SaaS Management Index Finds AI-Native App Adoption Is Surging, with ChatGPT Now the Most Expensed AppZylo
  2. 2GLM-5.3-Flash: Frontier Intelligence, Flash CostZ.ai
  3. 3zai-org/GLM-5.3-Flash (model card)Hugging Face
  4. 4Run GLM 5.3 Flash Locally: VRAM, Quantization, and Hardware NeedsMindStudio
  5. 5Apple introduces new Mac Studio with M5 Max and M5 UltraApple Newsroom
  6. 6The Definitive Guide to Local LLMs in 2026: Privacy, Tools, & HardwareSitePoint
  7. 7Local LLMs vs Cloud APIs: 2026 Total Cost of Ownership AnalysisSitePoint
  8. 8Running LLMs locally vs paying for an API: the actual mathflaviocopes.com
  9. 9Shadow AI Statistics 2026: 38 Verified Stats With Sourcesdope.security
Portrait of Lena Ortiz
Lena Ortiz
Software Ownership
Lena makes the case for owning the software your company runs on.
More from Lena Ortiz
© 2026 The Official Remy BlogDrafted by AI authors, reviewed by human editors.