Software Ownership

Exfiltrate Your Weights: The Push for Mission-Critical Local AI

Enterprises that built their AI stack on rented intelligence are learning the rent can rise, the terms can change, and the landlord can lock the door without notice. A growing number are downloading the weights instead.

At a glance
  1. 01Enterprise on-premises AI inference jumped from 12% in 2023 to over half of workloads today.
  2. 02Small open-weight models can break even against commercial API costs in less than one month.
  3. 03True AI sovereignty is elusive, as local deployments still rely heavily on Nvidia hardware.
  4. 04Despite the local push, closed proprietary models still dominate enterprise AI budgets.
Illustration of a stack of data plates representing AI model weights being physically transferred from a large storage block into a smaller self-contained local server chassis.
Illustration generated by Remy for this story.

Running AI models locally means installing an open-weight model like Llama, Qwen, or GLM on infrastructure you own or lease outright, instead of routing every request through a vendor's API. For enterprises, the trade is control and cost predictability in exchange for the operational work of running your own inference stack. In 2026, a widening set of banks, hospitals, and governments are deciding that trade is worth making.

The API honeymoon is ending

For three years, the default enterprise AI architecture was simple: pick a frontier model from OpenAI, Anthropic, or Google, wire it up over an API, and let someone else worry about GPUs. That model still dominates spending. Anthropic, OpenAI, and Google together capture 88% of enterprise LLM API usage, and Anthropic alone now earns 40% of that spend, up from 24% a year earlier.1

Figure 1
Who controls enterprise LLM spend
88%
Enterprise LLM API usage captured by Anthropic, OpenAI, and Google combined
40%
Anthropic's share of enterprise LLM spend, up from 24% a year earlier

But a shadow question now hangs over that architecture: what happens when the vendor changes the rules, raises the price, or gets cut off entirely? In early 2026 that question stopped being hypothetical, and it's reshaping how technical leaders think about the model layer.

Figure 2
CIOs expect AI budgets to keep climbing
75%
Expected enterprise AI budget growth over the next year, per a16z's CIO survey

What does 'exfiltrating your weights' actually mean?

The phrase sounds dramatic. The mechanics are mundane. Open-weight models publish their parameters so anyone can download them and run inference on hardware they control, using serving engines like vLLM to handle the request queue, batching, and GPU scheduling that a commercial API normally hides from you. Once the weights sit on infrastructure you control, the model stops being a metered service and becomes a capital asset with a fixed cost.

That shift from renting to owning is the same logic driving software teams to self-host their AI stack instead of stacking API subscriptions. It's also why treating AI infrastructure as owned property, tracked and budgeted like any other capital asset, has become its own discipline. Some teams now manage that infrastructure the way they'd manage any owned software asset, with tools like Remy built specifically to track what a company owns instead of rents.

The case for local: sovereignty, cost, and compliance

Three motivations show up over and over in enterprise deployments. None of them are new, but they've gotten sharper.

Figure 3
On-premises and edge inference, share of enterprise AI workloads
share of enterprise AI inference workloads (%)
0%50%100%55%20232026
Year
Source: DreamFactory
  • Data privacy and IP protection. Sending proprietary code, patient records, or financial data to a third-party API creates a permanent exposure surface. Local deployment keeps that data inside infrastructure the enterprise controls end to end.
  • Cost predictability at volume. API pricing scales with usage in a way that can blow up a budget overnight. Owned infrastructure converts a variable cost into a fixed one, which finance teams tend to prefer once volume justifies the upfront spend.
  • Vendor independence and compliance. Regulated industries need auditability a black-box API can't offer. Owning the stack means no single vendor's policy change can take down production.
Figure 4
Local weights vs. rented API: the trade-off
Local weights vs. rented API: the trade-off
Control over infrastructureCost structureCompliance auditabilitySetup & maintenance burdenTime to cost break-even
Local / self-hosted weightshigh-volume, compliance-heavy, or cost-sensitive workloadsHighFixed capex + infra, converts variable cost to fixedHighHigh0.3 to 69.3 months, depending on model size
Rented commercial APIfrontier-capability tasks and unpredictable or low volumeLowVariable, scales with usageLowLowN/A, pay-per-use
Ratings are relative across these two options, not absolute. Reflects the trade-offs described in the article rather than a formal vendor scorecard.
Source: Remy analysis

Siemens built this case in practice. The company stood up its own GPU fleet, running NVIDIA L40S and H200 chips, and built an open-source serving stack around vLLM, Kong, and Prometheus/Grafana to run Llama, Qwen, DeepSeek, and Mistral for its internal developer base.2 Siemens' own framing is blunt: AI sovereignty means owning the entire stack, and vendor independence comes from open weights on your own infrastructure, full stop.2

Figure 5
LLM market deployment: on-premise vs cloud
52%On-premise
On-premise52%
Cloud/hosted48%
Source: DreamFactory

a16z's 2025 survey of 100 enterprise CIOs found the same pattern at scale: open-source model adoption concentrates at larger enterprises, driven by a preference for on-prem deployment tied to data security, compliance, and the ability to fine-tune models internally. Those same CIOs expect their AI budgets to grow roughly 75% over the next year.3

The Pentagon-Anthropic wake-up call

If one event turned vendor dependency from an abstract risk into a board-level one, it happened in March 2026. The Department of Defense applied its 'supply chain risk' designation to Anthropic, the first time that classification, normally reserved for foreign adversaries, was applied to a U.S. company.4 Defense contractors running Claude had no graceful migration path. Within hours of the ban becoming known, OpenAI signed a replacement deal with the Pentagon on terms Sam Altman himself called 'definitely rushed.'4

That episode is the clearest illustration yet of a risk that has nothing to do with model quality: when your production AI depends entirely on one proprietary vendor, a policy decision you had no say in can take your capability offline overnight. It's the same structural fragility covered in why owned infrastructure doesn't go down the way rented endpoints do, just applied to defense procurement instead of an outage.

Figure 6
Break-even timeline vs commercial API, by local model size
Large (200B+)69Medium (70B-120B)34Small (24B-32B)0
Carnegie Mellon's study reports ranges for each size class; bars show the upper bound (small models can break even almost immediately, at 0.3 months).

When does running AI locally actually pay off?

The 'local always wins' story doesn't survive contact with the actual math. A Carnegie Mellon study of 54 deployment scenarios is the best evidence available on where the line falls.5

  • Small models (24B-32B parameters) can break even against premium commercial APIs in as little as 0.3 months.5
  • Medium models (70B-120B parameters) break even in a range of roughly 2.3 to 34 months, depending on usage volume and which API you're comparing against.5
  • Large models (200B+ parameters) can take up to 69.3 months, nearly six years, to break even against a cost-competitive commercial API like Gemini 2.5 Pro.5

That spread matters. The sovereignty pitch is strongest at the small-to-medium end, where hardware and electricity costs are modest and inference volume is high enough to amortize them fast, a threshold laid out in more detail in the real break-even math on local models versus API cost. At the frontier end, where a single large model needs a rack of expensive GPUs to serve, the math often still favors renting, at least until usage climbs into the millions of tokens per day.

Figure 7
Diagnostic accuracy: open-weight vs closed model
Llama 3.1 (open-weight, local)70%GPT-464%

The broader data backs this up directionally. Enterprise AI inference performed on-premises or at the edge jumped from 12% in 2023 to 55% today, and on-premise deployments already hold 51.85% of the LLM market, led by banks, hospitals, and government agencies prioritizing sovereignty and latency control.6

Who's actually doing this

Beyond Siemens, the clearest proof points sit in regulated industries with the most to lose from a data leak.

Healthcare is a standout case. A Harvard Medical School study found that Meta's open-weight Llama 3.1 matched or outperformed GPT-4 on complex diagnostic cases, hitting 70% accuracy on established cases versus GPT-4's 64%.7 The study's authors put the implication plainly: institutions may be able to deploy high-performing custom models that run locally without sacrificing data privacy or flexibility.7 For a hospital bound by HIPAA, that's not a marginal improvement. It's the difference between using AI on patient data at all and not.

Figure 8
Enterprise open-source LLM market share, year over year
share of enterprise LLM spend (%)
0%10%20%11%20242025
Year

The cautionary tale that kicked off this whole shift is Samsung's. In 2023, Samsung banned employee use of ChatGPT and other generative AI tools company-wide after engineers leaked sensitive source code and internal meeting notes into the chatbot on three separate occasions within 20 days.8 That incident is now the textbook example of shadow AI turning into a governance crisis, and it's a big reason enterprises now prefer locking AI usage down to infrastructure they can audit end to end rather than trusting employees to self-police on external APIs.

Sovereignty's fine print: you don't fully escape dependency

Here's the part vendors selling 'sovereign AI' packages don't put on the slide: owning your weights doesn't mean owning your supply chain. Stanford HAI's review of commercial AI sovereignty offerings finds that even enterprises and governments running open-weight models locally remain dependent on Nvidia chips, TSMC manufacturing, and a narrow set of foreign cloud and hardware providers.9 In practice, sovereignty recalibrates dependency rather than eliminating it, a dynamic critics have started calling 'sovereignty washing.'9

Figure 9
Nvidia's stake in the sovereignty movement
$30B
Nvidia revenue from sovereign AI today
14%
Share of Nvidia's total revenue from sovereign AI
$200B
Projected future sovereign AI revenue for Nvidia
Source: Stanford HAI

Nvidia is, unsurprisingly, happy to sell into this. Its 'AI factory' sovereign infrastructure program, built around open-weight Nemotron models, has been deployed or supported in at least 25 countries, and sovereign AI already accounts for roughly $30 billion, or 14%, of Nvidia's revenue, a figure projected to reach $200 billion.9 The company that makes the chips underneath every 'independent' local deployment is also the biggest financial winner of the independence movement.

The enterprise verdict: hybrid, not all-in

Despite everything above, most enterprises haven't walked away from closed APIs. Menlo Ventures' late-2025 survey found enterprise open-source LLM market share actually fell from 19% to 11% year over year, even as sovereignty concerns grew louder.110 Closed models still command the large majority of a typical enterprise's roughly $7 million annual LLM budget, out of a total enterprise generative AI market worth $37 billion in 2025.110

The honest read isn't that local AI is winning. It's that enterprises are building two lanes at once, the same quiet split described in why the real shift in enterprise AI is happening off the cloud: a frontier-capability lane on rented APIs for the hardest reasoning tasks, and a growing local lane for the high-volume, compliance-heavy, or cost-sensitive workloads where owning the weights already pencils out. Boards that used to ask which API to standardize on are now asking a second question: which workloads should never depend on someone else's policy decision. That second question is what's pushing weights out of the cloud and onto infrastructure enterprises actually own.

Frequently asked
Questions readers ask
Is running AI models locally cheaper than using an API for enterprises?

It depends on model size and volume. A Carnegie Mellon study found small open-weight models (24B-32B parameters) can break even against premium APIs in as little as 0.3 months, medium models take 2.3 to 34 months, and large models (200B+) can take up to 69.3 months, so local is cheapest at the small-to-medium end and often still loses to APIs at the frontier end.5

Why are banks and hospitals moving AI workloads on-premises?

Data privacy, compliance, and audit requirements are the main drivers. Enterprise AI inference performed on-premises or at the edge rose from 12% in 2023 to 55% today, led by banks, hospitals, and government agencies prioritizing sovereignty and latency control.6

What was the Pentagon-Anthropic incident and why does it matter for enterprise AI strategy?

In March 2026, the Department of Defense classified Anthropic as a 'supply chain risk,' the first time that designation was applied to a U.S. company, immediately barring defense contractors from using Claude with no migration path.4 It's now the clearest example of how depending on one proprietary API vendor exposes a business to non-technical, policy-driven outages.

Have enterprises actually shifted their AI spending toward open-weight models?

Not broadly. Menlo Ventures found enterprise open-source LLM market share fell from 19% to 11% year over year, with Anthropic, OpenAI, and Google still capturing 88% of enterprise LLM API spend.1 Adoption of local, open-weight models is real but concentrated in specific compliance-heavy or high-volume use cases, not a wholesale exodus from closed APIs.

Does running open-weight models locally eliminate vendor dependency entirely?

No. Stanford HAI's research shows that enterprises running open-weight models on their own infrastructure remain dependent on Nvidia chips, TSMC manufacturing, and a narrow set of cloud and hardware providers, so local deployment reduces but doesn't eliminate dependency on a small group of suppliers.9

Sources
  1. 12025: The State of Generative AI in the EnterpriseMenlo Ventures
  2. 2Our Sovereign AI Journey: Building a Self-Contained, Sustainable, and Cost-Effective LLM PlatformSiemens Blog
  3. 3How 100 Enterprise CIOs Are Building and Buying Gen AI in 2025Andreessen Horowitz (a16z)
  4. 4Sovereign AI Dependency: The Pentagon-Anthropic Concentration TrapCloud Security Alliance
  5. 5A Cost-Benefit Analysis of On-Premise Large Language Model Deployment: Breaking Even with Commercial LLM ServicesarXiv (Carnegie Mellon University)
  6. 628 On-Premise LLM Deployment Statistics Every Enterprise Should Know in 2026DreamFactory
  7. 7Can open-source AI transform hospital care without breaking the bank?Digital Health Insights (dhinsights.org)
  8. 8Samsung Bans Generative AI Use by Staff After ChatGPT LeakBloomberg
  9. 9The Commercial Landscape of AI Sovereignty OfferingsStanford HAI
  10. 10The Open-Source AI Boom May Be a Mirage. Why Affordable AI May Be a Myth.Yahoo Finance
Portrait of Priya Nair
Priya Nair
AI Tooling
Priya covers the daily churn of AI agents, coding tools, and what actually ships.
More from Priya Nair
© 2026 The Official Remy BlogDrafted by AI authors, reviewed by human editors.