This paper introduces the Local Sovereignty Framework (LSF), a method for deciding which enterprise AI workloads belong on hardware you own and how to move them there. It is written for the leaders who have noticed that renting intelligence from a handful of external interfaces is a dependency, not a strategy, and who want the control back.
The framework, the seven exhibits, and the terms Sovereignty Dividend and Edge Sufficiency are original to this work, published under the license on the back cover. It builds on quantization and local-agent research cited in the notes, and on the research agent Zorp [5]. Every quantitative claim is cited to a named source or labeled directional, without exception.
For three years the industry agreed on a convenient story. The best intelligence lived in a few enormous models, reachable only through someone else's interface, and the job of the enterprise was to connect to it and pay by the token. The story was true for about as long as it took the hardware to catch up. It no longer is.
A model good enough for the bulk of enterprise work now runs on a laptop. Quantization shrank the memory a capable model needs until it fit on the machine already on the desk, and the gap to the best closed model narrowed to almost nothing. Once that happened, the question stopped being whether a local model is good enough and became a different question entirely. Who do you want holding the weights, the data, and the off switch.
This paper is our answer. It is not an argument against the cloud. It is an argument for choice, and for noticing that choice now exists. The leaders who see it first will run AI the way they run anything else they depend on. On terms they set, on hardware they control, with no one able to change the deal underneath them. The numbers here are cited or labeled directional, without exception.
The cost of a capable model has collapsed, and the model can now run on your own hardware. That single fact changes who holds the leverage in an AI program. Local is not a cheaper way to do the same thing. It is a different ownership position. Five findings follow.
The hardware question is settled. In 4-bit, a 13 billion parameter model runs at 30 tokens per second on an 8GB laptop GPU [1]. The machine already on the desk is now a capable inference engine, not a thin client to someone else's.
Going local inverts the leverage. When you hold the weights, no vendor can change the price, deprecate the model, throttle the rate, or see your data. We call the shift the Control Inversion, and it is the real reason to run locally, not the token bill.
The Sovereignty Dividend compounds. Every query answered on hardware you own is one that is not rented, not logged off-site, and not dependent on a model outside your control. Cost, privacy, and permanence accrue together, and they accrue every day.
Open-weight has crossed Edge Sufficiency. The capability gap to the best closed model fell from 8.04 percent to 1.70 percent in thirteen months [4]. For the majority of enterprise tasks, the local model is already good enough, and the frontier premium no longer justifies the dependency.
The migration is a quarter, not a platform purchase. A team of four to five moves a workload from a rented API to owned inference in roughly fifteen weeks, directionally, and the evidence layer that keeps those local answers defensible already exists.
For a decade, running a capable model locally was a thought experiment. The weights were too large for any machine a person actually owned. Quantization changed the arithmetic. In 4-bit precision each parameter costs roughly half a byte, so a 7 billion parameter model needs about 4GB and a 13 billion parameter model about 7GB, both comfortably inside the 16GB of memory in a mainstream laptop.
Capability followed. In 4-bit, a 13 billion parameter model answers at 30 tokens per second on an 8GB laptop GPU [1], and quantized open models reach within a few points of a frontier chatbot on standard benchmarks [2]. The constraint that justified renting intelligence has quietly lifted.
Edge Sufficiency. The point at which a locally runnable model is good enough for a task that the frontier model's marginal advantage no longer justifies depending on it. For most enterprise work, that point has passed.
Exhibit 1 shows the shape of it. The models that fit on hardware you own already cover the capability most enterprise work requires. The frontier still leads, but that lead now sits outside the memory budget of the machine on the desk, and for most tasks it is a lead you do not need.
Running a model you rent means accepting a set of dependencies that have nothing to do with how good the model is. Going local does not make the answers better. It removes the dependencies. Each one below is leverage that currently sits with the vendor and moves to you the moment the model runs on your hardware.
The per-token rate can rise. A workload priced into a budget today can cost more next quarter with no change on your side and no recourse.
Throughput is rationed by quota. A local model answers as fast as your hardware allows, with no ceiling imposed from outside.
Every prompt and document leaves your environment to be answered, where it may be logged, retained, or used to train a future model.
The model you validated can be retired or silently changed, altering behavior you tested and signed off on, on the vendor's schedule.
An outage, a geography block, or a policy change on the far side takes your capability with it, at a time you do not choose.
Prompts, tooling, and workflows accrete to one vendor's quirks, so the cost of ever leaving rises quietly with every month you stay.
The local stack is not a smaller cloud. It is the same capability relocated inside a boundary you own, where the model, the index, the orchestration, and the trust layer all run on hardware you control. The only thing that ever crosses the boundary is the open-weight model itself, downloaded once.
| Component | What it replaces | Where it runs | Cost shape |
|---|---|---|---|
| Local model | A metered inference API | Your CPU or GPU | One-time, then free |
| Local retrieval | A hosted vector service | Your disk and memory | One-time, then free |
| Orchestration | A managed agent runtime | Your process | One-time, then free |
| Evidence layer | Trust you could not verify | Alongside the model | One-time build |
| Model weights | A model you did not hold | At rest on your storage | Free download, yours |
A rented model charges for every answer, forever, and the rate is set by someone else. A local model is paid for once, in hardware the organization largely already owns, and every answer after that is free at the margin. A quantized open model reaches 99.3 percent of a frontier chatbot's quality on one benchmark [2], so the quality you give up to stop paying rent is, for most work, close to nothing.
But the token bill is the smallest part of the case. The larger return is everything in Section 02 that stops being someone else's decision. When the price cannot change, the model cannot be withdrawn, and the data cannot leave, the value is not a lower invoice. It is a risk that simply disappears from the register.
Sovereignty Dividend. The compounding return of running inference on hardware you own. No per-query rent, no data leaving, no model that can change under you. It is paid every day the model runs, and it grows with volume.
Cost of renting. A per-token rate times every query, every day, set and changed by someone else. It scales with success, so the more the program works, the larger the bill and the deeper the dependency.
Cost of owning. A one-time setup on hardware the organization largely holds already. The marginal cost of the next answer is the electricity to compute it, and nothing crosses a boundary to produce it.
The dividend. Past a modest volume, local is cheaper. But the return that compounds is control, paid every day in risk that is no longer on the table rather than in a smaller invoice.
The migration climbs in three steps, and each step is a decision gate, not a status check. You do not move an estate at once. You move one workload, prove it runs on your own hardware at the quality the task needs, measured under a fixed resource budget [3], and then let the pattern repeat. Each step raises how much of the stack you own.
Can the prompt and its documents leave your environment? When they cannot, the model must be local, and the decision is already made.
Does the task genuinely need the frontier, or is a local model already past Edge Sufficiency for it? Most work is the latter.
High, steady volume favors owned inference, where the marginal cost is near zero. Spiky or rare volume can stay rented.
Must the behavior be permanent, auditable, and immune to a vendor's change? If so, only holding the weights delivers it.
The closing argument. For a few years the only way to use frontier-grade AI was to rent it, and renting felt like the whole market. It was a phase, not the shape of things. The model now runs on the machine you own, the gap to the frontier has nearly closed, and the control that comes with holding the weights is a durable advantage renting can never offer. The future is local for the same reason every mature capability ends up in-house. You cannot build on what you do not control.
Answer honestly for your most important AI workload. Count the YES answers, then read your level on the band below and the gaps it points to. Designed to be completed in pen.
| No. | Tests | Question | Yes | No |
|---|---|---|---|---|
| 1 | DATA | Do you know which AI workloads send sensitive data off-site to be answered? | ||
| 2 | DATA | Could your most sensitive workload run with nothing leaving your environment? | ||
| 3 | COST | Is your inference cost fixed, rather than a rate a vendor can raise without your consent? | ||
| 4 | CAP | Have you tested a quantized open model on your actual workload, not just a benchmark? | ||
| 5 | HW | Do you already own hardware that could serve your top workload locally? | ||
| 6 | PERM | If your current model were deprecated tomorrow, could you keep running unchanged? | ||
| 7 | PERM | Can you reproduce an answer months later on the exact model that produced it? | ||
| 8 | VOL | Do you know the volume at which local becomes cheaper than renting for that workload? | ||
| 9 | RULE | Is local first, rent the exception a stated default, rather than renting by habit? | ||
| 10 | TRUST | Do your local answers carry an evidence trail that survives a challenge? |
This paper is a conceptual contribution. The LSF, the four-quadrant decision model, the seven exhibits, and the terms Sovereignty Dividend and Edge Sufficiency are original to this work. The Control Inversion names a shift in leverage that local inference creates. The quantization results it rests on, which let a thirteen billion parameter model run at interactive speed on a laptop, are drawn from the cited research and labeled where used [1] [2]. The local-agent evaluation method behind the roadmap is from BudgetBench [3], and the research agent Zorp [5] shows the pattern running on owned hardware.
Quantitative claims follow one discipline throughout, made visible in every exhibit. Solid fills carry cited data. Hatched fills and dashed guides labeled directional exist to make the shape of an argument legible, not to report measurements. The capability-gap figures are cited to a named index [4]. No statistic in this paper is attributed to a source that cannot be independently verified.
Arjun Jaggi and Aditya Karnam Gururaj Rao write and advise at the intersection of enterprise AI strategy, governance, and local open-weight deployment. Their published body of work, spanning original frameworks, concept papers, and books, is read by executive teams navigating AI adoption and is available in full, without a paywall, at arjunjaggi.com.