↓ Download PDF
RENTED · RECEDING LOCAL · IN YOUR CONTROL THE WEIGHTS, THE DATA, THE DECISION
Arjun Jaggi
Enterprise AI Research · White Paper No. 03
October 2026
The Future Is
Local
Open-weight models, your own hardware, and the end of renting intelligence
Authors
Arjun Jaggi
Aditya Karnam Gururaj Rao
Framework
LSF, Version 1.0
Discipline
Enterprise AI Sovereignty
Readership
Boards · C-suite · Infrastructure
The Future Is LocalArjun Jaggi & Aditya Karnam Gururaj Rao
Contents
·
Foreword
A letter to the reader
03
·
Executive Summary
Five findings for the board
04
01
The Hardware Question Is Settled
Why a capable model now fits on the machine you own
05
02
The Control Inversion
The six dependencies that going local removes
06
03
The Local Stack
Everything the workload needs, inside the box
07
04
The Economics
The Sovereignty Dividend and the cost of renting
08
05
The Migration Roadmap
From rented API to owned inference in one quarter
10
06
The Decision Framework
Which workloads belong on your own hardware
11
·
The LSF on a Page
The complete framework, one spread, desk-ready
12
·
Self-Assessment Scorecard
Ten questions place your estate on the sovereignty curve
13
·
Methodology, Notes, and the Authors
14

About This Paper

This paper introduces the Local Sovereignty Framework (LSF), a method for deciding which enterprise AI workloads belong on hardware you own and how to move them there. It is written for the leaders who have noticed that renting intelligence from a handful of external interfaces is a dependency, not a strategy, and who want the control back.

The framework, the seven exhibits, and the terms Sovereignty Dividend and Edge Sufficiency are original to this work, published under the license on the back cover. It builds on quantization and local-agent research cited in the notes, and on the research agent Zorp [5]. Every quantitative claim is cited to a named source or labeled directional, without exception.

How to read the exhibits
Solid fill, cited data
Hatched fill, directional illustration
Red, a dependency or egress path
Arjun Jaggi · Aditya Karnam Gururaj Rao02
The Future Is LocalForeword
Foreword
A letter to the reader

For three years the industry agreed on a convenient story. The best intelligence lived in a few enormous models, reachable only through someone else's interface, and the job of the enterprise was to connect to it and pay by the token. The story was true for about as long as it took the hardware to catch up. It no longer is.

A model good enough for the bulk of enterprise work now runs on a laptop. Quantization shrank the memory a capable model needs until it fit on the machine already on the desk, and the gap to the best closed model narrowed to almost nothing. Once that happened, the question stopped being whether a local model is good enough and became a different question entirely. Who do you want holding the weights, the data, and the off switch.

This paper is our answer. It is not an argument against the cloud. It is an argument for choice, and for noticing that choice now exists. The leaders who see it first will run AI the way they run anything else they depend on. On terms they set, on hardware they control, with no one able to change the deal underneath them. The numbers here are cited or labeled directional, without exception.

Arjun Jaggi · Aditya Karnam Gururaj Rao
Enterprise AI Research · arjunjaggi.com

At a Glance

30/s
Tokens per second for a 13B model on an 8GB laptop GPU, in 4-bit [1]
99.3%
Quality of a quantized open model vs a frontier chatbot, one benchmark [2]
1.70%
Open vs closed capability gap, down from 8.04% a year earlier [4]
15 wks
From a rented API to owned inference, directional
Arjun Jaggi · Aditya Karnam Gururaj Rao03
The Future Is LocalExecutive Summary
Executive Summary
Five findings for the board

The cost of a capable model has collapsed, and the model can now run on your own hardware. That single fact changes who holds the leverage in an AI program. Local is not a cheaper way to do the same thing. It is a different ownership position. Five findings follow.

1

The hardware question is settled. In 4-bit, a 13 billion parameter model runs at 30 tokens per second on an 8GB laptop GPU [1]. The machine already on the desk is now a capable inference engine, not a thin client to someone else's.

2

Going local inverts the leverage. When you hold the weights, no vendor can change the price, deprecate the model, throttle the rate, or see your data. We call the shift the Control Inversion, and it is the real reason to run locally, not the token bill.

3

The Sovereignty Dividend compounds. Every query answered on hardware you own is one that is not rented, not logged off-site, and not dependent on a model outside your control. Cost, privacy, and permanence accrue together, and they accrue every day.

4

Open-weight has crossed Edge Sufficiency. The capability gap to the best closed model fell from 8.04 percent to 1.70 percent in thirteen months [4]. For the majority of enterprise tasks, the local model is already good enough, and the frontier premium no longer justifies the dependency.

5

The migration is a quarter, not a platform purchase. A team of four to five moves a workload from a rented API to owned inference in roughly fifteen weeks, directionally, and the evidence layer that keeps those local answers defensible already exists.

Arjun Jaggi · Aditya Karnam Gururaj Rao04
The Future Is LocalSection 01
01
Section 01 · The Hardware Question
The model you need already fits on the machine you own

For a decade, running a capable model locally was a thought experiment. The weights were too large for any machine a person actually owned. Quantization changed the arithmetic. In 4-bit precision each parameter costs roughly half a byte, so a 7 billion parameter model needs about 4GB and a 13 billion parameter model about 7GB, both comfortably inside the 16GB of memory in a mainstream laptop.

Capability followed. In 4-bit, a 13 billion parameter model answers at 30 tokens per second on an 8GB laptop GPU [1], and quantized open models reach within a few points of a frontier chatbot on standard benchmarks [2]. The constraint that justified renting intelligence has quietly lifted.

Framework Term

Edge Sufficiency. The point at which a locally runnable model is good enough for a task that the frontier model's marginal advantage no longer justifies depending on it. For most enterprise work, that point has passed.

EXHIBIT 1
The models that fit on a laptop already cover most of the capability
Model class by 4-bit memory footprint and relative capability
16 GB FITS ON A LAPTOP 7B 13B30 tok/s 30B 70B MEMORY FOOTPRINT IN 4-BIT, GB → CAPABILITY →
Source. Capability point for the 13B model and the 30 tokens per second figure, AWQ [1]; footprints derived from 4-bit quantization arithmetic. Relative capability is indicative.

Exhibit 1 shows the shape of it. The models that fit on hardware you own already cover the capability most enterprise work requires. The frontier still leads, but that lead now sits outside the memory budget of the machine on the desk, and for most tasks it is a lead you do not need.

Arjun Jaggi · Aditya Karnam Gururaj Rao05
The Future Is LocalSection 02
02
Section 02 · The Inversion
Local does not add a feature, it removes six dependencies

Running a model you rent means accepting a set of dependencies that have nothing to do with how good the model is. Going local does not make the answers better. It removes the dependencies. Each one below is leverage that currently sits with the vendor and moves to you the moment the model runs on your hardware.

I
Price Risk

The per-token rate can rise. A workload priced into a budget today can cost more next quarter with no change on your side and no recourse.

II
Rate Limits

Throughput is rationed by quota. A local model answers as fast as your hardware allows, with no ceiling imposed from outside.

III
Data Egress

Every prompt and document leaves your environment to be answered, where it may be logged, retained, or used to train a future model.

IV
Model Deprecation

The model you validated can be retired or silently changed, altering behavior you tested and signed off on, on the vendor's schedule.

V
Availability

An outage, a geography block, or a policy change on the far side takes your capability with it, at a time you do not choose.

VI
Lock-In

Prompts, tooling, and workflows accrete to one vendor's quirks, so the cost of ever leaving rises quietly with every month you stay.

EXHIBIT 2
In the hosted path your data leaves to be answered; in the local path nothing crosses
YOUR BOUNDARY INSIDE YOUR ENVIRONMENT OUTSIDE · VENDOR CONTROLLED HOSTED DATA + PROMPTyours, sensitive your data leaves VENDOR MODELlogged, maybe trained on answer returns, metered LOCAL DATA + PROMPTstays put LOCAL MODELon your hardware ANSWER nothing crosses the boundary
Source. LSF data-flow model, original to this work. The red path marks data that leaves your environment; the local path keeps the prompt, the document, and the answer on hardware you control.
Arjun Jaggi · Aditya Karnam Gururaj Rao06
The Future Is LocalSection 03
03
Section 03 · The Architecture
Everything the workload needs runs inside the box you control

The local stack is not a smaller cloud. It is the same capability relocated inside a boundary you own, where the model, the index, the orchestration, and the trust layer all run on hardware you control. The only thing that ever crosses the boundary is the open-weight model itself, downloaded once.

EXHIBIT 3
The entire stack sits inside one boundary, and only the model crosses it, once
YOUR MACHINE ORCHESTRATIONthe agent loop and tools, run locally LOCAL MODELquantized open weights, the inference engine LOCAL RETRIEVALyour documents, indexed on your disk EVIDENCE LAYERpre-registration and provenance, see White Paper No. 02 STORAGE · the weights and the data, at rest on hardware you own OPEN-WEIGHTMODELpublic, free to download ONCE
Source. LSF local stack, original to this work. The evidence layer is the Defensible Discovery Framework of White Paper No. 02; the model download is the only transfer across the boundary, and it happens one time.
ComponentWhat it replacesWhere it runsCost shape
Local modelA metered inference APIYour CPU or GPUOne-time, then free
Local retrievalA hosted vector serviceYour disk and memoryOne-time, then free
OrchestrationA managed agent runtimeYour processOne-time, then free
Evidence layerTrust you could not verifyAlongside the modelOne-time build
Model weightsA model you did not holdAt rest on your storageFree download, yours
Arjun Jaggi · Aditya Karnam Gururaj Rao07
The Future Is LocalSection 04
04
Section 04 · The Economics
Rented intelligence recurs; owned intelligence is paid once

A rented model charges for every answer, forever, and the rate is set by someone else. A local model is paid for once, in hardware the organization largely already owns, and every answer after that is free at the margin. A quantized open model reaches 99.3 percent of a frontier chatbot's quality on one benchmark [2], so the quality you give up to stop paying rent is, for most work, close to nothing.

But the token bill is the smallest part of the case. The larger return is everything in Section 02 that stops being someone else's decision. When the price cannot change, the model cannot be withdrawn, and the data cannot leave, the value is not a lower invoice. It is a risk that simply disappears from the register.

99.3%
The quality a quantized open model reaches against a frontier chatbot on one benchmark, at zero marginal cost once it runs locally [2]
EXHIBIT 4
On every dimension that is not raw capability, control moves from the vendor to you
Degree of control over each dimension, hosted versus local, directional
LOWFULL DEGREE OF CONTROL IN YOUR HANDS Cost predictability Data control Availability Model permanence Throughput ceiling Hosted Local
Source. Directional illustration, original to this work. Positions reflect the structural difference between renting and owning inference, not a measured index.
Arjun Jaggi · Aditya Karnam Gururaj Rao08
The Future Is LocalSection 04 · Continued
EXHIBIT 5
The memory a capable model needs fell below the laptop line as precision dropped
Memory footprint of a 13B model by numeric precision, derived from quantization arithmetic
16 GB LAPTOP LIMIT FITS ON A LAPTOP 0 16 30 GB 26 GB 13 GB 6.5 GB FP16 8-BIT 4-BIT
Source. Footprints derived from precision arithmetic (bytes per parameter times 13 billion). The 4-bit footprint is what makes the 13B model in AWQ [1] run on an 8GB laptop GPU.
Framework Term

Sovereignty Dividend. The compounding return of running inference on hardware you own. No per-query rent, no data leaving, no model that can change under you. It is paid every day the model runs, and it grows with volume.

The Return on Control

Cost of renting. A per-token rate times every query, every day, set and changed by someone else. It scales with success, so the more the program works, the larger the bill and the deeper the dependency.

Cost of owning. A one-time setup on hardware the organization largely holds already. The marginal cost of the next answer is the electricity to compute it, and nothing crosses a boundary to produce it.

The dividend. Past a modest volume, local is cheaper. But the return that compounds is control, paid every day in risk that is no longer on the table rather than in a smaller invoice.

Arjun Jaggi · Aditya Karnam Gururaj Rao09
The Future Is LocalSection 05
05
Section 05 · The Roadmap
From a rented API to owned inference in one quarter

The migration climbs in three steps, and each step is a decision gate, not a status check. You do not move an estate at once. You move one workload, prove it runs on your own hardware at the quality the task needs, measured under a fixed resource budget [3], and then let the pattern repeat. Each step raises how much of the stack you own.

EXHIBIT 6
Each phase raises how much of the stack you own, gated by a measurable exit
SOVEREIGNTY → PHASE 1 · PILOTone workload, local model, weeks 1 to 5 GATE 1 PHASE 2 · HARDENevidence layer, throughput, weeks 6 to 11 GATE 2 PHASE 3 · OWNestate-wide, owned by default, weeks 12+ TIME, AND SHARE OF THE STACK YOU OWN →
Source. LSF migration model, original to this work. Each gate is a measurable exit condition; a program that cannot pass Gate 1 on one workload does not proceed to the estate.
1
Weeks 1-5
Pilot on your hardware
  • Pick one data-sensitive workload
  • Stand up a quantized open model
  • Match the task quality bar
  • Gate 1. The workload runs locally at acceptable quality, nothing leaves
2
Weeks 6-11
Harden and make defensible
  • Add the evidence layer
  • Tune throughput for real load
  • Run a cutover and a fallback drill
  • Gate 2. The local path meets the service bar under real volume
3
Weeks 12+
Own it by default
  • Extend to every fitting workload
  • Make local the default, rent the exception
  • Set a model-refresh cadence
  • Success. New workloads start local unless a reason says otherwise
Arjun Jaggi · Aditya Karnam Gururaj Rao10
The Future Is LocalSection 06
06
Section 06 · The Decision
Four variables decide whether a workload belongs on your own hardware
1
Data Sensitivity

Can the prompt and its documents leave your environment? When they cannot, the model must be local, and the decision is already made.

2
Capability Ceiling

Does the task genuinely need the frontier, or is a local model already past Edge Sufficiency for it? Most work is the latter.

3
Volume

High, steady volume favors owned inference, where the marginal cost is near zero. Spiky or rare volume can stay rented.

4
Control Requirement

Must the behavior be permanent, auditable, and immune to a vendor's change? If so, only holding the weights delivers it.

EXHIBIT 7
Most enterprise work lands in a quadrant where local is the right call
LOCAL BY DEFAULTsensitive, routine capability LOCAL, RENT THE EDGEsensitive, needs the frontier LOCAL BY DEFAULTopen data, routine capability HOSTED OR HYBRIDopen data, needs the frontier SENSITIVITY AND CONTROL NEED → CAPABILITY THE TASK NEEDS → HIGH LOW
Source. LSF decision model, original to this work. Three of the four quadrants favor local; only open, non-sensitive work that genuinely needs the frontier stays hosted by default.

The closing argument. For a few years the only way to use frontier-grade AI was to rent it, and renting felt like the whole market. It was a phase, not the shape of things. The model now runs on the machine you own, the gap to the frontier has nearly closed, and the control that comes with holding the weights is a durable advantage renting can never offer. The future is local for the same reason every mature capability ends up in-house. You cannot build on what you do not control.

Arjun Jaggi · Aditya Karnam Gururaj Rao11
The Future Is LocalThe LSF on a Page
The LSF on a Page
Pin This Page
01 · Six dependencies local removes
I
Price Risk
II
Rate Limits
III
Data Egress
IV
Model Deprecation
V
Availability
VI
Lock-In
02 · The local stack, inside one boundary
Orchestration
Local Model
Retrieval
Evidence
Storage
03 · The decision, four quadrants
Local by default
Sensitive, routine.
Local, rent the edge
Sensitive, frontier.
Local by default
Open, routine.
Hosted or hybrid
Open, frontier.
04 · One quarter, two gates
Phase 1 · Pilot
GATE 1
Phase 2 · Harden
GATE 2
Phase 3 · Own
Framework vocabulary
Sovereignty Dividend Edge Sufficiency Control Inversion LSF
The future is local.
Pin this page, then run the
scorecard on page 13 to place your estate
Arjun Jaggi · Aditya Karnam Gururaj Rao12
The Future Is LocalSelf-Assessment
Self-Assessment · Ten Questions
Where does your estate stand on sovereignty?

Answer honestly for your most important AI workload. Count the YES answers, then read your level on the band below and the gaps it points to. Designed to be completed in pen.

No. Tests Question Yes No
1DATADo you know which AI workloads send sensitive data off-site to be answered?
2DATACould your most sensitive workload run with nothing leaving your environment?
3COSTIs your inference cost fixed, rather than a rate a vendor can raise without your consent?
4CAPHave you tested a quantized open model on your actual workload, not just a benchmark?
5HWDo you already own hardware that could serve your top workload locally?
6PERMIf your current model were deprecated tomorrow, could you keep running unchanged?
7PERMCan you reproduce an answer months later on the exact model that produced it?
8VOLDo you know the volume at which local becomes cheaper than renting for that workload?
9RULEIs local first, rent the exception a stated default, rather than renting by habit?
10TRUSTDo your local answers carry an evidence trail that survives a challenge?
Your score, your sovereignty level
0-2 YES
Level 1 · Rented
3-4 YES
Level 2 · Aware
5-6 YES
Level 3 · Piloting
7-8 YES
Level 4 · Hybrid
9-10 YES
Level 5 · Sovereign
Arjun Jaggi · Aditya Karnam Gururaj Rao13
The Future Is LocalMethodology · Notes · Authors

About the Research

This paper is a conceptual contribution. The LSF, the four-quadrant decision model, the seven exhibits, and the terms Sovereignty Dividend and Edge Sufficiency are original to this work. The Control Inversion names a shift in leverage that local inference creates. The quantization results it rests on, which let a thirteen billion parameter model run at interactive speed on a laptop, are drawn from the cited research and labeled where used [1] [2]. The local-agent evaluation method behind the roadmap is from BudgetBench [3], and the research agent Zorp [5] shows the pattern running on owned hardware.

Quantitative claims follow one discipline throughout, made visible in every exhibit. Solid fills carry cited data. Hatched fills and dashed guides labeled directional exist to make the shape of an argument legible, not to report measurements. The capability-gap figures are cited to a named index [4]. No statistic in this paper is attributed to a source that cannot be independently verified.

Notes

About the Authors

AJ
Arjun Jaggi and Aditya Karnam Gururaj Rao

Arjun Jaggi and Aditya Karnam Gururaj Rao write and advise at the intersection of enterprise AI strategy, governance, and local open-weight deployment. Their published body of work, spanning original frameworks, concept papers, and books, is read by executive teams navigating AI adoption and is available in full, without a paywall, at arjunjaggi.com.

Engage · arjunjaggi.com · calendly.com/arjunjaggi
Arjun Jaggi · Aditya Karnam Gururaj Rao14
YOUR MACHINE
Arjun Jaggi
Enterprise AI Research
The model you need already fits on the machine you own. The next decade of enterprise AI belongs to those who run it locally.
Sovereignty Dividend Edge Sufficiency Control Inversion LSF
© 2026 Arjun Jaggi and Aditya Karnam Gururaj Rao. All rights reserved. Academic citation permitted with attribution; commercial use and derivative frameworks require written permission.
White Paper No. 03