The premise is wrong, and that is the point. The internet is not free. You pay for bandwidth, hosting, CDNs, and DNS. What is free is the protocol layer. TCP/IP, HTTP, and DNS are royalty-free open standards that anyone can implement. AI has the equivalent: open-weight models. What it lacks is the neutral middle layer , the Inference Commons, that would make AI access economics look like internet access economics.
The question "if the internet is free, why not AI?" contains a wrong premise and a correct intuition. The wrong premise is that the internet is free. It is not. Broadband costs money. Hosting costs money. CDNs, load balancers, and data centers cost money. The internet's infrastructure layer is priced, metered, and sold commercially, just like AI inference. The correct intuition is that there is something structurally different about how the internet democratized access to information, and that AI has not replicated that structure. Understanding what that difference actually is changes what you think needs to be built.
The internet's access revolution happened at the protocol layer. TCP/IP, HTTP, and DNS are open standards maintained by the IETF and IANA with no licensing fees, no usage restrictions, and no ownership by any commercial entity. A nonprofit in Lagos, a startup in Ho Chi Minh City, and a Fortune 500 in New York all send HTTP requests under identical terms. The protocol layer is a commons. AI has produced the equivalent for models, specifically open-weight releases that anyone can download and run, with no per-inference fee. What AI has not produced is the equivalent of internet exchange points (IXPs): neutral infrastructure where organizations connect to the network at cost rather than at margin. That gap is what this post names and maps.
A neutral AI compute layer that operates on cost-recovery pricing rather than commercial margin pricing, enabling organizations to access AI inference at the marginal cost of computation rather than at the price of market concentration. The Inference Commons would stand in relation to AI capability as internet exchange points stand in relation to internet bandwidth: not free, but priced at cost, accessible regardless of organizational budget, and governed as public infrastructure rather than private service. The term originates with this work.
The structural absence, in the current AI stack, of a neutral inference layer equivalent to the internet's open protocol layer (TCP/IP, HTTP). Unlike the internet, where the protocol layer creates roughly equal access at the connection level regardless of organizational size or budget, AI inference is commercial at every layer above the model weights themselves. The Access Layer Gap means that AI capability scales with budget in ways that internet access never has, creating compounding inequality between AI-funded and AI-unfunded organizations. The term originates with this work.
The internet's protocol layer did not become open by accident. TCP/IP was developed under ARPA (Advanced Research Projects Agency) with public funding and released without restriction. HTTP was developed at CERN, also with public institutional funding, and released under royalty-free terms in 1991. DNS was specified in IETF RFCs 882 and 883 and has been maintained as an open standard since 1983. None of these protocols had a commercial owner at the point of release. The commons was built by public institutions before the commercial opportunity was visible.
What followed was a structural property that most people now take for granted. Because the protocol layer was royalty-free and unowned, no single organization could extract rent from the act of connecting. Every organization that wanted to build on the internet could do so under identical access terms. The IXP model extended this to the physical layer: rather than requiring every organization to pay a single carrier for transit, IXPs created neutral peering points where networks could exchange traffic at cost, collapsing the price of bandwidth toward marginal cost. The result was that internet access economics became progressively less dependent on organizational size.
The internet's democratization of access was not the result of the bandwidth being free. It was the result of the protocol layer being unowned and the infrastructure layer being governed as shared cost rather than extracted margin. Both conditions were necessary. Neither exists today for AI inference.
AI has made real progress on the model layer equivalent. The release of open-weight foundation models, including the work documented in Touvron et al. [1], which means that the model itself, the "protocol" equivalent, is increasingly available without per-inference fees. An organization with access to compute can run a capable model with no ongoing licensing cost to a model vendor. This mirrors the internet's royalty-free protocol layer more closely than the pre-open-weights era.
The divergence is at the compute layer. Running inference on a capable foundation model requires significant GPU capacity, either owned or rented. Organizations without the capital to own GPU clusters must rent inference capacity from a small set of cloud providers or specialized inference APIs. Those providers operate on commercial margin pricing, not cost-recovery pricing. The result is that AI capability access is gated by the ability to pay commercial rates, with no neutral alternative. There is no AI equivalent of an IXP: no shared infrastructure where inference cost approaches marginal compute cost.
The Access Layer Gap compounds differently for different organization types. For a hyperscaler, it is irrelevant: they own the compute. For a large enterprise, it is a cost management problem: the per-inference price is affordable, but at scale it becomes meaningful, and the work documented in FrugalGPT [2] specifically addresses how to compress it. For a mid-market organization, it is a prioritization constraint: AI use cases are evaluated partly on whether the inference cost can be justified at expected volumes. For a nonprofit or a public institution, it is often a categorical barrier: commercial inference pricing makes many high-value AI applications economically impossible at the scale where they would be most impactful.
This creates an access profile that has no equivalent in internet infrastructure. A healthcare nonprofit serving rural populations cannot afford the per-query cost to run AI-assisted clinical decision support at the volume required. A municipal government trying to automate permit processing at scale faces inference costs that public budgets cannot absorb. A community college building AI tutoring tools for students who cannot pay runs into per-student inference economics that make the service dependent on philanthropic subsidy rather than sustainable infrastructure. None of these organizations face a comparable barrier when sending HTTP requests.
The Inference Commons is not a demand for free compute. Compute costs money regardless of who operates it. The argument is about governance model and pricing structure, not about eliminating cost. An Inference Commons would be analogous to a cooperative utility: member organizations contribute to a shared compute pool and access it at the marginal cost of their usage, with governance distributed among contributors rather than concentrated in a commercial operator.
Three structural models could produce something like this. The first is public infrastructure: governments fund shared AI inference capacity the same way they fund internet backbone, and provide it to public institutions (hospitals, universities, municipalities, NGOs) at or near cost. The second is the cooperative model: a consortium of organizations pools GPU capacity and operates shared inference infrastructure on a cost-recovery basis, similar to how research networks (Internet2 in the US, GÉANT in Europe) operate neutral infrastructure for academic institutions. The third is an open-source inference layer: a protocol for federated inference that allows any compute provider to contribute capacity, with routing optimized for cost and latency rather than for a single operator's margin. Practitioners building in this space, including the team at Zorp.dev (Aviskaar) [4], are actively mapping what this looks like in practice.
A development finance institution funds a global NGO to verify supplier compliance across a three-tier supply chain in five manufacturing regions. The NGO wants to use AI to analyze supplier documents, flag anomalies, and generate audit summaries at a pace that matches the supply chain's velocity. At commercial inference API rates, the per-document cost makes the program dependent on grant funding for every query rather than operating as sustainable infrastructure. The NGO's program team lacks GPU infrastructure and has no path to a cost-recovery alternative. The correct accountability tier for this gap is not the NGO. It is the absence of a neutral compute layer that would allow mission-driven organizations to access inference at marginal cost. The corrective action is procurement of inference through a cooperative or public infrastructure model if one exists in the relevant region, combined with aggressive inference optimization using cascade approaches documented in FrugalGPT [2] to maximize the coverage achievable within budget.
A community bank with $2.4B in assets wants to deploy AI-assisted underwriting to close the loan processing speed gap with fintech competitors. The use case is high-value per decision but low-volume relative to a large institution. At commercial API rates, the per-query inference cost is justified at the application review stage but not earlier in the pipeline, where exploratory queries would add value but cannot be cost-justified. The bank's technology team evaluates self-hosted inference on cloud GPU instances but finds the minimum-viable infrastructure cost exceeds what the pilot budget can support. The Access Layer Gap is manifest here as a scale minimum: inference economics only work at volumes that smaller institutions cannot organically reach. The corrective action is a phased inference budget that reserves AI queries for the decision-critical stages of the underwriting workflow, with open-weight models on shared regional infrastructure as the path to cost reduction.
A public hospital network with fourteen facilities wants to deploy AI-assisted clinical documentation and differential diagnosis support for emergency department physicians. The program is clinically compelling and the organizational readiness is high. The barrier is per-patient inference economics. At the emergency department volume across fourteen facilities, commercial inference pricing exceeds what the network's operating budget can absorb without eliminating another program. The network's legal team identifies the use case as high-risk under EU AI Act Article 6 [see Who Answers When Your AI Gets It Wrong], requiring conformity assessment regardless of the cost question. The corrective action is engagement with regional public health AI infrastructure initiatives. Several EU member states are evaluating public inference capacity for healthcare, combined with open-weight deployment on existing data center infrastructure to bring inference cost below the clinical value threshold.
| Scenario | Volume | Budget Sensitivity | Recommended Path | Inference Commons Fit |
|---|---|---|---|---|
| High-volume, cost-sensitive, mission-driven | Millions of queries | Critical | Cooperative or public infrastructure. Open weights on shared compute. FrugalGPT cascade [2] | High. Primary use case for a neutral layer |
| Low-volume, high-value, commercial enterprise | Thousands of queries | Low | Commercial API. Budget management through cascade and caching | Low. Commercial pricing is manageable at this volume |
| Mid-market, scaling volume | Tens of thousands | Moderate | Hybrid. Open weights for bulk processing, commercial API for latency-sensitive queries | Medium. Becomes high as volume grows |
| Public institution, fixed budget | Variable | Very high | Public infrastructure if available. Cooperative model. Open weights on owned compute as last resort | Very high. Public institutions are the primary constituency for a neutral layer |
| Startup, early-stage | Low to moderate | High | Commercial API with aggressive cost management. Open weights when engineering capacity allows | Medium. Depends on use case and growth trajectory |
Own GPU infrastructure for high-volume, cost-sensitive workloads where the inference economics are calculable and the volume is committed. This path requires capital expenditure, ML infrastructure engineering, and ongoing operations. It is the right choice for organizations that have crossed the volume threshold where ownership is cheaper than rental, typically well above what most enterprises will reach in the next two years. The build path also requires model maintenance capabilities: open-weight models need updating as capability improves.
Use commercial inference APIs for low-to-moderate volume, latency-sensitive, or rapidly evolving use cases. This is the right path for most enterprises today. The cost management levers are documented and the tooling is mature. FrugalGPT-style cascade approaches [2] can reduce commercial API costs by up to 98% on appropriate use cases by routing simpler queries to cheaper models. The risk is dependency on commercial pricing, which can change, and on provider roadmaps, which cannot be controlled.
Participate in or advocate for cooperative and public inference infrastructure as it develops. Several regional and sectoral initiatives are building shared AI compute capacity for healthcare, education, and public administration. Configuration investment is advocacy, procurement, and integration rather than infrastructure ownership. This is the medium-term path for organizations whose use case economics make commercial APIs prohibitive. Platforms like Zorp.dev (Aviskaar) [4] are building practitioner infrastructure in this space worth monitoring.
Map inference costs by use case. Identify volume thresholds where open-weight deployment becomes cheaper than commercial APIs. Implement cascade and caching on the top three workloads by cost. Go/no-go gate: cost per outcome is tracked and optimization levers are documented.
Stand up open-weight inference for bulk, cost-sensitive workloads. Evaluate cooperative or public infrastructure options in your region and sector. Reduce single-provider dependency to under 60% of inference spend. Go/no-go gate: hybrid inference is operational and cost reduction is measured.
Actively monitor and participate in cooperative AI infrastructure initiatives relevant to your sector. Contribute engineering capacity or governance voice to neutral layer development. Treat infrastructure policy as a technology strategy input, not a background noise item. Success criteria: at least one deferred use case becomes viable through lower-cost infrastructure access.
Commercial inference pricing is not regulated and can change unilaterally. Organizations with no alternative path face pricing volatility that can invalidate AI program economics retroactively.
Use cases deferred on inference economics are compounding opportunity costs. The organizations building access resilience now will operate those use cases when infrastructure options mature. The ones that don't will be rebuilding the case from scratch.
Inference cost reduction through cascade and caching has documented potential of up to 98% for appropriate workloads [2]. The optimization investment is typically recoverable within the first month of operation at any meaningful scale.
Organizations that participate in cooperative infrastructure governance early gain disproportionate influence over pricing structures and access terms. The window to shape those terms is before the infrastructure is built, not after.
The accountability structure for AI systems deployed on commercial inference infrastructure is covered in Who Answers When Your AI Gets It Wrong. If your inference provider changes model behavior unilaterally and a downstream harm results, the accountability chain starts at the Enterprise Deployer tier, not the provider. Understanding who owns what in the access stack is as important as understanding what access costs.
Investigation processes that rely on AI inference face compounding costs that are often invisible to the research function. The Investigation Debt framework in Investigation Debt addresses how to structure AI research queries to maximize signal per inference dollar, which is directly applicable to any organization managing inference costs against a fixed research budget.