Research Translation · Enterprise AI · arXiv:2609.13149

Your LLM Memory Strategy Is Tested at Full Context. It Runs at a Budget.

BudgetBench, a new open research protocol co-authored by Aditya Karnam Gururaj Rao and Arjun Jaggi, is the first harness to benchmark memory strategies at realistic per-call token budgets. What it found exposes a systematic blind spot in how the industry evaluates local LLM deployments.

Arjun Jaggi  ·  September 15, 2026  ·  12 min read
5 budget tiers tested: 2K, 4K, 8K, 16K, 32K tokens per call [1]
3 pilot benchmarks: SWE-bench Verified, LongBench v2, LongMemEval oracle [1]
1st open protocol treating per-call token budget as the independent variable [1]

The Problem No Standard Benchmark Sees

When your team evaluates a memory strategy for a local LLM agent, the benchmark you run almost certainly tests it at full context: the model gets everything it needs, memory retrieval has no ceiling, and quality scores reflect an environment that does not exist in your deployment.

In production, every call runs under a budget. Prefill latency grows with context length. Cache storage costs accumulate per session. Service latency objectives constrain how many tokens you can afford on each turn. Your memory strategy does not operate in an unconstrained sandbox. It operates inside these limits, every single time.

The gap between benchmark conditions and deployment reality is not a minor calibration issue. It is a structural measurement failure. A memory strategy that scores well at full context can fail silently under a realistic budget, and you will not know it until your agent starts producing degraded outputs in a live environment.

BudgetBench is the protocol that closes this gap. It treats the per-call input-token budget as the independent variable, sweeps it across five tiers from 2K to 32K tokens, and reports not just quality scores but budget-compliance rates: the frequency with which a memory strategy violates its own allocation. It is the first open, reusable measurement surface designed specifically for this class of problem. [1]

Definition: Budget-Blind Evaluation

Budget-Blind Evaluation is the practice of benchmarking LLM memory strategies at unconstrained context windows, producing quality scores that do not transfer to enterprise deployments operating under per-call token budgets. The term originates with this work and identifies a systematic gap in current evaluation methodology that no existing benchmark standard addresses.

Definition: Compliance Cliff

Compliance Cliff is the budget threshold below which a memory strategy begins violating its token allocation on a non-trivial fraction of calls, producing outputs that appear complete but reflect constraint violations rather than full memory access. The Compliance Cliff is invisible in single-budget evaluation and only surfaces when budget is swept as an independent variable. The term originates with this work.

What BudgetBench Actually Measures

The protocol introduces four first-class outcome metrics that standard benchmarks do not track together:

The architecture is deliberately modular. A swappable MemoryStrategy contract lets any team plug in their own retrieval or summarization approach. Explicit budget enforcement ensures violation events are captured, not silently truncated. Prompt-audit metadata and reproducibility artifacts let external teams replicate and extend every result. [1]

Key Finding

Across pilot studies, the harness exposed budget-compliance failures, non-monotonic quality curves, and operating points that single-budget evaluation hides entirely. A non-monotonic quality curve means more context budget does not reliably produce better outputs: quality can peak at an intermediate budget and decline as the budget grows, a phenomenon that only becomes visible when budget is swept systematically. [1]

The Architecture of the Protocol

Fig. 1: BudgetBench Harness Architecture
BUDGET SWEEP 2K / 4K / 8K / 16K / 32K MEMORY STRATEGY Swappable Contract LOCAL LLM Fixed: model, sampler, decode BUDGET ENFORCER Explicit token ceiling PROMPT AUDIT Metadata + reproducibility OUTCOME METRICS Quality  ·  Budget Utilization  ·  Latency  ·  Budget-Violation Rate Controlled Evaluation Surface

What the Pilot Studies Found

The paper reports three pilot studies, each deliberately scoped as pilots rather than final rankings. This is important: the authors distinguish between a protocol contribution and a claim-bearing result. The protocol is the durable contribution. The pilots demonstrate it.

The local pilot used Qwen2.5:1.5B across 89 items each on SWE-bench Verified and LongBench v2. The hosted replication used Qwen3 30B-A3B on 50 LongBench items with exact tokenization. The oracle study ran 500 items on LongMemEval scored by a reference evaluator. [1]

Across all three, the harness surfaced findings that would be invisible in a single-budget evaluation:

One finding deserves particular attention for enterprise teams: the early pilot's tokenizer approximation undercounts some served-model prompts, meaning violation rows in that slice are diagnostics about the approximation, not about the strategy itself. The authors surface this transparently. That kind of disclosure is the discipline that makes a protocol trustworthy. [1]

Budget-Violation Rate vs. Token Budget Tier (Directional Illustration)
Values are directional illustrations of the pattern the BudgetBench protocol is designed to surface. Actual results vary by memory strategy and task type. Source: directional, based on BudgetBench pilot findings [1].
Quality Score vs. Budget Tier: Non-Monotonic Pattern (Directional Illustration)
Illustrative representation of the non-monotonic quality curve the protocol exposes. A strategy that appears optimal at 32K tokens may underperform at 16K on specific task types. Directional illustration only.

Why This Matters for Enterprise AI Leaders

Most enterprise AI deployments that use local LLMs are making memory strategy decisions based on benchmarks that have nothing to do with their operational constraints. The budget tiers your engineers actually configure, 2K for fast retrieval agents, 8K for document assistants, 16K for analysis workflows, are not represented in the standard benchmarks your vendors cite.

BudgetBench changes that. It is not a vendor product. It is an open research protocol with a published harness, a swappable strategy contract, and reproducibility artifacts. Any team can run it against their own memory strategy at their own budget tiers. The paper is available at arXiv:2609.13149, and the code artifacts are referenced in the paper. [1]

For CTOs and Chief AI Officers

The question to ask your AI platform team: "What is our memory strategy's budget-violation rate at the token ceilings we actually enforce in deployment?" If the answer is "we don't track that," you are operating blind on a metric that directly predicts output quality under load.

Decision Framework: Should You Adopt BudgetBench?

Your Situation BudgetBench Applies? Priority
Local LLM agents with per-call token ceilings enforced in production Yes High
Comparing multiple memory strategies for a local deployment decision Yes High
Diagnosing unexplained quality degradation in a deployed local agent Yes High
API-only deployment with no per-call budget enforcement Partial Medium
Hosted model with unlimited context and cost-insensitive architecture No Low
Academic benchmark comparison with no operational constraints No Out of scope

Three Enterprise Scenarios

VP of Engineering · Financial Services

Compliance Document Agent Under Latency Constraints

A regional bank deploys a local LLM agent for regulatory document retrieval. Latency SLOs require a 4K token ceiling per call. The memory strategy was benchmarked at full context and passed quality gates. After deployment, analysts report inconsistent retrieval quality on multi-part queries. Running the BudgetBench protocol at the 4K tier reveals a substantial budget-violation rate on queries exceeding three clauses: the strategy silently exceeds its allocation, truncating context without logging a failure. The fix is a retrieval-first strategy with explicit chunk sizing, validated at the 4K tier before re-deployment.

Chief AI Officer · Healthcare System

Clinical Summary Agent with Non-Monotonic Quality Degradation

A health system runs a local summarization agent on discharge notes. The team increases the token budget from 8K to 16K expecting quality improvement, and instead observes degradation on structured note types. Standard benchmarks showed monotonic quality improvement with context length. The BudgetBench protocol reveals that the memory strategy's summarization approach introduces noise at higher budgets on short, structured inputs: the 8K tier outperforms 16K on this task type. The operating point is 8K, not 16K. Without budget-sweeping the evaluation, the team would have continued expanding budget and compounding the degradation.

ML Platform Lead · Enterprise Software

Memory Strategy Selection for a Multi-Tenant Agent Platform

A software company is choosing between three memory strategies for a multi-tenant agent platform where tenant tier determines token budget: 2K for free tier, 8K for professional, 32K for enterprise. Evaluating all three strategies at 32K only, as standard benchmarks do, produces one ranking. Running the BudgetBench protocol across all five tiers produces a different ranking at 2K and 8K: the strategy optimal for enterprise tenants is the worst performer for free-tier tenants. The team selects different strategies per tier, a decision invisible without budget-tiered evaluation.

How to Contribute and Build On This Work

BudgetBench is an open research protocol. The authors have published harness code, prompt-audit metadata, and reproducibility artifacts alongside the paper. There are several concrete ways enterprise AI teams and researchers can contribute:

Contribution Pathways

Implementation Roadmap for Enterprise Teams

Phase 1 · Weeks 1-3

Audit and Baseline

Document the token budgets actually enforced in each deployed local LLM agent. Identify which memory strategies are in use. Run the BudgetBench harness at your actual budget tiers. Establish a baseline violation rate for each strategy. Go/no-go gate: does any strategy show a violation rate above a threshold you define before hardening begins?

Phase 2 · Weeks 4-8

Strategy Selection

For agents where violation rates exceed tolerance, evaluate alternative memory strategies using the BudgetBench protocol at the same budget tiers. Use quality and violation rate together as the selection criteria: a strategy with lower quality scores but zero budget violations may outperform a higher-scoring strategy with a 20% violation rate in production.

Phase 3 · Weeks 9+

Operational Integration

Integrate budget-violation rate as a first-class production metric alongside latency and quality. Add alerting when violation rates exceed baseline. Re-run the BudgetBench protocol on any memory strategy change as standard practice before deployment. Contribute task-type graders back to the open harness.

Executive Checklist

Build vs. Buy vs. Configure

ComponentApproachRationale
BudgetBench harness Use open harness Protocol and code are published. Build on the existing harness rather than re-implementing. Extend with your task-specific graders.
Memory strategy implementation Build or source internally The MemoryStrategy contract is swappable. Plug your existing retrieval or summarization approach into the harness directly.
Graders for your task types Build Deterministic graders for enterprise tasks (document classification, structured extraction) do not yet exist in the harness. These are your contribution to the protocol.
Budget monitoring in production Configure from existing observability stack Budget-violation rate is a counter metric. Add it to your existing LLM observability tooling as a first-class signal.

BudgetBench is published on arXiv at arXiv:2609.13149. If your team works on local LLM deployment, memory architecture, or enterprise AI evaluation, read it and run the protocol against your own stack. The field benefits when more teams report budget-violation rates as a standard metric alongside quality scores.

For teams working on enterprise memory architecture, the protocol integrates directly with any retrieval layer. For teams thinking about the broader memory problem in enterprise AI, BudgetBench provides the measurement discipline that makes memory strategy comparisons credible. And for teams running local LLM agents at scale, budget-violation rate is now a metric you cannot afford not to track.

Excited about AI, innovation, and growth?

Start a conversation

References