How Much Memory Does Your Agent Actually Need?

2026-08-20 · Hugging Face

How Much Memory Does Your Agent Actually Need?

Introduction

In our previous post, we compared ALTK-Evolve with ACE and showed that the delivery method of an agent's self-distilled guidelines — retrieving a few per task versus injecting the entire set — significantly affects both accuracy and cost. This article steps back to address a more fundamental question: how much memory should you actually give your agent?

Equipping an agent with agentic memory sounds straightforward: distill lessons from past work, insert them back into context, and more experience should lead to better performance. However, it does not always work that way. After scaling evaluations to eight models ranging from a 30B dense model to frontier proprietary systems, one finding stood out clearly: Agentic memory is not a feature you simply switch on. It is a dose you must calibrate to the model.

TL;DR

ALTK-Evolve allows an agent to learn from its own past trajectories by distilling reusable behavioral guidelines and injecting them at inference time, with no weight updates and no human annotation.

The optimal dose varies by model capability: strong models with headroom benefit from the full guideline set; weaker models perform best with a compact high-confidence core plus per-task retrieval; saturated models show no measurable improvement.

Curated retrieval can be both the most accurate and cheapest option. For example, gpt-oss-120b gained +16.1 percentage points in task completion with only +5% additional tokens. Prompt caching makes even the full guideline set affordable in production.

The Key Insight: Dosage Depends on Capability

Not every model benefits equally from the same amount of memory. Across eight models spanning the capability spectrum, three recurring patterns emerged:

Strong models with headroom want the complete guideline set, including rare edge-case lessons. They possess the capacity to absorb and apply everything. DeepSeek-V3.2 (671B MoE) improved task completion by +9.5 percentage points when given its full self-mined guideline set.

Smaller or weaker models become overwhelmed by a large guideline set. For these, a tight, high-confidence core combined with a handful of task-relevant guidelines retrieved per task works best. gpt-oss-120b (117B MoE) achieved a +16.1pp gain with this selective approach, while the full set delivered smaller gains at roughly 50% higher token cost.

Saturated models exhibit no measurable gain from additional memory. This is labeled the "saturated pattern." GLM-5 (745B MoE) fell into this category. Possible reasons include the model already operating near its performance ceiling on these tasks, guidelines not addressing remaining failure modes, or ineffective application of the provided guidance.

The factor determining which pattern a model follows is not simply parameter count. Benchmark headroom, context window size, architecture, guideline quality, and task distribution all appear to influence the outcome. The practical takeaway remains: the right dose of memory must be calibrated to the specific model.

Learning Happens Around the Model, Not Inside It

Here, "memory" does not mean replaying past transcripts. It refers to a set of distilled behavioral guidelines — successful strategies, mistakes to avoid, and edge cases — extracted from the agent's own prior trajectories.

The ALTK-Evolve loop is straightforward:

  • The agent attempts tasks and generates trajectories.
  • Guidelines are extracted from both successful and unsuccessful runs.
  • These are consolidated into a reusable set.
  • At inference time, the agent receives either the full set or a task-relevant selection.

No model weights are updated. Learning occurs by changing the guidance available to the agent rather than modifying the underlying model. This is why the approach is inexpensive to adopt and portable across widely different models.

Evaluation on AppWorld

The study evaluated performance on AppWorld, which contains 585 multi-step tasks (168 test_normal + 417 test_challenge) across nine simulated applications including calendars, messaging, payments, and others.

Tasks are scored on two metrics: Task Goal Completion (TGC) — whether the agent fully completes the task — and the stricter Scenario Goal Completion (SGC), which requires every variant of a scenario to pass.

Three Clearly Defined Configurations

To eliminate ambiguity about context window content, the three configurations are defined as follows. Both memory configurations use the exact same guideline set mined exclusively from the training split of AppWorld. Test data is never used.

  • Baseline: No memory — the agent operates as originally provided.
  • Full guideline set: Every mined guideline is injected on every ReAct step.
  • Curated retrieval: A fixed high-confidence core plus a small number of task-relevant guidelines retrieved for each specific task.

Because the number of guidelines mined varies with model capability, results are reported by delivery strategy rather than raw counts.

Practical Takeaways

Curated retrieval frequently provides the best balance of accuracy and cost. Prompt caching makes even the full guideline set viable in production environments.

The central conclusion is that the optimal dosage of agentic memory must be carefully calibrated to each individual model. One size does not fit all.

Source