Thinking of ACE? We Can Do It with Fewer Tokens

2026-08-20 · Hugging Face

Thinking of ACE? We Can Do It with Fewer Tokens

Background

When given realistic multi-step tasks—splitting a bill, finding a song, or reconciling orders across nine simulated applications—LLM agents usually fail not because they lack knowledge of the APIs, but because they have not internalized how to use them reliably. These usage patterns are learnable from the agent’s own interaction history.

Two systems address this challenge: ACE (Agentic Context Engineering) and ALTK-Evolve from IBM Research. Both are forms of agentic memory that convert an agent’s past trajectories into reusable lessons and feed them back at inference time, without weight updates or human labels.

Core Agreement: Do Not Compress

Despite different terminology, both systems reach the same fundamental conclusion: lessons should not be compressed.

ACE identifies two key risks—brevity bias and context collapse—and counters them by maintaining a rich, itemized playbook where every bullet carries a helpful/harmful counter. The model is allowed to distill relevance at read time.

ALTK-Evolve reaches the same position from another direction: every distinct guideline retains a support count indicating how many independent episodes produced it. The system never collapses the entire store into a handful of rules. A lesson validated across five tasks is treated as a different object from one that appeared once; both are worth preserving.

Thus, on the central question of whether to compress an agent’s hard-won experience into tidy summaries, both systems answer “no.” Count the support, do not collapse the detail. ACE’s per-bullet counters and ALTK-Evolve’s support counts are two implementations of the same principle.

Key Differences

The systems diverge in two areas: how the memory store is built (Consolidation) and how it is delivered at inference time (Delivery). The latter is what drives the token cost difference.

Consolidation

ACE grows a single evolving playbook through a Generator → Reflector → Curator loop, using incremental delta updates and embedding-based deduplication. It clusters near-duplicate lessons and merges them while preserving total support counts.

ALTK-Evolve extracts typed guidelines (strategy, recovery, optimization) with causal attribution and provenance linked back to the source trajectories at subtask granularity, enabling cross-app transfer.

Delivery — The Token Deciding Factor

ACE injects its comprehensive playbook at every single step, regardless of model or task.

ALTK-Evolve treats delivery as a dial rather than a constant:

  • A small fixed core of high-support guidelines
  • Extended per task by a handful of selected guidelines (via cosine similarity or LLM-guided retrieval, priority-weighted)
  • Or, when the model has sufficient capacity, the full consolidated set

The same lessons exist in both systems. The difference is that ACE always sends everything, while ALTK-Evolve sends only what a given model can effectively utilize.

Experimental Results on AppWorld

Using the same base ReAct agent:

DeepSeek-V3.2

  • ACE: TGC 80.4 / SGC 73.2, 634K tokens/task
  • ALTK-Evolve: TGC 89.3 / SGC 80.4, 263K tokens/task

gpt-oss-120b

  • ACE: TGC 54.8 / SGC 35.7, 777K tokens/task
  • ALTK-Evolve: TGC 56.0 / SGC 37.5, 116K tokens/task

On the stronger model, ALTK-Evolve outperforms ACE on both accuracy and cost at roughly 40% of the token budget. On the weaker model, accuracy is comparable (or statistically tied within benchmark noise), while using approximately one-seventh the tokens.

Performance by Task Difficulty

Breakdown by difficulty reveals model-dependent patterns. On gpt-oss-120b, ACE’s full playbook helps more on Easy and Medium tasks where generic instruction following dominates. However, on Hard tasks—where the model must select the right lesson from many—curated retrieval pulls ahead, and these hard tasks determine the aggregate score.

On the stronger DeepSeek-V3.2, the stronger model can better absorb ACE’s full playbook, yet ALTK-Evolve still maintains an overall advantage.

Cost Perspective

While ACE focuses on building its context efficiently, ALTK-Evolve’s advantage lies in serving it. Retrieving a small number of relevant guidelines per task instead of injecting the entire playbook every step is where the substantial token savings occur. This is the direct consequence of the differing delivery strategies.

Both approaches demonstrate the power of learning from self-generated trajectories. ALTK-Evolve offers a more token-efficient path to similar or better performance by intelligently modulating how much learned context is provided at inference time.

Source