Files
Luna/Documentation/Luna AI Assistant/Context Compaction Research.md
T
darman 9929941748 Add project documentation and reference materials
Include Luna AI Assistant design docs covering channels, configuration,
core architecture, memory, scheduler, and skills. Add reference docs
from OpenClaw and ZeroClaw projects, plus Mistral and OpenAI API specs.
2026-04-04 04:14:06 +02:00

5.7 KiB

Context Compaction Research

This document captures a research analysis comparing Claude Code's compaction engine patterns against Luna's current implementation, with recommendations for adoption. Reference source: https://barazany.dev/blog/claude-codes-compaction-engine

Luna's Current Compaction Architecture

Component Current Approach
Trigger Token threshold: session.TokenAmount > MaxContextTokens * 0.75
Token estimation Naive text.Length / 4 heuristic (TokenEstimator)
Compaction method Single LLM call via LibrarianAgent (Mistral Small) — plain bullet-point summary
Retained context Last N messages (default: 2) carried forward
Post-compaction reconstruction Summary wrapped in MEMORY BEGIN/END markers as an Assistant message, then retained messages appended
Persistent memory File-based MemoryStore — raw conversation logs to disk, last 10 read back as system message
Tiers None — only full LLM summarization
Cache awareness None
Tool result management None — tool outputs accumulate until full compaction

Claude Code's Three-Tier Pattern

Claude Code employs a tiered strategy to manage context while maximizing cache efficiency:

  • Tier 1: Lightweight deterministic cleanup before every API call. This process clears old tool results (retaining only the last 5) and replaces them with placeholders. No LLM is involved at this stage.
  • Tier 2: API-level server-side strategies for token management using Anthropic-specific infrastructure.
  • Tier 3: Full LLM summarization as a last resort. This involves a structured 9-section summary with a chain-of-thought scratchpad. Post-compaction reconstruction includes a boundary marker, the summary, the 5 most recently read files (capped at 50K tokens), re-injected skills, tool definitions, and session hooks.

The key architectural insight is that cache economics drive every decision. This includes using cache_edits for surgical server-side deletions and ensuring the summarization call reuses the same cache key.

Pattern Evaluation

1. Tier 1 — Deterministic Tool Result Cleanup

Gap: Luna currently has zero tool result management. Recommendation: ADOPT. Effort: Low. Details: Implement a pre-call sanitizer that trims old tool results before each API call to prevent context bloating from large tool outputs.

2. Structured Compaction Prompt

Gap: Luna uses a simple "max 12 bullet points" prompt. Recommendation: ADOPT. Effort: Low. Details: Replace the current prompt with a structured template that forces categorized output, including user intent, key decisions, unresolved tasks, relevant facts, and technical context.

3. Tiered Compaction (Delay LLM Summarization)

Gap: Luna only supports full LLM summarization. Recommendation: ADOPT. Effort: Moderate. Details: Implement a 2-tier system. Tier 1 performs deterministic cleanup on every call, while Tier 2 triggers LLM summarization only when Tier 1 is insufficient. The Anthropic-specific server-side tier will be skipped.

4. Post-Compaction Reconstruction

Gap: Luna's reconstruction logic is basic and uses Assistant messages for summaries. Recommendation: PARTIALLY ADOPT. Effort: Moderate. Details: Place the summary as a System message instead of an Assistant message. Re-inject agent instructions and add a boundary marker with metadata (timestamp, pre-compaction message count). Include a continuation message so the agent does not treat the summary as something to respond to.

5. Improved Token Estimation

Gap: Luna relies on a text.Length / 4 heuristic. Recommendation: ADOPT. Effort: Low. Details: Replace the current heuristic with a proper tokenizer, such as Microsoft.ML.Tokenizers, or a significantly improved heuristic.

6. Autonomous Continuation After Compaction

Gap: Compaction can disrupt the conversation flow. Recommendation: ADOPT LIGHTLY. Effort: Trivial. Details: Prepend a brief system message after compaction, such as "Context was compacted. Continue naturally."

7. Cache-Aware cache_edits

Gap: This is specific to Anthropic's API. Recommendation: NOT APPLICABLE. Effort: N/A. Details: This could be revisited if an Anthropic provider is added to IProvider in the future.

8. Same-Cache-Key Summarization

Gap: Luna uses a cheaper model (Mistral Small) for compaction. Recommendation: NOT APPLICABLE. Effort: N/A. Details: Luna's current approach is effective when prompt caching is not a primary factor.

Priority Adoption Matrix

Priority Pattern Effort Impact
P0 Tier 1 — Deterministic tool result cleanup Low High
P0 Structured compaction prompt Low High
P1 Tiered compaction (delay LLM summarization) Moderate High
P1 Post-compaction reconstruction improvements Moderate Medium
P1 Improved token estimation Low Medium
P2 Continuation message after compaction Trivial Low-Medium
N/A Cache-aware cache_edits Not applicable (Mistral)
N/A Same-cache-key summarization Not applicable

Known Issues Found During Analysis

There is a bug in SessionManager.SaveSessionLogAsync() at line 97:

.Where(m => m.Role != ChatRole.System || m.Role != ChatRole.Tool)

This condition is always true due to De Morgan's law. It should use && instead of || to correctly filter out system and tool messages.

Cross-References

  • Core — SessionManager, compaction flow, token estimation
  • Memory — MemoryStore, persistent conversation logs
  • Configuration — SessionOptions (ContextTokenThreshold, RetainedMessagesAfterCompacting)