avatar
原创2026/06/22

深入理解 Claude Code /compact 命令的底层实现

AI 总结

深入分析 Claude Code /compact 命令的底层实现原理,涵盖 Context Window 限制、Token 计数机制、Anthropic Messages API 调用流程、会话压缩策略,以及 compact 摘要如何被整合回对话上下文。

/compact 是 Claude Code 中用于压缩对话历史的命令。当 context window 接近上限时,执行该命令会将当前完整的对话历史压缩成结构化摘要,释放大量空间使对话得以继续。

Context Window 与 Token

大模型在每次推理时能处理的文本量是有上限的,这个上限以 Token 为单位衡量。Token 是分词器(Tokenizer)切分后的最小语义单元。对于英文,大约 1 个单词 ≈ 1.3 个 Token;对于中文,1 个汉字 ≈ 1~2 个 Token。

Claude 3.5 / 3.7 系列模型支持最高 200K Token 的上下文窗口。这个窗口包含 System Prompt、所有历史对话消息、当前用户输入和模型即将生成的回复。

┌─────────────────────────────────────────────┐
│              Context Window (200K)           │
│                                              │
│  ┌──────────────┐                            │
│  │ System Prompt│  ~2K tokens                │
│  └──────────────┘                            │
│  ┌──────────────────────────────────────┐    │
│  │         Conversation History         │    │
│  │  turn 1: human  → assistant          │    │
│  │  turn 2: human  → assistant          │    │
│  │  ...                                 │    │
│  │  turn N: human  → assistant          │    │
│  └──────────────────────────────────────┘    │
│  ┌──────────────┐                            │
│  │ Current Input│                            │
│  └──────────────┘                            │
└─────────────────────────────────────────────┘

随着对话轮次增加,历史消息不断累积,最终逼近上限。此时模型要么被截断历史导致失忆,要么拒绝继续回复。

Anthropic Messages API

Claude Code 底层通过 Anthropic 的 Messages API 与模型交互。每次对话轮次都是一次 HTTP 请求:

POST https://api.anthropic.com/v1/messages
 
{
  "model": "claude-opus-4-5",
  "max_tokens": 8192,
  "system": "You are Claude Code, an AI assistant...",
  "messages": [
    { "role": "user",      "content": "帮我优化这个函数..." },
    { "role": "assistant", "content": "好的,我来分析一下..." },
    { "role": "user",      "content": "再加上错误处理" },
    { "role": "assistant", "content": "..." },
    { "role": "user",      "content": "现在的问题是..." }
  ]
}

messages 数组携带完整的对话历史,这就是 context window 被占满的直接原因——每次请求都要把所有历史消息原样传给 API。

/compact 的执行流程

执行 /compact 后,Claude Code 在本地发起一次专门的摘要请求:

构造摘要请求

Claude Code 将当前全部对话历史打包,加上系统指令,向 API 发送一次独立的摘要调用:

const compactResponse = await anthropic.messages.create({
  model: currentModel,
  max_tokens: 8192,
  system: `Your task is to create a comprehensive summary of the conversation so far.
The summary will be used as context for continuing the conversation.
 
Include:
1. What was being worked on (files, features, bugs)
2. Key decisions and approaches taken
3. Current state of any code changes
4. Outstanding tasks or issues
5. Any important context or constraints discovered
 
Be thorough but concise. The summary should allow seamless continuation.`,
  messages: currentConversationHistory
})

这次 API 调用会消耗相当数量的 token(输入 = 全部历史,输出 = 摘要),但这是一次性的代价。

替换对话历史

摘要生成后,Claude Code 用一条合成的对话轮次替换整个历史消息数组:

const summary = compactResponse.content[0].text
 
conversationHistory = [
  {
    role: 'user',
    content: `[Conversation compacted. Here is a summary of what happened so far:\n\n${summary}\n\nPlease continue from where we left off.]`
  },
  {
    role: 'assistant',
    content:
      'I understand. I have the full context from the summary and am ready to continue.'
  }
]

此后的所有 API 请求,messages 数组里只有这两条合成消息加上新的对话,token 用量骤降。

Token 计数

Claude Code 会在 compact 前后分别统计 token 用量。Token 计数通过 Anthropic API 响应中的 usage 字段获取:

{
  "id": "msg_01...",
  "usage": {
    "input_tokens": 142847,
    "output_tokens": 3621
  }
}

摘要的结构

一次典型的 compact 摘要输出大致如下:

## Session Summary

### Working Context
- Project: Next.js blog (bnotes)
- Primary task: Adding a new post

### Files Modified
- `contents/posts/claude-code-compact.mdx` — new file, currently being written

### Key Decisions
- Post category: `ai`, tags: [ai, claude, llm]

### Current State
- Frontmatter and first three sections completed
- Remaining sections to be finished

### Outstanding Tasks
- Complete remaining sections
- Verify frontmatter matches content-collections schema

摘要的质量直接决定了 compact 后对话的连续性——模型需要从这段文字中重建足够的上下文才能无缝继续。

--auto-compact 模式

除了手动执行 /compact,Claude Code 还支持 --auto-compact 标志:

Terminal
claude --auto-compact

开启后,Claude Code 会在后台监控 context window 的使用率。当使用率超过阈值(约 95%)时自动触发 compact。实现方式是在每次 API 响应返回后检查 usage.input_tokens 与模型最大 context 的比值:

const COMPACT_THRESHOLD = 0.95
 
function shouldAutoCompact(usage: Usage, modelContextWindow: number): boolean {
  const usageRatio = usage.input_tokens / modelContextWindow
  return usageRatio >= COMPACT_THRESHOLD
}
 
if (autoCompact && shouldAutoCompact(response.usage, MODEL_CONTEXT_WINDOW)) {
  await runCompact()
}

Compact 的局限性

信息损失不可避免。摘要本质上是有损压缩,细节性的代码片段、模糊的需求讨论、中途被否决的方案都可能在摘要中丢失。

摘要本身也消耗 token。compact 后的第一条合成消息虽然远小于原始历史,但随着对话继续,token 依然会再次累积。对于极长的工作会话,可能需要多次 compact。

跨文件状态难以保全。如果对话涉及多个文件的复杂修改,摘要可能无法精确还原每个文件的当前状态,需要在 compact 后手动确认关键文件内容。

总结

/compact 的本质是用模型自身来压缩上下文。整个流程可以概括为:

  • 将完整对话历史作为输入,发起一次独立的摘要 API 请求
  • 将摘要内容注入为一条合成的对话轮次,替换原始历史
  • 后续请求的 messages 数组从这条合成消息重新开始累积

这个设计完全在客户端完成,不依赖服务端会话状态管理,同时复用了模型本身的语言理解能力来完成压缩。

点击播放