代码仓库ChainReaction

模型 API 是无状态的。请求发出去,回复拿回来,服务端不留上下文。你平时看到的”多轮对话”,是客户端每次把完整历史重新贴一遍。LangChain 里被贴的那份历史,是一组消息对象。

工具调用走的是同一条路。模型说”调 get_weather,参数 Paris”,这是一条消息;工具函数执行完返回的 “Sunny, 22C”,是另一条消息;两条消息靠一个 ID 串起来。ID 对不上,整个请求被拒。

消息不只是提示词的包装类,它是模型和工具之间唯一的数据载体。这件事决定了很多日常问题长什么样。请求报 400,多半是消息顺序或配对不合法;模型答非所问,是历史里混进了该清掉的内容;上下文超限,问题变成该裁掉哪几条。你没法直接改 HTTP body,能改的只有这串对象。

这篇讲清四类消息的字段、工具调用的配对规则、content 和 content_blocks 的分工、元数据里能读到什么、怎么存怎么还原,以及长对话怎么裁。下面所有代码都用 DeepSeek 实跑过,输出是终端里的原文,脚本在仓库的 Messages/ 目录下。

1
2
3
4
5
6
7
8
9
10
from langchain_openai import ChatOpenAI
import os

model = ChatOpenAI(
api_key=os.getenv('DEEPSEEK_API_KEY'),
base_url="https://api.deepseek.com/v1",
model="deepseek-chat",
temperature=0.1,
max_tokens=1000,
)

四类消息

类 type 谁产生 关键字段
SystemMessage system 你的代码 content
HumanMessage human 你的代码 / 用户输入 content、name、id
AIMessage ai 模型返回 content、tool_calls、usage_metadata、response_metadata
ToolMessage tool 你的工具执行代码 content、tool_call_id、name、artifact

四类消息都继承自同一个基类,共同字段是 content、type 和 id,另外还有一个 additional_kwargs,默认是空字典,用来装 provider 特有的零碎数据。字段一样,才可能用一段代码处理整条历史。

SystemMessage 放初始指令,定角色和规矩。它得排在消息列表最前面,位置错了模型会把它当普通对话读。

HumanMessage 是用户输入。除了 content,还能带 name(区分不同用户)和 id(追踪用)。文档里专门提醒 name 的行为各家 provider 不一致,有的拿来识别用户,有的直接忽略。

AIMessage 是模型输出。纯文本回答只是它的常见形态,它同时承载 tool_calls、token 统计和 provider 原始元数据。文档也提到,可以手工造一个 AIMessage 插进历史里,装作是模型说过的,用来伪造上下文或补全被裁剪的历史。

ToolMessage 是单个工具的执行结果。必填 tool_call_id,它指向 AIMessage 里那次调用的 ID。artifact 字段是给程序看的,不会发给模型,适合塞检索到的文档 ID、页码、调试信息。文档给的检索例子很典型:content 放原文片段供模型引用,artifact 放 {"document_id": "doc_123", "page": 0},前端拿它去渲染原文位置,模型那边完全看不到这些。

四条消息在一次带工具调用的往返里是这样流动的:

注意第三步和第四步都在你这边。模型只负责说”要调什么”,执行是你的事,把结果包装成 ToolMessage 也是你的事。

手工构造四类消息

先用最笨的方式把四条消息造出来,看它们长什么样:

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
from langchain.messages import SystemMessage, HumanMessage, AIMessage, ToolMessage

system_msg = SystemMessage("You are a helpful assistant. Answer in one sentence.")
human_msg = HumanMessage(
content="What's the weather in Paris?",
name="alice",
id="msg_human_001",
)
ai_msg = AIMessage(
content=[],
tool_calls=[{
"name": "get_weather",
"args": {"location": "Paris"},
"id": "call_abc123",
}],
)
tool_msg = ToolMessage(
content="Sunny, 22C",
tool_call_id="call_abc123",
name="get_weather",
)

for m in (system_msg, human_msg, ai_msg, tool_msg):
print(type(m).__name__, "| type =", m.type, "| id =", m.id)
print(" content :", repr(m.content))
print(" content_blocks :", m.content_blocks)

真实输出:

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
---- SystemMessage ----
type : system
id : None
content : 'You are a helpful assistant. Answer in one sentence.'
content_blocks : [{'type': 'text', 'text': 'You are a helpful assistant. Answer in one sentence.'}]

---- HumanMessage ----
type : human
id : msg_human_001
content : "What's the weather in Paris?"
content_blocks : [{'type': 'text', 'text': "What's the weather in Paris?"}]

---- AIMessage ----
type : ai
id : None
content : []
content_blocks : [{'type': 'tool_call', 'id': 'call_abc123', 'name': 'get_weather', 'args': {'location': 'Paris'}}]
tool_calls : [{'name': 'get_weather', 'args': {'location': 'Paris'}, 'id': 'call_abc123', 'type': 'tool_call'}]

---- ToolMessage ----
type : tool
id : None
content : 'Sunny, 22C'
content_blocks : [{'type': 'text', 'text': 'Sunny, 22C'}]
tool_call_id : call_abc123
name : get_weather

两个细节值得停一下。

手工构造的消息 id 是 None,除非你自己传。模型返回的 AIMessage 会带一个 LangChain 生成的 ID,形如 lc_run--01a10783-e4b5-7903-90e7-a4da6961de2c-0。做追踪、做去重、写 RemoveMessage 的时候都要用到它,所以别指望手工消息有 ID。

AIMessage 的 content 是空列表 [],但 content_blocks 里有一个 tool_call 块。纯工具调用不带文本,这不影响发送,模型能正常读到调用意图。

工具调用:两个 ID 必须对上

真实场景里 tool_calls 不是你写的,是模型生成的。下面这段完整跑了一遍往返:

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
import os
from langchain_openai import ChatOpenAI
from langchain.tools import tool
from langchain.messages import HumanMessage, ToolMessage

@tool
def get_weather(location: str) -> str:
"""Get the current weather for a city."""
fake = {"Paris": "Sunny, 22C", "Tokyo": "Rainy, 18C"}
return fake.get(location, "Unknown city")

model_with_tools = model.bind_tools([get_weather])

messages = [HumanMessage("What's the weather in Paris?")]
ai_message = model_with_tools.invoke(messages)
messages.append(ai_message)

call = ai_message.tool_calls[0]
tool_result = get_weather.invoke(call["args"])

tool_message = ToolMessage(
content=tool_result,
tool_call_id=call["id"],
name=call["name"],
)
messages.append(tool_message)

print("ToolMessage.tool_call_id :", tool_message.tool_call_id)
print("AIMessage.tool_calls[0].id :", call["id"])
print("matched :", tool_message.tool_call_id == call["id"])

final = model_with_tools.invoke(messages)
print("final content:", final.content)

真实输出:

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
---- AIMessage with tool calls ----
content : "I'll check the current weather in Paris for you."
tool_calls: [{'name': 'get_weather', 'args': {'location': 'Paris'}, 'id': 'call_00_m5sEphpzABSo8NpneDTG4455', 'type': 'tool_call'}]
usage : {'input_tokens': 273, 'output_tokens': 48, 'total_tokens': 321, ...}

tool executed -> 'Sunny, 22C'

ToolMessage.tool_call_id : call_00_m5sEphpzABSo8NpneDTG4455
AIMessage.tool_calls[0].id: call_00_m5sEphpzABSo8NpneDTG4455
matched : True

---- final AIMessage ----
content : "The weather in Paris is currently **sunny with a temperature of 22°C** (about 72°F). It's a lovely day there!"
usage : {'input_tokens': 339, 'output_tokens': 35, 'total_tokens': 374, ...}
tool_calls: []

DeepSeek 生成的调用 ID 长这样:call_00_m5sEphpzABSo8NpneDTG4455。你把它原样抄进 ToolMessage.tool_call_id,配对就成立。手工构造时常用的 call_123 这种短 ID 只是示例写法,真实 API 返回的字符串不要改。

故意把 ID 写错会怎样,我试了:

1
2
3
4
5
6
7
8
9
10
11
bad_messages = [
HumanMessage("What's the weather in Paris?"),
ai_message,
ToolMessage(content="Sunny, 22C", tool_call_id="call_wrong_id", name="get_weather"),
]
try:
r = model_with_tools.invoke(bad_messages)
print("no error, content:", r.content)
except Exception as e:
print(type(e).__name__)
print(str(e)[:800])

真实报错:

1
2
OpenAIInvalidRequestError
Error code: 400 - {'error': {'message': "An assistant message with 'tool_calls' must be followed by tool messages responding to each 'tool_call_id', The following tool_call_ids did not have response messages: call_wrong_id (request_id: 39d6737c-8128-437c-ba84-41f7b65b2b7f)", 'type': 'invalid_request_error', 'param': None, 'code': 'invalid_request_error'}}

报错信息有点绕:它说的是 call_wrong_id 没有对应的响应消息。因为模型看到的历史里,call_wrong_id 这个 ID 从来没被调用过,你却在回应它;而真正被调用的那个 ID,你又没回。服务端两边都对不上,直接 400。

模型偶尔也会给出格式坏掉的调用。这类内容会落在 invalid_tool_calls 里,块类型是 invalid_tool_call,带上工具名、参数和出错原因。它和 tool_calls 是分开的两个字段,遍历工具调用时记得两个都看一眼,不然坏掉的那次调用会被静默跳过。

还有一种常见的错法:两个工具调用,你只回了一个 ToolMessage。结果一样是 400。文档在讲删除消息时也专门警告过这条规则,多数 provider 要求带工具调用的 assistant 消息后面必须跟着对应的 tool 结果。

content 和 content_blocks

content 是原始载荷,类型很松,可以是字符串,也可以是内容块列表。content_blocks 是 LangChain 在 v1 里加的标准视图,它会惰性地把 content 解析成统一的、带类型的块列表。文档强调 content_blocks 不替代 content,两者并存。

差异在跨 provider 时才明显。Anthropic 的 thinking 块和 OpenAI 的 reasoning 块格式完全不同,经过 content_blocks 都会变成 {'type': 'reasoning', ...}。你写业务代码只认标准块,换模型时不用改解析逻辑。

构造多模态输入有两条路。标准块写法:

1
2
3
4
5
6
from langchain.messages import HumanMessage

mm = HumanMessage(content_blocks=[
{"type": "text", "text": "Describe this picture in Chinese."},
{"type": "image", "url": "https://s2.loli.net/2025/11/15/NS4g8bRiYu7z31f.png"},
])

provider 原生写法(这里用 OpenAI 风格的 image_url):

1
2
3
4
mm2 = HumanMessage(content=[
{"type": "text", "text": "Describe this picture in Chinese."},
{"type": "image_url", "image_url": {"url": "https://s2.loli.net/2025/11/15/NS4g8bRiYu7z31f.png"}},
])

两者打印出来的 content_blocks 对比:

1
2
3
4
5
6
7
---- multimodal: standard content_blocks ----
content : [{'type': 'text', 'text': 'Describe this picture in Chinese.'}, {'type': 'image', 'url': 'https://s2.loli.net/2025/11/15/NS4g8bRiYu7z31f.png'}]
content_blocks : [{'type': 'text', 'text': 'Describe this picture in Chinese.'}, {'type': 'image', 'url': 'https://s2.loli.net/2025/11/15/NS4g8bRiYu7z31f.png'}]

---- multimodal: provider-native content ----
content : [{'type': 'text', 'text': 'Describe this picture in Chinese.'}, {'type': 'image_url', 'image_url': {'url': 'https://s2.loli.net/2025/11/15/NS4g8bRiYu7z31f.png'}}]
content_blocks : [{'type': 'text', 'text': 'Describe this picture in Chinese.'}, {'type': 'image', 'id': 'lc_4f10da0d-4c08-45f1-9244-fc1521628111', 'url': 'https://s2.loli.net/2025/11/15/NS4g8bRiYu7z31f.png'}]

标准块进、标准块出,URL 原样保留。原生块被翻译成了 image,LangChain 还补了一个 lc_ 开头的块 ID。用 content_blocks 初始化消息时会顺带填好 content,所以选哪种写法看你后面要拿哪个属性。

如果程序之外的系统也要读标准块,文档给了一个开关:把环境变量 LC_OUTPUT_VERSION 设成 v1,或者在初始化模型时传 output_version="v1",消息的 content 里存的就直接是标准块,不再需要惰性解析。

base64 图片要带 mime_type,文档写明了这一点:

1
2
3
4
5
6
mm3 = HumanMessage(content_blocks=[
{"type": "text", "text": "What is in this image?"},
{"type": "image", "base64": "iVBORw0KGgoAAAANSUhEUg==", "mime_type": "image/png"},
])
print(mm3.content_blocks)
# [{'type': 'text', 'text': 'What is in this image?'}, {'type': 'image', 'base64': 'iVBORw0KGgoAAAANSUhEUg==', 'mime_type': 'image/png'}]

图片这条链路能不能通,得看模型。/v1/models 里每个模型都带一个 input_modalities 字段,直接说明它收不收图:

1
2
{"id": "deepseek-flash", "name": "DeepSeek-V4.1-Flash", "context_window": 1048576, "input_modalities": ["text", "image"], "output_modalities": ["text"], ...}
{"id": "deepseek-v4-pro", "name": "DeepSeek-V4-Pro", "context_window": 1048576, "input_modalities": ["text"], "output_modalities": ["text"], ...}

我只留了 id、name、context_window 和两个模态字段,原返回里还有 effort、api_capabilities 这些,跟读图无关。整个列表里只有 deepseek-flash 的 input_modalities 带 image。

拿一张真实照片测。原图 1439×2326,缩到最长边 640 变成 396×640,JPEG quality 80,36KB,base64 之后 48072 个字符:

1
2
3
4
5
6
7
8
9
10
11
12
13
from PIL import Image
import base64, io

img = Image.open(r"D:\blog\source\images\IMG_20251026_131027.jpg").convert("RGB")
img.thumbnail((640, 640))
buf = io.BytesIO()
img.save(buf, format="JPEG", quality=80)
b64 = base64.b64encode(buf.getvalue()).decode()

msg = HumanMessage(content_blocks=[
{"type": "text", "text": "这张照片里有什么?用一句话描述主体和场景。"},
{"type": "image", "base64": b64, "mime_type": "image/jpeg"},
])

发给 deepseek-flash,标准块和 provider 原生块各跑一遍:

1
2
3
4
5
6
7
8
9
---- deepseek-flash 读图(标准 content_blocks)----
回答 : 画面主体是一个渺小的黑色背影,独自行走在蓝紫色调、长满茂密高草的广阔原野中,前景的虚化草穗增强了纵深感和孤寂的氛围。
usage_metadata : {'input_tokens': 248, 'output_tokens': 495, 'total_tokens': 743, 'input_token_details': {'cache_read': 0}, 'output_token_details': {'reasoning': 455}}
model_name : deepseek-flash

---- deepseek-flash 读图(provider 原生 image_url)----
回答 : 照片呈现的是一个孤独的背影,背对镜头站立在长满紫色野草的广阔田野之中。
usage_metadata : {'input_tokens': 248, 'output_tokens': 330, 'total_tokens': 578, 'input_token_details': {'cache_read': 0}, 'output_token_details': {'reasoning': 309}}
model_name : deepseek-flash

照片里就是一个渺小的背影站在长满芦苇的旷野上,两次都读对了。两种写法最终翻译成同一份 payload,input_tokens 都是 248。output_token_details 里还多出一个 reasoning,495 个输出 token 里 455 个花在推理上,这部分怎么流式吐出来放在 Streaming 那篇。

同一张图发给 deepseek-v4-pro,接口不报错,HTTP 200:

1
2
3
4
---- LangChain invoke: deepseek-v4-pro ----
没有报错,回答: 抱歉,我无法查看这张图片。请重新上传或确认图片格式,我就能帮你用一句话描述主体和场景。
model_name : deepseek-v4-pro
usage_metadata: {'input_tokens': 100, 'output_tokens': 128, 'total_tokens': 228, 'input_token_details': {'cache_read': 0}, 'output_token_details': {'reasoning': 101}}

让模型复述它收到的消息原文,能直接看到图片块被换成了占位符:

1
2
3
4
---- 追问 deepseek-v4-pro:你能看到的用户消息原文 ----
回答: 逐字复述本轮用户消息里的全部文本,包括任何方括号或尖括号标记,不要解释。

[Unsupported Image]

input_tokens 也印证了这件事:同一张图,flash 那边 248,v4-pro 只有 100,图根本没进 token 计数。纯文本模型遇到图不会拒绝请求,它只是看不见,然后在回答里道歉。这种失败最难查,没有异常栈,也没有 4xx,日志干净,用户拿到的是一句“我无法查看这张图片”。

最后说回 deepseek-chat。它现在是 deepseek-flash 的别名,传 deepseek-chat 进去,response_metadata 里的 model_name 会是 deepseek-flash(前面元数据那节就能看到),图也照读不误。我最早用三张 1×1 纯色 PNG 测“链路通不通”就是踩了这个坑,三次都答对,看着像 deepseek-chat 会看图,其实是背后的 flash 在看。纯色块太容易蒙对,验证视觉能力得用有真实内容的照片,答错和答对都要能分辨出来。官方模型列表里已经不列 deepseek-chat,选模型以 /v1/models 的 input_modalities 为准。

文档里也提醒过,不是所有模型都支持所有文件类型,PDF 这类还要看各家 provider 的具体要求。

元数据:id、usage_metadata、response_metadata

模型返回的 AIMessage 上挂了三个容易忽略的字段。跑一次普通对话看看:

1
2
3
4
5
6
7
8
9
10
11
12
13
messages = [
SystemMessage("You are a terse assistant. Answer in one short sentence."),
HumanMessage("What is the capital of France?"),
AIMessage("Paris."),
HumanMessage("And of Japan?"),
]
response = model.invoke(messages)

print("type :", type(response).__name__)
print("id :", response.id)
print("content :", repr(response.content))
print("usage_metadata :", response.usage_metadata)
print("response_metadata :", response.response_metadata)

真实输出:

1
2
3
4
5
type              : AIMessage
id : lc_run--01a10783-e4b5-7903-90e7-a4da6961de2c-0
content : 'Tokyo.'
usage_metadata : {'input_tokens': 35, 'output_tokens': 2, 'total_tokens': 37, 'input_token_details': {'cache_read': 0}, 'output_token_details': {}}
response_metadata : {'token_usage': {'completion_tokens': 2, 'prompt_tokens': 35, 'total_tokens': 37, 'prompt_tokens_details': {'cached_tokens': 0, ...}, 'prompt_cache_hit_tokens': 0, 'prompt_cache_miss_tokens': 35}, 'model_provider': 'openai', 'model_name': 'deepseek-flash', 'system_fingerprint': 'aeb56401ca74e127821c4f9126dcb669', 'id': 'f17d8ae8-582f-46e9-9a87-917f8d020db7', 'finish_reason': 'stop', 'logprobs': None}

id 是这条消息的标识,LangChain 生成,用于追踪和引用。

usage_metadata 是跨 provider 标准化过的 token 计数。input_tokens、output_tokens、total_tokens 三个数字最常用,input_token_details 和 output_token_details 放缓存命中、推理 token 之类的细项。DeepSeek 这里 input_token_details 只有 cache_read,没有音频字段。做成本统计就认这个字段,别去解析 response_metadata。

response_metadata 是 provider 原始返回,没做标准化。上面能看到 model_name 是 deepseek-flash(虽然我传的是 deepseek-chat),finish_reason 是 stop,还有服务端的请求 ID 和指纹。调优、排查、对账的时候有用,写通用代码时不要依赖它的键名。

有个细节容易踩:我传了 4 条消息进去,input_tokens 只有 35。手工插入的那条 AIMessage("Paris.") 也在历史里,只是这段对话本来就短。

流式输出拿到的是 AIMessageChunk,不是一个完整的 AIMessage。它和消息对象一样支持 +,把一路收到的 chunk 依次相加,最后得到的就是完整消息。要拿 token 统计和工具调用,得等拼完之后再读,中途的 chunk 上只有片段。

序列化与还原

存历史、断点续聊、把会话扔进数据库,都绕不开消息和普通数据之间的转换。

最简单的形式是 OpenAI chat completions 风格的字典,model.invoke 直接吃:

1
2
3
4
5
6
response = model.invoke([
{"role": "system", "content": "You are a terse assistant."},
{"role": "user", "content": "Reply with the single word: pong"},
])
print(type(response).__name__, repr(response.content))
# AIMessage 'pong'

反向转换用 convert_to_messages,它按 role 把字典变成对应的消息对象:

1
2
3
4
5
6
7
8
9
10
from langchain_core.messages import convert_to_messages

raw = [
{"role": "system", "content": "You are a helpful assistant."},
{"role": "user", "content": "Hi"},
{"role": "assistant", "content": "Hello! How can I help?"},
{"role": "tool", "content": "Sunny, 22C", "tool_call_id": "call_abc123"},
]
for o in convert_to_messages(raw):
print(type(o).__name__, "| type =", o.type, "| tool_call_id =", getattr(o, "tool_call_id", None))

真实输出:

1
2
3
4
SystemMessage | type= system | content= 'You are a helpful assistant.' | tool_call_id= None
HumanMessage | type= human | content= 'Hi' | tool_call_id= None
AIMessage | type= ai | content= 'Hello! How can I help?' | tool_call_id= None
ToolMessage | type= tool | content= 'Sunny, 22C' | tool_call_id= call_abc123

注意 tool 角色的字典里必须带 tool_call_id,否则还原出来的 ToolMessage 没法配对。assistant 带工具调用时,字典里还要有 tool_calls 字段。

消息对象转回普通字典用 model_dump:

1
{"content": "What is the capital of France?", "additional_kwargs": {}, "response_metadata": {}, "type": "human", "name": null, "id": "msg_human_001"}

要完整往返、连 tool_calls 一起保住,用文档给的 dumpd / load:

1
2
3
4
5
6
7
8
9
from langchain_core.load import dumpd, load

ai = AIMessage(
content="",
tool_calls=[{"name": "get_weather", "args": {"location": "Paris"}, "id": "call_abc123"}],
)
serialized = dumpd(ai)
restored = load(serialized)
print(type(restored).__name__, restored.tool_calls == ai.tool_calls)

真实输出:

1
2
3
4
5
dumpd type: dict
{"lc": 1, "type": "constructor", "id": ["langchain", "schema", "messages", "AIMessage"], "kwargs": {"content": "", "type": "ai", "tool_calls": [{"name": "get_weather", "args": {"location": "Paris"}, "id": "call_abc123", "type": "tool_call"}], "invalid_tool_calls": []}}
restored : AIMessage
tool_calls: [{'name': 'get_weather', 'args': {'location': 'Paris'}, 'id': 'call_abc123', 'type': 'tool_call'}]
equal to original: True

dumpd 出来的东西带 lc 标记和类路径,load 靠这些信息把类重新实例化。文档给了个明确警告:load() 会执行反序列化构造,可能触发副作用,别对不可信来源的数据调用它。运行时还会弹一条 beta 警告和一个关于 allowed_objects 默认值的弃用提醒,生产代码里最好显式传 allowed_objects。

整条链路可以画成这样:

三种形式的分工:

形式 用什么转 保留了什么 适用场景
{"role": ..., "content": ...} convert_to_messages 角色和正文,工具调用要手动补字段 和外部系统交换、手写历史
model_dump() load 不认,需要自己拼 字段完整但不含类信息 只读展示、存日志
dumpd / load 成对使用 类信息、tool_calls 全保留 持久化会话、断点续聊

长对话里的裁剪

历史只会变长,context window 不会。文档列了几种做法:裁剪、删除、摘要、自定义过滤。

langchain_core.messages 里有个 trim_messages 工具,按 token 预算保留尾部消息。我拿一段 8 条消息的历史试了试,用 token_counter=len 把”token”当成消息条数来数,方便观察:

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
from langchain.messages import SystemMessage, HumanMessage, AIMessage
from langchain_core.messages import trim_messages

history = [
SystemMessage("You are a helpful assistant."),
HumanMessage("Hi, my name is Bob."),
AIMessage("Nice to meet you, Bob."),
HumanMessage("Write a poem about cats."),
AIMessage("Cats sit on walls..."),
HumanMessage("Now do the same for dogs."),
AIMessage("Dogs run in parks..."),
HumanMessage("What is my name?"),
]

trimmed = trim_messages(
history,
max_tokens=5,
strategy="last",
token_counter=len,
include_system=True,
start_on="human",
)
print("original:", len(history), "-> trimmed:", len(trimmed))
for m in trimmed:
print(" ", type(m).__name__, repr(m.content)[:60])

真实输出:

1
2
3
4
5
6
7
8
9
10
11
12
13
original length: 8
trimmed length: 4 (token_counter=len, max_tokens=5, strategy=last)
SystemMessage 'You are a helpful assistant.'
HumanMessage 'Now do the same for dogs.'
AIMessage 'Dogs run in parks...'
HumanMessage 'What is my name?'

include_system=False -> length: 5
HumanMessage 'Write a poem about cats.'
AIMessage 'Cats sit on walls...'
HumanMessage 'Now do the same for dogs.'
AIMessage 'Dogs run in parks...'
HumanMessage 'What is my name?'

strategy="last" 从尾部往回留,include_system=True 保住系统指令,start_on="human" 保证结果以用户消息开头。上面第一次裁剪留下 4 条:系统消息加上末尾三轮对话。把 include_system 关掉就变成 5 条,系统指令被丢掉,从用户消息开始。

文档”Trim messages”那一节的示例代码用的其实是 RemoveMessage,不是 trim_messages。它挂在 @before_model 中间件上,每次调模型前重写消息列表:

1
2
3
4
5
6
7
8
9
10
11
12
13
from langchain.messages import RemoveMessage
from langgraph.graph.message import REMOVE_ALL_MESSAGES
from langchain.agents.middleware import before_model

@before_model
def trim_messages(state, runtime):
"""Keep only the last few messages to fit context window."""
messages = state["messages"]
if len(messages) <= 3:
return None
first_msg = messages[0]
recent_messages = messages[-3:] if len(messages) % 2 == 0 else messages[-4:]
return {"messages": [RemoveMessage(id=REMOVE_ALL_MESSAGES), first_msg, *recent_messages]}

RemoveMessage(id=REMOVE_ALL_MESSAGES) 是清空整个列表的哨兵值,后面跟着的就是新历史。要删单条,用 RemoveMessage(id=消息的 id)。这套机制依赖 add_messages reducer,默认的 AgentState 自带。

摘要走的是内置的 SummarizationMiddleware,文档里的用法是给它一个便宜些的模型、一个触发条件和一个保留量,比如 trigger=("tokens", 4000)、keep=("messages", 20),意思是在历史超过 4000 token 时触发,摘要后保留最近 20 条消息。

删除会丢信息,文档说得很直白:裁剪掉的消息里的信息就没了。上面那段 include_system=False 的输出里,Hi, my name is Bob. 被裁掉了,如果后面接着问”我叫什么”,模型只能回答不知道。信息必须留住,就换成 SummarizationMiddleware,让模型把早期历史压成摘要。

裁剪还有一条硬约束:结果历史必须合法。有的 provider 要求历史以 user 消息开头;带工具调用的 assistant 消息后面必须跟对应的 tool 结果。裁剪的边界要是正好切在这对中间,下一个请求就是 400。

直接调模型,还是交给 agent

前面所有例子都是直接调模型,工具执行、消息拼装全靠自己。换成 create_agent 会怎样:

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
from langchain.agents import create_agent
from langgraph.checkpoint.memory import InMemorySaver

agent = create_agent(
model=model,
tools=[get_weather],
system_prompt="You are a helpful assistant.",
checkpointer=InMemorySaver(),
)
config = {"configurable": {"thread_id": "demo-1"}}
result = agent.invoke(
{"messages": [{"role": "user", "content": "What's the weather in Paris?"}]},
config,
)
for m in result["messages"]:
print(type(m).__name__, "| content =", repr(m.content)[:60],
"| tool_call_id =", getattr(m, "tool_call_id", None))

真实输出:

1
2
3
4
5
6
7
8
9
10
11
12
---- direct model call ----
returned type : AIMessage
content : "I'll check the current weather in Paris for you."
tool_calls : [{'name': 'get_weather', 'args': {'location': 'Paris'}, 'id': 'call_00_QWOud95DNoG98VytHC2i0841', 'type': 'tool_call'}]

---- through an agent ----
state keys : ['messages']
message count : 4
HumanMessage | content= "What's the weather in Paris?" | tool_calls= None | tool_call_id= None
AIMessage | content= "I'll check the current weather in Paris for you." | tool_calls= [{...}] | tool_call_id= None
ToolMessage | content= 'Sunny, 22C' | tool_calls= None | tool_call_id= call_00_NRo1kEWP2HT5qlQDMdDm8013
AIMessage | content= "The weather in Paris is currently sunny with a temperature of 22°C (72°F)." | tool_calls= [] | tool_call_id= None

直接调模型,一次 invoke 换回一条 AIMessage,里面只有 tool_calls,没有工具结果。agent 一次 invoke 把整个循环跑完,返回的状态里躺着 4 条消息:用户输入、带调用的 AI 消息、工具结果、最终回答。ToolMessage 是 agent 自己生成的,tool_call_id 也是它填的,你不用管配对。

循环长这样:

配上 checkpointer 后,同一个 thread_id 的下一轮会接着旧历史。我接着问”And tomorrow?”,消息数从 4 变成 6。这一轮模型没调工具,直接回答自己只有当前天气数据、给不了预报,历史里那条 ToolMessage 还在。

1
2
3
---- follow-up turn on the same thread ----
message count : 6
last content : "I'm sorry, but I can only retrieve the **current** weather - I don't have access to forecast data for tomorrow. ..."

两条路的差别在于谁持有历史。直接调模型,历史是你手里的一个列表,随时能改、能存、能裁剪。用 agent,历史在图的 state 里,append-only,改动要走中间件和 RemoveMessage。要做精细控制,比如手工插入伪造的 AIMessage、按业务规则改写历史,直接调模型更顺手。要做多轮工具循环,agent 省掉的配对和循环代码不少。

小结

  • 消息是模型和工具之间唯一的数据载体。模型无状态,历史靠消息列表重建;工具调用的请求和结果也各是一条消息。
  • 四类消息的 type 是 system、human、ai、tool。AIMessage 承载 tool_calls,ToolMessage 用 tool_call_id 回应它,配不上就是 400。
  • content 是原始载荷,content_blocks 是标准化视图。多模态用 content_blocks 写更省事,base64 图片记得带 mime_type。
  • usage_metadata 做成本统计,response_metadata 留原始返回做排查。持久化用 dumpd / load,和外部系统交换用 {"role": ..., "content": ...}。
  • 长对话裁剪优先用 trim_messages 或 before_model 中间件加 RemoveMessage。裁之前确认历史合法,别把工具调用和它的结果切开。