上下文工程1-流式agent对话

上下文工程这一章主要涉及到流式输出和上下文的管理。这一节先做一个最小的 agent,它要能:

流式输出协议格式

先来看看 deepseek 的流式输出调用:

DeepSeek 流式输出调用

请求里还有一些前面没用过的参数,例如:是否关闭思考、推理等级。把 stream 设为 true 之后,这里以提问「杭州的天气怎么样?」,拿到的是一串 chunk。

流式输出是 step-by-step 的,每次一个或几个 token。客户端看起来就是打字机效果,首token出现会比非流式快很多。

{
  "id": "772281c9-56b1-4765-94b4-97f805e3067a",
  "object": "chat.completion.chunk",
  "created": 1791531311,
  "model": "deepseek-flash",
  "choices": [
    {
      "index": 0,
      "delta": {
        "role": "assistant",
        "content": null,
        "reasoning_content": ""
      },
      "finish_reason": null
    }
  ]
}

思考阶段,token落在 reasoning_content,content为空,如有思考展示需求需要特殊处理。

{"content": null, "reasoning_content": "The"}
{"content": null, "reasoning_content": " user"}
{"content": null, "reasoning_content": " asks"}
{"content": null, "reasoning_content": " about"}
{"content": null, "reasoning_content": " today"}
{"content": null, "reasoning_content": "'s"}
{"content": null, "reasoning_content": " weather"}
{"content": null, "reasoning_content": " in"}
{"content": null, "reasoning_content": " Hang"}
{"content": null, "reasoning_content": "zhou"}
{"content": null, "reasoning_content": "."}
{"content": null, "reasoning_content": " I"}
{"content": null, "reasoning_content": " should"}
{"content": null, "reasoning_content": " call"}
{"content": null, "reasoning_content": " the"}
{"content": null, "reasoning_content": " weather"}
{"content": null, "reasoning_content": " tool"}
{"content": null, "reasoning_content": ".\n\n"}
{"content": null, "reasoning_content": "Wait"}
{"content": null, "reasoning_content": ","}
{"content": null, "reasoning_content": " the"}
{"content": null, "reasoning_content": " system"}
{"content": null, "reasoning_content": " prompt"}
{"content": null, "reasoning_content": " says"}
{"content": null, "reasoning_content": " I"}
{"content": null, "reasoning_content": " can"}
{"content": null, "reasoning_content": " invoke"}
{"content": null, "reasoning_content": " tools"}
{"content": null, "reasoning_content": "."}
{"content": null, "reasoning_content": " Let"}
{"content": null, "reasoning_content": " me"}
{"content": null, "reasoning_content": " call"}
{"content": null, "reasoning_content": " get"}
{"content": null, "reasoning_content": "_"}
{"content": null, "reasoning_content": "weather"}
{"content": null, "reasoning_content": " with"}
{"content": null, "reasoning_content": " address"}
{"content": null, "reasoning_content": " "}
{"content": null, "reasoning_content": "杭州"}
{"content": null, "reasoning_content": "."}

推理结束后开始吐工具。第一块带上 id 和函数名,参数是空的:

{
  "choices": [
    {
      "index": 0,
      "delta": {
        "tool_calls": [
          {
            "index": 0,
            "id": "call_00_mE6AiuBEG3YwkCmN9Rur0420",
            "type": "function",
            "function": {
              "name": "get_weather",
              "arguments": ""
            }
          }
        ]
      },
      "finish_reason": null
    }
  ]
}

后面的块只有参数碎片,拼起来才是 {"address": "杭州"}:

{"tool_calls": [{"index": 0, "function": {"arguments": "{"}}]}
{"tool_calls": [{"index": 0, "function": {"arguments": "\""}}]}
{"tool_calls": [{"index": 0, "function": {"arguments": "address"}}]}
{"tool_calls": [{"index": 0, "function": {"arguments": "\""}}]}
{"tool_calls": [{"index": 0, "function": {"arguments": ": "}}]}
{"tool_calls": [{"index": 0, "function": {"arguments": "\""}}]}
{"tool_calls": [{"index": 0, "function": {"arguments": "杭州"}}]}
{"tool_calls": [{"index": 0, "function": {"arguments": "\""}}]}
{"tool_calls": [{"index": 0, "function": {"arguments": "}"}}]}

最后一块停住,finish_reason 是 tool_calls:

{
  "choices": [
    {
      "index": 0,
      "delta": {
        "content": "",
        "reasoning_content": null
      },
      "finish_reason": "tool_calls"
    }
  ],
  "usage": {
    "prompt_tokens": 306,
    "completion_tokens": 79,
    "total_tokens": 385,
    "completion_tokens_details": {
      "reasoning_tokens": 40
    }
  }
}

思考模式下,先走一轮思考:choices -> delta -> reasoning_content,这时 content 是空的。

推理完成后,模型判断要调用工具,就开始流式输出工具信息:函数名和参数。最后 finish_reason 停在 tool_calls,等待工具调用。

agent-loop实现

这里的实现和之前的非流式递归解析工具类似。我们需要一个循环,在每轮对话完成后调用对应的工具。

def event_stream(messages):
    response = requests.post(
        f"{base_url}/chat/completions",
        headers={"Authorization": f"Bearer {api_key}"},
        json={"model": model_name, "messages": messages, "tools": TOOLS, "stream": True},
        stream=True,
        timeout=120,
    )
    response.raise_for_status()
    try:
        for line in response.iter_lines():
            if not line:
                continue
            text = line.decode()
            if not text.startswith("data: "):
                continue
            data = text[6:]
            if data == "[DONE]":
                break
            yield json.loads(data)
    finally:
        response.close()


@app.post("/stream_chat")
def stream_chat(body: MessageIn):
    messages = context_list[body.context_id]["messages"]
    messages.append({"role": "user", "content": body.message})

    def generate():
        # 循环调用流式事件,有工具执行
        while True:
            # 单次调用的工具调用、推理内容、消息内容
            function_calls = []
            reasoning_content = ""
            message_content = ""
            for chunk in event_stream(messages):
                for choice in chunk["choices"]:
                    tool_calls = choice["delta"].get("tool_calls", [])
                    reasoning_content += choice["delta"].get("reasoning_content") or ""
                    message_content += choice["delta"].get("content") or ""
                    for tool_call in tool_calls:
                        if "id" in tool_call:
                            function_calls.append({
                                "id": tool_call["id"],
                                "name": "",
                                "arguments": ""
                            })
                        if "function" in tool_call:
                            function_calls[-1]["name"] += tool_call["function"].get("name") or ""
                            function_calls[-1]["arguments"] += tool_call["function"].get("arguments") or ""
                yield f"data: {json.dumps(chunk, ensure_ascii=False)}\n\n"
            # 如果没有工具需要执行,则sse结束
            if not function_calls:
                if message_content:
                    messages.append({"role": "assistant", "content": message_content})
                break
            # 工具消息组装进上下文
            messages.append({
                "role": "assistant",
                "content": "",
                "reasoning_content": reasoning_content,
                "tool_calls": [
                    {
                        "id": call["id"],
                        "type": "function",
                        "function": {"name": call["name"], "arguments": call["arguments"]},
                    }
                    for call in function_calls
                ],
            })
            # 工具执行结果组装进上下文
            for function_call in function_calls:
                content = "工具调用失败" if function_call["name"] != "get_weather" else get_weather(json.loads(function_call["arguments"]).get("address", ""))
                messages.append({
                    "role": "tool",
                    "tool_call_id": function_call["id"],
                    "content": content,
                })
                yield_data = {
                    "choices": [{
                        "delta": {
                            "tool_calls": [{
                                "id": function_call["id"],
                                "type": "function_result",
                                "function": {"name": function_call["name"], "arguments": function_call["arguments"]},
                                "content": content,
                            }]
                        }
                    }]}
                yield f"data: {json.dumps(yield_data, ensure_ascii=False)}\n\n"
        # 单次调用的结束,进入下一轮推理

    return StreamingResponse(generate(), media_type="text/event-stream")

这里的 while 循环,实际上就是 agent 的 react 过程:思考 -> 调用工具 -> 拿到结果 -> 继续思考。

我们这里维护了一个简单的对话列表,没有持久化,只在内存里。包含基本的创建、销毁、查询。每轮消息完成后,把完整的消息再添加进上下文的 messages:llm 回复、工具调用、工具结果。

这一轮 不需要调用工具,对话就结束了。有工具调用,就把工具调用和执行结果装进消息上下文,继续推理。

流式输出测试

到这里,整个 react-agent 已经基本完成了。缺少的上下文压缩,放到下一章。

现在要解析 SSE,测试连续对话。思考过程先不打印,只输出 content,再加上工具调用和执行结果:

def stream_chat(context_id, message):
    data = json.dumps({"context_id": context_id, "message": message}).encode()
    req = urllib.request.Request(
        base + "/stream_chat",
        data=data,
        headers={"Content-Type": "application/json"},
        method="POST",
    )
    with urllib.request.urlopen(req) as resp:
        for raw in resp:
            line = raw.decode().strip()
            if not line.startswith("data: "):
                continue
            chunk = json.loads(line[6:])
            for choice in chunk["choices"]:
                content = (choice.get("delta") or {}).get("content")
                if content:
                    print(content, end="", flush=True)
                # 输出工具调用
                tool_calls = choice.get("delta", {}).get("tool_calls", [])
                for tool_call in tool_calls:
                    tool_call_type = tool_call.get("type", "")
                    if tool_call_type == "function_result":
                        print("\nresult:", tool_call["id"], tool_call["function"]["name"], tool_call["function"]["arguments"], tool_call["content"])
                    else:
                        print(tool_call.get("id", ""), tool_call["function"].get("name", ""), tool_call["function"].get("arguments", ""), end="", flush=True)
    print()

测试一下输出:

DeepSeek 流式输出测试

这时 agent 已经有了上下文记忆、工具调用,以及执行完成后继续 react。 再下一章节,我们将会加上上下文压缩的相关内容,避免大模型超出上下文的限制。

仓库地址:https://github.com/qingpingwang/learn

交流群(QQ):图形/渲染/音视频/AI应用交流群(523219063)