function-calling的使用-agent工作的基础

前一章分别用 openai / anthropic 两套格式调用了一次底层 llm。后面的开发就固定用 openai 的格式。

这一章算是大模型发展史上的一次里程碑:function-call 出现了。在这之前,大模型只能推理、输出,并不能帮用户操作任何事情。写代码、操作浏览器、整理文件都做不了。这时候的 llm 还是一个没有肢体的超级大脑。

function-call 出现之后,llm 才和宿主的机器、代码产生关联。它可以调用本地的命令行和代码,去感知用户机器的状态,然后再思考、执行对应的函数(react模式)。

最早期不是所有 llm 都支持 function-call。现在这已经是大语言模型的基本能力了。不得不感叹,AI 的发展之快。

本章以查询天气为例,让 deepseek 调用天气查询,回复用户有关天气的问题。

工具怎么声明

openai 的官方协议格式可以在这里查。

我们定义的 tools 如下:

TOOLS = [
    {
        "type": "function",
        "function": {
            "name": "get_weather",
            "description": "按地址查询当前天气",
            "parameters": {
                "type": "object",
                "properties": {
                    "address": {
                        "type": "string",
                        "description": "地点,例如北京、上海市浦东新区",
                    }
                },
                "required": ["address"],
            },
        },
    }
]

对应的执行函数就简单点:

def get_weather(address):
    return f"{address}是晴天!"

带上工具再调用:

response = post_request(
    f"{base_url}/chat/completions",
    {"model": model_name, "messages": messages, "tools": TOOLS},
)

由于我们并没有接入类似FastMCP这样的库,帮我们去处理schema、调用。所以在调用完成后,我们需要自己去 解析结果,查看是否触发了tool_calls。

模型发起工具调用

以提问「今天杭州天气怎么样?」为例,追踪这次返回:

{
  "id": "684a9b1f-a7a0-4853-b80f-80f97caf23e6",
  "object": "chat.completion",
  "created": 1791521467,
  "model": "deepseek-flash",
  "choices": [
    {
      "index": 0,
      "message": {
        "role": "assistant",
        "content": "",
        "reasoning_content": "The user asks about Hangzhou weather today. I should call the weather tool.",
        "tool_calls": [
          {
            "index": 0,
            "id": "call_00_q4pSstRvTU14wNBWBER22266",
            "type": "function",
            "function": {
              "name": "get_weather",
              "arguments": "{\"address\": \"杭州\"}"
            }
          }
        ]
      },
      "logprobs": null,
      "finish_reason": "tool_calls"
    }
  ],
  "usage": {
    "prompt_tokens": 316,
    "completion_tokens": 55,
    "total_tokens": 371,
    "prompt_tokens_details": {
      "cached_tokens": 128
    },
    "completion_tokens_details": {
      "reasoning_tokens": 16
    },
    "prompt_cache_hit_tokens": 128,
    "prompt_cache_miss_tokens": 188
  },
  "system_fingerprint": "aeb56401ca74e127821c4f9126dcb669"
}

可以看到,基本的消息格式和前一章是一样的。不过这里停止推理的理由是:tool_calls 我们可以根据这个字段,或者直接解析tool_calls,循环将要对应的工具结果返回。注意:这里不能漏掉、或是跳过,对于llm来说,未拿到对应的tools_result则会被认为是调用错误。

返回工具调用结果

这块的代码belike:

def handle_tools(messages: list, response: dict):
    message = response["choices"][0]["message"]
    tool_calls = message.get("tool_calls") or []
    if not tool_calls:
        return response
    messages.append(message)
    for call in tool_calls:
        function_name = call["function"]["name"]
        args = json.loads(call["function"]["arguments"] or "{}")
        content = "工具调用失败" if function_name != "get_weather" else get_weather(args.get("address", ""))
        messages.append(
            {
                "role": "tool",
                "tool_call_id": call["id"],
                "content": content,
            }
        )
    # 递归调用
    return handle_tools(messages, post_request(
        f"{base_url}/chat/completions",
        {"model": model_name, "messages": messages, "tools": TOOLS},
    ))

这一块需要注意的是:调用工具返回后,进入大模型的react,推理出来的结果可能还需要继续工具调用(这种场景很常见,比如先查询完地址,再根据地址查询天气等)。所以这里处理工具调用要递归处理,直到llm认为不需要调用工具为止。这里可能大模型会抽风(早期),所以为了安全起见,可以设置一个最大递归深度,防止出现死循环。
由于我们只有get_weather函数,所以只执行这个(TOOLS告知了tool_list的schema,理论上不会调用额外的工具,这里只是安全性校验)。

工具调用到这里就结束了,但是工程上,有很多很多事情要去做,比如工具的权限校验(human in the loop),需要做的事情还非常多。

工具结果送回去之后,模型这次不再调用工具,直接给出最终回复:

{
  "id": "d25e81d5-c445-4e66-9fd6-59f85d745f9f",
  "object": "chat.completion",
  "created": 1791522161,
  "model": "deepseek-flash",
  "choices": [
    {
      "index": 0,
      "message": {
        "role": "assistant",
        "content": "今天杭州是晴天 ☀️,适合外出活动。",
        "reasoning_content": ""
      },
      "logprobs": null,
      "finish_reason": "stop"
    }
  ],
  "usage": {
    "prompt_tokens": 386,
    "completion_tokens": 13,
    "total_tokens": 399,
    "prompt_tokens_details": {
      "cached_tokens": 128
    },
    "completion_tokens_details": {
      "reasoning_tokens": 0
    },
    "prompt_cache_hit_tokens": 128,
    "prompt_cache_miss_tokens": 258
  },
  "system_fingerprint": "aeb56401ca74e127821c4f9126dcb669"
}

如果接入mcp的话,就能帮助我们处理schema的生成和函数的调用,以及外接mcp-server,但这不是我们章节的主要内容了。

有了function-call,便能让llm帮助你写代码、做ppt、操作电脑甚至是其它硬件设备,给了llm更多的想象空间,从此agent开发才真正进入百花齐放的时代。

在下一章节,我们将此工程改在成一个能agent-loop、维护上下文、支持流式聊天的最小agent项目

仓库地址:https://github.com/qingpingwang/learn

交流群(QQ):图形/渲染/音视频/AI应用交流群(523219063)