
Llama3-Instruct模型在启用工具调用(Function Calling)后,若未合理配置模板与推理逻辑,会将所有用户输入(包括“你是谁”等通用问题)强行解析为工具调用格式。本文详解根本原因、修复方案及完整可运行示例。
当llama3-instruct启用工具调用后,模型可能对所有输入(如“你是谁?”)都只生成函数调用json,而非自然语言回复——这并非模型能力缺陷,而是因chat template未区分“需调用工具”与“直接回答”两类意图,且缺少系统级指令约束与后处理机制。
在基于transformers本地加载Llama3-Instruct(如Llama-3.2-1B-Instruct)并集成工具调用时,出现“所有问题都只返回函数调用”是典型配置失配现象。其本质原因有三:
-
tokenizer.apply_chat_template(..., tools=tools)强制启用了工具感知模式:该方法会将工具定义注入系统提示,并默认引导模型优先考虑工具调用,即使当前query完全无需外部操作; -
缺失明确的系统指令(SYSTEM prompt)来平衡行为:原生Llama3-Instruct无内置工具调用意识,必须通过
system角色明确告知模型“仅在必要时调用工具,否则直接回答”; -
缺少调用判定与结果路由逻辑:模型本身从不真正执行函数——它仅生成符合规范的JSON片段(如
{"name": "get_current_temperature", "arguments": {"location": "Beijing"}}),后续需由Orchestrator(如LangChain、Llama-Worker或自定义解析器)识别、校验、执行并拼接结果。若跳过此步,仅打印原始输出,就会误以为“模型只会调用工具”。
✅ 正确做法:分阶段控制行为流
一、修复Chat Template与系统提示
不要依赖apply_chat_template(..., tools=...)的自动注入(它过于激进)。改为手动构造结构化提示,显式声明工具能力边界:
from transformers import AutoTokenizer, AutoModelForCausalLM
import json
import torch
checkpoint = "models/Llama-3.2-1B-Instruct"
tokenizer = AutoTokenizer.from_pretrained(checkpoint)
model = AutoModelForCausalLM.from_pretrained(
checkpoint,
torch_dtype=torch.bfloat16,
device_map="cpu"
)
# 显式定义工具描述(遵循OpenAI-style function schema)
tools = [
{
"type": "function",
"function": {
"name": "get_current_temperature",
"description": "Get the current temperature at a location.",
"parameters": {
"type": "object",
"properties": {
"location": {"type": "string", "description": "City, Country"},
"unit": {"type": "string", "enum": ["celsius", "fahrenheit"]}
},
"required": ["location", "unit"]
}
}
},
{
"type": "function",
"function": {
"name": "get_current_wind_speed",
"description": "Get the current wind speed in km/h at a given location.",
"parameters": {
"type": "object",
"properties": {
"location": {"type": "string", "description": "City, Country"}
},
"required": ["location"]
}
}
}
]
# 构造严格分层的messages —— 关键:system需强调「按需调用」
messages = [
{
"role": "system",
"content": (
"You are a helpful AI assistant. You have access to tools to fetch real-time data. "
"ONLY call a tool when the user's question explicitly requires external information (e.g., weather, time, live data). "
"For general knowledge, self-introduction, definitions, or reasoning tasks — answer directly in natural language. "
"Do NOT call any tool for questions like 'who are you?', 'what is apple?', or 'explain quantum physics'. "
f"Available tools: {json.dumps(tools, ensure_ascii=False)}"
)
},
{"role": "user", "content": "Hey, who are you ?"}
]
# 手动应用模板(不传tools参数!)
prompt = tokenizer.apply_chat_template(
messages,
tokenize=False,
add_generation_prompt=True,
# ⚠️ 注意:此处不传 tools=tools,避免自动注入强引导逻辑
)
inputs = tokenizer(prompt, return_tensors="pt").to(model.device)
# 生成时添加停止词,防止模型过度生成JSON
output = model.generate(
**inputs,
max_new_tokens=256,
do_sample=False,
temperature=0.1,
stop_strings=["<|eot_id|>", "```"], # Llama3常用结束符
tokenizer=tokenizer
)
response = tokenizer.decode(output[0][inputs.input_ids.shape[1]:], skip_special_tokens=True)
print("Raw model output:")
print(repr(response))二、添加智能解析层(必备后处理)
模型输出需经解析才能判断意图。以下是一个轻量级判定函数:
import re
import json
def parse_tool_call_or_answer(text: str) -> dict:
"""
尝试从模型输出中提取工具调用JSON,否则返回普通回答
返回: {"type": "tool_call" | "answer", "content": str}
"""
# 匹配标准function call JSON块(Llama3常用格式)
json_match = re.search(r'\{.*?"name"\s*:\s*".*?".*?\}', text, re.DOTALL)
if json_match:
try:
call = json.loads(json_match.group(0))
if isinstance(call, dict) and "name" in call and "arguments" in call:
return {"type": "tool_call", "content": call}
except (json.JSONDecodeError, KeyError):
pass
# 若无有效JSON,视为直接回答
# 清理可能的残留符号(如开头的```json、结尾的```)
clean_text = re.sub(r'^```(?:json)?\s*|\s*```$', '', text).strip()
return {"type": "answer", "content": clean_text}
# 使用示例
result = parse_tool_call_or_answer(response)
if result["type"] == "tool_call":
print("→ 模型建议调用工具:", result["content"])
# 此处应由Orchestrator执行:call_tool(result["content"]["name"], result["content"]["arguments"])
else:
print("→ 模型直接回答:", result["content"])三、关键注意事项总结
- ✅ 永远不要假设模型会“执行”工具:LLM仅生成文本,工具调用是应用层责任(LangChain、Llama-Worker、OpenClaw等均提供成熟Router);
- ✅ 系统提示(system prompt)比工具schema更重要:清晰界定“何时调用、何时直答”能显著降低误触发率;
- ✅ 避免滥用
tools=参数:apply_chat_template(tools=...)适用于OpenAI兼容API场景,本地transformers推理建议手动构造更可控; - ✅ 验证模型是否真支持工具调用:Llama3.2-1B-Instruct原生不支持Function Calling,需使用已微调/增强的版本(如Ollama中
llama3.2-tools镜像)或自行添加LoRA适配器; - ✅ 调试技巧:对简单问题(如“你好”)强制设置
temperature=0.0并观察输出,若仍返回JSON,则说明模板或权重存在强偏置,需检查训练/量化配置。
通过以上三步重构,你的Llama3将恢复理性:对“你是谁?”给出人格化回应,对“北京现在温度多少?”才严谨输出工具调用结构——这才是生产级Agent应有的行为范式。

















