一、什么是结构化输出?
LangChain 的结构化输出(Structured Output)指的是:
要求模型最终返回一个符合预定义结构的数据对象(比如固定字段的 JSON、Pydantic 模型、TypedDict),而不再是无格式的自然语言文本。
核心目标:把"自然语言回答"变成"程序可以稳定消费的数据"。
举个例子:
不是让模型输出:
盗梦空间在2010年上映,导演是克里斯托弗·诺兰,评分9.3。而是让它输出成这样的结构:
{
"title": "盗梦空间",
"year": 2010,
"director": "克里斯托弗·诺兰",
"rating": 9.3
}价值有三个方面:
更容易被代码处理:下游系统可以直接读字段,不用再从自然语言里做解析。
结果更稳定:减少"模型说法变了但意思差不多"导致的解析失败。
更适合工程化:适用于表单抽取、分类、路由、工具参数生成、工作流状态传递等场景。
二、传统方式 vs 结构化输出
传统方式(繁琐、不推荐)
# 1. 提示词要求 JSON
prompt = "以JSON格式返回:{name, age, occupation}"
response = model.invoke(prompt)
# 2. 手动解析
import json
data = json.loads(response.content)
# 3. 手动验证类型
if not isinstance(data['age'], int):
raise ValueError("age must be int")
# 4. 手动创建对象
person = Person(**data)结构化输出(一步到位)
structured_llm = model.with_structured_output(Person)
person = structured_llm.invoke("张三是一名 30 岁的软件工程师")
# ✅ 自动解析、验证、创建对象为什么结构化输出这么受欢迎?
在没有 Pydantic 等结构化方案之前,开发者需要写大量 Prompt 苦口婆心地求模型"请返回 JSON,不要带任何解释",然后自己写繁琐的 json.loads() 和 try...except。
有了 Pydantic 等方案结合 .with_structured_output() 之后:
Prompt 变干净了:字段的
description直接充当了 Prompt 的一部分。类型安全:编辑器能自动补全,运行前就能做类型检查。
极其稳定:依托模型厂商底层的 JSON 模式,输出错误率降到极低。
三、四种模式总览
LangChain 1.x 支持多种 Schema 与结构化输出方式:
关键区别:
只有 Pydantic 返回的是 Schema 类实例,其余三种都返回字典。
只有 Pydantic 在类型不匹配时会抛出异常(强校验)。
模型支持情况:大部分现代模型支持(OpenAI gpt-4 / gpt-3.5-turbo、Anthropic claude-3、Groq llama-3 等,通过函数调用)。某些旧模型不支持,此时 LangChain 会回退到"提示词 + JSON 解析"。
四、模式一:Pydantic(生产首选)
Pydantic 在运行时强制执行类型提示,确保数据正确性和一致性,是生产场景首选。
4.1 基本使用三要素
所有结构化输出的数据模型都必须继承
BaseModel。使用类型提示:
str、int、float、List[xxx]、Optional[xxx]等。用
Field()添加字段默认值和描述,帮助 LLM 理解字段含义。
完整示例:
# 1. 模型初始化
from langchain.chat_models import init_chat_model
from dotenv import load_dotenv
import os
load_dotenv(override=True)
CLOSEAI_API_KEY = os.getenv("CLOSEAI_API_KEY")
CLOSEAI_BASE_URL = os.getenv("CLOSEAI_BASE_URL")
model = init_chat_model(
model="gpt-5.4-mini",
model_provider="openai",
api_key=CLOSEAI_API_KEY,
base_url=CLOSEAI_BASE_URL
)
# 2. 定义 Pydantic 模型
from pydantic import BaseModel, Field
class Person(BaseModel):
"""人物信息"""
name: str = Field(description="姓名")
age: int = Field(description="年龄")
occupation: str = Field(description="职业")
# 3. 使用 with_structured_output
structured_llm = model.with_structured_output(Person)
result = structured_llm.invoke("张三是一名 30 岁的软件工程师")
print(result)
print(type(result))
# name='张三' age=30 occupation='软件工程师'
# <class '__main__.Person'>
print(result.name) # "张三"
print(result.age) # 30
print(result.occupation) # "软件工程师"⚠️ 没有描述,LLM 可能格式错误。 所以一定要写
description。
情感分析示例:
class SentimentAnalysis(BaseModel):
"""情感分析结果"""
sentiment: str = Field(description="情感倾向:positive/negative/neutral")
confidence: float = Field(description="置信度,0-1之间")
keywords: list[str] = Field(description="关键词列表")
structured_model = model.with_structured_output(SentimentAnalysis)
text = "这个课程内容很实用,学到了很多知识,强烈推荐!"
result = structured_model.invoke(f"分析以下文本的情感:\n{text}")
print(result.sentiment) # positive
print(result.confidence) # 0.99
print(result.keywords) # ['实用', '学到了很多知识', '强烈推荐']4.2 高级特性
① 可选字段(Optional)
LLM 未填充某些字段怎么办?用 Optional 指定字段可选。
from typing import Optional
class Person(BaseModel):
"""人物信息"""
name: str = Field(description="姓名")
age: Optional[int] = Field(description="年龄") # 可选
occupation: str = Field(description="职业")不用
Optional:Person(name='张三', age=0, occupation='医生')(会被填 0)用
Optional:Person(name='张三', age=None, occupation='医生')(填 None)
② 默认值
格式:Field(default="默认值", description="描述")
class Product(BaseModel):
"""产品信息"""
name: str = Field(description="产品名称")
price: float = Field(description="价格")
description: Optional[str] = Field(description="产品描述")
stock: int = Field(default=100, description="库存") # 默认 100⚠️ 注意:不同模型提供商对
default字段的支持是不同的,测试时要以实际平台为准。
③ 枚举类型(Enum / Literal)
用枚举限制字段可选值。
from enum import Enum
from typing import Optional
from pydantic import BaseModel, Field
class Priority(str, Enum):
LOW = "低"
MEDIUM = "中"
HIGH = "高"
class CustomerInfo(BaseModel):
"""客户信息"""
name: str = Field(description="客户姓名")
phone: str = Field(description="电话号码")
email: Optional[str] = Field(description="邮箱")
issue: str = Field(description="问题描述")
urgency: Priority = Field(description="紧急程度") # 只能是 低/中/高或者用 Literal 直接写死(更简单):
from typing import Literal
class CustomerInfo(BaseModel):
"""客户信息"""
name: str = Field(description="客户姓名")
urgency: Literal["低", "中", "高"] = Field(description="紧急程度")应用场景:自动填充 CRM、工单自动分类、客服辅助。
④ 列表提取(List)
from typing import List
class Person(BaseModel):
name: str
age: int
class PersonList(BaseModel):
"""人物列表信息"""
people: List[Person] # 多个 Person 对象
structured_llm = model.with_structured_output(PersonList)
result = structured_llm.invoke("张三 30岁,李四 25岁")
# people=[Person(name='张三', age=30), Person(name='李四', age=25)]应用场景:批量处理用户评论、自动生成分析报告、发现产品改进点、自动化财务处理、OCR 后结构化。
⑤ 嵌套结构
from pydantic import BaseModel, Field
from typing import List
class Actor(BaseModel):
"""演员信息"""
name: str = Field(description="演员姓名")
role: str = Field(description="饰演的角色")
class Movie(BaseModel):
"""电影信息"""
title: str = Field(description="电影标题")
year: int = Field(description="上映年份")
director: str = Field(description="导演")
cast: List[Actor] = Field(description="演员列表") # 嵌套列表
rating: float = Field(description="评分")
structured_model = model.with_structured_output(Movie)
response = structured_model.invoke("请介绍电影《盗梦空间》")
print(response.cast) # [Actor(name='莱昂纳多·迪卡普里奥', role='柯布'), ...]⚠️ LLM 能力有限,复杂嵌套结构可能会出错。 建议:
嵌套层级 ≤ 3 层(4 层以上容易出错)
使用清晰的 description
必要时拆分成多个调用
⑥ 限制条件(参数校验)
用 Field 的约束参数做字段级校验:
from pydantic import ValidationError
class User(BaseModel):
name: str = Field(min_length=2, max_length=20)
age: int = Field(ge=0, le=150) # 0~150
email: str
try:
user = User(name="李四", age=200, email="li@example.com")
except ValidationError as e:
print(e.errors()[0]['msg'])
# Input should be less than or equal to 150常用约束:min_length、max_length、ge(≥)、le(≤)、gt(>)。
五、模式二:TypedDict(轻量)
TypedDict 是 Python 3.8+ 引入的类型提示工具,即"带类型声明的字典结构"。适合快速定义字典结构、无需 Pydantic 重量级功能的场景。
普通 dict 没有类型信息,TypedDict 可以声明字段和类型。但它主要是类型声明,不是运行时强校验器。
from typing_extensions import TypedDict
class MovieDict(TypedDict):
title: str
year: int
director: str
rating: float
movie: MovieDict = {
"title1": "盗梦空间", # 字段名不一致,IDE 会标记,但运行不报错
"year": 2010,
"director": "克里斯托弗·诺兰",
"rating": 8.8,
}注意:字段名写错(如
title1)IDE 静态检查会提示,但不会导致运行时异常。
用 Annotated 附加描述
Annotated 在"类型"之外附加元数据(类似 Pydantic 的 Field)。
from typing_extensions import TypedDict, Annotated
from typing import List
class Actor(TypedDict):
"""演员情况"""
name: Annotated[str, "演员姓名"]
role: Annotated[str, "饰演的角色"]
class Movie(TypedDict):
"""电影情况"""
title: Annotated[str, "电影标题"]
year: Annotated[int, "上映年份"]
director: Annotated[str, "导演"]
cast: Annotated[List[Actor], "演员列表"] # 嵌套列表
rating: Annotated[float, "评分"]
structured_llm = model_with_closeai.with_structured_output(Movie)
resp = structured_llm.invoke("给我介绍下电影《盗梦空间》")
print(resp['title']) # 盗梦空间(返回的是字典,用 [] 访问)六、模式三:JSON Schema(跨语言通用)
JSON Schema 与前/后端、跨语言接口最通用,直接传符合 JSON Schema 标准的字典或字符串。
json_schema = {
"title": "Movie",
"description": "A movie with details",
"type": "object",
"properties": {
"title": {"type": "string", "description": "The title of the movie"},
"year": {"type": "integer", "description": "The year the movie was released"},
"director": {"type": "string", "description": "The director of the movie"},
"rating": {"type": "number", "description": "The movie's rating out of 10"}
},
"required": ["title", "year", "director", "rating"]
}
structured_model = model.with_structured_output(
json_schema,
method="json_schema"
)返回的是字典,不校验字段匹配。
七、模式四:dataclass
用标准库的 @dataclass 装饰器定义数据结构,让"数据结构定义"更简洁、清晰。
from dataclasses import dataclass
from pydantic import Field
@dataclass
class Movie:
"""
电影的详细信息
"""
title: str = Field(description="电影标题")
year: int = Field(description="电影上映年份")
director: str = Field(description="导演")
rating: float = Field(description="电影评分,满分十分")
structured_model = model.with_structured_output(Movie)
response = structured_model.invoke("给出盗梦空间的信息")
print(response) # {'title': '盗梦空间', 'year': 2010, ...}
print(type(response)) # <class 'dict'>⚠️ 注意:
@dataclass修饰的类是数据类,被标准库标记并携带字段元信息,可以作为 LangChain 的 Schema;而手写__init__等方法的普通类不能替代。另外它返回的是未经校验的字典。
八、四种模式的关键区别:类型校验
通过一个"fake server"(模拟 DeepSeek 服务端,故意返回字段不匹配的数据)可以直观地看到差异。
假设模型返回了字段名不匹配的数据(title1、year2),各自的处理如下:
Pydantic 的报错示例:
ValidationError: 2 validation errors for MovieModel
title
Field required [type=missing, ...]
year
Field required [type=missing, ...]结论:用 Pydantic 定义 schema,收到响应后会强制校验,字段不匹配则抛异常;其余三种方式不校验。
这也是为什么生产环境更推荐 Pydantic —— 它能第一时间帮你发现"模型吐错了字段"的问题。
九、获取结构化结果的方式
除了四种 Schema 模式,获取结构化结果本身也有两种方式。
方式 1:with_structured_output(推荐)
最新、最简洁的 API,直接让模型"理解"数据结构并返回解析好的对象。
还可以传 include_raw=True 参数,返回解析前的原始 AIMessage,从而访问令牌用量等元数据:
from pydantic import BaseModel, Field
from rich import print as rprint
class Movie(BaseModel):
"""电影信息"""
title: str = Field(description="电影标题")
year: int = Field(description="上映年份")
director: str = Field(description="导演")
rating: float = Field(description="评分(10分制)")
model_with_structure = model.with_structured_output(Movie, include_raw=True)
resp = model_with_structure.invoke("给我介绍下电影《星际穿越》")
print(type(resp)) # <class 'dict'>返回结果是一个字典,包含三个字段:
raw:返回的原始 AIMessage(含 token 用量等元数据)。parsed:解析后的输出(比如Movie(...)实例)。parsing_error:解析错误。当前用 Pydantic,格式不符合 schema 会报错;其余三种方式不会。
方式 2:输出解析器(传统,不推荐)
更传统的方法:在提示词里明确指示模型输出特定格式文本,再用解析器转换。
流程:提示词指导 → 模型生成文本 → 解析器转换。
from langchain_core.output_parsers import JsonOutputParser
from langchain_core.prompts import ChatPromptTemplate
from pydantic import BaseModel, Field
# 1. 提示词模板
prompt_template = ChatPromptTemplate.from_messages([
("system", "回答用户问题,必须始终输出一个包含title(电影标题)和year(上映年份)的 JSON 对象"),
("human", "问题:{question}")
])
# 2. 定义结构
class Movie(BaseModel):
"""电影信息"""
title: str = Field(description="电影标题")
year: int = Field(description="上映年份")
# 3. 创建解析器
parser = JsonOutputParser(pydantic_object=Movie)
# 4. 创建链(用 | 管道拼接)
chain = prompt_template | model | parser
# 5. 调用
response = chain.invoke({"question": "介绍电影《盗梦空间》"})
print(response) # {'title': '盗梦空间', 'year': 2010}这种方式的缺点是:依赖 Prompt"求"模型输出 JSON,稳定性不如 with_structured_output。
十、工作原理图解
以 Pydantic + with_structured_output 为例,内部流程分为四步:
第 1 步:定义结构
from pydantic import BaseModel, Field
class BookInfo(BaseModel):
title: str = Field(description="书名")
author: str = Field(description="作者名字")
tags: list[str] = Field(description="书籍的标签或分类")第 2 步:协议转换
LangChain 内部调用 Pydantic 的底层方法(如 model_json_schema()),把你写的 Python 代码自动转成标准 JSON Schema——详细描述有哪些字段、字段类型(string、array 等)以及字段描述。
第 3 步:模型交互与强约束
LangChain 把 JSON Schema 包装进模型 API 请求中。
现代方法(
.with_structured_output):现代大模型普遍支持"函数/工具调用"或"JSON Mode",LangChain 把 JSON Schema 作为 Tools 传入。大模型侧约束:像 OpenAI 的
strict=True参数会启动语法采样约束(Grammar-based sampling),模型解码生成 token 时严格按 JSON Schema 的语法树选择,从底层保证输出格式不走样。
第 4 步:自动解析与验证
模型返回 JSON 字符串后,PydanticStructuredOutputParser(解析器)接管:
解析(Parsing):把字符串解析为 Python 字典。
验证(Validation):把字典喂给 Pydantic 模型,自动检查类型。缺必填字段或类型错误会抛验证错误(或触发 LangChain 重试)。
返回(Return):通过验证后,你拿到的是可直接点出属性的 Pydantic 对象(如
result.title)。
评论区