SKILL.md 技能文档
能力目标
让一个会产出大块文本的工具不再把上下文撑爆:产物落进虚拟文件、上下文里只留一条带预览的引用,需要时再读回;并且能自己调阈值让自动卸载在中等体量输出上就现形、量出省下了多少、说清代价在哪。
前置
- 已能用
create_deep_agent构造智能体(见 bootstrapping-deepagents-env)。 - 文件工具
ls、read_file、write_file、edit_file、glob、grep由框架自动注入,不需要你写任何文件读写代码。 - 一个认知前提:这些「文件」不在磁盘上,而是状态里
files通道持有的一个扁平字典,键是绝对路径字符串、值含content、encoding、created_at、modified_at四个字段。ls看到的「目录」是按路径前缀做字符串匹配合成出来的视图,字典里并没有目录条目。去磁盘上找文件会一无所获。
实操流程
准备一个会产出大块文本的工具。要让机制可见,输出下限取 2048 字节:
def big_output(topic: str) -> str: """Generate a detailed multi-section report for the given topic.""" sections = [ f"## Section {i}: {topic} aspect {i}\n" + ("lorem ipsum detail line. " * 40) for i in range(1, 13) ] return "\n\n".join(sections)迁移到自己的业务时只改这个函数体(生成报告、汇总搜索结果、抓网页全文都行),框架层一行不动。
走第一条路径:用系统提示引导模型自己存盘再读回。不调用任何文件 API,只用自然语言告诉它怎么做:
agent = create_deep_agent( model=model, tools=[big_output], system_prompt=( "You are a helpful assistant. When given a topic, use the big_output " "tool to generate a report, then SAVE the full report to a file using " "write_file, and finally read it back with read_file and summarize the " "key points." ), ) result = agent.invoke({"messages": [HumanMessage(content=( "Generate a comprehensive report on 'Python async programming' using the " "big_output tool, save it to /reports/python_async.md, then read it back " "and summarize the top 3 key points." ))]}) print(list(result.get("files", {}).keys()))走第二条路径:让框架按阈值自动卸载。要自定义阈值就必须降到低层构造入口自己挂中间件——高层入口已内置了一个文件系统中间件,再传一个会因重复实例而断言失败:
from langchain.agents import create_agent from deepagents.middleware.filesystem import FilesystemMiddleware agent = create_agent( model, [big_output], middleware=[FilesystemMiddleware(tool_token_limit_before_evict=50)], ) result = agent.invoke({"messages": [HumanMessage(content=( "Generate a report on 'artificial intelligence' using big_output." ))]})算清阈值量纲再定值。触发条件是
len(content) > NUM_CHARS_PER_TOKEN × tool_token_limit_before_evict,其中每 token 按 4 字符估算、该上限默认 20000 token,即约 80000 字符(约 78 KB)。一个普通工具很难一次返回这么多,所以默认配置下你几乎看不到自动卸载——把这个值调小,中等输出就能触发:from deepagents.middleware.filesystem import NUM_CHARS_PER_TOKEN print(NUM_CHARS_PER_TOKEN * 20000) # 默认触发线,单位是字符清点卸载留下的两个产物:
from langchain_core.messages import ToolMessage for msg in result["messages"]: if isinstance(msg, ToolMessage) and "too large" in str(msg.content).lower(): print(msg.content) # 产物一:被替换成引用的工具消息 for path, fd in result["files"].items(): if "large_tool_results" in path: print(path, len(fd["content"])) # 产物二:持有全文的虚拟文件需要文件跨多次调用存活时,挂一个检查点保存器并复用同一个会话标识:
from langgraph.checkpoint.memory import MemorySaver agent = create_deep_agent(model=model, tools=[note_tool], system_prompt=..., checkpointer=MemorySaver()) cfg = {"configurable": {"thread_id": "workspace-1"}} agent.invoke({"messages": [...]}, config=cfg) # 第一次写 agent.invoke({"messages": [...]}, config=cfg) # 第二次读得到
校验回路
- 第一条路径:
result["files"]里出现你指定的那个路径键;消息链依次是提问、工具调用与返回、write_file、read_file、最终总结。 - 第二条路径:
result["files"]里出现/large_tool_results/<工具调用 ID>这个键、且它持有完整原文;同时上下文里那条工具消息被替换成Tool result too large, the result of this tool call ... was saved in the filesystem at this path: /large_tool_results/...,后面跟着前 5 行加后 5 行的预览和一句「用 read_file 分段取回」。 - 量收益:并排跑一次不挂文件系统中间件的对照,比较两边工具消息的总字节数。实测同一份约 12.4 KB 的输出,卸载后上下文里的工具消息降到约 4.9 KB、省约 60%,而全文仍完整存在状态里。
- 不挂检查点保存器时,第二次调用的
result["files"]为空字典、ls返回空列表——这是预期行为,不是配错了。
常见陷阱
- 把文件系统中间件塞进高层入口想改阈值:会抛
AssertionError: Please remove duplicate middleware instances.。这不是语法写错,是高层入口已内置一个同类中间件。要自定义参数就改用低层构造入口自己挂。 - 以为几 KB 输出就会被自动卸载:默认触发线约 80000 字符,3 KB 量级的输出连边都够不着,
result["files"]会是空的。想在小输出上看见机制,就把tool_token_limit_before_evict调小。 - 担心
read_file取回全文又超阈值、陷入死循环:不会。六个文件工具全在卸载排除清单里、永不被自动卸载。卸载闸门的顺序是先查排除清单命中则直接放行、再比内容长度与阈值。 - 同一路径第二次
write_file:报Cannot write to <path> because it already exists. Read and then make an edit, or write to a new path.——写入工具是只创建语义,改已有文件必须走edit_file。这条错误会作为工具消息回流给模型,模型通常会自己改用先读后编辑。 edit_file的待替换串在文件里出现多次:报String '...' appears N times in file. Use replace_all=True to replace all instances, or provide a more specific string.。要么加replace_all=True,要么把待替换串写得更长更唯一。- 把卸载当纯收益:卸载后上下文里只剩预览,智能体要拿全文必须多发一次
read_file,多一轮往返延迟;更要紧的是它若判断失误不去读回,就会在残缺预览上作答。省 token 换的是这笔开销加一份判断失误的风险。 - 把卸载与摘要压缩、子智能体隔离混为一谈:卸载只替换单条工具消息的内容、其余消息原封不动,且是无损的(全文仍在状态里、随时读回)。压缩对话历史是另一套有损机制,子智能体隔离解决的是中间过程不污染主上下文,三者不能互相替代。
适用范围与前置条件
- 已能用
create_deep_agent构造智能体(见 bootstrapping-deepagents-env)。 - 文件工具
ls、read_file、write_file、edit_file、glob、grep由框架自动注入,不需要你写任何文件读写代码。 - 一个认知前提:这些「文件」不在磁盘上,而是状态里
files通道持有的一个扁平字典,键是绝对路径字符串、值含content、encoding、created_at、modified_at四个字段。ls看到的「目录」是按路径前缀做字符串匹配合成出来的视图,字典里并没有目录条目。去磁盘上找文件会一无所获。
怎么使用
使用步骤
准备一个会产出大块文本的工具。要让机制可见,输出下限取 2048 字节:
def big_output(topic: str) -> str: """Generate a detailed multi-section report for the given topic.""" sections = [ f"## Section {i}: {topic} aspect {i}\n" + ("lorem ipsum detail line. " * 40) for i in range(1, 13) ] return "\n\n".join(sections)迁移到自己的业务时只改这个函数体(生成报告、汇总搜索结果、抓网页全文都行),框架层一行不动。
走第一条路径:用系统提示引导模型自己存盘再读回。不调用任何文件 API,只用自然语言告诉它怎么做:
agent = create_deep_agent( model=model, tools=[big_output], system_prompt=( "You are a helpful assistant. When given a topic, use the big_output " "tool to generate a report, then SAVE the full report to a file using " "write_file, and finally read it back with read_file and summarize the " "key points." ), ) result = agent.invoke({"messages": [HumanMessage(content=( "Generate a comprehensive report on 'Python async programming' using the " "big_output tool, save it to /reports/python_async.md, then read it back " "and summarize the top 3 key points." ))]}) print(list(result.get("files", {}).keys()))走第二条路径:让框架按阈值自动卸载。要自定义阈值就必须降到低层构造入口自己挂中间件——高层入口已内置了一个文件系统中间件,再传一个会因重复实例而断言失败:
from langchain.agents import create_agent from deepagents.middleware.filesystem import FilesystemMiddleware agent = create_agent( model, [big_output], middleware=[FilesystemMiddleware(tool_token_limit_before_evict=50)], ) result = agent.invoke({"messages": [HumanMessage(content=( "Generate a report on 'artificial intelligence' using big_output." ))]})算清阈值量纲再定值。触发条件是
len(content) > NUM_CHARS_PER_TOKEN × tool_token_limit_before_evict,其中每 token 按 4 字符估算、该上限默认 20000 token,即约 80000 字符(约 78 KB)。一个普通工具很难一次返回这么多,所以默认配置下你几乎看不到自动卸载——把这个值调小,中等输出就能触发:from deepagents.middleware.filesystem import NUM_CHARS_PER_TOKEN print(NUM_CHARS_PER_TOKEN * 20000) # 默认触发线,单位是字符清点卸载留下的两个产物:
from langchain_core.messages import ToolMessage for msg in result["messages"]: if isinstance(msg, ToolMessage) and "too large" in str(msg.content).lower(): print(msg.content) # 产物一:被替换成引用的工具消息 for path, fd in result["files"].items(): if "large_tool_results" in path: print(path, len(fd["content"])) # 产物二:持有全文的虚拟文件需要文件跨多次调用存活时,挂一个检查点保存器并复用同一个会话标识:
from langgraph.checkpoint.memory import MemorySaver agent = create_deep_agent(model=model, tools=[note_tool], system_prompt=..., checkpointer=MemorySaver()) cfg = {"configurable": {"thread_id": "workspace-1"}} agent.invoke({"messages": [...]}, config=cfg) # 第一次写 agent.invoke({"messages": [...]}, config=cfg) # 第二次读得到