返回资源广场

Skills 资源 / 技能包

offloading-large-tool-output-to-files

用智能体的虚拟文件系统把大块工具输出移出上下文——既走「模型主动 write_file 存盘、需要时 read_file 取回」,也走「框架按阈值自动卸载并留一条引用」。Use when 工具一次返回几十 KB 文本、上下文被大输出撑爆、账单随历史线性上涨、或要配置自动卸载阈值、排查 write_file 报已存在与 edit_file 报字符串出现多次这类文件工具报错时。涵盖两条卸载路径、阈值量纲与配置入口、文件生命周期、写入与编辑契约;不含对话历史摘要压缩(见 configuring-conversation-compaction)。

SKILL.md 技能文档

能力目标

让一个会产出大块文本的工具不再把上下文撑爆:产物落进虚拟文件、上下文里只留一条带预览的引用,需要时再读回;并且能自己调阈值让自动卸载在中等体量输出上就现形、量出省下了多少、说清代价在哪。

前置

  • 已能用 create_deep_agent 构造智能体(见 bootstrapping-deepagents-env)。
  • 文件工具 ls、read_file、write_file、edit_file、glob、grep 由框架自动注入,不需要你写任何文件读写代码。
  • 一个认知前提:这些「文件」不在磁盘上,而是状态里 files 通道持有的一个扁平字典,键是绝对路径字符串、值含 content、encoding、created_at、modified_at 四个字段。ls 看到的「目录」是按路径前缀做字符串匹配合成出来的视图,字典里并没有目录条目。去磁盘上找文件会一无所获。

实操流程

  1. 准备一个会产出大块文本的工具。要让机制可见,输出下限取 2048 字节:

    def big_output(topic: str) -> str:
        """Generate a detailed multi-section report for the given topic."""
        sections = [
            f"## Section {i}: {topic} aspect {i}\n" + ("lorem ipsum detail line. " * 40)
            for i in range(1, 13)
        ]
        return "\n\n".join(sections)
    

    迁移到自己的业务时只改这个函数体(生成报告、汇总搜索结果、抓网页全文都行),框架层一行不动。

  2. 走第一条路径:用系统提示引导模型自己存盘再读回。不调用任何文件 API,只用自然语言告诉它怎么做:

    agent = create_deep_agent(
        model=model,
        tools=[big_output],
        system_prompt=(
            "You are a helpful assistant. When given a topic, use the big_output "
            "tool to generate a report, then SAVE the full report to a file using "
            "write_file, and finally read it back with read_file and summarize the "
            "key points."
        ),
    )
    result = agent.invoke({"messages": [HumanMessage(content=(
        "Generate a comprehensive report on 'Python async programming' using the "
        "big_output tool, save it to /reports/python_async.md, then read it back "
        "and summarize the top 3 key points."
    ))]})
    print(list(result.get("files", {}).keys()))
    
  3. 走第二条路径:让框架按阈值自动卸载。要自定义阈值就必须降到低层构造入口自己挂中间件——高层入口已内置了一个文件系统中间件,再传一个会因重复实例而断言失败:

    from langchain.agents import create_agent
    from deepagents.middleware.filesystem import FilesystemMiddleware
    
    agent = create_agent(
        model,
        [big_output],
        middleware=[FilesystemMiddleware(tool_token_limit_before_evict=50)],
    )
    result = agent.invoke({"messages": [HumanMessage(content=(
        "Generate a report on 'artificial intelligence' using big_output."
    ))]})
    
  4. 算清阈值量纲再定值。触发条件是 len(content) > NUM_CHARS_PER_TOKEN × tool_token_limit_before_evict,其中每 token 按 4 字符估算、该上限默认 20000 token,即约 80000 字符(约 78 KB)。一个普通工具很难一次返回这么多,所以默认配置下你几乎看不到自动卸载——把这个值调小,中等输出就能触发:

    from deepagents.middleware.filesystem import NUM_CHARS_PER_TOKEN
    print(NUM_CHARS_PER_TOKEN * 20000)   # 默认触发线,单位是字符
    
  5. 清点卸载留下的两个产物:

    from langchain_core.messages import ToolMessage
    
    for msg in result["messages"]:
        if isinstance(msg, ToolMessage) and "too large" in str(msg.content).lower():
            print(msg.content)          # 产物一:被替换成引用的工具消息
    for path, fd in result["files"].items():
        if "large_tool_results" in path:
            print(path, len(fd["content"]))   # 产物二:持有全文的虚拟文件
    
  6. 需要文件跨多次调用存活时,挂一个检查点保存器并复用同一个会话标识:

    from langgraph.checkpoint.memory import MemorySaver
    
    agent = create_deep_agent(model=model, tools=[note_tool], system_prompt=..., checkpointer=MemorySaver())
    cfg = {"configurable": {"thread_id": "workspace-1"}}
    agent.invoke({"messages": [...]}, config=cfg)   # 第一次写
    agent.invoke({"messages": [...]}, config=cfg)   # 第二次读得到
    

校验回路

  • 第一条路径:result["files"] 里出现你指定的那个路径键;消息链依次是提问、工具调用与返回、write_file、read_file、最终总结。
  • 第二条路径:result["files"] 里出现 /large_tool_results/<工具调用 ID> 这个键、且它持有完整原文;同时上下文里那条工具消息被替换成 Tool result too large, the result of this tool call ... was saved in the filesystem at this path: /large_tool_results/...,后面跟着前 5 行加后 5 行的预览和一句「用 read_file 分段取回」。
  • 量收益:并排跑一次不挂文件系统中间件的对照,比较两边工具消息的总字节数。实测同一份约 12.4 KB 的输出,卸载后上下文里的工具消息降到约 4.9 KB、省约 60%,而全文仍完整存在状态里。
  • 不挂检查点保存器时,第二次调用的 result["files"] 为空字典、ls 返回空列表——这是预期行为,不是配错了。

常见陷阱

  • 把文件系统中间件塞进高层入口想改阈值:会抛 AssertionError: Please remove duplicate middleware instances.。这不是语法写错,是高层入口已内置一个同类中间件。要自定义参数就改用低层构造入口自己挂。
  • 以为几 KB 输出就会被自动卸载:默认触发线约 80000 字符,3 KB 量级的输出连边都够不着,result["files"] 会是空的。想在小输出上看见机制,就把 tool_token_limit_before_evict 调小。
  • 担心 read_file 取回全文又超阈值、陷入死循环:不会。六个文件工具全在卸载排除清单里、永不被自动卸载。卸载闸门的顺序是先查排除清单命中则直接放行、再比内容长度与阈值。
  • 同一路径第二次 write_file:报 Cannot write to <path> because it already exists. Read and then make an edit, or write to a new path.——写入工具是只创建语义,改已有文件必须走 edit_file。这条错误会作为工具消息回流给模型,模型通常会自己改用先读后编辑。
  • edit_file 的待替换串在文件里出现多次:报 String '...' appears N times in file. Use replace_all=True to replace all instances, or provide a more specific string.。要么加 replace_all=True,要么把待替换串写得更长更唯一。
  • 把卸载当纯收益:卸载后上下文里只剩预览,智能体要拿全文必须多发一次 read_file,多一轮往返延迟;更要紧的是它若判断失误不去读回,就会在残缺预览上作答。省 token 换的是这笔开销加一份判断失误的风险。
  • 把卸载与摘要压缩、子智能体隔离混为一谈:卸载只替换单条工具消息的内容、其余消息原封不动,且是无损的(全文仍在状态里、随时读回)。压缩对话历史是另一套有损机制,子智能体隔离解决的是中间过程不污染主上下文,三者不能互相替代。

适用范围与前置条件

  • 已能用 create_deep_agent 构造智能体(见 bootstrapping-deepagents-env)。
  • 文件工具 ls、read_file、write_file、edit_file、glob、grep 由框架自动注入,不需要你写任何文件读写代码。
  • 一个认知前提:这些「文件」不在磁盘上,而是状态里 files 通道持有的一个扁平字典,键是绝对路径字符串、值含 content、encoding、created_at、modified_at 四个字段。ls 看到的「目录」是按路径前缀做字符串匹配合成出来的视图,字典里并没有目录条目。去磁盘上找文件会一无所获。

怎么使用

使用步骤

  1. 准备一个会产出大块文本的工具。要让机制可见,输出下限取 2048 字节:

    def big_output(topic: str) -> str:
        """Generate a detailed multi-section report for the given topic."""
        sections = [
            f"## Section {i}: {topic} aspect {i}\n" + ("lorem ipsum detail line. " * 40)
            for i in range(1, 13)
        ]
        return "\n\n".join(sections)
    

    迁移到自己的业务时只改这个函数体(生成报告、汇总搜索结果、抓网页全文都行),框架层一行不动。

  2. 走第一条路径:用系统提示引导模型自己存盘再读回。不调用任何文件 API,只用自然语言告诉它怎么做:

    agent = create_deep_agent(
        model=model,
        tools=[big_output],
        system_prompt=(
            "You are a helpful assistant. When given a topic, use the big_output "
            "tool to generate a report, then SAVE the full report to a file using "
            "write_file, and finally read it back with read_file and summarize the "
            "key points."
        ),
    )
    result = agent.invoke({"messages": [HumanMessage(content=(
        "Generate a comprehensive report on 'Python async programming' using the "
        "big_output tool, save it to /reports/python_async.md, then read it back "
        "and summarize the top 3 key points."
    ))]})
    print(list(result.get("files", {}).keys()))
    
  3. 走第二条路径:让框架按阈值自动卸载。要自定义阈值就必须降到低层构造入口自己挂中间件——高层入口已内置了一个文件系统中间件,再传一个会因重复实例而断言失败:

    from langchain.agents import create_agent
    from deepagents.middleware.filesystem import FilesystemMiddleware
    
    agent = create_agent(
        model,
        [big_output],
        middleware=[FilesystemMiddleware(tool_token_limit_before_evict=50)],
    )
    result = agent.invoke({"messages": [HumanMessage(content=(
        "Generate a report on 'artificial intelligence' using big_output."
    ))]})
    
  4. 算清阈值量纲再定值。触发条件是 len(content) > NUM_CHARS_PER_TOKEN × tool_token_limit_before_evict,其中每 token 按 4 字符估算、该上限默认 20000 token,即约 80000 字符(约 78 KB)。一个普通工具很难一次返回这么多,所以默认配置下你几乎看不到自动卸载——把这个值调小,中等输出就能触发:

    from deepagents.middleware.filesystem import NUM_CHARS_PER_TOKEN
    print(NUM_CHARS_PER_TOKEN * 20000)   # 默认触发线,单位是字符
    
  5. 清点卸载留下的两个产物:

    from langchain_core.messages import ToolMessage
    
    for msg in result["messages"]:
        if isinstance(msg, ToolMessage) and "too large" in str(msg.content).lower():
            print(msg.content)          # 产物一:被替换成引用的工具消息
    for path, fd in result["files"].items():
        if "large_tool_results" in path:
            print(path, len(fd["content"]))   # 产物二:持有全文的虚拟文件
    
  6. 需要文件跨多次调用存活时,挂一个检查点保存器并复用同一个会话标识:

    from langgraph.checkpoint.memory import MemorySaver
    
    agent = create_deep_agent(model=model, tools=[note_tool], system_prompt=..., checkpointer=MemorySaver())
    cfg = {"configurable": {"thread_id": "workspace-1"}}
    agent.invoke({"messages": [...]}, config=cfg)   # 第一次写
    agent.invoke({"messages": [...]}, config=cfg)   # 第二次读得到
    

继续探索

全部资源
Skills 资源 / 技能包

bootstrapping-deepagents-env

在一台干净机器上装好 DeepAgents 运行环境、接入一个 OpenAI 兼容大模型凭证,并跑通第一个 create_deep_agent 工具调用闭环。Use when 需要初始化 DeepAgents 开发环境、系统 Python 版本不达标装不上包、不确定装到了哪个版本、接 DeepSeek 之类国产模型报 ImportError 或 404 这类环境层故障时。涵盖解释器版本核对、虚拟环境置备、主包与提供方包安装、版本核验、凭证注入、最小示例验收;不含 Agent 各项能力的用法(见 tracking-task-progress-with-todos 等能力型 skill)。

Skills 资源 / 技能包

inspecting-agent-graph-and-tools

把一个 create_deep_agent 建出来的智能体拆开看:列出执行图节点、列出实际挂载的工具、捕获框架预装的中间件清单、抓取每轮真正发给模型的工具集。Use when 需要确认某项能力是否真的挂上了、排查「我的工具去哪了 / 这些工具哪来的 / 内置工具到底几个」、验证自定义中间件是否进了图、或要在改配置前后做结构对照时。涵盖图节点自省、工具清单反查、中间件清单捕获、编译期与运行期工具集差异;不含具体能力的用法。

Skills 资源 / 技能包

tracking-task-progress-with-todos

让智能体把多步任务拆成结构化待办清单写进状态,并从调用结果里取出清单、渲染成实时进度、兜底检测「勾完清单却没给答案」的失败形态。Use when 需要给长任务做进度面板、想稳定触发 write_todos、发现规划没被触发、或要把 todos 推给前端 UI 与日志时。涵盖稳定触发写法、取清单的两条路径、三态进度渲染、失败形态检测;不含子任务委派(见 delegating-subtasks-to-subagents)。