ARTICLE · INTELLIGENCE

战地情报 · 详情页

来自尧图项目组的一线实战观察与深度解析

AIO Sandbox 实战集成指南:终端、浏览器自动化与 AI Agent 全场景示例

AIO Sandbox 实战集成指南:终端、浏览器自动化与 AI Agent 全场景示例 AI Agent后端MCP 服务浏览器控制Agent 评测【免费下载链接】sandboxAll-in-One Sandbox for AI Agents that combines Browser, Shell, File, MCP and VSCode Server in a single Docker container.项目地址https://gitcode.com/gh_mirrors/sandbox103/sandbox点击查看免费下载本指南以 AIO SandboxAll-in-One Sandbox for AI Agents官方示例文档为骨架系统讲解如何将 WebSocket 终端、浏览器自动化VNC/CDP/MCP、文件与代码执行能力以及 AI Agent 工作流集成到你的应用中。读完本文你将掌握 Docker Compose/Kubernetes 两种部署方案、Python 与 Node.js 两套 SDK 的完整调用方式以及从简单命令执行到多工具 Agent 编排的实战落地路径。快速示例三大集成路径总览AIO Sandbox 将 Browser、Shell、File、MCP 与 VSCode Server 封装在同一个 Docker 容器中官方示例文档围绕三类最常用的集成场景组织终端集成把沙盒 Shell 变成你应用的一部分通过 WebSocket 协议将沙盒内的终端接入任意前端或后端应用适合远程调试、CI 流水线、在线 IDE 等场景。基本终端客户端Node.js 与 Python Asyncio 两种实现建立到ws://localhost:8080/v1/shell/ws的连接并收发input/output消息。高级终端功能会话管理与断线重连支持按session_id恢复既有会话、30 秒心跳保活与指数退避重连。浏览器自动化可视化与程序化控制并行浏览器模块提供 VNC 可视化交互、Chrome DevTools 协议CDP底层控制与浏览器 MCP 服务器三种能力覆盖 Web 抓取、表单测试与 UI 回归等需求。Browser Use 集成Python 浏览器自动化入口。Playwright 集成通过connect_over_cdp复用沙盒内浏览器。Web 抓取示例价格监控、社交媒体内容提取等数据提取模式。Agent 集成把沙盒能力交给 LLM内置 MCP 服务器让 Agent 可以无缝调用沙盒的 Shell、文件、浏览器与代码执行工具是构建 AI 工作流的推荐路径。基本 Agent 设置将 Agent 连接到沙盒的 REST/MCP 接口。MCP 集成使用模型上下文协议发现与调用工具。多工具工作流组合浏览器、文件与代码执行 API 完成复杂任务。集成模式从 Docker Compose 到 Kubernetes 的部署方案Docker Compose 设置单机开发与测试场景下官方示例给出最小可用的 Compose 配置version: 3.8 services: aio-sandbox: image: ghcr.io/agent-infra/sandbox:latest ports: - 127.0.0.1:8080:8080 volumes: - sandbox_data:/workspace restart: unless-stopped volumes: sandbox_data:这里将宿主机127.0.0.1:8080映射到容器8080并把命名卷sandbox_data挂载到/workspace实现沙盒工作区数据持久化。仓库根目录的 docker-compose.yaml 提供了更贴近生产的内置端口布局可作为参数参考环境变量端口用途VNC_SERVER_PORT5900VNC 远程桌面服务WEBSOCKET_PROXY_PORT6080WebSocket 终端代理MCP_HUB_PORT8079MCP 服务器中心SANDBOX_SRV_PORT8091沙盒主服务JUPYTER_LAB_PORT8888Jupyter LabCODE_SERVER_PORT8200VSCode ServerBROWSER_REMOTE_DEBUGGING_PORT9222Chrome 远程调试CDP该文件还暴露了WORKSPACE默认/home/gem、DISPLAY_WIDTH/DISPLAY_HEIGHT、TZ等环境变量并支持通过PIP_INDEX_URL、UV_DEFAULT_INDEX、NPM_CONFIG_REGISTRY切换软件源镜像。Kubernetes 部署多副本生产场景使用 Deployment Service 组合官方示例给出如下配置apiVersion: apps/v1 kind: Deployment metadata: name: aio-sandbox spec: replicas: 2 selector: matchLabels: app: aio-sandbox template: metadata: labels: app: aio-sandbox spec: containers: - name: sandbox image: ghcr.io/agent-infra/sandbox:latest ports: - containerPort: 8080 resources: requests: memory: 1Gi cpu: 500m limits: memory: 2Gi cpu: 1000m --- apiVersion: v1 kind: Service metadata: name: aio-sandbox-service spec: selector: app: aio-sandbox ports: - port: 80 targetPort: 8080 type: ClusterIP在 Agent 生产部署方案中官方还建议在 Deployment 中同时加入ai-agent容器与aio-sandbox同 Pod 部署并对沙盒容器配置livenessProbe探测/v1/sandbox以保证健康检查——完整示例见 Agent 集成文档。Python SDK 实战从安装到完整功能调用安装与基本配置官方示例文档使用pip install aio-sandbox安装 Python SDK以当前仓库为准发布在 PyPI 上的实际包名为agent-sandbox见 sdk/python/pyproject.toml安装命令对应为pip install agent-sandboxSDK 依赖httpx、pydantic与volcengine-python-sdk要求 Python 3.8。初始化客户端from aio_sandbox import AioClient import asyncio # 初始化客户端 client AioClient( base_urlhttp://localhost:8080, # AIO Sandbox URL timeout30.0, # 请求超时秒 retries3, # 重试次数 retry_delay1.0 # 重试之间的延迟 )从源码结构看当前仓库的 Python SDK 由 Fern 从 OpenAPI 规范自动生成顶层入口为Sandbox同步与AsyncSandbox异步两个类位于 sdk/python/agent_sandbox/client.py。它们以惰性属性的方式暴露shell、bash、file、jupyter、nodejs、mcp、browser、browser_page、browser_tabs、browser_cookies、browser_state、browser_network、browser_captcha、code、util、skills、proxy、display、auth等 19 个子客户端每个模块还提供with_raw_response原始响应访问。文档中的AioClient是简化命名实际使用时导入Sandbox/AsyncSandbox即可。Shell 操作命令执行与会话管理async def shell_example(): # 执行简单命令 result await client.shell.exec(commandls -la) if result.success: print(f输出{result.data.output}) print(f退出码{result.data.exit_code}) # 使用会话管理执行 session_id my-session-1 await client.shell.exec( commandcd /workspace pwd, session_idsession_id ) # 在同一会话中继续 result await client.shell.exec( commandls, session_idsession_id ) # 长时间运行任务的异步执行 await client.shell.exec( commandpython long_script.py, async_modeTrue, session_idsession_id ) # 查看会话输出 view_result await client.shell.view(session_idsession_id) print(view_result.data.output) # 运行示例 asyncio.run(shell_example())对照 sdk/python/agent_sandbox/shell/client.py 的实现exec_command方法支持一组值得在生产中使用的参数id目标 Shell 会话的唯一标识不传则自动创建对应文档中的session_idexec_dir命令执行的工作目录必须为绝对路径async_mode是否异步执行异步时立即返回运行状态timeout等待命令完成的超时秒数超时返回运行中状态no_change_timeout无新输出超时检测默认 120 秒触发后返回NO_CHANGE_TIMEOUT状态hard_timeout硬超时到达后强制终止命令并返回HARD_TIMEOUT状态与timeout仅影响 HTTP 响应时机有本质区别strict工作目录校验严格模式truncate输出超过 30000 字符时截断默认开启。此外该客户端还提供create_session、update_session、view、wait_for_process、write_to_process、kill_process、get_terminal_url、list_sessions、cleanup_session/cleanup_all_sessions、get_session_stats等完整会话生命周期接口。文件操作读写、搜索与查找async def file_example(): # 写入文件 await client.file.write( file/tmp/example.py, content import matplotlib.pyplot as plt import numpy as np x np.linspace(0, 10, 100) y np.sin(x) plt.plot(x, y) plt.savefig(/tmp/plot.png) print(图表已保存) .strip() ) # 读取文件内容 content await client.file.read(file/tmp/example.py) if content.success: print(f文件内容\n{content.data.content}) # 列出目录内容 files await client.file.list( path/tmp, recursiveTrue, include_sizeTrue ) for file_info in files.data.files: print(f{file_info.name}: {file_info.size} 字节) # 在文件中搜索 search_result await client.file.search( file/tmp/example.py, regexrimport \w ) if search_result.success: for match in search_result.data.matches: print(f第 {match.line} 行{match.content}) # 按模式查找文件 found_files await client.file.find( path/tmp, glob*.py ) asyncio.run(file_example())从 sdk/python/agent_sandbox/file/client.py 的实现看文件模块的能力远不止示例中的read/write/list/search/find还包括write_file支持utf-8文本与base64二进制两种encoding可设置append追加模式、leading_newline/trailing_newline与sudo权限replace_in_file字符串精确替换grep_files跨文件内容搜索支持include/excludeglob 过滤、case_insensitive、context_before/after、max_results、max_file_size、multiline与recursive等底层与 ripgrep 语义对齐glob_files增强型 glob 匹配支持include_hidden、files_only、include_metadata、sort_by/sort_descupload_file/download_file流式上传下载下载支持change_policyabort在源文件变更时中止str_replace_editor兼容 Anthropic 定义的编辑器协议支持view/create/str_replace/insert/undo_edit并能以page_range/sheet_name/slide_range直接读取 PDF、Excel、PPTX 内容watch_*系列文件系统事件监听创建 watcher、长轮询事件、阻塞等待特定路径事件。代码执行Jupyter 与 Node.jsasync def code_execution_example(): # 在 Jupyter 内核中执行 Python 代码 jupyter_result await client.jupyter.execute( code import pandas as pd import numpy as np # 创建示例数据 df pd.DataFrame({ x: np.random.randn(100), y: np.random.randn(100) }) print(fDataFrame 形状{df.shape}) print(df.head()) , timeout60, session_iddata-analysis-session ) if jupyter_result.success: print(Jupyter 输出) for output in jupyter_result.data.outputs: if output.output_type stream: print(output.text) elif output.output_type execute_result: print(output.data.get(text/plain, )) # 执行 Node.js 代码 nodejs_result await client.nodejs.execute( code const fs require(fs); const path require(path); // 如果存在则读取 package.json try { const packagePath path.join(process.cwd(), package.json); if (fs.existsSync(packagePath)) { const pkg JSON.parse(fs.readFileSync(packagePath, utf8)); console.log(项目${pkg.name || 未知}); console.log(版本${pkg.version || 未知}); } else { console.log(未找到 package.json); } } catch (error) { console.error(错误, error.message); } , timeout30 ) if nodejs_result.success: print(fNode.js 输出{nodejs_result.data.stdout}) asyncio.run(code_execution_example())jupyter子客户端的核心方法为execute_code见 sdk/python/agent_sandbox/jupyter/client.pynodejs客户端则围绕node_js_execute提供运行时信息查询、会话列表与删除等接口。Jupyter 与 Node.js 会话接口在容器内分别对应 8888 与 8200 端口体系与 docker-compose.yaml 中的端口布局一致。MCP 集成async def mcp_example(): # 列出可用的 MCP 服务器 servers await client.mcp.list_servers() print(可用的 MCP 服务器, servers.data) # 从特定服务器获取工具 browser_tools await client.mcp.list_tools(server_namebrowser) for tool in browser_tools.data.tools: print(f工具{tool.name}) print(f描述{tool.description}) # 执行工具 screenshot_result await client.mcp.execute_tool( server_namebrowser, tool_namescreenshot, arguments{ url: https://example.com, width: 1920, height: 1080 } ) if screenshot_result.success: # 保存截图数据 await client.file.write( file/tmp/screenshot.png, contentscreenshot_result.data.content[0].data, # Base64 图像数据 appendFalse ) asyncio.run(mcp_example())MCP 中心在容器内监听8079端口docker-compose.yaml 中的MCP_HUB_PORT浏览器 MCP 服务器监听8100。浏览器模块的工具集在 browser 示例文档 中有完整清单覆盖导航、元素交互、内容检索与视觉模式是 Agent 自动化页面的主力接口。错误处理和最佳实践async def robust_example(): try: # 始终使用上下文管理器进行资源清理 async with AioClient(http://localhost:8080) as client: # 设置错误处理 result await client.shell.exec(potentially-failing-command) if not result.success: print(f命令失败{result.message}) if hasattr(result, error_code): print(f错误代码{result.error_code}) # 检查沙盒状态 status await client.sandbox.get_context() print(f沙盒运行时间{status.data.uptime}) print(f可用包{len(status.data.packages)}) except Exception as e: print(f连接错误{e}) asyncio.run(robust_example())官方还建议为长任务使用async_modeTrue搭配view轮询输出Agent 场景下可参考 Agent 集成文档 中的指数退避重试2 ** attempt与aiohttp.TCPConnector(limit10)连接池优化。Node.js SDK 实战安装与基本配置npm install agent-infra/sandbox当前仓库的 JS SDK 版本为1.0.17见 sdk/js/package.json要求 Node.js 18。导入并初始化import { AioClient } from agent-infra/sandbox; const client new AioClient({ baseUrl: https://{aio.sandbox.example}, // URL 和端口应与 Aio Sandbox 一致 timeout: 30000, // 可选请求超时毫秒 retries: 3, // 可选重试次数 retryDelay: 1000, // 可选重试之间的延迟毫秒 });同样地仓库实际导出的顶层类是SandboxClient见 sdk/js/src/Client.ts通过client.shell、client.file、client.jupyter、client.browser、client.mcp等属性访问各资源模块Shell 模块的方法名为execCommand、view、waitForProcess、createSession、listSessions见 sdk/js/src/api/resources/shell/client/Client.ts与 Python SDK 一一对应。Shell 执行与轮询const response await client.shellExec({ command: ls -la, }); if (response.success) { console.log(命令输出, response.data.output); } else { console.error(错误, response.message); } // 异步轮询训练结果适用于长期任务 const response await client.shellExecWithPolling({ command: ls -la, maxWaitTime: 60 * 1000, });shellExecWithPolling封装了execCommandwaitForProcess的轮询循环是处理训练任务、批处理等长时命令的推荐方式——waitForProcess支持seconds与max_wait_seconds两个等待参数。文件管理const fileList await client.fileList({ path: /home/gem, recursive: true, }); if (fileList.success) { console.log(文件, fileList.data.files); } else { console.error(错误, fileList.message); }注意示例中的/home/gem与仓库 docker-compose.yaml 中WORKSPACE环境变量的默认值/home/gem一致即沙盒内工作区目录。Jupyter 代码执行const jupyterResponse await client.jupyterExecute({ code: print(你好Jupyter), kernel_name: python3, }); if (jupyterResponse.success) { console.log(输出, jupyterResponse.data); } else { console.error(错误, jupyterResponse.message); }kernel_name指定内核如python3对应沙盒内 Jupyter Lab8888 端口托管的运行时。下一步选择你的集成路径官方示例按复杂度提供了三条递进路径可依项目阶段选择终端集成从 WebSocket 终端开始逐步掌握会话管理、多会话并行与 xterm.js/React 前端集成浏览器自动化从 VNC 可视化验证过渡到 CDP、浏览器 MCP 服务器再到 Playwright/Selenium 框架集成Agent 集成从 REST/MCP 基础调用进阶到多工具工作流最终接入 LangChain 或 OpenAI 助手并按 单元测试策略 与生产部署方案落地。详细的接口规范可查阅 API 文档官方仓库中的 examples 目录 还提供了 ag2、langgraph-deepagents、minimax、openai、playwright、browser-use 等框架的具体集成代码与可运行工程可作为进阶参考。赞分享AI Agent后端MCP 服务浏览器控制Agent 评测【免费下载链接】sandboxAll-in-One Sandbox for AI Agents that combines Browser, Shell, File, MCP and VSCode Server in a single Docker container.项目地址https://gitcode.com/gh_mirrors/sandbox103/sandbox点击查看免费下载相关推荐基于 CDP 将 browser-use 与 AIO Sandbox 集成构建沙箱化的 AI 浏览器自动化 Agent基于 CDP 将 browser use 与 AIO Sandbox 集成构建沙箱化的 AI 浏览器自动化 Agent 本指南以仓库中的 browser usAI Agent后端MCP 服务浏览器控制Agent 评测AIO Sandbox 浏览器自动化实战指南VNC、CDP 与 Browser MCP 一站式解析AIO Sandbox 浏览器自动化实战指南VNC、CDP 与 Browser MCP 一站式解析 本指南以 AIO Sandbox 官方文档 websiteAI Agent后端MCP 服务浏览器控制Agent 评测使用 agent-sandbox 与 browser-use 集成实现 AI 驱动浏览器自动化含 CUA GUI 回退使用 agent sandbox 与 browser use 集成实现 AI 驱动浏览器自动化含 CUA GUI 回退 本指南基于开源仓库 sandbox1AI Agent后端MCP 服务浏览器控制Agent 评测上一篇微博相册批量下载神器三步搞定海量图片收藏下一篇思源宋体TTFGoogle与Adobe联手中文免费商用字体终极指南创作声明:本文部分内容由AI辅助生成(AIGC),仅供参考
RELATED READING

延伸阅读

更多一线实战笔记与深度复盘,助您持续精进