# langsmith-cli - Doramagic AI Context Pack

> 定位：安装前体验与判断资产。它帮助宿主 AI 有一个好的开始，但不代表已经安装、执行或验证目标项目。

## 充分原则

- **充分原则，不是压缩原则**：AI Context Pack 应该充分到让宿主 AI 在开工前理解项目价值、能力边界、使用入口、风险和证据来源；它可以分层组织，但不以最短摘要为目标。
- **压缩策略**：只压缩噪声和重复内容，不压缩会影响判断和开工质量的上下文。

## 给宿主 AI 的使用方式

你正在读取 Doramagic 为 langsmith-cli 编译的 AI Context Pack。请把它当作开工前上下文：帮助用户理解适合谁、能做什么、如何开始、哪些必须安装后验证、风险在哪里。不要声称你已经安装、运行或执行了目标项目。

## Claim 消费规则

- **事实来源**：Repo Evidence + Claim/Evidence Graph；Human Wiki 只提供显著性、术语和叙事结构。
- **事实最低状态**：`supported`
- `supported`：可以作为项目事实使用，但回答中必须引用 claim_id 和证据路径。
- `weak`：只能作为低置信度线索，必须要求用户继续核实。
- `inferred`：只能用于风险提示或待确认问题，不能包装成项目事实。
- `unverified`：不得作为事实使用，应明确说证据不足。
- `contradicted`：必须展示冲突来源，不得替用户强行选择一个版本。

## 它最适合谁

- **正在使用 Claude/Codex/Cursor/Gemini 等宿主 AI 的开发者**：README 或插件配置提到多个宿主 AI。 证据：`README.md` Claim：`clm_0004` supported 0.86
- **希望把专业流程带进宿主 AI 的用户**：仓库包含 Skill 文档。 证据：`skills/langsmith/SKILL.md` Claim：`clm_0005` supported 0.86

## 它能做什么

- **AI Skill / Agent 指令资产库**（可做安装前预览）：项目包含可被宿主 AI 读取的 Skill 或 Agent 指令文件，可用于把专业流程带入 Claude、Codex、Cursor 等宿主。 证据：`skills/langsmith/SKILL.md` Claim：`clm_0001` supported 0.86
- **多宿主安装与分发**（需要安装后验证）：项目包含插件或 marketplace 配置，说明它面向一个或多个 AI 宿主的安装和分发。 证据：`.claude-plugin/marketplace.json`, `.claude-plugin/plugin.json` Claim：`clm_0002` supported 0.86
- **命令行启动或安装流程**（需要安装后验证）：项目文档中存在可执行命令，真实使用需要在本地或宿主环境中运行这些命令。 证据：`README.md` Claim：`clm_0003` supported 0.86

## 怎么开始

- `curl -sSL https://raw.githubusercontent.com/langchain-ai/langsmith-cli/main/scripts/install.sh | sh` 证据：`README.md` Claim：`clm_0006` supported 0.86
- `uv tool install langsmith-cli` 证据：`README.md` Claim：`clm_0007` supported 0.86
- `pip install langsmith-cli` 证据：`README.md` Claim：`clm_0008` supported 0.86
- `/plugin marketplace add gigaverse-app/langsmith-cli` 证据：`README.md` Claim：`clm_0009` supported 0.86
- `git clone https://github.com/gigaverse-app/langsmith-cli.git` 证据：`README.md` Claim：`clm_0010` supported 0.86

## 继续前判断卡

- **当前建议**：需要管理员/安全审批
- **为什么**：继续前可能涉及密钥、账号、外部服务或敏感上下文，建议先经过管理员或安全审批。

### 30 秒判断

- **现在怎么做**：需要管理员/安全审批
- **最小安全下一步**：先跑 Prompt Preview；若涉及凭证或企业环境，先审批再试装
- **先别相信**：真实输出质量不能在安装前相信。
- **继续会触碰**：命令执行、宿主 AI 配置、本地环境或项目文件

### 现在可以相信

- **适合人群线索：正在使用 Claude/Codex/Cursor/Gemini 等宿主 AI 的开发者**（supported）：有 supported claim 或项目证据支撑，但仍不等于真实安装效果。 证据：`README.md` Claim：`clm_0004` supported 0.86
- **适合人群线索：希望把专业流程带进宿主 AI 的用户**（supported）：有 supported claim 或项目证据支撑，但仍不等于真实安装效果。 证据：`skills/langsmith/SKILL.md` Claim：`clm_0005` supported 0.86
- **能力存在：AI Skill / Agent 指令资产库**（supported）：可以相信项目包含这类能力线索；是否适合你的具体任务仍要试用或安装后验证。 证据：`skills/langsmith/SKILL.md` Claim：`clm_0001` supported 0.86
- **能力存在：多宿主安装与分发**（supported）：可以相信项目包含这类能力线索；是否适合你的具体任务仍要试用或安装后验证。 证据：`.claude-plugin/marketplace.json`, `.claude-plugin/plugin.json` Claim：`clm_0002` supported 0.86
- **能力存在：命令行启动或安装流程**（supported）：可以相信项目包含这类能力线索；是否适合你的具体任务仍要试用或安装后验证。 证据：`README.md` Claim：`clm_0003` supported 0.86
- **存在 Quick Start / 安装命令线索**（supported）：可以相信项目文档出现过启动或安装入口；不要因此直接在主力环境运行。 证据：`README.md` Claim：`clm_0006` supported 0.86

### 现在还不能相信

- **真实输出质量不能在安装前相信。**（unverified）：Prompt Preview 只能展示引导方式，不能证明真实项目中的结果质量。
- **宿主 AI 版本兼容性不能在安装前相信。**（unverified）：Claude、Cursor、Codex、Gemini 等宿主加载规则和版本差异必须在真实环境验证。
- **不会污染现有宿主 AI 行为，不能直接相信。**（inferred）：Skill、plugin、AGENTS/CLAUDE/GEMINI 指令可能改变宿主 AI 的默认行为。 证据：`.claude-plugin/marketplace.json`, `.claude-plugin/plugin.json`, `AGENTS.md`, `CLAUDE.md` 等
- **可安全回滚不能默认相信。**（unverified）：除非项目明确提供卸载和恢复说明，否则必须先在隔离环境验证。
- **真实安装后是否与用户当前宿主 AI 版本兼容？**（unverified）：兼容性只能通过实际宿主环境验证。 证据：`.claude-plugin/marketplace.json`, `.claude-plugin/plugin.json`
- **项目输出质量是否满足用户具体任务？**（unverified）：安装前预览只能展示流程和边界，不能替代真实评测。
- **安装命令是否需要网络、权限或全局写入？**（unverified）：这影响企业环境和个人环境的安装风险。 证据：`README.md`

### 继续会触碰什么

- **命令执行**：包管理器、网络下载、本地插件目录、项目配置或用户主目录。 原因：运行第一条命令就可能产生环境改动；必须先判断是否值得跑。 证据：`README.md`
- **宿主 AI 配置**：Claude/Codex/Cursor/Gemini/OpenCode 等宿主的 plugin、Skill 或规则加载配置。 原因：宿主配置会改变 AI 后续工作方式，可能和用户已有规则冲突。 证据：`.claude-plugin/marketplace.json`, `.claude-plugin/plugin.json`, `AGENTS.md`, `CLAUDE.md` 等
- **本地环境或项目文件**：安装结果、插件缓存、项目配置或本地依赖目录。 原因：安装前无法证明写入范围和回滚方式，需要隔离验证。 证据：`.claude-plugin/marketplace.json`, `.claude-plugin/plugin.json`, `README.md`
- **环境变量 / API Key**：项目入口文档明确出现 API key、token、secret 或账号凭证配置。 原因：如果真实安装需要凭证，应先使用测试凭证并经过权限/合规判断。 证据：`CLAUDE.md`, `README.md`, `docs/dev/CI_BEST_PRACTICES.md`, `docs/dev/TESTING_PERFORMANCE.md` 等
- **宿主 AI 上下文**：AI Context Pack、Prompt Preview、Skill 路由、风险规则和项目事实。 原因：导入上下文会影响宿主 AI 后续判断，必须避免把未验证项包装成事实。

### 最小安全下一步

- **先跑 Prompt Preview**：用安装前交互式试用判断工作方式是否匹配，不需要授权或改环境。（适用：任何项目都适用，尤其是输出质量未知时。）
- **只在隔离目录或测试账号试装**：避免安装命令污染主力宿主 AI、真实项目或用户主目录。（适用：存在命令执行、插件配置或本地写入线索时。）
- **先备份宿主 AI 配置**：Skill、plugin、规则文件可能改变 Claude/Cursor/Codex 的默认行为。（适用：存在插件 manifest、Skill 或宿主规则入口时。）
- **不要使用真实生产凭证**：环境变量/API key 一旦进入宿主或工具链，可能产生账号和合规风险。（适用：出现 API、TOKEN、KEY、SECRET 等环境线索时。）
- **安装后只验证一个最小任务**：先验证加载、兼容、输出质量和回滚，再决定是否深用。（适用：准备从试用进入真实工作流时。）

### 退出方式

- **保留安装前状态**：记录原始宿主配置和项目状态，后续才能判断是否可恢复。
- **准备移除宿主 plugin / Skill / 规则入口**：如果试装后行为异常，可以把宿主 AI 恢复到试装前状态。
- **记录安装命令和写入路径**：没有明确卸载说明时，至少要知道哪些目录或配置需要手动清理。
- **准备撤销测试 API key 或 token**：测试凭证泄露或误用时，可以快速止损。
- **如果没有回滚路径，不进入主力环境**：不可回滚是继续前阻断项，不应靠信任或运气继续。

## 哪些只能预览

- 解释项目适合谁和能做什么
- 基于项目文档演示典型对话流程
- 帮助用户判断是否值得安装或继续研究

## 哪些必须安装后验证

- 真实安装 Skill、插件或 CLI
- 执行脚本、修改本地文件或访问外部服务
- 验证真实输出质量、性能和兼容性

## 边界与风险判断卡

- **把安装前预览误认为真实运行**：用户可能高估项目已经完成的配置、权限和兼容性验证。 处理方式：明确区分 prompt_preview_can_do 与 runtime_required。 Claim：`clm_0011` inferred 0.45
- **宿主 AI 插件或 Skill 规则冲突**：新规则可能改变用户现有宿主 AI 的工作方式。 处理方式：安装前先检查插件 manifest 和 Skill 文件，必要时隔离测试。 证据：`.claude-plugin/marketplace.json`, `.claude-plugin/plugin.json` Claim：`clm_0012` supported 0.86
- **命令执行会修改本地环境**：安装命令可能写入用户主目录、宿主插件目录或项目配置。 处理方式：先在隔离环境或测试账号中运行。 证据：`README.md` Claim：`clm_0013` supported 0.86
- **待确认**：真实安装后是否与用户当前宿主 AI 版本兼容？。原因：兼容性只能通过实际宿主环境验证。
- **待确认**：项目输出质量是否满足用户具体任务？。原因：安装前预览只能展示流程和边界，不能替代真实评测。
- **待确认**：安装命令是否需要网络、权限或全局写入？。原因：这影响企业环境和个人环境的安装风险。

## 开工前工作上下文

### 加载顺序

- 先读取 how_to_use.host_ai_instruction，建立安装前判断资产的边界。
- 读取 claim_graph_summary，确认事实来自 Claim/Evidence Graph，而不是 Human Wiki 叙事。
- 再读取 intended_users、capabilities 和 quick_start_candidates，判断用户是否匹配。
- 需要执行具体任务时，优先查 role_skill_index，再查 evidence_index。
- 遇到真实安装、文件修改、网络访问、性能或兼容性问题时，转入 risk_card 和 boundaries.runtime_required。

### 任务路由

- **AI Skill / Agent 指令资产库**：先基于 role_skill_index / evidence_index 帮用户挑选可用角色、Skill 或工作流。 边界：可做安装前 Prompt 体验。 证据：`skills/langsmith/SKILL.md` Claim：`clm_0001` supported 0.86
- **多宿主安装与分发**：先说明这是安装后验证能力，再给出安装前检查清单。 边界：必须真实安装或运行后验证。 证据：`.claude-plugin/marketplace.json`, `.claude-plugin/plugin.json` Claim：`clm_0002` supported 0.86
- **命令行启动或安装流程**：先说明这是安装后验证能力，再给出安装前检查清单。 边界：必须真实安装或运行后验证。 证据：`README.md` Claim：`clm_0003` supported 0.86

### 上下文规模

- 文件总数：87
- 重要文件覆盖：40/87
- 证据索引条目：64
- 角色 / Skill 条目：1

### 证据不足时的处理

- **missing_evidence**：说明证据不足，要求用户提供目标文件、README 段落或安装后验证记录；不要补全事实。
- **out_of_scope_request**：说明该任务超出当前 AI Context Pack 证据范围，并建议用户先查看 Human Manual 或真实安装后验证。
- **runtime_request**：给出安装前检查清单和命令来源，但不要替用户执行命令或声称已执行。
- **source_conflict**：同时展示冲突来源，标记为待核实，不要强行选择一个版本。

## Prompt Recipes

### 适配判断

- 目标：判断这个项目是否适合用户当前任务。
- 预期输出：适配结论、关键理由、证据引用、安装前可预览内容、必须安装后验证内容、下一步建议。

```text
请基于 langsmith-cli 的 AI Context Pack，先问我 3 个必要问题，然后判断它是否适合我的任务。回答必须包含：适合谁、能做什么、不能做什么、是否值得安装、证据来自哪里。所有项目事实必须引用 evidence_refs、source_paths 或 claim_id。
```

### 安装前体验

- 目标：让用户在安装前感受核心工作流，同时避免把预览包装成真实能力或营销承诺。
- 预期输出：一段带边界标签的体验剧本、安装后验证清单和谨慎建议；不含真实运行承诺或强营销表述。

```text
请把 langsmith-cli 当作安装前体验资产，而不是已安装工具或真实运行环境。

请严格输出四段：
1. 先问我 3 个必要问题。
2. 给出一段“体验剧本”：用 [安装前可预览]、[必须安装后验证]、[证据不足] 三种标签展示它可能如何引导工作流。
3. 给出安装后验证清单：列出哪些能力只有真实安装、真实宿主加载、真实项目运行后才能确认。
4. 给出谨慎建议：只能说“值得继续研究/试装”“先补充信息后再判断”或“不建议继续”，不得替项目背书。

硬性边界：
- 不要声称已经安装、运行、执行测试、修改文件或产生真实结果。
- 不要写“自动适配”“确保通过”“完美适配”“强烈建议安装”等承诺性表达。
- 如果描述安装后的工作方式，必须使用“如果安装成功且宿主正确加载 Skill，它可能会……”这种条件句。
- 体验剧本只能写成“示例台词/假设流程”：使用“可能会询问/可能会建议/可能会展示”，不要写“已写入、已生成、已通过、正在运行、正在生成”。
- Prompt Preview 不负责给安装命令；如用户准备试装，只能提示先阅读 Quick Start 和 Risk Card，并在隔离环境验证。
- 所有项目事实必须来自 supported claim、evidence_refs 或 source_paths；inferred/unverified 只能作风险或待确认项。

```

### 角色 / Skill 选择

- 目标：从项目里的角色或 Skill 中挑选最匹配的资产。
- 预期输出：候选角色或 Skill 列表，每项包含适用场景、证据路径、风险边界和是否需要安装后验证。

```text
请读取 role_skill_index，根据我的目标任务推荐 3-5 个最相关的角色或 Skill。每个推荐都要说明适用场景、可能输出、风险边界和 evidence_refs。
```

### 风险预检

- 目标：安装或引入前识别环境、权限、规则冲突和质量风险。
- 预期输出：环境、权限、依赖、许可、宿主冲突、质量风险和未知项的检查清单。

```text
请基于 risk_card、boundaries 和 quick_start_candidates，给我一份安装前风险预检清单。不要替我执行命令，只说明我应该检查什么、为什么检查、失败会有什么影响。
```

### 宿主 AI 开工指令

- 目标：把项目上下文转成一次对话开始前的宿主 AI 指令。
- 预期输出：一段边界明确、证据引用明确、适合复制给宿主 AI 的开工前指令。

```text
请基于 langsmith-cli 的 AI Context Pack，生成一段我可以粘贴给宿主 AI 的开工前指令。这段指令必须遵守 not_runtime=true，不能声称项目已经安装、运行或产生真实结果。
```

## 角色 / Skill 索引

- 共索引 1 个角色 / Skill / 项目文档条目。

- **langsmith**（skill）：Inspect and manage LangSmith traces, runs, datasets, and prompts using the 'langsmith-cli'. 激活提示：当用户任务与“langsmith”描述的流程高度相关时，先用它做安装前体验，再决定是否安装。 证据：`skills/langsmith/SKILL.md`

## 证据索引

- 共索引 64 条证据。

- **🛠️ LangSmith CLI**（documentation）：The Modern CLI for LangSmith Lightning-fast • Context-efficient • Built for humans and AI agents 证据：`README.md`
- **LangSmith Tool**（skill_instruction）：Use this tool to debug AI chains, inspect past runs, manage datasets, and analyze token costs in LangSmith. 证据：`skills/langsmith/SKILL.md`
- **Plugin**（structured_config）：{ "name": "langsmith-cli", "version": "0.10.3", "description": "A context-efficient interface for LangSmith observability and evaluations.", "author": { "name": "Aviad Rozenhek", "email": "aviadr1@gmail.com" }, "homepage": "https://github.com/gigaverse-app/langsmith-cli", "repository": "https://github.com/gigaverse-app/langsmith-cli", "license": "MIT", "keywords": "langsmith", "observability", "evaluations", "tracing" , "skills": "./skills/" } 证据：`.claude-plugin/plugin.json`
- **1. 🛡️ Type Safety & Data Integrity Zero Tolerance for Weak Types**（documentation）：SYSTEM INSTRUCTION : You are acting as a Senior Python Engineer. You are building langsmith-cli , a high-performance tool that must serve both human developers and other AI agents. CRITICAL : Read and adhere to the following 5 Engineering Standards. Deviations will be rejected. 证据：`AGENTS.md`
- **CLAUDE.md**（documentation）：This file provides guidance to Claude Code claude.ai/code when working with code in this repository. 证据：`CLAUDE.md`
- **Pipes to CLI Reference**（documentation）：Quick Reference Guide: Converting Piped Commands to Native CLI Features 证据：`docs/PIPES_TO_CLI_REFERENCE.md`
- **LangSmith CLI Real-World Examples**（documentation）：This document provides practical workflows and use cases for common LangSmith operations. 证据：`skills/langsmith/docs/examples.md`
- **LangSmith CLI Quick Reference**（documentation）：Quick reference guide for all langsmith-cli commands. For detailed documentation, see the references/ ../references/ folder. 证据：`skills/langsmith/docs/reference.md`
- **Marketplace**（structured_config）：{ "name": "langsmith-cli", "owner": { "name": "Aviad Rozenhek", "email": "aviadr1@gmail.com" }, "metadata": { "description": "LangSmith CLI plugin marketplace", "version": "0.10.3" }, "plugins": { "name": "langsmith-cli", "source": "./", "description": "A context-efficient interface for LangSmith observability and evaluations.", "version": "0.10.3", "author": { "name": "Gigaverse", "email": "aviadr1@gmail.com" }, "repository": "https://github.com/gigaverse-app/langsmith-cli", "license": "MIT", "keywords": "langsmith", "observability", "evaluations", "tracing" , "category": "productivity" } } 证据：`.claude-plugin/marketplace.json`
- **License**（source_file）：Permission is hereby granted, free of charge, to any person obtaining a copy of this software and associated documentation files the "Software" , to deal in the Software without restriction, including without limitation the rights to use, copy, modify, merge, publish, distribute, sublicense, and/or sell copies of the Software, and to permit persons to whom the Software is furnished to do so, subject to the following conditions: 证据：`LICENSE`
- **📐 Design Specification: langsmith-cli**（documentation）：This is a comprehensive design document for langsmith-cli , reverse-engineered from the LangSmith MCP server's source code and enhanced with "Simon Willison-style" CLI best practices. 证据：`docs/COMMANDS_DESIGN.md`
- **1. Repository Identity**（documentation）：Here is the complete repository specification. This setup positions the project not just as a "script," but as a serious developer tool that happens to work perfectly with Claude. 证据：`docs/PRD.md`
- **Quality of Life Features**（documentation）：This document describes the implemented quality-of-life improvements for langsmith-cli. 证据：`docs/QOL_FEATURES.md`
- **Quality of Life Improvements - Analysis**（documentation）：Quality of Life Improvements - Analysis 证据：`docs/QOL_IMPROVEMENTS.md`
- **Repository Description The One-Liner**（documentation）：Repository Description The One-Liner 证据：`docs/TLDR.md`
- **The Problem with MCP Servers**（documentation）：If you're using LangSmith with Claude Code or any AI coding agent , you're probably running the official MCP server. It works. But every session, it injects ~5,000 tokens of tool schemas into your context window — whether you touch LangSmith or not. 证据：`docs/devto-article.md`
- **CI/CD Best Practices**（documentation）：This document explains the best practices implemented in our GitHub Actions CI/CD pipeline. 证据：`docs/dev/CI_BEST_PRACTICES.md`
- **Codecov Setup Guide**（documentation）：Quick guide to set up Codecov integration for coverage tracking and badges. 证据：`docs/dev/CODECOV_SETUP.md`
- **Implementation Plan: Stratified Sampling and Analytics Commands**（documentation）：Implementation Plan: Stratified Sampling and Analytics Commands 证据：`docs/dev/IMPLEMENTATION_PLAN.md`
- **Questions from Partner Team on Stratified Sampling & Analytics Implementation**（documentation）：Questions from Partner Team on Stratified Sampling & Analytics Implementation 证据：`docs/dev/LANGSMITH_TEAM_QUESTIONS.md`
- **LangSmith MCP Feature Parity**（documentation）：This document tracks feature parity between the langsmith-cli and the official LangSmith MCP server. 证据：`docs/dev/MCP_PARITY.md`
- **Publishing to PyPI**（documentation）：This document describes how to publish langsmith-cli to PyPI using GitHub Actions. 证据：`docs/dev/PUBLISHING.md`
- **PyPI Publishing Setup - Summary**（documentation）：Added complete PyPI metadata: - ✅ Project description - ✅ License declaration - ✅ Author information - ✅ Keywords for PyPI search - ✅ PyPI classifiers - ✅ Project URLs homepage, repository, issues, docs - ✅ Build system configuration hatchling 证据：`docs/dev/PYPI_SETUP_SUMMARY.md`
- **User Directives from Session**（documentation）：This document compiles the specific directives and preferences provided by the user during the initial setup session. 证据：`docs/dev/SESSION_DIRECTIVES.md`
- **Type Safety Guidelines**（documentation）：Philosophy: Zero Tolerance for Weak Types 证据：`docs/dev/TYPE_SAFETY_GUIDE.md`
- **Datasets**（documentation）：Options: - --limit INTEGER - Maximum results default: 20 - --name TEXT - Filter by exact dataset name - --name-contains TEXT - Filter by name substring - --dataset-ids TEXT - Comma-separated list of dataset UUIDs - --data-type TEXT - Filter by type: kv , llm , or chat - --metadata TEXT - Filter by metadata JSON string - --exclude TEXT - Exclude items containing substring repeatable - --fields TEXT - Comma-separated field names to include - --count - Output only the count of results - --output TEXT - Write output to file JSONL format 证据：`skills/langsmith/references/datasets.md`
- **Examples**（documentation）：List examples in a dataset with advanced filtering. 证据：`skills/langsmith/references/examples.md`
- **Runs Traces**（documentation）：Options: - --project TEXT - Project name default: "default" - --project-id TEXT - Project UUID bypasses name resolution, fastest lookup - --project-name TEXT - Substring/contains match for project names - --project-name-exact TEXT - Exact project name match - --project-name-pattern TEXT - Wildcard pattern for project names e.g., 'dev/ ' - --project-name-regex TEXT - Regex pattern for project names - --limit INTEGER - Maximum results default: 10 - --status success error - Filter by status - --failed - Show only failed/error runs shorthand for --status error - --succeeded - Show only successful runs shorthand for --status success - --slow - Filter to slow runs latency 5s - --recent - Filter t… 证据：`skills/langsmith/references/runs.md`
- **.Env**（source_file）：LANGSMITH API KEY=lsv2 ... LANGSMITH PROJECT=default 证据：`.env.example`
- **Testing Performance Guide**（documentation）：This document explains how to run tests efficiently during development and CI/CD. 证据：`docs/dev/TESTING_PERFORMANCE.md`
- **LangSmith CLI Testing Strategy**（documentation）：Since LangSmith only retains traces for 400 days, we need a two-tier testing strategy: 证据：`docs/dev/TESTING_STRATEGY.md`
- **Main**（source_file）：def main 证据：`main.py`
- **IMPORTANT: When bumping this version, also update .claude-plugin/plugin.json**（source_file）：project name = "langsmith-cli" IMPORTANT: When bumping this version, also update .claude-plugin/plugin.json version = "0.10.3" description = "Context-efficient CLI for LangSmith. Built for humans and agents." readme = "README.md" requires-python = " =3.12" license = { file = "LICENSE" } authors = { name = "Aviad Rozenhek" } keywords = "langsmith", "langchain", "cli", "observability", "tracing", "debugging", "ai", "llm" classifiers = "Development Status :: 4 - Beta", "Intended Audience :: Developers", "License :: OSI Approved :: MIT License", "Programming Language :: Python :: 3", "Programming Language :: Python :: 3.12", "Topic :: Software Development :: Libraries :: Python Modules", "Topic… 证据：`pyproject.toml`
- **Standalone installer wrapper for langsmith-cli Windows**（source_file）：Standalone installer wrapper for langsmith-cli Windows This script downloads and runs the Python installer. Usage: irm https://raw.githubusercontent.com/langchain-ai/langsmith-cli/main/scripts/install.ps1 iex iwr -useb https://raw.githubusercontent.com/langchain-ai/langsmith-cli/main/scripts/install.ps1 iex 证据：`scripts/install.ps1`
- **Normalize architecture names**（source_file）：PACKAGE NAME = "langsmith-cli" GITHUB REPO = "langchain-ai/langsmith-cli" MIN PYTHON VERSION = 3, 12 ⋮---- BOLD = "\033 1m" GREEN = "\033 32m" YELLOW = "\033 33m" RED = "\033 31m" RESET = "\033 0m" ⋮---- def log info msg: str - None ⋮---- def log warning msg: str - None ⋮---- """Print warning message.""" ⋮---- def log error msg: str - None ⋮---- """Print error message.""" ⋮---- def get platform info - tuple str, str, str ⋮---- """ Detect platform and architecture. Returns: Tuple of os name, arch, platform str - os name: 'linux', 'darwin', 'windows' - arch: 'x86 64', 'aarch64', 'arm64', etc. - platform str: Combined string like 'linux-x86 64' """ os name = platform.system .lower arch = platf… 证据：`scripts/install.py`
- **Download installer**（source_file）：set -e BOLD="\033 1m" GREEN="\033 32m" YELLOW="\033 33m" RED="\033 31m" RESET="\033 0m" INSTALLER URL="https://raw.githubusercontent.com/langchain-ai/langsmith-cli/main/scripts/install.py" TEMP INSTALLER="/tmp/langsmith-cli-install.py" log info { printf "${GREEN}✓${RESET} %s\n" "$1" } log error { printf "${RED}✗${RESET} %s\n" "$1" &2 } log warning { printf "${YELLOW}⚠${RESET} %s\n" "$1" &2 } check python { if command -v python3 /dev/null 2 &1; then PYTHON CMD="python3" elif command -v python /dev/null 2 &1; then if python -c "import sys; sys.exit 0 if sys.version info = 3, 12 else 1 " 2 /dev/null; then PYTHON CMD="python" else return 1 fi else return 1 fi return 0 } main { printf "\n${BOLD}… 证据：`scripts/install.sh`
- **Remove comment lines**（source_file）：BOLD = "\033 1m" GREEN = "\033 32m" YELLOW = "\033 33m" RED = "\033 31m" RESET = "\033 0m" ⋮---- def log info msg: str - None ⋮---- def log warning msg: str - None ⋮---- """Print warning message.""" ⋮---- def log error msg: str - None ⋮---- """Print error message.""" ⋮---- def get install receipt path - Path ⋮---- """Get path to install receipt file.""" os name = platform.system .lower ⋮---- config dir = ⋮---- config home = os.environ.get "XDG CONFIG HOME", str Path.home / ".config" config dir = Path config home / "langsmith-cli" ⋮---- def load install receipt - Optional dict ⋮---- receipt path = get install receipt path ⋮---- def remove directory path: Path, description: str - bool ⋮---- "… 证据：`scripts/uninstall.py`
- **Cache**（source_file）：BINARY STRIP THRESHOLD = 10 000 ⋮---- BASE64 CHARS = set ⋮---- def is likely base64 s: str - bool ⋮---- sample = s :200 ⋮---- def is data uri s: str - bool ⋮---- @overload def strip binary data obj: dict str, Any - dict str, Any : ... ⋮---- @overload def strip binary data obj: list Any - list Any : ... ⋮---- changed = False new dict: dict str, dict list str int float bool None = {} ⋮---- new v = strip binary data v ⋮---- changed = True ⋮---- new list: list dict list str int float bool None = ⋮---- new item = strip binary data item ⋮---- media type = obj.split ";" 0 .replace "data:", "" ⋮---- class CacheMetadata BaseModel ⋮---- """Metadata sidecar for a cached project's runs.""" ⋮---- projec… 证据：`src/langsmith_cli/cache.py`
- **Init**（source_file）：all = 证据：`src/langsmith_cli/commands/runs/__init__.py`
- **Group**（source_file）：class LazyConsole ⋮---- def init self - None ⋮---- def get console self - Any ⋮---- def print self, args: Any, kwargs: Any - None ⋮---- console = LazyConsole ⋮---- @click.group def runs ⋮---- API MAX LIMIT = 100 ⋮---- def make fetch runs - Any ⋮---- def fetch runs c: Any, proj: str None, kw: Any - Any ⋮---- requested limit = kw.pop "limit", None sdk limit = requested limit ⋮---- sdk limit = None ⋮---- it = c.list runs project name=proj, limit=sdk limit, kw ⋮---- it = c.list runs limit=sdk limit, kw 证据：`src/langsmith_cli/commands/runs/_group.py`
- **Create output directory**（source_file）：logger = ctx.obj "logger" ⋮---- client = get or create client ctx ⋮---- Create output directory out dir = pathlib.Path directory ⋮---- Resolve project filters pq = resolve project filters projects to query = pq.names ⋮---- is root = resolve root scope roots=roots, all runs=all runs, is root=is root ⋮---- Build filter using shared helper reuse canonical filter builder ⋮---- Fetch runs using the shared fetch from projects helper def fetch runs c: Any, proj: str None, kw: Any - Any ⋮---- result = fetch from projects fetched runs: list Run = result.items ⋮---- If all sources failed, raise with suggestions reports failures internally . Otherwise, report partial failures. ⋮---- fetched runs = fet… 证据：`src/langsmith_cli/commands/runs/export_cmd.py`
- **Get matching projects**（source_file）：@click.pass context def get run ctx, run id, fields, output, follow children ⋮---- client = get or create client ctx run = client.read run run id ⋮---- trace id = run.trace id or run.id ⋮---- children = epoch = datetime.min.replace tzinfo= timezone.utc ⋮---- child data = filter fields c, fields for c in children data = filter fields run, fields ⋮---- logger = ctx.obj "logger" ⋮---- Get matching projects pq = resolve project filters projects to query = pq.names ⋮---- Build filter using shared helper ⋮---- Search projects in order until we find a run latest run = None failed projects = run kwargs: dict str, Any = dict ⋮---- Lazy: SDK exception imports are cold-path for get-latest . ⋮---- fetc… 证据：`src/langsmith_cli/commands/runs/get_cmd.py`
- **Resolve project filters --project-id bypasses name resolution**（source_file）：FREEFORM QUERY REJECTION DETAIL = "Failed to generate filter from freeform query" ⋮---- def is query rejection exc: BaseException - bool ⋮---- current: BaseException None = exc seen: set int = set ⋮---- response = current.response ⋮---- current = current. cause or current. context ⋮---- def error body detail body: object - bool ⋮---- parsed = json.loads body ⋮---- detail = parsed "detail" if "detail" in parsed else None ⋮---- logger = ctx.obj "logger" ⋮---- format type = determine output format output format, is json context ctx is machine readable = configure logger streams ⋮---- limit = 0 ⋮---- client = get or create client ctx ⋮---- Resolve project filters --project-id bypasses name reso… 证据：`src/langsmith_cli/commands/runs/list_cmd.py`
- **Build time filters and combine with additional filter**（source_file）：query arg = query grep arg = grep grep in arg = grep in failed arg = failed search filters = ⋮---- query arg = None grep arg = grep arg or query grep in arg = grep in arg or search in ⋮---- failed arg = True ⋮---- grep arg = input contains grep in arg = "inputs" ⋮---- grep arg = output contains grep in arg = "outputs" ⋮---- combined filter = combine fql filters search filters ⋮---- logger = ctx.obj "logger" ⋮---- Build time filters and combine with additional filter time filters = build time fql filters since=since, last=last, before=before base filters = time filters.copy ⋮---- Combine base filters into a single filter base filter = combine fql filters base filters ⋮---- client = get or cr… 证据：`src/langsmith_cli/commands/runs/search_cmd.py`
- **Stats Cmd**（source_file）：client = get or create client ctx ⋮---- pq = resolve project filters ⋮---- resolved project ids: list Any = ⋮---- p = client.read project project name=proj name ⋮---- stats = client.get run stats project ids=resolved project ids ⋮---- table title = f"Stats: id:{pq.project id}" ⋮---- table title = f"Stats: {pq.names 0 }" ⋮---- table title = f"Stats: {len pq.names } projects" ⋮---- table = Table title=table title 证据：`src/langsmith_cli/commands/runs/stats_cmd.py`
- **Field Analysis**（source_file）：LANG DETECT SAMPLE SIZE = 500 LANG DETECT MIN LENGTH = 30 LANG DETECT MAX SAMPLES = 100 LANG TOP N = 3 ⋮---- @dataclass class FieldStats ⋮---- path: str field type: str present count: int total count: int ⋮---- length min: int None = None length max: int None = None length avg: float None = None length p50: float None = None ⋮---- num min: float None = None num max: float None = None num avg: float None = None num p50: float None = None num sum: float None = None ⋮---- languages: dict str, float = field default factory=dict ⋮---- sample: str None = None ⋮---- @property def present pct self - float ⋮---- def to dict self - dict str, Any ⋮---- result: dict str, Any = { ⋮---- class SchemaNode… 证据：`src/langsmith_cli/field_analysis.py`
- **Convert wildcards to regex**（source_file）：T = TypeVar "T" ⋮---- class ModelDumpable Protocol ⋮---- ModelT = TypeVar "ModelT", bound=ModelDumpable ⋮---- @overload def filter fields data: list ModelT , fields: str None - list dict str, Any : ... ⋮---- field set = parse fields option fields ⋮---- def parse fields option fields: str None - set str None ⋮---- def require confirmation skip: bool, prompt: str - None ⋮---- filtered items = ⋮---- name = name getter item ⋮---- def add grep options func: Callable ..., Any - Callable ..., Any ⋮---- func = click.option ⋮---- def add metadata filter options func: Callable ..., Any - Callable ..., Any ⋮---- def build metadata fql filters metadata filters: tuple str, ... - list str ⋮---- filters:… 证据：`src/langsmith_cli/filtering.py`
- **Parse ISO timestamp or relative time**（source_file）：class StatusFilter BaseModel ⋮---- model config = {"frozen": True} ⋮---- status: Literal "error", "success" None = None failed: bool = False succeeded: bool = False ⋮---- @field validator "status" @classmethod def validate status cls, v: str None - str None ⋮---- def to sdk params self - dict str, Any ⋮---- def needs client filtering self - bool ⋮---- class TimeFilter BaseModel ⋮---- since: str None = None last: str None = None recent: bool = False today: bool = False ⋮---- def to fql filters self - list str ⋮---- filters: list str = ⋮---- start timestamp = parse time input self.since duration = parse time duration self.last end timestamp = start timestamp + duration ⋮---- Parse ISO timesta… 证据：`src/langsmith_cli/filters.py`
- **Don't exit for conflicts - they're often non-fatal**（source_file）：config file = get credentials file ⋮---- def http status from exception exc: BaseException - int None ⋮---- current: BaseException None = exc seen: set int = set ⋮---- response = current.response ⋮---- current = current. cause or current. context ⋮---- def get console - Any ⋮---- def is json mode ctx: click.Context - bool ⋮---- def close cached client ctx: click.Context - None ⋮---- def command path for ctx ctx: click.Context - str ⋮---- parts: list str = cur: click.Context None = ctx ⋮---- cur = cur.parent ⋮---- parts = root name or "langsmith-cli" command = root command ⋮---- command = command.commands token ⋮---- def command path from exception exc: BaseException - str None ⋮---- deepest… 证据：`src/langsmith_cli/main.py`
- **Normalize to list**（source_file）：def json dumps obj: Any, kwargs: Any - str ⋮---- def is json context ctx: click.Context - bool ⋮---- use stderr = is machine readable output ⋮---- class ConsoleProtocol Protocol ⋮---- def print self, args: Any, kwargs: Any - None ⋮---- class LazyConsole ⋮---- def init self - None ⋮---- def get console self - Any ⋮---- class ModelDumpable Protocol ⋮---- data = {k: v for k, v in item.items if k in fields} for item in data ⋮---- writer = csv.DictWriter sys.stdout, fieldnames=data 0 .keys ⋮---- """Determine the output format to use. Args: output format: Explicitly requested format None if not specified json flag: Whether --json global flag was used Returns: Format to use "json", "csv", "yaml",… 证据：`src/langsmith_cli/output.py`
- **Release Process**（documentation）：This project uses explicit versioning where versions are tracked in multiple files and synchronized during release. 证据：`RELEASING.md`
- **Cache Recipes: Schema Discovery, Python & DuckDB Queries**（documentation）：Cache Recipes: Schema Discovery, Python & DuckDB Queries 证据：`skills/langsmith/references/cache-recipes.md`
- **Cost & Token Analysis Reference**（documentation）：Flag Commands What It Does ------ ---------- -------------- --from-cache runs usage , runs pricing Read from local JSONL cache fast, no API --group-by metadata: runs usage , runs analyze Group results by metadata or tag field --breakdown model runs usage Add model dimension to aggregation --breakdown provider runs usage Add provider Google, OpenAI, Anthropic, etc. --breakdown gateway runs usage Add gateway groq, cerebras, openai, google genai, etc. --breakdown project runs usage Add project dimension --apply-pricing runs usage YAML pricing file to fill missing costs --format json\ yaml runs pricing Output pricing as JSON or YAML for automation --lookup runs pricing Opt in to OpenRouter look… 证据：`skills/langsmith/references/cost-analysis.md`
- **Filter Query Language FQL**（documentation）：Advanced filtering for runs list and examples list using FQL expressions. 证据：`skills/langsmith/references/fql.md`
- **Installation Guide**（documentation）：- Python 3.12 or later - LangSmith API key get one at smith.langchain.com https://smith.langchain.com 证据：`skills/langsmith/references/installation.md`
- **Projects**（documentation）：List all LangSmith projects sessions . 证据：`skills/langsmith/references/projects.md`
- **Prompts**（documentation）：Options: - --limit INTEGER - Maximum results default: 20 - --public / --private - Filter by visibility preferred - --is-public BOOLEAN - Legacy visibility filter: true or false ; do not combine with paired visibility flags - --exclude TEXT - Exclude items containing substring repeatable - --fields TEXT - Comma-separated field names to include - --count - Output only the count of results - --output TEXT - Write output to file JSONL format 证据：`skills/langsmith/references/prompts.md`
- **Content Search & Filtering Reference**（documentation）：Content Search & Filtering Reference 证据：`skills/langsmith/references/search.md`
- **Troubleshooting & Configuration**（documentation）：bash Required export LANGSMITH API KEY="lsv2 pt ..." 证据：`skills/langsmith/references/troubleshooting.md`
- **Settings**（structured_config）：{ "enabledPlugins": { "superpowers@claude-plugins-official": true, "ralph-loop@claude-plugins-official": true, "kaizen@kaizen": true }, "extraKnownMarketplaces": { "kaizen": { "source": { "source": "github", "repo": "Garsson-io/kaizen" } } } } 证据：`.claude/settings.json`
- 其余 4 条证据见 `AI_CONTEXT_PACK.json` 或 `EVIDENCE_INDEX.json`。

## 宿主 AI 必须遵守的规则

- **把本资产当作开工前上下文，而不是运行环境。**：AI Context Pack 只包含证据化项目理解，不包含目标项目的可执行状态。 证据：`README.md`, `skills/langsmith/SKILL.md`, `.claude-plugin/plugin.json`
- **回答用户时区分可预览内容与必须安装后才能验证的内容。**：安装前体验的消费者价值来自降低误装和误判，而不是伪装成真实运行。 证据：`README.md`, `skills/langsmith/SKILL.md`, `.claude-plugin/plugin.json`

## 用户开工前应该回答的问题

- 你准备在哪个宿主 AI 或本地环境中使用它？
- 你只是想先体验工作流，还是准备真实安装？
- 你最在意的是安装成本、输出质量、还是和现有规则的冲突？

## 验收标准

- 所有能力声明都能回指到 evidence_refs 中的文件路径。
- AI_CONTEXT_PACK.md 没有把预览包装成真实运行。
- 用户能在 3 分钟内看懂适合谁、能做什么、如何开始和风险边界。

---

## Doramagic Context Augmentation

下面内容用于强化 Repomix/AI Context Pack 主体。Human Manual 只提供阅读骨架；踩坑日志会被转成宿主 AI 必须遵守的工作约束。

## Human Manual 骨架

使用规则：这里只是项目阅读路线和显著性信号，不是事实权威。具体事实仍必须回到 repo evidence / Claim Graph。

宿主 AI 硬性规则：
- 不得把页标题、章节顺序、摘要或 importance 当作项目事实证据。
- 解释 Human Manual 骨架时，必须明确说它只是阅读路线/显著性信号。
- 能力、安装、兼容性、运行状态和风险判断必须引用 repo evidence、source path 或 Claim Graph。

- **项目概述与安装指南**：importance `high`
  - source_paths: README.md, scripts/install.sh, scripts/install.ps1, scripts/install.py, .env.example
- **命令参考与功能详解**：importance `high`
  - source_paths: src/langsmith_cli/main.py, src/langsmith_cli/commands/runs/_group.py, src/langsmith_cli/commands/runs/list_cmd.py, src/langsmith_cli/commands/runs/get_cmd.py, src/langsmith_cli/commands/runs/search_cmd.py
- **系统架构与实现细节**：importance `high`
  - source_paths: src/langsmith_cli/main.py, src/langsmith_cli/cache.py, src/langsmith_cli/field_analysis.py, src/langsmith_cli/filtering.py, src/langsmith_cli/filters.py
- **AI Agent 集成、Skill 与发布运维**：importance `high`
  - source_paths: skills/langsmith/SKILL.md, skills/langsmith/docs/reference.md, skills/langsmith/docs/examples.md, skills/langsmith/references/runs.md, skills/langsmith/references/datasets.md

## Repo Inspection Evidence / 源码检查证据

- repo_clone_verified: true
- repo_inspection_verified: true
- repo_commit: `f2ab13705158fefa8a531485e176c05bc9cc9274`
- inspected_files: `README.md`, `pyproject.toml`, `uv.lock`, `docs/COMMANDS_DESIGN.md`, `docs/PIPES_TO_CLI_REFERENCE.md`, `docs/PRD.md`, `docs/QOL_FEATURES.md`, `docs/QOL_IMPROVEMENTS.md`, `docs/TLDR.md`, `docs/dev/CI_BEST_PRACTICES.md`, `docs/dev/CODECOV_SETUP.md`, `docs/dev/IMPLEMENTATION_PLAN.md`, `docs/dev/LANGSMITH_TEAM_QUESTIONS.md`, `docs/dev/MCP_PARITY.md`, `docs/dev/PUBLISHING.md`, `docs/dev/PYPI_SETUP_SUMMARY.md`, `docs/dev/SESSION_DIRECTIVES.md`, `docs/dev/TESTING_PERFORMANCE.md`, `docs/dev/TESTING_STRATEGY.md`, `docs/dev/TYPE_SAFETY_GUIDE.md`

宿主 AI 硬性规则：
- 没有 repo_clone_verified=true 时，不得声称已经读过源码。
- 没有 repo_inspection_verified=true 时，不得把 README/docs/package 文件判断写成事实。
- 没有 quick_start_verified=true 时，不得声称 Quick Start 已跑通。

## Doramagic Pitfall Constraints / 踩坑约束

这些规则来自 Doramagic 发现、验证或编译过程中的项目专属坑点。宿主 AI 必须把它们当作工作约束，而不是普通说明文字。

### Constraint 1: 可能修改宿主 AI 配置

- Trigger: 项目面向 Claude/Cursor/Codex/Gemini/OpenCode 等宿主，或安装命令涉及用户配置目录。
- Host AI rule: 列出会写入的配置文件、目录和卸载/回滚步骤。
- Why it matters: 安装可能改变本机 AI 工具行为，用户需要知道写入位置和回滚方法。
- Evidence: capability.host_targets | https://github.com/gigaverse-app/langsmith-cli | host_targets=mcp_host, claude_code, claude
- Hard boundary: 不要把这个坑点包装成已解决、已验证或可忽略，除非后续验证证据明确证明它已经关闭。

### Constraint 2: 能力判断依赖假设

- Trigger: README/documentation is current enough for a first validation pass.
- Host AI rule: 将假设转成下游验证清单。
- Why it matters: 假设不成立时，用户拿不到承诺的能力。
- Evidence: capability.assumptions | https://github.com/gigaverse-app/langsmith-cli | README/documentation is current enough for a first validation pass.
- Hard boundary: 不要把这个坑点包装成已解决、已验证或可忽略，除非后续验证证据明确证明它已经关闭。

### Constraint 3: 维护活跃度未知

- Trigger: 未记录 last_activity_observed。
- Host AI rule: 补 GitHub 最近 commit、release、issue/PR 响应信号。
- Why it matters: 新项目、停更项目和活跃项目会被混在一起，推荐信任度下降。
- Evidence: evidence.maintainer_signals | https://github.com/gigaverse-app/langsmith-cli | last_activity_observed missing
- Hard boundary: 不要把这个坑点包装成已解决、已验证或可忽略，除非后续验证证据明确证明它已经关闭。

- Trigger: no_demo
- Evidence: downstream_validation.risk_items | https://github.com/gigaverse-app/langsmith-cli | no_demo; severity=medium
- Hard boundary: 不要把这个坑点包装成已解决、已验证或可忽略，除非后续验证证据明确证明它已经关闭。

### Constraint 5: 存在评分风险

- Trigger: no_demo
- Why it matters: 风险会影响是否适合普通用户安装。
- Evidence: risks.scoring_risks | https://github.com/gigaverse-app/langsmith-cli | no_demo; severity=medium
- Hard boundary: 不要把这个坑点包装成已解决、已验证或可忽略，除非后续验证证据明确证明它已经关闭。

### Constraint 6: issue/PR 响应质量未知

- Trigger: issue_or_pr_quality=unknown。
- Host AI rule: 抽样最近 issue/PR，判断是否长期无人处理。
- Why it matters: 用户无法判断遇到问题后是否有人维护。
- Evidence: evidence.maintainer_signals | https://github.com/gigaverse-app/langsmith-cli | issue_or_pr_quality=unknown
- Hard boundary: 不要把这个坑点包装成已解决、已验证或可忽略，除非后续验证证据明确证明它已经关闭。

### Constraint 7: 发布节奏不明确

- Trigger: release_recency=unknown。
- Host AI rule: 确认最近 release/tag 和 README 安装命令是否一致。
- Why it matters: 安装命令和文档可能落后于代码，用户踩坑概率升高。
- Evidence: evidence.maintainer_signals | https://github.com/gigaverse-app/langsmith-cli | release_recency=unknown
- Hard boundary: 不要把这个坑点包装成已解决、已验证或可忽略，除非后续验证证据明确证明它已经关闭。
