# web-search-plus-mcp - Doramagic AI Context Pack

> 定位：安装前体验与判断资产。它帮助宿主 AI 有一个好的开始，但不代表已经安装、执行或验证目标项目。

## 充分原则

- **充分原则，不是压缩原则**：AI Context Pack 应该充分到让宿主 AI 在开工前理解项目价值、能力边界、使用入口、风险和证据来源；它可以分层组织，但不以最短摘要为目标。
- **压缩策略**：只压缩噪声和重复内容，不压缩会影响判断和开工质量的上下文。

## 给宿主 AI 的使用方式

你正在读取 Doramagic 为 web-search-plus-mcp 编译的 AI Context Pack。请把它当作开工前上下文：帮助用户理解适合谁、能做什么、如何开始、哪些必须安装后验证、风险在哪里。不要声称你已经安装、运行或执行了目标项目。

## Claim 消费规则

- **事实来源**：Repo Evidence + Claim/Evidence Graph；Human Wiki 只提供显著性、术语和叙事结构。
- **事实最低状态**：`supported`
- `supported`：可以作为项目事实使用，但回答中必须引用 claim_id 和证据路径。
- `weak`：只能作为低置信度线索，必须要求用户继续核实。
- `inferred`：只能用于风险提示或待确认问题，不能包装成项目事实。
- `unverified`：不得作为事实使用，应明确说证据不足。
- `contradicted`：必须展示冲突来源，不得替用户强行选择一个版本。

## 它最适合谁

- **正在使用 Claude/Codex/Cursor/Gemini 等宿主 AI 的开发者**：README 或插件配置提到多个宿主 AI。 证据：`README.md` Claim：`clm_0002` supported 0.86

## 它能做什么

- **命令行启动或安装流程**（需要安装后验证）：项目文档中存在可执行命令，真实使用需要在本地或宿主环境中运行这些命令。 证据：`README.md` Claim：`clm_0001` supported 0.86

## 怎么开始

- `pip install web-search-plus-mcp` 证据：`README.md` Claim：`clm_0003` supported 0.86

## 继续前判断卡

- **当前建议**：需要管理员/安全审批
- **为什么**：继续前可能涉及密钥、账号、外部服务或敏感上下文，建议先经过管理员或安全审批。

### 30 秒判断

- **现在怎么做**：需要管理员/安全审批
- **最小安全下一步**：先跑 Prompt Preview；若涉及凭证或企业环境，先审批再试装
- **先别相信**：工具权限边界不能在安装前相信。
- **继续会触碰**：命令执行、本地环境或项目文件、环境变量 / API Key

### 现在可以相信

- **适合人群线索：正在使用 Claude/Codex/Cursor/Gemini 等宿主 AI 的开发者**（supported）：有 supported claim 或项目证据支撑，但仍不等于真实安装效果。 证据：`README.md` Claim：`clm_0002` supported 0.86
- **能力存在：命令行启动或安装流程**（supported）：可以相信项目包含这类能力线索；是否适合你的具体任务仍要试用或安装后验证。 证据：`README.md` Claim：`clm_0001` supported 0.86
- **存在 Quick Start / 安装命令线索**（supported）：可以相信项目文档出现过启动或安装入口；不要因此直接在主力环境运行。 证据：`README.md` Claim：`clm_0003` supported 0.86

### 现在还不能相信

- **工具权限边界不能在安装前相信。**（unverified）：MCP/tool 类项目通常会触碰文件、网络、浏览器或外部 API，必须真实检查权限和日志。
- **真实输出质量不能在安装前相信。**（unverified）：Prompt Preview 只能展示引导方式，不能证明真实项目中的结果质量。
- **宿主 AI 版本兼容性不能在安装前相信。**（unverified）：Claude、Cursor、Codex、Gemini 等宿主加载规则和版本差异必须在真实环境验证。
- **不会污染现有宿主 AI 行为，不能直接相信。**（inferred）：Skill、plugin、AGENTS/CLAUDE/GEMINI 指令可能改变宿主 AI 的默认行为。
- **可安全回滚不能默认相信。**（unverified）：除非项目明确提供卸载和恢复说明，否则必须先在隔离环境验证。
- **真实安装后是否与用户当前宿主 AI 版本兼容？**（unverified）：兼容性只能通过实际宿主环境验证。
- **项目输出质量是否满足用户具体任务？**（unverified）：安装前预览只能展示流程和边界，不能替代真实评测。
- **安装命令是否需要网络、权限或全局写入？**（unverified）：这影响企业环境和个人环境的安装风险。 证据：`README.md`

### 继续会触碰什么

- **命令执行**：包管理器、网络下载、本地插件目录、项目配置或用户主目录。 原因：运行第一条命令就可能产生环境改动；必须先判断是否值得跑。 证据：`README.md`
- **本地环境或项目文件**：安装结果、插件缓存、项目配置或本地依赖目录。 原因：安装前无法证明写入范围和回滚方式，需要隔离验证。 证据：`README.md`
- **环境变量 / API Key**：项目入口文档明确出现 API key、token、secret 或账号凭证配置。 原因：如果真实安装需要凭证，应先使用测试凭证并经过权限/合规判断。 证据：`CHANGELOG.md`, `README.md`
- **宿主 AI 上下文**：AI Context Pack、Prompt Preview、Skill 路由、风险规则和项目事实。 原因：导入上下文会影响宿主 AI 后续判断，必须避免把未验证项包装成事实。

### 最小安全下一步

- **先跑 Prompt Preview**：用安装前交互式试用判断工作方式是否匹配，不需要授权或改环境。（适用：任何项目都适用，尤其是输出质量未知时。）
- **只在隔离目录或测试账号试装**：避免安装命令污染主力宿主 AI、真实项目或用户主目录。（适用：存在命令执行、插件配置或本地写入线索时。）
- **不要使用真实生产凭证**：环境变量/API key 一旦进入宿主或工具链，可能产生账号和合规风险。（适用：出现 API、TOKEN、KEY、SECRET 等环境线索时。）
- **安装后只验证一个最小任务**：先验证加载、兼容、输出质量和回滚，再决定是否深用。（适用：准备从试用进入真实工作流时。）

### 退出方式

- **保留安装前状态**：记录原始宿主配置和项目状态，后续才能判断是否可恢复。
- **记录安装命令和写入路径**：没有明确卸载说明时，至少要知道哪些目录或配置需要手动清理。
- **准备撤销测试 API key 或 token**：测试凭证泄露或误用时，可以快速止损。
- **如果没有回滚路径，不进入主力环境**：不可回滚是继续前阻断项，不应靠信任或运气继续。

## 哪些只能预览

- 解释项目适合谁和能做什么
- 基于项目文档演示典型对话流程
- 帮助用户判断是否值得安装或继续研究

## 哪些必须安装后验证

- 真实安装 Skill、插件或 CLI
- 执行脚本、修改本地文件或访问外部服务
- 验证真实输出质量、性能和兼容性

## 边界与风险判断卡

- **把安装前预览误认为真实运行**：用户可能高估项目已经完成的配置、权限和兼容性验证。 处理方式：明确区分 prompt_preview_can_do 与 runtime_required。 Claim：`clm_0004` inferred 0.45
- **命令执行会修改本地环境**：安装命令可能写入用户主目录、宿主插件目录或项目配置。 处理方式：先在隔离环境或测试账号中运行。 证据：`README.md` Claim：`clm_0005` supported 0.86
- **待确认**：真实安装后是否与用户当前宿主 AI 版本兼容？。原因：兼容性只能通过实际宿主环境验证。
- **待确认**：项目输出质量是否满足用户具体任务？。原因：安装前预览只能展示流程和边界，不能替代真实评测。
- **待确认**：安装命令是否需要网络、权限或全局写入？。原因：这影响企业环境和个人环境的安装风险。

## 开工前工作上下文

### 加载顺序

- 先读取 how_to_use.host_ai_instruction，建立安装前判断资产的边界。
- 读取 claim_graph_summary，确认事实来自 Claim/Evidence Graph，而不是 Human Wiki 叙事。
- 再读取 intended_users、capabilities 和 quick_start_candidates，判断用户是否匹配。
- 需要执行具体任务时，优先查 role_skill_index，再查 evidence_index。
- 遇到真实安装、文件修改、网络访问、性能或兼容性问题时，转入 risk_card 和 boundaries.runtime_required。

### 任务路由

- **命令行启动或安装流程**：先说明这是安装后验证能力，再给出安装前检查清单。 边界：必须真实安装或运行后验证。 证据：`README.md` Claim：`clm_0001` supported 0.86

### 上下文规模

- 文件总数：52
- 重要文件覆盖：40/52
- 证据索引条目：51
- 角色 / Skill 条目：3

### 证据不足时的处理

- **missing_evidence**：说明证据不足，要求用户提供目标文件、README 段落或安装后验证记录；不要补全事实。
- **out_of_scope_request**：说明该任务超出当前 AI Context Pack 证据范围，并建议用户先查看 Human Manual 或真实安装后验证。
- **runtime_request**：给出安装前检查清单和命令来源，但不要替用户执行命令或声称已执行。
- **source_conflict**：同时展示冲突来源，标记为待核实，不要强行选择一个版本。

## Prompt Recipes

### 适配判断

- 目标：判断这个项目是否适合用户当前任务。
- 预期输出：适配结论、关键理由、证据引用、安装前可预览内容、必须安装后验证内容、下一步建议。

```text
请基于 web-search-plus-mcp 的 AI Context Pack，先问我 3 个必要问题，然后判断它是否适合我的任务。回答必须包含：适合谁、能做什么、不能做什么、是否值得安装、证据来自哪里。所有项目事实必须引用 evidence_refs、source_paths 或 claim_id。
```

### 安装前体验

- 目标：让用户在安装前感受核心工作流，同时避免把预览包装成真实能力或营销承诺。
- 预期输出：一段带边界标签的体验剧本、安装后验证清单和谨慎建议；不含真实运行承诺或强营销表述。

```text
请把 web-search-plus-mcp 当作安装前体验资产，而不是已安装工具或真实运行环境。

请严格输出四段：
1. 先问我 3 个必要问题。
2. 给出一段“体验剧本”：用 [安装前可预览]、[必须安装后验证]、[证据不足] 三种标签展示它可能如何引导工作流。
3. 给出安装后验证清单：列出哪些能力只有真实安装、真实宿主加载、真实项目运行后才能确认。
4. 给出谨慎建议：只能说“值得继续研究/试装”“先补充信息后再判断”或“不建议继续”，不得替项目背书。

硬性边界：
- 不要声称已经安装、运行、执行测试、修改文件或产生真实结果。
- 不要写“自动适配”“确保通过”“完美适配”“强烈建议安装”等承诺性表达。
- 如果描述安装后的工作方式，必须使用“如果安装成功且宿主正确加载 Skill，它可能会……”这种条件句。
- 体验剧本只能写成“示例台词/假设流程”：使用“可能会询问/可能会建议/可能会展示”，不要写“已写入、已生成、已通过、正在运行、正在生成”。
- Prompt Preview 不负责给安装命令；如用户准备试装，只能提示先阅读 Quick Start 和 Risk Card，并在隔离环境验证。
- 所有项目事实必须来自 supported claim、evidence_refs 或 source_paths；inferred/unverified 只能作风险或待确认项。

```

### 角色 / Skill 选择

- 目标：从项目里的角色或 Skill 中挑选最匹配的资产。
- 预期输出：候选角色或 Skill 列表，每项包含适用场景、证据路径、风险边界和是否需要安装后验证。

```text
请读取 role_skill_index，根据我的目标任务推荐 3-5 个最相关的角色或 Skill。每个推荐都要说明适用场景、可能输出、风险边界和 evidence_refs。
```

### 风险预检

- 目标：安装或引入前识别环境、权限、规则冲突和质量风险。
- 预期输出：环境、权限、依赖、许可、宿主冲突、质量风险和未知项的检查清单。

```text
请基于 risk_card、boundaries 和 quick_start_candidates，给我一份安装前风险预检清单。不要替我执行命令，只说明我应该检查什么、为什么检查、失败会有什么影响。
```

### 宿主 AI 开工指令

- 目标：把项目上下文转成一次对话开始前的宿主 AI 指令。
- 预期输出：一段边界明确、证据引用明确、适合复制给宿主 AI 的开工前指令。

```text
请基于 web-search-plus-mcp 的 AI Context Pack，生成一段我可以粘贴给宿主 AI 的开工前指令。这段指令必须遵守 not_runtime=true，不能声称项目已经安装、运行或产生真实结果。
```

## 角色 / Skill 索引

- 共索引 3 个角色 / Skill / 项目文档条目。

- **🔍 web-search-plus-mcp**（project_doc）：! PyPI version https://img.shields.io/pypi/v/web-search-plus-mcp.svg https://pypi.org/project/web-search-plus-mcp/ ! Python 3.10+ https://img.shields.io/badge/python-3.10%2B-blue.svg https://www.python.org/downloads/ ! MCP https://img.shields.io/badge/MCP-compatible-green.svg https://modelcontextprotocol.io/ ! CI https://github.com/robbyczgw-cla/web-search-plus-mcp/actions/workflows/ci.yml/badge.svg https://github.c… 激活提示：当用户需要理解项目结构、安装方式或边界时参考。 证据：`README.md`
- **Migrating web-search-plus-mcp 0.x to 1.0**（project_doc）：Migrating web-search-plus-mcp 0.x to 1.0 激活提示：当用户需要理解项目结构、安装方式或边界时参考。 证据：`docs/MIGRATION_1_0.md`
- **Changelog**（project_doc）：Added - Sync the portable Web Search Plus v3.1.1 policy layer: budget preflight, diversity scoring/reranking, self-hosted profiles, shadow-policy observations, semantic extraction spans, extraction cache identity v6, and SQLite state schema v3. - Add the public wsp sdk package plus fail-closed providers.d discovery, startup diagnostics, and network-free provider conformance checks. - Expose deterministic semantic sp… 激活提示：当用户需要理解项目结构、安装方式或边界时参考。 证据：`CHANGELOG.md`

## 证据索引

- 共索引 51 条证据。

- **🔍 web-search-plus-mcp**（documentation）：! PyPI version https://img.shields.io/pypi/v/web-search-plus-mcp.svg https://pypi.org/project/web-search-plus-mcp/ ! Python 3.10+ https://img.shields.io/badge/python-3.10%2B-blue.svg https://www.python.org/downloads/ ! MCP https://img.shields.io/badge/MCP-compatible-green.svg https://modelcontextprotocol.io/ ! CI https://github.com/robbyczgw-cla/web-search-plus-mcp/actions/workflows/ci.yml/badge.svg https://github.com/robbyczgw-cla/web-search-plus-mcp/actions/workflows/ci.yml ! License: MIT https://img.shields.io/badge/License-MIT-yellow.svg ./LICENSE ! Glama https://glama.ai/mcp/servers/robbyczgw-cla/web-search-plus-mcp/badge https://glama.ai/mcp/servers/robbyczgw-cla/web-search-plus-mcp 证据：`README.md`
- **License**（source_file）：Permission is hereby granted, free of charge, to any person obtaining a copy of this software and associated documentation files the "Software" , to deal in the Software without restriction, including without limitation the rights to use, copy, modify, merge, publish, distribute, sublicense, and/or sell copies of the Software, and to permit persons to whom the Software is furnished to do so, subject to the following conditions: 证据：`LICENSE`
- **Migrating web-search-plus-mcp 0.x to 1.0**（documentation）：Migrating web-search-plus-mcp 0.x to 1.0 证据：`docs/MIGRATION_1_0.md`
- **Request.Schema**（structured_config）：{ "$schema": "https://json-schema.org/draft/2020-12/schema", "$id": "https://websearchplus.xyz/schema/v3/request.schema.json", "title": "RequestV3", "type": "object", "additionalProperties": false, "required": "contract version", "capability", "input" , "properties": { "contract version": { "const": "3.0" }, "request id": { "type": "string", "minLength": 1 }, "capability": { "$ref": " /$defs/Capability" }, "input": { "type": "object" }, "options": { "type": "object" }, "cache": { "$ref": " /$defs/CacheRequest" }, "routing": { "$ref": " /$defs/RoutingRequest" }, "budget": { "$ref": " /$defs/BudgetRequest" }, "client": { "$ref": " /$defs/ClientNegotiation" } }, "allOf": { "if": { "properties"… 证据：`schemas/v3/request.schema.json`
- **Pyproject**（source_file）：build-system requires = "hatchling" build-backend = "hatchling.build" 证据：`pyproject.toml`
- **Attempt Engine V3**（source_file）：@dataclass frozen=True class AttemptContext ⋮---- provider: str capability: Capability endpoint: str credential fingerprint: str budget scope: str budget window: str budget units: int = 1 budget limit units: int = 3 daily budget scope: str None = None daily budget window: str None = None daily budget limit units: int None = None deadline monotonic: float None = None ⋮---- def post init self - None ⋮---- daily values = ⋮---- @property def circuit key self - CircuitKey ⋮---- @dataclass frozen=True class AttemptExecution ⋮---- payload: Optional Dict receipt: ProviderAttemptV3 ⋮---- class AttemptEngine ⋮---- @staticmethod def attempt id context: AttemptContext, started: int - str ⋮---- @staticm… 证据：`web_search_plus_mcp/attempt_engine_v3.py`
- **A concurrent writer of the same immutable content won.**（source_file）：DEFAULT MAX URLS = 10 HARD MAX URLS = 50 DEFAULT MAX CONTEXT CHARS = 60 000 MIN CONTEXT CHARS = 1 000 MAX CONTEXT CHARS = 200 000 DEFAULT FULL TEXT TTL SECONDS = 604 800 DEFAULT FULL TEXT MAX BYTES = 268 435 456 STORE NAME = "web text v3" MEDIA TYPE = "text/markdown" OWNED MARKER = " Dict str, Any : ... ⋮---- def integer option value: Any, name: str, default: int - int ⋮---- policy = config.get "bounded context" or {} ⋮---- policy = {} ⋮---- requested max urls = integer option operator max urls = integer option max urls = min ⋮---- operator default chars = integer option requested chars = integer option max context chars = min ⋮---- urls = list request.input "urls" processed urls = urls :ma… 证据：`web_search_plus_mcp/bounded_context_v3.py`
- **Budget Preflight V3**（source_file）：CHECK NAMES = ACTIONS = {"proceed", "degrade", "abort"} ON EXCEED = {"degrade", "abort"} ABORT REASONS = { ⋮---- @dataclass frozen=True class PreflightCheck ⋮---- check: str limit: int None observed: int None verdict: Literal "ok", "exceeded" ⋮---- def post init self - None ⋮---- def to dict self - Dict str, Any ⋮---- @dataclass frozen=True class PreflightDecision ⋮---- action: Literal "proceed", "degrade", "abort" checks: tuple PreflightCheck, ... = adjustments: Dict str, int = field default factory=dict reason: str None = None ⋮---- payload: Dict str, Any = { ⋮---- def section config: Mapping str, Any - Mapping str, Any ⋮---- section = config.get "budget preflight" or {} ⋮---- def limit s… 证据：`web_search_plus_mcp/budget_preflight_v3.py`
- **Cache Identity V3**（source_file）：EXTRACTION CACHE IDENTITY VERSION = 6 ⋮---- def canonical value value: Any - Any ⋮---- normalized = {} ⋮---- def canonical json bytes value: Mapping str, Any - bytes ⋮---- """Serialize an identity in the one form used for its SHA-256 key.""" ⋮---- @dataclass frozen=True class ExtractionCacheIdentityV3 ⋮---- requested urls: tuple str, ... attempt budget: Mapping str, Any effective context limits: Mapping str, Any output format: str include images: bool include raw html: bool render js: bool semantic spans: Mapping str, Any provider selection: Mapping str, Any provider endpoint config: Mapping str, Any url policy: Mapping str, Any storage policy: Mapping str, Any identity version: int = EXTRA… 证据：`web_search_plus_mcp/cache_identity_v3.py`
- **Cache V3**（source_file）：CACHE SCHEMA VERSION = 3 CACHE OWNER = "web-search-plus:v3" NORMALIZER VERSION = "runtime-v3-amendment-002" EXTRACTION IDENTITY VARY KEY = "extraction cache identity" ⋮---- @dataclass frozen=True class CacheLookupV3 ⋮---- disposition: str payload: Optional Dict str, Any = None entry id: Optional str = None age seconds: Optional int = None source contract version: Optional str = None legacy payload: Optional Dict str, Any = None ⋮---- attempts = payload.get "provider attempts" or successful = next routing = payload.get "routing receipt" or {} capability = str payload.get "capability" or "" extract legacy = legacy payload or {} if capability == "extract" else {} legacy provider = legacy paylo… 证据：`web_search_plus_mcp/cache_v3.py`
- **Compat V3**（source_file）：capability = Capability capability provider = str payload.get "provider" or "auto" default fallback = capability is Capability.EXTRACT or provider == "auto" routing = { cache = { client = {"accept contract versions": "3.0", "2.x" } ⋮---- query = unicodedata.normalize "NFC", str payload.get "query" or "" .strip mode = str payload.get "mode", "normal" ⋮---- options: Dict str, Any = { ⋮---- value = payload.get key ⋮---- locale = { ⋮---- urls = payload.get "urls" or ⋮---- urls = urls extract options: Dict str, Any = { ⋮---- spans query = payload.get "spans query" ⋮---- spans query = payload.get "query" ⋮---- def project execution: ExecutedV3, capability: Capability - Dict str, Any ⋮---- def v3… 证据：`web_search_plus_mcp/compat_v3.py`
- **Check config.json first**（source_file）：CONFIG ENV VAR = "WEB SEARCH PLUS CONFIG" ⋮---- class SelfHostedProfileError ProviderConfigError ⋮---- error type = "self hosted profile unavailable" ⋮---- SUPPORTED PROFILES = frozenset {"standard", "self hosted"} SELF HOSTED SEARCH PROVIDER IDS = "searxng", KEYLESS PROVIDER IDS SELF HOSTED EXTRACT PROVIDER IDS = tuple KEYLESS EXTRACT PROVIDER IDS ⋮---- def is placeholder env value value: str - bool ⋮---- def clean env value value: str - Optional str ⋮---- def load env file ⋮---- DEFAULT CONFIG = { ⋮---- def deepcopy default config - Dict str, Any ⋮---- ROUTING PROVIDER NAMES = set PROVIDER SPECS VALID PROVIDERS = ROUTING PROVIDER NAMES ⋮---- def normalize routing provider config provider:… 证据：`web_search_plus_mcp/config.py`
- **Contract V3**（source_file）：CONTRACT VERSION = "3.0" ⋮---- class StrEnum str, Enum ⋮---- def str self - str ⋮---- class Capability StrEnum ⋮---- SEARCH = "search" EXTRACT = "extract" ⋮---- class ResponseStatus StrEnum ⋮---- OK = "ok" DEGRADED = "degraded" FAILED = "failed" ⋮---- class DegradedReason StrEnum ⋮---- SERVED STALE = "wsp.cache.served stale" CONTENT TRUNCATED = "wsp.content.truncated" URLS OMITTED = "wsp.extract.urls omitted" PARTIAL EXTRACTION = "wsp.extract.partial" BUDGET LIMITED = "wsp.budget.limited" FINGERPRINTING REDUCED = "wsp.independence.method degraded" ⋮---- class ErrorClass StrEnum ⋮---- INVALID REQUEST = "invalid request" UNSUPPORTED = "unsupported" CONFIG = "config" AUTH = "auth" QUOTA = "quo… 证据：`web_search_plus_mcp/contract_v3.py`
- **Diversity V3**（source_file）：MULTI LABEL SUFFIXES = frozenset ⋮---- TRACKING PARAMETER NAMES = frozenset ⋮---- DOMAIN DIVERSITY WEIGHT = 0.40 URL DUPLICATION WEIGHT = 0.30 CONTENT DIVERSITY WEIGHT = 0.20 PROVIDER MIX WEIGHT = 0.10 ⋮---- DEFAULT NEAR DUPLICATE THRESHOLD = 0.60 ⋮---- class DiversityComponents TypedDict ⋮---- domain diversity: float url duplication: float content diversity: float provider mix: float ⋮---- class DuplicateCandidate TypedDict ⋮---- kind: str kept: int dropped candidate: int ⋮---- class DominantDomain TypedDict ⋮---- domain: str share: float ⋮---- class DiversityReport TypedDict ⋮---- score: float components: DiversityComponents duplicates: List DuplicateCandidate dominant domain: Optional Do… 证据：`web_search_plus_mcp/diversity_v3.py`
- **Env Loader**（source_file）：def is placeholder env value value: str - bool ⋮---- stripped = value or "" .strip .strip '"' .strip "'" ⋮---- def clean env value value: str - Optional str ⋮---- """Return a real env value, or None for empty/template placeholders.""" ⋮---- TRUTHY VALUES = {"1", "true", "yes", "on"} ⋮---- def is truthy value: object - bool ⋮---- def get hermes env path - Path ⋮---- def candidate env paths anchor file: Union str, Path - List Path ⋮---- plugin dir = Path anchor file .resolve .parent paths = deduped: List Path = seen: Set str = set ⋮---- key = str path.expanduser ⋮---- def load env files anchor file: Union str, Path , environ: Optional MutableMapping str, str = None - List Path ⋮---- target en… 证据：`web_search_plus_mcp/env_loader.py`
- **Extract**（source_file）：EXTRACT PROVIDER PRIORITY = list EXTRACT PROVIDER IDS ⋮---- def daily preflight budget config: Dict str, Any - Dict str, Any ⋮---- raw off = os.environ.get "WSP BUDGET PREFLIGHT OFF" ⋮---- section = config.get "budget preflight" or {} limit = section.get "max daily provider calls" if isinstance section, dict else None ⋮---- def resolve extract provider priority config: Optional Dict str, Any = None - List str ⋮---- auto config = config.get "auto routing", {} if isinstance config, dict else {} ⋮---- auto config = {} raw priority = auto config.get "extract provider priority" ⋮---- raw values = raw priority.split "," ⋮---- raw values = raw priority ⋮---- raw values = ⋮---- providers: List str… 证据：`web_search_plus_mcp/extract.py`
- **Orchestrator V3**（source_file）：PIPELINE STAGES: Tuple str, ... = ⋮---- @dataclass frozen=True class ProviderPlan ⋮---- candidate order: Tuple str, ... selected provider: str routing metadata: Dict str, Any = field default factory=dict mode: str = "classic" execution id: str = field ⋮---- def post init self - None ⋮---- PlanFn = Callable RequestV3, Dict str, Any , ProviderPlan ⋮---- @dataclass frozen=True class CapabilityExecution ⋮---- payload: Dict str, Any provider attempts: Tuple Any, ... = stages: Tuple str, ... = "provider attempt", ⋮---- unknown = set self.stages - set PIPELINE STAGES ⋮---- positions = PIPELINE STAGES.index stage for stage in self.stages ⋮---- ExecuteFn = Callable NormalizeFn = Callable RequestV3,… 证据：`web_search_plus_mcp/orchestrator_v3.py`
- **Provider Adapter Protocol**（source_file）：SEARCH ADAPTER PARAMETERS = EXTRACT ADAPTER PARAMETERS = ⋮---- @runtime checkable class SearchAdapter Protocol ⋮---- @runtime checkable class ExtractAdapter Protocol ⋮---- def signature matches adapter: Any, expected: tuple str, ... - bool ⋮---- parameters = tuple inspect.signature adapter .parameters.values ⋮---- errors: list str = contracts = ⋮---- actual providers = set dispatch ⋮---- """Fail module initialization when registry and adapters drift apart.""" ⋮---- errors = dispatch conformance errors ⋮---- def contract failure code: str - ProviderContractFailure ⋮---- """Validate a source-only provider envelope before orchestration consumes it. Errors contain stable codes only. Provider pa… 证据：`web_search_plus_mcp/provider_adapter_protocol.py`
- **Provider Dispatch**（source_file）：def resolve namespace: Any, name: str - Callable ..., Dict str, Any ⋮---- def locale prov: str, args: Any, config: Dict str, Any ⋮---- def call serper search search module, prov, args, key, config, routing info ⋮---- def call serpbase search search module, prov, args, key, config, routing info ⋮---- serpbase config = config.get "serpbase", {} ⋮---- def call brave search search module, prov, args, key, config, routing info ⋮---- brave config = config.get "brave", {} ⋮---- def call tavily search search module, prov, args, key, config, routing info ⋮---- def call linkup search search module, prov, args, key, config, routing info ⋮---- linkup config = config.get "linkup", {} ⋮---- def call quer… 证据：`web_search_plus_mcp/provider_dispatch.py`
- **Provider Registry**（source_file）：PROVIDERS DIRECTORY = Path file .resolve .with name "providers.d" NON PRODUCTION DISCOVERY ENV = "WSP SDK ALLOW NON PRODUCTION" ⋮---- def non production discovery allowed - bool ⋮---- value = os.environ.get NON PRODUCTION DISCOVERY ENV, "" ⋮---- PROVIDER ID PATTERN = re.compile r"^ a-z0-9 + ?:- a-z0-9 + $" ENV VAR PATTERN = re.compile r"^ A-Z A-Z0-9 $" SEARCH PARAMETERS = EXTRACT PARAMETERS = ⋮---- BUILTIN PROVIDER SPECS = ⋮---- BUILTIN EXTRACT PROVIDER IDS = BUILTIN DEFAULT PROVIDER PRIORITY = provider specs = list BUILTIN PROVIDER SPECS discovered provider ids: set str = set ⋮---- PROVIDER SPECS: Dict str, ProviderSpec = {} SEARCH PROVIDER IDS: tuple str, ... = EXTRACT PROVIDER IDS: tuple… 证据：`web_search_plus_mcp/provider_registry.py`
- **Parse results**（source_file）：BATCH TIMEOUT GRACE SECONDS = 5 ⋮---- FRESHNESS VALUES = "day", "week", "month", "year" ⋮---- PROVIDER FRESHNESS FORMATS: Dict str, Dict str, str = { ⋮---- def normalize freshness value: Optional str - Optional str ⋮---- normalized = str value .strip .lower ⋮---- def provider supports freshness provider: str - bool ⋮---- def map freshness for provider provider: str, freshness: Optional str - Optional str ⋮---- def freshness metadata provider: str, requested: str - Dict str, Any ⋮---- native = map freshness for provider provider, requested ⋮---- SEARCH TYPE VALUES = "search", "news" ⋮---- PROVIDER SEARCH TYPES: Dict str, Dict str, str = { ⋮---- def normalize search type value: Optional str -… 证据：`web_search_plus_mcp/providers.py`
- **Request Gate V3**（source_file）：SOURCE ONLY SEMANTICS = frozenset {"source results", "source text"} BANNED BODY KEYS = frozenset BANNED INSTRUCTION FRAGMENTS = ⋮---- def validate provider mode provider: str, capability: str - str ⋮---- spec = PROVIDER SPECS.get provider ⋮---- supported = spec.supports search semantics = spec.search output semantics ⋮---- supported = spec.supports extract semantics = spec.extract output semantics ⋮---- def walk value: Any ⋮---- def validate outbound body provider: str, body: Mapping str, Any - None ⋮---- lowered = key.lower ⋮---- text = value.lower 证据：`web_search_plus_mcp/request_gate_v3.py`
- **Conservative class-aware provider boosts positive and penalties negative**（source_file）：def provider configured provider: str, config: Dict str, Any None = None - bool ⋮---- ROUTING POLICY = "routing-v2" ⋮---- def iter all selectable provider modes - Tuple str, ... ⋮---- ROUTING CLASS RULES: List Tuple str, Tuple str, ... = ⋮---- MULTILINGUAL ROUTING CLASS = "multilingual current" DEFAULT ROUTING CLASS = "general" ⋮---- LANGUAGE HINT PROVIDER BOOSTS: Dict str, List Tuple str, float = { ⋮---- LANGUAGE INFERENCE MIN MATCHES = 2 ⋮---- LANGUAGE INFERENCE STOPWORDS: Dict str, frozenset = { ⋮---- LANGUAGE INFERENCE CHAR HINTS: Dict str, str = { ⋮---- def infer query language query: str - Optional str ⋮---- lowered = query.lower words = set re.findall r"\w+", lowered counts: Dict str… 证据：`web_search_plus_mcp/routing.py`
- **Runtime V3**（source_file）：def canonical url value: str - str ⋮---- parsed = urlsplit value host = parsed.hostname or "" .lower port = parsed.port netloc = host ⋮---- netloc = f"{host}:{port}" path = parsed.path or "/" ⋮---- def stable id prefix: str, parts: object - str ⋮---- raw = "\x1f".join str part for part in parts ⋮---- def valid rfc3339 value: object - str None ⋮---- parsed = datetime.fromisoformat value.replace "Z", "+00:00" ⋮---- def error message: str, provider: str None = None - ErrorV3 ⋮---- lowered = message.lower ⋮---- error class = ErrorClass.CONFIG code = "wsp.config.self hosted profile unavailable" ⋮---- code = "wsp.config.missing credentials" ⋮---- error class = ErrorClass.TIMEOUT code = "wsp.provi… 证据：`web_search_plus_mcp/runtime_v3.py`
- **Server**（source_file）：version = "1.1.0" ⋮---- SEARCH SCRIPT = Path file .parent / "search.py" app = Server "web-search-plus", version= version ⋮---- SEARCH PROVIDERS = { EXTRACT PROVIDERS = list EXTRACT PROVIDER IDS PRESETS = { ⋮---- CONFIG ENV VAR = "WEB SEARCH PLUS CONFIG" PROVIDER ALIASES = {"kilo perplexity": "kilo-perplexity"} RETIRED ANSWER PROVIDERS = {"perplexity", "kilo-perplexity"} ROUTING PROVIDER ORDER = list DEFAULT PROVIDER PRIORITY DEFAULT SEARCH SUBPROCESS TIMEOUT SECONDS = 75 RESEARCH SUBPROCESS GRACE SECONDS = 10 ⋮---- def canonical provider provider: str - str ⋮---- value = provider or "" .strip .lower ⋮---- def research subprocess timeout value: Any - int ⋮---- """Keep the outer process alive… 证据：`web_search_plus_mcp/server.py`
- **Span Extraction V3**（source_file）：Span = Dict str, Union int, float, str Ranker = Callable str, str , float ⋮---- TOKEN RE = re.compile r" ^\W + ?: '\N{RIGHT SINGLE QUOTATION MARK} ^\W + ?", re.UNICODE PARAGRAPH BREAK RE = re.compile r" ?:\r?\n \t \f\v {2,}" SENTENCE END RE = re.compile r" ? str ⋮---- """Return the canonical NFC string used for all span offsets.""" ⋮---- def tokens text: str - List str ⋮---- def trimmed candidate text: str, start: int, end: int - Candidate None ⋮---- def split long segment text: str, start: int, end: int, limit: int - List Candidate ⋮---- pieces: List Candidate = cursor = start ⋮---- boundary = min end, cursor + limit ⋮---- whitespace = text.rfind " ", cursor + max 1, limit // 2 , boundary… 证据：`web_search_plus_mcp/span_extraction_v3.py`
- **State Store V3**（source_file）：SCHEMA VERSION = 3 SHADOW EVALUATION RETENTION SECONDS = 30 24 60 60 SHADOW EVALUATION MAX ROWS = 10 000 DEFAULT OPEN SECONDS = { ⋮---- material = secret if secret else " " ⋮---- @dataclass frozen=True class CircuitKey ⋮---- provider: str capability: Capability endpoint: str credential fingerprint: str ⋮---- def values self - tuple str, str, str, str ⋮---- @dataclass frozen=True class CircuitRecord ⋮---- state: CircuitState = CircuitState.CLOSED failure count: int = 0 open until: Optional int = None updated at: int = 0 ⋮---- @dataclass frozen=True class AdmissionDecision ⋮---- allowed: bool circuit state: CircuitState skip reason: Optional SkipReason = None store available: bool = True bloc… 证据：`web_search_plus_mcp/state_store_v3.py`
- **Changelog**（documentation）：Added - Sync the portable Web Search Plus v3.1.1 policy layer: budget preflight, diversity scoring/reranking, self-hosted profiles, shadow-policy observations, semantic extraction spans, extraction cache identity v6, and SQLite state schema v3. - Add the public wsp sdk package plus fail-closed providers.d discovery, startup diagnostics, and network-free provider conformance checks. - Expose deterministic semantic spans through the MCP web extract schema and CLI projection. 证据：`CHANGELOG.md`
- **Glama**（structured_config）：{ "$schema": "https://glama.ai/mcp/schemas/server.json", "maintainers": "robbyczgw-cla" , "name": "web-search-plus-mcp", "description": "Source-only multi-provider web search and bounded extraction with registry-backed routing, quality reports, and operator onboarding.", "repository": "https://github.com/robbyczgw-cla/web-search-plus-mcp", "homepage": "https://github.com/robbyczgw-cla/web-search-plus-mcp", "license": "MIT", "runtime": "python", "transport": "stdio" , "install": { "uvx": "uvx web-search-plus-mcp", "pip": "pip install web-search-plus-mcp" }, "mcp": { "command": "uvx", "args": "web-search-plus-mcp" }, "tools": { "name": "web search", "description": "Source-only web search thro… 证据：`glama.json`
- **Response.Schema**（structured_config）：{ "$schema": "https://json-schema.org/draft/2020-12/schema", "$id": "https://websearchplus.xyz/schema/v3/response.schema.json", "title": "ResponseV3", "type": "object", "additionalProperties": false, "required": "contract version", "request id", "execution id", "capability", "status", "results", "observations", "policy actions", "source diversity", "provider attempts", "routing receipt", "cache status", "limits applied", "dedup clusters", "warnings" , "properties": { "contract version": { "const": "3.0" }, "request id": { "type": "string", "minLength": 1 }, "execution id": { "type": "string", "minLength": 1 }, "capability": { "$ref": " /$defs/Capability" }, "status": { "$ref": " /$defs/Resp… 证据：`schemas/v3/response.schema.json`
- **Python**（source_file）：Python pycache / .py cod $py.class .so .egg .egg-info/ dist/ build/ eggs/ parts/ var/ sdist/ develop-eggs/ .installed.cfg lib/ lib64/ wheels/ share/python-wheels/ .whl 证据：`.gitignore`
- **Environment variables optional, for tool inspection**（source_file）：LABEL org.opencontainers.image.title="web-search-plus-mcp" LABEL org.opencontainers.image.description="Multi-provider web search MCP server" LABEL org.opencontainers.image.source="https://github.com/robbyczgw-cla/web-search-plus-mcp" 证据：`Dockerfile`
- **Gen Contract V3 Schemas**（source_file）：ROOT = Path file .resolve .parents 1 ⋮---- OUT = ROOT / "schemas" / "v3" ⋮---- def enum schema enum cls ⋮---- def obj properties, required= , , additional=False ⋮---- value = { ⋮---- error = obj ⋮---- attempt = obj ⋮---- provider try error = obj provider try = obj ⋮---- request defs = { ⋮---- request schema = { ⋮---- response defs = { ⋮---- RECEIPT COMPLETION FIELDS = ⋮---- response schema = { ⋮---- parser = argparse.ArgumentParser ⋮---- args = parser.parse args ⋮---- stale = ⋮---- path = OUT / name rendered = json.dumps schema, indent=2, ensure ascii=False + "\n" 证据：`scripts/gen_contract_v3_schemas.py`
- **Init**（source_file）：PACKAGE DIR = str Path file .resolve .parent ⋮---- version = "1.1.0" 证据：`web_search_plus_mcp/__init__.py`
- **Cache**（source_file）：CACHE DIR = Path os.environ.get "WSP CACHE DIR", os.path.join os.path.dirname os.path.dirname os.path.abspath file , ".cache" DEFAULT CACHE TTL = 3600 PROVIDER HEALTH FILENAME = "provider health.json" ⋮---- WEB TEXT CACHE DIRNAME = "web" MAX STORED TEXT CHARS = 2 000 000 ⋮---- def atomic write text path: Path, text: str - None ⋮---- def get web text cache path url: str - Path ⋮---- key = hashlib.sha256 url.encode "utf-8" .hexdigest ⋮---- def iter web text cache files ⋮---- web dir = CACHE DIR / WEB TEXT CACHE DIRNAME ⋮---- def iter web text temp files ⋮---- def web text cache stats - Dict str, Any ⋮---- entries = total size = 0 ⋮---- def store web text url: str, text: str, max chars: int =… 证据：`web_search_plus_mcp/cache.py`
- **Daemon Tasks**（source_file）：class DaemonTask ⋮---- def init self, fn: Callable ..., Any , args: Any, kwargs: Any ⋮---- def run self, fn: Callable ..., Any , args: tuple, kwargs: dict - None ⋮---- def done self - bool ⋮---- def result self, timeout: Optional float = None - Any 证据：`web_search_plus_mcp/daemon_tasks.py`
- **Errors V3**（source_file）：MESSAGES = { ⋮---- CODES = { ⋮---- def classify provider error error: BaseException, , provider: str - ErrorV3 ⋮---- status = getattr error, "status code", None retry after = getattr error, "retry after", None class name = type error . name ⋮---- error class = ErrorClass.CONFIG ⋮---- error class = ErrorClass.TIMEOUT ⋮---- error class = ErrorClass.AUTH ⋮---- error class = ErrorClass.QUOTA ⋮---- error class = ErrorClass.RATE LIMIT ⋮---- error class = ErrorClass.TRANSIENT ⋮---- error class = ErrorClass.PROVIDER CONTRACT ⋮---- error class = ErrorClass.INTERNAL ⋮---- retryable = error class in { 证据：`web_search_plus_mcp/errors_v3.py`
- **Ensure User-Agent is set required by some APIs like Exa/Cloudflare**（source_file）：TRANSIENT HTTP CODES = {429, 503} ⋮---- version = "1.1.0" ⋮---- DEFAULT USER AGENT = f"ClawdBot-WebSearchPlus-MCP/{ version }" ⋮---- class ProviderRequestError Exception ⋮---- """Structured provider error with retry/cooldown metadata.""" ⋮---- def response header response, name: str - str ⋮---- """Return an HTTP response header from urllib response/error objects.""" ⋮---- value = response.getheader name ⋮---- headers = getattr response, "headers", None ⋮---- value = headers.get name ⋮---- value = None ⋮---- def read response body response - bytes ⋮---- """Read an urllib response body and decode supported Content-Encoding values.""" raw = response.read encoding = response header response, "C… 证据：`web_search_plus_mcp/http_client.py`
- **Independence V3**（source_file）：TRACKING KEYS = {"fbclid", "gclid", "dclid", "msclkid"} TOKEN RE = re.compile r" \w +", re.UNICODE MINHASH PERMUTATIONS = 32 MINHASH THRESHOLD = 0.65 ⋮---- def canonicalize url value: str - str ⋮---- parsed = urlsplit value scheme = parsed.scheme.lower host = parsed.hostname or "" .lower port = parsed.port netloc = host ⋮---- netloc = f"{host}:{port}" path = parsed.path or "/" ⋮---- path = path.rstrip "/" or "/" query items = query = urlencode sorted query items ⋮---- def tokens result: Dict str, Any - List str ⋮---- text = " ".join normalized = unicodedata.normalize "NFKC", text .casefold ⋮---- def shingles result: Dict str, Any - set str ⋮---- tokens = tokens result ⋮---- def minhash shin… 证据：`web_search_plus_mcp/independence_v3.py`
- **Provider Health**（source_file）：PROVIDER HEALTH FILE = CACHE DIR / "provider health.json" COOLDOWN STEPS SECONDS = 60, 300, 1500, 3600 RETRY BACKOFF SECONDS = 1, 3, 9 ⋮---- RETRY JITTER FRACTION = 0.5 ⋮---- FAILURE DECAY SECONDS = 1800 ⋮---- RATE LIMIT MAX ATTEMPTS = 2 ⋮---- MAX RETRY AFTER WAIT SECONDS = 30.0 ⋮---- HEALTH LOCK = threading.Lock ⋮---- def retry delay attempt: int - float ⋮---- base = RETRY BACKOFF SECONDS min attempt, len RETRY BACKOFF SECONDS - 1 ⋮---- def ensure parent path: Path - None ⋮---- def load provider health - Dict str, Any ⋮---- data = json.load f ⋮---- def save provider health state: Dict str, Any - None ⋮---- def provider in cooldown provider: str - Tuple bool, int ⋮---- state = load provider… 证据：`web_search_plus_mcp/provider_health.py`
- **Provider Stats**（source_file）：PROVIDER STATS FILE = CACHE DIR / "provider stats.json" ⋮---- MAX SAMPLES PER PROVIDER = 50 ⋮---- SAMPLE MAX AGE SECONDS = 7 24 3600 ⋮---- MIN SAMPLES FOR ADJUSTMENT = 5 ⋮---- MAX SCORE ADJUSTMENT = 1.0 ⋮---- LATENCY CEILING SECONDS = 8.0 ⋮---- PERFORMANCE BASELINE = 0.75 ⋮---- STATS LOCK = threading.Lock ⋮---- def load stats - Dict str, Any ⋮---- data = json.load f ⋮---- def save stats state: Dict str, Any - None ⋮---- sample = { ⋮---- state = load stats samples = state.get provider ⋮---- samples = ⋮---- def fresh samples samples: Any, now: float - List Dict str, Any ⋮---- cutoff = now - SAMPLE MAX AGE SECONDS ⋮---- def get provider performance provider: str, now: Optional float = None - O… 证据：`web_search_plus_mcp/provider_stats.py`
- **Use last meaningful path segment as context**（source_file）：ROUTING POLICY = "routing-v2" ⋮---- def title from url url: str - str ⋮---- parsed = urlparse url domain = parsed.netloc.replace "www.", "" Use last meaningful path segment as context segments = s for s in parsed.path.strip "/" .split "/" if s ⋮---- last = segments -1 .replace "-", " " .replace " ", " " ⋮---- last = re.sub r'\.\w{2,4}$', '', last ⋮---- def normalize result url url: str - str ⋮---- parsed = urlparse url.strip netloc = parsed.netloc or "" .lower ⋮---- netloc = netloc 4: path = parsed.path.rstrip "/" ⋮---- def deduplicate results across providers results by provider: List Tuple str, Dict str, Any , max results: int - Tuple List Dict str, Any , int ⋮---- deduped = seen = set de… 证据：`web_search_plus_mcp/quality.py`
- **Research**（source_file）：RESULT GRACE SECONDS = 0.25 ⋮---- provider errors: List Dict str, Any = now = now fn or time.monotonic start = now ⋮---- def remaining budget - Optional float ⋮---- pending: List Tuple int, str = tasks: Dict int, DaemonTask = {} workers = max workers or max 1, len research providers gate = threading.Semaphore workers ⋮---- def run gated provider name: str - Dict str, Any ⋮---- remaining = remaining budget ⋮---- results by index: Dict int, Tuple str, Dict str, Any = {} ⋮---- timeout = RESULT GRACE SECONDS ⋮---- timeout = remaining ⋮---- provider results: List Tuple str, Dict str, Any = ⋮---- diversity duplicates = ⋮---- merged candidates: List Dict str, Any = ⋮---- candidate = item.copy ⋮---… 证据：`web_search_plus_mcp/research.py`
- **Determine provider**（source_file）：http make get request = make get request http make request = make request ⋮---- def make get request args, kwargs ⋮---- def make request args, kwargs ⋮---- get cached result = cache get cache search result = cache put clear cache = cache clear get cache stats = cache stats ⋮---- def load env file ⋮---- ROUTING POLICY = "routing-v2" ⋮---- COMPATIBILITY SHIM DEPRECATION = { ⋮---- def get compatibility shim policy - Dict str, Any ⋮---- def sync routing dependencies - None ⋮---- class QueryAnalyzer routing.QueryAnalyzer ⋮---- def init self, args, kwargs ⋮---- def auto route provider args, kwargs ⋮---- def explain routing args, kwargs ⋮---- def provider auto allowed args, kwargs ⋮---- def sync p… 证据：`web_search_plus_mcp/search.py`
- **Search Locale**（source_file）：LANGUAGE INFERENCE MIN MATCHES = 2 LANGUAGE INFERENCE STOPWORDS: Dict str, frozenset str = { LANGUAGE INFERENCE CHAR HINTS: Dict str, str = { ⋮---- def infer query language query: str - Optional str ⋮---- lowered = query.lower words = set re.findall r"\w+", lowered counts: Dict str, int = {} ⋮---- count = len words & stopwords ⋮---- ranked = sorted counts.items , key=lambda item: -item 1 , item 0 ⋮---- FALLBACK COUNTRY = "us" FALLBACK LANGUAGE = "en" ⋮---- AUTO LANGUAGE = "auto" ⋮---- PROVIDER LOCALE CONFIG KEYS: Dict str, Tuple Optional str , Optional str = { ⋮---- LOCATION COUNTRY HINTS: Dict str, str = { ⋮---- LOCATION HINT PATTERNS: Tuple Tuple Any, str , ... = tuple ⋮---- def provider… 证据：`web_search_plus_mcp/search_locale.py`
- **Shadow Policy V3**（source_file）：POLICY ID = "shadow-quality" POLICY REVISION = "3.1" ⋮---- query = request.input "query" analysis = QueryAnalyzer dict config .analyze query scores = analysis "provider scores" candidates = tuple dict.fromkeys plan.candidate order ranked = shadow provider = 证据：`web_search_plus_mcp/shadow_policy_v3.py`
- **State Migration V3**（source_file）：BACKUP OWNER = "web-search-plus-state-migration-v3" BACKUP SCHEMA VERSION = 1 MIGRATION ID = "legacy-json-v1" MAX SOURCE BYTES = 8 1024 1024 MAX PROVIDERS = 256 MAX SAMPLES PER PROVIDER = 1000 PROVIDER RE = re.compile r"^ a-z0-9 a-z0-9 - {0,63}$" BACKUP ID RE = re.compile r"^ 0-9 {8}T 0-9 {6}Z- a-f0-9 {12} ?:- 0-9 + ?$" ⋮---- @dataclass frozen=True class MigrationReport ⋮---- action: str status: str dry run: bool sqlite available: bool health providers: int = 0 adaptive providers: int = 0 adaptive samples: int = 0 source digest: str None = None backup id: str None = None error code: str None = None ⋮---- def to dict self - dict str, Any ⋮---- def render migration report report: MigrationRep… 证据：`web_search_plus_mcp/state_migration_v3.py`
- **Init**（source_file）：all = 证据：`wsp_sdk/__init__.py`
- **Api**（source_file）：SearchExecute = Callable Any, str, Any, str None, Mapping str, Any , Mapping str, Any , dict str, Any ExtractExecute = Callable ⋮---- @dataclass frozen=True, init=False class ProviderSpec ⋮---- provider: str env var: str display name: str description: str config section: str supports search: bool supports extract: bool capability labels: tuple str, ... auto allowed by default: bool recommended: bool free tier: str signup url: str upstream capabilities: tuple str, ... keyless: bool search output semantics: str None extract output semantics: str None provider fields allowlist: tuple str, ... rejected reason: str None execute search: SearchExecute None execute extract: ExtractExecute None supp… 证据：`wsp_sdk/api.py`
- **Conformance**（source_file）：def provider conformance errors - tuple str, ... ⋮---- errors = list dispatch conformance errors SEARCH DISPATCH, EXTRACT DISPATCH, PROVIDER SPECS ⋮---- payload = json.loads str exc ⋮---- def assert provider conformance - None ⋮---- errors = provider conformance errors 证据：`wsp_sdk/conformance.py`
- **Errors**（source_file）：class ProviderSDKError Exception ⋮---- class ProviderConfigError ProviderSDKError ⋮---- class ProviderContractFailure ProviderSDKError ⋮---- class ProviderRegistrationError ProviderSDKError ⋮---- class DuplicateProviderError ProviderRegistrationError ⋮---- class ProviderDiscoveryError ProviderSDKError ⋮---- class ProviderStartupDiagnostic ProviderDiscoveryError ⋮---- def init self, module: str, code: str - None 证据：`wsp_sdk/errors.py`

## 宿主 AI 必须遵守的规则

- **把本资产当作开工前上下文，而不是运行环境。**：AI Context Pack 只包含证据化项目理解，不包含目标项目的可执行状态。 证据：`README.md`, `LICENSE`, `docs/MIGRATION_1_0.md`
- **回答用户时区分可预览内容与必须安装后才能验证的内容。**：安装前体验的消费者价值来自降低误装和误判，而不是伪装成真实运行。 证据：`README.md`, `LICENSE`, `docs/MIGRATION_1_0.md`

## 用户开工前应该回答的问题

- 你准备在哪个宿主 AI 或本地环境中使用它？
- 你只是想先体验工作流，还是准备真实安装？
- 你最在意的是安装成本、输出质量、还是和现有规则的冲突？

## 验收标准

- 所有能力声明都能回指到 evidence_refs 中的文件路径。
- AI_CONTEXT_PACK.md 没有把预览包装成真实运行。
- 用户能在 3 分钟内看懂适合谁、能做什么、如何开始和风险边界。

---

## Doramagic Context Augmentation

下面内容用于强化 Repomix/AI Context Pack 主体。Human Manual 只提供阅读骨架；踩坑日志会被转成宿主 AI 必须遵守的工作约束。

## Human Manual 骨架

使用规则：这里只是项目阅读路线和显著性信号，不是事实权威。具体事实仍必须回到 repo evidence / Claim Graph。

宿主 AI 硬性规则：
- 不得把页标题、章节顺序、摘要或 importance 当作项目事实证据。
- 解释 Human Manual 骨架时，必须明确说它只是阅读路线/显著性信号。
- 能力、安装、兼容性、运行状态和风险判断必须引用 repo evidence、source path 或 Claim Graph。

- **项目概述与 v3 合约架构**：importance `high`
  - source_paths: README.md, web_search_plus_mcp/server.py, web_search_plus_mcp/contract_v3.py, web_search_plus_mcp/orchestrator_v3.py, web_search_plus_mcp/runtime_v3.py
- **提供商系统与自动路由（含 v1.0 回落修复）**：importance `high`
  - source_paths: web_search_plus_mcp/providers.py, web_search_plus_mcp/provider_registry.py, web_search_plus_mcp/provider_dispatch.py, web_search_plus_mcp/provider_adapter_protocol.py, web_search_plus_mcp/routing.py
- **提取、有界上下文、语义跨度与缓存**：importance `high`
  - source_paths: web_search_plus_mcp/extract.py, web_search_plus_mcp/bounded_context_v3.py, web_search_plus_mcp/span_extraction_v3.py, web_search_plus_mcp/cache_v3.py, web_search_plus_mcp/cache_identity_v3.py
- **配置、SDK、入门与运维（含 v1.0 迁移）**：importance `high`
  - source_paths: web_search_plus_mcp/config.py, web_search_plus_mcp/env_loader.py, web_search_plus_mcp/compat_v3.py, web_search_plus_mcp/request_gate_v3.py, web_search_plus_mcp/budget_preflight_v3.py

## Repo Inspection Evidence / 源码检查证据

- repo_clone_verified: true
- repo_inspection_verified: true
- repo_commit: `f9fe6737baf5fe8cab3c5eb9746fa6ac6a765fd7`
- inspected_files: `Dockerfile`, `README.md`, `pyproject.toml`, `uv.lock`, `docs/MIGRATION_1_0.md`

宿主 AI 硬性规则：
- 没有 repo_clone_verified=true 时，不得声称已经读过源码。
- 没有 repo_inspection_verified=true 时，不得把 README/docs/package 文件判断写成事实。
- 没有 quick_start_verified=true 时，不得声称 Quick Start 已跑通。

## Doramagic Pitfall Constraints / 踩坑约束

这些规则来自 Doramagic 发现、验证或编译过程中的项目专属坑点。宿主 AI 必须把它们当作工作约束，而不是普通说明文字。

### Constraint 1: 失败模式：installation: web-search-plus-mcp v0.11.0

- Trigger: Developers should check this installation risk before relying on the project: web-search-plus-mcp v0.11.0
- Host AI rule: Before packaging this project, run the relevant install/config/quickstart check for: web-search-plus-mcp v0.11.0. Context: Observed when using python
- Why it matters: Upgrade or migration may change expected behavior: web-search-plus-mcp v0.11.0
- Evidence: failure_mode_cluster:github_release | https://github.com/robbyczgw-cla/web-search-plus-mcp/releases/tag/v0.11.0 | web-search-plus-mcp v0.11.0
- Hard boundary: 不要把这个坑点包装成已解决、已验证或可忽略，除非后续验证证据明确证明它已经关闭。

### Constraint 2: 失败模式：installation: web-search-plus-mcp v0.17.0

- Trigger: Developers should check this installation risk before relying on the project: web-search-plus-mcp v0.17.0
- Host AI rule: Before packaging this project, run the relevant install/config/quickstart check for: web-search-plus-mcp v0.17.0. Context: Observed when using python
- Why it matters: Upgrade or migration may change expected behavior: web-search-plus-mcp v0.17.0
- Evidence: failure_mode_cluster:github_release | https://github.com/robbyczgw-cla/web-search-plus-mcp/releases/tag/v0.17.0 | web-search-plus-mcp v0.17.0
- Hard boundary: 不要把这个坑点包装成已解决、已验证或可忽略，除非后续验证证据明确证明它已经关闭。

### Constraint 3: 失败模式：installation: web-search-plus-mcp v1.0.0

- Trigger: Developers should check this installation risk before relying on the project: web-search-plus-mcp v1.0.0
- Host AI rule: Before packaging this project, run the relevant install/config/quickstart check for: web-search-plus-mcp v1.0.0. Context: Observed when using python
- Why it matters: Upgrade or migration may change expected behavior: web-search-plus-mcp v1.0.0
- Evidence: failure_mode_cluster:github_release | https://github.com/robbyczgw-cla/web-search-plus-mcp/releases/tag/v1.0.0 | web-search-plus-mcp v1.0.0
- Hard boundary: 不要把这个坑点包装成已解决、已验证或可忽略，除非后续验证证据明确证明它已经关闭。

### Constraint 4: 失败模式：installation: web-search-plus-mcp v1.1.0

- Trigger: Developers should check this installation risk before relying on the project: web-search-plus-mcp v1.1.0
- Host AI rule: Before packaging this project, run the relevant install/config/quickstart check for: web-search-plus-mcp v1.1.0. Context: Observed when using python
- Why it matters: Upgrade or migration may change expected behavior: web-search-plus-mcp v1.1.0
- Evidence: failure_mode_cluster:github_release | https://github.com/robbyczgw-cla/web-search-plus-mcp/releases/tag/v1.1.0 | web-search-plus-mcp v1.1.0
- Hard boundary: 不要把这个坑点包装成已解决、已验证或可忽略，除非后续验证证据明确证明它已经关闭。

### Constraint 5: 来源证据：v1.0.0: auto routing builds single-candidate plans — no fallback on quota/rate-limit

- Trigger: GitHub 社区证据显示该项目存在一个安装相关的待验证问题：v1.0.0: auto routing builds single-candidate plans — no fallback on quota/rate-limit
- Host AI rule: 来源显示可能已有修复、规避或版本变化，说明书中必须标注适用版本。
- Why it matters: 可能增加新用户试用和生产接入成本。
- Evidence: community_evidence:github | https://github.com/robbyczgw-cla/web-search-plus-mcp/issues/27 | 来源类型 github_issue 暴露的待验证使用条件。
- Hard boundary: 不要把这个坑点包装成已解决、已验证或可忽略，除非后续验证证据明确证明它已经关闭。

### Constraint 6: 可能修改宿主 AI 配置

- Trigger: 项目面向 Claude/Cursor/Codex/Gemini/OpenCode 等宿主，或安装命令涉及用户配置目录。
- Host AI rule: 列出会写入的配置文件、目录和卸载/回滚步骤。
- Why it matters: 安装可能改变本机 AI 工具行为，用户需要知道写入位置和回滚方法。
- Evidence: capability.host_targets | https://github.com/robbyczgw-cla/web-search-plus-mcp | host_targets=mcp_host, claude, cursor
- Hard boundary: 不要把这个坑点包装成已解决、已验证或可忽略，除非后续验证证据明确证明它已经关闭。

### Constraint 7: 失败模式：configuration: v0.12.0

- Trigger: Developers should check this configuration risk before relying on the project: v0.12.0
- Host AI rule: Before packaging this project, run the relevant install/config/quickstart check for: v0.12.0. Context: Source discussion did not expose a precise runtime context.
- Why it matters: Upgrade or migration may change expected behavior: v0.12.0
- Evidence: failure_mode_cluster:github_release | https://github.com/robbyczgw-cla/web-search-plus-mcp/releases/tag/v0.12.0 | v0.12.0
- Hard boundary: 不要把这个坑点包装成已解决、已验证或可忽略，除非后续验证证据明确证明它已经关闭。

### Constraint 8: 失败模式：configuration: v1.0.0: auto routing builds single-candidate plans — no fallback on quota/rate-limit

- Trigger: Developers should check this configuration risk before relying on the project: v1.0.0: auto routing builds single-candidate plans — no fallback on quota/rate-limit
- Host AI rule: Before packaging this project, run the relevant install/config/quickstart check for: v1.0.0: auto routing builds single-candidate plans — no fallback on quota/rate-limit. Context: Source discussion did not expose a precise runtime context.
- Why it matters: Developers may misconfigure credentials, environment, or host setup: v1.0.0: auto routing builds single-candidate plans — no fallback on quota/rate-limit
- Evidence: failure_mode_cluster:github_issue | https://github.com/robbyczgw-cla/web-search-plus-mcp/issues/27 | v1.0.0: auto routing builds single-candidate plans — no fallback on quota/rate-limit
- Hard boundary: 不要把这个坑点包装成已解决、已验证或可忽略，除非后续验证证据明确证明它已经关闭。

### Constraint 9: 失败模式：configuration: web-search-plus-mcp v0.13.0

- Trigger: Developers should check this configuration risk before relying on the project: web-search-plus-mcp v0.13.0
- Host AI rule: Before packaging this project, run the relevant install/config/quickstart check for: web-search-plus-mcp v0.13.0. Context: Source discussion did not expose a precise runtime context.
- Why it matters: Upgrade or migration may change expected behavior: web-search-plus-mcp v0.13.0
- Evidence: failure_mode_cluster:github_release | https://github.com/robbyczgw-cla/web-search-plus-mcp/releases/tag/v0.13.0 | web-search-plus-mcp v0.13.0
- Hard boundary: 不要把这个坑点包装成已解决、已验证或可忽略，除非后续验证证据明确证明它已经关闭。

### Constraint 10: 失败模式：configuration: web-search-plus-mcp v0.14.0

- Trigger: Developers should check this configuration risk before relying on the project: web-search-plus-mcp v0.14.0
- Host AI rule: Before packaging this project, run the relevant install/config/quickstart check for: web-search-plus-mcp v0.14.0. Context: Observed when using python
- Why it matters: Upgrade or migration may change expected behavior: web-search-plus-mcp v0.14.0
- Evidence: failure_mode_cluster:github_release | https://github.com/robbyczgw-cla/web-search-plus-mcp/releases/tag/v0.14.0 | web-search-plus-mcp v0.14.0
- Hard boundary: 不要把这个坑点包装成已解决、已验证或可忽略，除非后续验证证据明确证明它已经关闭。
