# routellm - Doramagic AI Context Pack

> Positioning: a pre-install experience and judgment asset. It helps the host AI get off to a good start, but it does not mean the project has already been installed, run, or validated.

## Sufficiency Principle

- **Sufficiency over compression**: The AI Context Pack should be sufficient for the host AI to understand the project's value, capability boundaries, entrypoints, risks, and evidence sources before starting work; it may be layered, but it does not aim for the shortest possible summary.
- **Compression policy**: Compress only noise and duplication, never context that affects judgment or the quality of the work.

## How the Host AI Should Use This

You are reading the AI Context Pack that Doramagic compiled for routellm. Treat it as pre-work context: help the user understand who it fits, what it can do, how to start, what must be verified after install, and where the risks are. Do not claim that you have already installed, run, or executed the target project.

## Claim Consumption Rules

- **Fact source**: Repo Evidence + Claim/Evidence Graph; the Human Wiki only supplies salience, terminology, and narrative structure.
- **Minimum status for a fact**: `supported`
- `supported`: May be used as a project fact, but the answer must cite the claim_id and evidence path.
- `weak`: Usable only as a low-confidence lead; the user must be asked to keep verifying.
- `inferred`: Usable only for risk notes or open questions; must not be packaged as a project fact.
- `unverified`: Must not be used as fact; state clearly that evidence is insufficient.
- `contradicted`: Must show the conflicting sources and must not force a single version on the user's behalf.

## Who It Fits Best

- **AI researchers or builders of research-oriented Agents**: The README clearly centers on research, experiment, or paper workflows. Evidence: `README.md` Claim: `clm_0002` supported 0.86
- **Developers already using host AIs such as Claude/Codex/Cursor/Gemini**: The README or plugin config mentions multiple host AIs. Evidence: `README.md` Claim: `clm_0003` supported 0.86

## What It Can Do

- **Command-Line Startup or Install Flow** (Verify after install): The project documentation contains runnable commands; real use requires running them in a local or host environment. Evidence: `README.md` Claim: `clm_0001` supported 0.86

## How to Start

- `pip install "routellm[serve,eval]"` Evidence: `README.md` Claim: `clm_0004` supported 0.86
- `git clone https://github.com/lm-sys/RouteLLM.git` Evidence: `README.md` Claim: `clm_0005` supported 0.86
- `pip install -e .[serve,eval]` Evidence: `README.md` Claim: `clm_0006` supported 0.86

## Continue-or-Stop Decision Card

- **Current recommendation**: Needs admin / security approval
- **Why**: Continuing may involve secrets, accounts, external services, or sensitive context; get admin or security approval first.

### 30-Second Read

- **What to do now**: Needs admin / security approval
- **Minimum safe next step**: Run Prompt Preview first; if credentials or an enterprise environment are involved, get approval before trialing
- **Do not trust yet**: Real output quality cannot be trusted before install.
- **Continuing will touch**: Command execution, Local environment or project files, Environment variables / API keys

### What You Can Trust Now

- **Target-audience signal: AI researchers or builders of research-oriented Agents** (supported): Backed by a supported claim or project evidence, but that still is not the same as real install results. Evidence: `README.md` Claim: `clm_0002` supported 0.86
- **Target-audience signal: Developers already using host AIs such as Claude/Codex/Cursor/Gemini** (supported): Backed by a supported claim or project evidence, but that still is not the same as real install results. Evidence: `README.md` Claim: `clm_0003` supported 0.86
- **Capability exists: Command-Line Startup or Install Flow** (supported): You can trust that the project contains signals of this capability; whether it fits your specific task still needs trial or after-install verification. Evidence: `README.md` Claim: `clm_0001` supported 0.86
- **There are Quick Start / install-command signals** (supported): You can trust that the docs mention a startup or install entrypoint; do not run it directly in your primary environment because of that. Evidence: `README.md` Claim: `clm_0004` supported 0.86

### What You Cannot Trust Yet

- **Real output quality cannot be trusted before install.** (unverified): Prompt Preview can only show how it guides you; it cannot prove result quality in the real project.
- **Host AI version compatibility cannot be trusted before install.** (unverified): Host loading rules and version differences across Claude, Cursor, Codex, Gemini, and others must be verified in a real environment.
- **That it will not pollute your existing host AI's behavior cannot be trusted directly.** (inferred): Skill, plugin, and AGENTS/CLAUDE/GEMINI instructions may change the host AI's default behavior.
- **Safe rollback cannot be assumed by default.** (unverified): Unless the project clearly provides uninstall and recovery instructions, verify in an isolated environment first.
- **After a real install, is it compatible with the user's current host AI version?** (unverified): Compatibility can only be verified in the actual host environment.
- **Does the project's output quality meet the user's specific task?** (unverified): The pre-install preview can only show flow and boundaries; it cannot replace real evaluation.
- **Do the install commands require network access, permissions, or global writes?** (unverified): This affects install risk in both enterprise and personal environments. Evidence: `README.md`

### What Continuing Will Touch

- **Command execution**: Package managers, network downloads, the local plugin directory, project config, or the user's home directory. Why: Running the very first command can already change your environment; decide whether it is worth running first. Evidence: `README.md`
- **Local environment or project files**: Install results, plugin caches, project config, or local dependency directories. Why: The write scope and rollback path cannot be proven before install and need isolated verification. Evidence: `README.md`
- **Environment variables / API keys**: Project entry docs explicitly showing API key, token, secret, or account credential configuration. Why: If a real install needs credentials, use test credentials first and go through a permission/compliance review. Evidence: `README.md`, `examples/python_sdk.md`, `examples/routing_to_local_models.md`
- **Host AI context**: The AI Context Pack, Prompt Preview, Skill routing, risk rules, and project facts. Why: Importing context affects the host AI's later judgment, so avoid packaging unverified items as facts.

### Minimum Safe Next Steps

- **Run Prompt Preview first**: Use a pre-install interactive trial to judge whether the way of working fits; it needs no authorization or environment change. (applies when: Applies to any project, especially when output quality is unknown.)
- **Trial-install only in an isolated directory or a test account**: Avoid letting install commands pollute your primary host AI, real projects, or home directory. (applies when: When there are signals of command execution, plugin config, or local writes.)
- **Do not use real production credentials**: Once an environment variable / API key enters the host or toolchain, it can create account and compliance risk. (applies when: When environment signals like API, TOKEN, KEY, or SECRET appear.)
- **After install, verify just one minimal task**: Verify loading, compatibility, output quality, and rollback first, then decide whether to use it deeply. (applies when: When moving from a trial into a real workflow.)

### Exit Plan

- **Preserve the pre-install state**: Record the original host config and project state so you can later judge whether it is recoverable.
- **Record the install commands and written paths**: Without clear uninstall instructions, you at least need to know which directories or configs to clean up manually.
- **Be ready to revoke test API keys or tokens**: If test credentials leak or are misused, you can cut losses quickly.
- **If there is no rollback path, do not enter your primary environment**: No rollback is a blocker before continuing; do not proceed on trust or luck.

## What Can Only Be Previewed

- Explain who the project fits and what it can do
- Demonstrate a typical conversation flow based on project docs
- Help the user decide whether it is worth installing or researching further

## What Must Be Verified After Install

- Actually installing the Skill, plugin, or CLI
- Running scripts, modifying local files, or accessing external services
- Verifying real output quality, performance, and compatibility

## Boundary & Risk Decision Card

- **Mistaking the pre-install preview for a real run**: The user may overestimate how much configuration, permission, and compatibility verification the project has already done. Mitigation: Clearly separate prompt_preview_can_do from runtime_required. Claim: `clm_0007` inferred 0.45
- **Command execution will modify the local environment**: Install commands may write to the user's home directory, the host plugin directory, or project configuration. Mitigation: Run in an isolated environment or a test account first. Evidence: `README.md` Claim: `clm_0008` supported 0.86
- **To confirm**: After a real install, is it compatible with the user's current host AI version?. Why: Compatibility can only be verified in the actual host environment.
- **To confirm**: Does the project's output quality meet the user's specific task?. Why: The pre-install preview can only show flow and boundaries; it cannot replace real evaluation.
- **To confirm**: Do the install commands require network access, permissions, or global writes?. Why: This affects install risk in both enterprise and personal environments.

## Pre-Work Working Context

### Loading Order

- First read how_to_use.host_ai_instruction to establish the boundaries of this pre-install judgment asset.
- Read claim_graph_summary to confirm facts come from the Claim/Evidence Graph, not the Human Wiki narrative.
- Then read intended_users, capabilities, and quick_start_candidates to judge whether the user is a match.
- When you need to carry out a concrete task, check role_skill_index first, then evidence_index.
- For real install, file modification, network access, performance, or compatibility questions, turn to risk_card and boundaries.runtime_required.

### Task Routes

- **Command-Line Startup or Install Flow**: State that this is an after-install capability first, then give a pre-install checklist. Boundary: Must be verified after a real install or run. Evidence: `README.md` Claim: `clm_0001` supported 0.86

### Context Scale

- Total files: 99
- Important-file coverage: 21/99
- Evidence index entries: 20
- Role / Skill entries: 4

### Handling Insufficient Evidence

- **missing_evidence**: State that evidence is insufficient and ask the user for the target file, a README section, or after-install verification records; do not fill in facts.
- **out_of_scope_request**: State that the task is beyond the current AI Context Pack's evidence scope and suggest the user check the Human Manual or verify after a real install.
- **runtime_request**: Provide a pre-install checklist and command sources, but do not run commands for the user or claim they have been run.
- **source_conflict**: Show the conflicting sources side by side, mark them as unverified, and do not force a single version.

## Prompt Recipes

### Fit assessment

- Goal: Judge whether this project fits the user's current task.
- Expected output: A fit conclusion, key reasons, evidence citations, what can be previewed before install, what must be verified after install, and a next-step recommendation.

```text
Based on the AI Context Pack for routellm, ask me 3 necessary questions first, then judge whether it fits my task. The answer must cover: who it fits, what it can do, what it cannot do, whether it is worth installing, and where the evidence comes from. Every project fact must cite evidence_refs, source_paths, or a claim_id.
```

### Pre-install experience

- Goal: Let the user feel the core workflow before installing, while avoiding packaging the preview as real capability or a marketing promise.
- Expected output: An experience script with boundary labels, an after-install verification checklist, and a cautious recommendation; with no real-run promises or strong marketing language.

```text
Treat routellm as a pre-install experience asset, not an already-installed tool or a real runtime environment.

Output exactly four parts:
1. Ask me 3 necessary questions first.
2. Give an "experience script": use the three labels [Previewable before install], [Must verify after install], and [Insufficient evidence] to show how it might guide the workflow.
3. Give an after-install verification checklist: list which capabilities can only be confirmed after a real install, real host loading, and a real project run.
4. Give a cautious recommendation: only "worth researching/trialing further", "add information before deciding", or "not recommended to continue"; do not endorse the project.

Hard boundaries:
- Do not claim you have installed, run, executed tests, modified files, or produced real results.
- Do not write promise-like phrasing such as "auto-adapts", "guarantees passing", "perfect fit", or "strongly recommend installing".
- If you describe how it works after install, you must use a conditional such as "if installed successfully and the host loads the Skill correctly, it might...".
- The experience script may only be written as "example lines / hypothetical flow": use "might ask / might suggest / might show", not "has written, has generated, has passed, is running, is generating".
- Prompt Preview does not hand out install commands; if the user is ready to trial, only prompt them to read Quick Start and the Risk Card first and to verify in an isolated environment.
- Every project fact must come from a supported claim, evidence_refs, or source_paths; inferred/unverified items can only be risks or open questions.

```

### Role / Skill selection

- Goal: Pick the best-matching asset from the project's roles or Skills.
- Expected output: A list of candidate roles or Skills, each with an applicable scenario, evidence paths, risk boundary, and whether after-install verification is needed.

```text
Read role_skill_index and recommend 3-5 of the most relevant roles or Skills for my target task. For each recommendation, state the applicable scenario, likely output, risk boundary, and evidence_refs.
```

### Risk pre-check

- Goal: Identify environment, permission, rule-conflict, and quality risks before installing or adopting.
- Expected output: A checklist of environment, permission, dependency, license, host-conflict, quality risk, and unknown items.

```text
Based on risk_card, boundaries, and quick_start_candidates, give me a pre-install risk pre-check list. Do not run commands for me; only explain what I should check, why, and what impact a failure would have.
```

### Host AI kickoff instruction

- Goal: Turn the project context into a host AI instruction for the start of a conversation.
- Expected output: A pre-work instruction with clear boundaries and clear evidence citations, suitable to copy to a host AI.

```text
Based on the AI Context Pack for routellm, generate a pre-work instruction I can paste to my host AI. This instruction must obey not_runtime=true and must not claim the project has been installed, run, or produced real results.
```

## Role / Skill Index

- Indexed 4 role / Skill / project-doc entries.

- **RouteLLM** (project_doc): RouteLLM is a framework for serving and evaluating LLM routers. Activation hint: Reference this when the user needs to understand the project's structure, install path, or boundaries. Evidence: `README.md`
- **Benchmarks** (project_doc): In our blog https://lmsys.org/blog/2024-07-01-routellm/ , we report MT Bench results comparing our routers to Martian https://withmartian.com and Unify AI https://unify.ai , two commercial offerings for routing. We conduct these benchmarks using the official FastChat https://github.com/lm-sys/FastChat/tree/main/fastchat/llm judge repository, replacing model calls to either Martian or Unify AI using an OpenAI-compati… Activation hint: Reference this when the user needs to understand the project's structure, install path, or boundaries. Evidence: `benchmarks/README.md`
- **RouteLLM Python SDK** (project_doc): You can interface with RouteLLM in Python via the Controller class. Activation hint: Reference this when the user needs to understand the project's structure, install path, or boundaries. Evidence: `examples/python_sdk.md`
- **Routing to Local Models with RouteLLM and Ollama** (project_doc): Routing to Local Models with RouteLLM and Ollama Activation hint: Reference this when the user needs to understand the project's structure, install path, or boundaries. Evidence: `examples/routing_to_local_models.md`

## Evidence Index

- Indexed 20 evidence entries.

- **RouteLLM** (documentation): RouteLLM is a framework for serving and evaluating LLM routers. Evidence: `README.md`
- **Benchmarks** (documentation): In our blog https://lmsys.org/blog/2024-07-01-routellm/ , we report MT Bench results comparing our routers to Martian https://withmartian.com and Unify AI https://unify.ai , two commercial offerings for routing. We conduct these benchmarks using the official FastChat https://github.com/lm-sys/FastChat/tree/main/fastchat/llm judge repository, replacing model calls to either Martian or Unify AI using an OpenAI-compatible server. Evidence: `benchmarks/README.md`
- **License** (source_file): Apache License Version 2.0, January 2004 http://www.apache.org/licenses/ Evidence: `LICENSE`
- **RouteLLM Python SDK** (documentation): You can interface with RouteLLM in Python via the Controller class. Evidence: `examples/python_sdk.md`
- **Routing to Local Models with RouteLLM and Ollama** (documentation): Routing to Local Models with RouteLLM and Ollama Evidence: `examples/routing_to_local_models.md`
- **Config.Example** (source_file): sw ranking: arena battle datasets: - lmsys/lmsys-arena-human-preference-55k - routellm/gpt4 judge battles arena embedding datasets: - routellm/arena battles embeddings - routellm/gpt4 judge battles embeddings causal llm: checkpoint path: routellm/causal llm gpt4 augmented bert: checkpoint path: routellm/bert gpt4 augmented mf: checkpoint path: routellm/mf gpt4 augmented Evidence: `config.example.yaml`
- **Create and launch a chat interface with Gradio** (source_file): parser = argparse.ArgumentParser description="Chatbot Interface for RouteLLM" ⋮---- args = parser.parse args ⋮---- openai api key = "EMPTY" openai api base = args.model url ⋮---- client = OpenAI ⋮---- def predict message, history, threshold, router, temperature ⋮---- history openai format = ⋮---- stream = client.chat.completions.create ⋮---- model=f"router-{router}-{threshold}", Model name to use messages=history openai format, Chat history temperature=temperature, Temperature for text generation stream=True, Stream response ⋮---- partial message = "" ⋮---- model prefix = f" {chunk.model} \n" ⋮---- Create and launch a chat interface with Gradio Evidence: `examples/router_chat.py`
- **Pyproject** (source_file): build-system requires = "setuptools" build-backend = "setuptools.build meta" Evidence: `pyproject.toml`
- **Calibrate Threshold** (source_file): parser = argparse.ArgumentParser ⋮---- args = parser.parse args ⋮---- battles df = load dataset args.battles dataset, split="train" .to pandas controller = Controller ⋮---- win rates = controller.batch calculate win rate ⋮---- thresholds df = load dataset ⋮---- threshold = thresholds df router .quantile q=1 - args.strong model pct Evidence: `routellm/calibrate_threshold.py`
- **Some Python magic to match the OpenAI Python SDK** (source_file): GPT 4 AUGMENTED CONFIG = { ⋮---- class RoutingError Exception ⋮---- @dataclass class ModelPair ⋮---- strong: str weak: str ⋮---- class Controller ⋮---- config = GPT 4 AUGMENTED CONFIG ⋮---- router pbar = None ⋮---- router pbar = tqdm routers ⋮---- Some Python magic to match the OpenAI Python SDK ⋮---- def parse model name self, model: str ⋮---- threshold = float threshold ⋮---- Look at the last turn for routing. Our current routers were only trained on first turn data, so more research is required here. prompt = messages -1 "content" routed model = self.routers router .route prompt, threshold, self.model pair ⋮---- Mainly used for evaluations ⋮---- router instance = self.routers router ⋮---… Evidence: `routellm/controller.py`
- **Openai Server** (source_file): CONTROLLER = None ⋮---- openai client = AsyncOpenAI count = defaultdict lambda: defaultdict int ⋮---- @asynccontextmanager async def lifespan app ⋮---- CONTROLLER = Controller ⋮---- app = fastapi.FastAPI lifespan=lifespan ⋮---- class ErrorResponse BaseModel ⋮---- object: str = "error" message: str ⋮---- class UsageInfo BaseModel ⋮---- prompt tokens: int = 0 total tokens: int = 0 completion tokens: Optional int = 0 ⋮---- class ChatCompletionRequest BaseModel ⋮---- model: str messages: Union frequency penalty: Optional float = 0.0 logit bias: Optional Dict int, float = None logprobs: Optional bool = None top logprobs: Optional int = None max tokens: Optional int = None n: Optional int = 1 pre… Evidence: `routellm/openai_server.py`
- **Llm Utils** (source_file): def load model config yaml path: str ⋮---- yaml data = yaml.safe load file ⋮---- def load prompt format model id ⋮---- prompt format dict = PROMPT FORMAT CONFIGS model id ⋮---- def get model config: RouterModelConfig, model ckpt: str, pad token id: int = 2 ⋮---- tokenizer = AutoTokenizer.from pretrained ⋮---- def to openai api messages system message, classifier message, messages ⋮---- ret = {"role": "system", "content": system message} Evidence: `routellm/routers/causal_llm/llm_utils.py`
- **assert that config has 5 special tokens with the format rating** (source_file): class CausalLLMClassifier ⋮---- s = time.time ⋮---- assert that config has 5 special tokens with the format rating ⋮---- model = get model config=config, model ckpt=ckpt local path ⋮---- def preprocess self, row ⋮---- """prepare each prompt before feeding it to the model""" add additional fields to the final output e.g. for later evaluation data row = {} ⋮---- select turns from the prompt field openai messages = convert openai messages formot to model's prompt format text = self.prompt format.generate prompt openai messages tokenize and encode ⋮---- def call self, row ⋮---- row = self.preprocess row input ids = torch.as tensor row "input ids" .to "cuda" .reshape 1, -1 ⋮---- output new = sel… Evidence: `routellm/routers/causal_llm/model.py`
- **Model** (source_file): MODEL IDS = { ⋮---- class MFModel torch.nn.Module, PyTorchModelHubMixin ⋮---- def get device self ⋮---- def forward self, model id, prompt ⋮---- model id = torch.tensor model id, dtype=torch.long .to self.get device ⋮---- model embed = self.P model id model embed = torch.nn.functional.normalize model embed, p=2, dim=1 ⋮---- prompt embed = prompt embed = torch.tensor prompt embed, device=self.get device prompt embed = self.text proj prompt embed ⋮---- @torch.no grad def pred win rate self, model a, model b, prompt ⋮---- logits = self.forward model a, model b , prompt winrate = torch.sigmoid logits 0 - logits 1 .item ⋮---- def load self, path Evidence: `routellm/routers/matrix_factorization/model.py`
- **Train Matrix Factorization** (source_file): class PairwiseDataset Dataset ⋮---- def init self, data ⋮---- def len self ⋮---- def getitem self, index ⋮---- def get dataloaders self, batch size, shuffle=True ⋮---- class MFModel Train torch.nn.Module ⋮---- embeddings = np.load npy path ⋮---- def get device self ⋮---- def forward self, model win, model loss, prompt, test=False, alpha=0.05 ⋮---- model win = model win.to self.get device model loss = model loss.to self.get device prompt = prompt.to self.get device ⋮---- model win embed = self.P model win model win embed = F.normalize model win embed, p=2, dim=1 model loss embed = self.P model loss model loss embed = F.normalize model loss embed, p=2, dim=1 prompt embed = self.Q prompt ⋮----… Evidence: `routellm/routers/matrix_factorization/train_matrix_factorization.py`
- **Routers** (source_file): def no parallel cls ⋮---- class Router abc.ABC ⋮---- NO PARALLEL = False ⋮---- @abc.abstractmethod def calculate strong win rate self, prompt ⋮---- def route self, prompt, threshold, routed pair ⋮---- def str self ⋮---- @no parallel class CausalLLMRouter Router ⋮---- model config = RouterModelConfig prompt format = load prompt format model config.model id ⋮---- system message = hf hub download classifier message = hf hub download ⋮---- system message = pr.read ⋮---- classifier message = pr.read ⋮---- def calculate strong win rate self, prompt ⋮---- input = {} ⋮---- output = self.router model input ⋮---- @no parallel class BERTRouter Router ⋮---- inputs = self.tokenizer ⋮---- outputs = self.… Evidence: `routellm/routers/routers.py`
- **Generate Embeddings** (source_file): def get embeddings battles df ⋮---- battles df = preprocess battles battles df ⋮---- client = openai.OpenAI ⋮---- batch size = 2000 embeddings = user prompts = battles df "first turn" .tolist ⋮---- battles = user prompts i : i + batch size responses = client.embeddings.create ⋮---- embeddings = torch.tensor embeddings embeddings = embeddings.numpy ⋮---- parser = argparse.ArgumentParser ⋮---- args = parser.parse args battles df = load dataset args.battles dataset, split="train" .to pandas ⋮---- embeddings = get embeddings battles df ⋮---- embeddings dataset = Dataset.from dict {"embeddings": embeddings} Evidence: `routellm/routers/similarity_weighted/generate_embeddings.py`
- **Utils** (source_file): choices = "A", "B", "C", "D" OPENAI CLIENT = OpenAI ⋮---- def compute tiers model ratings, num tiers ⋮---- n = len model ratings m = num tiers ⋮---- model ratings list = list model ratings.values ⋮---- dp = np.zeros n, n, m dp split = np.zeros n, n, m ⋮---- cur n = n split idx = ⋮---- split = int dp split 0 cur n - 1 tier ⋮---- cur n = split + 1 ⋮---- split idx = split idx ::-1 + n - 1 model2tier = {} cur idx = 0 ⋮---- model name = list model ratings.keys j ⋮---- cur idx = split idx i + 1 ⋮---- models = pd.concat df "model a" , df "model b" .unique models = pd.Series np.arange len models , index=models ⋮---- df = pd.concat df, df , ignore index=True p = len models.index n = df.shape 0 ⋮----… Evidence: `routellm/routers/similarity_weighted/utils.py`
- **RouteLLM specific** (source_file): RouteLLM specific routellm/evals/ /cache.npy Evidence: `.gitignore`
- **.Pre Commit Config** (source_file): repos: - repo: https://github.com/psf/black rev: 24.4.2 hooks: - id: black - repo: https://github.com/pycqa/isort rev: 5.13.2 hooks: - id: isort name: isort python Evidence: `.pre-commit-config.yaml`

## Rules the Host AI Must Follow

- **Treat this asset as pre-work context, not a runtime environment.**: The AI Context Pack contains only an evidence-backed understanding of the project, not the project's executable state. Evidence: `README.md`, `benchmarks/README.md`, `LICENSE`
- **When answering the user, distinguish what can be previewed from what can only be verified after install.**: The consumer value of the pre-install experience comes from reducing bad installs and misjudgments, not from pretending to be a real run. Evidence: `README.md`, `benchmarks/README.md`, `LICENSE`

## Questions the User Should Answer First

- Which host AI or local environment do you plan to use it in?
- Do you just want to experience the workflow first, or are you ready to actually install?
- What matters most to you: install cost, output quality, or conflicts with your existing rules?

## Acceptance Checks

- Every capability claim can be traced back to a file path in evidence_refs.
- AI_CONTEXT_PACK.md does not package previews as a real run.
- The user can understand who it fits, what it can do, how to start, and the risk boundaries within 3 minutes.

---

## Doramagic Context Augmentation

The following sections strengthen the repository context for a host AI. Human Manual data is a reading route, and pitfall notes become operating constraints.

## Human Manual Outline

Usage rule: this is only a reading route and salience signal, not factual authority. Concrete claims must still return to repo evidence or Claim Graph.

Host AI hard rules:
- Do not treat page titles, section order, summaries, or importance values as factual project evidence.
- When explaining the Human Manual outline, state that it is only a reading route or salience signal.
- Capability, installation, compatibility, runtime state, and risk claims must cite repo evidence, source paths, or Claim Graph.

- **Overview & System Architecture**: importance `high`
  - source_paths: routellm/__init__.py, routellm/controller.py, routellm/openai_server.py, README.md
- **Router Implementations & Algorithms**: importance `high`
  - source_paths: routellm/routers/routers.py, routellm/routers/matrix_factorization/model.py, routellm/routers/matrix_factorization/train_matrix_factorization.py, routellm/routers/similarity_weighted/utils.py, routellm/routers/similarity_weighted/generate_embeddings.py
- **Deployment, Configuration & Threshold Calibration**: importance `high`
  - source_paths: routellm/openai_server.py, routellm/calibrate_threshold.py, config.example.yaml, examples/router_chat.py, examples/python_sdk.md
- **Extensibility, Training & Custom Integrations**: importance `high`
  - source_paths: routellm/routers/routers.py, routellm/routers/matrix_factorization/train_matrix_factorization.py, routellm/controller.py, routellm/routers/similarity_weighted/generate_embeddings.py, pyproject.toml

## Repo Inspection Evidence

- repo_clone_verified: true
- repo_inspection_verified: true
- repo_commit: `0b64fdafe049e596a3f5657c219329f24af24198`
- inspected_files: `README.md`, `pyproject.toml`, `examples/python_sdk.md`, `examples/router_chat.py`, `examples/routing_to_local_models.md`

Host AI hard rules:
- Without repo_clone_verified=true, do not claim that the source code has been read.
- Without repo_inspection_verified=true, do not write README, docs, or package-file conclusions as facts.
- Without quick_start_verified=true, do not claim that the Quick Start path has run successfully.

## Doramagic Pitfall Constraints

These rules come from Doramagic discovery, validation, or compilation findings. The host AI must treat them as operating constraints, not background notes.

### Constraint 1: Capability evidence risk requires verification

- Trigger: README/documentation is current enough for a first validation pass.
- Host AI rule: Reproduce the official install and quickstart path in an isolated environment.
- Why it matters: May increase setup, validation, or first-run risk for the user.
- Evidence: capability.assumptions | https://github.com/lm-sys/RouteLLM
- Hard boundary: Do not present this pitfall as solved, verified, or ignorable unless later evidence explicitly closes it.

### Constraint 2: Security or permission risk requires verification

- Trigger: no_demo
- Host AI rule: Reproduce the official install and quickstart path in an isolated environment.
- Why it matters: May increase setup, validation, or first-run risk for the user.
- Evidence: downstream_validation.risk_items | https://github.com/lm-sys/RouteLLM
- Hard boundary: Do not present this pitfall as solved, verified, or ignorable unless later evidence explicitly closes it.

### Constraint 3: Security or permission risk requires verification

- Trigger: no_demo
- Host AI rule: Reproduce the official install and quickstart path in an isolated environment.
- Why it matters: May increase setup, validation, or first-run risk for the user.
- Evidence: risks.scoring_risks | https://github.com/lm-sys/RouteLLM
- Hard boundary: Do not present this pitfall as solved, verified, or ignorable unless later evidence explicitly closes it.

### Constraint 4: Maintenance risk requires verification

- Trigger: issue_or_pr_quality=unknown。
- Host AI rule: Reproduce the official install and quickstart path in an isolated environment.
- Why it matters: May increase setup, validation, or first-run risk for the user.
- Evidence: evidence.maintainer_signals | https://github.com/lm-sys/RouteLLM
- Hard boundary: Do not present this pitfall as solved, verified, or ignorable unless later evidence explicitly closes it.
