Stop sending every prompt to the frontier model. A cost-aware gateway: a semantic cache + difficulty-based routing that serves near-duplicates for free and sends simple prompts to a small model, ~66% cheaper and ~57% faster than always using the big one, measured against an explicit baseline and gated in CI. Fully offline.
Token cost is the line item that surprises teams in production (every 2026 agent playbook flags it). The two highest-leverage fixes are routing (use the cheapest model that can do the job) and caching (don't pay twice for the same question). This repo implements both, and, true to the rest of my portfolio, measures the result against a baseline and gates it, rather than asserting a number.
| Lever | Where |
|---|---|
| Difficulty-based model routing (cheap vs frontier) | difficulty.py |
| Semantic cache (near-duplicate reuse) | cache.py |
| Transparent per-tier price sheet | models.py |
| Cost/latency measured vs an explicit baseline + gate | benchmark.py |
flowchart TD
P([prompt]) --> C{cache hit?}
C -->|yes| HIT[cached answer · $0 · instant]
C -->|no| D[classify difficulty]
D -->|simple| CHEAP[small model]
D -->|complex| BIG[frontier model]
CHEAP --> STORE[store] --> R([response + cost + latency])
BIG --> STORE
make dev # venv + install -e ".[dev]"
llm-router route "What is the capital of France?" # cheap tier
llm-router route "Explain why consensus is hard" # capable tier
llm-router benchmark # cost/routing/cache gateNo keys, no network. Set ROUTER_PROVIDER=openai to run the same logic against real models.
llm-router benchmark runs a workload and compares against always use the frontier model,
no cache (report):
| metric | value | threshold |
|---|---|---|
| cost_savings | 66.0% | ≥ 40% |
| routing_accuracy | 1.000 | ≥ 0.85 |
| cache_hit_rate | 22.2% | ≥ 15% |
| latency_savings | 56.7% | n/a |
Total spend $0.0185 vs a baseline of $0.0544, a 66% reduction, from routing simple prompts to a model ~20x cheaper per token and serving near-duplicates from cache. The model calls are mocked, but the routing and cache decisions are real and the price sheet is transparent, so the savings are the genuine effect of the policy, not a hardcoded number.
- Cheapest capable model, not cheapest model. The router reserves the frontier model for genuine reasoning (explain / design / debug / derive) and sends lookups and classification to the small one, the difference that actually moves the bill.
- Cache is semantic, not exact. Lightly-reworded prompts ("…to miles" vs "…into miles") still hit, via cosine similarity over prompt embeddings with a tunable threshold.
- Baseline-relative, gated. Savings are measured against an explicit baseline and fail the build if the policy stops paying off, the same honest-metrics discipline as my eval and benchmark repos.
src/llm_router/ embeddings · models · difficulty · cache · router · provider · benchmark · cli
data/ workload.jsonl
reports/ benchmark_report_example.md
Part of a portfolio on production ML & LLM engineering:
- ai-harness · llm-eval-observability · llm-guardrails-redteam · hybrid-graph-rag
- support-copilot · invoice-ap-agent · deep-research-agent · vendor-risk-agent
- timeseries-forecasting · tabular-ml · timeseries-classification
- llm-router: this repo.
MIT © 2026 Taha Siddiqui