Skip to content

Repository files navigation

llm-router

Stop sending every prompt to the frontier model. A cost-aware gateway: a semantic cache + difficulty-based routing that serves near-duplicates for free and sends simple prompts to a small model, ~66% cheaper and ~57% faster than always using the big one, measured against an explicit baseline and gated in CI. Fully offline.

CI python license

Token cost is the line item that surprises teams in production (every 2026 agent playbook flags it). The two highest-leverage fixes are routing (use the cheapest model that can do the job) and caching (don't pay twice for the same question). This repo implements both, and, true to the rest of my portfolio, measures the result against a baseline and gates it, rather than asserting a number.

What this demonstrates

Lever Where
Difficulty-based model routing (cheap vs frontier) difficulty.py
Semantic cache (near-duplicate reuse) cache.py
Transparent per-tier price sheet models.py
Cost/latency measured vs an explicit baseline + gate benchmark.py

Architecture

flowchart TD
    P([prompt]) --> C{cache hit?}
    C -->|yes| HIT[cached answer · $0 · instant]
    C -->|no| D[classify difficulty]
    D -->|simple| CHEAP[small model]
    D -->|complex| BIG[frontier model]
    CHEAP --> STORE[store] --> R([response + cost + latency])
    BIG --> STORE
Loading

Quickstart

make dev            # venv + install -e ".[dev]"

llm-router route "What is the capital of France?"     # cheap tier
llm-router route "Explain why consensus is hard"      # capable tier
llm-router benchmark                                  # cost/routing/cache gate

No keys, no network. Set ROUTER_PROVIDER=openai to run the same logic against real models.

The benchmark

llm-router benchmark runs a workload and compares against always use the frontier model, no cache (report):

metric value threshold
cost_savings 66.0% ≥ 40%
routing_accuracy 1.000 ≥ 0.85
cache_hit_rate 22.2% ≥ 15%
latency_savings 56.7% n/a

Total spend $0.0185 vs a baseline of $0.0544, a 66% reduction, from routing simple prompts to a model ~20x cheaper per token and serving near-duplicates from cache. The model calls are mocked, but the routing and cache decisions are real and the price sheet is transparent, so the savings are the genuine effect of the policy, not a hardcoded number.

Design decisions

  • Cheapest capable model, not cheapest model. The router reserves the frontier model for genuine reasoning (explain / design / debug / derive) and sends lookups and classification to the small one, the difference that actually moves the bill.
  • Cache is semantic, not exact. Lightly-reworded prompts ("…to miles" vs "…into miles") still hit, via cosine similarity over prompt embeddings with a tunable threshold.
  • Baseline-relative, gated. Savings are measured against an explicit baseline and fail the build if the policy stops paying off, the same honest-metrics discipline as my eval and benchmark repos.

Layout

src/llm_router/  embeddings · models · difficulty · cache · router · provider · benchmark · cli
data/  workload.jsonl
reports/  benchmark_report_example.md

Related repositories

Part of a portfolio on production ML & LLM engineering:

License

MIT © 2026 Taha Siddiqui

About

Cost-aware LLM gateway: semantic cache + difficulty-based model routing that cuts spend (~66%) and latency vs always using the frontier model, measured against a baseline and gated in CI. Fully offline.

Topics

Resources

Stars

2 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages