A German-language RAG agent over automotive technical documents — with a measured evaluation harness, a working automation workflow, and a container image small enough to deploy for €0.00.
▶ Live demo: docubot-pb.streamlit.app — no signup, no key needed to see it refuse an out-of-scope question.
Ask "Was bedeutet der Fehlercode P0420?" and get a grounded German answer cited to the exact record it came from. Ask something the corpus does not cover and get an explicit refusal instead of a confident invention — and that refusal costs no LLM call at all.
- Projektübersicht (DE)
- Why this is not another "chat with your PDFs" demo
- Architecture
- Run it locally
- Evaluation results
- Optimization: PyTorch → ONNX int8
- API
- Automation workflow
- Deployment
- Design decisions
- Known limitations
- Tests
- Project status
- Licence and document provenance
DocuBot ist ein deutschsprachiger RAG-Agent für Kfz-Technik: OBD-II-Fehlercodes, Wartungsintervalle und der Prüfumfang von HU und AU nach StVZO § 29. Der Korpus besteht aus vier eigens verfassten deutschen Fachdokumenten sowie einem 52-seitigen Auszug der StVZO — zusammen 502 indexierte Chunks.
Die Einbettung erfolgt mit einem mehrsprachigen Modell, dessen Kontextfenster bei 128 Tokens endet. Diese Grenze bestimmt das gesamte Chunking: Fließtext wird tokenweise statt zeichenweise geteilt, und jeder Fehlercode bildet einen eigenen Chunk. Fragen dürfen auf Deutsch oder Englisch gestellt werden — abgerufen wird immer gegen den deutschen Index, die Antwortsprache ist davon unabhängig.
Kernergebnisse der Evaluation: Trefferquote 100 %, MRR 0,979, Context Recall 93,2 % bei k=8, sowie 5 von 5 korrekten Verweigerungen bei Fragen außerhalb des Korpus. Das Container-Image wurde von 934 MB auf 251 MB reduziert, der Kaltstart von 17,5 s auf 5,0 s — ohne messbaren Qualitätsverlust beim Retrieval.
Die Demo läuft unter https://docubot-pb.streamlit.app/ — kostenfrei, ohne Zahlungsmittel und ohne hinterlegtes Abrechnungskonto.
- German-language corpus and prompting, using a multilingual embedding model — not English documents behind a German UI label.
- Cross-lingual retrieval: ask in English, get an English answer cited to the original German sources. One prompt per answer language, because an 8B model follows the dominant language of its instructions and will answer a German-instructed prompt in German no matter what you ask it.
- A real evaluation harness — hit rate, MRR, context recall, answer similarity and abstention accuracy on a hand-written 29-pair German QA set, wired in as a Docker build gate rather than a script nobody remembers to run.
- The embedding model's limits are designed around, not ignored. It truncates silently at 128 tokens, so chunking is measured in tokens and enforced by a build-failing assertion.
- Dense retrieval's classic failure mode is fixed, not hidden. Pure vector search scored 0/6 top-1 on fault-code lookup; a deterministic lexical path took it to 8/8. See Design decisions.
- It abstains, two independent ways, and both are measured.
flowchart TD
subgraph build["docker build"]
DOCS["documents/<br/>German automotive .md + .pdf"] --> SPLIT["per-source splitter<br/>prose: token-aware · fault codes: 1 chunk/record"]
SPLIT --> EMB["multilingual MiniLM<br/>128-token ceiling, asserted"]
EMB --> CHROMA[("ChromaDB<br/>baked in, read-only")]
CHROMA --> GATE{"retrieval eval gate<br/>fails the build under 80% hit rate"}
end
subgraph api["docubot-api"]
GATE --> RET["hybrid retrieval<br/>lexical fault codes + dense"]
RET --> FLOOR{"relevance floor"}
FLOOR -->|below| ABST["German abstention<br/>no LLM call"]
FLOOR -->|above| LCEL["LCEL chain<br/>bilingual prompt + grounding rules"]
LCEL --> GROQ["Llama 3.1 8B"]
GROQ --> EP["/chat · /health · /metrics<br/>/ingest → 403 in production"]
ABST --> EP
end
EP --> UI["Streamlit UI<br/>DE/EN toggle · bring-your-own-key"]
EP --> N8N["automation workflow<br/>webhook → ingest → chat → Discord"]
EP --> EVAL["eval stage B<br/>on demand, temperature 0"]
The index is built at image-build time and shipped read-only. A container filesystem is ephemeral, so anything written at runtime disappears on the next instance recycle — a runtime-only ingest would mean the first request after every cold start hits an empty vector store.
The same chain runs in two shapes. Locally and in the container it sits behind the HTTP API above. On the live demo the UI calls it in-process, because the only free host with enough memory runs Streamlit apps and nothing else — see Deployment. The retrieval, prompt and abstention logic is one module either way; only the transport differs.
Everything runs on CPU. No GPU, no paid service, no vector-database account.
git clone https://github.com/Prennoy99/DocuBot.git && cd DocuBot
cp .env.example .env # optional: add a free Groq key for generation
docker compose up # api on :8000, UI on :8501Open http://localhost:8501. Without a key it still works: out-of-scope questions are refused
(that path needs no LLM), and in-scope questions return 401 with instructions for pasting your own
key into the sidebar. Add --profile n8n to bring up the automation stack as well.
To work on it without Docker:
python -m venv .venv && source .venv/bin/activate
pip install -r requirements.txt # PyTorch backend, for local development
python -m app.ingest # build the index (~41 s)
uvicorn app.main:app --reload # :8000
streamlit run frontend/app.py # :8501To run the embedded single-process app — the shape that is deployed — instead:
pip install -r deploy/streamlit/requirements.txt # torch-free; no FastAPI, no uvicorn
streamlit run deploy/streamlit/streamlit_app.py # fetches the model and index on first bootTwo stages. Stage A is deterministic, costs nothing, and gates every Docker build. Stage B goes over
HTTP at temperature=0 and spends one API call per answerable pair.
Stage A — retrieval only, 24 answerable pairs, zero LLM calls:
k hit rate MRR context recall
3 100.0% 0.979 88.6%
5 100.0% 0.979 88.6%
8 100.0% 0.979 93.2%
Stage B — end to end, 29 pairs:
keyword recall 0.722
answer similarity (cosine) 0.778
correct abstentions 5/5
grounded, no false abstention 22/24
source hit rate 0.958
Per-pair output is committed: assets/eval_retrieval.csv and
assets/eval_results.csv.
python -m eval.evaluate --validate # check the ground truth itself
python -m eval.evaluate --retrieval-only # stage A
python -m eval.evaluate --full # stage A + B, needs the API runningHit rate is saturated at 100% and is therefore a weak regression signal. Eight of the 24 answerable pairs are fault-code lookups that the deterministic lexical path cannot get wrong. MRR and context recall carry the real signal; a reviewer who reads a saturated metric as success is being misled, so it is stated here instead of buried.
Context recall is the metric that earned its place. Hit rate says the right chunk came back;
context recall says the retrieved text actually contains the answer. It inverted the reading of
the first Stage B run: three apparent false abstentions came back hit=True, rank=1, context_recall=0.0 — the model was correct to refuse, and loosening the prompt would have
traded a right answer for a hallucination. It also settled k: raising 5 → 8 improves context
recall with no loss in hit rate, MRR, or abstention accuracy.
Validating the ground truth caught three bad pairs before they inflated the numbers. The worst:
the keyword "12" matched "Nummer 1.2.1.1" in unrelated legal text, scoring a false-positive
context recall of 1.0. All numeric keywords are now bound to their unit. The honest figure is 93.2%,
not the 97.7% an earlier run reported.
One Stage B failure is infrastructure, not the system — a 503 from the LLM provider throttling
under sustained load, which is the graceful-degradation path working as designed. On requests that
completed, it is 22/23.
Latency figures are load artefacts. p50 14.2 s / p95 17.5 s across a full run, but the first four requests land in 0.37–0.53 s before the free tier throttles. Single-request and sustained-load latency differ by roughly 30×. Quote both or neither.
The production image ships no PyTorch. The embedding model is exported to ONNX and dynamically
quantized to int8 at build time, and a ~90-line Embeddings implementation replaces the
sentence-transformers stack: tokenize → encode → mean-pool over real tokens → L2 normalise.
| PyTorch fp32 | ONNX int8 | |
|---|---|---|
| compressed image | 934 MB | 251 MB |
| unpacked | 3.26 GB | 1.02 GB |
| Python venv layer | 449 MB | 131 MB |
| model weights layer | 438 MB | 74 MB |
| cold start → first answer | 17.5 s | 5.0 s |
| — of which model load | 15.2 s | 3.0 s |
Why quantization pays off so well here: the multilingual model carries the XLM-R vocabulary of 250k tokens, so the embedding matrix alone is 250 002 × 384 ≈ 96 M parameters — 384 MB of the 470 MB total. The transformer body is the small part. Dynamic quantization acts on the dominant term, which would not be true of a monolingual model of the same depth.
Quality was verified, not assumed. Comparing raw vectors against the PyTorch reference gives a cosine of 0.975 worst-case — a real perturbation, not rounding, and not on its own a reason to ship. So both backends built a full index and Stage A ran against each:
k hit rate MRR context recall
3 100% → 100% 0.979 → 0.979 88.6% → 86.4%
5 100% → 100% 0.979 → 0.979 88.6% → 90.9%
8 100% → 100% 0.979 → 0.979 93.2% → 93.2%
hit and rank are identical on all 72 per-pair rows, and at k=8 — the configured operating
point — all three metrics match exactly. Only three pairs shift at all, and all three sit in the one
topic where the two corpus sources genuinely overlap.
One thing got worse, and it belongs next to the wins: a full re-ingest takes 44.6 s against 26 s under PyTorch. Dynamic quantization dequantizes per operation at inference time, so bulk throughput is not what it optimises. Irrelevant in practice — ingestion runs once, at build time — but "ONNX made everything faster" would be false.
python scripts/export_onnx.py # exports + quantizes + verifies against PyTorch| Endpoint | Purpose |
|---|---|
POST /chat |
Ask a question. Returns the answer, ranked sources, and whether it abstained. |
GET /health |
Index facts read from the build manifest, cross-checked against the store on disk. |
GET /metrics |
Prometheus text: retrieval vs LLM latency split, chunks retrieved, abstentions, quota. |
POST /ingest |
Rebuilds the index. Enabled locally, 403 in production. |
curl -s -X POST localhost:8000/chat -H 'Content-Type: application/json' \
-d '{"question":"Was bedeutet der Fehlercode P0420?","language":"de"}'{
"answer": "Der Fehlercode P0420 bedeutet, dass der Katalysator-Wirkungsgrad unter dem Schwellenwert liegt (Bank 1).",
"sources": [{ "source": "obd2_fehlercodes.md", "section": "P0420 — Katalysator-Wirkungsgrad unter Schwellenwert (Bank 1)", "relevance": 0.4364 }],
"abstained": false,
"language": "de",
"model": "llama-3.1-8b-instant"
}Ask in English and the German sources still come back:
curl -s -X POST localhost:8000/chat -H 'Content-Type: application/json' \
-d '{"question":"When must the timing belt be replaced?","language":"en"}'
# → "The timing belt must be replaced every 90,000 to 180,000 km or every 5 to 8 years..."
# sources: wartungsintervalle.mdOut of scope, and note model is empty — no LLM call was made:
{ "answer": "Diese Information ist in den vorliegenden Dokumenten nicht enthalten.",
"abstained": true, "model": "" }Key handling. A caller may send their own key as X-Groq-Key; otherwise a capped demo allowance
is used, and past that the API returns 429 pointing at the bring-your-own-key path. Keys are never
logged, never persisted, and never echoed in a response — all three are asserted in the test suite.
When the provider's quota is exhausted, /chat returns 503 including the retrieved sources, so
the retrieval half stays visibly functional.
webhook → POST /ingest → POST /chat → Discord. Runs locally only; it is never deployed. The
committed export at n8n/workflow_export.json contains no credentials —
not even a credential reference — so importing it prompts you to attach your own.
That screenshot is the most useful one in this repo, because it is three behaviours in one frame: a
grounded maintenance answer, a fault-code answer resolved by the lexical path, and a refusal that
reports Modell: keins — Abbruch vor dem LLM-Aufruf. The sink says outright that no model was
invoked, which is the clearest evidence that the relevance floor is a real cost saving and not just
a phrase in a prompt.
The red run visible on the canvas is honest and left in place: a transient DNS failure resolving Discord from inside the container. Ingest and chat had both already succeeded in that run, and the next run 77 seconds later passed.
Live at https://docubot-pb.streamlit.app/, on a free tier, with no payment card and no billing account attached to this project anywhere.
The whole system runs as one Streamlit app with the RAG in-process: the same hybrid retrieval, the same relevance floor, the same bilingual prompt and abstention, called directly instead of over HTTP. What is reimplemented for the embedded mode is only the HTTP-shaped shell — key resolution, the demo allowance, the status codes. Everything that decides an answer stays in one module and is shared, because a divergence in the shell is cosmetic while a divergence in the chain would mean the deployed demo and the eval harness were measuring different systems.
The UI file is shared verbatim between the two modes. It gained exactly one seam, use_backend(),
which swaps the HTTP client for the in-process backend; the two return the same payload keys, and a
test asserts that parity. Branching inside each renderer is how the two modes would quietly drift.
The model and index cannot live in git — the quantized encoder is 118 MB, past GitHub's 100 MB
file ceiling. They ship as a single GitHub Release asset, fetched on first boot and verified
against a checksum committed in deploy/artifacts.json before anything is
extracted; a truncated download otherwise surfaces as an opaque protobuf parse error. Index and model
travel as one artifact under one checksum because they are only equivalent when they came from the
same export. Git LFS was the obvious alternative and is the wrong one: its free tier is 1 GB of
bandwidth a month, every cold boot would spend 81 MB of it, and the app would break after a handful
of restarts in a way that looks like a code bug.
| measured on the deployed app | |
|---|---|
| peak memory | 630 MB of the 1 GB limit (62%) |
| cold start to first answer, including model load | 5.6 s |
| warm answer | 2.7 s |
| artifact bundle fetched once per container | 81 MB compressed (139 MB raw) |
Stated plainly, because the alternative is implying otherwise: you cannot curl a deployed
/chat, /health or /metrics. The FastAPI service, its image, the build-time index, the eval
gate and CI publishing to a registry are all real and all verifiable by anyone with Docker — but
they are not hosted.
The API needs roughly 600 MB of RAM. Every remaining container host that takes no payment card
offers 512 MB. Run under a hard 512 MB cgroup cap, the service starts, answers /health, and is
then OOM-killed on the first query — exit 137. Where the memory goes:
| step | peak RSS |
|---|---|
| python + numpy + ONNX Runtime + tokenizers | 44 MB |
| inference session created | 260 MB |
| first inference, batch of one | 467 MB |
| + vector store imported | 502 MB |
The +207 MB on the first inference is dynamic int8 quantization dequantizing weights per operation — the same mechanism behind the one thing the ONNX migration made worse (re-ingest is slower). It is inherent, not a tunable: disabling the CPU memory arena moved it by 0.3 MB. Serving a public demo at 98% of a ceiling that OOM-kills would be worse than not serving one — a reviewer clicks the link and gets a dead app.
So the deployment went to the only free Python runtime left with real headroom, which runs Streamlit apps and nothing else. Hence one app, RAG in-process.
The API still runs, in one command:
docker run -p 8000:8000 ghcr.io/prennoy99/docubot-api:latest
# → /health, /chat, /metrics on :8000; /ingest returns 403, as it does in productionThat image is built and published by CI on every push to main, with ingestion and the retrieval
eval gate running inside the build.
The image was designed for Cloud Run and is still deployable there unchanged. It is not deployed there because Cloud Run's Always Free tier requires a billing account to exist, and creating one requires a payment card. Zero cost is a hard constraint on this project rather than a preference, and Google offers no hard spending cap — a budget alert is an email, not a limit. Given a real financial limit, an architecture that cannot bill beats one that merely shouldn't.
gcloud run deploy docubot-api \
--memory=2Gi --cpu=1 --cpu-boost --concurrency=8 --timeout=120 \
--min-instances=0 --max-instances=1 \
--set-secrets=GROQ_API_KEY=groq-key:latest \
--set-env-vars=ALLOW_INGEST=false--min-instances=0— mandatory, and the single largest cost risk in the project. One always-on instance burns roughly 15× the free vCPU allowance. This is why cold starts are unavoidable and why shrinking the image was worth the effort.--concurrency=8— the default of 80 lets one instance thrash inference threads across dozens of simultaneous requests.--cpu-boost— free, and roughly halves model load on a cold start.--memory=2Gi— headroom; 1 GiB risks an OOM-kill that surfaces as an opaque 503, which is exactly what the 512 MB measurement above demonstrates.- Cloud Run cannot pull from GHCR — it accepts only Artifact Registry or GCR. The registry copy would be the deploy source and GHCR the CI artifact, which are different things.
Two portability details fell out of designing for that host and cost nothing to keep: the container
reads $PORT rather than hard-coding 8000, and the runtime stage already runs as a non-root uid.
Designing around one host's constraints is what made the image portable to hosts that were never
under consideration.
Code comments reference decisions by ID. The ones that shaped the system most:
| ID | Decision | Why |
|---|---|---|
| D-01 | Index built at image-build time, shipped read-only | Cloud Run's filesystem is ephemeral; a runtime-only ingest means an empty store after every cold start |
| D-03 | Chunking budgeted in tokens, per source type | German prose runs 4.01 chars/token but the fault-code reference 3.43 — no single character budget fits both, and the model truncates silently |
| D-07 | Explicit German abstention instruction in the prompt | Without it the model answers from parametric knowledge: plausible, unsourced, sometimes wrong |
| D-09 | language selects the answer language only; retrieval is always cross-lingual |
The embedding model is trained for cross-lingual alignment, so an English question genuinely retrieves the right German chunks |
| D-14 | Two-stage eval: deterministic retrieval, then end-to-end | Stage A is free and gates every build; Stage B costs real API calls and runs on demand |
| D-15 | Stage A is a Docker build gate | Makes the eval a quality gate rather than a script someone remembers to run |
| D-21 | Production runs ONNX int8; local development runs PyTorch | See Optimization. PyTorch stays locally as the reference implementation the ONNX path is measured against |
| D-30 | A retrieval relevance floor gates abstention, not the prompt alone | Cheaply refuses obviously off-topic questions with no LLM call — but see the limitation below, it is necessary and not sufficient |
| D-31 | Corpus vocabulary is part of retrieval design | Record headings read ### Fehlercode P0420 — …, not ### P0420 — …, because users search with the word "Fehlercode" |
| D-32 | Hybrid retrieval: deterministic lexical path for fault codes | The single biggest finding in the project — see below |
Pure dense retrieval scored 0/6 top-1 at record level. "Was bedeutet der Fehlercode P0420?"
returned the P0700 record — and so did every other code query. Two inherent causes: this is a
paraphrase model that encodes sentence semantics, and an identifier like P0420 fragments into
subword pieces carrying almost no distinguishing signal; meanwhile the P0700 record happens to
describe fault codes generically, so it matches the shape of the question better than any specific
record matches its content. Exact identifier lookup is the classic failure mode of dense retrieval,
and no amount of chunk tuning fixes it.
After adding an exact metadata-match path for [PBCU]\d{4}: 8/8 top-1, invented codes correctly
find no match and fall through to the floor, and non-code questions are unregressed.
The process lesson is worth as much as the fix: an earlier spot-check scored the source file and reported 10/11, which hid a total record-level failure — every code query did land in the right file, just on the wrong record. The eval harness scores at record level for exactly this reason.
These are measured and documented rather than quietly worked around.
The relevance floor does not separate hard negatives, and must not be raised to try. Across five out-of-scope questions including deliberately hard ones — in-domain, German, plausibly present but absent — scores were 0.653, 0.632, 0.504, 0.232, 0.167, while the lowest in-scope score was 0.492. The ranges overlap, so no threshold separates them. Raising the floor to catch 0.653 would reject valid questions. The floor is necessary and not sufficient; questions that are in-domain but absent can only be refused by the model's own grounding. Two independent defences, each covering what the other cannot.
A chunk can be the right chunk and still not contain the answer. Asking which defect classes
exist at a vehicle inspection retrieves the correct file and the correct heading at rank 1 — but
that 120-token chunk holds only the heading and its introductory sentence. The four actual classes
fell into sibling chunks, which lost the remaining slots to legal text. So hit=True, rank=1, context_recall=0, and the model correctly abstained. Fixing it means section-aware chunking, which
would change the index and invalidate every number published above; that is a deliberate trade, not
an oversight.
One eval pair fails on German lexical ambiguity. The corpus plainly states the 12-month inspection interval for taxis, but the question asks "In welchem Abstand…" — and Abstand means both a temporal interval and a physical distance. The legal excerpt discusses distances between inspection bays, so dense retrieval resolves the wrong sense. Adding vocabulary improved the retrieved context without fixing the answer.
Cross-lingual retrieval is phrasing-sensitive. "How often must brake fluid be replaced?" retrieves correctly; "How often must brake fluid be replaced and why?" does not. The model then correctly abstained on context that genuinely lacked the answer — so the abstention was right while retrieval was wrong. The claim holds, but weakly and inconsistently in English.
The authored German content has not been reviewed by a domain expert. The four hand-written
documents are original technical prose based on generally documented, non-proprietary knowledge, and
they are fit for demonstrating a retrieval system. They are not a reference for actual vehicle
work: limit values, intervals and dates have not been checked against current regulations. See
documents/PROVENANCE.md. This does not affect any number above — the
eval measures faithfulness to the corpus, not to reality.
pytest -q # 68 tests, ~2 sContract tests only: the retriever and the LLM are swapped out via FastAPI dependency overrides, so the suite imports no model, opens no database, and makes no network call. Two of the tests exist because casual testing already missed the bug once — the English-prompt test asserts on the rendered prompt rather than the answer, and the hybrid-retrieval test reproduces the exact broken ranking so that deleting the lexical path fails the test.
A separate set of tests guards the dependency files themselves, so the torch-free production set cannot silently regain PyTorch — which would not break anything visible, it would just start costing money.
CI runs lint and the suite on every push, and builds and publishes both images on main — with
ingestion and the retrieval eval gate running inside the API build, so a retrieval regression fails
the build rather than shipping.
The fast gate installs a torch-free subset of the requirements, and the absence is the
enforcement. A tokenizer import buried in a LangChain module pulls in the entire PyTorch stack
whenever it happens to be installed, so if a lazy import in app/ ever regresses to a module-level
one, CI fails with an ImportError instead of quietly getting minutes slower. It is also a free
rehearsal of the production image, which genuinely has no PyTorch in it.
| Corpus, ingestion, RAG chain, API | done |
| Streamlit UI, evaluation harness | done |
| Tests, CI | done, green on every push |
| Docker images, local compose stack | done |
| Automation workflow | done, verified end to end |
| ONNX int8 optimization | done, measured |
| Live deployment | done — docubot-pb.streamlit.app; no hosted API endpoint, see Deployment |
| Domain-expert review of the German corpus | pending |
LangChain (LCEL) · ChromaDB · Groq / Llama 3.1 8B · sentence-transformers multilingual MiniLM · ONNX Runtime with int8 dynamic quantization · FastAPI · Streamlit · Docker & Compose · GitHub Actions → GHCR · Google Cloud Run · n8n · Prometheus instrumentation · pytest · ruff
PdM-Lite — predictive-maintenance API on the same deployment stack (FastAPI, Docker, Cloud Run, GitHub Actions). DocuBot deliberately does not repeat its Prometheus/Grafana monitoring stack; the observability here is a metrics endpoint with RAG-specific series and no extra containers.
Source code is MIT. NOTICE records what the MIT licence does not cover.
The corpus is not uniformly licensed, and documents/PROVENANCE.md
records the origin and legal basis of every file:
- Four authored German documents — written for this project, MIT along with the code.
- A 52-page StVZO excerpt — official German legislation, outside copyright under § 5 UrhG. Page selection only; both the excerpt and the full original are checksummed in the provenance file so anyone can verify the pages are unmodified.
No OEM workshop manual content is included. Being findable online is not a licence to redistribute, and this repository's index is baked into a published container image.





