-
Notifications
You must be signed in to change notification settings - Fork 0
Expand file tree
/
Copy path.env.example
More file actions
61 lines (53 loc) · 3.01 KB
/
Copy path.env.example
File metadata and controls
61 lines (53 loc) · 3.01 KB
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
# DocuBot — copy to .env and fill in. .env is gitignored; never commit a real key.
# cp .env.example .env
# --- Groq (required from M3 onward) ---
# Free key, no credit card: https://console.groq.com
# On Cloud Run this is injected from Secret Manager, never as a plain env var (D-10).
# Locally it funds your own testing; in the deployment it funds the capped demo
# allowance, and callers may override it per-request with an X-Groq-Key header.
GROQ_API_KEY=
GROQ_MODEL=llama-3.1-8b-instant
# --- ingestion (D-02) ---
# true locally and in docker compose. Set false on Cloud Run, where the index is
# baked into the image at build time and runtime writes would be discarded.
ALLOW_INGEST=true
# --- embedding runtime (D-21) ---
# The embedding *model* is locked (§3); how it is executed is not.
# torch -- HuggingFaceEmbeddings, fp32. The local default: boring and known-good, so
# a retrieval bug is obviously ours rather than the runtime's.
# onnx -- onnxruntime with an int8 model. What the API image runs, and what makes the
# project free: 934 MB -> 251 MB compressed, cold start 17.5 s -> 5.0 s, with
# Stage A hit rate and MRR identical at every k (measured M9a).
# Using onnx locally first requires `python scripts/export_onnx.py` (~2 min, needs torch);
# the Docker build does its own export in a stage that is thrown away.
EMBEDDING_BACKEND=torch
ONNX_MODEL_DIR=models/onnx
# 0 lets onnxruntime choose, which is right on a multi-core dev box. The image sets 1,
# because Cloud Run gives the container 1 vCPU (D-23).
ONNX_THREADS=0
# --- chunking, measured in TOKENS not characters (D-03) ---
# The embedding model truncates at 128 tokens and truncation is silent. German prose
# runs ~4.0 chars/token but the fault-code reference ~3.4, so no single character
# budget is safe for both -- hence a token budget. 120 leaves room for the 2 special
# tokens plus margin. Do not raise above 126.
CHUNK_TOKENS=120
CHUNK_OVERLAP_TOKENS=24
# --- retrieval ---
# Raised 5 -> 8 at M5 on eval evidence: hit rate (100%) and MRR (0.979) were identical
# at k=3/5/8, but context recall -- whether the retrieved text actually contains the
# answer -- rose 88.6% -> 93.2%. See `python -m eval.evaluate --retrieval-only`.
RETRIEVAL_K=8
# Below this cosine relevance, /chat abstains with no LLM call (D-30). Necessary but
# NOT sufficient: measured out-of-scope scores reach 0.653 while in-scope drops to
# 0.492, so the ranges overlap and no threshold separates them. Raising this would
# reject valid questions; hard negatives are caught by the model's grounding instead.
RELEVANCE_FLOOR=0.35
# --- rate limiting and demo quota (D-12) ---
# Counters are in-process and reset on cold start, so these caps are approximate
# by design. That is accepted; do not add Redis for it.
RATE_LIMIT_PER_MINUTE=10
DEMO_QUOTA_PER_DAY=50
# --- frontend -> API ---
# The Streamlit service calls the API server-side, so no CORS config is needed.
# docker compose: http://api:8000 Cloud Run: the docubot-api service URL
API_URL=http://localhost:8000