Statistical inference for AI evaluations, from model comparisons to corrections for LLM judge bias, optimized for small sample sizes. All defaults battle-tested in Monte Carlo simulations.
-
Updated
Sep 24, 2026 - Python
Statistical inference for AI evaluations, from model comparisons to corrections for LLM judge bias, optimized for small sample sizes. All defaults battle-tested in Monte Carlo simulations.
Blind A/B testing for AI agent skills, .cursorrules, CLAUDE.md, and system prompts. Detects placebo and harmful instructions.
Measure prompt and skill improvements with blind A/B comparison.
Another day, another Awesome List repo. A comprehensive list of Chainforge-related content
Local-first prompt library and version-control system for AI prompts (Desktop App, CLI, MCP Server, and AI execution engine).
Official implementation for "GLaPE: Gold Label-agnostic Prompt Evaluation and Optimization for Large Language Models" (stay tuned & more will be updated)
The prompt engineering, prompt management, and prompt evaluation tool for Python
Who evaluates the evaluator? Judicator audits LLM-as-a-Judge systems for 7 documented bias types. Zero config. Works with any LLM.
pi extension for fixed-task-set eval runs and prompt/system comparisons with reproducible reports
The prompt engineering, prompt management, and prompt evaluation tool for TypeScript, JavaScript, and NodeJS.
A project to take a suboptimal prompt from Langsmith, enhance it, submit it again, and then reevaluate the results. #LangSmith #PromptEngineer
Production-grade prompt engineering, migration audits, and eval gates for GPT-5.6 Sol, Terra, and Luna.
Prompt Architect: reusable prompt patterns, approval-gated workflows, and verification checklists for AI-assisted development.
A Simple Prompt Optimization Using 3 different algorithms for testing.
Audit and score Codex skills with a transparent, evidence-based custom rubric
Benchmark and continuously improve your Superwhisper custom modes against your own voice recording history.
A pure-Python dependency graph for composable LLM prompts — section-level change tracking that separates the full dependency reach of an edit from the smaller set that actually needs re-evaluation.
Declarative LLM prompt evaluation harness, easy to use UI, all in one binary.
A lightweight CLI tool for evaluating LLM prompts. Run prompts side-by-side to instantly compare token usage, tone, reading level, and API costs. Natively supports OpenAI, Anthropic, Gemini, Groq, and Ollama. Perfect for safely refactoring prompts, tuning bot personalities, or proving that "be concise" actually lowers your bill.
A few prompts that I am storing in a repo for the purpose of running controlled experiments comparing and benchmarking different LLMs for defined use-cases
To associate your repository with the prompt-evaluation topic, visit your repo's landing page and select "manage topics."