The LLM Evaluation Framework
-
Updated
Sep 28, 2026 - Python
The LLM Evaluation Framework
[NeurIPS D&B '25] The one-stop repository for LLM unlearning
LangFair is a Python library for conducting use-case level LLM bias and fairness assessments
[ACL'24] A Knowledge-grounded Interactive Evaluation Framework for Large Language Models
Run a prompt against all, or some, of your models running on Ollama. Creates web pages with the output, performance statistics and model info. All in a single Bash shell script.
Measure of estimated confidence for non-hallucinative nature of outputs generated by Transformer-based Language Models.
Create an evaluation framework for your LLM based app. Incorporate it into your test suite. Lay the monitoring foundation.
Tools for systematic large language model evaluations
Your prompts need tests too. Run prompts against real datasets, score outputs with LLM judges, version everything, and compare runs to see what got better.
insideLLMs is a Python library and CLI for comparing LLM behaviour across models using shared probes and datasets. The harness is deterministic by design, so you can store run artefacts and reliably diff behaviour in CI.
VerifyAI is a verification harness for testing, auditing, and evaluating responses from AI assistants and coding agents.
In this we evaluate the LLM responses and find accuracy
Evaluating a RAG pipeline with DeepEval — Answer Relevancy, Faithfulness, Contextual Precision, Recall & Relevancy scored by an LLM judge
Label every claim in AI output as (u) given, (m) checked or (g) generated, in plain text, so a guess can't quietly become a "fact" when text passes between people and AI agents. Includes a swarm test (the spoke and wheel test), a parser and a gate. Early findings.
Offical implementation for Spectral Scaling Laws (EMNLP 2025)
Deepeval AI context pack for Claude Code, Codex, Cursor, and Aider: AGENTS.md, CLAUDE.md, prompts, evals, pitfalls, and verification notes for confident-ai/deepeval.
This repo is for an streamlit application that provides a user-friendly interface for evaluating large language models (LLMs) using the beyondllm package.
Benchmark LLM accuracy, latency, cost, and hallucination rates across models with this open-source evaluation suite.
To associate your repository with the llm-evaluation-metrics topic, visit your repo's landing page and select "manage topics."