Skip to content
#

llm-evaluation-metrics

Here are 22 public repositories matching this topic...

Label every claim in AI output as (u) given, (m) checked or (g) generated, in plain text, so a guess can't quietly become a "fact" when text passes between people and AI agents. Includes a swarm test (the spoke and wheel test), a parser and a gate. Early findings.

  • Updated Sep 26, 2026
  • HTML

Add this topic to your repo

To associate your repository with the llm-evaluation-metrics topic, visit your repo's landing page and select "manage topics."

Learn more