A few prompts that I am storing in a repo for the purpose of running controlled experiments comparing and benchmarking different LLMs for defined use-cases
-
Updated
Dec 4, 2024 - Python
A few prompts that I am storing in a repo for the purpose of running controlled experiments comparing and benchmarking different LLMs for defined use-cases
Prompt Evaluator & AI Configuration Advisor — An LLM prompt benchmarking platform to analyze prompts, benchmark models (OpenAI, Anthropic, Gemini, Ollama), track via MLflow, and get cost/latency-optimized recommendations.
To associate your repository with the prompt-benchmarking topic, visit your repo's landing page and select "manage topics."