LLM Performance Dashboard

Daily quality benchmarks on Google Cloud Vertex AI ·

Latest Results

Overall Quality Index

Mean of all benchmark scores per model, latest run.

Latest Scores by Benchmark

Model Comparison

Latest run vs. previous run (Δ shows change).

Radar

Run Details

About the Benchmarks

Each benchmark uses a fixed, frozen subset of questions so daily scores stay comparable over time. All scoring is fully programmatic (no LLM judge): exact-match, code execution, or rule checks.

Questions & Expected Answers