Benchmarking Model Prompt Tests
Introducing the first benchmark runner for model-prompt-tests: a small harness for running prompt suites across model providers, scoring outputs against rubrics, and publishing comparable reports.
Series
Recurring model comparisons on the model-prompt-tests harness: shared methodology, then each X vs Y run with scores, caveats, and links back to the suite.
5 articles in reading order
Introducing the first benchmark runner for model-prompt-tests: a small harness for running prompt suites across model providers, scoring outputs against rubrics, and publishing comparable reports.
A benchmark writeup comparing Grok 4.5 and Sonnet 5 across 13 practical prompt tests, with methodology, caveats, and raw artifacts preserved in model-prompt-tests.
A model-prompt-tests run that put a local Gemma candidate next to Grok 4.5, Sonnet 5, and GPT-5.5: peer scores, latency reality, and what a high local score does not prove.
DeepSeek V4 Pro scored 4.50 to Flash's 4.38 in one practical prompt-test batch, but cost about 21 times as much for candidate answers.
GLM-5.3-Flash led this four-model benchmark at 4.83 and had the lowest completed-response latency, while Qwen3.8-Flash had the lowest recorded cost.