Four Flash Models on Practical Prompt Tests
GLM-5.3-Flash led this four-model benchmark at 4.83 and had the lowest completed-response latency, while Qwen3.8-Flash had the lowest recorded cost.
Multi-agent systems · Model evaluation
Building multi-agent systems, frontier model evaluation tooling, and honest findings from work that has to survive outside the demo.
bshp.io · research · systems · shipping
I build multi-agent systems and tools for evaluating language models. Most of it starts as something I needed myself; I open-source or write up the useful bits.
I'm biased toward small systems I can run and measure - demos are easy, keeping something useful is harder. The projects and articles here are notes from that process: what worked, what broke, and what I'd try next.
Content-collection-first llms.txt for Astro
Technology stack: TypeScript · Astro · Zod
Claude Code review and rescue workflows backed by Grok Build
Technology stack: JavaScript · Shell · Claude Code Plugin SDK · Grok Build CLI
Unified AI usage dashboard across coding agents and local inference
Technology stack: React · TypeScript · Vite · Tailwind · Express · SQLite · Bun · TanStack Query · SSE
GLM-5.3-Flash led this four-model benchmark at 4.83 and had the lowest completed-response latency, while Qwen3.8-Flash had the lowest recorded cost.
I reconstructed one LHR to SFO flight from Mission Control, git, and Linear: 5 hours 46 minutes of active work, 27 merged PRs, and several limits the logs cannot explain.
What it took to use a full $30 SuperGrok weekly allowance: 31 coding-agent sessions, 9.66 million input tokens, and a lot of deliberate work.