The bench.

Local LLM inference, measured on hardware I own. Every run comes with its methodology and run log, numbers you can reproduce.

rig: NVIDIA GeForce RTX 5090 32 GB · AMD Ryzen 9 9950X3D · 64 GB DDR5

At a glance

1,632
Commits, past year
every project, deduped by hash
103,459
Tokens in one prompt
all of Frankenstein, one 5090
96.8%
gsm8k, strict match
27B NVFP4 quant · n=250
350
Fastest decode, tok/s
Qwen3-32B + n-gram speculation

Runs

The 5090 receipt run

I fed a 27B model all of Frankenstein in one prompt on a single RTX 5090. It took four undocumented fixes, and every number is sourced to a log file.

Read more Aug 31, 2026

Qwen3-32B, three configs

One vLLM server, three configs, the same 24 prompts across code, creative and RAG workloads. Determinism checked rather than assumed.

Read more Jul 21, 2026