The 5090 receipt run
I fed a 27B model all of Frankenstein in one prompt on a single RTX 5090. It took four undocumented fixes, and every number is sourced to a log file.
Local LLM inference, measured on hardware I own. Every run comes with its methodology and run log, numbers you can reproduce.
I fed a 27B model all of Frankenstein in one prompt on a single RTX 5090. It took four undocumented fixes, and every number is sourced to a log file.
One vLLM server, three configs, the same 24 prompts across code, creative and RAG workloads. Determinism checked rather than assumed.