How I Got Qwen3.8-27B from 18.66 to 192.40 tok/s on One L40S
It started with our internal model feeling way too slow and turned into a rabbit hole through FP8, Marlin, MTP, DFlash2 and a custom vLLM build.
Aug 31, 202613 min read83

Search for a command to run...