DEV Community

Jahn profile picture

Jahn

Inference performance engineering on NVIDIA Blackwell. Conatus AI.

Joined Joined on 
Name the Blackwell serving cell you are actually in

Name the Blackwell serving cell you are actually in

Comments
5 min read
DGX Spark (GB10) memory sizing for LLM serving: the numbers

DGX Spark (GB10) memory sizing for LLM serving: the numbers

Comments
7 min read
DGX Spark (GB10) bare-metal vLLM: the install that works, two landmines, measured timings

DGX Spark (GB10) bare-metal vLLM: the install that works, two landmines, measured timings

Comments
2 min read
SGLang outputs endless repetition on NVFP4 models: the FP8 lm_head bug

SGLang outputs endless repetition on NVFP4 models: the FP8 lm_head bug

2
Comments 1
3 min read
The sm_120 shared-memory cliff: why FP8 KV cache crashes vLLM on workstation Blackwell

The sm_120 shared-memory cliff: why FP8 KV cache crashes vLLM on workstation Blackwell

Comments
3 min read
Did FP8 make the model dumber? A per-prompt regression check for quantized serving

Did FP8 make the model dumber? A per-prompt regression check for quantized serving

Comments
3 min read
Qwen3-8B on workstation Blackwell: vLLM vs SGLang vs llama.cpp, plus an FP8 pass

Qwen3-8B on workstation Blackwell: vLLM vs SGLang vs llama.cpp, plus an FP8 pass

1
Comments 1
3 min read
loading...