DEV Community

Local Model Benchmarks Series' Articles

Back to Mihai Perdum's Series
The best 4-bit format in MLX is the one its own converter sets up to lose
Cover image for The best 4-bit format in MLX is the one its own converter sets up to lose

The best 4-bit format in MLX is the one its own converter sets up to lose

Comments
7 min read
MLX memory ceiling on a 96 GB M3 Ultra
Cover image for MLX memory ceiling on a 96 GB M3 Ultra

MLX memory ceiling on a 96 GB M3 Ultra

Comments
15 min read
llama.cpp vs MLX on Qwen3.6-27B: MTP is 1.04x here, not 1.85x
Cover image for llama.cpp vs MLX on Qwen3.6-27B: MTP is 1.04x here, not 1.85x

llama.cpp vs MLX on Qwen3.6-27B: MTP is 1.04x here, not 1.85x

Comments
18 min read
Ling 3.0 Flash won't load on a Mac Studio. Qwen3-Coder-Next does, at 73 tok/s.
Cover image for Ling 3.0 Flash won't load on a Mac Studio. Qwen3-Coder-Next does, at 73 tok/s.

Ling 3.0 Flash won't load on a Mac Studio. Qwen3-Coder-Next does, at 73 tok/s.

Comments
18 min read
MXFP8 vs Q8: 10x the weight error, 1% the perplexity
Cover image for MXFP8 vs Q8: 10x the weight error, 1% the perplexity

MXFP8 vs Q8: 10x the weight error, 1% the perplexity

Comments
14 min read
MLX vs GGUF on Apple Silicon: Benchmarking the Same Local Model Two Ways
Cover image for MLX vs GGUF on Apple Silicon: Benchmarking the Same Local Model Two Ways

MLX vs GGUF on Apple Silicon: Benchmarking the Same Local Model Two Ways

Comments
7 min read
Qwen3-Coder-Next + MXFP8: The 128GB Local LLM That Runs Predictably
Cover image for Qwen3-Coder-Next + MXFP8: The 128GB Local LLM That Runs Predictably

Qwen3-Coder-Next + MXFP8: The 128GB Local LLM That Runs Predictably

Comments
25 min read