Skip to content
Navigation menu
Search
Powered by Algolia
Search
Log in
Create account
DEV Community
Close
Local Model Benchmarks Series' Articles
Back to Mihai Perdum's Series
The best 4-bit format in MLX is the one its own converter sets up to lose
Mihai Perdum
Mihai Perdum
Mihai Perdum
Follow
Sep 18
The best 4-bit format in MLX is the one its own converter sets up to lose
#
aicoding
Comments
Add Comment
7 min read
MLX memory ceiling on a 96 GB M3 Ultra
Mihai Perdum
Mihai Perdum
Mihai Perdum
Follow
Sep 19
MLX memory ceiling on a 96 GB M3 Ultra
#
aicoding
#
localai
Comments
Add Comment
15 min read
llama.cpp vs MLX on Qwen3.6-27B: MTP is 1.04x here, not 1.85x
Mihai Perdum
Mihai Perdum
Mihai Perdum
Follow
Sep 20
llama.cpp vs MLX on Qwen3.6-27B: MTP is 1.04x here, not 1.85x
#
aicoding
#
localai
Comments
Add Comment
18 min read
Ling 3.0 Flash won't load on a Mac Studio. Qwen3-Coder-Next does, at 73 tok/s.
Mihai Perdum
Mihai Perdum
Mihai Perdum
Follow
Sep 20
Ling 3.0 Flash won't load on a Mac Studio. Qwen3-Coder-Next does, at 73 tok/s.
#
aicoding
#
localai
Comments
Add Comment
18 min read
MXFP8 vs Q8: 10x the weight error, 1% the perplexity
Mihai Perdum
Mihai Perdum
Mihai Perdum
Follow
Sep 20
MXFP8 vs Q8: 10x the weight error, 1% the perplexity
#
aicoding
#
localai
Comments
Add Comment
14 min read
MLX vs GGUF on Apple Silicon: Benchmarking the Same Local Model Two Ways
Mihai Perdum
Mihai Perdum
Mihai Perdum
Follow
Sep 20
MLX vs GGUF on Apple Silicon: Benchmarking the Same Local Model Two Ways
#
ai
#
mlx
#
benchmark
#
localai
Comments
Add Comment
7 min read
Qwen3-Coder-Next + MXFP8: The 128GB Local LLM That Runs Predictably
Mihai Perdum
Mihai Perdum
Mihai Perdum
Follow
Sep 20
Qwen3-Coder-Next + MXFP8: The 128GB Local LLM That Runs Predictably
#
aicoding
#
localai
#
llm
#
mlx
Comments
Add Comment
25 min read
We're a place where coders share, stay up-to-date and grow their careers.
Log in
Create account