DEV Community

#benchmarking

Posts

đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.
Eggs, Cholesterol, and GPU Flags

Eggs, Cholesterol, and GPU Flags

Comments
3 min read
Measure the Binary You Run

Measure the Binary You Run

Comments 1
2 min read
Same Model, 13.3% to 38.3%

Same Model, 13.3% to 38.3%

Comments 3
7 min read
The sleep loop is the tell: agents that pay per action optimize to do nothing

The sleep loop is the tell: agents that pay per action optimize to do nothing

Comments 1
2 min read
I benchmarked Dragonfly vs Redis vs Valkey. First, let me show you how I kept it honest.

I benchmarked Dragonfly vs Redis vs Valkey. First, let me show you how I kept it honest.

1
Comments
6 min read
Our benchmark was leaking the answers to the model. The numbers looked fine the whole time.

Our benchmark was leaking the answers to the model. The numbers looked fine the whole time.

Comments
3 min read
Fairness Under the Microscope: Why HY3 Beats Nemotron 3 Ultra on lforla's Bias Stereotypes Audit

Fairness Under the Microscope: Why HY3 Beats Nemotron 3 Ultra on lforla's Bias Stereotypes Audit

1
Comments
3 min read
The Compiler Got 5% Slower. The Benchmark Called It a 10% Regression a Quarter of the Time.

The Compiler Got 5% Slower. The Benchmark Called It a 10% Regression a Quarter of the Time.

Comments 1
6 min read
I benchmarked 5 managed graph databases — and the "obvious" winner changed depending on what I measured

I benchmarked 5 managed graph databases — and the "obvious" winner changed depending on what I measured

Comments
5 min read
The Model Reading My Benchmark Mattered More Than the Memory System Did

The Model Reading My Benchmark Mattered More Than the Memory System Did

Comments
7 min read
I benchmarked CognoDB against four other graph databases. The most interesting result had nothing to do with CognoDB.

I benchmarked CognoDB against four other graph databases. The most interesting result had nothing to do with CognoDB.

Comments 1
5 min read
I benchmarked 5 graph databases. The first four hours measured the Indian Ocean.

I benchmarked 5 graph databases. The first four hours measured the Indian Ocean.

Comments
6 min read
Don't Trust a New Model's Benchmarks Until You Run Your Own 30-Minute Smoke Test

Don't Trust a New Model's Benchmarks Until You Run Your Own 30-Minute Smoke Test

Comments
5 min read
Before You Adopt MiniMax H3, Run a Twenty-Minute Model Audit

Before You Adopt MiniMax H3, Run a Twenty-Minute Model Audit

Comments
3 min read
DeepSeek Harness: What "Everything is a Plugin" Actually Means for Agent Frameworks

DeepSeek Harness: What "Everything is a Plugin" Actually Means for Agent Frameworks

Comments
2 min read
đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.