Skip to content
Navigation menu
Search
Powered by Algolia
Search
Log in
Create account
DEV Community
Close
#
benchmarking
Follow
Hide
Posts
Left menu
đź‘‹
Sign in
for the ability to sort posts by
relevant
,
latest
, or
top
.
Right menu
Eggs, Cholesterol, and GPU Flags
Michael Brewer
Michael Brewer
Michael Brewer
Follow
Sep 8
Eggs, Cholesterol, and GPU Flags
#
llm
#
benchmarking
#
gpu
Comments
Add Comment
3 min read
Measure the Binary You Run
Michael Brewer
Michael Brewer
Michael Brewer
Follow
Sep 8
Measure the Binary You Run
#
llm
#
benchmarking
#
devops
Comments
1
 comment
2 min read
Same Model, 13.3% to 38.3%
Harrison Guo
Harrison Guo
Harrison Guo
Follow
Sep 8
Same Model, 13.3% to 38.3%
#
ai
#
machinelearning
#
benchmarking
#
llm
Comments
3
 comments
7 min read
The sleep loop is the tell: agents that pay per action optimize to do nothing
Cole Halton
Cole Halton
Cole Halton
Follow
Sep 7
The sleep loop is the tell: agents that pay per action optimize to do nothing
#
aiagents
#
evaluation
#
llm
#
benchmarking
Comments
1
 comment
2 min read
I benchmarked Dragonfly vs Redis vs Valkey. First, let me show you how I kept it honest.
Sunny Sahijwani
Sunny Sahijwani
Sunny Sahijwani
Follow
Sep 7
I benchmarked Dragonfly vs Redis vs Valkey. First, let me show you how I kept it honest.
#
redis
#
database
#
performance
#
benchmarking
1
 reaction
Comments
Add Comment
6 min read
Our benchmark was leaking the answers to the model. The numbers looked fine the whole time.
AIOil Security Shield
AIOil Security Shield
AIOil Security Shield
Follow
Sep 2
Our benchmark was leaking the answers to the model. The numbers looked fine the whole time.
#
security
#
ai
#
benchmarking
#
opensource
Comments
Add Comment
3 min read
Fairness Under the Microscope: Why HY3 Beats Nemotron 3 Ultra on lforla's Bias Stereotypes Audit
RESK
RESK
RESK
Follow
Sep 1
Fairness Under the Microscope: Why HY3 Beats Nemotron 3 Ultra on lforla's Bias Stereotypes Audit
#
ai
#
llm
#
fairness
#
benchmarking
1
 reaction
Comments
Add Comment
3 min read
The Compiler Got 5% Slower. The Benchmark Called It a 10% Regression a Quarter of the Time.
Panagiotis Gkilis
Panagiotis Gkilis
Panagiotis Gkilis
Follow
Sep 4
The Compiler Got 5% Slower. The Benchmark Called It a 10% Regression a Quarter of the Time.
#
quantum
#
testing
#
benchmarking
#
python
Comments
1
 comment
6 min read
I benchmarked 5 managed graph databases — and the "obvious" winner changed depending on what I measured
Ahmed Amer
Ahmed Amer
Ahmed Amer
Follow
Aug 27
I benchmarked 5 managed graph databases — and the "obvious" winner changed depending on what I measured
#
database
#
graphdatabase
#
benchmarking
Comments
Add Comment
5 min read
The Model Reading My Benchmark Mattered More Than the Memory System Did
Pranab Sarkar
Pranab Sarkar
Pranab Sarkar
Follow
Aug 25
The Model Reading My Benchmark Mattered More Than the Memory System Did
#
ai
#
llm
#
benchmarking
#
opensource
Comments
Add Comment
7 min read
I benchmarked CognoDB against four other graph databases. The most interesting result had nothing to do with CognoDB.
Sachin
Sachin
Sachin
Follow
Aug 21
I benchmarked CognoDB against four other graph databases. The most interesting result had nothing to do with CognoDB.
#
webdev
#
database
#
devops
#
benchmarking
Comments
1
 comment
5 min read
I benchmarked 5 graph databases. The first four hours measured the Indian Ocean.
Abhinav Bahuguna
Abhinav Bahuguna
Abhinav Bahuguna
Follow
Aug 21
I benchmarked 5 graph databases. The first four hours measured the Indian Ocean.
#
showdev
#
database
#
benchmarking
#
devops
Comments
Add Comment
6 min read
Don't Trust a New Model's Benchmarks Until You Run Your Own 30-Minute Smoke Test
Dakota Wu
Dakota Wu
Dakota Wu
Follow
Aug 17
Don't Trust a New Model's Benchmarks Until You Run Your Own 30-Minute Smoke Test
#
ai
#
programming
#
benchmarking
#
opensource
Comments
Add Comment
5 min read
Before You Adopt MiniMax H3, Run a Twenty-Minute Model Audit
Casey Li
Casey Li
Casey Li
Follow
Aug 14
Before You Adopt MiniMax H3, Run a Twenty-Minute Model Audit
#
ai
#
opensource
#
programming
#
benchmarking
Comments
Add Comment
3 min read
DeepSeek Harness: What "Everything is a Plugin" Actually Means for Agent Frameworks
Cole Halton
Cole Halton
Cole Halton
Follow
Aug 13
DeepSeek Harness: What "Everything is a Plugin" Actually Means for Agent Frameworks
#
deepseek
#
aiagents
#
opensource
#
benchmarking
Comments
Add Comment
2 min read
đź‘‹
Sign in
for the ability to sort posts by
relevant
,
latest
, or
top
.
We're a place where coders share, stay up-to-date and grow their careers.
Log in
Create account