DEV Community

#benchmarks

Posts

đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.
GPT-6 Astra scores 95% on one robot task, 10% on another

GPT-6 Astra scores 95% on one robot task, 10% on another

5
Comments
4 min read
Claude Formalized Fermat in 11 Days. The Math Isn't New.

Claude Formalized Fermat in 11 Days. The Math Isn't New.

Comments
3 min read
13 of 14 Models Write Messier Code Than the Human Who Fixed the Same Bug

13 of 14 Models Write Messier Code Than the Human Who Fixed the Same Bug

Comments
6 min read
DeepSeek's Vision Flash Plays the Agent Game

DeepSeek's Vision Flash Plays the Agent Game

Comments
2 min read
Frontier Models Fail at Research-Level Reasoning

Frontier Models Fail at Research-Level Reasoning

Comments
2 min read
Temperature 0 is not reproducible. I measured 30 percent of my output changing between identical runs.

Temperature 0 is not reproducible. I measured 30 percent of my output changing between identical runs.

1
Comments 1
3 min read
Our 4B beat Claude Opus on a 440K-token corpus. Then it came last on the public benchmark.

Our 4B beat Claude Opus on a 440K-token corpus. Then it came last on the public benchmark.

2
Comments
4 min read
A 4B on a 6GB laptop matched frontier-model accuracy on aggregation — except when the answer is a number

A 4B on a 6GB laptop matched frontier-model accuracy on aggregation — except when the answer is a number

1
Comments
4 min read
Your agent truncates the corpus and answers anyway. Two harnesses, and a router that picks between them.

Your agent truncates the corpus and answers anyway. Two harnesses, and a router that picks between them.

1
Comments
5 min read
Frontier Models Hit a Wall on Research Thinking

Frontier Models Hit a Wall on Research Thinking

Comments
2 min read
Ranking Language Models by How Well They Spot Liars

Ranking Language Models by How Well They Spot Liars

Comments
9 min read
The Benchmarkpocalypse: Why AI Benchmarks Are Broken — and What Dan Luu Says We Should Do About It

The Benchmarkpocalypse: Why AI Benchmarks Are Broken — and What Dan Luu Says We Should Do About It

1
Comments
4 min read
Labs are ditching factual knowledge for reasoning speed

Labs are ditching factual knowledge for reasoning speed

Comments
2 min read
Why AI Benchmarks Mean Less Than You Think

Why AI Benchmarks Mean Less Than You Think

Comments
6 min read
An AI Capture-the-Flag Tournament: What the Scoreboard Counted

An AI Capture-the-Flag Tournament: What the Scoreboard Counted

Comments
6 min read
đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.