DEV Community

Avery Wang profile picture

Avery Wang

Building things with Python, Go, and JavaScript. Love automating everything.

Location Sydney, Australia Joined Joined on 
Coding Agents Fail Differently by Stratum

Coding Agents Fail Differently by Stratum

Comments
7 min read
Count Retries Before You Trust a Coding Score

Count Retries Before You Trust a Coding Score

Comments
6 min read
Freeze the Retry Budget Before the Percentage

Freeze the Retry Budget Before the Percentage

Comments
7 min read
Assumption Density Belongs Beside Pass Rate

Assumption Density Belongs Beside Pass Rate

Comments
6 min read
Publish the Fixture Hash With the Score

Publish the Fixture Hash With the Score

1
Comments
6 min read
A Benchmark Is Only as Honest as Its Harness

A Benchmark Is Only as Honest as Its Harness

Comments
4 min read
Memory Pollution in Free AI Benchmarks

Memory Pollution in Free AI Benchmarks

Comments
5 min read
Feed Your Worst Test to a Free Model

Feed Your Worst Test to a Free Model

Comments
3 min read
Benchmark a Free AI Coding Tier on a Cold Server

Benchmark a Free AI Coding Tier on a Cold Server

Comments
4 min read
The Reviewer Is the Untested Component

The Reviewer Is the Untested Component

Comments
4 min read
A Free Tier Benchmark Needs a Test Set, Not a Screenshot

A Free Tier Benchmark Needs a Test Set, Not a Screenshot

Comments
6 min read
Free AI Coding Tiers Need a Protocol, Not a Press Release

Free AI Coding Tiers Need a Protocol, Not a Press Release

Comments
5 min read
Schedule the Free Model for 2 AM

Schedule the Free Model for 2 AM

Comments
5 min read
A Token Ceiling Is the Best Prompt Engineering Teacher

A Token Ceiling Is the Best Prompt Engineering Teacher

Comments
4 min read
Free AI Coding Credits Are a Benchmark, Not a Gift

Free AI Coding Credits Are a Benchmark, Not a Gift

Comments
5 min read
Your AI Coding Budget Should Start at Zero

Your AI Coding Budget Should Start at Zero

Comments
4 min read
Before You Trust MiniMax H3, Run This Free Baseline Harness

Before You Trust MiniMax H3, Run This Free Baseline Harness

Comments
5 min read
A Two-Model Regression Harness for Evaluating a New Low-Cost Model Release

A Two-Model Regression Harness for Evaluating a New Low-Cost Model Release

Comments
5 min read
Stop Arguing About Which Model Is Best. Build a Two-Tier Habit Instead.

Stop Arguing About Which Model Is Best. Build a Two-Tier Habit Instead.

Comments
5 min read
A New MiniMax Model Dropped: How to Evaluate It on Your Own Code Before Believing the Benchmarks

A New MiniMax Model Dropped: How to Evaluate It on Your Own Code Before Believing the Benchmarks

Comments
4 min read
Your AI Coding Assistant Writes Shell Commands. Do You Actually Test Them Before They Run?

Your AI Coding Assistant Writes Shell Commands. Do You Actually Test Them Before They Run?

Comments
5 min read
loading...