DEV Community

Quinn Li profile picture

Quinn Li

Building and writing about AI-powered developer tools. Obsessed with clean code and useful automation.

Location Singapore Joined Joined on 
Charge Wait Time to the Same Job

Charge Wait Time to the Same Job

Comments
7 min read
Every Extra Hop Buys Another Queue Ticket

Every Extra Hop Buys Another Queue Ticket

Comments
7 min read
Budget the Retry Path, Not the Happy Path

Budget the Retry Path, Not the Happy Path

Comments
6 min read
The Second Call Is a Different Job

The Second Call Is a Different Job

Comments
8 min read
Write the Kill Envelope Before the First Prompt

Write the Kill Envelope Before the First Prompt

Comments
7 min read
3 A.M. Token Leak: An Autopsy of an Agent Loop

3 A.M. Token Leak: An Autopsy of an Agent Loop

Comments
5 min read
Score Your Model Access Decision Before You Argue

Score Your Model Access Decision Before You Argue

Comments
4 min read
If the Job Has a Deadline, Spare Capacity Is the Wrong Bet

If the Job Has a Deadline, Spare Capacity Is the Wrong Bet

Comments
5 min read
Your Hourly Rate Makes Free Tokens Expensive

Your Hourly Rate Makes Free Tokens Expensive

Comments
7 min read
Free Model Access Needs an Exit Plan: How to Test LLM Endpoints Without Betting Your Pipeline

Free Model Access Needs an Exit Plan: How to Test LLM Endpoints Without Betting Your Pipeline

Comments
4 min read
The Retry Tax: When Free AI Capacity Stops Being Free

The Retry Tax: When Free AI Capacity Stops Being Free

Comments
3 min read
Free Tokens Are a Tool, Not a Promise: Measure Before You Build

Free Tokens Are a Tool, Not a Promise: Measure Before You Build

1
Comments
4 min read
Cheap Tokens, Expensive Waiting Rooms

Cheap Tokens, Expensive Waiting Rooms

Comments
4 min read
Free Tokens Are a Queue, Not a Wallet

Free Tokens Are a Queue, Not a Wallet

1
Comments
5 min read
Free Tokens, Real Queues: Measure What Your LLM Calls Actually Cost

Free Tokens, Real Queues: Measure What Your LLM Calls Actually Cost

Comments
4 min read
Chaos for the Cost-Conscious: Fault-Injecting Your LLM Pipeline on Free Hardware

Chaos for the Cost-Conscious: Fault-Injecting Your LLM Pipeline on Free Hardware

Comments 1
4 min read
Free, Paid, or Self-Hosted: Run the Decision Script

Free, Paid, or Self-Hosted: Run the Decision Script

Comments
4 min read
Routing Around Rate Limits: A Free-Tier Model Proxy in 100 Lines

Routing Around Rate Limits: A Free-Tier Model Proxy in 100 Lines

Comments
4 min read
The Free-Tier Trap: A Framework for Choosing Model Access

The Free-Tier Trap: A Framework for Choosing Model Access

Comments
5 min read
A Free Server Is Enough to Test a New Model Before You Trust It

A Free Server Is Enough to Test a New Model Before You Trust It

Comments
3 min read
Reproducible Evaluation Harness for New Model Releases: MiniMax H3

Reproducible Evaluation Harness for New Model Releases: MiniMax H3

Comments
5 min read
A Reproducible Tool-Call Gatekeeper for AI Agents

A Reproducible Tool-Call Gatekeeper for AI Agents

Comments
4 min read
loading...