Skip to content
Navigation menu
Search
Powered by Algolia
Search
Log in
Create account
DEV Community
Close
#
evaluation
Follow
Hide
Posts
Left menu
đź‘‹
Sign in
for the ability to sort posts by
relevant
,
latest
, or
top
.
Right menu
El modelo encontró evidencia relevante y aun asà falló: por qué “alucinación” se me quedó corta
Cristian Gormaz
Cristian Gormaz
Cristian Gormaz
Follow
Sep 7
El modelo encontró evidencia relevante y aun asà falló: por qué “alucinación” se me quedó corta
#
ai
#
testing
#
llm
#
evaluation
Comments
Add Comment
4 min read
The sleep loop is the tell: agents that pay per action optimize to do nothing
Cole Halton
Cole Halton
Cole Halton
Follow
Sep 7
The sleep loop is the tell: agents that pay per action optimize to do nothing
#
aiagents
#
evaluation
#
llm
#
benchmarking
Comments
1
 comment
2 min read
The model did the reverse-engineering. The validator was the hard part.
Cole Halton
Cole Halton
Cole Halton
Follow
Sep 7
The model did the reverse-engineering. The validator was the hard part.
#
aicoding
#
agents
#
llm
#
evaluation
Comments
1
 comment
2 min read
Judging AI hackathon projects: what to check when every team says 'we used AI'
PRANJUL RATHOUR
PRANJUL RATHOUR
PRANJUL RATHOUR
Follow
Sep 6
Judging AI hackathon projects: what to check when every team says 'we used AI'
#
hackathonjudging
#
ai
#
evaluation
#
rubric
Comments
Add Comment
3 min read
How to evaluate a RAG system: recall, faithfulness and the questions that matter
PRANJUL RATHOUR
PRANJUL RATHOUR
PRANJUL RATHOUR
Follow
Sep 6
How to evaluate a RAG system: recall, faithfulness and the questions that matter
#
rag
#
evaluation
#
metrics
#
guide
Comments
Add Comment
3 min read
How to Evaluate AI Agents
Quantiles.io
Quantiles.io
Quantiles.io
Follow
Sep 4
How to Evaluate AI Agents
#
ai
#
agentskills
#
agents
#
evaluation
2
 reactions
Comments
Add Comment
7 min read
JuryTrace: make agent-judge failures inspectable
Ama Senevirathne
Ama Senevirathne
Ama Senevirathne
Follow
Sep 5
JuryTrace: make agent-judge failures inspectable
#
ai
#
llm
#
evaluation
#
python
Comments
1
 comment
4 min read
My board never scored an outage as a regression. My evidence couldn't prove it.
Erik Hill
Erik Hill
Erik Hill
Follow
Aug 31
My board never scored an outage as a regression. My evidence couldn't prove it.
#
evaluation
#
testing
#
llm
#
opensource
Comments
Add Comment
5 min read
We spent two days bisecting a prompt change. The regression was noise.
Muhammad Waqas
Muhammad Waqas
Muhammad Waqas
Follow
Aug 30
We spent two days bisecting a prompt change. The regression was noise.
#
mlops
#
regression
#
evaluation
#
agents
Comments
Add Comment
1 min read
Production-Ready Multi-Turn Evaluation
Humza Tareen
Humza Tareen
Humza Tareen
Follow
Aug 25
Production-Ready Multi-Turn Evaluation
#
multiturn
#
evaluation
#
python
#
docker
Comments
Add Comment
7 min read
A Free Server Is Enough to Test a New Model Before You Trust It
Quinn Li
Quinn Li
Quinn Li
Follow
Aug 17
A Free Server Is Enough to Test a New Model Before You Trust It
#
ai
#
python
#
evaluation
#
agents
Comments
Add Comment
3 min read
Writing the code is no longer the bottleneck
The Engineering Manager’s Desk
The Engineering Manager’s Desk
The Engineering Manager’s Desk
Follow
Aug 15
Writing the code is no longer the bottleneck
#
engineeringculture
#
evaluation
#
ai
#
techdebt
Comments
Add Comment
3 min read
Why AI Benchmarks Mean Less Than You Think
The AI Downside
The AI Downside
The AI Downside
Follow
Aug 15
Why AI Benchmarks Mean Less Than You Think
#
benchmarks
#
llms
#
evaluation
#
hype
Comments
Add Comment
6 min read
One Quality Score Is a Lie: Split Your RAG Judge Into Retrieval, Groundedness, and Relevance
Saurav Bhattacharya
Saurav Bhattacharya
Saurav Bhattacharya
Follow
Aug 19
One Quality Score Is a Lie: Split Your RAG Judge Into Retrieval, Groundedness, and Relevance
#
ai
#
evaluation
#
dotnet
#
testing
1
 reaction
Comments
1
 comment
5 min read
We Almost Deployed a Temporal Knowledge Graph. The Eval Said No.
Guatu
Guatu
Guatu
Follow
Aug 14
We Almost Deployed a Temporal Knowledge Graph. The Eval Said No.
#
aiagents
#
knowledgegraph
#
evaluation
#
rag
Comments
Add Comment
7 min read
đź‘‹
Sign in
for the ability to sort posts by
relevant
,
latest
, or
top
.
We're a place where coders share, stay up-to-date and grow their careers.
Log in
Create account