Skip to content
Navigation menu
Search
Powered by Algolia
Search
Log in
Create account
DEV Community
Close
#
evaluation
Follow
Hide
Posts
Left menu
đź‘‹
Sign in
for the ability to sort posts by
relevant
,
latest
, or
top
.
Right menu
AI Model Evaluation: Best Practices for Testing and Validation
ryan2run
ryan2run
ryan2run
Follow
Sep 14
AI Model Evaluation: Best Practices for Testing and Validation
#
ai
#
evaluation
#
machinelearning
#
testing
Comments
Add Comment
2 min read
Stop asking the model that wrote the code to review it
Cole Halton
Cole Halton
Cole Halton
Follow
Sep 12
Stop asking the model that wrote the code to review it
#
aicodereview
#
evaluation
#
aiagents
#
codereview
Comments
1
 comment
2 min read
El modelo encontró evidencia relevante y aun asà falló: por qué “alucinación” se me quedó corta
Cristian Gormaz
Cristian Gormaz
Cristian Gormaz
Follow
Sep 7
El modelo encontró evidencia relevante y aun asà falló: por qué “alucinación” se me quedó corta
#
ai
#
testing
#
llm
#
evaluation
Comments
Add Comment
4 min read
The sleep loop is the tell: agents that pay per action optimize to do nothing
Cole Halton
Cole Halton
Cole Halton
Follow
Sep 7
The sleep loop is the tell: agents that pay per action optimize to do nothing
#
aiagents
#
evaluation
#
llm
#
benchmarking
Comments
1
 comment
2 min read
The model did the reverse-engineering. The validator was the hard part.
Cole Halton
Cole Halton
Cole Halton
Follow
Sep 7
The model did the reverse-engineering. The validator was the hard part.
#
aicoding
#
agents
#
llm
#
evaluation
Comments
1
 comment
2 min read
Judging AI hackathon projects: what to check when every team says 'we used AI'
PRANJUL RATHOUR
PRANJUL RATHOUR
PRANJUL RATHOUR
Follow
Sep 6
Judging AI hackathon projects: what to check when every team says 'we used AI'
#
hackathonjudging
#
ai
#
evaluation
#
rubric
Comments
Add Comment
3 min read
How to evaluate a RAG system: recall, faithfulness and the questions that matter
PRANJUL RATHOUR
PRANJUL RATHOUR
PRANJUL RATHOUR
Follow
Sep 6
How to evaluate a RAG system: recall, faithfulness and the questions that matter
#
rag
#
evaluation
#
metrics
#
guide
Comments
Add Comment
3 min read
The benchmark that disproved its own result
the kilted dev
the kilted dev
the kilted dev
Follow
Sep 10
The benchmark that disproved its own result
#
localmodels
#
llm
#
evaluation
#
buildinpublic
Comments
1
 comment
6 min read
How to Evaluate AI Agents
Quantiles.io
Quantiles.io
Quantiles.io
Follow
Sep 4
How to Evaluate AI Agents
#
ai
#
agentskills
#
agents
#
evaluation
2
 reactions
Comments
Add Comment
7 min read
JuryTrace: make agent-judge failures inspectable
Ama Senevirathne
Ama Senevirathne
Ama Senevirathne
Follow
Sep 5
JuryTrace: make agent-judge failures inspectable
#
ai
#
llm
#
evaluation
#
python
Comments
1
 comment
4 min read
My board never scored an outage as a regression. My evidence couldn't prove it.
Erik Hill
Erik Hill
Erik Hill
Follow
Aug 31
My board never scored an outage as a regression. My evidence couldn't prove it.
#
evaluation
#
testing
#
llm
#
opensource
Comments
Add Comment
5 min read
We spent two days bisecting a prompt change. The regression was noise.
Muhammad Waqas
Muhammad Waqas
Muhammad Waqas
Follow
Aug 30
We spent two days bisecting a prompt change. The regression was noise.
#
mlops
#
regression
#
evaluation
#
agents
Comments
Add Comment
1 min read
They knew it wasn't the model. They patched it anyway.
JaviMaligno
JaviMaligno
JaviMaligno
Follow
Sep 12
They knew it wasn't the model. They patched it anyway.
#
ai
#
agents
#
evaluation
#
research
Comments
6
 comments
9 min read
Nineteen of twenty patched the symptom
JaviMaligno
JaviMaligno
JaviMaligno
Follow
Sep 11
Nineteen of twenty patched the symptom
#
ai
#
agents
#
evaluation
#
research
Comments
2
 comments
9 min read
Production-Ready Multi-Turn Evaluation
Humza Tareen
Humza Tareen
Humza Tareen
Follow
Aug 25
Production-Ready Multi-Turn Evaluation
#
multiturn
#
evaluation
#
python
#
docker
Comments
Add Comment
7 min read
đź‘‹
Sign in
for the ability to sort posts by
relevant
,
latest
, or
top
.
We're a place where coders share, stay up-to-date and grow their careers.
Log in
Create account