Skip to content
Navigation menu
Search
Powered by Algolia
Search
Log in
Create account
DEV Community
Close
#
llmevaluation
Follow
Hide
Posts
Left menu
đź‘‹
Sign in
for the ability to sort posts by
relevant
,
latest
, or
top
.
Right menu
Building a Production-Grade Eval Pipeline for Your Agent, Not Just a Demo
Ali Suleyman TOPUZ
Ali Suleyman TOPUZ
Ali Suleyman TOPUZ
Follow
Sep 13
Building a Production-Grade Eval Pipeline for Your Agent, Not Just a Demo
#
llmevaluation
#
microsoftagentframew
#
dotnet
#
agents
Comments
Add Comment
12 min read
Your .NET Agent Is in Production, Your Engineering Discipline Isn’t.
Ali Suleyman TOPUZ
Ali Suleyman TOPUZ
Ali Suleyman TOPUZ
Follow
Sep 13
Your .NET Agent Is in Production, Your Engineering Discipline Isn’t.
#
microsoftagentframew
#
llmevaluation
#
dotnet
#
agenticai
Comments
Add Comment
13 min read
AX-RAY: VIDRAFT's Open AI Safety Diagnostic Leaderboard & Dataset Now on Hugging Face
AI OpenFree
AI OpenFree
AI OpenFree
Follow
Aug 19
AX-RAY: VIDRAFT's Open AI Safety Diagnostic Leaderboard & Dataset Now on Hugging Face
#
aisafety
#
causalleakage
#
llmevaluation
#
huggingfaceleaderboa
Comments
Add Comment
4 min read
Your eval monitor fired on four days this week. At your sample size, that was the most likely count
Maya Andersson
Maya Andersson
Maya Andersson
Follow
Aug 7
Your eval monitor fired on four days this week. At your sample size, that was the most likely count
#
llmevaluation
#
monitoring
#
statistics
#
observability
1
 reaction
Comments
Add Comment
8 min read
LLM Evaluation System Prompts Scored Rubrics Runtime Guardrails: A Practical Guide for Production
Christopher Hoeben
Christopher Hoeben
Christopher Hoeben
Follow
Jul 14
LLM Evaluation System Prompts Scored Rubrics Runtime Guardrails: A Practical Guide for Production
#
llmevaluation
#
systemprompts
#
scoredrubrics
#
runtimeguardrails
Comments
Add Comment
7 min read
Try It: A Working Assessment-First Course
Michael Tuszynski
Michael Tuszynski
Michael Tuszynski
Follow
Jul 13
Try It: A Working Assessment-First Course
#
aieducation
#
llmevaluation
#
opensource
#
developertools
Comments
Add Comment
4 min read
Your LLM Judge Needs a Test Suite
Michael Tuszynski
Michael Tuszynski
Michael Tuszynski
Follow
Jul 8
Your LLM Judge Needs a Test Suite
#
llmevaluation
#
aiengineering
#
softwaretesting
#
generativeai
Comments
Add Comment
4 min read
I reviewed six "operator-ready" checklists for AI agents. None of them define the problem correctly.
Maya Andersson
Maya Andersson
Maya Andersson
Follow
Jul 1
I reviewed six "operator-ready" checklists for AI agents. None of them define the problem correctly.
#
agents
#
llmevaluation
#
mlops
#
agentreliability
Comments
1
 comment
5 min read
How to Add Evals to an LLM Feature
techpotions
techpotions
techpotions
Follow
Jul 11
How to Add Evals to an LLM Feature
#
llmevaluation
#
evals
#
llmfeatures
#
aitesting
Comments
Add Comment
5 min read
đź‘‹
Sign in
for the ability to sort posts by
relevant
,
latest
, or
top
.
We're a place where coders share, stay up-to-date and grow their careers.
Log in
Create account