Source

ai-evals-course/evals-skills

10 skills · 2.5K combined installs

Skills from this source

#
Skill
Source
8W Activity
Installs
1
error-discovery Run error analysis on a dataset. Build a review UI, select diverse samples, monitor annotations, and organize failure…
ai-evals-course/evals-skills
324
2
eval-audit Audit an LLM eval pipeline and surface problems: missing error analysis, unvalidated judges, vanity metrics, etc. Use…
ai-evals-course/evals-skills
321
3
validate-evaluator Calibrate an LLM judge against human labels using data splits, TPR/TNR, and bias correction. Use after writing a judg…
ai-evals-course/evals-skills
315
4
write-judge-prompt Design LLM-as-Judge evaluators for subjective criteria that code-based checks cannot handle. Use when a failure mode …
ai-evals-course/evals-skills
315
5
generate-synthetic-data Create diverse synthetic test inputs for LLM pipeline evaluation using dimension-based tuple generation. Use when boo…
ai-evals-course/evals-skills
310
6
evaluate-rag Guides evaluation of RAG pipeline retrieval and generation quality. Use when evaluating a retrieval-augmented generat…
ai-evals-course/evals-skills
308
7
build-review-interface Build a custom browser-based annotation interface tailored to your data for reviewing LLM traces and collecting struc…
ai-evals-course/evals-skills
304
8
start Entry point for evals. Use when the user asks for help with evals, does not know where to begin, or asks for somethin…
ai-evals-course/evals-skills
227
9
evals-start Entry point for evals. Use when the user asks for help with evals, does not know where to begin, or asks for somethin…
ai-evals-course/evals-skills
87
10
error-analysis Entry point for "error analysis" requests. Routes to the targeted skill that matches the user's situation. Use when t…
ai-evals-course/evals-skills
2