davila7/claude-code-templates
nemo-evaluator-sdk
Evaluates LLMs across 100+ benchmarks from 18+ harnesses (MMLU, HumanEval, GSM8K, safety, VLM) with multi-backend execution. Use when needing scalable…
Installation
npx skills add https://github.com/davila7/claude-code-templates
Similar popular skills
Related neighbors and high-traction skills in the same topics — useful to compare before installing.
Evaluates LLMs across 100+ benchmarks from 18+ harnesses (MMLU, HumanEval, GSM8K, safety, VLM) …
741 installsINVOKE THIS SKILL when building evaluation pipelines for LangSmith. Covers three core component…
4.2K installsEvaluate models, datasets, and agents with the NeMo Evaluator plugin. Use for metric selection,…
1.8K installsHandles LLM-as-judge evaluation workflows on Arize including creating/updating evaluators, runn…
1.1K installsHandles LLM-as-judge and code evaluator workflows on Arize including creating/updating evaluato…
2.5K installsTechnology stack evaluation and comparison with TCO analysis, security assessment, and ecosyste…
904 installsAlso in this package
Other skills from davila7/claude-code-templates · top by installs.
npx skills add https://github.com/davila7/claude-code-templates
More details
Agent compatibility
Declared targets from SKILL.md / docs. Unmarked agents are not listed — the skill may still install via the CLI.
Repository health
main
Skill metadata
Parsed from SKILL.md frontmatter.
History
- First seen on skills.sh
- First recorded snapshot · 86 installs