davila7/claude-code-templates
evaluating-code-models
Evaluates code generation models across HumanEval, MBPP, MultiPL-E, and 15+ benchmarks with pass@k metrics. Use when benchmarking code models, comparing coding…
Installation
npx skills add https://github.com/davila7/claude-code-templates
Similar popular skills
Related neighbors and high-traction skills in the same topics — useful to compare before installing.
Evaluates code generation models across HumanEval, MBPP, MultiPL-E, and 15+ benchmarks with pas…
744 installsHelp users navigate complex product decisions by weighing competing priorities and calculating …
1.9K installsHelp users make better hiring decisions. Use when someone is evaluating job candidates, making …
1.6K installsHelp users evaluate emerging technologies. Use when someone is assessing new tools, making buil…
1.6K installsEvaluates LLMs across 60+ academic benchmarks (MMLU, HumanEval, GSM8K, TruthfulQA, HellaSwag). …
762 installsEvaluates NVIDIA Cosmos Policy on LIBERO and RoboCasa simulation environments. Use when setting…
656 installsAlso in this package
Other skills from davila7/claude-code-templates · top by installs.
npx skills add https://github.com/davila7/claude-code-templates
More details
Agent compatibility
Declared targets from SKILL.md / docs. Unmarked agents are not listed — the skill may still install via the CLI.
Repository health
main
Skill metadata
Parsed from SKILL.md frontmatter.
History
- First seen on skills.sh
- First recorded snapshot · 359 installs