akillness/ai-research-skills · Archived
evaluating-code-models
Evaluates code generation models across HumanEval, MBPP, MultiPL-E, and 15+ benchmarks with pass@k metrics. Use when benchmarking code models, comparing coding…
Installation
npx skills add https://github.com/akillness/ai-research-skills
Stronger alternatives
This repository is archived — consider an actively maintained alternative.
Runs LLM inference on CPU, Apple Silicon, and consumer GPUs without NVIDIA hardware. Use for ed…
2 installsExtend context windows of transformer models using RoPE, YaRN, ALiBi, and position interpolatio…
2 installsWrite publication-ready ML/AI/Systems papers for NeurIPS, ICML, ICLR, ACL, AAAI, COLM, OSDI, NS…
1 installsPyTorch library for audio generation including text-to-music (MusicGen) and text-to-sound (Audi…
1 installsSimilar popular skills
Related neighbors and high-traction skills in the same topics — useful to compare before installing.
Help users navigate complex product decisions by weighing competing priorities and calculating …
1.9K installsHelp users make better hiring decisions. Use when someone is evaluating job candidates, making …
1.6K installsHelp users evaluate emerging technologies. Use when someone is assessing new tools, making buil…
1.6K installsEvaluates LLMs across 60+ academic benchmarks (MMLU, HumanEval, GSM8K, TruthfulQA, HellaSwag). …
762 installsEvaluates code generation models across HumanEval, MBPP, MultiPL-E, and 15+ benchmarks with pas…
744 installsEvaluates NVIDIA Cosmos Policy on LIBERO and RoboCasa simulation environments. Use when setting…
656 installsAlso in this package
Other skills from akillness/ai-research-skills · top by installs.
npx skills add https://github.com/akillness/ai-research-skills
More details
Agent compatibility
Declared targets from SKILL.md / docs. Unmarked agents are not listed — the skill may still install via the CLI.
Repository health
main
History
- First seen on skills.sh
- First recorded snapshot · 1 installs