itsmostafa/llm-engineering-skills
rlhf
Understanding Reinforcement Learning from Human Feedback (RLHF) for aligning language models. Use when learning about preference data, reward modeling, policy…
Installation
$
npx skills add https://github.com/itsmostafa/llm-engineering-skills
Similar popular skills
Related neighbors and high-traction skills in the same topics — useful to compare before installing.
openrlhf-training
Similar name
orchestra-research/ai-research-skills
High-performance RLHF framework with Ray+vLLM acceleration. Use for PPO, GRPO, RLOO, DPO traini…
752 installs
openrlhf-training
Similar name
davila7/claude-code-templates
High-performance RLHF framework with Ray+vLLM acceleration. Use for PPO, GRPO, RLOO, DPO traini…
308 installsAlso in this package
Other skills from itsmostafa/llm-engineering-skills.
npx skills add https://github.com/itsmostafa/llm-engineering-skills
More details
Agent compatibility
Declared targets from SKILL.md / docs. Unmarked agents are not listed — the skill may still install via the CLI.
Claude Code
Not declared
Cursor
Not declared
Codex
Not declared
GitHub Copilot
Not declared
Windsurf
Not declared
Gemini CLI
Not declared
Cline
Not declared
OpenCode
Not declared
Repository health
Stars
23
License
LICENSE
Default branch
main
Open issues
0
Status
Active
History
- First seen on skills.sh
- First recorded snapshot · 21 installs