openbmb/minicpm · Archived

minicpm5-deploy

Pick the right inference backend for a MiniCPM5-1B or MiniCPM5-2B checkpoint and route to a backend-specific cookbook skill.

First seen Jun 1, 2026

Installation

$ npx skills add openbmb/minicpm --skill minicpm5-deploy

Summary

  • Pick the right inference backend for a MiniCPM5-1B or MiniCPM5-2B checkpoint and route to a backend-specific cookbook skill.
  • Use when the user wants to deploy / serve / chat-with / benchmark a MiniCPM5 model and has not yet committed to a specific engine, or when they say "deploy MiniCPM5", "run MiniCPM5", "serve MiniCPM5", "MiniCPM5 推理", "部署 MiniCPM5".

Stronger alternatives

This repository is archived — consider an actively maintained alternative.

Also in this package

Other skills from openbmb/minicpm.

npx skills add openbmb/minicpm

Browse all from openbmb/minicpm

More details

Agent compatibility

Declared targets from SKILL.md / docs. Unmarked agents are not listed — the skill may still install via the CLI.

Claude Code Not declared
Cursor Not declared
Codex Not declared
GitHub Copilot Not declared
Windsurf Not declared
Gemini CLI Not declared
Cline Not declared
OpenCode Not declared

Repository health

Stars 10.5K
License LICENSE
Default branch main
Open issues 7
Status Archived

Package contents

Files included with this skill beyond the listing page.

  • skill md SKILL.md 5,678 B
  • docs SUMMARY.md 386 B

History

  1. First seen on skills.sh
  2. First recorded snapshot · 1 installs

SKILL.md

Deploy MiniCPM5-1B and MiniCPM5-2B — backend router

You're being asked to deploy / serve / chat-with a MiniCPM5-1B or MiniCPM5-2B checkpoint. Your job is to pick exactly one backend skill below based on the user's hardware, format, and goal, then invoke that skill rather than improvising.

1. Required input from the user

Before picking a backend, you MUST know:

Variable Example Where to ask
MODEL_PATH HF id openbmb/MiniCPM5-2B or openbmb/MiniCPM5-1B or a local path "Which checkpoint? HF id or local path?"
Hardware NVIDIA GPU / Apple Silicon / CPU only infer from context, otherwise ask
Goal "interactive chat" / "OpenAI server" / "Python script" / "benchmark" infer from context

Available checkpoints on Hugging Face

Variant HF repo Use with
HF fp16 (recommended) openbmb/MiniCPM5-2B or openbmb/MiniCPM5-1B transformers / vllm (no --quantization) / vllm-ascend / sglang / any minicpm5-finetune-*
GGUF F16 / Q80 / Q4K_M openbmb/MiniCPM5-2B-GGUF or openbmb/MiniCPM5-1B-GGUF minicpm5-deploy-llama-cpp / -ollama / -lmstudio
MLX (Apple Silicon) openbmb/MiniCPM5-2B-MLX or openbmb/MiniCPM5-1B-MLX minicpm5-deploy-mlx

If the user has a local copy, accept any directory path that contains config.json and model.safetensors (or the equivalent GGUF / MLX layout).

2. Decision matrix — pick exactly one

User says / wants Hardware Format → Skill to invoke
"Quick Python script" / "one-shot generation" / "no server" any GPU or CPU HF safetensors minicpm5-deploy-transformers
"OpenAI server" / "production serving" / "high QPS" NVIDIA GPU HF safetensors minicpm5-deploy-vllm
"vLLM-Ascend" / "Ascend NPU" / "CANN" / "torch_npu" Huawei Ascend NPU HF safetensors minicpm5-deploy-vllm-ascend
"RadixAttention" / "prefix cache" / "batched eval" NVIDIA GPU HF safetensors minicpm5-deploy-sglang
"GGUF" / "llama.cpp" / "llama-cli" / "CPU only" any CPU + optional GPU GGUF minicpm5-deploy-llama-cpp
"Ollama" / "ollama run" / "Modelfile" macOS / Linux laptop GGUF minicpm5-deploy-ollama
"LM Studio" / "desktop GUI" macOS / Windows / Linux GGUF or MLX minicpm5-deploy-lmstudio
"MLX" / "Apple Silicon native" / "fastest on Mac" Apple Silicon MLX minicpm5-deploy-mlx

If the user has not specified any of the above and asks "how do I run this?":

  • CUDA box, want fastest server: pick minicpm5-deploy-vllm.
  • Ascend NPU, want an OpenAI-compatible server: pick minicpm5-deploy-vllm-ascend.
  • CUDA box, want minimal Python: pick minicpm5-deploy-transformers.
  • Apple Silicon laptop: pick minicpm5-deploy-ollama (easiest) or minicpm5-deploy-mlx (fastest).
  • CPU only / Windows / low-VRAM: pick minicpm5-deploy-llama-cpp (Q4KM).

3. Invocation contract

Once you've picked a backend skill, invoke that skill with MODEL_PATH set. Do NOT inline the backend's commands here — each backend has its own pitfalls (mandatory flags, env-var pins, install order) that the dedicated skill handles. The user explicitly does NOT want you to "improvise" — read the picked sub-skill in full first.

4. Sanity check after deploy

Whichever backend you pick, after launch run this universal sanity check:

# Replace localhost:PORT with the backend's actual port (default in each sub-skill)
curl http://localhost:PORT/v1/chat/completions \
    -H "Content-Type: application/json" \
    -d '{
        "model": "MiniCPM5-2B",
        "messages": [{"role":"user","content":"1+1=?"}],
        "temperature": 1.0, "top_p": 0.95, "max_tokens": 64,
        "chat_template_kwargs": {"enable_thinking": true}
    }'

Expected: HTTP 200 with choices[0].message.content containing "2".

5. Known cross-backend pitfalls

These are common to multiple backends — surface to the user up front:

  • Think vs no-think: MiniCPM5-2B uses Think mode with temperature=1.0, topp=0.95. MiniCPM5-1B also supports No-think with enablethinking=false + temperature=0.7, top_p=0.95.
  • 128 K context: maxpositionembeddings=131072, rope_theta=5e6, no rope-scaling. Pass --max-model-len 131072 (vLLM) / --context-length 131072 (SGLang) / -c 131072 (llama.cpp) to use the full window. Lower if VRAM is tight.
  • Untied lmhead: tiewordembeddings=false. Tools that assume the Llama tied default (e.g. mlxlm.convert < 0.31) will silently drop lm_head → output collapses to random tokens. The MLX skill bakes in the fix.

6. Don't reinvent: link to the cookbook

Each sub-skill is paired with a one-page cookbook in [docs/deployment/](../../docs/deployment/). The skill is the machine-readable shortcut; the cookbook is the human-readable reference. Both are kept in sync.