flagos-ai/skills

model-migrate-flagos

Migrate a model from the latest vLLM upstream repository into the vllm-plugin-FL project (pinned at vLLM v0.13.0). Use this skill whenever someone wants to add support for a new model to vllm-plugin-FL, port model code from upstream vLLM, or backport a newly released model. Trigger when the user says things like "migrate X model", "add X model support", "port X from upstream vLLM", "make X work with the FL plugin", or simply "/model-migrate-flagos model_name". The model_name argument uses snake…

First seen Mar 22, 2026

Installation

$ npx skills add flagos-ai/skills --skill model-migrate-flagos

Similar popular skills

Related neighbors and high-traction skills in the same topics — useful to compare before installing.

Also in this package

Other skills from flagos-ai/skills · top by installs.

npx skills add flagos-ai/skills

Browse all from flagos-ai/skills

More details

Agent compatibility

Declared targets from SKILL.md / docs. Unmarked agents are not listed — the skill may still install via the CLI.

Claude Code Not declared
Cursor Not declared
Codex Not declared
GitHub Copilot Not declared
Windsurf Not declared
Gemini CLI Not declared
Cline Not declared
OpenCode Not declared

Also listed on

Alternate registries and mirrors of this skill.

Repository health

Stars 19
License LICENSE
Default branch main
Open issues 1
Status Active

Skill metadata

Parsed from SKILL.md frontmatter.

Version1.0.0
CompatibilityRequires vLLM 0.13.0, Python 3.8+, GPU with CUDA, vllm-plugin-FL installed via pip install -e
Allowed toolsBash(pytest:*) Bash(python3:*) Bash(git:*) Bash(vllm:*) Bash(cp:*) Bash(curl:*) Bash(bash:*) Bash(test:*) Bash(ls:*) Bash(pip:*) Bash(cd:*) Read Edit Write Glob Grep AskUserQuestion TaskCreate TaskUpdate TaskList TaskGet
More metadata
version
1.0.0
author
flagos-ai
category
workflow-automation
tags
["model-migration","vllm","backport","e2e-verification"]

Package contents

Files included with this skill beyond the listing page.

  • skill md SKILL.md 9,123 B
  • docs SUMMARY.md 695 B

History

  1. First seen on skills.sh
  2. First recorded snapshot · 49 installs

SKILL.md

<!-- Copyright 2026 FlagOS Contributors

Licensed under the Apache License, Version 2.0 (the "License"); you may not use this file except in compliance with the License. You may obtain a copy of the License at

http://www.apache.org/licenses/LICENSE-2.0

Unless required by applicable law or agreed to in writing, software distributed under the License is distributed on an "AS IS" BASIS, WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. See the License for the specific language governing permissions and limitations under the License. -->

FL Plugin — Model Migration Skill

Usage

/model-migrate-flagos <model_name> [upstream_folder] [plugin_folder]
Argument Required Default
model_name Yes —
upstream_folder No /tmp/vllm-upstream-ref
plugin_folder No current working directory

Execution

Step 1: Parse arguments and validate paths

Extract from user input:

  • {{modelname}} = first argument (required, snakecase)
  • {{upstream_folder}} = second argument or /tmp/vllm-upstream-ref
  • {{plugin_folder}} = third argument or current working directory

If {{upstreamfolder}} doesn't exist, ask user whether to clone it. If {{pluginfolder}} doesn't exist, error out.

→ Tell user: Confirm parsed model name and paths.

Step 2: Load references and resolve placeholders

Read these files (relative to this SKILL.md):

  • references/procedure.md — step-by-step migration procedure
  • references/compatibility-patches.md — 0.13.0 patch catalog
  • references/operational-rules.md — communication, TaskList, bash rules, resilience

The procedure references executable scripts in scripts/:

  • scripts/validate_migration.py — automated code review (Step 6)
  • scripts/benchmark.sh — benchmark verification (Step 9)
  • scripts/serve.sh — serve model locally (Step 10.1, also used for E2E)
  • scripts/request.sh — test request (Step 10.2)
  • scripts/e2e_eval.py — E2E correctness verification (Step 11)
  • scripts/e2etestprompts.json — test prompts for E2E (5 text + 5 multimodal)
  • scripts/e2econfig.template.json — E2E config template (copy to e2econfig.json and fill in)
  • scripts/e2eremoteserve.sh — manage GT server on remote machine via SSH

Then investigate upstream source + HuggingFace to resolve all placeholders:

Placeholder How to derive
{{model_name}} Direct from argument
{{modelnamelower}} Lowercase of modelname (usually identical, e.g. qwen35) — used in file paths
{{MODELDISPLAYNAME}} From upstream code or HF model card
{{ModelClassName}} From upstream model class (PascalCase)
{{model_type}} From HF config.json model_type field
{{ConfigClassName}} From upstream or derive from model_type
{{skill_root}} Absolute path to this skill's folder (the directory containing this SKILL.md)

Naming conventions vary per model — always verify from actual source, never guess.

→ Tell user: Present all resolved values. Use AskUserQuestion if anything is ambiguous.

Step 3: Execute procedure

With placeholders resolved, execute every step in procedure.md sequentially. Apply patches from compatibility-patches.md during the copy-then-patch step. Follow operational-rules.md throughout.

→ Tell user: Before starting, output a numbered plan. Report progress at each step boundary.

Scripts Reference

Script Step Description
validate_migration.py 6 Automated import/API/registration checks
benchmark.sh 9 vllm bench throughput with dummy weights
serve.sh 10, 11 Start local vLLM server (port 8122, VLLMFLPREFER_ENABLED=false)
request.sh 10 Quick smoke-test request
e2e_eval.py 11 Token-level comparison vs upstream GT server
e2etestprompts.json 11 5 text + 5 multimodal test prompts
e2e_config.template.json 11 Config template (GT machine, local port, eval params)
e2eremoteserve.sh 11 SSH-based GT server lifecycle (start/stop/status/logs)

Examples

Example 1: Typical new model

User says: "/model-migrate-flagos kimi_k25"
Actions:
  1. Parse → model_name=kimi_k25, defaults for upstream/plugin paths
  2. Clone upstream, find vllm/model_executor/models/kimi_k25.py
  3. Discover it wraps DeepseekV2 → follow kimi_k25 (wrapper) pattern
  4. Copy file, apply P1+P2 patches, create config bridge
  5. Register, validate, test, benchmark, serve+request
  6. E2E verification against upstream GT
Result: kimi_k25 fully working in plugin, all 11 steps passed

Example 2: Re-run after upstream update

User says: "migrate qwen3_5 again, upstream updated"
Actions:
  1. Idempotent re-run — overwrite existing files with fresh upstream copy
  2. Re-apply patches, re-validate, re-test
  3. Re-run E2E to confirm no regression
Result: qwen3_5 updated to match latest upstream, no regressions

Troubleshooting

General principle: When any runtime error occurs, first compare vLLM upstream code against both the plugin adaptation and the installed 0.13.0 environment. The diff is the fastest path to root cause. See operational-rules.md § Debugging Priority: Upstream-First for the full protocol.

Problem Typical Cause Fix
ImportError after copy-then-patch Missing P1 fix (relative→absolute imports) Verify all from .xxx converted to from vllm. or from vllm_fl.
AttributeError: module 'vllm' has no attribute X API doesn't exist in 0.13.0 Check P3 in compatibility-patches.md; stub or remove
Config not recognized by vLLM model_type mismatch or config bridge missing Verify CONFIGREGISTRY[model_type] matches HF config.json exactly
Registration has no effect Class name or import path typo Compare with existing registrations in init.py
Benchmark KeyError on config field Config bridge missing a field Compare upstream config class vs bridge; add missing fields with defaults
Benchmark/Serve fails with OOM or "insufficient memory" GPUs occupied by other processes Kill GPU processes: `nvidia-smi --query-compute-apps=pid --format=csv,noheader \ xargs -r kill -9` then retry. Never skip these steps.
Model outputs garbled/gibberish text ColumnParallelLinear used for merged projections with different sub-dimensions (TP sharding mismatch) Override init to use MergedColumnParallelLinear(output_sizes=[...]). See P8 in compatibility-patches.md
AssertionError: Duplicate op name Child class imports custom op from different module path than parent Use same import path as parent module (e.g. vllmfl.ops.fla not vllmfl.models.fla_ops). See P11
AttributeError on fusedrecurrent* during CUDA graph warmup init override with nn.Module.init(self) missed attributes used by inherited forwardcore Create ALL attributes from parent's init, especially custom ops. See P12
E2E: local server not reachable serve.sh port doesn't match e2e_config.json local port Ensure both use same port (default 8122)
E2E: GT server not reachable GT machine down or docker/conda env wrong Check e2eremoteserve.sh status or SSH manually
E2E: early token divergence (first 5 tokens) Weight loading bug, TP sharding error Check loadweights, stackedparams_mapping, MergedColumnParallelLinear
E2E: late minor divergence (token #15+) Numerical noise from different op implementations Usually acceptable; document in report
resolveop fails with VLLMFLPREFERENABLED=false Op not registered in dispatch, no fallback Add try/except fallback to flag_gems in op import code