npx skills add https://modelscope.cn/collections/FlagRelease/FlagOS-Skills
flagos-ai/skills
perf-test-flagos
Run accuracy benchmarks (FlagEval, when available) and performance benchmarks (vllm bench serve) against a served model. Covers 5 workload profiles: short/long prefill x short/long decode + high concurrency. Collects throughput, latency, TTFT, TPOT metrics.
Installation
npx skills add flagos-ai/skills --skill perf-test-flagos
Similar popular skills
Related neighbors and high-traction skills in the same topics — useful to compare before installing.
Browser automation CLI for AI agents. Use when the user needs to interact with websites, includ…
810.4K installsDebug Azure production issues on Azure using AppLens, Azure Monitor, resource health, and safe …
568.9K installsPre-deployment validation for Azure readiness. Run deep checks on configuration, infrastructure…
567.7K installsConfigure Azure API Management as an AI Gateway for AI models, MCP tools, and agents. WHEN: sem…
566.3K installsAzure VM/VMSS router. WHEN: create / provision / deploy / spin-up VM, recommend VM size, compar…
510K installsPostgres best practices maintained by Supabase, for Postgres running anywhere. Load this skill …
391.6K installsAlso in this package
Other skills from flagos-ai/skills · top by installs.
npx skills add flagos-ai/skills
More details
Agent compatibility
Declared targets from SKILL.md / docs. Unmarked agents are not listed — the skill may still install via the CLI.
Also listed on
Alternate registries and mirrors of this skill.
Repository health
main
Skill metadata
Parsed from SKILL.md frontmatter.
Bash(*) Read Edit Write Glob Grep WebSearch WebFetch AskUserQuestionPackage contents
Files included with this skill beyond the listing page.
-
skill md
SKILL.md6,031 B -
docs
SUMMARY.md281 B
History
- First seen on skills.sh
- First recorded snapshot · 41 installs
SKILL.md
<!-- Copyright 2026 FlagOS Contributors
Licensed under the Apache License, Version 2.0 (the "License"); you may not use this file except in compliance with the License. You may obtain a copy of the License at
http://www.apache.org/licenses/LICENSE-2.0
Unless required by applicable law or agreed to in writing, software distributed under the License is distributed on an "AS IS" BASIS, WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. See the License for the specific language governing permissions and limitations under the License. -->
Accuracy + Performance Test
Start vLLM serve with the target model, run accuracy benchmarks (when FlagEval is available) and performance benchmarks (vllm bench serve) across multiple profiles.
Skill Components
perf-test/
├── SKILL.md # This file — execution flow
├── scripts/
│ ├── run_benchmark.py # Run single benchmark profile (JSON output)
│ └── run_all_benchmarks.py # Run all 5 profiles, collect + summarize (JSON)
└── references/
└── benchmark-profiles.md # Profile definitions, metrics, vllm bench usage
Reused from env-verify:
env-verify/scripts/testservemode.py— can be used to verify server is healthy
before benchmarking (optional pre-check)
Prerequisites
- Running container with software stack installed
- model-verify completed — know which stack to use (
fullvsbase) - Model path, TP size, and recommended stack from model-verify
If invoked standalone, ask for container name, model path, TP size, and stack config. If invoked from /flagrelease, these are passed as context.
Execution Flow
Step 1: Start vLLM Server
Use the stack recommended by model-verify. Read references/benchmark-profiles.md for the vllm serve command pattern.
docker exec -d <CONTAINER> bash -c '
export USE_FLAGGEMS=<0|1>
export FLAGCX_PATH=<path_or_unset>
export VLLM_PLUGINS=<fl_or_unset>
vllm serve <MODEL_PATH> \
--tensor-parallel-size <TP_SIZE> \
--max-num-batched-tokens 4096 \
--max-num-seqs 256 \
--trust-remote-code \
--port 8000 \
<EXTRA_ARGS>
'
Wait for server ready (poll /health, timeout 300s):
docker exec <CONTAINER> bash -c '
for i in $(seq 1 150); do
if curl -s http://localhost:8000/health 2>/dev/null | grep -qE "ok|200|\{\}"; then
echo "SERVER_READY"; break
fi
sleep 2
done
'
If server doesn't start, report error and exit.
Step 2: Get Model Name from Server
docker exec <CONTAINER> bash -c '
curl -s http://localhost:8000/v1/models | python3 -c "
import json, sys; print(json.load(sys.stdin)[\"data\"][0][\"id\"])"
'
Part A: Accuracy Test (FlagEval) — PLACEHOLDER
STATUS: FlagEval test client not yet available.
When FlagEval becomes available, update this section with:
- Docker image URL or pip package name
- Supported benchmarks (MMLU, GSM8K, HumanEval, etc.)
- Required arguments and configuration
- Expected output format
- Pass/fail criteria (accuracy thresholds)
Current behavior: Report accuracy test as SKIPPED.
Part B: Performance Benchmarks
Step 3: Run All Benchmark Profiles
Copy scripts into the container and run:
docker cp <SKILL_DIR>/scripts/run_benchmark.py <CONTAINER>:/tmp/
docker cp <SKILL_DIR>/scripts/run_all_benchmarks.py <CONTAINER>:/tmp/
docker exec <CONTAINER> python3 /tmp/run_all_benchmarks.py \
--model <MODEL_NAME> \
--tokenizer <MODEL_PATH> \
--port 8000 \
--output-dir /data/results/perf
The script runs all 5 default profiles (see references/benchmark-profiles.md), saves per-profile JSON to /data/results/perf/, and outputs a combined JSON report with a summary table.
Important: One profile failure does NOT skip remaining profiles.
Step 4: Stop Server
docker exec <CONTAINER> bash -c 'pkill -f "vllm serve" || true'
Step 5: Produce Report
{
"status": "PASS | PARTIAL | FAIL",
"stage": "perf-test",
"model": "<MODEL_PATH>",
"tensor_parallel_size": 8,
"flags": {"USE_FLAGGEMS": "1|0", "FLAGCX_PATH": "..."},
"accuracy": {
"status": "SKIPPED",
"reason": "FlagEval test client not yet available"
},
"performance": {
"status": "PASS | PARTIAL | FAIL",
"profiles_passed": "5/5",
"profiles": [ "...per-profile results..." ],
"summary_table": "...markdown table..."
}
}
Present the summary table to the user:
| Profile | Input | Output | Prompts | Req/s | Tok/s | TTFT(ms) | TPOT(ms) | P99(ms) | Status |
|---------|-------|--------|---------|-------|-------|----------|----------|---------|--------|
| ... | ... | ... | ... | ... | ... | ... | ... | ... | ... |
Status logic:
PASS— all profiles completedPARTIAL— some passed, some failedFAIL— server didn't start or all profiles failed
Error Handling
| Failure | Behavior |
|---|---|
| Server fails to start | Report error; exit |
vllm bench serve not found |
Report vllm version issue |
| Single profile fails | Report error, continue remaining profiles |
| Single profile times out | Kill after 600s, report partial, continue |
| Server crashes mid-benchmark | Capture logs, report which profile caused crash |
| OOM during high concurrency | Report, suggest reducing num_prompts |
Timeout Rules
| Operation | Timeout |
|---|---|
| Server startup | 300s |
| Per profile benchmark | 600s |