smithery/jiunbae

ml-benchmark

Guides ML model benchmarks and evaluations. Measures inference speed, memory usage, and accuracy metrics. Use for "벤치마크", "모델 평가", "성능 ?

Installation

$ npx skills add smithery/jiunbae --skill benchmarking-ml-models

Similar popular skills

Related neighbors and high-traction skills in the same topics — useful to compare before installing.

Also in this package

Other skills from smithery/jiunbae · top by installs.

npx skills add smithery/jiunbae

Browse all from smithery/jiunbae

More details

Agent compatibility

Declared targets from SKILL.md / docs. Unmarked agents are not listed — the skill may still install via the CLI.

Claude Code Not declared
Cursor Not declared
Codex Not declared
GitHub Copilot Not declared
Windsurf Not declared
Gemini CLI Not declared
Cline Not declared
OpenCode Not declared

Package contents

Files included with this skill beyond the listing page.

  • skill md SKILL.md 2,417 B
  • docs SUMMARY.md 223 B

History

  1. First recorded snapshot · 0 installs

SKILL.md

ML Benchmark

Model performance benchmarking.

Quick Benchmark

import time
import torch

# Warmup
for _ in range(10):
    model(sample_input)

# CUDA calls are asynchronous. Finish warmup before starting the clock.
use_cuda = torch.cuda.is_available()
if use_cuda:
    torch.cuda.synchronize()

# Benchmark
start = time.perf_counter()
for _ in range(100):
    model(sample_input)
if use_cuda:
    torch.cuda.synchronize()
elapsed = time.perf_counter() - start

print(f"Avg latency: {elapsed/100*1000:.2f}ms")

Metrics

Metric Description Command
Latency Inference time time.time()
Throughput Samples/sec samples / elapsed
Memory VRAM usage torch.cuda.maxmemoryallocated()
Accuracy Model quality accuracyscore(ytrue, y_pred)

Benchmark Script

The bundled helper has no inference or dataset adapter. Its run and evaluate commands intentionally exit nonzero before printing or saving results. Do not treat endpoint health, sleeps, or generated labels as model measurements.

# From this Skill's directory; currently verifies the fail-closed contract.
bash scripts/ml-benchmark.sh run \
  --url localhost:8001 \
  --model langdetector \
  --input ./sample.wav \
  --batch-size 32 \
  --runs 100

Before enabling that command, add a real, versioned inference adapter and test that every recorded sample comes from a completed model response. Accuracy also requires a dataset adapter that reads ground-truth labels and predictions.

Output Format

## Benchmark Results: {model_name}

| Metric | Value |
|--------|-------|
| Latency (p50) | 15.2ms |
| Latency (p99) | 22.1ms |
| Throughput | 65 samples/sec |
| Memory | 4.2 GB |
| Accuracy | 92.3% |

### Configuration
- GPU: NVIDIA A100
- Batch size: 32
- Precision: FP16

Compare Models

results = {}
for model_name, model in models.items():
    results[model_name] = benchmark(model)

# Generate comparison table

Best Practices

  • Always warmup before measuring
  • Use torch.cuda.synchronize() for GPU
  • Report p50/p99 latencies
  • Document hardware configuration