smithery/sunholo-data

model-manager

Test, validate, and add new AI models to the eval suite. Use when user asks to add new models, test model access, check pricing, or update models.yml.

Installation

$ npx skills add smithery/sunholo-data --skill model-manager

Similar popular skills

Related neighbors and high-traction skills in the same topics — useful to compare before installing.

Also in this package

Other skills from smithery/sunholo-data · top by installs.

npx skills add smithery/sunholo-data

Browse all from smithery/sunholo-data

More details

Agent compatibility

Declared targets from SKILL.md / docs. Unmarked agents are not listed — the skill may still install via the CLI.

Claude Code Not declared
Cursor Not declared
Codex Not declared
GitHub Copilot Not declared
Windsurf Not declared
Gemini CLI Declared
Cline Not declared
OpenCode Not declared

Skill metadata

Parsed from SKILL.md frontmatter.

Declared agents gemini

Package contents

Files included with this skill beyond the listing page.

  • skill md SKILL.md 28,727 B
  • docs SUMMARY.md 171 B

History

  1. First recorded snapshot · 0 installs

SKILL.md

Model Manager

Test API access, validate configurations, and add new AI models to the AILANG eval suite.

Quick Start

Most common usage:

# User says: "Can we add GPT-5.1 to the eval suite?"
# This skill will:
# 1. Test API access to GPT-5.1
# 2. Find the correct API model name
# 3. Look up pricing information
# 4. Update models.yml configuration
# 5. Run a test benchmark to verify

When to Use This Skill

Invoke this skill when:

  • User asks to "add a new model" to eval suite
  • User mentions checking if a model is "accessible" or "available"
  • User wants to "test API access" to a model
  • User asks to "update models.yml" or "check pricing"
  • User says "can we use [model name]?" for evaluations

Available Scripts

scripts/testmodelaccess.sh <provider> <model-name>

Test API access to a model and display authentication status.

Usage:

# Test OpenAI model
scripts/test_model_access.sh openai gpt-5.1

# Test Anthropic model
scripts/test_model_access.sh anthropic claude-sonnet-4-5-20250929

# Test Google Gemini via Vertex AI
scripts/test_model_access.sh google gemini-3-pro-preview-11-2025

Output:

Testing: openai/gpt-5.1
✓ OPENAI_API_KEY found
✓ API call successful
✓ Model: gpt-5.1-2025-11-13
✓ Tokens: 13 input, 10 output (10 reasoning)
Ready to add to models.yml

scripts/findmodelinfo.sh <model-keywords>

Search for model information using web search and return API names + pricing.

Usage:

# Find GPT-5.1 info
scripts/find_model_info.sh "GPT-5.1 API model name pricing"

# Find Gemini 3 Pro info
scripts/find_model_info.sh "Gemini 3 Pro API documentation"

Output:

Searching for: GPT-5.1 API model name pricing
✓ Found API names:
  - gpt-5.1 (Thinking mode)
  - gpt-5.1-chat-latest (Instant mode)
✓ Pricing:
  Input: $1.25 per 1M tokens
  Output: $10.00 per 1M tokens
  Cached: $0.125 per 1M tokens

scripts/updatemodelsyml.sh <friendly-name> <api-name> <provider> <input-price> <output-price>

Add a new model to models.yml configuration.

Usage:

# Add GPT-5.1
scripts/update_models_yml.sh \
  gpt5-1 \
  "gpt-5.1" \
  openai \
  0.00125 \
  0.01

Output:

Adding model to models.yml:
  Friendly name: gpt5-1
  API name: gpt-5.1
  Provider: openai
  Pricing: $0.00125 / $0.01 per 1K tokens

✓ Updated models.yml
✓ Validated YAML syntax
✓ Ready to test

scripts/verifyvertexmodel.sh <model-name>

Check if a Gemini model is available in Vertex AI.

Usage:

# Check if Gemini 3 Pro is available
scripts/verify_vertex_model.sh gemini-3-pro-preview-11-2025

Output:

Checking Vertex AI for: gemini-3-pro-preview-11-2025
✓ GCP project: multivac-internal-prod
✓ Access token obtained
✗ Model not found (404)
Recommendation: Monitor for availability, check again in 1-2 weeks

scripts/runtestbenchmark.sh <model-name>

Run a small test benchmark to verify model works end-to-end.

Usage:

# Test GPT-5.1 with fizzbuzz benchmark
scripts/run_test_benchmark.sh gpt5-1

Output:

Running test benchmark: fizzbuzz
Model: gpt5-1
✓ Benchmark completed
✓ Result: PASS (100%)
✓ Tokens: 245 input, 89 output
✓ Cost: $0.002
Model is ready for production use

Workflow

1. Test API Access

First, verify you can call the model:

# Use test_model_access.sh
scripts/test_model_access.sh openai gpt-5.1

What to check:

  • API key is set (OPENAIAPIKEY, ANTHROPICAPIKEY, or gcloud auth)
  • API call succeeds (not 401/403/404)
  • Model returns expected structure
  • Token usage is reported

For Gemini models:

  • Uses Vertex AI (not public API)
  • Requires gcloud auth application-default login
  • Check availability with verifyvertexmodel.sh

2. Find Model Information

Search for official documentation:

# Find API model name and pricing
scripts/find_model_info.sh "GPT-5.1 API documentation pricing"

What to gather:

  • Exact API model name (e.g., gpt-5.1 not GPT-5.1)
  • Provider (openai, anthropic, google)
  • Input price per 1K tokens
  • Output price per 1K tokens
  • Context limits (if relevant)
  • Special features (adaptive reasoning, caching, etc.)

Reference: See [resources/providerendpoints.md](resources/providerendpoints.md)

2a. Reasoning-Model Check (REQUIRED before gating any verdict)

The GLM-5.2 lesson (v0.30.0, 2026-07-19): GLM-5.2 was rejected in June as "worse than 5.1" — but the regression was OUR truncation: its always-on thinking phase (28–32K tokens) shared a 32,768 maxoutputtokens budget with content, and the harness recorded neither reasontokens nor finishreason, so the guillotine was invisible. With 64K headroom, 5.2 beats 5.1 on every axis. Kimi K3 nearly repeated this (thinks by default, same 32K cap, unmeasured).

Before writing any gate verdict for a new model:

  1. Probe default thinking — one small OpenRouter/API call, no reasoning params:

``bash # reasoningtokens > 0 → the model thinks by default curl -s https://openrouter.ai/api/v1/chat/completions -H "Authorization: Bearer $OPENROUTERAPIKEY" \ -H "Content-Type: application/json" \ -d '{"model":"<apiname>","messages":[{"role":"user","content":"Prove there are infinitely many primes, then state the 10th prime."}],"maxtokens":3000}' \ | jq '{reasoning: .usage.completiontokensdetails.reasoningtokens, provider}' ``

  1. **If it thinks: ground maxoutputtokens in the DECLARED provider ceilings,

not a guess. Context window ≠ output cap ≠ thinking budget — a "1M context" model may declare no completion cap at all (Kimi K3 via Moonshot: undeclared → generation is bounded only by request maxtokens + remaining window). ``bash # Per-provider documented ceilings for this model (machine-readable): curl -s "https://openrouter.ai/api/v1/models/<apiname>/endpoints" \ | jq -r '.data.endpoints[] | "\(.providername): ctx=\(.contextlength) maxcompletion=\(.maxcompletiontokens // "undeclared")"' ` Providers DIVERGE (glm-5.2: 32768 on DeepInfra vs 131072 on Z.AI first-party vs 1M on others) — routing roulette means the effective cap varies per request unless you pin provider order. Set maxoutputtokens from the first-party / dominant provider's declared ceiling, sized for thinking + content (never the fleet-default 32768), and record the provenance in the models.yml comment. Also check the vendor's own docs for the intended control surface: some models bound thinking by an effort dial** (K3: Low/Standard/High/Max via reasoningeffort in models.yml), Anthropic by budgettokens, not by output headroom. Reasoning bills as OUTPUT tokens — at K3's $15/M a chatty trace can cost more than the visible answer, so an eval that doesn't pin effort isn't reproducible; leave vendor-default unless deviating, and if you deviate, record it in models.yml (reasoningeffort`).

  1. Leave default thinking ON. A thinking-tuned model with thinking suppressed

is not the model you're gating. The knob that is actually enforced is maxtokens headroom — OpenRouter third-party upstreams (Baidu, StreamLake) ignore reasoning: {maxtokens: N} (probed 2026-07-19); reasoningmaxtokens in models.yml is best-effort only.

  1. **Read finishreason + reasontokens in every failure before concluding

capability** (recorded on standard results since v0.30.0). A finish=length failure with huge reason_tokens is truncation, not weakness. Results banked before v0.30.0 cannot show this — never re-derive a verdict from them alone.

  1. Cost note: thinking bills as output tokens. Budget/cost projections for a

reasoning model must use output + reasoning, and expect ~2-10x the completion tokens of a non-thinking peer.

3. Update models.yml

Add the model configuration:

# Add to models.yml
scripts/update_models_yml.sh \
  <friendly-name> \
  <api-name> \
  <provider> \
  <input-per-1k> \
  <output-per-1k>

Naming conventions:

  • Friendly name: gpt5-1, claude-sonnet-4-5, gemini-3-pro
  • API name: Exact string for API calls
  • Use hyphens, lowercase

Also update:

  • Model suites (benchmarksuite, extendedsuite, dev_models)
  • Add notes about special features
  • Document agent CLI support (if available)

4. Run Test Benchmark

Verify end-to-end:

# Test with a simple benchmark
scripts/run_test_benchmark.sh <model-name>

What to verify:

  • Benchmark completes successfully
  • Results are reasonable (not garbage output)
  • Token usage matches expectations
  • Cost calculation works
  • No errors in logs

5. Apply the Smoke-Test Gate (HARD RULE)

Rule of thumb (project-wide):

Rule out adding a model to our eval suite if it can't pass ALL the smoke tests.

The smoke test is the canonical smoke tier — benchmarks tagged tier: smoke in their YAML spec, selected with --tier smoke (NOT a hand-typed --benchmarks list). These are the fundamental "can it speak AILANG at all" tests that the established frontier tier passes cleanly. If a candidate fails them, the failure is on the model, not the harness or the benchmark.

The smoke tier is the source of truth — do NOT hardcode a benchmark list. Run ailang eval-suite --tier smoke --dry-run to see the current set (23 as of 2026-06-16, up from 17 — it grows, so always derive it, never trust this number). It includes fizzbuzz, adtoption, gcdlcm, nestedrecords, recordupdate, recursionfibonacci, typesaferecordaccess, balanced_parens, etc. — all fundamental.

⚠️ csvtojson_converter is tier: core, NOT tier: smoke. An earlier
version of this skill hardcoded fizzbuzz,adtoption,csvtojsonconverter as
"the smoke set." That was wrong: csvtojson is a core-tier discriminator that
the majority of frontier models fail in standard mode (gpt5 base, gemini-3-pro,
gemini-3-flash, sonnet-4-5, gpt5-mini all FAIL it; only the top tier —
opus-4-6/4-7, sonnet-4-6, gemini-3-1-pro, gpt5-2-codex/gpt5-4 — pass). Gating OS
models on csvtojson means "be top-3-tier or be cut," which unfairly excludes
viable models. Keep csvtojson in --tier core runs for ranking, never as
an include/exclude gate. (Empirically verified against eval baselines 2026-06-02.)

Run smoke against a candidate (canonical tier):

Standard mode accepts --tier smoke directly. Agent mode requires an explicit --benchmarks list (deliberate cost guardrail), so derive it from the tier: tags — never hardcode (the list drifts):

# Derive the smoke set from the tier tags (works for both modes, stays in sync)
SMOKE=$(grep -l 'tier: smoke' benchmarks/*.yml | xargs -n1 basename | sed 's/\.yml$//' | paste -sd, -)

# Standard mode:
ailang eval-suite --models <candidate>,claude-sonnet-4-6 --tier smoke \
  --langs ailang --output /tmp/smoke_<candidate> --parallel 2

# Agent mode (must pass the derived list explicitly):
ailang eval-suite --agent --models <candidate>,claude-sonnet-4-6 \
  --benchmarks "$SMOKE" --langs ailang --output /tmp/smoke_<candidate> --parallel 2

# Tabulate pass/fail (agent mode → results land under /agent, standard → /standard)
for f in /tmp/smoke_<candidate>/*/*.json; do
  name=$(basename "$f" .json | sed 's/_[0-9]*$//')
  jq -r --arg name "$name" '"\($name)\t\(if .compile_ok and .runtime_ok and .stdout_ok then "PASS" else "FAIL" end)\t\(.err_code // .error_category // "—")"' "$f"
done | column -ts $'\t'

Decision tree (N = number of benchmarks in the smoke tier, 23 as of 2026-06-16 — derive it, don't assume):

  1. claude-sonnet-4-6 fails any smoke-tier benchmark — smoke tier is broken;

fix the benchmark (or its tier: tag) before evaluating candidates.

  1. Candidate fails most of the tier — CUT. Do not add to models.yml. Note the

failure types in the cut commit message (WRONG_LANG, syntax, runtime, wrong-output) for future reference.

  1. Candidate fails 1–2 (near-clean) — NEAR-MISS. Optionally keep with a

"near-miss" comment block in models.yml (precedent: motoko-or-gemma-4-26b, motoko-or-qwen3-5-35b-a3b). Re-run periodically; if it starts passing the tier clean, that's a signal stdlib/prompt has improved. Note: agent-mode failures with errorcategory: apierror + "step budget exhausted" are a harness step-budget cap, not a model gap — don't count them as capability failures (bump the motoko v2 step budget instead).

  1. Candidate passes the tier clean — it has cleared the FLOOR, nothing more.

Smoke qualifies a model; it does NOT rank it (see step 5.5). Add the opt-in models.yml entry now (gate PASSED, promotion PENDING), then go to step 5.5 to decide whether it actually earns a suite slot / replaces an incumbent.

Failure-mode taxonomy (worth capturing in the cut commit message):

Failure Meaning Likely cause
WRONG_LANG Model produced Python/JS instead of AILANG Prompt-following gap; small/MoE models lose plot at 23k-token system prompt
syntax-error (no WRONG_LANG) Invented AILANG syntax (e.g. let rec, \n. lambda) Model hasn't seen enough AILANG in training
wrong-output Compiled and ran, wrong stdout Spec-following gap, not language gap
runtime-error Compiled, crashed at runtime Logic bug

2026-05-04 finding (precedent): Tested 6 SOTA OS models (Gemma 4 26B, Qwen3 30B-A3B, Qwen3 235B-A22B, DeepSeek V4 Flash, Kimi K2.6, Qwen3 Coder Flash) against this smoke set. Proprietary baselines passed 3/3; zero OS models passed all 3. Most common failure: WRONG_LANG (model produced Python). Even frontier-class OS models fall back on training-corpus patterns when given AILANG's 23k-token teaching prompt — they've seen plenty of Python but very little AILANG. Two near-misses (or-gemma-4-26b, or-qwen3-coder-flash) retained on the watchlist; rest cut.

Implication for stdlib/prompt work: the smoke test doubles as a language-improvement metric. Re-run it after stdlib changes or prompt revisions; if the near-miss watchlist starts passing the third benchmark, the language has become more "trainable-feel."

Caveat — agent mode is a separate gate: the smoke set above runs in standard (single-shot API generation) mode. Models that fail standard mode may still perform usefully in agent mode (--agent flag, opencode/pi harnesses) where they get multi-turn iteration. If a candidate fails standard smoke, run ailang eval-suite --agent --models <candidate> ... separately before fully cutting it. Agent mode results don't override the standard-mode gate but can justify adding the model under a different harness entry (e.g. opencode-<candidate>, pi-<candidate>).

2026-05-04 agent-mode smoke finding (precedent): Tested 9 OS-via-OR candidates through opencode harness. Cross-mode behaviour:

| Model | Standard | Agent | Δ | |-------|---------:|------:|--:| | GLM 5 (z.ai) | not tested | 3/3 ✅ | — first OS model to pass | | Gemma 4 26B | 2/3 | 2/3 | 0 (same near-miss) | | DeepSeek V4 Flash | 0/3 | 2/3 | +2 (agent unlock) | | GLM 4.7 Flash | not tested | 2/3 | — near-miss | | Kimi K2.6 | 1/3 | 1/3 | 0 | | Qwen3 30B-A3B | 1/3 | 1/3 | 0 | | Qwen3 Coder Flash | 2/3 | 1/3 | -1 (agent regressed) | | DeepSeek V4 Pro | not tested | 1/3 | Pro under-performed Flash | | Qwen3 235B-A22B | 0/3 | 0/3 | 0 |

Key takeaways for the model-manager workflow:

  1. Agent mode is not a universal fix. Most models that fail standard

smoke also fail agent smoke. Multi-turn helps when the model can read compile errors and adjust; it hurts when the model interprets tool-call setup as the answer (Qwen3 Coder Flash regression).

  1. Pro tier ≠ better. DeepSeek V4 Pro (1/3) under-performed V4 Flash

(2/3) on AILANG smoke. The Pro reasoning/long-output overhead can hurt simple-task accuracy. Test both tiers when available.

  1. csvtojson_converter is a core-tier DISCRIMINATOR, not a smoke gate.

Of the 27 benchmark runs (9 models × 3), csvtojson was the single most-failed test — only GLM 5 passed it among OS candidates. ⚠️ CORRECTION (2026-06-02): this is exactly why it must NOT gate inclusion — it's failed by the majority of frontier models (gpt5 base, gemini-3-pro, gemini-3-flash, sonnet-4-5, gpt5-mini all FAIL; only opus-4-6/4-7, sonnet-4-6, gemini-3-1-pro, gpt5-2-codex/gpt5-4 pass). It lives in tier: core, not tier: smoke. Use it as a high-signal ranking/discriminator metric in --tier core runs and as a language- improvement tracker — never as an OS-model include/exclude gate. The gate is --tier smoke.

  1. GLM 5 is genuinely cost-competitive frontier OS. $0.60/$2.08 per 1M

tokens, ~5–7× cheaper than Claude Sonnet 4.6 on input. Worth standing inclusion in eval rotation alongside frontier proprietary models.

  1. Vendor-prefix wiring is forward-compat infrastructure. When adding

models from a new vendor (e.g. z-ai/, moonshotai/, microsoft/, minimax/), add the prefix to internal/ai/config.go::openrouterVendorPrefixes so future ad-hoc ailang run --ai vendor/model invocations work without needing a models.yml entry.

  1. Per-benchmark timeouts can be tighter than agent-mode needs. The

csvtojsonconverter.yml spec has timeout: 90s baked in (set to match Claude Sonnet 4.6's ~43s typical solve time). OS models in agent mode routinely need 90–180s of iteration on csvtojson — they CAN solve it but get killed by the timeout. Two follow-up models that demonstrated this on 2026-05-04: - Kimi K2.6 (Moonshot): fizzbuzz✅ 119s, adtoption✅ 47s, csvtojson❌ (timeout — initial run also had apierrors) - MiniMax M2.7: fizzbuzz✅ 46s, adtoption✅ 42s, csvtojson❌ (timeout, not capability) Both are effectively 2/3 near-misses pending a benchmark timeout bump. When investigating a model that fails only csvtojson with errorcategory=apierror and stderr saying "exceeded hard timeout (1m30s)", the failure is the benchmark spec, not the model.

  1. apierror vs syntax-error vs WRONGLANG matters. When tabulating

smoke results, always check errorcategory: - apierror — infrastructure issue (rate limit, timeout, network). Re-run before counting against the model. - compileerror (no errcode) — syntax-error: model produced AILANG that doesn't parse. Genuine model gap. - WRONGLANG — model produced Python/JS/etc. instead of AILANG. Genuine prompt-following gap. - runtimeerror — compiled but crashed. Logic bug in generation.

5.5 Smoke is a FLOOR, not a RANKING — place the candidate on the ANCHORED ELO series (HARD RULE)

**Passing smoke is necessary but NOT sufficient. Smoke says "this model can
speak AILANG at all"; it does NOT say "this model is good enough to add" or
"this model beats the incumbent." Those are RANKING questions, and smoke is
saturated — every frontier-class model scores ~the same on it. Never make an
add/keep/replace decision on smoke numbers.**

Why: the smoke tier is deliberately fundamental ("can it speak AILANG"), so any viable model passes ~all of it. A smoke tie is the expected outcome, not a signal — it carries zero ranking information. The discriminator is --tier core (~26 benchmarks incl. csvtojson_converter, the contract/state-machine tests), plus --tier frontier when you want placement at the hard end.

⛔ Run the CANDIDATE ALONE. Do not re-run incumbents or anchors.

M-EVAL-ROLLING-ELO (landed 2026-08-27, PRs #939/#942) changed this. ELO ratings used to be incomparable across fits — FitFromTrials seeded every model and benchmark at 1500 with no scale anchor, so the same rows produced different absolute numbers in different pools (measured: or-glm-5-3-flash rated 2763 in one pool vs 1995 in another over comparable rows). That is why the old protocol re-ran candidate + incumbent + anchor together — a shared pool was the only way to make numbers mean anything.

That is no longer true, and doing it now is pure waste:

  • internal/evalharness/anchorv1.json freezes the fitted difficulties of the

discriminating standard benchmarks. Standard-mode fits hold that panel fixed and let model ratings move ([cmd/ailang/evalelo.go:172](../../../cmd/ailang/evalelo.go)).

  • So a candidate run alone is placed onto the same scale as every model

ever measured. Anchored drift is 31.2 ELO vs 311.7 unanchored.

  • D3 retired full baselines as the default release measurement. The full run

is demoted to quarterly re-anchoring (make eval-baseline FULL=true).

# CORRECT — candidate only; the anchored fit places it against banked history.
ailang eval-suite --models <candidate> --tier core,frontier --langs ailang \
  --output /tmp/cf_<candidate> --parallel 4

ailang eval-elo /tmp/cf_<candidate> --json     # read-only ANCHORED placement fit
go run ./tools/eval-elo --persist /tmp/cf_<candidate>   # bank it into the series

ailang eval-elo does NOT persist — it refuses --persist and tells you to use tools/eval-elo. Persisting through the cmd path silently no-ops (it swallowed --persist for weeks; see projectagentratingsseedingevalelo).

Compare the resulting rating to the banked ratings already in observatory.db (LoadModelRatings) — that is what the series is for.

⚠️ CHECK THE BANKED SERIES IS ACTUALLY ANCHORED BEFORE COMPARING. An
anchored placement and a pre-anchor banked rating are on DIFFERENT SCALES, and
nothing in the output warns you. Measured 2026-09-01 on the dev box: every
model_ratings row for mode='standard' was stamped 2026-08-03 — before
the anchor landed (2026-08-28) — so those 19 values are unanchored, and
trial_history held agent-mode rows only (739 rows, 3 models), meaning
there were no banked standard trials to re-level them from. Reading Hy4's
anchored 1915.5 against that table would have been exactly the 2763-vs-1995
error the anchor exists to prevent.

```bash
sqlite3 ~/.ailang/state/observatory.db \
"SELECT modelid, ROUND(rating,1), ntrials, substr(last_updated,1,10)
FROM model_ratings WHERE mode='standard' ORDER BY rating DESC;"
sqlite3 ~/.ailang/state/observatory.db \
"SELECT mode, COUNT(*), COUNT(DISTINCT modelid) FROM trialhistory GROUP BY mode;"
```

If lastupdated predates the anchor, or trialhistory has no rows for the
mode you are placing in, you have a placement but no valid comparison set.
Say so plainly rather than ranking against stale numbers. Report the candidate's
pass profile against the anchored benchmark difficulties (which the fit does
give you) and treat the leaderboard position as unavailable until the series is
re-fit. Do NOT "fix" this by re-running comparators — that is the anti-pattern
above; the fix is banking standard-mode trials so the series can accumulate.

When you DO still co-run models

Three cases, and only these:

  1. Agent mode. There is no agent-mode anchor yet — eval_elo.go:171-172

anchors standard mode only, agent fits are unanchored. For an agent-mode ranking question the old same-pool rule still holds.

  1. The incumbent's banked rating predates a baseline-moving language change.

The anchor pins benchmark difficulty, not the harness or stdlib. A change that moves what models can do (e.g. the 2026-07-29 extension fix) invalidates older per-benchmark rates. Check the banked rating's version/date provenance first; if it is stale, re-run just the incumbent — never the whole panel.

  1. A paired/discordant analysis, where ailang eval-paired <on> <off> needs

both arms from the same run by construction.

When you genuinely do run several models, they must still go in ONE eval-suite command — it overwrites its output directory (.claude/rules/eval.md). That rule is about not clobbering results; it is not a reason to add models.

Deciding promote / replace

  • Match-or-beat the incumbent to promote. A tie at higher cost keeps the

incumbent; a tie at equal-or-lower cost favours the newer generation.

  • Close call on N=1 — core is ~26 single-shot runs and OS-model variance is

real. Within 1–2 benchmarks (or overlapping ELO bands), escalate to N≥3 before deciding. Never flip an incumbent on a 1-benchmark N=1 delta.

  • Read finishreason + reasontokens on every failure before calling it

capability (§2a). A finish=length with large reasoning is truncation.

⚠️ Anti-pattern (2026-06-16, GLM-5.2 vs GLM-5.1): GLM-5.2 cleared standard
smoke at 22/23 — an exact tie with GLM-5.1 (both failed only
denseoperatorprogram, which the anchor ALSO failed → a benchmark/harness
issue, not a model gap). The first-pass conclusion was *"tie at +43% cost →
keep GLM-5.1."* That was WRONG. A smoke tie is meaningless because smoke is
saturated — it proves only that the candidate cleared the floor and QUALIFIES.
When a candidate ties the incumbent on smoke, that is your cue to run core, NOT
your answer.

⚠️ Anti-pattern (2026-09-01, Hy4 preview): after Hy4 passed smoke 23/23, the
placement run was launched as six models × 31 benchmarks = 186 runs, ~$3.75 —
candidate plus incumbent plus four comparators, on the pre-rolling-ELO reflex
that a shared pool was needed. It was not: five of those six models already had
banked anchored ratings, and the candidate alone costs $0.23. Killed at
$0.24. A comparator you re-run is a comparator you pay for twice.

6. Document the Model

Update relevant documentation:

  • Add model to this skill's resource guide
  • Note any special parameters (e.g., maxcompletiontokens for GPT-5.1)
  • Document authentication requirements
  • Add to teaching prompts if needed

7. Bank the placement — do NOT run a full baseline

Once the candidate has its core/frontier placement, persist it into the anchored series so the next question can be answered from banked data instead of a re-run:

go run ./tools/eval-elo --persist /tmp/cf_<candidate>

make eval-baseline FULL=true is NOT part of adding a model. D3 of M-EVAL-ROLLING-ELO demoted the full baseline to a quarterly re-anchoring event (and longitudinal spot-checks). Running one to place a new model costs $5-25 + hours of wall clock to produce a number the anchored fit already gives you for well under a dollar. If you think you need a full baseline, you almost certainly need a linking run instead — re-read §5.5.

Resources

Provider Endpoints

See [resources/providerendpoints.md](resources/providerendpoints.md) for:

  • API endpoint URLs for each provider
  • Authentication methods
  • How to test access manually
  • Common errors and fixes

Pricing Guide

See [resources/pricingguide.md](resources/pricingguide.md) for:

  • How to find official pricing
  • Price conversion (per 1M → per 1K)
  • Cost calculation verification
  • Caching and discounts

Progressive Disclosure

This skill loads information progressively:

  1. Always loaded: This SKILL.md file (workflow and script descriptions)
  2. Execute as needed: Scripts in scripts/ (testing, updating, verification)
  3. Load on demand: Resources (detailed endpoint docs, pricing references)

Notes

Important:

  • Always test API access BEFORE updating models.yml
  • Vertex AI (Gemini) requires gcloud auth, not API key
  • GPT-5.1+ uses maxcompletiontokens instead of max_tokens
  • New models may not be available in all regions immediately
  • Check for preview/beta status before adding to production suites

Prerequisites:

  • API keys set in environment (OPENAIAPIKEY, ANTHROPICAPIKEY)
  • For Gemini: gcloud CLI installed and authenticated
  • For Gemini: GCP project set (gcloud config set project PROJECT_ID)
  • curl, python3, and jq available in PATH

Files modified by this skill:

  • internal/eval_harness/models.yml - Model configurations
  • (Optional) prompts/vX.Y.Z.md - Teaching prompts
  • (Optional) .claude/skills/model-manager/resources/ - Local model database