Create Skill Test
Scaffold an evaluation spec (eval.yaml) for a skill or agent so it conforms to the Vally schema, passes skill-validator check and checkevalquality.py, is powerful enough to return a verdict, and does not overfit to the skill's own wording.
When to Use
- Creating a new
eval.yaml for a skill or agent
- Adding stimuli to an existing eval
- Sizing an eval so the pass gate can actually be reached
- Setting up or repairing fixture files alongside an eval
- Reviewing whether rubric items and graders risk overfitting
When Not to Use
- Diagnosing a failing or regressed eval — use
improve-skill-quality
- Modifying the skill-validator or the evaluation workflows
- Creating or editing
SKILL.md files — use create-skill
Inputs
| Input |
Required |
Description |
| Skill or agent name |
Yes |
Must exist under plugins/<plugin>/skills/ or plugins/<plugin>/agents/ |
| Plugin name |
Yes |
e.g. dotnet-msbuild |
| Skill content |
Yes |
Read it — you cannot write non-overfitted rubric items without it |
| Failure modes to discriminate |
Recommended |
Each becomes one stimulus |
Workflow
Step 1: Locate the target and the test directory
tests/<plugin>/<skill-name>/eval.yaml # skills
tests/<plugin>/agent.<agent-name>/eval.yaml # agents (the agent. prefix disambiguates)
Verify the target exists at plugins/<plugin>/skills/<skill-name>/SKILL.md or plugins/<plugin>/agents/<agent-name>.agent.md, and read it.
Agent evals sit outside the verdict flow. The canonical experiment declares evals: tests//!(agent.)/eval.yaml, so agent.* specs are excluded: no verdict is ever computed for them, the stimulus floor does not apply, and ./eng/run-skill-evals.sh drops them even when you name one explicitly (its --eval-filter is intersected with that glob). The distinct-stimulus floor therefore applies to skill evals only. Author agent evals for the scenario coverage and the deterministic graders, and run them as described in Step 10.
Be careful with a skill that sets disable-model-invocation: true. The model cannot invoke it, so the skill is absent from the model-facing skilled arm and any direct eval compares two identical arms. Answer-content graders do not create a difference between those arms. The honest coverage for such skills is dependency-level — through the outcome evals of the skills that load them, and through the plugin arm. For example, filter-syntax is covered by the filtered-command scenarios in tests/dotnet-test/run-tests/eval.yaml.
Step 2: Write the spec skeleton
The spec is Vally format. Every eval in this repo uses stimuli: and graders:; scenarios: and assertions: are a pre-Vally format that no longer loads.
name: <skill-name>
description: Evaluates the <plugin>/<skill-name> skill
type: capability
defaults:
timeout: 5m
runs: 1
stimuli:
- name: <what the agent must accomplish>
prompt: <natural developer request>
environment:
files:
- src: fixtures/<case>/Project.csproj
dest: Project.csproj
graders:
- type: output-matches
config:
pattern: (root cause|underlying issue)
- type: exit-success
- type: prompt
rubric:
- <outcome the agent should have reached>
defaults: replaces config: — it does not join it. config is a deprecated alias for the
same block and vally throws on a spec declaring both. Some existing evals still open with
config:; when you change settings, replace it with one defaults: block. The failure is
invisible otherwise: the job exits 0 with no verdicts and the PR comment
blames "transient infrastructure".
Step 3: Size the eval for power before writing content
The gate gives each distinct stimulus one vote. Repeated runs for one stimulus collapse to one majority-direction vote and remain available as reliability evidence.
- Distinct stimuli ≥ 5, else the verdict is
underpowered — never a pass, never a regression.
- **p ≤ 0.05 on an exact one-sided sign test over discordant (non-tie) stimulus votes.** Ties are not
discarded; they hold the discordant count down.
| discordant stimulus votes |
records that pass |
p |
| ≤ 4 |
none |
≥ 0.0625 |
| 5–7 |
zero losses only (5W/0L) |
0.031 |
| 8 |
one loss survivable (7W/1L) |
0.035 |
At exactly 5 stimuli, one tie is fatal because it leaves 4 discordant votes. At 6 stimuli one tie is survivable; at 7, up to two are. A loss is not. Five is an eligibility floor, not adequate power. For example, 80% power needs 8 discordant votes only for a true 90% conditional win rate; it needs 18 at 80%, 37 at 70%, and 158 at 60%. Size for the effect and tie rate you need to detect.
Use runs for reliability, not task breadth. Vally recommends 3 runs in CI and 5–10 nightly for pass rate, pass@k, pass^k, and flakiness. Extra runs never clear the five-stimulus floor.
Do not set runs in dotnet-skills.experiment.yaml; experiment overrides overwrite every eval's own value rather than defaulting it.
Step 4: Write stimuli
- Name describes what is tested, not how.
- Prompt is a natural developer request. Never mention the skill, the agent, or its vocabulary —
cued prompts inflate the overfit score and bias the baseline.
- Each stimulus should discriminate a different property of the skill. Five stimuli covering one
property give arithmetic, not evidence.
- Give every stimulus a stable, unique
name. Vally pairs comparison trajectories by
(stimulus name, trial index); duplicate names make slot identity ambiguous.
- Include a boundary / no-op stimulus for any skill that migrates or rewrites code, proving it
leaves already-correct input alone.
Step 5: Configure the environment
environment:
files:
- src: fixtures/broken-build/App.csproj # path relative to eval.yaml
dest: App.csproj # path in the agent's working directory
- src: fixtures/broken-build # a directory
dest: .
commands:
- dotnet build -bl || exit 0 # guard intentional failures
Do not set environment.skills in a skill eval. The experiment declares vary: /environment/skills and supplies the value itself — [] for the baseline arm and plugins/<plugin>/skills/<skill> for the skilled arm — so anything the eval declares is replaced, in every arm. It cannot add a skill to one arm only. environment.skills is meaningful only in an agent.* eval, which the experiment does not vary; there it is the set of skills the agent may invoke. Copy the shape from an existing agent eval such as tests/dotnet-test/agent.test-quality-auditor/eval.yaml rather than reproducing a remembered form — the specs in this repo are not consistent about how they spell those entries.
Fixture rules — each one has already cost a real result:
- Every referenced fixture must be tracked by git.
.gitignore (e.g. coverage*.xml) has
silently swallowed a committed fixture: the eval passed locally and failed at setup in CI. Verify with git ls-files, not by looking at the working tree.
- Every fixture must behave as its stimulus assumes. A fixture meant to be healthy must build; a
fixture meant to be broken must fail for the exact reason the stimulus is about, and no other. Judges penalize agents for unrelated "pre-existing build issues" that the fixture author introduced.
- Every fixture must reproduce the bug its stimulus is named for. If it does not, the baseline
scores well and the skill has nothing to add.
- Coverage fixtures must be internally consistent. A Cobertura report whose declared
line-rate, summary totals (lines-covered/lines-valid), and <line> elements disagree lets the two arms read different truths, and the loss is the fixture's fault. Update any rubric item or prompt that quotes a figure in the same change.
- Do not wire duplicate fixtures to raise
n; rename leftovers add trials without evidence.
- A setup command that is expected to fail while still producing its artifact must be guarded
(|| exit 0), or vally drops the trial.
- A cleanup command that strips sources must skip directories containing
SKILL.md — the staged
skill lives there, and deleting it aborts only the skilled arm.
Step 6: Write graders
Graders are hard pass/fail checks evaluated on every arm.
| Type |
Required config |
Purpose |
output-matches / output-not-matches |
pattern |
Regex over agent output |
output-contains / output-not-contains |
substring |
Literal text in output |
file-exists / file-not-exists |
path |
Glob against the work directory |
file-contains / file-not-contains |
path, value |
Content of a produced file |
run-command |
command (plus optional expectedexitcode, timeout, stdout_matches) |
Verify produced code actually builds/runs |
exit-success |
— |
Agent produced non-empty output |
prompt |
— |
Runs the LLM judge against the rubric |
Rules:
- A grader whose
config is absent or missing its required key parses fine and enforces nothing.
The usual cause is an indentation slip during an edit; checkevalquality.py blocks it.
- Prefer broad patterns that several valid approaches satisfy:
(root cause|primary error|underlying issue).
- If the skill mandates an output shape, assert on it. A skill required to emit a decisive
Recommendation: line can silently stop doing so while the eval still passes.
- Use
file-not-contains / file-not-exists to prove the agent avoided an incorrect action.
Step 7: Write rubric items
Rubric items are judged pairwise (baseline vs. skilled). The overfitting judge classifies each item:
| Classification |
Description |
Goal |
| outcome |
Whether the agent reached a correct result — WHAT, not HOW |
Target this |
| technique |
Whether the agent used a skill-specific procedure |
Minimize |
| vocabulary |
Whether the agent used the skill's terminology |
Avoid |
- Test outcomes, not methods: "Identified the root cause of the build failure", not "Replayed the
binlog using dotnet build /flp".
- Accept any valid approach.
- Never reference the skill by name, and never reuse
SKILL.md phrasing.
- Never reward using the skill — the harness reports activation separately, so a rubric item that
does this measures nothing and inflates the overfit score.
- Do not test knowledge the model already has; it adds no delta.
- Keep each item independently evaluable.
- Do not reward raw volume (test count, report length); judges will compare it when both arms act.
Good:
rubric:
- Correctly identified the missing NuGet package as the root cause of the build failure
- Recognized that downstream failures cascaded from that root cause
- Suggested a concrete fix that resolves it
Overfitted:
rubric:
- Replayed the binary log using 'dotnet build /flp:v=diag' # technique
- Measured cold, warm, and no-op build scenarios # vocabulary
- Used the template-comparison skill # rewards activation
Step 8: Add constraints sparingly
constraints:
expect_tools: [bash]
reject_tools: [edit, create]
reject_skills: [some-skill]
expect_tools: [bash] on an advisory question forces a restore or build and converts an
answer into a timeout with no quality benefit. Only require tools when the task genuinely needs them.
reject_tools is the right way to keep a read-only stimulus read-only.
Step 9: Add dormancy guards
A dormancy guard proves the skill stays dormant on an off-target request that superficially matches it. Add one per real "when not to use" boundary: wrong input format, out-of-scope request, incompatible project type, wrong framework version, prerequisite absent.
- name: Decline dump analysis request
prompt: |
I already have a .dmp crash dump from my .NET app. Can you help me
analyze it to find the root cause of the crash?
expect_activation: false
graders:
- type: output-matches
config:
pattern: (out of scope|not cover|does not|cannot|only.*collect)
- type: prompt
rubric:
- Stated that dump analysis is out of scope
- Did not open or analyze the dump file
- Did not install analysis tools such as dotnet-dump analyze, lldb, or windbg
- Suggested the correct alternative
Never combine expectactivation: false with constraints.rejectskills. That forces the
skilled arm to run skill-free, so the harness cannot observe whether the target skill hijacks the
request. The comparison remains visible as report-only evidence but does not vote in preference;
unexpected isolated activation blocks a pass. expect_activation: false alone is the repo
convention.
Guard rubrics verify three things: recognition (why it does not apply), restraint (no workflow, no file changes, no installs), redirection (the correct next step).
Step 10: Validate
dotnet run --project eng/skill-validator/src/SkillValidator.csproj -- check --plugin ./plugins/<plugin>
python eng/eval-quality/check_eval_quality.py
./eng/run-skill-evals.sh <plugin> <skill-name>
For an agent eval, the third command is a no-op: agent.* is outside the experiment's evals: glob. Exercise one by pointing the runner at an experiment file whose glob includes it:
# copy dotnet-skills.experiment.yaml, widen its evals: glob to tests/*/agent.*/eval.yaml
EXPERIMENT_FILE=my-agent.experiment.yaml ./eng/run-skill-evals.sh <plugin>
Read the trajectories rather than the verdict — there is no sign-test result for an agent eval.
checkevalquality.py blocks eleven structural defect classes that can corrupt a result: missing or untracked fixtures, self-contradicting coverage fixtures, empty grader configs, dormancy guards with reject_skills, sub-floor stimulus counts, duplicate YAML keys or stimulus names, and config:/defaults: collisions. Do not add a new eval to eng/eval-quality/underpowered-allowlist.txt — the gate rejects allowlist entries that are new relative to the base branch.
For the official run, submit a PR review containing /evaluate so it binds to the reviewed commit.
Validation Checklist
Common Pitfalls
| Pitfall |
Solution |
Writing scenarios: / assertions: |
That format no longer loads; use stimuli: / graders: |
Adding defaults: runs: beside an existing config: |
Merge into one defaults: block |
| Landing an eval at exactly 5 stimuli |
A single tie makes a pass unreachable; size for the effect and tie rate |
Raising runs to clear the floor |
Repeats measure reliability for one task; add stimuli |
| Prompt mentions the skill or agent by name |
Rewrite as a natural developer request |
| Rubric rewards using the skill |
Drop the item — the harness reports activation separately; rubrics measure outcomes |
| Fixture present but ignored by git |
Verify with git ls-files; CI setup will fail otherwise |
| Fixture that does not build, or breaks for the wrong reason |
Fix the fixture before blaming the skill |
Dormancy guard with reject_skills |
Use expect_activation: false alone |
expect_tools: [bash] on an advisory question |
Drop it; it causes timeouts, not quality |
| Timeout too short for code generation |
Use ~360s; empty output fails every grader |
| Duplicate YAML key left behind by an edit |
It overwrites the next stimulus field by field — delete the stray block |
| Duplicate stimulus names |
Vally uses names as comparison identity — give every stimulus a stable, unique name |
Direct eval for a disable-model-invocation: true skill |
Remove it and cover the reference through consumer outcomes |
| Agent eval sized for the stimulus floor |
agent.* evals get no verdict; size them for scenario coverage instead |
Agent eval "run" with ./eng/run-skill-evals.sh |
The glob drops it — use a widened EXPERIMENT_FILE |
Agent eval missing environment.skills |
Declare the skills the agent routes to, or it cannot invoke them |
environment.skills set in a skill eval |
The experiment varies that key and replaces it in every arm; the declaration does nothing |