SKILL.md
Testing Procedures
Invoke with /tzurot-testing for test-related procedures.
Testing patterns are in .claude/rules/02-code-standards.md - they apply automatically.
Running Tests
# Run all tests
pnpm test
# Run specific service
pnpm --filter @tzurot/ai-worker test
# Run specific file
pnpm test -- MyService.test.ts
# Run the component tier (PGLite) — NOT included in `pnpm test`
pnpm test:component
# Run the integration + contract tiers (real Redis/Postgres) — NOT in `pnpm test`
pnpm test:integration
# Run with coverage
pnpm test:coverage
# Run only changed packages
pnpm focus:test
Always invoke through the package's own config. A root-invoked npx vitest run <pkg path> resolves the ROOT config, which has no setupFiles — the package's setup (dotenv, the global afterEach(vi.clearAllMocks)) never loads, hoisted-mock call history accumulates across tests in a file, and expect(mock).not.toHaveBeenCalled() fails on calls leaked from earlier tests. It masquerades as flaky or pre-existing breakage; it is an invocation artifact. Use pnpm --filter @tzurot/<pkg> test, or cd services/<pkg> && npx vitest run <relative-path> — and never diagnose "pre-existing failures" from a root-invoked run.
Coverage Audit Procedure
# Run unified audit (CI does this automatically)
pnpm ops test:audit
# Filter by category
pnpm ops test:audit --category=services
pnpm ops test:audit --category=contracts
# Update baseline (after closing gaps)
pnpm ops test:audit --update
# Strict mode (fails on ANY gap)
pnpm ops test:audit --strict
Unified Baseline: test-coverage-baseline.json (project root)
Test File Types
Tiers are canonically defined in one place — [Test Tier Taxonomy](../../../docs/reference/guides/TESTING.md#test-tier-taxonomy). This table is the suffix→tier quick reference; don't re-define the tiers here (the pnpm ops guard:test-taxonomy gate enforces the single-source link).
| Suffix / location | Tier ([taxonomy](../../../docs/reference/guides/TESTING.md#test-tier-taxonomy)) | Infrastructure | Location |
|---|---|---|---|
*.test.ts |
Unit | Fully mocked | Next to source |
*.component.test.ts |
Component (one service, PGLite) | PGLite | Next to source |
*.integration.test.ts |
Integration (real DB+Redis) | Real services | tests/e2e/ |
*.contract.test.ts |
Contract (provider↔consumer) | Real services | tests/e2e/contracts/ |
Each suffix names its tier directly. Schema test ≠ contract test: a Zod
schema test (a plain*.test.ts) validates one type's own rules (unit-tier); a
contract test verifies two services agree. Runpnpm ops test:tiersfor the
per-package tier distribution.
Verify with the config that runs the tier (CRITICAL)
The unit pnpm test config EXCLUDES the .component / .integration / .contract suffixes. A contract/integration/component test "passes" under pnpm test by NOT running at all — so a green unit run is NOT evidence it works.
| Suffix | Verify with |
|---|---|
*.test.ts (unit) |
pnpm test |
*.component.test.ts |
pnpm test:component |
.integration.test.ts / .contract.test.ts |
pnpm test:integration |
When you touch a .contract / .integration test, run pnpm test:integration (Redis + Postgres up: podman start tzurot-redis tzurot-postgres); a .component test, pnpm test:component. pnpm test / pnpm --filter <pkg> test is the unit tier ONLY. CI's component-integration-tests job runs test:component + test:integration, so it catches what a unit-only local run silently skipped — don't let CI be the first thing that actually executes your contract test.
Gotcha for mention/ID fixtures: tests need a valid 17–19 digit Discord snowflake — isValidDiscordId silently drops toy ids like 555 before resolution, making the assertion pass/fail for the wrong reason.
Debugging Test Failures
1. Run Specific Test
pnpm test -- MyService.test.ts --reporter=verbose
2. Check for Fake Timer Issues
// ❌ WRONG - Promise rejection warning
const promise = asyncFunction();
await vi.runAllTimersAsync(); // Rejection happens here!
await expect(promise).rejects.toThrow(); // Too late
// ✅ CORRECT - Attach handler BEFORE advancing
const promise = asyncFunction();
const assertion = expect(promise).rejects.toThrow('Error');
await vi.runAllTimersAsync();
await assertion;
3. Reset Mock State
beforeEach(() => {
vi.clearAllMocks(); // Clear call history, keep impl
});
afterEach(() => {
vi.restoreAllMocks(); // Restore originals (spies only)
});
Creating Mock Factories
// Use async factory for vi.mock hoisting
vi.mock('./MyService.js', async () => {
const { mockMyService } = await import('../test/mocks/MyService.mock.js');
return mockMyService;
});
// Import accessors after vi.mock
import { getMyServiceMock } from '../test/mocks/index.js';
it('should call service', () => {
expect(getMyServiceMock().someMethod).toHaveBeenCalled();
});
Probe fixtures come from the real corpus
At the moment of writing a probe or test fixture — a filename, a payload, a URL, a config value — do not invent a minimal one: copy a real specimen out of the corpus and mutate only what the test needs. An input built to isolate one property is exactly the input that cannot reveal an interaction with a second property, and its green is indistinguishable from a green that covered the real case (a space-free probe filename passed end-to-end while every real tracker/ path — all space-bearing — would have been quote-mangled). If a synthetic input is unavoidable, enumerate the properties real specimens carry (spaces, non-ASCII, length, punctuation, nesting) and give it ALL of them. The refutation direction inverts this: the probe that settles a disputed finding varies a single property at a time.
Mock-reachability check (before writing assertions)
Before asserting on a subject that sits behind a mocked collaborator, answer two questions: (a) which mock makes this code path reachable, and (b) what would a wiring bug at that seam look like in this test's output? If the answer to (b) is "the test can't see it," the test asserts through the seam it mocked — add the seam assertion per 02-code-standards.md § Assert what crosses a mocked seam before writing more cases. This is the coverage-illusion class: a mock missing one property leaves the dependent branch green and untested; a feature flag can no-op silently under a fully green suite.
Stacked gates: make upstream gates inert
When asserting gate N in a multi-gate pipeline — a hook with several checks, a validator chain, a new guard stacked in front of an older one — an exit code or shared failure result is NOT attribution: the new gate inherits the older gate's failure, so assert_exit 2 (or .toThrow(), or "returns false") passes whether YOUR gate fired or the one behind it did. Two obligations:
- Construct the fixture so every upstream gate is inert (would pass on its
own), leaving the gate under test as the only possible failure source. Asserting a distinctive banner/message helps but is secondary — an inert- upstream fixture proves attribution even when messages change.
- Canary THIS gate: mutate the gate under test and confirm the specific
assertion reddens (Core Principle 9, 02-code-standards.md). A canary against the pipeline as a whole proves nothing about which gate you pinned.
Anatomy to fear: a cluster of vacuous assertions shipped in one hook PR (#2078) with exactly this shape — including one whose test NAME was false, passing only because a different gate blocked first. Reading had already passed every one of them; per-gate canaries are what caught them.
Integration Tests with PGLite
describe('UserService', () => {
let pglite: PGlite;
let prisma: PrismaClient;
beforeAll(async () => {
pglite = new PGlite({ extensions: { vector, citext } });
await pglite.exec(loadPGliteSchema());
prisma = new PrismaClient({ adapter: new PrismaPGlite(pglite) });
});
it('should create user', async () => {
const service = new UserService(prisma);
const userId = await service.getOrCreateUser('123', 'testuser');
expect(userId).toBeDefined();
});
});
⚠️ ALWAYS use loadPGliteSchema() - NEVER create tables manually!
Component Test Triggers
Component tests (*.component.test.ts) run separately from unit tests and are not included in pnpm test or pre-push hooks.
Always run pnpm test:component after:
| Change | Why |
|---|---|
| Add/remove slash command options | CommandHandler.component.test.ts snapshots capture full command structure |
| Add/remove subcommands | Same snapshot tests |
| Restructure command directories | getCommandFiles() discovery changes affect command loading |
| Change component prefix routing | Component tests verify button/select menu routing |
Update snapshots with: pnpm vitest run --config vitest.component.config.ts <file> --update
Human-Verification Requests (manual testing by the user)
The user is the ONLY manual-QA executor — usually on a phone, across session crashes and compactions. Every request for manual verification MUST be a complete, self-contained instruction with all five parts:
- Exact repro path — including the axis that keeps getting asked back:
regular message/reply vs extended context, which channel type, DM vs guild, which command variant. If the axis matters, say which; if it doesn't, say "either."
- The invariant under test — what property is being verified, so an
equivalent action counts ("any persona without an override," not "use persona X"). The user shouldn't have to ask "does that not count?"
- Masking state — what cached/DB state could make the test falsely pass
(e.g., an already-described image pulls from the DB and never exercises the new path). Name the reset step if one is needed.
- Expected observable — exactly what the user should see if it works, and
what failure looks like. Verify the expectation against the CODE (buildModelFooterText, not a persona's in-character explanation) before stating it.
- How to report — screenshot, paste, or a simple pass/fail.
Checklists are durable artifacts, not chat. A release smoke-test checklist is written into CURRENT.md (with per-item status), so it survives compaction and the user can re-consult it from a phone. Track progress there as results come in ("from my quick tests how much of the checklist did that get us?" must be answerable from the file). Justify every case — a bloated matrix wastes the user's time; each case states what it uniquely proves. After the user reports, close the loop with runtime evidence (logs) when the user-visible outcome can't prove the new code path ran.
Each checklist item carries a confidence tier — high (CI + review + blast-radius cover it; ships without a manual round) or needs-smoke (with the one-line reason confidence is limited: runtime-unverified path, mobile rendering, a failure sequence tests can't reach). Surface only the needs-smoke tier for the owner to run — they don't want to hand-test high-confidence changes ("what's the level of confidence… I don't feel like testing them unless confidence is limited").
Batch smoke asks into ONE numbered pass at release kickoff — including deferrals carried from previous releases. The owner's explicit ask: "I would like to know everything you need me to smoke test in dev so I can just knock it all out in one go. maybe including the deferred items still hanging around from previous releases too." Sweep CURRENT.md for carried-over deferred items before presenting; drip-feeding asks across the session forces repeat phone sessions. Two gates before an item enters the batch: (a) retro-log-close first — check whether prod/dev logs can already prove the item ran correctly (a 15-deployment sweep has closed several items the owner would otherwise have hand-tested; the owner's "I feel like I probably did 8 already" is a search order); (b) the item is only askable for work that is merged to develop — dev deploys from develop, so "smoke test my branch" is not an executable request.
Definition of Done
- New service files have
.component.test.ts - New API schemas have a colocated
.test.ts(unit-tier; no dedicated schema suffix) - Coverage doesn't drop (Codecov enforces 80%)
- Run
pnpm ops test:auditto verify no new gaps
References
- Full testing guide:
docs/reference/guides/TESTING.md - Mock factories:
services/*/src/test/mocks/ - PGLite setup:
docs/reference/testing/PGLITE_SETUP.md - Coverage audit:
docs/reference/testing/COVERAGEAUDITSYSTEM.md - Rules:
.claude/rules/02-code-standards.md