smithery/lbds137

tzurot-testing

Testing procedures. Invoke with /tzurot-testing for test execution, coverage audits, and debugging test failures.

Installation

$ npx skills add smithery/lbds137 --skill tzurot-testing

Similar popular skills

Related neighbors and high-traction skills in the same topics — useful to compare before installing.

Also in this package

Other skills from smithery/lbds137.

npx skills add smithery/lbds137

Browse all from smithery/lbds137

More details

Agent compatibility

Declared targets from SKILL.md / docs. Unmarked agents are not listed — the skill may still install via the CLI.

Claude Code Declared
Cursor Not declared
Codex Not declared
GitHub Copilot Not declared
Windsurf Not declared
Gemini CLI Not declared
Cline Not declared
OpenCode Not declared

Skill metadata

Parsed from SKILL.md frontmatter.

Declared agents claude-code

Package contents

Files included with this skill beyond the listing page.

  • skill md SKILL.md 14,119 B
  • docs SUMMARY.md 202 B

History

  1. First recorded snapshot · 0 installs

SKILL.md

Testing Procedures

Invoke with /tzurot-testing for test-related procedures.

Testing patterns are in .claude/rules/02-code-standards.md - they apply automatically.

Running Tests

# Run all tests
pnpm test

# Run specific service
pnpm --filter @tzurot/ai-worker test

# Run specific file
pnpm test -- MyService.test.ts

# Run the component tier (PGLite) — NOT included in `pnpm test`
pnpm test:component

# Run the integration + contract tiers (real Redis/Postgres) — NOT in `pnpm test`
pnpm test:integration

# Run with coverage
pnpm test:coverage

# Run only changed packages
pnpm focus:test

Always invoke through the package's own config. A root-invoked npx vitest run <pkg path> resolves the ROOT config, which has no setupFiles — the package's setup (dotenv, the global afterEach(vi.clearAllMocks)) never loads, hoisted-mock call history accumulates across tests in a file, and expect(mock).not.toHaveBeenCalled() fails on calls leaked from earlier tests. It masquerades as flaky or pre-existing breakage; it is an invocation artifact. Use pnpm --filter @tzurot/<pkg> test, or cd services/<pkg> && npx vitest run <relative-path> — and never diagnose "pre-existing failures" from a root-invoked run.

Coverage Audit Procedure

# Run unified audit (CI does this automatically)
pnpm ops test:audit

# Filter by category
pnpm ops test:audit --category=services
pnpm ops test:audit --category=contracts

# Update baseline (after closing gaps)
pnpm ops test:audit --update

# Strict mode (fails on ANY gap)
pnpm ops test:audit --strict

Unified Baseline: test-coverage-baseline.json (project root)

Test File Types

Tiers are canonically defined in one place — [Test Tier Taxonomy](../../../docs/reference/guides/TESTING.md#test-tier-taxonomy). This table is the suffix→tier quick reference; don't re-define the tiers here (the pnpm ops guard:test-taxonomy gate enforces the single-source link).

Suffix / location Tier ([taxonomy](../../../docs/reference/guides/TESTING.md#test-tier-taxonomy)) Infrastructure Location
*.test.ts Unit Fully mocked Next to source
*.component.test.ts Component (one service, PGLite) PGLite Next to source
*.integration.test.ts Integration (real DB+Redis) Real services tests/e2e/
*.contract.test.ts Contract (provider↔consumer) Real services tests/e2e/contracts/

Each suffix names its tier directly. Schema test ≠ contract test: a Zod
schema test (a plain *.test.ts) validates one type's own rules (unit-tier); a
contract test verifies two services agree. Run pnpm ops test:tiers for the
per-package tier distribution.

Verify with the config that runs the tier (CRITICAL)

The unit pnpm test config EXCLUDES the .component / .integration / .contract suffixes. A contract/integration/component test "passes" under pnpm test by NOT running at all — so a green unit run is NOT evidence it works.

Suffix Verify with
*.test.ts (unit) pnpm test
*.component.test.ts pnpm test:component
.integration.test.ts / .contract.test.ts pnpm test:integration

When you touch a .contract / .integration test, run pnpm test:integration (Redis + Postgres up: podman start tzurot-redis tzurot-postgres); a .component test, pnpm test:component. pnpm test / pnpm --filter <pkg> test is the unit tier ONLY. CI's component-integration-tests job runs test:component + test:integration, so it catches what a unit-only local run silently skipped — don't let CI be the first thing that actually executes your contract test.

Gotcha for mention/ID fixtures: tests need a valid 17–19 digit Discord snowflake — isValidDiscordId silently drops toy ids like 555 before resolution, making the assertion pass/fail for the wrong reason.

Debugging Test Failures

1. Run Specific Test

pnpm test -- MyService.test.ts --reporter=verbose

2. Check for Fake Timer Issues

// ❌ WRONG - Promise rejection warning
const promise = asyncFunction();
await vi.runAllTimersAsync(); // Rejection happens here!
await expect(promise).rejects.toThrow(); // Too late

// ✅ CORRECT - Attach handler BEFORE advancing
const promise = asyncFunction();
const assertion = expect(promise).rejects.toThrow('Error');
await vi.runAllTimersAsync();
await assertion;

3. Reset Mock State

beforeEach(() => {
  vi.clearAllMocks(); // Clear call history, keep impl
});
afterEach(() => {
  vi.restoreAllMocks(); // Restore originals (spies only)
});

Creating Mock Factories

// Use async factory for vi.mock hoisting
vi.mock('./MyService.js', async () => {
  const { mockMyService } = await import('../test/mocks/MyService.mock.js');
  return mockMyService;
});

// Import accessors after vi.mock
import { getMyServiceMock } from '../test/mocks/index.js';

it('should call service', () => {
  expect(getMyServiceMock().someMethod).toHaveBeenCalled();
});

Probe fixtures come from the real corpus

At the moment of writing a probe or test fixture — a filename, a payload, a URL, a config value — do not invent a minimal one: copy a real specimen out of the corpus and mutate only what the test needs. An input built to isolate one property is exactly the input that cannot reveal an interaction with a second property, and its green is indistinguishable from a green that covered the real case (a space-free probe filename passed end-to-end while every real tracker/ path — all space-bearing — would have been quote-mangled). If a synthetic input is unavoidable, enumerate the properties real specimens carry (spaces, non-ASCII, length, punctuation, nesting) and give it ALL of them. The refutation direction inverts this: the probe that settles a disputed finding varies a single property at a time.

Mock-reachability check (before writing assertions)

Before asserting on a subject that sits behind a mocked collaborator, answer two questions: (a) which mock makes this code path reachable, and (b) what would a wiring bug at that seam look like in this test's output? If the answer to (b) is "the test can't see it," the test asserts through the seam it mocked — add the seam assertion per 02-code-standards.md § Assert what crosses a mocked seam before writing more cases. This is the coverage-illusion class: a mock missing one property leaves the dependent branch green and untested; a feature flag can no-op silently under a fully green suite.

Stacked gates: make upstream gates inert

When asserting gate N in a multi-gate pipeline — a hook with several checks, a validator chain, a new guard stacked in front of an older one — an exit code or shared failure result is NOT attribution: the new gate inherits the older gate's failure, so assert_exit 2 (or .toThrow(), or "returns false") passes whether YOUR gate fired or the one behind it did. Two obligations:

  1. Construct the fixture so every upstream gate is inert (would pass on its

own), leaving the gate under test as the only possible failure source. Asserting a distinctive banner/message helps but is secondary — an inert- upstream fixture proves attribution even when messages change.

  1. Canary THIS gate: mutate the gate under test and confirm the specific

assertion reddens (Core Principle 9, 02-code-standards.md). A canary against the pipeline as a whole proves nothing about which gate you pinned.

Anatomy to fear: a cluster of vacuous assertions shipped in one hook PR (#2078) with exactly this shape — including one whose test NAME was false, passing only because a different gate blocked first. Reading had already passed every one of them; per-gate canaries are what caught them.

Integration Tests with PGLite

describe('UserService', () => {
  let pglite: PGlite;
  let prisma: PrismaClient;

  beforeAll(async () => {
    pglite = new PGlite({ extensions: { vector, citext } });
    await pglite.exec(loadPGliteSchema());
    prisma = new PrismaClient({ adapter: new PrismaPGlite(pglite) });
  });

  it('should create user', async () => {
    const service = new UserService(prisma);
    const userId = await service.getOrCreateUser('123', 'testuser');
    expect(userId).toBeDefined();
  });
});

⚠️ ALWAYS use loadPGliteSchema() - NEVER create tables manually!

Component Test Triggers

Component tests (*.component.test.ts) run separately from unit tests and are not included in pnpm test or pre-push hooks.

Always run pnpm test:component after:

Change Why
Add/remove slash command options CommandHandler.component.test.ts snapshots capture full command structure
Add/remove subcommands Same snapshot tests
Restructure command directories getCommandFiles() discovery changes affect command loading
Change component prefix routing Component tests verify button/select menu routing

Update snapshots with: pnpm vitest run --config vitest.component.config.ts <file> --update

Human-Verification Requests (manual testing by the user)

The user is the ONLY manual-QA executor — usually on a phone, across session crashes and compactions. Every request for manual verification MUST be a complete, self-contained instruction with all five parts:

  1. Exact repro path — including the axis that keeps getting asked back:

regular message/reply vs extended context, which channel type, DM vs guild, which command variant. If the axis matters, say which; if it doesn't, say "either."

  1. The invariant under test — what property is being verified, so an

equivalent action counts ("any persona without an override," not "use persona X"). The user shouldn't have to ask "does that not count?"

  1. Masking state — what cached/DB state could make the test falsely pass

(e.g., an already-described image pulls from the DB and never exercises the new path). Name the reset step if one is needed.

  1. Expected observable — exactly what the user should see if it works, and

what failure looks like. Verify the expectation against the CODE (buildModelFooterText, not a persona's in-character explanation) before stating it.

  1. How to report — screenshot, paste, or a simple pass/fail.

Checklists are durable artifacts, not chat. A release smoke-test checklist is written into CURRENT.md (with per-item status), so it survives compaction and the user can re-consult it from a phone. Track progress there as results come in ("from my quick tests how much of the checklist did that get us?" must be answerable from the file). Justify every case — a bloated matrix wastes the user's time; each case states what it uniquely proves. After the user reports, close the loop with runtime evidence (logs) when the user-visible outcome can't prove the new code path ran.

Each checklist item carries a confidence tier — high (CI + review + blast-radius cover it; ships without a manual round) or needs-smoke (with the one-line reason confidence is limited: runtime-unverified path, mobile rendering, a failure sequence tests can't reach). Surface only the needs-smoke tier for the owner to run — they don't want to hand-test high-confidence changes ("what's the level of confidence… I don't feel like testing them unless confidence is limited").

Batch smoke asks into ONE numbered pass at release kickoff — including deferrals carried from previous releases. The owner's explicit ask: "I would like to know everything you need me to smoke test in dev so I can just knock it all out in one go. maybe including the deferred items still hanging around from previous releases too." Sweep CURRENT.md for carried-over deferred items before presenting; drip-feeding asks across the session forces repeat phone sessions. Two gates before an item enters the batch: (a) retro-log-close first — check whether prod/dev logs can already prove the item ran correctly (a 15-deployment sweep has closed several items the owner would otherwise have hand-tested; the owner's "I feel like I probably did 8 already" is a search order); (b) the item is only askable for work that is merged to develop — dev deploys from develop, so "smoke test my branch" is not an executable request.

Definition of Done

  • New service files have .component.test.ts
  • New API schemas have a colocated .test.ts (unit-tier; no dedicated schema suffix)
  • Coverage doesn't drop (Codecov enforces 80%)
  • Run pnpm ops test:audit to verify no new gaps

References

  • Full testing guide: docs/reference/guides/TESTING.md
  • Mock factories: services/*/src/test/mocks/
  • PGLite setup: docs/reference/testing/PGLITE_SETUP.md
  • Coverage audit: docs/reference/testing/COVERAGEAUDITSYSTEM.md
  • Rules: .claude/rules/02-code-standards.md