smithery.ai

qa-quarto

Adversarial Quarto-vs-Beamer parity QA. A critic agent compares the Quarto HTML render to the Beamer PDF benchmark for content/visual parity; a fixer agent applies fixes; loops until APPROVED (max 5 rounds). Use when user says "qa the quarto", "check parity", "does the html match the pdf?", "quarto matches beamer?", or after a translate-to-quarto run. Requires both the `.qmd` rendered and a `.pdf` benchmark.

First seen Apr 17, 2026

Installation

$ npx skills add https://smithery.ai

Summary

  • Adversarial Quarto-vs-Beamer parity QA.
  • A critic agent compares the Quarto HTML render to the Beamer PDF benchmark for content/visual parity; a fixer agent applies fixes; loops until APPROVED (max 5 rounds).
  • Use when user says "qa the quarto", "check parity", "does the html match the pdf?", "quarto matches beamer?", or after a translate-to-quarto run.
  • Requires both the `.qmd` rendered and a `.pdf` benchmark.

Similar popular skills

Related neighbors and high-traction skills in the same topics — useful to compare before installing.

Also in this package

Other skills from smithery.ai · top by installs.

npx skills add https://smithery.ai

Browse all from smithery.ai

More details

Agent compatibility

Declared targets from SKILL.md / docs. Unmarked agents are not listed — the skill may still install via the CLI.

Claude Code Not declared
Cursor Not declared
Codex Not declared
GitHub Copilot Not declared
Windsurf Not declared
Gemini CLI Not declared
Cline Not declared
OpenCode Not declared

Skill metadata

Parsed from SKILL.md frontmatter.

Allowed toolsRead, Grep, Glob, Write, Edit, Bash, Agent, Task

Package contents

Files included with this skill beyond the listing page.

  • skill md SKILL.md 4,569 B
  • docs SUMMARY.md 199 B

History

  1. First seen on skills.sh
  2. First recorded snapshot · 1 installs

SKILL.md

Adversarial Quarto vs Beamer QA Workflow

Compare Quarto HTML slides against their Beamer PDF benchmark using an iterative critic/fixer loop.

Philosophy: The Beamer PDF is the gold standard. The Quarto translation must be at least as good in every dimension.


Workflow

Phase 0: Pre-flight → Phase 1: Critic audit → Phase 2: Fixer → Phase 3: Re-audit → Loop until APPROVED (max 5 rounds)

Hard Gates (Non-Negotiable)

Gate Condition
Overflow NO content cut off
Plot Quality Interactive charts >= static plots
Content Parity No missing slides/equations/text
Visual Regression Quarto >= Beamer in all dimensions
Slide Centering Content centered, no jumping
Notation Fidelity All math verbatim from Beamer

Phase 0: Pre-flight

  1. Locate Beamer (.tex/.pdf) and Quarto (.qmd/.html) files
  2. Check freshness (re-render if QMD newer than HTML)
  3. Verify TikZ SVGs if applicable

Phase 1: Initial Audit

Launch the quarto-critic agent to compare Beamer vs Quarto comprehensively. Report saved to qualityreports/[Lecture]qacriticround1.md.

Phase 2: Fix Cycle

If not APPROVED, launch quarto-fixer agent to apply fixes (Critical → Major → Minor), re-render, and verify.

Phase 3: Re-Audit

Re-launch critic to verify fixes. Loop back to Phase 2 if needed.

Iteration Limits — loop-until-dry

This is the loop-until-dry primitive from [orchestrator-protocol.md](../../rules/orchestrator-protocol.md): the critic returns FINDINGs (the hard-gate table is the CRITICAL roll-up, per [orchestration-schemas.md](../../references/orchestration-schemas.md)); the loop converges when a round adds 0 new CRITICAL/MAJOR findings (deduped on id = sha1(file:line:locus)), not at a fixed round count.

  • Fallback cap: 5 rounds bounds a non-converging loop, then escalate to the user with remaining issues.
  • Two-strikes: the same gate failing in rounds N and N+2 is flagged for the user, not patched again ([summary-parity.md](../../rules/summary-parity.md)).
  • APPROVED iff every hard gate passes (zero CRITICAL).

Final Report

Save to qualityreports/[Lecture]qa_final.md with hard gate status, iteration summary, and remaining issues.

Findings are validated, not just written (v2.5)

This skill's reviewers emit findings under the machine-checked contract in [finding-schema.json](../../references/finding-schema.json). Reports are JSON arrays.

Smoke-test the harness before spending review effort — a run that fans out reviewers and then cannot write a valid report has wasted the whole pass:

echo '[]' | python3 scripts/validate-findings.py

Then, before presenting any summary:

python3 scripts/validate-findings.py <report>.json   # exit 0 required

What the contract forces, and why:

  • rule — the documented rule or standard violated. A finding citing no rule is an

opinion, and opinions do not gate a commit.

  • failing_case — a concrete configuration under which the claim breaks, or the exact

missing hypothesis. "This could be clearer" does not validate.

  • id = sha1("<file>:<line>:<locus>") — deterministic, so dedup across rounds is

exact and the two-strikes rule is checkable rather than eyeballed.

  • mechanicaltrue only for fixes that cannot change a result (typo, cross-reference,

formatting, label). Never for an estimand, assumption, specification, inference procedure, sample definition, or reporting language: those return to the researcher.

Apply the per-lens evidence burdens and the "does NOT count" filters in [orchestration-schemas.md §7](../../references/orchestration-schemas.md) before verification, so known false alarms never reach the judge. The verifier pass is refute-biased: only verdict: "confirmed" findings ship; anything it cannot ground is dropped, not downgraded to a warning.