mblode/ghostwriter · Archived

evaluate-ghostwriter

Runs blind, human-labeled baseline-versus-profile evaluations through a locally authenticated Codex or Claude Code CLI.

First seen Jul 23, 2026

Installation

$ npx skills add mblode/ghostwriter --skill evaluate-ghostwriter

Summary

  • Runs blind, human-labeled baseline-versus-profile evaluations through a locally authenticated Codex or Claude Code CLI.
  • Measures whether the exact ghostwriter skill and private platform profiles improve resemblance to held-out writing.
  • Use when asked to "evaluate my voice", "test my voice profile", compare writing with and without the ghostwriter skill, review blind A/B candidates, or report voice evaluation results.
  • Use train-ghostwriter to create the profiles and held-out cases this measures.

Stronger alternatives

This repository is archived — consider an actively maintained alternative.

Similar popular skills

Related neighbors and high-traction skills in the same topics — useful to compare before installing.

Also in this package

Other skills from mblode/ghostwriter.

npx skills add mblode/ghostwriter

Browse all from mblode/ghostwriter

More details

Agent compatibility

Declared targets from SKILL.md / docs. Unmarked agents are not listed — the skill may still install via the CLI.

Claude Code Declared
Cursor Not declared
Codex Declared
GitHub Copilot Not declared
Windsurf Not declared
Gemini CLI Not declared
Cline Not declared
OpenCode Not declared

Repository health

Stars 7
License LICENSE
Default branch main
Open issues 0
Status Archived

Skill metadata

Parsed from SKILL.md frontmatter.

Declared agents claude-code codex

Package contents

Files included with this skill beyond the listing page.

  • skill md SKILL.md 7,729 B
  • docs SUMMARY.md 527 B

History

  1. First seen on skills.sh
  2. First recorded snapshot · 8 installs

SKILL.md

Evaluate ghostwriter

Measure a fixed runtime skill and fixed private profiles against real held-out writing.

  • IS: clean candidate generation, deterministic blinding, human review, and descriptive reporting.
  • IS NOT: training profiles, modifying corpora, grading with another model, or exposing the treatment before a human choice. Use train-ghostwriter to create profiles and held-out cases.

The scripts execute the deterministic parts of this workflow. Do not reproduce their logic manually. scripts/run-eval.ts generates candidates without opening references. scripts/review-eval.ts joins references only after generation and records human choices. The scripts use [assets/candidate-output.schema.json](assets/candidate-output.schema.json) for structured runner output; do not edit it per run.

Both branches of a pair get the same CLI, model, case bytes, and output contract in fresh non-persistent sessions. The treatment alone also receives the complete runtime SKILL.md and platform profile, encoded losslessly as JSON strings. Never summarize or selectively copy either file; manifest.json pins hashes of their original bytes.

This measures the whole ghostwriter skill (anti-AI-prose pass plus strategy layer) and profile bundle against a raw-model baseline with no style guidance. A treatment win cannot be attributed to the profile alone, since it does not isolate the profile's own marginal effect.

Workflow

Copy this checklist and work top to bottom; each item is a section below.

  • 1. Pin inputs and runner (cases, runtime skill, profiles dir, runs dir, runner, model)
  • 2. Preview the provider boundary and get confirmation
  • 3. Generate matched pairs with run-eval.ts
  • 4. Review blind with review-eval.ts, choosing a/b/tie/invalid per case
  • 5. Verify counts reconcile against labels.jsonl and preserve the immutable run

1. Pin inputs and runner

Resolve these paths explicitly:

  • evals/cases.jsonl, containing no held-out responses
  • the exact installed ghostwriter/SKILL.md
  • the private profile directory
  • soul.md in the profiles dir (optional, cross-platform voice core), sent with the treatment when present and hashed into the manifest as soulHash
  • evals/runs
  • one runner (codex or claude) and an explicit model

Keep evals/references.jsonl separate. Do not pass it, quote it, or read it while generating candidates. Validate the contracts in the reference before invoking a model.

Cases, profiles, and the runtime skill must not contain absolute or traversal-based local image paths. The evaluator rejects them because Codex retains a residual view_image tool even after its generation-capable tools are disabled.

Existing profiles with unknown training provenance are useful for drafting, but they do not support an unbiased claim against examples they may have seen.

2. Preview the provider boundary

Tell the user exactly what will happen before the first model call:

  • Cases, facts, and constraints are sent to the selected model provider twice.
  • The treatment also sends the exact runtime skill, the selected private profile, and soul.md when present, since it is part of what the runtime drafts with.
  • Real held-out responses are not sent during candidate generation.
  • The repository adds no telemetry or direct API request, but the local agent CLI still contacts its provider.

Get confirmation. Do not print private profile content as part of the preview.

3. Generate matched pairs

Run from this skill directory:

node scripts/run-eval.ts \
  --cases <home>/evals/cases.jsonl \
  --runtime-skill <installed-ghostwriter>/SKILL.md \
  --profiles-dir <home> \
  --runs-dir <home>/evals/runs \
  --runner codex \
  --model <model>

Use --runner claude for Claude Code. The user may set GHOSTWRITER_HOME; otherwise <home> is ~/.config/ghostwriter.

The command prints the new run directory. On failure, preserve that directory and rerun the same command with its --run-id plus --resume. Resume is rejected if inputs, runner policy, or stored candidate metadata changed.

4. Review blind

Do not open manifest.json or candidates.jsonl before choices are complete. They contain branch identity. Run:

node scripts/review-eval.ts \
  --run-dir <run-directory> \
  --references <home>/evals/references.jsonl

For each case, choose a, b, tie, or invalid. Judge which candidate sounds more like the real held-out response while preserving all required facts. Use tie when both are comparably close and factually valid, and invalid when the case, reference, or both candidates make a fair comparison impossible. Do not label on polish alone: a fluent candidate that alters a required fact loses. Labels are saved after every choice, and the blind and reference file hashes are pinned when review begins.

For an externally completed blind review, provide a JSONL choices file through --labels-file. It must contain only id and choice.

5. Verify and preserve

Confirm:

  • labels.jsonl has one record per blind case.
  • Overall and per-platform report counts reconcile with those labels.
  • Invalid labels are separate from valid win/tie rates.
  • The report contains raw observations only, with no automated quality claim.

Keep the immutable run as evidence. Start a new run after changing any skill, profile, case, runner, or model. Inspect every treatment loss and invalid case and group recurring causes before changing anything.

Gotchas

  • ERRUNKNOWNFILE_EXTENSION: Unknown file extension ".ts" means the local Node is older than 22.18 and cannot run these scripts. Report the version and ask the user to upgrade; nothing in this skill works around it.
  • Never add --references to run-eval.ts; the option is intentionally unsupported so held-out answers cannot enter candidate prompts.
  • Never weaken Codex's flag set. --disable shelltool matters because a read-only sandbox blocks writes but still allows private file reads; the apps, multi-agent, image-generation and web-search disables each close an uncontrolled generation path; and -c skills.includeinstructions=false matters because --ignore-user-config alone still exposes globally installed skill descriptions and can contaminate the baseline.
  • Codex still registers updateplan, requestuserinput, applypatch, and viewimage. Read-only mode blocks applypatch; prompt rules prohibit all tools; local image-path rejection limits view_image reads.
  • Never open manifest.json before labeling, and never edit references, blind candidates, mappings, or stored labels after review begins. blindMapping identifies the treatment, and pinned hashes make an altered run fail closed.
  • Leave --seed unset for real reviews; the script creates a private random seed so the public assignment algorithm does not reveal A/B.
  • Never compare runs that changed both the profile and model; the result cannot isolate the profile's effect.
  • Never regenerate successful branches during resume; the script preserves them so retries do not silently change the pair.
  • Never treat a small win rate as proof.