Files included with this skill beyond the listing page.
skill mdSKILL.md5,003 B
docsSUMMARY.md338 B
History
First seen on skills.sh
First recorded snapshot · 1 installs
SKILL.md
Result-to-Claim Gate
Experiments produce numbers; this gate decides what those numbers mean. Collect results from available sources, get a Codex judgment, then auto-route based on the verdict.
Context: $ARGUMENTS
When to Use
After a set of experiments completes (main results, not just sanity checks)
Before committing to claims in a paper or review response
When results are ambiguous and you need an objective second opinion
Workflow
Step 1: Collect Results
Gather experiment data from whatever sources are available in the project:
W&B (preferred): wandb.Api().run("<entity>/<project>/<run_id>").history() — metrics, training curves, comparisons
EXPERIMENT_LOG.md: full results table with baselines and verdicts
EXPERIMENT_TRACKER.md: check which experiments are DONE vs still running
Log files: ssh server "tail -100 /path/to/training.log" if no other source
docs/research_contract.md: intended claims and experiment design
Assemble the key information:
What experiments were run (method, dataset, config)
Main metrics and baseline comparisons (deltas)
The intended claim these experiments were designed to test
Any known confounds or caveats
Step 2: Codex Judgment
Send the collected results to Codex for objective evaluation:
mcp__codex__codex:
config: {"model_reasoning_effort": "xhigh"}
prompt: |
RESULT-TO-CLAIM EVALUATION
I need you to judge whether experimental results support the intended claim.
Intended claim: [the claim these experiments test]
Experiments run:
[list experiments with method, dataset, metrics]
Results:
[paste key numbers, comparison deltas, significance]
Baselines:
[baseline numbers and sources — reproduced or from paper]
Known caveats:
[any confounding factors, limited datasets, missing comparisons]
Please evaluate:
1. claim_supported: yes | partial | no
2. what_results_support: what the data actually shows
3. what_results_dont_support: where the data falls short of the claim
4. missing_evidence: specific evidence gaps
5. suggested_claim_revision: if the claim should be strengthened, weakened, or reframed
6. next_experiments_needed: specific experiments to fill gaps (if any)
7. confidence: high | medium | low
Be honest. Do not inflate claims beyond what the data supports.
A single positive result on one dataset does not support a general claim.
Step 3: Parse and Normalize
Extract structured fields from Codex response:
- claim_supported: yes | partial | no
- what_results_support: "..."
- what_results_dont_support: "..."
- missing_evidence: "..."
- suggested_claim_revision: "..."
- next_experiments_needed: "..."
- confidence: high | medium | low
Step 4: Route Based on Verdict
no — Claim not supported
Record postmortem in findings.md (Research Findings section):
- What was tested, what failed, hypotheses for why - Constraints for future attempts (what NOT to try again)
Update CLAUDE.md Pipeline Status
Decide whether to pivot to next idea from IDEA_CANDIDATES.md or try an alternative approach
partial — Claim partially supported
Update the working claim to reflect what IS supported
Record the gap in findings.md
Design and run supplementary experiments to fill evidence gaps
Re-run result-to-claim after supplementary experiments complete
Multiple rounds of partial on the same claim → record analysis in findings.md, consider whether to narrow the claim scope or switch ideas
yes — Claim supported
Record confirmed claim in project notes
If ablation studies are incomplete → trigger /ablation-planner
If all evidence is in → ready for paper writing
Rules
Codex is the judge, not CC. CC collects evidence and routes; Codex evaluates. This prevents post-hoc rationalization.
Do not inflate claims beyond what the data supports. If Codex says "partial", do not round up to "yes".
A single positive result on one dataset does not support a general claim. Be honest about scope.
If confidence is low, treat the judgment as inconclusive and add experiments rather than committing to a claim.
If Codex MCP is unavailable (call fails), CC makes its own judgment and marks it [pending Codex review] — do not block the pipeline.
Always record the verdict and reasoning in findings.md, regardless of outcome.