Summary
创建符合 CAC 评测系统规范的标准化测试题目。当用户需要添加新题目、创建测试题、或提到"创建题目"、"新建题目"、"添加题目"、"add question"、"create question"时使用此 Skill。支持代码、数理、逻辑、综合四类题库。
smithery.ai
创建符合 CAC 评测系统规范的标准化测试题目。当用户需要添加新题目、创建测试题、或提到"创建题目"、"新建题目"、"添加题目"、"add question"、"create question"时使用此 Skill。支持代码、数理、逻辑、综合四类题库。
创建符合 CAC 评测系统规范的标准化测试题目。当用户需要添加新题目、创建测试题、或提到"创建题目"、"新建题目"、"添加题目"、"add question"、"create question"时使用此 Skill。支持代码、数理、逻辑、综合四类题库。
Related neighbors and high-traction skills in the same topics — useful to compare before installing.
Turn a decision you can't fully answer into a questionnaire for someone else to fill in.
274.6K installs把模糊问题改写成 Agent 可推理、可批评、可验证的问题说明书,并判断自动化解决程度。用户要求把问题…
18.1K installsClarify requirements before implementing. Use when serious doubts arise.
6.9K installs把你无法完整回答的 decision 转成一份交给他人填写的 questionnaire。
1.6K installsCreate or update a Sumsub KYC questionnaire definition. POST `/resources/api/agent/questionnair…
20 installsWrites and executes SQL queries against the data warehouse using dbt's Semantic Layer or ad-hoc…
803 installsOther skills from smithery.ai · top by installs.
npx skills add https://smithery.ai
Declared targets from SKILL.md / docs. Unmarked agents are not listed — the skill may still install via the CLI.
Files included with this skill beyond the listing page.
SKILL.md
3,135 B
SUMMARY.md
306 B
创建符合 CAC 评测系统规范的标准化测试题目。
创建题目前,请提供:
NNN-problem-name 格式python scripts/validate_questions.py{题库名}/{difficulty}-test/NNN-problem-name/
├── README.md # 人类阅读的完整文档
├── meta.yaml # 机器读取的元数据
├── prompt.md # 发给被测模型的 prompt
├── reference.md # 标准答案/评判依据
└── test-results/ # 测试结果目录
详细格式规范见 [questionformatguide.md](references/questionformatguide.md)。 编号规则见 [numberingrules.md](references/numberingrules.md)。 文件模板见 [templates/](templates/) 目录。
id: {category}-{difficulty}-{number}
brief: 题目简短描述
category: math | code | logic | comprehensive
difficulty: base | advanced | final | final+
scoring_std:
max_score: 10
indicators:
- accuracy # 准确性
- completeness # 完整性
| 类型 | 可用指标 |
|---|---|
| 代码题 | anscorrect, codequality, efficiency, robustness |
| 理论题 | completeness, accuracy, clarity, depth |
| 设计题 | anscorrect, examplequality, completeness, practicality |
纯净的题目文本,直接发给被测模型,不含元数据。
标准答案和评判依据,供评判模型参考。
| 类别 | 路径 | category |
|---|---|---|
| 代码 | 代码能力基准测试题库/ |
code |
| 数理 | 数理能力基准测试题库/ |
math |
| 逻辑 | 自然语言与逻辑能力基准测试题库/ |
logic |
| 综合 | 综合能力测评/ |
comprehensive |
参考现有题目:
数理能力基准测试题库/base-test/001-chicken-rabbit-cage/代码能力基准测试题库/base-test/001-simple-calculator/自然语言与逻辑能力基准测试题库/base-test/001-age-multiple-reasoning/