开源 · MIT · 兼容 Claude Code、DeepSeek Harness 等 AI 编程助手 Open source · MIT · works with Claude Code, DeepSeek Harness and other AI coding agents

让测试工程Make test engineering
成为 AI 的本能an instinct for AI

一套面向 AI Agent 的测试工程 Skill 框架。
装进编程助手,AI 就能按资深工程师的方式做测试:
读需求、排风险、写用例、跑自动化、出报告。

A QA engineering skill framework for AI agents.
Installed into your coding agent, the AI tests like a senior engineer:
read requirements, map risk, write cases, run automation, report.

实测:用例合格率 26% → 98%·已上架 skills.sh / npm·MIT 开源免费 Measured: case pass rate 26% → 98%·on skills.sh / npm·MIT, free & open source

01 · Why

软件发布之前,Before any release ships,
总得有人先把住质量关someone has to hold the line on quality

每一款软件上线前,都要回答成百上千个「如果」: 并发下单会不会超卖?优惠券到期会不会自动失效?接口异常时数据会不会错乱?

Before launch, every piece of software has to answer hundreds of “what ifs”: Will concurrent orders oversell? Will coupons expire on their own? Will data corrupt when APIs fail?

干这件事的,是软件测试工程师: 在缺陷到达用户之前发现它、定位它,再确认它被修好。

That job belongs to the QA engineer: find defects before users do, localize them, and verify the fix.

qa-skills 做的,就是让 AI 助手学会这套本事。qa-skills teaches that craft to your AI agent.

02 · Problem

能写,更要能执行。Writing cases is easy. Executing them is the bar.

通用大模型写出的用例,常常看着专业、实际没法执行:
判定模糊、占位符没替换、时限没说清、入口是编的。
这是能力问题,更是纪律问题。

Cases from general-purpose models often look professional but can't run:
vague verdicts, placeholders left in, unstated timeouts, made-up entry points.
That's a capability gap — and a discipline gap.

✗ 无 Skill 的模型产出✗ Model output, no skills CASE-001

判定模糊,无法执行Vague verdict, unrunnable

预期结果不可判定,等于没有验证:An undecidable expected result is no verification at all:

「检查优惠券功能是否正常」
「正常」是什么标准?谁判定?等多久算超时?
“Check that the coupon feature works correctly”
What's the bar for “correctly”? Who decides? How long before timeout?

判定不了的用例,覆盖再全也计零分。A case that can't be judged scores zero — however broad its coverage.

✓ qa-skills 的真实产出✓ Real qa-skills output TC-03-05

步骤精确,判定可量化Precise steps, measurable verdicts

前置、步骤、预期、时限,四要素齐备:Preconditions, steps, expected results, timeouts — all four present:

「选一张 10 分钟后到期 的已发布优惠券,等待到期。
预期:1 小时内状态自动变为「已结束」,
超过 1 小时即判失败。」
“Pick a published coupon expiring in 10 minutes, and wait.
Expected: status flips to “Ended” automatically within 1 hour;
any longer counts as a failure.”

任何测试人员拿到就能执行,不用讲解。Any tester can execute it as-is, no briefing needed.

产出标准只有一条:
任何一个没接触过这个需求的人,拿到文件就能直接执行。 背后是 8 条可执行性硬标准,评测中一票否决。
One output bar only:
anyone who never touched this requirement can pick up the file and run it. Backed by 8 hard executability standards — veto power in evaluation.

03 · What

不是软件,Not software —
是一套装进 AI 的工作方法。a working method installed into AI.

由 12 个 QA Skill 和一个共享知识库(core/)组成,全部是 Markdown 指令文件。 放对目录就能生效。

12 QA skills plus a shared knowledge base (core/), all plain Markdown instruction files. Drop them in the right directory and they take effect.

qa-skills/ · Markdown 指令qa-skills/ · Markdown instructions
skills/
├─ requirement-analysis/
├─ test-strategy/
├─ test-case-writing/
├─ automated-e2e-testing/
├─ bug-analysis/ · regression-testing/ …
core/ 共享知识库 · 8 条硬标准shared knowledge · 8 hard standards
AI Agent · 装载后获得AI Agent · what it gains
先澄清什么,歧义点逐条向人确认What to clarify first — ambiguities confirmed with a human
依据什么排风险,每个评级给证据How risks are ranked — every rating shows evidence
哪些类型必测,哪些明确不测Which test types are in, which are explicitly out
每条结论标注几级证据,可追溯Every conclusion tagged with evidence level, traceable
产出写成文件,中断可接力Outputs written to files — resume across sessions
modular01

12 个 QA Skill + 共享知识库12 QA skills + shared knowledge base

需求分析、策略、用例、评审、执行、缺陷、回归,每个环节一个专属 Skill,按需加载。

One dedicated skill per stage — requirements, strategy, cases, review, execution, defects, regression — loaded on demand.

file-based02

一条自动接力的流水线A pipeline that hands off automatically

每个阶段的产出都写成文件,不靠对话记忆。中断了,新开一个对话接着跑。

Every stage writes its output to files, not chat memory. Interrupted? Start a new session and continue.

human-gated03

关键决定,由你拍板Key decisions stay with you

需求没写清的,先问你;Bug 怎么定性、预算设多少,AI 助手只提案、不代答。你的裁决会记录在案,后续不得推翻。

Unclear requirements? It asks you first. Bug severity, budget caps — the agent proposes, never decides. Your rulings are recorded and never overridden later.

04 · How

一句指令,从需求到报告。One instruction, requirement to report.

对 AI 助手说「帮我测试这个需求」,从需求到报告各环节自动接力,每个环节都有明确的产出物(完整流水线共九个阶段,含探索旁路与回归验证)。

Say “test this requirement” to your agent and the stages chain automatically, each with a concrete deliverable (the full pipeline has nine stages, including the exploratory bypass and regression verification).

1

需求分析Requirement analysis

不懂就先问Ask before assuming
2

测试策略Test strategy

锁定高危区域Lock the high-risk zones
3

用例设计Case design

人和机器都能读Readable by humans and machines
4

用例评审Case review

再挑一遍错One more pass for defects
5

自动执行Automation

浏览器 + 接口Browser + API
6

缺陷分析Defect analysis

追到出错代码Down to the failing line
7

回归 + 报告Regression + report

证据齐全Evidence included
05 · Skills

从用例到回归,From cases to regression,
每道工序都有对应的 Skill。every stage has its own skill.

每个 Skill 单独可用,也能串成完整流水线。需要哪一段,就用哪一段。

Each skill works standalone or as part of the full pipeline. Use exactly the stage you need.

test-case-writing01

用例设计Case design

代码优先:先读实现、审出潜在缺陷,再落笔。产出两份,markmap 脑图给人读,schema.yaml 给机器校验。

Code-first: read the implementation, hunt latent bugs, then write. Two deliverables — a markmap for humans, schema.yaml for machines.

test-strategy02

测试策略Test strategy

风险评级必须给出证据;性能、安全、并发等十个类型,逐一决定测还是不测,每个决定都留下记录。说不清依据的策略,等于没有策略。

Risk ratings must cite evidence; ten test types — performance, security, concurrency and more — each explicitly tested or not, every decision on record. A strategy that can't explain itself is no strategy.

test-case-review03

用例评审Case review

把所有能测的点列全,作为基准逐条复审:覆盖够不够、能不能执行,直接修订文件并留下审查记录。

Enumerate every testable point as a baseline, then re-review line by line: coverage enough? executable? Files get revised in place with a review record.

automated-e2e · api-testing04

自动化执行Automation

用例转成可运行的 Playwright / pytest 代码:按 Page Object 规范组织,测试数据自己造、跑完自己清。

Cases become runnable Playwright / pytest code: Page Object conventions, self-built test data, self-cleaning after runs.

bug-analysis05

缺陷分析Defect analysis

先复现,再读代码定位到出错的那一行,然后分析影响面、给出回归建议。产出的 Bug 条目带根因、带证据。

Reproduce first, then read the code down to the failing line, analyze blast radius, recommend regression scope. Bug entries carry root cause and evidence.

regression-testing06

回归测试Regression testing

从代码改动推算波及范围,生成回归清单。防止修一个、坏三个。

Derive regression scope from the code diff and generate the retest list. Stop fixing one thing and breaking three.

06 · Proof

每一个数字,都来自实测。Every number comes from measurement.

同一批任务,同一个模型,唯一差异是装没装 qa-skills。
数字如实披露,包括不好看的那些。

Same tasks, same model; the only difference is whether qa-skills is installed.
All numbers disclosed as measured — the unflattering ones too.

98%

用例合格率Case pass rate

格式与内容逐条核对format & content checked item by item

不装时 26%26% without
75%

植入缺陷检出率Seeded-bug detection

代码审查任务实测measured on code-review tasks

换一套裁判评:100%100% under an alternate judge
7/ 9

E2E 用例真实执行通过E2E cases passing real runs

单任务×3 采样真浏览器实测real-browser runs, 1 task × 3 samples

不装时 0/90/9 without
0.88

类型查全率Test-type recall

最弱模型上实测measured on the weakest model

不装时为 00 without

Token 成本约 3.3 倍。覆盖提升:写用例 +8.7pp、全任务 +13.2pp、找缺陷 +9.7pp。 完整数据见 README,里程碑版 Release 附跨模型增益矩阵快照。

Token cost ≈ 3.3×. Coverage gains: case writing +8.7pp, all tasks +13.2pp, defect finding +9.7pp. Full data in the README; milestone releases ship a cross-model gain-matrix snapshot.

它知道,什么该由你决定。It knows what's yours to decide.

需求有歧义,先澄清;执行走人工还是自动化,你选; Bug 怎么定性、预算上限设多少,你说了算。 这四类检查点,AI 助手只提案、不代答。

Ambiguous requirements get clarified first; manual or automated execution, you pick; bug severity, budget caps — you decide. At these four checkpoints the agent proposes, never decides.

每一次裁决都记录在案,后续阶段照办,不得推翻。 拍板的,始终是你。

Every ruling is recorded; later stages comply, no override. The call is always yours.

Decision required · checkpoint 2/4
Bug 定性Bug callP1 逻辑缺陷 · 影响结算主流程 · 附出错代码行P1 logic defect · hits the main settlement flow · failing line attached
回归建议Regression补充结算边界用例 ×3,纳入本轮回归Add 3 settlement-edge cases to this regression round
批准执行Approve 驳回,换个方案Reject, propose another
14:32 裁决已记录 → decisions.yaml
后续阶段照办,不得推翻
14:32 ruling recorded → decisions.yaml
later stages comply, no override
07 · FAQ

装之前的疑虑,Every pre-install doubt,
逐条如实回答。answered honestly.

包括不那么好听的答案——Token 更贵、Skill 正文还没双语,都直接写在这里。

Including the unflattering ones — tokens cost more, skill bodies aren't bilingual yet.

qa-skills 是什么?要花钱吗?What is qa-skills? Does it cost money?
一套装进 AI 编程助手的测试工程方法:12 个 Skill + 一个共享知识库,全部是 Markdown 指令文件。不是软件,没有账号和订阅,MIT 协议免费开源。它确实会多花 Token——约为不装时的 3.3 倍,这笔账我们在首页如实写着。
A QA engineering method installed into your AI coding assistant: 12 skills + a shared knowledge base, all Markdown instruction files. Not software — no account, no subscription, MIT-licensed and free. It does cost more tokens — about 3.3× the unskilled runs — and we state that openly on this page.
我的 AI 助手能用吗?Will it work with my AI assistant?
Skill 是纯 Markdown(frontmatter + 相对路径引用),不依赖宿主特性:Claude Code 端到端实测通过,npx skills add fishzjp/qa-skills --skill '*' 一行装到 Claude Code / Cursor / Codex / OpenCode 等 70+ 宿主。宿主不支持子代理时,流水线自动退化为顺序会话 + 文件衔接,正确性不受影响。
Skills are plain Markdown (frontmatter + relative-path references) with no host-specific dependencies: verified end-to-end on Claude Code, and npx skills add fishzjp/qa-skills --skill '*' installs into Claude Code / Cursor / Codex / OpenCode and 70+ other hosts. Where the host lacks subagents, the pipeline degrades to sequential sessions joined by files — correctness is unaffected.
AI 产出的用例,真的能直接执行吗?Can the AI's cases really run as-is?
这是本框架唯一的产出标准:没读过需求、没人讲解的人,拿着文件能直接开工。背后是 8 条可执行性硬标准,评测中一票否决——判定模糊、占位符未替换、虚构入口的用例,覆盖再全也计零分。实测:无 Skill 时用例合格率 26%,装后 98%。
That is the framework's one output bar: someone who never read the requirement and got no walkthrough can pick up the file and start working. Behind it are 8 hard executability standards with veto power in evaluation — vague verdicts, leftover placeholders or made-up entry points score zero no matter the coverage. Measured: 26% case pass rate without skills, 98% with.
装完还要学一遍怎么用吗?Is there a setup tutorial to go through?
不需要。对 AI 助手说「帮我测试这个需求:{需求描述 + 仓库地址}」,流水线就从需求理解跑到测试报告;只要其中一段产出(写用例 / 审查 / 回归范围)时,直接说需求。需求歧义、执行策略、Bug 定性、预算上限这四类决策由你拍板,AI 只提案、不代答。
No. Say “Test this requirement: {description + repo URL}” to your agent and the pipeline runs from requirement understanding to test report; when you need just one stage (case writing / review / regression scope), state the need directly. Four decision types stay with you — requirement ambiguities, execution strategy, bug severity, budget caps; the AI proposes, never decides.
我的代码和数据会被上传吗?Will my code and data get uploaded?
qa-skills 本身零运行时、零遥测:装进来的只有 Markdown 文件,没有任何网络调用;测试产出与 .qa/ 项目知识库都存在你自己的仓库里,随你的 git 管控。(AI 助手本身如何处理数据,取决于你选的宿主与模型服务商,与本框架无关。)
qa-skills itself is zero-runtime and zero-telemetry: what gets installed is only Markdown files, with no network calls; test outputs and the .qa/ project knowledge base live in your own repo under your own git. (How your AI assistant handles data depends on the host and model provider you choose — outside this framework.)
支持英文吗?Does it support English?
落地页与 README 已支持英文;12 个 Skill 中 11 个的 description 已双语(bug-analysis 暂缓),Skill 指令正文当前仍是中文。需求描述与被测代码为中英混合时可正常工作。希望 Skill 正文也完整双语?欢迎到 Discussions 告诉我们你的场景,这会直接影响优先级。
This page and the README are available in English; 11 of the 12 skill descriptions are bilingual (bug-analysis pending), while the skill instruction bodies are still Chinese. Mixed Chinese/English requirements and code under test work fine. Want fully bilingual bodies? Tell us your use case in Discussions — it directly moves the priority.
08 · Install

两步接入你的 AI 助手Two steps to your AI assistant

不需要账号,不需要订阅。两条命令,一句话。

No account, no subscription. Two commands, one sentence.

01

一行命令,装进 AI 编程助手One command, installed into your coding agent

Claude Code / Cursor / Codex / OpenCode 等 70+ 宿主通用,全量安装(12 个 Skill + core 共享知识库):

Works across 70+ hosts — Claude Code / Cursor / Codex / OpenCode and more; full install (12 skills + the core shared knowledge base):

$ npx skills add fishzjp/qa-skills --skill '*'

不用 npx?克隆仓库跑安装脚本(自动检测宿主目录),也可按 README 手动复制文件:

No npx? Clone the repo and run the install script (auto-detects host directories), or copy files manually per the README:

$ git clone https://github.com/fishzjp/qa-skills.git
$ cd qa-skills && ./install.sh --auto

用 DeepSeek Harness(dsh)?插件已上架 npm,一条命令装好:dsh plugin --profile web add dsh-qa-skills

On DeepSeek Harness (dsh)? The plugin is on npm, one command away: dsh plugin --profile web add dsh-qa-skills

无论哪种方式,core/ 必须一起装(方式一已自带)——单装某个 skill 不带 core,相对引用会断。

Whichever way you install, core/ must come along (Option 1 includes it) — a single skill without core breaks the relative references.

02

对 AI 助手说一句话Say one sentence to your agent

像给测试同事派活一样:

Like handing work to a QA colleague:

你,对 AI 助手说You, to your AI agent
「帮我测试这个需求:{需求描述 + 仓库地址}」“Test this requirement: {description + repo URL}”

需求分析、测试策略、用例设计、评审、执行、缺陷分析、回归与报告, 流水线就此跑通。只需要其中某一步,直接说需求就行,不必走全流程。

Requirement analysis, test strategy, case design, review, execution, defect analysis, regression and report — the pipeline runs end to end. Need just one stage? Say so and skip the rest.