软件发布之前,Before any release ships,
总得有人先把住质量关someone has to hold the line on quality
每一款软件上线前,都要回答成百上千个「如果」: 并发下单会不会超卖?优惠券到期会不会自动失效?接口异常时数据会不会错乱?
Before launch, every piece of software has to answer hundreds of “what ifs”: Will concurrent orders oversell? Will coupons expire on their own? Will data corrupt when APIs fail?
干这件事的,是软件测试工程师: 在缺陷到达用户之前发现它、定位它,再确认它被修好。
That job belongs to the QA engineer: find defects before users do, localize them, and verify the fix.
qa-skills 做的,就是让 AI 助手学会这套本事。qa-skills teaches that craft to your AI agent.
能写,更要能执行。Writing cases is easy. Executing them is the bar.
通用大模型写出的用例,常常看着专业、实际没法执行:
判定模糊、占位符没替换、时限没说清、入口是编的。
这是能力问题,更是纪律问题。
Cases from general-purpose models often look professional but can't run:
vague verdicts, placeholders left in, unstated timeouts, made-up entry points.
That's a capability gap — and a discipline gap.
判定模糊,无法执行Vague verdict, unrunnable
预期结果不可判定,等于没有验证:An undecidable expected result is no verification at all:
「正常」是什么标准?谁判定?等多久算超时?
What's the bar for “correctly”? Who decides? How long before timeout?
判定不了的用例,覆盖再全也计零分。A case that can't be judged scores zero — however broad its coverage.
步骤精确,判定可量化Precise steps, measurable verdicts
前置、步骤、预期、时限,四要素齐备:Preconditions, steps, expected results, timeouts — all four present:
预期:1 小时内状态自动变为「已结束」,
超过 1 小时即判失败。」
Expected: status flips to “Ended” automatically within 1 hour;
any longer counts as a failure.”
任何测试人员拿到就能执行,不用讲解。Any tester can execute it as-is, no briefing needed.
产出标准只有一条:
任何一个没接触过这个需求的人,拿到文件就能直接执行。
背后是 8 条可执行性硬标准,评测中一票否决。
One output bar only:
anyone who never touched this requirement can pick up the file and run it.
Backed by 8 hard executability standards — veto power in evaluation.
不是软件,Not software —
是一套装进 AI 的工作方法。a working method installed into AI.
由 12 个 QA Skill 和一个共享知识库(core/)组成,全部是 Markdown 指令文件。 放对目录就能生效。
12 QA skills plus a shared knowledge base (core/), all plain Markdown instruction files. Drop them in the right directory and they take effect.
12 个 QA Skill + 共享知识库12 QA skills + shared knowledge base
需求分析、策略、用例、评审、执行、缺陷、回归,每个环节一个专属 Skill,按需加载。
One dedicated skill per stage — requirements, strategy, cases, review, execution, defects, regression — loaded on demand.
一条自动接力的流水线A pipeline that hands off automatically
每个阶段的产出都写成文件,不靠对话记忆。中断了,新开一个对话接着跑。
Every stage writes its output to files, not chat memory. Interrupted? Start a new session and continue.
关键决定,由你拍板Key decisions stay with you
需求没写清的,先问你;Bug 怎么定性、预算设多少,AI 助手只提案、不代答。你的裁决会记录在案,后续不得推翻。
Unclear requirements? It asks you first. Bug severity, budget caps — the agent proposes, never decides. Your rulings are recorded and never overridden later.
一句指令,从需求到报告。One instruction, requirement to report.
对 AI 助手说「帮我测试这个需求」,从需求到报告各环节自动接力,每个环节都有明确的产出物(完整流水线共九个阶段,含探索旁路与回归验证)。
Say “test this requirement” to your agent and the stages chain automatically, each with a concrete deliverable (the full pipeline has nine stages, including the exploratory bypass and regression verification).
需求分析Requirement analysis
不懂就先问Ask before assuming测试策略Test strategy
锁定高危区域Lock the high-risk zones用例设计Case design
人和机器都能读Readable by humans and machines用例评审Case review
再挑一遍错One more pass for defects自动执行Automation
浏览器 + 接口Browser + API缺陷分析Defect analysis
追到出错代码Down to the failing line回归 + 报告Regression + report
证据齐全Evidence included从用例到回归,From cases to regression,
每道工序都有对应的 Skill。every stage has its own skill.
每个 Skill 单独可用,也能串成完整流水线。需要哪一段,就用哪一段。
Each skill works standalone or as part of the full pipeline. Use exactly the stage you need.
用例设计Case design
代码优先:先读实现、审出潜在缺陷,再落笔。产出两份,markmap 脑图给人读,schema.yaml 给机器校验。
Code-first: read the implementation, hunt latent bugs, then write. Two deliverables — a markmap for humans, schema.yaml for machines.
测试策略Test strategy
风险评级必须给出证据;性能、安全、并发等十个类型,逐一决定测还是不测,每个决定都留下记录。说不清依据的策略,等于没有策略。
Risk ratings must cite evidence; ten test types — performance, security, concurrency and more — each explicitly tested or not, every decision on record. A strategy that can't explain itself is no strategy.
用例评审Case review
把所有能测的点列全,作为基准逐条复审:覆盖够不够、能不能执行,直接修订文件并留下审查记录。
Enumerate every testable point as a baseline, then re-review line by line: coverage enough? executable? Files get revised in place with a review record.
自动化执行Automation
用例转成可运行的 Playwright / pytest 代码:按 Page Object 规范组织,测试数据自己造、跑完自己清。
Cases become runnable Playwright / pytest code: Page Object conventions, self-built test data, self-cleaning after runs.
缺陷分析Defect analysis
先复现,再读代码定位到出错的那一行,然后分析影响面、给出回归建议。产出的 Bug 条目带根因、带证据。
Reproduce first, then read the code down to the failing line, analyze blast radius, recommend regression scope. Bug entries carry root cause and evidence.
回归测试Regression testing
从代码改动推算波及范围,生成回归清单。防止修一个、坏三个。
Derive regression scope from the code diff and generate the retest list. Stop fixing one thing and breaking three.
每一个数字,都来自实测。Every number comes from measurement.
同一批任务,同一个模型,唯一差异是装没装 qa-skills。
数字如实披露,包括不好看的那些。
Same tasks, same model; the only difference is whether qa-skills is installed.
All numbers disclosed as measured — the unflattering ones too.
用例合格率Case pass rate
格式与内容逐条核对format & content checked item by item
不装时 26%26% without植入缺陷检出率Seeded-bug detection
代码审查任务实测measured on code-review tasks
换一套裁判评:100%100% under an alternate judgeE2E 用例真实执行通过E2E cases passing real runs
单任务×3 采样真浏览器实测real-browser runs, 1 task × 3 samples
不装时 0/90/9 without类型查全率Test-type recall
最弱模型上实测measured on the weakest model
不装时为 00 withoutToken 成本约 3.3 倍。覆盖提升:写用例 +8.7pp、全任务 +13.2pp、找缺陷 +9.7pp。 完整数据见 README,里程碑版 Release 附跨模型增益矩阵快照。
Token cost ≈ 3.3×. Coverage gains: case writing +8.7pp, all tasks +13.2pp, defect finding +9.7pp. Full data in the README; milestone releases ship a cross-model gain-matrix snapshot.
它知道,什么该由你决定。It knows what's yours to decide.
需求有歧义,先澄清;执行走人工还是自动化,你选; Bug 怎么定性、预算上限设多少,你说了算。 这四类检查点,AI 助手只提案、不代答。
Ambiguous requirements get clarified first; manual or automated execution, you pick; bug severity, budget caps — you decide. At these four checkpoints the agent proposes, never decides.
每一次裁决都记录在案,后续阶段照办,不得推翻。 拍板的,始终是你。
Every ruling is recorded; later stages comply, no override. The call is always yours.
后续阶段照办,不得推翻14:32 ruling recorded → decisions.yaml
later stages comply, no override
装之前的疑虑,Every pre-install doubt,
逐条如实回答。answered honestly.
包括不那么好听的答案——Token 更贵、Skill 正文还没双语,都直接写在这里。
Including the unflattering ones — tokens cost more, skill bodies aren't bilingual yet.
qa-skills 是什么?要花钱吗?What is qa-skills? Does it cost money?
我的 AI 助手能用吗?Will it work with my AI assistant?
npx skills add fishzjp/qa-skills --skill '*' 一行装到 Claude Code / Cursor / Codex / OpenCode 等 70+ 宿主。宿主不支持子代理时,流水线自动退化为顺序会话 + 文件衔接,正确性不受影响。npx skills add fishzjp/qa-skills --skill '*' installs into Claude Code / Cursor / Codex / OpenCode and 70+ other hosts. Where the host lacks subagents, the pipeline degrades to sequential sessions joined by files — correctness is unaffected.AI 产出的用例,真的能直接执行吗?Can the AI's cases really run as-is?
装完还要学一遍怎么用吗?Is there a setup tutorial to go through?
我的代码和数据会被上传吗?Will my code and data get uploaded?
支持英文吗?Does it support English?
两步接入你的 AI 助手Two steps to your AI assistant
不需要账号,不需要订阅。两条命令,一句话。
No account, no subscription. Two commands, one sentence.
一行命令,装进 AI 编程助手One command, installed into your coding agent
Claude Code / Cursor / Codex / OpenCode 等 70+ 宿主通用,全量安装(12 个 Skill + core 共享知识库):
Works across 70+ hosts — Claude Code / Cursor / Codex / OpenCode and more; full install (12 skills + the core shared knowledge base):
$ npx skills add fishzjp/qa-skills --skill '*'
不用 npx?克隆仓库跑安装脚本(自动检测宿主目录),也可按 README 手动复制文件:
No npx? Clone the repo and run the install script (auto-detects host directories), or copy files manually per the README:
$ git clone https://github.com/fishzjp/qa-skills.git $ cd qa-skills && ./install.sh --auto
用 DeepSeek Harness(dsh)?插件已上架 npm,一条命令装好:dsh plugin --profile web add dsh-qa-skills
On DeepSeek Harness (dsh)? The plugin is on npm, one command away: dsh plugin --profile web add dsh-qa-skills
无论哪种方式,core/ 必须一起装(方式一已自带)——单装某个 skill 不带 core,相对引用会断。
Whichever way you install, core/ must come along (Option 1 includes it) — a single skill without core breaks the relative references.
对 AI 助手说一句话Say one sentence to your agent
像给测试同事派活一样:
Like handing work to a QA colleague:
需求分析、测试策略、用例设计、评审、执行、缺陷分析、回归与报告, 流水线就此跑通。只需要其中某一步,直接说需求就行,不必走全流程。
Requirement analysis, test strategy, case design, review, execution, defect analysis, regression and report — the pipeline runs end to end. Need just one stage? Say so and skip the rest.