- Python 90.1%
- Rust 7.1%
- Shell 2.8%
| Filename | Latest commit message | Latest commit date |
|---|---|---|
clean-api-docs 1.1.0: version-pinned official sources move to rung 2; a local doc index becomes a rung-3 fallback for when they are unreachable or lack the resolved version. Adding docs to an index is an optional cache, not a second read. Decided on freshness and not measured; disclosed in the design and validation records. Crate 0.3.1. |
||
| docs | ||
| kit | ||
| skills | ||
| src | ||
| tests | ||
| tools | ||
| .gitignore | ||
| build.rs | ||
| Cargo.lock | ||
| Cargo.toml | ||
| CHANGELOG.md | ||
| LICENSE | ||
| README.md | ||
clean-skills
Agent Skills for AI-native engineering with an executable Definition of Done. They are portable across Claude Code and Codex, and work with any host that reads the Agent Skills format.
| Skill | Use it when |
|---|---|
clean-api-docs |
An answer or change depends on the current, version-specific behaviour of a library, framework, SDK, CLI or hosted API. It resolves the version the project actually uses, reads the installed code and then that version's official docs, stops early and labels the source |
clean-debate |
You want several strong models to work out a task independently and correct each other, then hand a qualified contract to a capable (not necessarily the strongest) executor. Fits feature work, migrations, decisions and documents |
clean-refactoring |
A large application should be rebuilt from first principles instead of refactored in place |
The idea
The strongest reasoning goes into the contract, cheaper execution goes into the work, and a fail-closed gate decides what is done. Agreement between models is not evidence. A plan is good enough to delegate when its acceptance checks are qualified:
- Every mandatory check accepts a bound known-good specimen and genuinely rejects a bound defective one.
- Every declared mutant is killed.
- Selected probes have resolved receipts.
The probes include two kinds:
- an interpretation probe, where independent interpreters predict observable outcomes;
- an adversarial probe, a lazy-compliance attempt that counts as a bypass only when it demonstrably violates a requirement.
Both skills share one verification kit (kit/scripts/cr.py, Python 3.9+ standard
library). It binds evidence to exact tree, suite and dependency digests. It enforces:
- verifier-only runners, repetitions (pass^k), flaky-history rejection, a holdout, suite locks and reviews with veto;
- prerequisite receipts and the control matrix.
Missing, stale, skipped, crashed or builder-produced evidence fails. The kit's self-test runs 113 end-to-end negative controls.
What the kit does not claim (see "Known limits" in the kit reference):
- runner and reviewer identities are labels, not authentication;
- a passing gate proves conformance to the contract in a stated environment, not that defects are absent;
- the qualification probes are bounded diagnostics, not certificates.
Install
With Rust's cargo (the skills themselves need Python 3.9+ to run the kit):
cargo install --locked clean-skills
clean-skills install --host codex # dry run: shows what it would write to ~/.agents/skills
clean-skills install --host codex --apply
clean-skills install --host claude --apply # writes to ~/.claude/skills
clean-skills copies the skills bundled in its release and leaves a marker in each
folder it writes, then prints the command that runs the kit's self-test from the
installed copy. Re-running it upgrades a copy only while that copy is unchanged since
install. An edited, extra or missing file, an unmarked folder or a symlink is reported
as a conflict and left alone. clean-skills uninstall --host <host> --apply removes
only its own unchanged copies. --dir PATH targets another folder, such as a
project's .claude/skills. If ~/.claude or its skills folder is a symlink (a
dotfiles setup, say), it works in the real folder the link points to.
From a checkout instead, linking the skills to the repository:
git clone https://x.vibewait.ing/x/clean-skills.git
cd clean-skills
tools/install.sh --host codex # dry run: shows the links it would create in ~/.agents/skills
tools/install.sh --host codex --apply
tools/install.sh --host claude --apply # links into ~/.claude/skills
That installer only creates and removes its own symlinks. Neither installer
overwrites anything else, and neither edits hooks, permissions or settings. Invoke a skill
with /clean-debate in Claude Code, or $clean-debate
in Codex. Then check the kit from the checkout:
python3 skills/clean-debate/scripts/cr.py selftest
Layout
kit/ single source (kit/SKILLS lists the skills that receive a copy): scripts/ (cr.py, crkit/, stop_gate.py, tests/), shared references and templates, VERSION
skills/ self-contained skills (clean-api-docs, clean-debate, clean-refactoring); the two listed in kit/SKILLS get a generated kit copy (scripts/KIT_VERSION records the kit digest)
src/ the clean-skills installer crate (embeds skills/ at build time; see build.rs)
tools/ sync-kit.sh (generate copies; --check detects drift), assemble.sh (dist/), install.sh, check.sh
docs/ design records, research and validation evidence
Edit the kit only under kit/, then run tools/sync-kit.sh. Before publishing, run
tools/check.sh. It:
- assembles
dist/; - runs the kit self-test and template manifest check on every skill listed in
kit/SKILLS, unit tests on every skill that has them, and skill validation on all; - runs lint and a publication scan;
- runs negative controls against the tooling itself: the drift check, an unlisted skill surviving the kit sync, symlinks in the kit source and copies, the publication scan, the installer and the gate;
- runs
cargo fmt,clippyand the installer's tests, then builds the packaged crate, checks it shipsskills/byte for byte, installs from it into a scratch home, runs the kit self-test there and uninstalls.
A missing required tool makes the result INCOMPLETE, never a pass.
Status
clean-api-docs1.1.0. Its helper reads local files only and never uses the network.- 1.1.0 puts version-pinned official sources before a local doc index, which is now a fallback. That order was decided on freshness and has not been measured; the numbers below are for 1.0.0.
- Designed and built with
clean-debate(two vendors' sealed briefs, a verifier, cheaper executors, four audits); see the design record. - 1.0.0 was measured against the previous skill on a frozen 12-item eval, run twice per arm: equal accuracy (23/24 each), 1.15x total tokens (+3.5% uncached) and equal safety flags. That failed the pre-registered token gate, and it was published on an explicit decision with these numbers. See the validation record for what the sample can and cannot show.
clean-refactoring2.1.0.- Kit 2.0 was reviewed in three cross-model cycles.
- Kit 2.1 was reviewed in further independent rounds; every finding became a regression scenario or a tooling control. See the validation record.
- A fresh agent's documentation smoke test passed end to end.
- Mutation testing caught every deliberate gate break.
clean-debate2.1.0 (nameddeliberate-delegate-verifyuntil 2.0.0). The design was settled with an external reviewer model after three cycles (seedocs/clean-debate/). It was dogfooded end to end to buildclean-api-docs; the 17 friction items that run surfaced are folded into 2.1.0.clean-refactoringhas not yet been run as a full multi-model workflow on a production application.
MIT licensed. See LICENSE.