Agent Skills for AI-native engineering with an executable Definition of Done: clean-api-docs, clean-debate and clean-refactoring, sharing a fail-closed verification kit. Install with: cargo install --locked clean-skills
  • Python 90.1%
  • Rust 7.1%
  • Shell 2.8%
Find a file
Repository files (latest commit first)
Filename Latest commit message Latest commit date
x f4715d9000 feat: put official pinned sources before the local doc index
clean-api-docs 1.1.0: version-pinned official sources move to rung 2; a local doc index becomes a rung-3 fallback for when they are unreachable or lack the resolved version. Adding docs to an index is an optional cache, not a second read. Decided on freshness and not measured; disclosed in the design and validation records. Crate 0.3.1.
2026-10-02 16:21:26 +09:30
docs feat: put official pinned sources before the local doc index 2026-10-02 16:21:26 +09:30
kit feat: add clean-api-docs 1.0.0 and clean-debate 2.1.0 2026-10-02 15:42:35 +09:30
skills feat: put official pinned sources before the local doc index 2026-10-02 16:21:26 +09:30
src feat: add clean-api-docs 1.0.0 and clean-debate 2.1.0 2026-10-02 15:42:35 +09:30
tests feat: add clean-api-docs 1.0.0 and clean-debate 2.1.0 2026-10-02 15:42:35 +09:30
tools feat: add clean-api-docs 1.0.0 and clean-debate 2.1.0 2026-10-02 15:42:35 +09:30
.gitignore feat: add the clean-skills installer crate 2026-10-01 14:35:18 +09:30
build.rs feat: add the clean-skills installer crate 2026-10-01 14:35:18 +09:30
Cargo.lock feat: put official pinned sources before the local doc index 2026-10-02 16:21:26 +09:30
Cargo.toml feat: put official pinned sources before the local doc index 2026-10-02 16:21:26 +09:30
CHANGELOG.md feat: put official pinned sources before the local doc index 2026-10-02 16:21:26 +09:30
LICENSE feat: publish clean-refactoring 2.1.0 and deliberate-delegate-verify 1.0.0 2026-10-01 13:25:58 +09:30
README.md feat: put official pinned sources before the local doc index 2026-10-02 16:21:26 +09:30

clean-skills

Agent Skills for AI-native engineering with an executable Definition of Done. They are portable across Claude Code and Codex, and work with any host that reads the Agent Skills format.

Skill Use it when
clean-api-docs An answer or change depends on the current, version-specific behaviour of a library, framework, SDK, CLI or hosted API. It resolves the version the project actually uses, reads the installed code and then that version's official docs, stops early and labels the source
clean-debate You want several strong models to work out a task independently and correct each other, then hand a qualified contract to a capable (not necessarily the strongest) executor. Fits feature work, migrations, decisions and documents
clean-refactoring A large application should be rebuilt from first principles instead of refactored in place

The idea

The strongest reasoning goes into the contract, cheaper execution goes into the work, and a fail-closed gate decides what is done. Agreement between models is not evidence. A plan is good enough to delegate when its acceptance checks are qualified:

  • Every mandatory check accepts a bound known-good specimen and genuinely rejects a bound defective one.
  • Every declared mutant is killed.
  • Selected probes have resolved receipts.

The probes include two kinds:

  • an interpretation probe, where independent interpreters predict observable outcomes;
  • an adversarial probe, a lazy-compliance attempt that counts as a bypass only when it demonstrably violates a requirement.

Both skills share one verification kit (kit/scripts/cr.py, Python 3.9+ standard library). It binds evidence to exact tree, suite and dependency digests. It enforces:

  • verifier-only runners, repetitions (pass^k), flaky-history rejection, a holdout, suite locks and reviews with veto;
  • prerequisite receipts and the control matrix.

Missing, stale, skipped, crashed or builder-produced evidence fails. The kit's self-test runs 113 end-to-end negative controls.

What the kit does not claim (see "Known limits" in the kit reference):

  • runner and reviewer identities are labels, not authentication;
  • a passing gate proves conformance to the contract in a stated environment, not that defects are absent;
  • the qualification probes are bounded diagnostics, not certificates.

Install

With Rust's cargo (the skills themselves need Python 3.9+ to run the kit):

cargo install --locked clean-skills
clean-skills install --host codex            # dry run: shows what it would write to ~/.agents/skills
clean-skills install --host codex --apply
clean-skills install --host claude --apply   # writes to ~/.claude/skills

clean-skills copies the skills bundled in its release and leaves a marker in each folder it writes, then prints the command that runs the kit's self-test from the installed copy. Re-running it upgrades a copy only while that copy is unchanged since install. An edited, extra or missing file, an unmarked folder or a symlink is reported as a conflict and left alone. clean-skills uninstall --host <host> --apply removes only its own unchanged copies. --dir PATH targets another folder, such as a project's .claude/skills. If ~/.claude or its skills folder is a symlink (a dotfiles setup, say), it works in the real folder the link points to.

From a checkout instead, linking the skills to the repository:

git clone https://x.vibewait.ing/x/clean-skills.git
cd clean-skills
tools/install.sh --host codex     # dry run: shows the links it would create in ~/.agents/skills
tools/install.sh --host codex --apply
tools/install.sh --host claude --apply   # links into ~/.claude/skills

That installer only creates and removes its own symlinks. Neither installer overwrites anything else, and neither edits hooks, permissions or settings. Invoke a skill with /clean-debate in Claude Code, or $clean-debate in Codex. Then check the kit from the checkout:

python3 skills/clean-debate/scripts/cr.py selftest

Layout

kit/        single source (kit/SKILLS lists the skills that receive a copy): scripts/ (cr.py, crkit/, stop_gate.py, tests/), shared references and templates, VERSION
skills/     self-contained skills (clean-api-docs, clean-debate, clean-refactoring); the two listed in kit/SKILLS get a generated kit copy (scripts/KIT_VERSION records the kit digest)
src/        the clean-skills installer crate (embeds skills/ at build time; see build.rs)
tools/      sync-kit.sh (generate copies; --check detects drift), assemble.sh (dist/), install.sh, check.sh
docs/       design records, research and validation evidence

Edit the kit only under kit/, then run tools/sync-kit.sh. Before publishing, run tools/check.sh. It:

  • assembles dist/;
  • runs the kit self-test and template manifest check on every skill listed in kit/SKILLS, unit tests on every skill that has them, and skill validation on all;
  • runs lint and a publication scan;
  • runs negative controls against the tooling itself: the drift check, an unlisted skill surviving the kit sync, symlinks in the kit source and copies, the publication scan, the installer and the gate;
  • runs cargo fmt, clippy and the installer's tests, then builds the packaged crate, checks it ships skills/ byte for byte, installs from it into a scratch home, runs the kit self-test there and uninstalls.

A missing required tool makes the result INCOMPLETE, never a pass.

Status

  • clean-api-docs 1.1.0. Its helper reads local files only and never uses the network.
    • 1.1.0 puts version-pinned official sources before a local doc index, which is now a fallback. That order was decided on freshness and has not been measured; the numbers below are for 1.0.0.
    • Designed and built with clean-debate (two vendors' sealed briefs, a verifier, cheaper executors, four audits); see the design record.
    • 1.0.0 was measured against the previous skill on a frozen 12-item eval, run twice per arm: equal accuracy (23/24 each), 1.15x total tokens (+3.5% uncached) and equal safety flags. That failed the pre-registered token gate, and it was published on an explicit decision with these numbers. See the validation record for what the sample can and cannot show.
  • clean-refactoring 2.1.0.
    • Kit 2.0 was reviewed in three cross-model cycles.
    • Kit 2.1 was reviewed in further independent rounds; every finding became a regression scenario or a tooling control. See the validation record.
    • A fresh agent's documentation smoke test passed end to end.
    • Mutation testing caught every deliberate gate break.
  • clean-debate 2.1.0 (named deliberate-delegate-verify until 2.0.0). The design was settled with an external reviewer model after three cycles (see docs/clean-debate/). It was dogfooded end to end to build clean-api-docs; the 17 friction items that run surfaced are folded into 2.1.0. clean-refactoring has not yet been run as a full multi-model workflow on a production application.

MIT licensed. See LICENSE.