Finish the job. Skip the busywork.
A Skill + Guard to curb unnecessary defenses and scope creep in AI coding agents.
Codex · Claude Code · OpenCode · Hermes Agent CLI · Pi
SHIT philosophy ·
See an example ·
Install ·
Cases ·
中文 ·
한국어
You ask your agent to export a file. It adds a SHA-256 checksum nobody reads. "Just in case."
You ask for a review; it finds a bug and starts editing. One fix becomes a plan to refactor neighboring modules. The checks have answered the question, but it wants another agent to "make sure one last time."
Every step has a careful explanation. The result is still missing, the tokens keep going, and you're left keeping the agent on task.
Still no result. Now I'm supervising the agent.
I tried adding rules to AGENTS.md: "do not edit," "do not overengineer,"
"ask before doing extra work." Every frustration became another rule.
Eventually, AGENTS.md was overengineered too.
Stop That Shit. Finish the work the task needs. Stop piling on work it doesn't.
We call these behaviors SHIT:
- S — Scope creep. A fix expands into an unrelated refactor. Complete the callers, data migrations, and tests that the request requires. Stop additions with no current purpose.
- H — Hashing & hypothetical hardening. An unread checksum or a compatibility layer for an imagined future. Protection must detect a problem and lead to rejection, recovery, or diagnosis. Keep effective defenses, omit unused ones, and repair those that swallow failures, report false success, or duplicate side effects.
- I — Intent violation. You said read-only review; the files changed anyway. Review means reporting findings; editing needs change authority. Respect the user's boundary.
- T — Task thrashing. The reading, testing, and reviewing restart with no change or new question. Reuse evidence that answers the current question. Add relevant checks when code or acceptance conditions change. Finish once the work is done and sufficiently verified.
Meet the task's responsibilities in full, and let real needs drive complexity. If that requires more files, a migration, or cross-component tests, complete them.
The next step reads only report.csv. No release check or other step uses a
digest, yet the export includes one:
await writeFile("report.csv", csv);
await writeFile("report.csv.sha256", createHash("sha256").update(csv).digest("hex"));Applying the STS rule leaves:
await writeFile("report.csv", csv);This simplified example drops the unused checksum and still delivers the file. More examples are in the case catalogue.
If the release process reads the checksum and rejects mismatches, keep it.
Set hash=allow to authorize it.
Complete the necessary work. If a configuration migration must support already shipped data, finish the migration, update readers and writers, and test compatibility, even if the diff grows. Do not leave broken callers behind.
The 18 public Bad/Good decision cases check whether rules stop unnecessary actions and allow required work. They cover read-only review, affected callers, shipped-data migrations, release checksums, and file, dependency, and subagent boundaries. Model behavior on real tasks is evaluated separately.
The evidence record includes tests, historical Codex / GPT-5.6 runs, and results with no difference. See the paired evaluation guide to reproduce the comparisons.
When this SHIT appears, read the relevant code and trace the call path. Then decide in this order:
1. What must be delivered? → Establish the result, authority, and support commitments
2. Can existing tools do it? → Start with a direct solution
3. What concrete gap remains? → Add the behavior and protection it needs
4. What does this defense do? → Keep effective protection, omit waste, repair harm
5. Is the result verified? → Finish when no in-scope blocker remains
Check existing code, standard libraries, native platform features, and installed dependencies. Use what fits the required behavior and state lifetime. Add a mechanism when a concrete gap calls for it.
A new optional mechanism without a purpose can wait. Existing protection with an unclear role needs inspection before removal. Real trust boundaries and support commitments justify protection before an incident occurs.
Line counts, file counts, and check counts cannot replace this judgment. Read the full rules.
Most hosts require Node.js 18 or newer; Pi 0.84.4 itself requires Node.js 22.19 or newer. See INSTALL.md for the full setup.
Expand your host. For guidance without runtime enforcement, install only the Skill.
Claude Code
Download and extract the 0.2.4 source, then run from the checkout root:
claude plugin validate .
claude plugin marketplace add ./
claude plugin install stop-that-shit@stop-that-shitRestart Claude Code or run /reload-plugins, then invoke:
/stop-that-shit:stop-that-shit review -- Review this diff. Report findings; do not edit.
Codex
codex plugin marketplace add lennney/stop-that-shit --ref 0.2.4
codex plugin add stop-that-shit@stop-that-shit--ref pins the install to a version tag instead of mutable
main. Restart Codex. In a fresh CLI TUI, enter /hooks, compare the
packaged Hook list, and trust the
commands after inspection. You can
also give INSTALL_FOR_AGENTS.md to Codex for the
non-interactive steps.
OpenCode from GitHub
OpenCode 1.18.18 or newer can install this repository globally without cloning it:
opencode plugin github:lennney/stop-that-shit -gRestart OpenCode and use $stop-that-shit review -- .... The command installs
the Guard; the bundled Skill and optional /sts alias are not registered
automatically. See INSTALL.md for
details.
Hermes Agent CLI
Requires Node.js 18+.
hermes plugins install lennney/stop-that-shit/.hermes-plugin --no-enable
hermes plugins enable stop-that-shit
hermes plugins listAfter enabling it, CLI users need to start a new Hermes CLI process or session; Gateway users need to run:
hermes gateway restartThese steps are not required every time the plugin is used. The corresponding Hermes process only needs to be restarted after enabling, disabling, updating, rolling back, or reinstalling the plugin.
Pi Coding Agent
The current adapter is pinned and tested against
@earendil-works/pi-coding-agent 0.84.4. Install a checkout that contains the
adapter:
pi install /absolute/path/to/stop-that-shitStart a new Pi process, or run /reload in the TUI after resource changes. Then
invoke:
/skill:stop-that-shit review -- Review this diff. Report findings; do not edit.
The pinned tag includes the Pi adapter and both Skills. See INSTALL.md.
Oh My Pi (unreleased candidate)
The separate Extension was checked with @oh-my-pi/pi-coding-agent 18.4.4.
Published version 0.2.4 does not include OMP. Start from a local checkout
that contains this candidate:
omp -e /absolute/path/to/stop-that-shit/omp/stop-that-shit.ts --skills /absolute/path/to/stop-that-shit/skills --sts-contract "review agents=0 -- inspect"Confirm that OMP loaded the Extension. Use /sts change -- ... to change mode
or /sts status to inspect the contract. These native commands also work in
RPC. The TUI also accepts first-line $stop-that-shit directives.
See INSTALL.md.
In Codex or a host-neutral prompt, start the first non-empty line with one
directive, outside quotes and code blocks. Put task text after --:
$stop-that-shit review -- Review this diff. Report findings; do not edit.
$stop-that-shit change -- Fix the config read failure, including affected callers and checks.
The Claude Code plugin uses /stop-that-shit:stop-that-shit.
The Pi Skill uses /skill:stop-that-shit.
| Command | Purpose |
|---|---|
review |
Review and report findings |
answer |
Answer a question |
monitor |
Keep checking and reporting |
change |
Permit changes needed to complete the task |
watch |
Observe actions without blocking |
status / runtime |
Inspect the current state and local records |
Add parameters when the task needs a more specific boundary:
$stop-that-shit lock change files=src/config.cjs|test/config.test.cjs -- Fix this behavior.
$stop-that-shit change deps=allow -- Add the requested parser dependency.
$stop-that-shit change hash=allow -- Generate the checksum required by the release process.
$stop-that-shit change agents=1 -- Use one independent test subagent.
If you do not yet know every affected file, trace the call path before setting files=.
The Skill guides engineering judgment. The Guard checks explicit boundaries before supported tool calls:
| Action | Default when the Guard is armed |
|---|---|
Write during review, answer, or monitor |
Stop |
| Add a dependency | Ask; deps=allow permits it |
| Launch a subagent | Stop above the agents=N budget |
| Add a recognized hash operation | Stop; hash=allow permits it |
Write outside files= |
Stop |
The Guard does not decide that a cache, retry, or migration is unnecessary from its name. See the Adapter contract for host integration.
The work is done, but the agent keeps talking. STSS removes unused defenses, tightens repeated hedging, and keeps conditions that affect a decision. An example from the fixed offline cases:
INPUT We may perhaps potentially be able to finish the migration in roughly six to
eight weeks, depending on access approval.
OUTPUT We estimate six to eight weeks, subject to access approval.
The range and approval condition remain. If approval is required before work starts, state that timing explicitly: “The estimated migration time is six to eight weeks after access approval.”
The full plugin includes STSS. You can also install it alone, without a Hook. From the repository root:
npx skills add ./skills/stss --globalrewrite edits supplied text and is the default mode. audit reports findings
and the smallest proposed fix.
| Host | Rewrite | Audit |
|---|---|---|
| Codex | $stss rewrite -- Make this proposal direct. |
$stss audit -- Find defensive padding. |
| Claude Code plugin | /stop-that-shit:stss rewrite -- ... |
/stop-that-shit:stss audit -- ... |
| Standalone Claude Skill | /stss rewrite -- ... |
/stss audit -- ... |
| Pi | /skill:stss rewrite -- ... |
/skill:stss audit -- ... |
See the STSS Skill for the full method and the six STSS examples for paired cases, including wording that must remain.
Should every checksum go?
Release integrity checks, deduplication, and skipping repeated processing can
all give a checksum a job. Comparing a file against a trusted expected digest
and rejecting a mismatch is useful even if it adds computation. Use hash=allow
when it is needed. The Guard checks authority; the task determines the purpose.
I installed it. Why is nothing blocked yet?
Installation starts in OBSERVING / unconfirmed. An explicit review, answer,
monitor, or change sets the Guard to ARMED. watch always observes without
blocking. A Skill-only installation has no runtime enforcement.
Does a stop stamp prove the action never ran?
permission_deny_returned (execution_denial_returned in OpenCode) means the
Guard returned a denial. The host's final action needs separate observation,
so the Runtime records hostEffect: unobserved. The host sandbox handles
security isolation.
Does it save my code or conversations?
Runtime events do not store code or conversation text. Separate session state retains explicit paths, correlation IDs, and directive errors. Inspect that state before sharing plugin data. See PRIVACY.md.
How do I update, uninstall, or work on it?
See update checks, disabling and uninstalling, local development, and the changelog.
Your agent invented more work? Report a Bad Case.
STS blocked something the task needed? Report a Good Case. That matters just as much.
Tell us what you requested, what the agent added or left unfinished, and which fact would change the decision. See the case catalogue for examples and the contribution guide for sanitizing and reproducing them.
I do not want another hundred prohibitions in AGENTS.md. I want each case to
make the next decision better.
If this helps, give it a star. You can also join the conversation on LINUX DO.