1 comments

  • LambdaLogic an hour ago

    AI slop in code has a pretty recognizable shape: code that was never actually verified, merged because the agent sounded confident.

    “Done, all tests pass.”

    Except they don’t.

    It’s usually not malicious. The agent just can’t reliably tell the difference between what it intended to do and what it actually did. So the fix isn’t necessarily a better prompt. It’s taking verification authority away from the agent itself.

    That’s what The Executor does.

    It’s a set of Markdown skills + POSIX bash scripts that takes a feature through:

    intake → spec → plan → execution → review → verification → handoff

    The important part is that state transitions are enforced by scripts that fail with a non-zero exit code, rather than by instructions the agent is expected to follow.

    It works with Claude Code, Codex, or pretty much any harness that can read skill files. No daemon. No SaaS.

    How it makes agents prove their work:

    Reviewers see the diff, not the implementer’s claims.

    The implementation agent’s “looks good” report is not evidence. Review verdicts are stored as files, and a FAIL verdict physically blocks exec-branch merge. The merge command simply refuses to run.

    Every finding has to point somewhere specific.

    There’s one ID namespace per initiative, and every finding references the exact spec requirement it violates.

    For example:

    INIT-0004-P01-T03-R02 fails INIT-0004-SPEC-01-R07

    No vague “this doesn’t look right” findings with nowhere to attach them.

    Verification happens from scratch.

    Each spec criterion gets its verification command run against the current commit, producing one of:

    PROVEN / FAILED / NOT-RUN / UNAVAILABLE

    And this matters: a single NOT-RUN means you cannot call the feature “complete.” Nothing gets upgraded by inference.

    Fix loops don’t continue forever.

    At round 4, the process escalates to a different model. Hitting the cap requires an explicit recorded ruling, so findings can’t just disappear because the agent got tired of fixing them.

    Context resets don’t reset reality.

    After a context reset or model switch, the controller re-reads the ledger instead of trusting the model to remember what happened.

    exec-run check audits the registry, ledger, and verdict files against each other and exits with 1 while naming the inconsistency.

    There is a cost, though: ceremony and tokens.

    This isn’t really for “change one line and ship.” It’s aimed at features that live for days and pass through multiple rounds of implementation, review, and verification.

    MIT licensed.

    Roast it. Where does this break down at your scale?