Do the Thing

An open-source agent skill

Work, with receipts.

Give your AI a task. Get it back with evidence for every claim — what was asked, what was decided, and what was actually checked.

npx skills add beamonic/do-the-thing

Claude Code, Codex, and any agent that reads SKILL.md. MIT licensed.

Receipt example

Finish the retention job · SPEC-31 rev 2

  1. AC-02 Log how many transcripts were deleted Pass ran report_removed(3), count logged
  2. AC-01 Delete transcripts past the window Held waits on Q-03 · job deletes nothing until set
  3. Q-03 SPEC-31 says 1 year. Legal says 30 days. Which wins? Yours recommended: 30 days — Legal covers every copy
  4. Review Matches the spec · meets the standards Pass findings recorded per axis

Next Answer Q-03, set one line, rerun AC-01.

Handing work to an AI starts another job.

You supply the context. You answer the same questions twice. You check whether “done” means done, and when the session ends you piece together what happened.

Do the Thing turns that into one workflow your agent follows — and leaves a record you can check instead of a summary you have to trust.

How it works

One entry point. Say “do the thing”, “take this issue and work it”, or “resume the stopped work”.

  1. Read the sources

    The issue, the code, and what is already decided. Facts are the agent’s job, not yours.

  2. Interview

    One question at a time, each with a recommended answer. A question nobody answered stays open, and nothing is built on a guess.

  3. Big spec

    The user outcome, shared rules, and acceptance criteria, each with a stable ID.

  4. Mini specs

    The plan is critiqued, then split into tasks that each cite the criteria they serve.

  5. Execute and prove

    Tests are shown to catch the defect they guard against — not just shown to pass.

  6. Review on two axes

    Does it match the spec? Does it meet the standards? Recorded separately, with what was not reviewed.

  7. Checkpoint and resume

    A record good enough to pick up cold. A damaged checkpoint is kept as it is and worked around, never overwritten.

Three rules it does not bend

Facts are its job. Decisions are yours.

It looks things up instead of asking you. When two sources disagree, it does not pick one — not even a “safe default”. It asks, recommends, and keeps working on everything that does not depend on the answer.

Green tests are not “done”.

Done means every acceptance criterion has evidence. Thirty-one passing tests and one unchecked criterion is not done, and it will say which criterion.

A session ending loses nothing.

Checkpoints record the checks that actually ran and the revisions they ran against, so the next session starts where the last one stopped.

Fits what you already use

  • Any skills-aware agent. Claude Code and Codex are tested; the folder is self-contained.
  • Tracker optional. Linear has a contract. Without a tracker, a file in your repository is the work record.
  • Spec Kit aware. When the repository carries GitHub’s Spec Kit, its commands are called at every step.

What it is not

Not a hosted app, an OAuth integration, a background worker, or an autonomous webhook service. It is a folder of instructions, templates, and a small validator that your agent reads.

The name comes from the Korean 일 좀 해 — “do some work”. It answers to “두더띵으로 해줘” too.

Try it on your next task.

npx skills add beamonic/do-the-thing