Facts are its job. Decisions are yours.
It looks things up instead of asking you. When two sources disagree, it does not pick one — not even a “safe default”. It asks, recommends, and keeps working on everything that does not depend on the answer.
An open-source agent skill
Give your AI a task. Get it back with evidence for every claim — what was asked, what was decided, and what was actually checked.
npx skills add beamonic/do-the-thing
Claude Code, Codex, and any agent that reads SKILL.md. MIT licensed.
Finish the retention job · SPEC-31 rev 2
report_removed(3), count logged
Next Answer Q-03, set one line, rerun AC-01.
You supply the context. You answer the same questions twice. You check whether “done” means done, and when the session ends you piece together what happened.
Do the Thing turns that into one workflow your agent follows — and leaves a record you can check instead of a summary you have to trust.
One entry point. Say “do the thing”, “take this issue and work it”, or “resume the stopped work”.
The issue, the code, and what is already decided. Facts are the agent’s job, not yours.
One question at a time, each with a recommended answer. A question nobody answered stays open, and nothing is built on a guess.
The user outcome, shared rules, and acceptance criteria, each with a stable ID.
The plan is critiqued, then split into tasks that each cite the criteria they serve.
Tests are shown to catch the defect they guard against — not just shown to pass.
Does it match the spec? Does it meet the standards? Recorded separately, with what was not reviewed.
A record good enough to pick up cold. A damaged checkpoint is kept as it is and worked around, never overwritten.
It looks things up instead of asking you. When two sources disagree, it does not pick one — not even a “safe default”. It asks, recommends, and keeps working on everything that does not depend on the answer.
Done means every acceptance criterion has evidence. Thirty-one passing tests and one unchecked criterion is not done, and it will say which criterion.
Checkpoints record the checks that actually ran and the revisions they ran against, so the next session starts where the last one stopped.
The skill ships with a re-runnable eval suite. Pass criteria are committed before the runs that judge them.
Samples are small, and each results file says what it does not show. Read them.
Not a hosted app, an OAuth integration, a background worker, or an autonomous webhook service. It is a folder of instructions, templates, and a small validator that your agent reads.
The name comes from the Korean 일 좀 해 — “do some work”. It answers to “두더띵으로 해줘” too.
npx skills add beamonic/do-the-thing