Skill · newcomer · finished artifact
Check whether a newcomer can use and understand what you are about to ship.
An installable skill for AI agents. Cold-eye checks whether a newcomer can understand and use what you are about to ship. Give it a finished artifact: a guide, a README, a skill, a spec, a site, a package, or a repo.
Your agent follows the skill and writes a critique file. Verdict first, then a ranked list of missing instructions, contradictions, and unsupported claims. The subject does not change. Hostile is the stance after you know the job: no credit for intent.
One example (illustrative)
You wrote: “Install the package and launch the app.”
Cold-eye finds: the guide names the package but never gives the launch command. A new user cannot finish setup from the instructions given.
You get: a ranked finding that quotes the line, says what a newcomer does instead, and says what to add.
# Cold-eye — setup guide
**Verdict:** close
## Ranked changes
1. **F-001** · test 2 — invented
> Install the package and launch the app.
Cold reader: installs the package, then guesses a launch command or stops.
Put: The exact launch command on the next line, and what the reader sees when the app is running.
Test 2 asks what a newcomer still has to invent. Here it is the launch command. One missing command is a small edit, so the verdict is close. The guide itself is unchanged.
Install in your agent · See a clean excerpt · See the first-run sample
What it reads, writes, and changes
| Reads | A finished skill, spec, page, package, site, or repo |
| Writes | <name>.cold-eye.md next to a file, or cold-eye.md at a system root. Chat only, if you ask |
| Changes | Nothing, unless you separately ask for an edit |
| Verdict | Means |
|---|---|
| holds | A newcomer can run the job from the file |
| close | Almost. A few edits to the file would close the gap |
| fails a hostile read | The reader still has to invent too much |
A repo, a package, or a site may get a split verdict: one for the main file read alone, one for all the files read together.
Clean verdict (illustrative)
Labeled example. A one-page checklist that names the file, the output path, and when to stop.
# Cold-eye — clean checklist
**Verdict:** holds
## Ranked changes
None.
## Protect
The four steps, and the line that says write `None.` when nothing is missing.
No invented faults. The file already tells a newcomer how to start, what to read, what to write, and when to stop.
The first-run sample
Get started has you run Cold-eye on a four-step sample checklist, until-ready.md. Step 4 says to repeat the review “until ready” and never says what ready means. A run should produce something with this shape:
# Cold-eye — review until ready
**Verdict:** fails a hostile read
## Ranked changes
1. **F-001** · test 10 — no_close
Absent: the procedure, what “ready” means
Cold reader: keeps repeating with no exit.
Put: Name the readiness criterion, then hand the file over and stop.
How to check your result:
- The finding should point at step 4, the unresolved “until ready” condition.
- The checklist should be unchanged.
- Wording varies by model. Check the shape, not an exact match. A second finding, such as
notes.mdnever being described, is fair if it points at the file. - A corrected copy, until-ready-fixed.md, defines readiness and says when to stop. It should not get the step 4 finding.
- A failing verdict alone does not prove your agent loaded the skill. Check the agent's skill list, or ask it to quote the first heading of
SKILL.md.
The repo's fixtures/ folder holds more sample subjects. Each has an expected.md naming the injected fault and the test that should catch it. The clean sample has no fault, so a wording-only finding on it is a miss.
What it checks
Ten questions, in order. Rank by what a newcomer hits first.
- Can the reader run the job as steps: start, read, write, stop? Stages and principles are not a procedure.
- Are required names present: schema, filenames, write-back, and what done looks like? A placeholder the procedure never binds is a guess.
- Do the files match each other, and does what shipped match the page?
- If a claimed check needs a second command, is that command next to the claim?
- Does a definition include a worked example and a do-not-emit case?
- Would the discovery copy fire on a different job?
- Are the same rules copied until they drift?
- Are maintainer notes sitting on a buyer page?
- Does the procedure say write X while the named tool destroys X?
- Does the procedure say when to stop?
Leave sentence polish. Cheerleading and unfinished plans are a different review.
Cold-eye is a readiness pass. It is not a code audit, a security audit, or a test run, unless those actions are the subject's own claimed checks. Detangler finds structural tangles after edits. Smell Check reviews prose register. Misemphasis reviews likely readings. A failed contract is not a preference about wording or layout.
Built by Catalyst Forge LLC. MIT.