grimoire

A book of tests, and the arguments against each one.

Half of published psychology papers using significance tests carry a p-value that does not match its own test statistic; one in eight carries an error large enough to move the conclusion. You do not need the raw data to find those. A mean, a standard deviation and a sample size can already contradict one another.

The hard part is not finding contradictions. It is that every one of these checks flags honest papers for documented, innocent reasons — a different rounding convention, a multi-item scale, a corrected p-value, a sample that shrank after exclusions. So nothing here reaches you until the tool has tried, mechanically, to destroy it.

1 · The text

Paste a results section, a table, or a single sentence. Nothing leaves this page — there is no network call in this program, which matters when the manuscript is under review.

EXAMPLES

2 · What was read, and what was not

The best-known tool in this family fails silently: results whose formatting deviates slightly are never seen, and the output says nothing about them. A hit count without a coverage figure is as unreadable as a survey without a response rate.

3 · The findings, with the case against each

Surviving and dissolved findings share one stream on purpose. Two sections would teach the eye to read the first and skip the second, and the second is where most of the work is. The row of marks under each finding is every alternative reading that was tried.

4 · Before any of this leaves your machine

Export is refused until all four are ticked, and the four travel inside the export so that whoever receives it can see what the sender was told.

5 · The tool, measured against itself

Accuracy is a claim until something measures it. These cases have known answers, and the two rates below are the ones that matter — the second is not usually published.