Psychology
Psychology research, evaluated after the replication crisis
How psychological judgment shows outside a journal's supplementary materials, and the standard the founding cohort will hold it to. Everything below is a draft in public, on purpose.
What counts as evidence of psychology research skill
Psychology reformed its evidence culture earlier and harder than most fields: preregistration, registered reports, open data and materials, and large multi-lab replication projects are now routine. A preregistration honored in the final write-up, or an analysis script that reproduces every reported number from raw data, is stronger evidence of skill than the venue the result eventually appeared in.
Outside academia, the same judgment runs UX research, clinical outcome measurement, and people-analytics teams. The researcher who designed a usability study that actually isolated the variable it claimed to, or who ran a clinical measure validation that a product team could trust, is doing psychology research whether or not it is called that.
The hardest thing to see from a publication list is what did not make it in: a well-powered study that found nothing, a manipulation check that failed and was reported anyway. Evaluators read for exactly that candor, because it is where the field's credibility was rebuilt.
Review work is evidence too. Registered-report reviews, replication attempts of a specific claim, and public critiques that identified a p-hacked analysis all sample the same reading skill this community is built to score — and they are visible without anyone's permission.
The psychology evaluation rubric, first draft
Written in the shadow of the replication crisis, this rubric treats a preregistered null as stronger evidence of skill than a flexible positive.
- Measurement validity
- Instruments and manipulations measure the construct they claim to, with evidence for it, and the construct does not quietly drift between method and conclusion.
- Analytic discipline
- Analyses are pre-specified or transparently exploratory, robustness is shown across reasonable forks, and the garden of forking paths is closed, not hidden.
- Sample-to-claim fit
- Power, sampling frame, and population match what the conclusion claims, and limitations are priced into the interpretation rather than parked in a footnote.
- Replication transparency
- Data, materials, and code are shared in a form that lets someone else actually check the result, and prior failed or null attempts are disclosed rather than filed away.
Founding psychology evaluators will pressure-test this draft first — including against their own past work.
What founding psychology evaluators will do
First, take the rubric apart: where it rewards preregistration theater over real rigor, where it punishes exploratory work that says so honestly, where a dimension cannot actually be scored from a real artifact.
Then run calibration rounds on public studies — preregistrations checked against outcomes, replication targets, applied evaluations — scoring independently and comparing spreads, so the first published scores carry a known uncertainty.
As the community opens, those calibration records will seed the weighting system: the first psychologists whose evaluation history is itself part of the instrument.
Who this is for
The founding cohort is looking for psychologists whose judgment is already in daily use, wherever they happen to practice it:
- Quantitative and clinical researchers who read the preregistration before the abstract.
- UX researchers and people-analytics scientists whose best studies never reach a journal.
- Replicators and meta-scientists who have done the field's least rewarded, most informative work.
- Early-career researchers trained on open-science norms who want that training to count.
Who this is not for
Self-selection matters more than any filter we could write, so here is the honest version:
- Anyone after a quick credential — the founding stage produces standards, not badges.
- Researchers looking to promote their own findings; evaluation here is of other people's artifacts, under a published rubric.
- Anyone uncomfortable having their evaluation accuracy tracked — the calibration record is the point of the design.
- Anyone who needs a live scoring platform today; the mechanics on this page are in design, and the tense is deliberate.
Apply to evaluate psychology
Psychology is pre-selected on the application. Link an OSF page, ORCID profile, or repository we can read.