Biology

Biology research, evaluated by biologists

From bench protocols to preprints, biological evidence has a shape of its own. This page describes what founding evaluators will look for in it, and the draft rubric they will hold it against.

What counts as evidence of biology skill

Biology's public record has widened fast — bioRxiv turned preprinting from a physics habit into a life-science norm, protocols are shared on dedicated platforms, and datasets land in repositories like GEO and the PDB. A methods section written to be repeated, a deposited dataset with usable metadata, or a preprint whose figures survive a careful read are all direct evidence of skill.

Industry biology produces artifacts academia rarely sees: assay validation reports, target-triage documents, tech-transfer packages that make a protocol work in a second lab. The scientist who designed the dose-response study at a biotech, or made a CRO's readout trustworthy, has research judgment worth evaluating — usually with nothing citable to show for it.

The field's chronic failure mode is the gap between what an experiment shows and what its abstract claims. Reading for that gap — controls, replicates, effect sizes, the difference between correlation and perturbation — is a learnable, scoreable skill, and it is the one this community most needs from its biologists.

Negative and confirmatory results carry evaluation weight here that journals rarely grant them. A well-powered replication that failed, a validation study showing a reagent does not do what its datasheet claims, or a preprint documenting why a published effect would not reproduce are exactly the artifacts a careful evaluator learns most from — and producing them is some of the strongest evidence of judgment the field offers.

The biology evaluation rubric, first draft

This rubric leans hard on the distance between data and claim, because preclinical biology's replication record makes that distance the field's central evaluation problem.

Experimental design
Controls, randomization, blinding, and sample sizes chosen before the experiment rather than after it. The design would convince a skeptic who has seen the failure modes.
Claim discipline
Conclusions stay within what the data support. Mechanistic language is earned by perturbation experiments, not borrowed from correlation.
Method transparency
Protocols detailed enough to rerun, reagents and cell lines identified, and data deposited where the field can reach them.
Quantitative rigor
Statistics match the design, effect sizes are reported alongside significance, and there are no signatures of outcome-dependent analysis.

Founding biology evaluators will revise this draft against real preprints before it is used to score anyone.

What founding biology evaluators will do

Begin by stress-testing the rubric against the artifacts biologists actually produce — a blot-heavy preprint reads differently from a genomics pipeline, and the standard has to survive both.

Run calibration rounds on public work: independent scoring of the same preprints and datasets, then structured comparison of where trained readers disagree and why.

Carry a per-discipline calibration record into the open community, seeding the weighting system with biologists whose judgment has a measured track record.

Who this is for

The cohort needs biologists who already read other people's work the hard way, across the bench and the pipeline:

  • Bench scientists, in academia or industry, who read methods sections the way referees are supposed to.
  • Computational biologists whose pipelines and deposited datasets are their real publication record.
  • Biotech and CRO scientists whose best work lives in validation reports no journal will ever see.
  • Early-career researchers who learned rigorous design in the shadow of the replication crisis and want that skill to count.

Who this is not for

Just as real, and worth reading before you apply:

  • Anyone hoping to accredit their own research program — evaluators will score other people's artifacts here.
  • Researchers who want a credential faster than a standard can be built honestly.
  • Anyone unwilling to have their scoring accuracy measured and visible; calibration is the community's anchor.
  • Anyone who needs the platform to be live today — evaluation mechanics are in design, and this page says so on purpose.

Apply to evaluate biology

Biology is pre-selected on the application. Link an ORCID profile, bioRxiv listing, or repository we can read.

VocaidDeep