Mathematics
Mathematics research, evaluated on the proof
Mathematics is the one field whose artifacts can, in principle, be checked completely. The draft standard below asks how much of that checkability a piece of work actually delivers.
What counts as evidence of mathematics skill
A proof is a complete research artifact on its own: no dataset, apparatus, or replication cohort stands between the reader and the claim. arXiv holds most of the field's working literature, and an evaluator can go arbitrarily deep — from the architecture of an argument down to whether a critical lemma actually holds.
Formalization has added a second kind of evidence. Contributions to proof libraries such as Lean's mathlib are machine-checked mathematics in public, and the judgment involved — stating definitions so theorems remain provable, structuring arguments so they compose — is research skill in a form no referee report can overstate.
Beyond academia, mathematical research runs through cryptography, optimization, and quantitative finance, where a wrong proof has a price. Expository writing counts too: making a hard argument genuinely readable is evidence of the same understanding that produces new ones.
The field's collaborative record is also readable: answers that settled questions on public research forums, referee reports where authors have made them public, errata that fixed real gaps, lecture notes an area actually uses. Each is a sample of mathematical judgment applied to someone else's argument — which is the exact activity this community formalizes and the first thing evaluators will be asked to do.
The mathematics evaluation rubric, first draft
Mathematics can be checked all the way down, so this rubric scores how much of that checkability the work delivers to its reader.
- Correctness and completeness
- The argument holds, gaps are named as gaps, and lemmas are proved or precisely cited — never waved through.
- Verifiable exposition
- The proof is written to be verified by a human reader, not merely believed: definitions are exact, dependencies explicit, notation stable.
- Checkability
- The work is structured so that its critical steps could be machine-checked or independently audited, whether or not a formalization exists yet.
- Problem placement
- The statement's significance is argued — its connection to a program of ideas, what its truth or falsity moves — rather than assumed from difficulty.
Founding mathematics evaluators will sharpen this draft on published proofs and formalized libraries before anyone is scored against it.
What founding mathematics evaluators will do
Probe the rubric with the field's edge cases: a computer-assisted proof, an announcement without full details, a deep expository survey.
Run calibration rounds on public proofs and formalization contributions, scoring independently and dissecting where careful readers differ.
Begin the calibration records that will give mathematical evaluation its weighting when the community opens beyond the founding cohort.
Who this is for
The cohort needs mathematicians who read proofs the way they were meant to be read — all the way down:
- Research mathematicians, in academia or industry, who referee with care and want that judgment on the record.
- Formalizers and proof-assistant contributors whose commits are checkable mathematics.
- Applied mathematicians in cryptography, optimization, and finance whose proofs face consequences.
- Graduate students and postdocs whose depth outruns their publication count.
Who this is not for
With the same precision the field expects elsewhere:
- Anyone seeking validation for their own theorems — evaluation here reads the work of others.
- Mathematicians for whom only their own subfield counts as real mathematics; the rubric must serve the whole field.
- Anyone who would rather not have their evaluation accuracy measured.
- Anyone who needs the scoring platform to exist today — it is being designed, and this page says so plainly.
Apply to evaluate mathematics
Mathematics is pre-selected on the application. Link an arXiv listing, ORCID profile, or formalization work we can read.