Computer science

Computer science research, evaluated by peers

No field publishes more of its evidence in runnable form. This page describes how that evidence will be read here, and the rubric the founding cohort will sharpen before anyone is scored.

What counts as evidence of computer science skill

Computer science is the one field where the artifact can be executed. Code that reproduces the paper's table, benchmarks with honest baselines, an ablation that shows which component actually carries the gain — these are checkable claims, and the field has already normalized checking them through artifact-evaluation tracks at its major conferences.

Outside papers entirely, sustained open-source maintainership is a research artifact in its own right: design decisions defended in public, performance claims tested by strangers, regressions owned. So is a production system whose design notes explain what was traded away and why. Both are evaluable evidence, and neither shows up in citation counts.

The skill hardest to fake is problem selection: work aimed at a real question rather than shaped to fit an existing benchmark. Evaluators will read for it explicitly — it is where strong researchers separate from strong paper-writers.

Theory and systems earn their evidence differently — a proof, a measured artifact — but both leave public trails: technical reports, dissertations, talks whose slides show the failure cases as well as the wins. Review work counts too. Program-committee service, artifact-evaluation reviews, and public reproduction attempts are direct samples of the judgment this community is built to measure, and they are readable by anyone.

The computer science evaluation rubric, first draft

Computer science already runs artifact evaluation for single papers; this rubric extends the habit from the artifact to the researcher behind it.

Baseline honesty
Comparisons use strong, tuned baselines and report them fairly. A win over a strawman is scored as what it is.
Ablation discipline
The work isolates which component produces the claimed improvement, and negative or neutral ablations are reported rather than trimmed.
Artifact quality
The code runs, the environment is specified, and the headline result reproduces within stated tolerance from what is released.
Problem framing
The question matters beyond the benchmark that measures it, and its framing does not smuggle in the conclusion.

Founding evaluators in computer science will run this draft against public repositories and released artifacts before it is used in earnest.

What founding computer science evaluators will do

Break the rubric on hard cases first: a systems paper without a benchmark, a theory contribution with no artifact, a model release with weights but no training code.

Run calibration rounds on public artifacts — repositories, papers with released code, benchmark claims — scoring independently and studying the disagreements.

Establish calibration records that will weight computer-science evaluations when the platform opens, so accuracy earns influence before reputation does.

Who this is for

The cohort needs computer scientists who run the code before they cite the paper:

  • Researchers and engineers who reproduce results before they believe them.
  • Open-source maintainers whose review judgment has been sharpened by years of pull requests.
  • Industry researchers whose strongest work ships as systems and internal reports rather than papers.
  • PhD students and postdocs who have served on artifact-evaluation committees and want the habit to matter more widely.

Who this is not for

The other side of the filter, stated plainly:

  • Anyone seeking a badge for their own repositories — evaluation here reads other people's work.
  • Leaderboard chasers; the rubric is designed to price benchmark-shaped work accurately, which cuts both ways.
  • Anyone who prefers their evaluation accuracy untracked — visible calibration is the core mechanism.
  • Anyone who needs a live platform this quarter rather than a standard built correctly first.

Apply to evaluate computer science

Computer science is pre-selected on the application. Link a repository, an ORCID or Scholar profile, or a published artifact.

VocaidDeep