Physical AI

Embodied AI research, evaluated after it leaves the simulator

How research skill shows when a learned policy has to run on hardware that breaks, and the standard the founding cohort will hold it to. Everything below is a draft in public, on purpose.

What counts as evidence of physical AI research skill

The field's defining problem is that its headline numbers are cheap. A success rate in simulation costs nothing to inflate — change the reset distribution, tune the reward, report the best seed — and the same policy on real hardware may not clear a tenth of it. Evidence of research skill is therefore evidence about transfer: what the policy did on the physical system, over how many trials, under conditions the author did not choose in advance.

Trial counts and failure taxonomies carry more signal than benchmark scores. Twenty real-robot episodes with every failure classified tells an evaluator more than a leaderboard entry with none, because the classification is where judgment lives: whether a miss was perception, control, calibration or a mechanical fault is a claim the author had to earn.

Reproducibility in this discipline is unusually hard and unusually informative. Hardware drifts, calibration ages, and a policy that works on one arm may not work on its twin. Researchers who report cross-platform or cross-session results — including the ones that degraded — are giving an evaluator exactly the artifact the field's incentives discourage.

The hardest skill to see is honest safety and reset accounting. Interventions per hour, resets excluded from the numbers, the episodes cut short because something was about to break: these are the details that separate a demonstration from a result, and they are almost never in the abstract.

The physical AI evaluation rubric, first draft

This rubric reads embodied work the way a skeptical replicator reads a demo video: for what happened off-camera, and for how much of the claim survives leaving the lab it was made in.

Sim-to-real accounting
Simulated and physical results are reported separately and comparably, with the transfer gap stated as a finding rather than smoothed over. Randomization and tuning applied to close the gap are disclosed.
Trial protocol integrity
Trial counts, seeds, reset procedures, human interventions and excluded episodes are reported. Conditions are fixed before evaluation rather than selected after, and the selection rule is stated either way.
Failure characterization
Failures are classified by subsystem and mechanism rather than counted, and the evidence separating perception failures from control, calibration or hardware failures is given.
Embodiment generalization
Claims are bounded by the platforms and conditions actually tested. Transfer across robots, sessions, or environments is demonstrated rather than asserted, and degradation is reported where it occurred.

Founding physical AI evaluators will attack this draft first — including against their own demonstrations and benchmark entries.

What founding physical AI evaluators will do

Take the rubric apart: where it penalizes work that legitimately cannot reach hardware, where it over-credits expensive robot fleets, where a dimension cannot be scored from a paper and a video alone.

Run calibration rounds on public work — open benchmark entries, released policies and datasets, reproduction attempts, conference demonstrations — scoring independently and comparing spreads, so the first published scores carry known uncertainty rather than the field's usual confidence.

Establish what a credible real-robot claim must report, and publish it as a standard the community can hold entries to — including its own.

Who this is for

The founding cohort is looking for researchers whose systems have to survive contact with the world, wherever they work:

  • Robot-learning researchers who report real-hardware results and want the reporting itself judged.
  • Control and perception engineers whose deployment work never becomes a paper.
  • Reproduction and benchmark maintainers who have already tried to rerun other people's claims.
  • Embodied-AI practitioners in industry whose evidence is fleet data rather than a leaderboard.

Who this is not for

Self-selection matters more than any filter we could write, so here is the honest version:

  • Anyone after a credential for its own sake — the founding stage produces standards, not badges.
  • Anyone whose evidence is a demonstration video with no protocol behind it; this rubric is built to ask what the video omits.
  • Anyone uncomfortable having their evaluation accuracy tracked — the calibration record is the point of the design.
  • Anyone who needs a live scoring platform today; the mechanics on this page are in design, and the tense is deliberate.

Apply to evaluate physical AI

Physical AI is pre-selected on the application. Link to work we can read — a real-robot evaluation, a released policy or dataset, a reproduction report, an ORCID or repository profile.

VocaidDeep