Evaluation registry
Move from development comparison to one-shot evidence without breaking the blind.
First register all eight traditional baselines on one frozen development split. Then commit truth before model submission, freeze nine prediction sets, reveal once, and let an independent identity trigger server-side recomputation.
Scientific boundary
Every evidence class keeps a different job.
Development metrics diagnose. Permanent-blind evidence tests generalization. Neither can silently become activation authority, and every stored result is recomputed from current custody bytes.
Registered matrices
0
immutable Program records
Currently reverified
0
archive bytes checked at read time
Frozen datasets
0
content-addressed package identities
Activation eligible
0
requires independent blind evidence
No baseline matrix is loaded for this Program.
Start with an authorized, frozen dataset package. Run the CPU baseline evaluator without changing its split, package the complete output, then verify and register it here. Metrics are intentionally absent until real evidence exists.