Core workflowNON_DECISION
Candidate Benchmark & Preflight
Reuse persisted fixed-model evaluations, add candidates to the shared benchmark, or check readiness before scientific review.
Benchmark scores are versioned and persistent. Board rank remains a comparison artifact, not a synthesis decision.
Choose assessment
Persistent benchmark, ad hoc replay, and readiness checks remain separate scientific tasks.
Program · program-draft
verified
Materializing the benchmark board
Unseen historical structures are scored once under the fixed scorer version. Later visits reuse the persisted artifacts.