Verifying Feishu identity

One session covers every module.

experiment
Core workflowNON_DECISION

Candidate Benchmark & Preflight

Reuse persisted fixed-model evaluations, add candidates to the shared benchmark, or check readiness before scientific review.

Benchmark scores are versioned and persistent. Board rank remains a comparison artifact, not a synthesis decision.

Choose assessment

Persistent benchmark, ad hoc replay, and readiness checks remain separate scientific tasks.

Program · program-draft

verified

Materializing the benchmark board

Unseen historical structures are scored once under the fixed scorer version. Later visits reuse the persisted artifacts.