08 · EVIDENCE
PlannedDataset and Evaluation Discipline
Do not claim quality without an independent test set, stable schema, and safety examples.
Target mix: 55% standard, 10% paraphrase, 15% missing information, 10% negative, and 10% escalation examples. Freeze a test set from a different asset or source.
A 100-question benchmark weighs domain accuracy, format, safety, uncertainty, and general ability retention. Reusing training data for evaluation is not generalization evidence.
01
First thought
Loss on the train split is independent quality evidence.
02
Correction
Separate validation and test sets are required; loss alone does not represent the real objective.
03
Decision rule
Freeze the benchmark before training and test base/adapted models with identical generation settings.