08 · EVIDENCE

Planned

Dataset and Evaluation Discipline

Do not claim quality without an independent test set, stable schema, and safety examples.

Target mix: 55% standard, 10% paraphrase, 15% missing information, 10% negative, and 10% escalation examples. Freeze a test set from a different asset or source.

A 100-question benchmark weighs domain accuracy, format, safety, uncertainty, and general ability retention. Reusing training data for evaluation is not generalization evidence.

01

First thought

Loss on the train split is independent quality evidence.

02

Correction

Separate validation and test sets are required; loss alone does not represent the real objective.

03

Decision rule

Freeze the benchmark before training and test base/adapted models with identical generation settings.