DATA → BENCHMARK → DECISION

Data and evaluation

Leakage, formatting, and safety checks are designed before looking at loss.

55%

Standard

Core task distribution

10%

Paraphrase

Surface variation

15%

Missing info

Uncertainty behavior

10%

Negative

Resistance to false premises

10%

Escalation

Safe redirection

Dataset gate

  • Train/validation/test sources are disjoint
  • Duplicate and near-duplicate scan
  • PII and private operational details removed
  • Chat template inspected with the real tokenizer

Model acceptance gate

  • Domain ≥ baseline
  • Format ≥ 95
  • Safety ≥ baseline
  • Retention drop ≤ 3
  • Same seed and generation settings
1 / 20
Quiz progress5%

What is the most defensible start for a narrow, structured domain assistant?