12 WEEKS

Roadmap

Every week ends with a deliverable and gate. Source state and your personal marks stay separate.

01
7H · COMPLETED

Training Foundations

Source100%

Build a precise decision vocabulary by separating model, data, and training layers.

Tasks
  • Base, Instruct, and Reasoning
  • Tokens, context, and attention
  • LoRA/QLoRA
  • Loss and generalization
Deliverable / gate

Concept check and seven-day review plan

02
6H · ACTIVE

Setup and Verification

Source45%

Verify a reproducible Studio environment on CachyOS that genuinely uses the GPU.

Tasks
  • NVIDIA/CUDA environment table
  • 4-bit path
  • GPU inference evidence
  • Turkish token measurement
Deliverable / gate

Working Studio, verified GPU, token measurement

03
6H · TODO

Studio Workflow

Source0%

Learn every main Studio surface from model discovery through adapter and export.

Tasks
  • Model and dataset
  • Data Recipes QA
  • Checkpoint
  • Export options
Deliverable / gate

End-to-end Studio dry run

04
7H · TODO

First Controlled Training

Source0%

Prove the complete training pipeline without aiming for model quality.

Tasks
  • 50–100 steps
  • Peak VRAM
  • Adapter reload
  • 10 fixed prompts
Deliverable / gate

0.8B smoke test

05
8H · TODO

Dataset Engineering

Source0%

Prepare the first dataset version with correct train, validation, and independent test splits.

Tasks
  • Duplicates/leakage
  • Schema consistency
  • Anonymization
  • Chat template
Deliverable / gate

Dataset v1 and independent test set

06
8H · TODO

First Real QLoRA

Source0%

Produce the first measurable domain adapter on a 4B Instruct model.

Tasks
  • 4-bit baseline
  • Validation loss
  • Base/adapted/export
  • Duration and tokens/s
Deliverable / gate

4B domain adapter

07
8H · TODO

Controlled Experiments

Source0%

Change only one variable per run to establish causal evidence.

Tasks
  • Rank 8/16
  • LR 1e-4/2e-4
  • Epoch 1/3
  • Context 1024/2048
Deliverable / gate

Parameter comparison report

08
8H · TODO

Evaluation System

Source0%

Measure quality independently of loss with a repeatable benchmark.

Tasks
  • Domain
  • Format
  • Safety
  • Retention
Deliverable / gate

100-question benchmark and scorecard

09
8H · TODO

Scaling to 9B and 14B

Source0%

Measure how model scale changes quality, speed, and VRAM on the same data and benchmark.

Tasks
  • Micro batch 1
  • OOM sequence
  • Tokens/s
  • System RAM bottleneck
Deliverable / gate

4B/9B/14B scaling report

10
10H · TODO

Industrial Capstone

Source0%

Design a safe, structured, and measurable Condition Monitoring adapter.

Tasks
  • 2,000–5,000 examples
  • Missing information
  • Safety note
  • JSON schema
Deliverable / gate

Industrial Condition Monitoring Adapter

11
8H · TODO

GRPO and Advanced Training

Source0%

Run an automatically verifiable reward experiment only after the SFT pipeline is reliable.

Tasks
  • Reward hacking
  • KL divergence
  • Length bias
  • SFT/GRPO comparison
Deliverable / gate

Small-model GRPO experiment

12
8H · TODO

Export and Local Use

Source0%

Compare adapters, merged models, and GGUF formats on the same benchmark.

Tasks
  • Safetensors
  • GGUF Q4/Q5/Q8
  • OpenAI-compatible API
  • Ollama/LM Studio
Deliverable / gate

Adapter, GGUF, and API comparison