Overview
This chapter explores a two-step approach for fine-tuning generative LLMs:- Supervised Fine-Tuning (SFT): Teach the model to follow instructions
- Direct Preference Optimization (DPO): Align outputs with human preferences
Use a GPU with at least 15GB VRAM. In Google Colab, select Runtime > Change runtime type > Hardware accelerator > GPU > GPU type > T4.
When to Fine-Tune Generation Models
Fine-tune generative models when:- You need specific output formats or styles
- Domain-specific knowledge is required
- You want to align with specific preferences or guidelines
- Base models don’t follow instructions well enough
- You have quality instruction-response pairs
Step 1: Supervised Fine-Tuning (SFT)
Data Preprocessing
Format data using the model’s chat template:Model Setup with QLoRA
LoRA Configuration
Training Configuration
SFT Training
Step 2: Preference Tuning (DPO)
DPO Dataset Preparation
Load SFT Model
DPO Training Configuration
DPO loss starts higher than SFT and decreases more gradually as the model learns preferences.
Merge Adapters
Inference
Two-Step Training Comparison
1
Supervised Fine-Tuning
Purpose: Learn instruction following and task completionData: Instruction-response pairsLoss: Cross-entropy on next token predictionResult: Model can follow instructions but may not align with preferences
2
Direct Preference Optimization
Purpose: Align with human preferences and quality standardsData: Prompt with chosen/rejected response pairsLoss: DPO loss encouraging chosen over rejectedResult: Model generates preferred outputs matching human judgment
QLoRA Benefits
Hyperparameters Guide
SFT Hyperparameters
DPO Hyperparameters
Best Practices
1
Data Quality
- Use diverse, high-quality instruction data
- Ensure chat template consistency
- Filter low-quality responses
- Balance different task types
2
QLoRA Configuration
3
Training Stability
- Monitor loss curves for smooth descent
- Use gradient checkpointing for memory
- Enable fp16 mixed precision
- Start with lower learning rates if unstable
4
Evaluation
- Test on held-out examples
- Compare base vs SFT vs DPO outputs
- Evaluate instruction following quality
- Check for catastrophic forgetting
Memory Optimization
For Limited VRAM
For Faster Training
Training Time Estimates
SFT (3,000 examples on T4):Next Steps
- Experiment with larger models (7B, 13B parameters)
- Try different LoRA ranks (8, 16, 32, 64)
- Implement PPO for more complex preference learning
- Use RLHF for human-in-the-loop refinement
- Evaluate with LM Eval Harness or MT-Bench
