Skip to main content
Open In Colab

Overview

This chapter explores methods for both training and fine-tuning embedding models. Text embedding models convert text into dense vector representations that capture semantic meaning, enabling similarity search, clustering, and retrieval tasks.
Use a GPU for training embedding models. In Google Colab, select Runtime > Change runtime type > Hardware accelerator > GPU > GPU type > T4.

When to Train Embedding Models

Consider training or fine-tuning embedding models when:
  • You need domain-specific embeddings (medical, legal, technical documentation)
  • Existing models don’t capture your data’s nuances
  • You want to optimize for specific similarity tasks
  • You have labeled pairs or triplets of similar/dissimilar texts

Dataset Preparation

We’ll use the MNLI (Multi-Genre Natural Language Inference) dataset from GLUE, which contains premise-hypothesis pairs with entailment labels.
Example data point:

Training from Scratch

Model Initialization

Start with a base BERT model without sentence-transformers weights:

Loss Functions

The choice of loss function depends on your data format and task requirements.

Softmax Loss

Best for classification-style training with labeled categories:

Cosine Similarity Loss

For training with similarity scores between sentence pairs:
Training results:

Multiple Negatives Ranking Loss

Most effective for retrieval tasks with anchor-positive-negative triplets:
Training results:

Training Configuration

Evaluation Setup

Fine-Tuning Existing Models

Starting from Pre-trained Model

Fine-tuning results:

Evaluation with MTEB

Massive Text Embedding Benchmark (MTEB) provides standardized evaluation:
Results:

Best Practices

1

Choose the Right Loss Function

  • SoftmaxLoss: Classification tasks with categories
  • CosineSimilarityLoss: Pairwise similarity learning
  • MultipleNegativesRankingLoss: Retrieval and semantic search (recommended)
2

Prepare Quality Data

  • Use domain-specific text pairs
  • Include hard negatives for better discrimination
  • Balance positive and negative examples
3

Optimize Training

  • Start with learning rate 2e-5
  • Use warmup steps (10% of total)
  • Enable fp16 for faster training
  • Monitor evaluation metrics during training
4

Evaluate Thoroughly

  • Test on multiple downstream tasks
  • Use MTEB for standardized benchmarking
  • Compare with baseline models
Clear VRAM between training runs to avoid out-of-memory errors:

Training Metrics Comparison

Next Steps

  • Experiment with different base models (RoBERTa, MPNet)
  • Try domain adaptation with your specific data
  • Implement hard negative mining for better performance
  • Evaluate on multiple MTEB tasks for comprehensive assessment