Skip to main content
Open In Colab

Overview

Text classification is one of the most fundamental tasks in Natural Language Processing. This chapter explores how to classify text using both representation models (encoder-based) and generative models (decoder and encoder-decoder models). You’ll learn multiple approaches ranging from zero-shot classification to fine-tuned models, and understand when to use each technique.

What You’ll Learn

1

Representation Models for Classification

Learn how to use pre-trained models like RoBERTa and sentence embeddings for classification tasks
2

Embedding-based Approaches

Discover how to leverage embeddings with traditional ML classifiers and zero-shot techniques
3

Generative Models for Classification

Use encoder-decoder models like FLAN-T5 and ChatGPT for text classification
4

Performance Comparison

Compare different approaches and understand their trade-offs

Use Cases

Text classification powers numerous real-world applications:
  • Sentiment Analysis: Classify movie reviews, product feedback, or social media posts as positive or negative
  • Content Moderation: Automatically detect toxic, spam, or inappropriate content
  • Customer Support: Route support tickets to the appropriate department
  • Document Organization: Categorize emails, news articles, or research papers
  • Intent Detection: Classify user queries in chatbots and virtual assistants

Dataset: Rotten Tomatoes Movie Reviews

Throughout this chapter, we use the Rotten Tomatoes dataset, which contains movie reviews labeled as positive (1) or negative (0).
Output:
Example reviews:

Text Classification with Representation Models

Representation models (encoder-based models like BERT, RoBERTa) excel at understanding and encoding the meaning of text into numerical vectors.

Approach 1: Using a Task-Specific Model

The simplest approach is to use a model that’s already fine-tuned for sentiment analysis.
This model is based on RoBERTa and has been fine-tuned specifically for sentiment analysis on Twitter data. It can classify text into negative, neutral, and positive categories.
Run inference on the test set:
Evaluate performance:
Results:
The model achieves 80% accuracy on movie reviews, despite being trained on Twitter data!

Approach 2: Supervised Classification with Embeddings

Instead of using a task-specific classifier, we can:
  1. Convert text to embeddings using a sentence transformer
  2. Train a traditional ML classifier on these embeddings
The embeddings have shape (8530, 768) - each text is represented as a 768-dimensional vector. Train a logistic regression classifier:
Results:
This achieves 85% accuracy - better than the task-specific model!
Alternative Approach: Instead of using a classifier, you can average the embeddings per class and use cosine similarity:
This achieves 84% accuracy without training any classifier!

Approach 3: Zero-Shot Classification

Zero-shot classification doesn’t require any training data - you just provide label descriptions!
Results:
Achieves 78% accuracy with zero training! The label descriptions you choose matter significantly.
Experiment with different descriptions: Try using "A very negative movie review" and "A very positive movie review" to see how results change!

Classification with Generative Models

Generative models can also perform classification by generating the class label as text.

Encoder-Decoder Models (FLAN-T5)

FLAN-T5 is a text-to-text model that can follow instructions and generate responses.
Prepare the data with a prompt:
Run inference:
Results:
FLAN-T5-small achieves 84% accuracy with simple prompting!

ChatGPT for Classification

Large language models like ChatGPT can perform classification through conversational prompting.
Create a structured prompt:
Run on the entire test set (requires API credits):
Results:
ChatGPT achieves 91% accuracy - the best performance of all methods!

Performance Comparison

Practical Applications

Zero-Shot Classification
  • Limited or no training data available
  • Need quick prototyping
  • Working with many different categories
  • Label definitions are clear and distinguishable
Task-Specific Models
  • Domain matches your use case
  • Need fast, consistent performance
  • Have computational constraints
  • Don’t have resources to train custom models
Embeddings + Classifier
  • Have sufficient labeled training data (hundreds to thousands of examples)
  • Need good balance of accuracy and speed
  • Want to use lightweight traditional ML models
  • Need model interpretability
Generative Models (FLAN-T5, ChatGPT)
  • Need highest possible accuracy
  • Have complex classification tasks with nuanced categories
  • Can afford API costs or computation time
  • Working with evolving categories or requirements

Key Takeaways

  1. Multiple paths to classification: Representation models, embeddings, and generative models all offer viable approaches
  2. Zero-shot is powerful: Modern embeddings enable decent classification without any training
  3. Trade-offs matter: Balance accuracy, speed, cost, and training requirements for your use case
  4. Prompt engineering helps: For generative models, well-crafted prompts significantly impact performance
  5. Embeddings are versatile: Sentence embeddings can power multiple approaches (supervised, zero-shot, similarity-based)

Next Steps

In Chapter 5, we’ll explore Text Clustering and Topic Modeling, where you’ll learn to discover patterns and topics in unlabeled text collections.
Try the notebook yourself: Open In Colab