Skip to main content
Open In Colab

Overview

While text classification requires labeled data, text clustering and topic modeling discover patterns and themes in unlabeled document collections. This chapter explores how to group similar documents together and extract meaningful topics using modern embedding-based approaches. You’ll learn how to build a complete clustering pipeline using embeddings, dimensionality reduction, and clustering algorithms, then extend it to topic modeling with BERTopic - a modular framework that combines the best of classical and modern NLP techniques.

What You’ll Learn

1

Document Embeddings

Convert text into high-quality vector representations using sentence transformers
2

Dimensionality Reduction

Use UMAP to reduce embedding dimensions while preserving semantic structure
3

Clustering Algorithms

Apply HDBSCAN to discover document clusters of varying densities
4

Topic Modeling with BERTopic

Extract interpretable topics using c-TF-IDF and representation models
5

Visualization & Exploration

Visualize document clusters and topic relationships interactively

Use Cases

Text clustering and topic modeling power numerous applications:
  • Research Analysis: Discover themes in academic papers, patents, or scientific literature
  • Customer Feedback: Identify common themes in product reviews or support tickets
  • Content Organization: Automatically categorize news articles, blog posts, or documents
  • Social Media Monitoring: Track trending topics in tweets, forums, or discussions
  • Document Discovery: Enable exploratory search in large text collections
  • Knowledge Management: Organize and navigate corporate knowledge bases

Dataset: ArXiv NLP Papers

We’ll work with abstracts from ArXiv papers in the Computation and Language (cs.CL) category - real research papers from the NLP community.
This dataset contains 44,949 abstracts from NLP research papers, making it ideal for discovering research themes and trends.

A Common Pipeline for Text Clustering

Text clustering typically follows a three-step pipeline:
1

1. Embed Documents

Convert text into dense vector representations
2

2. Reduce Dimensionality

Reduce high-dimensional embeddings to lower dimensions for clustering
3

3. Cluster Documents

Group similar documents using clustering algorithms

Step 1: Embedding Documents

First, we convert each abstract into a numerical vector that captures its semantic meaning.
Check the dimensions:
Output:
Each of the 44,949 abstracts is now represented as a 384-dimensional vector.
We use thenlper/gte-small, a compact but powerful embedding model. It balances quality and speed, making it suitable for large document collections.

Step 2: Reducing the Dimensionality of Embeddings

High-dimensional embeddings (384 dimensions) can be difficult for clustering algorithms. We use UMAP (Uniform Manifold Approximation and Projection) to reduce dimensions while preserving semantic relationships.
Why 5 dimensions? This strikes a balance:
  • More than 2-3 dimensions preserves more semantic information
  • Fewer than 10-20 dimensions makes clustering more effective
  • We can reduce to 2 dimensions later for visualization

Step 3: Cluster the Reduced Embeddings

Now we apply HDBSCAN (Hierarchical Density-Based Spatial Clustering of Applications with Noise) to discover clusters.
Output:
We discovered 156 distinct clusters in the NLP research literature!
Why HDBSCAN?
  • Doesn’t require specifying the number of clusters upfront
  • Can identify clusters of varying densities
  • Automatically identifies outliers (labeled as -1)
  • Works well with the output of UMAP

Inspecting the Clusters

Let’s manually examine the first three documents in cluster 0:
Output:
All three abstracts are about sign language translation - the clustering worked!

Visualizing Clusters

To visualize our clusters, we reduce the embeddings to 2 dimensions:
Create a static plot:
The visualization shows clear clusters of related papers, with outliers in grey!
Outliers (cluster -1) represent papers that don’t fit well into any cluster. This is normal and often represents either very unique papers or papers that bridge multiple topics.

From Text Clustering to Topic Modeling

While clustering groups similar documents, topic modeling goes further by:
  1. Identifying what makes each cluster unique
  2. Extracting representative keywords
  3. Providing interpretable topic descriptions

BERTopic: A Modular Topic Modeling Framework

BERTopic combines our clustering pipeline with topic representation techniques to create interpretable topics.
Training output:
Modularity is key: BERTopic allows you to swap out any component:
  • Use different embedding models (OpenAI, Cohere, local models)
  • Try different dimensionality reduction techniques
  • Experiment with clustering algorithms
  • Apply various representation models

Exploring Topics

View all topics with their counts and representations:
Sample output: The model discovered 156 topics covering different areas of NLP research!

Examining Topic Keywords

Get the top 10 keywords for a specific topic with their c-TF-IDF weights:
Output:
Topic 0 is clearly about speech recognition and ASR (Automatic Speech Recognition)!
c-TF-IDF (class-based TF-IDF) is like TF-IDF but treats each cluster as a single document. It identifies words that are frequent within a topic but rare across other topics, making topic representations more distinctive.

Searching for Topics

Find topics related to a search term:
Output:
Topic 22 has 95.5% similarity to “topic modeling”! Let’s inspect it:
Output:
This topic includes classic topic modeling terms like LDA (Latent Dirichlet Allocation)! Verify the BERTopic paper is in this topic:
Output:
Perfect! The BERTopic paper itself was correctly assigned to the topic modeling topic.

Visualizations

BERTopic provides rich interactive visualizations to explore topics.

Visualize Documents

Create an interactive plot showing all documents and their topics:
This creates an interactive scatter plot where you can:
  • Hover over points to see document titles
  • Zoom into specific regions
  • Filter by topic
  • Explore the relationship between different topics

Additional Visualizations

The hierarchy visualization is particularly useful - it shows how topics can be merged into higher-level themes, revealing the hierarchical structure of your document collection.

Representation Models

BERTopic’s modularity shines when using representation models to improve topic descriptions beyond c-TF-IDF.

Available Representation Models

Uses embeddings to select the most representative keywords for each topic.
Selects diverse keywords that are relevant but not redundant.
Uses large language models to generate descriptive topic labels.
Integrates with LangChain for custom LLM-based representations.

Updating Topics After Training

You can update topic representations after training, allowing quick iteration:
Combine multiple representation models for richer topic descriptions:
This creates multiple representations for each topic, giving you different perspectives!

Practical Applications

Scenario: A research lab has thousands of papers and wants to organize them by theme.Solution:
  1. Extract paper abstracts
  2. Generate embeddings and cluster using BERTopic
  3. Use OpenAI representation model for human-readable topic names
  4. Create interactive visualizations for exploration
  5. Build a search interface using topic assignments
Benefits: Researchers can quickly find related papers, identify research gaps, and track field evolution.
Scenario: An e-commerce company receives thousands of product reviews daily.Solution:
  1. Collect review text
  2. Apply BERTopic to discover common themes
  3. Track topic prevalence over time
  4. Alert teams when new issues emerge (new topics)
  5. Generate automated reports on customer concerns
Benefits: Product teams can prioritize improvements, customer service can proactively address issues, and executives get data-driven insights.
Scenario: A news website wants to recommend related articles.Solution:
  1. Model topics across all articles
  2. Assign new articles to topics in real-time
  3. Recommend articles from the same or related topics
  4. Use topic hierarchy for broader recommendations
Benefits: Increased engagement, longer session times, and better content discovery.

Key Differences: Clustering vs Topic Modeling

Performance Considerations

Scalability Tips:
  1. Large datasets (>100K documents):
    • Use approximate nearest neighbors for UMAP
    • Consider batching embeddings
    • Use low_memory=True in HDBSCAN
  2. Speed optimization:
    • Use smaller embedding models (e.g., all-MiniLM-L6-v2)
    • Pre-compute and save embeddings
    • Reduce UMAP dimensions to 3-5 instead of 5-10
  3. Quality improvement:
    • Use larger embedding models (e.g., all-mpnet-base-v2)
    • Increase min_cluster_size for more coherent topics
    • Experiment with different UMAP parameters

Key Takeaways

1

Embeddings are foundational

High-quality embeddings are crucial for both clustering and topic modeling
2

Pipeline matters

The combination of embedding → dimensionality reduction → clustering works well across domains
3

BERTopic is modular

Swap components to customize for your specific use case
4

Representation models enhance topics

Use LLMs or specialized models to generate better topic descriptions
5

Visualization aids understanding

Interactive plots help explore and validate discovered topics

Next Steps

Now that you understand text clustering and topic modeling, you can:
  • Apply these techniques to your own document collections
  • Experiment with different embedding models and parameters
  • Build search and recommendation systems using topics
  • Track topic evolution over time in dynamic corpora

Try the notebook yourself: Open In Colab