Overview
Building production LLM applications requires more than just good prompts. This chapter covers:- Chains: Composing multiple LLM calls into pipelines
- Memory: Maintaining context across conversations
- Agents: Building autonomous systems that use tools
Setting Up
To run the examples in this chapter, you’ll need a GPU. In Google Colab, go to Runtime > Change runtime type > Hardware accelerator > GPU > GPU type > T4.
Loading the LLM with LangChain
Chains
Chains allow you to compose multiple operations together, creating reusable pipelines for LLM interactions.Basic Chain with Prompt Template
1
Create Prompt Template
2
Chain with LLM
3
Invoke Chain
Multiple Chains: Building a Story Generator
Chain multiple LLM calls together to create complex workflows:Memory
By default, LLMs are stateless—they don’t remember previous interactions. Memory systems solve this problem.The Problem: No Memory
ConversationBufferMemory
Store all conversation history:1
Update Prompt Template
2
Add Memory
3
Test Memory
ConversationBufferWindowMemory
Retain only recent conversations to limit context size:Window memory is useful for managing context length while maintaining recent conversation history. Adjust
k based on your needs.ConversationSummaryMemory
Summarize conversation history to save tokens:1
Create Summary Prompt
2
Add Summary Memory
3
Test Summarization
Agents
Agents can autonomously decide which tools to use and in what order, enabling them to solve complex tasks that require multiple steps and external information.Setting Up an Agent
1
Configure OpenAI
2
Create ReAct Prompt Template
3
Prepare Tools
4
Create Agent Executor
Running an Agent
Ask the agent a complex question requiring multiple tools:The agent autonomously decided to:
- Search the web for MacBook Pro prices
- Use the calculator tool to perform currency conversion
- Combine results into a final answer
Comparison of Techniques
- Chains
- Memory
- Agents
When to use:
- Predictable, sequential workflows
- Multiple steps that always execute in the same order
- Building complex outputs from simpler components
- Story generation (title → character → story)
- Document processing pipelines
- Multi-step transformations
Best Practices
1
Start with Chains
Begin with simple chains for predictable workflows before adding complexity.
2
Choose Appropriate Memory
- Use Buffer for short conversations
- Use Window for medium-length interactions
- Use Summary for long-running conversations
3
Design Clear Tool Descriptions
Agents rely on tool descriptions to make decisions. Make them clear and specific.
4
Monitor Agent Behavior
Use
verbose=True during development to understand agent decision-making.5
Handle Errors Gracefully
Set
handle_parsing_errors=True for agents to manage unexpected outputs.6
Test Incrementally
Build and test each component (prompt, chain, memory, tool) independently before combining.
Key Takeaways
- Chains compose multiple LLM calls into reusable pipelines with prompt templates
- Memory systems enable stateful conversations:
- Buffer memory stores full history
- Window memory retains recent interactions
- Summary memory compresses long conversations
- Agents autonomously select and use tools based on ReAct prompting
- LangChain provides a unified framework for building these advanced patterns
- Choose the right tool for your use case: chains for predictable workflows, memory for conversations, agents for dynamic problem-solving
Next Steps
Now that you understand chains, memory, and agents, you can:- Build multi-step LLM applications with chains
- Create conversational interfaces with memory
- Develop autonomous systems with agents and tools
- Combine these techniques for sophisticated LLM-powered applications
