Skip to main content
Open In Colab This chapter explores prompt engineering techniques that help you get better results from Large Language Models. You’ll learn how to structure prompts effectively, provide context, and use advanced strategies to improve model outputs.

Overview

Prompt engineering is the practice of designing and optimizing inputs to LLMs to achieve desired outputs. This chapter covers:
  • Basic ingredients of effective prompts
  • Advanced prompt engineering techniques
  • Reasoning strategies for complex tasks
  • Output verification and formatting

Setting Up

To run the examples in this chapter, you’ll need a GPU. In Google Colab, go to Runtime > Change runtime type > Hardware accelerator > GPU > GPU type > T4.

Basic Prompt Engineering

Simple Prompts

The most basic form of prompting involves asking a direct question:

Understanding Chat Templates

Models use specific formatting for chat interactions. You can view the template being applied:
Output:

Temperature and Sampling

Control output randomness with temperature and top_p parameters:

Advanced Prompt Engineering

Complex Prompt Structure

Build comprehensive prompts with multiple components:
Experiment with removing or adding components to see their impact on generated outputs. Each element serves a specific purpose in guiding the model.

In-Context Learning: Few-Shot Prompting

Provide examples to guide the model’s behavior:

Chain Prompting: Breaking Down Complex Tasks

Split complex tasks into smaller, manageable steps:
1

Create Product Name and Slogan

2

Generate Sales Pitch from Product

Chain prompting allows you to maintain quality at each step while building complex outputs incrementally. Each step’s output becomes input for the next.

Reasoning with Generative Models

Chain-of-Thought (CoT) Prompting

Enable better reasoning by showing step-by-step thinking:

Zero-Shot Chain-of-Thought

Trigger reasoning without examples using magic phrases:
The phrase “Let’s think step-by-step” is remarkably effective at triggering reasoning behavior in LLMs without requiring examples.

Tree-of-Thought: Multiple Reasoning Paths

Simulate multiple experts reasoning together:

Output Verification

Structured Output with Examples

Guide the model to produce specific formats:

Grammar: Constrained Sampling

Force valid JSON output using constrained sampling with llama-cpp-python:
Constrained sampling guarantees valid JSON output but may require additional computational overhead. Use it when format compliance is critical.

Best Practices

1

Start Simple

Begin with clear, direct prompts before adding complexity.
2

Add Context Gradually

Include persona, instructions, context, format requirements, and tone as needed.
3

Use Examples

Few-shot prompting is powerful for demonstrating desired behavior and output formats.
4

Break Down Complex Tasks

Use chain prompting to split multi-step problems into manageable pieces.
5

Enable Reasoning

For mathematical or logical problems, use Chain-of-Thought prompting with “Let’s think step-by-step.”
6

Constrain When Necessary

Use constrained sampling or detailed format examples when you need specific output structures.

Key Takeaways

  • Prompt structure matters: Persona, instructions, context, format, audience, and tone all influence outputs
  • Examples are powerful: Few-shot learning can dramatically improve results
  • Chain complex tasks: Break down multi-step problems into sequential prompts
  • Trigger reasoning: Use CoT prompting for mathematical and logical tasks
  • Control output format: Use examples or constrained sampling for structured outputs
  • Experiment iteratively: Test different approaches and refine based on results

Next Steps

Continue to Chapter 7: Advanced Text Generation Techniques to learn about chaining, memory, and agents that extend beyond prompt engineering.