Hybrid RAG: Combining Dense and Sparse Retrieval
A linear, one-concept-per-lesson path through Hybrid RAG, dense (embedding) retrieval combined with sparse (keyword) retrieval, from hand-rolled BM25 to a complete FastAPI hybrid retrieval service. 26 lessons · 3 tiers.
Prerequisites: Comfortable writing basic Python. Assumes you've done this site's Naive RAG course (or otherwise know what dense retrieval, cosine similarity, and embeddings are) - Lesson 2 here is a compressed recap, not a full re-teach.
Lessons use Google's Gemini free tier (gemini-embedding-001 for embeddings, gemini-3.5-flash-lite for chat), the same models as this site's other RAG courses. No Docker, no database, no extra account - GOOGLE_API_KEY in a .env file is the only setup.
Built on rank_bm25, the BM25 library this course graduates to in the Advanced tier.

Beginner
Sparse retrieval by hand, then fusing it with dense using Reciprocal Rank Fusion.
Intermediate
Making the fusion robust: filters, persistence, and knowing which retriever to trust.
- 10Normalizing Scores Before FusionGitHub
- 11Tuning RRF's kGitHub
- 12Metadata-Aware Hybrid RetrievalGitHub
- 13Persisting Both IndexesGitHub
- 14Where Sparse WinsGitHub
- 15Where Dense WinsGitHub
- 16Failure Modes of Hybrid RetrievalGitHub
- 17Minimal Evaluation: Dense vs. Sparse vs. HybridGitHub
- 18Checkpoint: Persisted, Filterable Hybrid AssistantGitHub
Advanced
Graduating both retrievers to rank_bm25 and chromadb, and a capstone.
- 19Where Hand-Rolled Retrieval Breaks DownGitHub
- 20Introducing rank_bm25GitHub
- 21Repointing at chromadb and rank_bm25GitHub
- 22Alpha-Blended Fusion, RevisitedGitHub
- 23Refactoring Into ingest() and ask()GitHub
- 24Wrapping It as a ServiceGitHub
- 25Capstone: A Complete Hybrid RAG ServiceGitHub
- 26Where Hybrid RAG Hits a WallGitHub