Splitting the pipeline in half
Lesson 5's query_engine.query() did two things in one call: retrieve the most relevant Nodes, then ask the LLM to synthesize an answer from them. This lesson isolates the first half. index.as_retriever() returns a Retriever, an object whose only job is finding relevant Nodes, no LLM call involved at all, just the one embedding call needed to turn the question itself into a vector.
This matters for two practical reasons:
- Debugging is cheaper and clearer. If a
QueryEnginegives a bad answer, the first question is always "did it retrieve the right Nodes?" Calling a retriever directly answers that without spending an LLM call, and without the LLM's synthesized wording obscuring what was actually found. - Retrieval is a separate concern from synthesis. You can tune how Nodes are found (how many, what filters, what strategy) independently of how an answer gets written from them. A
QueryEngineis really just aRetrieverplus a response synthesizer wired together.
similarity_top_k
as_retriever(similarity_top_k=2) controls how many Nodes come back, ranked by how similar their embedding is to the question's embedding. This is the same knob a QueryEngine uses internally (Lesson 5's .as_query_engine() used the library default, 2, unstated), made explicit here. A larger similarity_top_k retrieves more context at the cost of a longer, more expensive synthesis step later, this is exactly the setting that makes Lesson 7's response_mode differences start to matter on real datasets.
.retrieve() returns raw evidence
results = retriever.retrieve(question)Returns a plain list of NodeWithScore objects, the same type found inside response.source_nodes in Lesson 5, but this time it's the whole result, not a side channel next to a synthesized answer. Each one carries the Node's full .text, its .metadata (source file), and a .score.
The code, piece by piece
retriever = index.as_retriever(similarity_top_k=2)results = retriever.retrieve(question)Builds a retriever from the same VectorStoreIndex Lesson 4 built, and retrieves the top 2 most similar Nodes to question. No Settings.llm call happens anywhere in this line.
for i, node in enumerate(results): source = Path(node.metadata["file_name"]).name print(f"--- Result {i} (from {source}, score={node.score:.4f}) ---") print(node.text.strip())Prints each retrieved Node's full text alongside its source file and score, exactly what a QueryEngine would have handed to the LLM to synthesize from, visible here directly instead of buried inside response.source_nodes.
Checkpoint
index.as_retriever(similarity_top_k=k): builds aRetriever, retrieval only, no LLM call, no synthesized answer..retrieve(question): returns a list ofNodeWithScoreobjects ranked by similarity, the raw evidence aQueryEnginewould otherwise hand straight to the LLM.similarity_top_k: how many Nodes come back; aQueryEngineuses this same setting internally, just with a default value if you don't set it yourself.- Retrieve vs query: a
Retrieverfinds relevant Nodes; aQueryEngine(Lesson 5) is aRetrieverplus a synthesis step that turns those Nodes into a written answer.
If anything here still feels unclear, ask before moving to Lesson 7.