The last step of the core loop
Lesson 1 named it: Document -> Node -> Index -> QueryEngine. Lesson 4 built the Index and ended on a cliffhanger: an Index stores and embeds, but it has no .query() method, it can't answer anything by itself. This lesson wraps that same index in a QueryEngine, the piece that actually answers questions.
index.as_query_engine() is the one-line way to get one, and calling .query(question) on it runs two steps behind the scenes:
1. Retrieve: embed the question, find the Nodes in the index whose vectors are most similar to it (Lesson 6 isolates this step alone). 2. Synthesize: hand those retrieved Nodes to Settings.llm and ask it to answer the question using only that text.
This is the RAG pattern in full, and it's the same two-step shape lessons/langchain's RAG lessons (27-29) built by hand with a retriever and a chain. LlamaIndex just gives it one call.
The Response object
.query() doesn't return a plain string, it returns a Response object with two attributes worth knowing:
| What it is | |
|---|---|
.response | The synthesized answer text, what you'd show a user |
.source_nodes | The NodeWithScore objects actually handed to the LLM |
Each item in .source_nodes wraps a Node (same object from Lesson 2, carrying .metadata like file_name) with a .score, how similar that Node's embedding was to the question's embedding. Printing .source_nodes is how you'd cite sources in a real app, or debug a wrong answer by seeing exactly what text the LLM was actually shown.
The code, piece by piece
query_engine = index.as_query_engine()Wraps the VectorStoreIndex from Lesson 4 in a QueryEngine. No arguments needed for the default behavior, Lesson 7 shows the response_mode argument that controls how synthesis works.
response = query_engine.query(question)Runs retrieve -> synthesize and returns a Response, not a string.
print(f"A: {response.response}")for node in response.source_nodes: source = Path(node.metadata["file_name"]).name print(f" - {source} (score={node.score:.4f})")Prints the answer, then walks .source_nodes to show which files fed the answer and how confident the retrieval step was about each one.
Checkpoint
index.as_query_engine(): wraps anIndexin aQueryEngine, the piece that can actually answer questions, anIndexalone cannot..query(question)runs retrieve (find similar Nodes) then synthesize (ask the LLM to answer from them), and returns aResponse, not a plain string..response: the synthesized answer text..source_nodes: theNodeWithScoreobjects actually used, each carrying the source Node's.metadataand a similarity.score, useful for citations and debugging.
If anything here still feels unclear, ask before moving to Lesson 6.