Learning project · Sole Author · Apr 2025 — Apr 2025
Smart PDF Assistant
Where the retrieval quietly stopped answering
Study project, never deployed
What mattered
Retrieval was capped at the top three of fifteen documents, which is where the answers stopped.
System map
Context
A document pipeline over invoices, US tax returns and property contracts: PyMuPDF pulls the text, a few-shot prompt classifies the document, a schema prompt extracts its fields as JSON, everything is embedded into a Chroma collection, and questions are answered from what the retriever returns. Built as the capstone for Google's Gen AI Intensive course.
The fifteen documents are synthetic, because real invoices and tax returns cannot be published, and generating them meant every document also had a ground-truth record to check against. That detail matters later.
The extraction half works: given the invoice above, the model returns every field correctly, line items included.
The retriever silently truncated the answer
The notebook's showcased question is "List the property buyers and their sale prices." It answers with three buyers. There are five.
Karen White's $550,186 — the second-largest sale in the corpus — is missing, and
so is Maria Lopez's. The model was not wrong; it answered faithfully from what
it was handed. ask_question_rag defaults to k=3:
results = collection.query(query_texts=[question], n_results=k)
context = "\n\n".join(results["documents"][0])
This is the failure mode that makes RAG demos dangerous. "List all" and "which is highest" are aggregation queries, and a retriever ranking by relevance has no way to report that it came up short. The answer arrives fluent, well formatted, and missing 40% of the data, with nothing in the pipeline raising a warning.
The fix is not a larger k — that just moves the cliff. It is routing
aggregation queries at a different layer: a metadata filter on
type == "property_contract", or a query over the extracted records, which the
pipeline already produces and then never uses again.
An embedding model that was constructed and never called
collection = chroma_client.create_collection(name="documents")
embedding_fn = GoogleGenerativeAiEmbeddingFunction(api_key=api_key, ...)
collection.add(documents=[doc["text"]], ids=[...], metadatas=[...])
create_collection is never passed embedding_function=, so Chroma falls back
to its own default sentence-transformer and every retrieval in the notebook ran
on that. The write-up claims Gemini embeddings. Nothing errored, and the results
looked plausible, so it went unnoticed for a year.
The check that was one loop away
The retry predicate is scoped correctly: 429 and 503 retried,
everything else raised, rather than a blanket catch. But the extraction result
is printed, not parsed:
try:
response = model.generate_content(prompt)
print(response.text.strip())
except Exception as e:
print("Quota hit or error:", str(e))
No json.loads, no validation against the schema, no comparison to the
ground-truth record that exists for every one of the fifteen documents. So "JSON
mode" is a prompt convention here rather than an enforced contract, and a field
accuracy number was a single loop away. The bare except also reports every
failure, malformed JSON included, as "Quota hit".
predict_label(doc) has the same shape of problem: it takes a doc parameter
and never reads it, using the loop's prompt closure instead. It works because
the loop rebuilds prompt before each call, but it is a retry-decorated
function reading mutable state from outside itself.
Result
The extraction half works; the retrieval half answered confidently and incompletely, and nothing in the pipeline could tell the difference. That missing check is the same gap I found in the signature project: a pipeline that runs, with nothing measuring whether it is right. Never deployed, and the numbers it originally reported should not be trusted.
There is also a write-up on Medium covering the build as it happened.
Screens

