Amir TabatabaeiWork with me
← All work

Learning project · Sole Author · Apr 2025 — Apr 2025

Smart PDF Assistant

Where the retrieval quietly stopped answering

Study project, never deployed

GeminiPyMuPDFChromaDBPythonRAG

What mattered

Retrieval was capped at the top three of fifteen documents, which is where the answers stopped.

System map

EXTRACTFEW-SHOTSCHEMAEMBEDQUERYCONTEXTGROUNDEDPDFPyMuPDF textGeminiField recordsChromaTop 3 of 15Answer

Context

A document pipeline over invoices, US tax returns and property contracts: PyMuPDF pulls the text, a few-shot prompt classifies the document, a schema prompt extracts its fields as JSON, everything is embedded into a Chroma collection, and questions are answered from what the retriever returns. Built as the capstone for Google's Gen AI Intensive course.

The fifteen documents are synthetic, because real invoices and tax returns cannot be published, and generating them meant every document also had a ground-truth record to check against. That detail matters later.

The extraction half works: given the invoice above, the model returns every field correctly, line items included.

The retriever silently truncated the answer

The notebook's showcased question is "List the property buyers and their sale prices." It answers with three buyers. There are five.

Karen White's $550,186 — the second-largest sale in the corpus — is missing, and so is Maria Lopez's. The model was not wrong; it answered faithfully from what it was handed. ask_question_rag defaults to k=3:

results = collection.query(query_texts=[question], n_results=k)
context = "\n\n".join(results["documents"][0])

This is the failure mode that makes RAG demos dangerous. "List all" and "which is highest" are aggregation queries, and a retriever ranking by relevance has no way to report that it came up short. The answer arrives fluent, well formatted, and missing 40% of the data, with nothing in the pipeline raising a warning.

The fix is not a larger k — that just moves the cliff. It is routing aggregation queries at a different layer: a metadata filter on type == "property_contract", or a query over the extracted records, which the pipeline already produces and then never uses again.

An embedding model that was constructed and never called

collection = chroma_client.create_collection(name="documents")
embedding_fn = GoogleGenerativeAiEmbeddingFunction(api_key=api_key, ...)
collection.add(documents=[doc["text"]], ids=[...], metadatas=[...])

create_collection is never passed embedding_function=, so Chroma falls back to its own default sentence-transformer and every retrieval in the notebook ran on that. The write-up claims Gemini embeddings. Nothing errored, and the results looked plausible, so it went unnoticed for a year.

The check that was one loop away

The retry predicate is scoped correctly: 429 and 503 retried, everything else raised, rather than a blanket catch. But the extraction result is printed, not parsed:

try:
    response = model.generate_content(prompt)
    print(response.text.strip())
except Exception as e:
    print("Quota hit or error:", str(e))

No json.loads, no validation against the schema, no comparison to the ground-truth record that exists for every one of the fifteen documents. So "JSON mode" is a prompt convention here rather than an enforced contract, and a field accuracy number was a single loop away. The bare except also reports every failure, malformed JSON included, as "Quota hit".

predict_label(doc) has the same shape of problem: it takes a doc parameter and never reads it, using the loop's prompt closure instead. It works because the loop rebuilds prompt before each call, but it is a retry-decorated function reading mutable state from outside itself.

Result

The extraction half works; the retrieval half answered confidently and incompletely, and nothing in the pipeline could tell the difference. That missing check is the same gap I found in the signature project: a pipeline that runs, with nothing measuring whether it is right. Never deployed, and the numbers it originally reported should not be trusted.

There is also a write-up on Medium covering the build as it happened.

Screens

An invoice PDF beside the JSON the model returned from it: invoice number, dates, client name and address, two line items with quantity and unit price, subtotal, tax and total — every field matching the ground-truth record
Five property contracts listed by sale price. Three are marked retrieved and appear in the answer; two, including the second-largest sale at $550,186, are marked never retrieved and missing from the answer