Amir TabatabaeiWork with me
← All work

Learning project · Sole Author · Jun 2026 — Jun 2026

EDIP

Hybrid retrieval with permissions in the query

Study project, never deployed

FastAPIpgvectorCeleryRedisMinIOHTMX

What mattered

Permission predicates are composed into retrieval queries before chunks reach the model.

System map

EMBEDTSQUERYIN (...)IN (...)TOP 50TOP 501/(k+r)TOP 8QuerypgvectorBM25ACL subqueryRRF mergeRerank + capSSE stream

Context

Semantic search and chat over a company's internal documents, with per-employee permissions. FastAPI on async SQLAlchemy, Postgres 16 with pgvector, Celery workers for parsing, embedding and OCR, MinIO for blobs, HTMX for the UI.

I built it right after the Kaggle capstone, which was a notebook with k=3 and no measurement. This is the same problem class taken seriously.

Permissions belong inside the query

The retrieval path is hybrid — vector top-50 and BM25 top-50, fused, reranked, truncated to eight chunks for the prompt. The question is where permissions enter it.

packages/rbac/can.py opens with the rule I held to:

NEVER filter ACL in Python after fetching — always in SQL.

So both arms take a user_id and inline an AND d.id IN (...) subquery over the user's roles and per-resource ACL entries. The vector top-50 is already that user's top-50.

Retrieving globally and filtering afterwards is broken rather than merely slow. A user with narrow permissions gets a short list, or an empty one, while documents they are allowed to read sit at rank 51. Ranking has to happen inside the permitted set, not be trimmed to it afterwards.

Fusing on rank, not score

Cosine similarity and BM25 scores are not on comparable scales, and normalising one into the other needs a calibration set I did not have. Reciprocal Rank Fusion sidesteps it by throwing the scores away and using position:

def _rrf_score(rank: int, k: int = 60) -> float:
    return 1.0 / (k + rank)

Each chunk's score is the sum of its RRF contributions from both lists, so a chunk ranked well by either method surfaces, and one ranked well by both wins.

Ordering the tail of the pipeline

Three steps that only work in one order: rerank, then diversify, then truncate.

Reranking runs a cross-encoder over the merged list, falling back to RRF order if the model is unavailable. Diversification then caps any single document at two chunks, preserving rank order within each document:

def _diversify(chunks, max_per_doc: int = 2):
    ...

Without that cap one long document fills all eight slots and the answer comes from a single file with citations that look plausible and are all the same source. Capping before reranking would throw away chunks the reranker might have promoted; truncating before diversifying makes the cap a no-op.

Chunks that can be found

A chunk reading "the figure rose 12%" names neither the figure nor the year, so embedding it alone makes it unretrievable. The embedding worker prepends a context line derived from a document-level summary before embedding each chunk. That summary is itself stored as a chunk at chunk_index = -1 and excluded from results by AND c.chunk_index >= 0, so it helps retrieval without ever being quoted back as a source.

A streaming bug worth writing down

The SSE generator opens its own database session rather than the request-scoped one, and I left the reason in the docstring:

FastAPI tears down Depends(yield ...) dependencies (commit + close) as soon as the route handler returns the StreamingResponse object, which happens before this generator body actually executes.

Reusing the request session means every write during the stream lands in a transaction that is never committed — messages appear to save and quietly do not. Nothing raises.

What I would fix before trusting it

The chunker estimates tokens from word count rather than running the embedder's tokenizer — fast and dependency-free, but a "512-token" chunk is 512 estimated tokens, and the estimate drifts on dense or non-Latin text.

And the ACL fragment is assembled with f-strings around user_id and company_id. The user-supplied search filters are correctly parameterised, and those two ids come from the authenticated session, so there is no live injection path — but the most security-sensitive predicate in the system is safe only because every caller passes server-derived UUIDs, and no test enforces that. It should be bound parameters.