Where should the documents live?
Wherever they already are. The index is itself a set of Apache Parquet files, so it can sit in the same bucket as the documents: connect to ans3:// URI instead of a local
path and nothing else in this guide changes. There is no separate database to load the
chunks into and no second service for keyword search.
What the index does not do is read PDFs. Text extraction, and OCR for scanned pages, is a
step you run first. For a large archive it is the slow step, so run it as its own batch
job and index what it produces. The loader below reads Markdown directly and extracts the
text layer of a PDF with pypdf.
Prepare the documents
Write a few sample documents to a folder. PointDOCS at your own folder to use your
documents instead.
Python
Python
Index keyword and vector search together
Create one table whose rows are chunks: the text, the section it came from, the document path, and the embedding. Declare a full-text index on the text and a vector index on the embedding, and both are built as the rows are written.Python
- The English analyzer. Questions are written in plain language, and without
stopwords="english"the keyword half matches words like “to”, “a”, and “do”, which appear in every chunk and pull unrelated passages up the ranking.stemmer="english"lets “claims” match “claim”. Identifiers such asERR-4012still index as their own terms. - The embedding text. Each chunk is embedded with its section heading in front, so a short passage under “Travel” still carries what it is about.
Retrieve with hybrid search
hybrid_search runs BM25 and vector search over the same rows and fuses the two ranked
lists with reciprocal rank fusion, in one call.
Python
ERR-4012 means, and BM25 finds it by the term. The second shares no words with the
answer (“lost or stolen, report it to security”), and vector search finds it by meaning.
One pipeline answers both, which is the case for hybrid retrieval in RAG.
Filter by folder, team, or date in SQL
The same search is a SQL table function, so a question can be restricted to part of the corpus, such as the policies folder, in one statement. AWHERE on hybrid_search is
applied to both the keyword and the vector half before they are fused, so the filter
ranks within the matching chunks rather than trimming a list that was ranked across
everything.
Python
doc_path. That is also how to keep a user’s retrieval inside the
documents they are allowed to read. The
SQL analytics guide covers joins, grouping, and
windows over search results.
Rerank and build the prompt
Fusion decides which chunks are candidates. A cross-encoder reranker then reads the question and each candidate together and orders them by how well the passage answers it, which is slower per chunk and more precise. Retrieve a generous set, rerank it, and keep the top few for the prompt.Python
prompt to the model of your choice. Each source carries its chunk_id, so an answer
that cites [policies/expenses.md#0] points back to the exact file and section, and a
reader can check it.
Keep the index current
When a document changes, replace its chunks: delete the rows for that path and append the new ones. Updates and deletes need durable storage, a local path or ans3:// URI, which
is why this guide connects to ./rag-index rather than memory://.
Python
append is a commit that writes new files, so write in batches, a document or a
folder at a time rather than a chunk at a time. optimize() is where maintenance happens:
it merges small files and rebuilds the table-wide vector index. Nothing runs in the
background inside your process, so call it after a batch of changes or on a schedule.
Scaling to a larger archive
- Separate extraction from indexing. Run text extraction and OCR as a batch job that writes plain text, then index its output. Re-indexing never repeats the OCR.
- Batch the writes. Append thousands of chunks per call and run
optimize()after each large batch. - Keep the index next to the data. Connect to
s3://oraz://with a local cache directory, so repeated queries read from disk instead of the bucket. See Connect and storage. - Re-embed into a new table when you change models. Vectors from two models are not comparable, so build the new table alongside the old one and switch when it is complete.
See also
- Search: every retrieval mode, including how hybrid fusion works
- Implementing hybrid search on Apache Parquet files: keyword, vector, and hybrid retrieval compared on 10,003 real questions
- Integrating SQL analytics with vector and full-text search
- Agent memory: the same table pattern for an agent’s history
