Skip to main content
To build RAG over company documents with hybrid search, split each document into chunks, store every chunk’s text and embedding on the same row of one table, and retrieve with a single query that runs keyword (BM25) and vector search together and fuses the two rankings. Keyword search finds exact identifiers such as an error code or a policy number, and vector search finds passages that answer the question in other words. Company documents are full of both, which is why either retriever alone misses answers. This guide builds that pipeline end to end over a small folder of policies, a runbook, and a handbook, then keeps it current as documents change. Run the blocks in order in one Python session. Nothing needs an account, a key, or a server.

Where should the documents live?

Wherever they already are. The index is itself a set of Apache Parquet files, so it can sit in the same bucket as the documents: connect to an s3:// URI instead of a local path and nothing else in this guide changes. There is no separate database to load the chunks into and no second service for keyword search. What the index does not do is read PDFs. Text extraction, and OCR for scanned pages, is a step you run first. For a large archive it is the slow step, so run it as its own batch job and index what it produces. The loader below reads Markdown directly and extracts the text layer of a PDF with pypdf.

Prepare the documents

Write a few sample documents to a folder. Point DOCS at your own folder to use your documents instead.
Python
Split each document at its headings, and split any long section at paragraph breaks. A chunk that follows the document’s own structure keeps one topic per chunk, and its heading becomes a label the answer can cite.
Python

Index keyword and vector search together

Create one table whose rows are chunks: the text, the section it came from, the document path, and the embedding. Declare a full-text index on the text and a vector index on the embedding, and both are built as the rows are written.
Python
Two choices here matter for RAG:
  • The English analyzer. Questions are written in plain language, and without stopwords="english" the keyword half matches words like “to”, “a”, and “do”, which appear in every chunk and pull unrelated passages up the ranking. stemmer="english" lets “claims” match “claim”. Identifiers such as ERR-4012 still index as their own terms.
  • The embedding text. Each chunk is embedded with its section heading in front, so a short passage under “Travel” still carries what it is about.
Embed questions with the same model at query time. Infino stores and searches the vectors you give it and does not call an embedding model itself. See Embeddings. hybrid_search runs BM25 and vector search over the same rows and fuses the two ranked lists with reciprocal rank fusion, in one call.
Python
The first question is an exact identifier. An embedding model has no reliable sense of what ERR-4012 means, and BM25 finds it by the term. The second shares no words with the answer (“lost or stolen, report it to security”), and vector search finds it by meaning. One pipeline answers both, which is the case for hybrid retrieval in RAG.

Filter by folder, team, or date in SQL

The same search is a SQL table function, so a question can be restricted to part of the corpus, such as the policies folder, in one statement. A WHERE on hybrid_search is applied to both the keyword and the vector half before they are fused, so the filter ranks within the matching chunks rather than trimming a list that was ranked across everything.
Python
A column you filter on often, such as a team, a customer, or an access group, belongs on the chunk row next to doc_path. That is also how to keep a user’s retrieval inside the documents they are allowed to read. The SQL analytics guide covers joins, grouping, and windows over search results.

Rerank and build the prompt

Fusion decides which chunks are candidates. A cross-encoder reranker then reads the question and each candidate together and orders them by how well the passage answers it, which is slower per chunk and more precise. Retrieve a generous set, rerank it, and keep the top few for the prompt.
Python
Send prompt to the model of your choice. Each source carries its chunk_id, so an answer that cites [policies/expenses.md#0] points back to the exact file and section, and a reader can check it.

Keep the index current

When a document changes, replace its chunks: delete the rows for that path and append the new ones. Updates and deletes need durable storage, a local path or an s3:// URI, which is why this guide connects to ./rag-index rather than memory://.
Python
Every append is a commit that writes new files, so write in batches, a document or a folder at a time rather than a chunk at a time. optimize() is where maintenance happens: it merges small files and rebuilds the table-wide vector index. Nothing runs in the background inside your process, so call it after a batch of changes or on a schedule.

Scaling to a larger archive

  • Separate extraction from indexing. Run text extraction and OCR as a batch job that writes plain text, then index its output. Re-indexing never repeats the OCR.
  • Batch the writes. Append thousands of chunks per call and run optimize() after each large batch.
  • Keep the index next to the data. Connect to s3:// or az:// with a local cache directory, so repeated queries read from disk instead of the bucket. See Connect and storage.
  • Re-embed into a new table when you change models. Vectors from two models are not comparable, so build the new table alongside the old one and switch when it is complete.

See also

Last modified on October 1, 2026