You don’t need two databases for hybrid search

· engineering

Keyword search and vector search can run in one engine over one copy of your data. Infino is an open source retrieval library that keeps the text, the embeddings, and both the BM25 and vector indexes in one table of Parquet files on object storage, and a single hybrid query returns one ranked result set. There is no second database to feed and no sync job between them.

Most teams arrive at the same setup by stages.

They start with embeddings and a vector database because semantic search feels like the new part. Then the misses show up. Error codes disappear. Part numbers blur together. Proper names drift. So they add keyword search beside it, and now retrieval works well enough.

It also means they are running two search systems for one job.

That part gets waved away too easily. People spend their time on fusion weights and ranking tricks because that work is visible. The expensive part sits underneath it. Two services to run. Two copies of the corpus. Two indexing paths. One sync loop that nobody wants to own and everybody depends on.

Hybrid search is not a luxury feature. It is what you end up building once you learn where each method breaks.

Embeddings are good at meaning. They are weak at exact strings. BM25 is good at exact terms and bad at semantic drift. If your corpus contains manuals, tickets, logs, or internal docs, you usually need both. We ran both retrievers over 10,003 real banking questions in where keyword search and vector search each fail, and what is hybrid search covers the fusion math.

The older assumption was that this meant two different systems. That assumption came from the period when vector search had to be bolted onto the stack from the outside. It no longer needs to be.

What gets expensive in the usual stack

Look at the operational shape of the common design.

You keep one system for vectors and another for full-text. You write every document into both. You keep two schemas in step. You watch two dashboards. You debug two failure modes. When one side lags, retrieval quality drops in ways that are hard to spot and annoying to explain.

The cost model is just as awkward. Both systems charge you for stored corpus size. The sync path costs you on every change. Query traffic matters less than people expect. Double the corpus and you often double the retrieval bill, even if traffic stays flat.

The ranking code is not the real burden here. Most teams can fuse two result lists in an afternoon. The part that lingers is everything around it.

Open source hybrid search on object storage

Infino takes the obvious route and keeps one copy. It is an open source retrieval library under Apache-2.0, embedded in your process rather than run as a cluster.

The table holds the text, the embeddings, and the indexes for both BM25 and vector search. A hybrid query runs over the same rows and returns one ranked result set. There is no second database to feed and no merge step in the application. That matters less because it saves a few milliseconds and more because it removes an entire layer of operational drag.

There is another problem teams feel later. Once the index format is proprietary, the data becomes sticky. Leaving means re-indexing everything, moving the corpus, and often rebuilding the serving path around it. The migration cost grows with time.

Infino cuts against that from both directions. The engine is open source, so nothing in the read path is a black box, and the storage format is one you already use. A superfile is a search index that is a valid Parquet file: the data and the BM25 and vector indexes in one file (core concepts). Those files sit in your own bucket, and any Parquet reader opens one as an ordinary table.

Storage is priced like object storage, not like hot memory reserved for the life of the corpus.

Readers do not need to hold the full dataset. They scale with query load rather than raw corpus size. A top-k lookup turns into a small set of offsets and range reads, with local caching where it helps.

This is the right shape for a lot of RAG systems. The corpus grows steadily. Queries are small. Most documents sit cold most of the time.

What it looks like in code

retrieval.py
import infino
import pyarrow as pa

db = infino.connect("s3://my-bucket/docs")

schema = pa.schema([
    pa.field("body", pa.large_utf8(), nullable=False),
    pa.field("embedding", pa.list_(pa.float32(), 384), nullable=False),
])

docs = db.create_table(
    "docs", schema,
    infino.IndexSpec().fts("body").vector("embedding", 384, "cosine"),
)

docs.append(rows)

hits = docs.hybrid_search(
    "body", "E-4021 after rotating a key",
    "embedding", query_vec,
    k=10,
    projection=["_id", "body", "score"],
)

The text and the embeddings live on the same rows. BM25 and vector search run in one place. The caller gets back one ranked set of Arrow rows, with projection choosing the columns (it returns _id and score alone if you leave it off). The search guide has every retrieval mode with runnable code, including how RRF fuses the two rankings.

The environment switch is small:

connect.py
infino.connect("memory://")              # tests
infino.connect("./data")                 # laptop, no server
infino.connect("s3://my-bucket/docs")    # your bucket, queried in place

Infino does not generate embeddings for you. Bring your own model, write the vectors you want, and query with the vectors you already trust.

If you use a reranker, keep using it. Retrieval and reranking are different jobs. Pull back a wider candidate set, rerank it with a cross-encoder, and pass the short list to the model. Nothing about a single-engine hybrid design changes that.

This is not a claim that every search workload belongs on object storage.

If you need an in-memory serving tier with every knob exposed and corpus size is not the problem, there are good tools for that. If your pain comes from running two systems for one retrieval path, the trade-off looks different. Then the question is less about peak benchmark numbers and more about how much machinery you are carrying just to issue a search.

That is the case Infino is built for.

One table. One copy of the corpus. Hybrid retrieval over your own files. SQL in the same process when you need it.

You can try it locally with a folder and move the same data to object storage later.

install.sh
pip install infino

The engine is on GitHub at infino-ai/infino.

Common questions

Do I need a separate vector database for hybrid search?

No. The vector index and the BM25 index can live in the same table as the text. Infino writes both into the same Parquet files and hybrid_search queries them together, so there is no second system to feed or keep in sync. The search guide.

What does running a vector database next to a search cluster cost?

Two copies of the corpus, two indexing paths, two schemas kept in step, and a sync job between them. Both systems usually charge by stored corpus size, so doubling the corpus tends to double the retrieval bill even when query traffic stays flat.

Where does the data live?

In your own bucket as Parquet files, or on local disk. The same code runs against memory:// for tests, a local folder, or an s3:// prefix, and any Parquet reader can open the files as an ordinary table. Parquet interop.

Does Infino generate embeddings?

No. Bring your own model, write the vectors with the rows, and query with a vector from the same model.

Is Infino open source?

Yes. The engine is Apache-2.0 licensed and runs as a library inside your process rather than as a cluster. infino-ai/infino on GitHub.