Answer any agent question directly on Parquet.

Unified retrieval infrastructure for AI workloads.

trusted by

built by

The creators of OpenSearch and engineering leaders across LinkedIn, Google, & Amazon.

infino diff --stack

Keep one copy of your data, in your own bucket.

replace elastic search + pinecone vectors + clickhouse analytics + snowflake warehouse with infino real-time ingest parquet on object storage s3 · gcs · azure blob iceberg · hudi table managers sql · full-text · vector · hybrid agents any question replace elastic search pinecone vectors clickhouse analytics snowflake warehouse with infino real-time ingest parquet on object storage s3 · gcs · azure blob iceberg · hudi table managers sql · full-text · vector · hybrid agents any question

architecture →  ·  integration →  ·  guide: hybrid search on Parquet files →

cat query.sql

Infino embeds functions in SQL so agents can express complex questions in a single query.

  1. An agent asks a question

    How manyA count. The agent wants a number, not a pile of log lines. SQL does this with COUNT. errors last nightFilters. Keep the error-level rows and the last 24 hours. mention disk fullKeyword search. Match those exact words in the log body. Today this is a call to a search engine. or look like a capacity incidentSemantic search. Match meaning even when the log never says disk full. Today this is a call to a vector database.? Which serviceA group on the same hits. SQL GROUP BY. Today a second warehouse query after the searches return. had them, and which team owns itA join. Match each service to its owner in a second table. Today the agent looks this up in another system and stitches the answers together.?

  2. It hides six jobs

    • countHow many
    • groupWhich service
    • joinwhich team owns it
    • filtererrors last night
    • keyword searchdisk full
    • semantic searchcapacity incident
  3. One SQL statement runs all six

    query.sql
    SELECT   o.team, h.service, count(*) AS hits
    FROM     hybrid_search(
               'logs', 'body', 'disk full',
               'embedding', :q, 1000
             ) AS h
    JOIN     owners o ON o.service = h.service
    WHERE    h.level = 'error'
      AND    h.ts > now() - interval '24 hours'
    GROUP BY o.team, h.service
    ORDER BY hits DESC;

hybrid search →

infino cost

Query your object storage directly to reduce costs.

storage $0.30 per GB-month i Stored documents follow monthly writes, since a month of ingest has to land somewhere. Shorter retention lowers both. methodology →
10M docs · tracks writes 10 GB billed
$3
write tokens $0.15 per 1k WT i Write tokens measure ingest work. Batched writes assumed. methodology →
1M docs written · 0.38 WPS avg 0.013 WT each · 1 GB · 13K WT
$2
read tokens $0.05 per 1k RT i Read tokens measure query work. Pinned tables assumed at higher volumes. methodology →
1M queries · 0.38 QPS avg 1.40 RT each · 1.4M RT
$70
returned bytes $0.05 per GB returned
10 docs returned per query 10 GB returned
$1
infino $75/month account minimum
Elasticsearch $788

Elastic Cloud Hosted · Platinum

data nodes2 nodes · 2 zones × 4 GB · $0.09/GB-hr$526
Kibana2 × 2 GB$263
master quorumdata nodes + included 1 GB tiebreaker$0
total$788
methodology →
10.5× more
OpenSearch $1,751

OpenSearch Service

data nodes3 × r6g.2xlarge.search · $0.669/hr$1,465
masters3 × m6g.large.search$280
EBS gp344 GB$5
automated snapshotsincluded · 14-day retention$0
total$1,751
methodology →
23.2× more

Postgres omitted at this scale i One RDS node cannot satisfy the RAM, vCPU, or storage requirement. Standard PostgreSQL does not shard horizontally. methodology →

Pinecone omitted at this scale i No published Dedicated Read Nodes configuration covers this working set (largest published example is 1.4 billion vectors), and on-demand exceeds the 2,000 RU/s index limit. methodology →

Infino Cloud Pricing. methodology →

pricing →

infino bench --latency

Infino is built for the latency and scale that agents need.

Internal External
Workload Cohere, 768 dimensions, 10M docs, top-10
p50
5ms 314ms
p99
12ms 850ms
  • Warm
  • Cold
infino playground

Ask one question and watch keyword, semantic and hybrid search answer it side by side.

ask a question

Ask a question about arXiv abstracts, or pick a suggestion.

this is the retrieval step of an agent, made visible: one call each way, and hybrid decides when keyword and meaning disagree. getting it otherwise takes a search stack, a vector stack, and glue.

query …

three searches run at once

Keyword, semantic and hybrid, each timed end to end over one Parquet table. Click a lane to see the SQL it sent.

via sql table functions idle
encode · …
bm25… vector… hybrid…

times are full round trips to Infino Cloud, network included · the engine alone is faster

appears after the first run

see what each one found

Hover a dot to read the result. Click it to pull in its neighbours.

hover to inspect, click to bloom
dataset: arXiv abstracts · 45,012 rows · arXiv metadata snapshot, CC0 read-only sandbox · get started for free with your own

the full playground →

infino spec

Specs in the cloud or on-prem.

search
full-text · vector · hybrid · NL
index
BM25 (PFOR-delta, FST) · HNSW · OPANN + Sq16
engine
Rust + Small Language Models (SLMs)
language
SQL (Apache DataFusion) · REST · Query DSL · MCP · NL
storage
object storage, S3 · GCS · Azure Blob · on-prem
format
Apache Parquet
deploy
on-prem · self-hosted cloud · hosted cloud
core license
Apache-2.0
security
SOC 2 Type II
encryption
TLS in transit · AES-256 at rest

architecture →

→ index created on s3://your-bucket · apache-2.0

★ star us on github ↗