Answer any agent question on Parquet.

Scalable search + analytics infrastructure on object storage, for 10x cheaper.

trusted by

built by

The creators of OpenSearch and engineering leaders across LinkedIn, Google, & Amazon.

infino diff --stack

Keep one copy of your data, in your own bucket.

replace elastic search + pinecone vectors + neo4j graph + clickhouse analytics + snowflake warehouse with infino real-time ingest parquet on object storage s3 · gcs · azure blob iceberg · hudi table managers sql · full-text · vector · graph agents any question replace elastic search pinecone vectors neo4j graph clickhouse analytics snowflake warehouse with infino real-time ingest parquet on object storage s3 · gcs · azure blob iceberg · hudi table managers sql · full-text · vector · graph agents any question

architecture →  ·  integration →  ·  guide: hybrid search on Parquet files →

cat query.sql

Infino embeds functions in SQL so agents can express complex questions in a single query.

  1. An agent asks a question

    How manyA count. The agent wants a number, not a pile of log lines. SQL does this with COUNT. errors last nightFilters. Keep the error-level rows and the last 24 hours. mention disk fullKeyword search. Match those exact words in the log body. Today this is a call to a search engine. or look like a capacity incidentSemantic search. Match meaning even when the log never says disk full. Today this is a call to a vector database., and are tied to last week’s outageGraph search. Keep the errors linked to the outage’s record through a key they share, like a host or a service, even when they never mention it. Today this is a call to a graph database.? Which serviceA group on the same hits. SQL GROUP BY. Today a second warehouse query after the searches return. had them, and which team owns itA join. Match each service to its owner in a second table. Today the agent looks this up in another system and stitches the answers together.?

  2. It hides seven jobs

    • countHow many
    • groupWhich service
    • joinwhich team owns it
    • filtererrors last night
    • keyword searchdisk full
    • semantic searchcapacity incident
    • graph searchtied to last week’s outage
  3. One SQL statement runs all seven

    query.sql
    SELECT   o.team, h.service, count(*) AS hits
    FROM     hybrid_search(
               'logs', 'body', 'disk full',
               'embedding', :q, 1000
             ) AS h
    JOIN     graph_walk('logs', 'incidents', [4521], 2, 5000) AS g
               ON g._id = h._id
    JOIN     owners o ON o.service = h.service
    WHERE    h.level = 'error'
      AND    h.ts > now() - interval '24 hours'
    GROUP BY o.team, h.service
    ORDER BY hits DESC;

hybrid search →  ·  graph search →

infino explain accuracy

Combine modes to increase retrieval accuracy.

  1. keyword searcha row’s words
  2. semantic searchits meaning
  3. graph searchits links
  4. 100%retrieval accuracy

graph search →

infino cost

Retrieval is cheap and scalable.

storage $0.30 per GB-month i Stored documents follow monthly writes, since a month of ingest has to land somewhere. Shorter retention lowers both. methodology →
10M docs · tracks writes 10 GB billed
$3
write tokens $0.15 per 1k WT i Write tokens measure ingest work. Batched writes assumed. methodology →
1M docs written · 0.38 WPS avg 0.013 WT each · 1 GB · 13K WT
$2
read tokens $0.05 per 1k RT i Read tokens measure query work. Pinned tables assumed at higher volumes. methodology →
1M queries · 0.38 QPS avg 1.40 RT each · 1.4M RT
$70
returned bytes $0.05 per GB returned
10 docs returned per query 10 GB returned
$1
infino $75/month account minimum
Elasticsearch $788

Elastic Cloud Hosted · Platinum

data nodes2 nodes · 2 zones × 4 GB · $0.09/GB-hr$526
Kibana2 × 2 GB$263
master quorumdata nodes + included 1 GB tiebreaker$0
total$788
methodology →
10.5× more
OpenSearch $1,751

OpenSearch Service

data nodes3 × r6g.2xlarge.search · $0.669/hr$1,465
masters3 × m6g.large.search$280
EBS gp344 GB$5
automated snapshotsincluded · 14-day retention$0
total$1,751
methodology →
23.2× more

Postgres omitted at this scale i One RDS node cannot satisfy the RAM, vCPU, or storage requirement. Standard PostgreSQL does not shard horizontally. methodology →

Pinecone omitted at this scale i No published Dedicated Read Nodes configuration covers this working set (largest published example is 1.4 billion vectors), and on-demand exceeds the 2,000 RU/s index limit. methodology →

Infino Cloud Pricing. methodology →

pricing →

infino bench --latency

Infino is built for latency from the ground up.

Internal External
Workload Cohere, 768 dimensions, 10M docs, top-10
p50
5ms 314ms
p99
12ms 850ms
  • Warm
  • Cold

Internal 10M-document run. Its full results are not published yet. The published engine benchmarks, with a way to reproduce each one, are on the benchmarks page.

Published runs: infino.ai/benchmarks ↗
infino playground

Ask one question and watch keyword, semantic and hybrid search answer it side by side.

ask a question

Ask a question about arXiv abstracts, or pick a suggestion.

this is the retrieval step of an agent, made visible: one call each way, and hybrid decides when keyword and meaning disagree. getting it otherwise takes a search stack, a vector stack, and glue.

query …

three searches run at once

Keyword, semantic and hybrid, each timed end to end over one Parquet table. Click a lane to see the SQL it sent.

via sql table functions idle
encode · …
bm25… vector… hybrid…

times are full round trips to Infino Cloud, network included · the engine alone is faster

appears after the first run

see what each one found

Hover a dot to read the result. Click it to pull in its neighbours.

hover to inspect, click to bloom
dataset: arXiv abstracts · 45,012 rows · arXiv metadata snapshot, CC0 read-only sandbox · get started for free with your own

the full playground →

infino spec

Specs in the cloud or on-prem.

search
full-text · vector · hybrid · graph · NL
index
BM25 (PFOR-delta, FST) · HNSW · OPANN + Sq16
engine
Rust + Small Language Models (SLMs)
language
SQL (Apache DataFusion) · REST · Query DSL · MCP · NL
storage
object storage, S3 · GCS · Azure Blob · on-prem
format
Apache Parquet
deploy
on-prem · self-hosted cloud · hosted cloud
core license
Apache-2.0
security
SOC 2 Type II
encryption
TLS in transit · AES-256 at rest

architecture →

→ index created on s3://your-bucket · apache-2.0

★ star us on github ↗