Skip to main content

Databricks Brings Vector and BM25 Search to Lakebase Postgres

29 SEPTEMBER 2026·2 MIN READ·1 SOURCE·Official source

Databricks has released two generally available Lakebase Postgres extensions for vector and BM25 search, pairing vendor-reported benchmark gains with a customer deployment spanning more than 100 million rows.

Databricks Brings Vector and BM25 Search to Lakebase Postgres

Key takeaways · 3

  • 01

    Benchmark Lakebase against your own recall and latency requirements rather than relying only on vendor-reported aggregate throughput.

  • 02

    Evaluate database consolidation against the operational complexity of maintaining a separate search engine and ETL pipeline.

  • 03

    Check benchmark topology carefully because Databricks tested pgvector and DiskANN on a single large instance.

Two Search Extensions

Databricks has made two Lakebase Postgres extensions generally available on AWS and Azure: lakebase_vector for scalable approximate-neighbor search and lakebase_text for BM25 full-text search. [1] Databricks says meeting agent search requirements has traditionally meant connecting a standalone search engine to the primary database through an ETL pipeline. [1]

On Databricks' VectorDBBench test using the 100 million-item LAION dataset, lakebase_vector delivered twice the throughput of the next-best system and cost four times less than a cloud Postgres vendor running pgvector, before additional autoscaling savings. [1] Databricks noted that pgvector and DiskANN were tested only on a single large instance. [1]

Performance and Deployment

Databricks reported P99 latency of 71 milliseconds at 97% recall in its tests, meaning the system retrieved the true nearest neighbors 97% of the time. [1] Customer Conexiom ran hybrid search with BM25 across more than 100 million rows using half the compute footprint of its previous pgvector setup, while consolidating OLTP and search workloads in a serverless database that scales to its needs. [1]

What it means

Lakebase Search positions Postgres as both transactional store and retrieval layer, instead of pairing a primary database with a standalone search engine and ETL pipeline. The benchmark comparison is strongest against cloud Postgres using pgvector: Databricks reports twice the next-best throughput and four-times-lower cost, but its pgvector and DiskANN tests used one large instance.

Conexiom's deployment provides a production example for BM25 hybrid search at more than 100 million rows and half its former compute footprint. The practical test is whether similar economics and retrieval quality appear on other workloads. What the sources don't address: whether independent benchmarks across varied datasets, instance configurations, and cloud environments would reproduce Databricks' results.

Lakebase offers AI practitioners a way to keep transactional, vector, and BM25 search workloads in one serverless Postgres system. Teams should still validate vendor-reported latency, recall, cost, and scaling results against their own workloads and infrastructure.

Why it matters
Daily session

Put this to work — one session a day, built for your industry.

Create a free account for a daily session — eight questions and one real-work challenge, on the news that affects your role.

Start free

How this developed

  1. 29 September 2026

    Databricks Brings Vector and BM25 Search to Lakebase Postgres

  2. 29 September 2026

    Event created from source cluster.

Sources

AI fluency, one session a day, built for your work.