Skip to content
TECH CEO Daily
SaaSLaunch

Databricks Lakebase Search goes GA, putting vector and full-text search inside Postgres

Two new Postgres extensions bring vector, keyword and hybrid search into Databricks' operational database, with usage-based pricing and scale-to-zero.

By · Editor

· 3 min read · Fact-checked

The 60-second brief

  • 1Databricks made Lakebase Search generally available on AWS and Azure on September 28, 2026, after a June beta.
  • 2It adds vector and BM25 keyword search inside Postgres, so hybrid queries run in one SQL statement.
  • 3Databricks' own benchmark claims twice the throughput of the next-best system; buyers should test with their data.

The news

On September 28, 2026, Databricks made Lakebase Search generally available on AWS and Azure. Databricks Lakebase Search adds vector and keyword search inside Lakebase, the company's managed Postgres database, so AI agents can retrieve context without a separate search system.

The product consists of two Postgres extensions. The first, lakebase_vector, runs approximate nearest neighbor search, which finds the stored embeddings (numeric representations of text or images) closest to a query. The second, lakebase_text, adds BM25 full-text search, a standard method for ranking documents by keyword relevance. Developers can combine both in a single SQL query, alongside ordinary filters and joins against application tables.

Databricks first announced Lakebase Search as a beta on June 16, 2026. In its general availability post, the company said the vector index uses hierarchical inverted-file clustering with a compression technique called RaBitQ, keeps durable data in cloud object storage and caches data in memory and on local NVMe drives, so each query reads only the relevant blocks. Databricks said its BM25 engine is faster than Postgres's built-in text search indexes but did not give a figure.

The performance numbers come from Databricks' own testing. The company said that on the VectorDBBench 100M benchmark, lakebase_vector delivered twice the throughput of the next-best system and was four times cheaper than a cloud Postgres vendor using pgvector, the popular open-source vector extension. It reported 71-millisecond latency at the 99th percentile with 97% recall, and a 1.13-second 90th-percentile response for the first query after the database scales to zero.

On pricing, Databricks said customers pay only for storage while the system is idle and pay for query usage rather than data volume. It cited customer Conexiom, which it said cut infrastructure costs by a factor of three and raised throughput fivefold compared with pgvector.

Databricks' documentation says lakebase_vector works with pgvector's vector types, distance operators and query syntax without modification, and requires Postgres 16 or later. It also warns that enabling Lakebase Search restarts all computes in a project, dropping active connections, and cannot be turned off once enabled.

The numbers

Throughput on VectorDBBench 100M (Databricks test)
2x the next-best system
Cost vs. a cloud Postgres vendor using pgvector (Databricks claim)
4x cheaper
P99 latency at 97% recall (100M vectors)
71 ms
P90 first query after scale-to-zero
1.13 seconds
Conexiom results vs. pgvector (per Databricks)
3x lower infrastructure cost, 5x throughput
Availability
GA on AWS and Azure, September 28, 2026

Why CEOs should care

For data and platform leaders, the pitch is consolidation. Many teams building AI agents run an operational database, a vector store and a keyword search cluster, with pipelines copying data among them. Databricks is arguing that one Postgres database can do all three. Before retiring anything, run your own tests on your data: compare relevance quality for hybrid queries, latency under burst load and index freshness against your current setup. The headline benchmarks are Databricks' own and have not been independently reproduced.

For CFOs, the pricing model matters as much as the speed. Paying for queries instead of provisioned capacity can cut costs for spiky agent traffic, and Databricks' June beta post put object storage at about $20 per terabyte per month versus about $3,000 for memory-resident indexes. But usage-based bills are harder to forecast when a single agent workflow can fire thousands of retrieval requests. The posts and documentation we reviewed did not list per-query rates, so ask for them, along with spending caps and alerts.

For CISOs and architects, fewer systems means fewer copies of sensitive data and fewer access-control models to audit. The trade-off is deeper dependence on one vendor: the extensions are Databricks-specific, even though pgvector-compatible syntax eases migration in. Operations teams should plan a maintenance window, because enabling the feature restarts computes and cannot be reversed.

The bigger picture

Databricks introduced Lakebase in 2025 as a serverless Postgres architecture for transactional workloads, and it is now adding the retrieval features agents need to the same database. Unite.AI reported that pgvector is the most-installed extension on Lakebase Postgres, which explains why Databricks framed its benchmarks against it. That positions Lakebase Search as a direct alternative to standalone vector databases and search engines whose main selling point is retrieval at scale.

The broader pattern is that operational databases are absorbing search. Agents read and write in tight loops, retrieving context, acting and storing memory, and Databricks says that burst pattern strains architectures built from several separate systems.

What’s next

Watch for independent benchmark results, published per-query pricing and whether Databricks extends Lakebase Search beyond AWS and Azure. Customers already running pgvector on Lakebase are the obvious early testers, since the documentation says their existing queries should work unchanged.

What “Fact-checked” means

Fact-checking means testing a story’s facts against the evidence before it is published. This story went through at least two separate checks before this version was published.

What we checked
Its names, figures, dates, job titles, quotes and who said what were checked against the story’s sources, including its main source where it could be opened. The headline was checked for accuracy and overstatement.
How
A first check reviewed the whole story. If it passed, a second, skeptical check went back to the sources to look for mistakes in the most important facts. If a check flagged the story, it was edited to fix the problems found, and a separate re-check then reviewed the whole story again.
Who
The checks are made by our newsroom, as steps kept separate from the writing, under rules set by our editor, . A story the checks still flag is held for the editor, who decides whether it is fixed, published or dropped.
If something is wrong
“Fact-checked” does not mean error-free. If a material error is found after publication, we correct the story and add a note saying what changed. Report an error

How we fact-check →

Companies in this story

DatabricksPostgresVector searchAI agents

Earlier coverage of Databricks

All Databricks coverage →

Written by

Editor · Technology & Business Writer

Hussein is a writer and business technology enthusiast focused on the intersection of technology, entrepreneurship, finance, artificial intelligence, and digital innovation.

CoversAICybersecurityBig TechSaaSStartupsFintech

About this story. Researched from primary sources whenever they are available and fact-checked before publication.

Published by Tech CEO Daily, an independent publication. Masthead · Editorial standards

Follow Tech CEO Daily on Facebook for the day’s top stories in your feed.

Free newsletters

The technology briefing for people running businesses.

Daily, weekly, bi-weekly or monthly. You choose.

How often

The Daily Brief · Monday to Saturday, 7 a.m. ET

Free forever. One click to unsubscribe. We never sell your email.