Databricks develops adaptive AI retriever with 2x+ lower latency

There’s a new technical post from Databricks on Adaptive Instructed-Retriever, a retrieval model that adjusts how much search an AI agent performs based on the complexity of a query.

The model combines parallel retrieval with adaptive, multi-step search. It can stop once it has gathered sufficient evidence for a straightforward query, while more complex requests can trigger additional search steps.

Across seven held-out internal and external retrieval benchmarks, Databricks says the model performed comparably to Claude Sonnet 5, GPT-5.6 Luna and DeepSeek-V4-Flash, with average end-to-end latency of 5.8 seconds – more than twice as fast as the comparison models.

Databricks trained the model using reinforcement learning to balance retrieval performance against the cost of additional search. The approach is aimed at AI agents that need to find information across enterprise data, including tables, notebooks, dashboards and documents.

The blog can be found here: https://www.databricks.com/blog/adaptive-instructed-retriever-frontier-quality-search-2x-lower-latency

Leave a Reply

Discover more from The IT Nerd

Subscribe now to keep reading and get access to the full archive.

Continue reading