Skip to content
Praval Technologies

Case study

Retrieval was never the problem. Choosing where to retrieve from was.

Retrieval-augmented generation works well over one corpus and degrades sharply across several. Wrapping each source as a tool and letting an agent choose between them at runtime routes a query to the right system without a single hard-coded rule.

Heterogeneous source types unified
5Heterogeneous source types unifiedAnalytics feeds, a cloud CMS, static PDFs, SQL and SharePoint
Hard-coded routing rules
0Hard-coded routing rules
To add a new data source
1 toolTo add a new data sourceScaling by registration rather than redesign

Sample content: not published

This engagement is sample content. This page is excluded from the sitemap and search indexing until the details are confirmed.

One corpus is a solved problem. An enterprise estate is not

Retrieval-augmented generation improves a model's answers by pulling relevant context out of a vectorised document store. In a single-domain application that is enough. In a real organisation the knowledge a question touches is spread across systems that share no format, no access pattern and no owner: analytics feeds carrying revenue updates, a cloud CMS holding product documentation, static PDFs of legal policy, SQL tables of support tickets and CRM records, and HR and finance policies scattered across SharePoint directories.

Standard architectures are not designed for dynamic source selection. The result is predictable: irrelevant or incomplete results, no domain specificity, and a pipeline that turns brittle the moment a new domain arrives.

  • Nothing decides which source to search. A conventional pipeline searches the store it was pointed at, so answer quality depends on whether the question happened to match that corpus.
  • Hard-coded routing does not survive contact. Classification rules written against today's sources break as soon as a sixth appears. The routing layer becomes the thing nobody wants to touch.
  • Domain logic has nowhere to live. A legal PDF and a CRM record need different handling before retrieval is useful, and a single-corpus design has no natural home for that.
  • Scaling means redesign, not addition. The cost of the sixth source is as high as the cost of the first, which is what stops teams adding it.
  • Retrieval alone cannot plan. Some questions need more than one lookup, in an order that depends on what the first one returned.

Let the model choose the source, from the descriptions alone

The failure is not retrieval quality. It is that nothing in a standard pipeline decides which store should have been searched in the first place, and that decision is the part nobody had automated.

Each indexed source is wrapped as a tool: a modular abstraction over an external function or data source, independently callable, carrying both the logic needed to query that source and a semantic description of when it should be used. The description is not documentation. It is the routing signal, and it deserves the same care as a prompt.

A zero-shot ReAct agent is initialised with a base model and that list of tools, and nothing else. At runtime it interprets the query, selects the relevant tool or tools from their descriptions alone, invokes them, and integrates what comes back with its own reasoning, reasoning and acting in alternation until the answer is complete.

Asked for the total leasable space and capacity at a named site, the agent establishes what is being asked, matches that intent against the tool descriptions, calls only the tool exposing space and capacity metrics, and synthesises the figures with the intermediate steps visible behind the answer. Nothing about that is configured per query.

Three layers that can change without disturbing each other

Ingestion, enrichment and retrieval sit in Azure AI Search. Each dataset is uploaded to Azure Data Lake Storage; a data source connects to that container, and an indexer orchestrates extraction, applies the enrichments defined in a skillset, and pushes the processed output to a search index. The skillset covers language detection, key phrase extraction and embedding generation, and the resulting index supports hybrid keyword and vector search, secured with Azure RBAC.

The tool layer wraps each retrieval channel behind one standard interface, and is not limited to Azure AI Search. The same wrapper fits a vector database, SQL retrieval over a structured database, or a REST API for an external service. The agent layer reasons over those interfaces and decides what to call.

The design is effective, not free, and the costs are worth stating. Multi-step reasoning increases latency, because every hop is another model call before the user sees anything. Agent decisions are non-deterministic, so the same question can take a different route on a different run, which complicates both testing and trust. Tool descriptions carry the routing burden, and two tools with overlapping descriptions is the most common way this goes wrong in practice. Selection logic is opaque without deliberate logging, and there are no built-in confidence metrics to threshold on.

Latency and non-determinism are the price of routing you did not have to write. Whether that is a good trade depends entirely on how many sources you are spanning; this pattern is over-engineered for anything that genuinely lives in one corpus.