The brief sounded simple: put an AI assistant on an existing Umbraco 13 site that catalogs healthcare technology vendors, and have it answer questions from that catalog. Thousands of vendor records, organized into categories, already indexed in ElasticSearch for the site's own search.
The constraint that shaped everything was the data. The assistant had to answer from the catalog and nothing else, with no general knowledge about vendors filling the gaps, and two people asking the same question had to get the same answer. That ruled out the quick version (send the question to a model with a big prompt). It also pushed me away from the heavy version. A RAG pipeline with a vector database would have meant a second copy of the data, an embedding job to keep in sync, and new infrastructure for a site that already had a good search index.
What I built instead: custom Model Context Protocol (MCP) tools running inside the Umbraco app, sitting on the existing ElasticSearch index, called by a single n8n agent. This post covers how the tools are put together, what went into the system prompt, and why the multi-agent design I started with lost.
What the Assistant Had to Do
The business defined four use cases, and every design decision traced back to them:
- Discovery. A user describes a need in plain language, and the assistant finds the matching categories and vendors.
- Lookup. Questions about a specific vendor: what it does, what it works with, and a link to its profile.
- Recommendation. A ranked shortlist for a need, with the reasoning shown.
- Comparison. One vendor against its closest competitors, attribute by attribute.
The MCP Server Lives Inside the Umbraco App
The MCP server is a .NET 8 project that plugs into the existing Umbraco 13 solution, so the tools get the same services, configuration, and search client the website already uses. The SDK was still in preview at the time:
<!-- Key MCP dependencies. The ModelContextProtocol C# SDK was pre-1.0 when
this was built and shipped breaking changes between previews. Check its
NuGet page for the current version before starting a new project. -->
<PackageReference Include="ModelContextProtocol.AspNetCore" Version="0.3.0-preview.3" />
<PackageReference Include="Umbraco.Cms.Core" Version="13.*" />
<PackageReference Include="System.Text.Json" Version="9.*" />
Registration sits behind a config flag, so the server can be switched on per environment:
public static IServiceCollection AddMcpServices(
this IServiceCollection services,
IConfiguration configuration,
bool enableMcp = false)
{
// Register MCP configuration from appsettings
services.Configure<McpConfiguration>(/* configuration section */);
if (enableMcp)
{
// Configure MCP server with timeouts and HTTP transport
services.AddMcpServer(options => {
// Set initialization timeout, request scoping
// Configure server info and version
})
.WithHttpTransport(transportOptions => {
// Set idle timeout and session management
})
.WithToolsFromAssembly(); // Auto-discover tools
}
// Register domain-specific services
services.AddScoped<IMcpServerService, McpServerService>();
services.AddScoped<IMcpUtilityService, McpUtilityService>();
return services;
}WithToolsFromAssembly() picks up every method marked as an MCP tool, so adding a tool means adding a class. The /mcp endpoint sits behind a small middleware that checks an API key header before a request reaches any tool:
public class McpAuthenticationMiddleware
{
public async Task InvokeAsync(HttpContext context, RequestDelegate next)
{
if (context.Request.Path.StartsWithSegments("/mcp")
&& !await ValidateApiKey(context))
{
context.Response.StatusCode = 401;
return;
}
await next(context);
}
// ValidateApiKey: read the key from a request header, compare it to the
// configured keys, and log the attempt.
}Three Tools, Built on Search Instead of Vectors
Each tool is a static method marked with [McpServerTool], with a [Description] on the method and on every parameter. Those descriptions are what the agent reads to decide which tool to call, so they get the same care as the code. Because the methods are static, they resolve services from a fresh scope on each call instead of holding state, which keeps them safe to call concurrently.
Category search
ElasticSearch does the first pass because it's fast and indexed: text match, status filters, pagination. Business rules that don't translate into a search query run afterward in C#. Then the result is reshaped before it goes back to the agent: consistent field names, real types instead of strings, and the URLs a user would click.
// Trimmed skeleton of the category search tool
[McpServerTool]
[Description("Search and filter categories based on various criteria")]
public static async Task<string> SearchCategories(
[Description("Text search across names and descriptions")] string? searchTerm,
[Description("Filter by status, etc.")] string? filters,
[Description("Page size, 1-100")] int limit,
[Description("Results to skip")] int offset)
{
using var scope = Services.CreateScope();
var searchService = scope.ServiceProvider.GetRequiredService<ISearchService>();
limit = Math.Clamp(limit, 1, 100);
offset = Math.Max(0, offset);
searchTerm = NormalizeInput(searchTerm); // n8n sends "null" as a string
// First pass in ElasticSearch: indexed text match, filters, paging
var searchResults = await searchService.ExecuteSearch(
BuildSearchRequest(searchTerm, filters, limit, offset));
// Second pass in C#: business rules that don't fit a search query,
// then reshape each hit into the fields the agent needs
var response = ApplyBusinessFilters(searchResults, filters)
.Select(ToAgentShape);
return JsonSerializer.Serialize(response, JsonOptions);
}Vendor comparison
The comparison tool finds a vendor's closest competitors by scoring candidates on four weighted factors: category overlap (40%), capability similarity (30%), market status (15%), and overall market presence (15%). It pulls up to 500 vendors from ElasticSearch, scores them, and keeps the ones above a threshold.
The part I'd repeat on any project: a model doesn't write the comparison. The tool builds the comparison matrix and the insights from templates, and the agent's job is to present them. The numbers come out the same every time, and the model never gets a chance to invent a feature a vendor doesn't have.
Matching a plain-language question
The third tool maps a free-text question to categories without embeddings. It runs one weighted ElasticSearch query that blends direct keyword matching (40%), ElasticSearch's more-like-this similarity (30%), capability matching (20%), and market context (10%). Each match comes back with a confidence score and the reasons it matched, and a minimum score drops the weak ones so the agent sees fewer, better candidates.
For a catalog with well-written descriptions, that got close enough to semantic search without a second data store to keep in sync.
Handling What n8n Actually Sends
Two details mattered in practice:
- n8n sometimes passes the literal string
"null"for an empty optional parameter. Every tool turns"null", empty, and whitespace-only values into a real null before using them. - Agents send out-of-range values. Page sizes get clamped to 1 through 100, offsets can't go below zero, and search terms are capped at 500 characters.
When something fails, the tool returns structured JSON instead of throwing: what went wrong, a suggestion the agent can act on, and a retry hint. That gives the agent something to try next instead of a dead end.
public static async Task<string> SearchVendors(/* parameters */)
{
var stopwatch = Stopwatch.StartNew();
try
{
return await ExecuteSearch(searchRequest);
}
catch (ElasticsearchClientException)
{
return CreateErrorResponse("Search service error", "Try reducing filters", retryAfterSeconds: 30);
}
catch (TimeoutException)
{
return CreateErrorResponse("Search timeout", "Try more specific terms", retryAfterSeconds: 60);
}
finally
{
LogPerformanceMetrics(stopwatch.Elapsed);
}
}One Agent Beat Four
I started with a multi-agent design: one specialized agent per use case, with a router in front. After building and testing both versions, the single agent won on every measure I cared about:
- Context. Every handoff between agents dropped some of the conversation's context.
- Cost. Delegation burned tokens on agents briefing other agents.
- Debugging. A wrong answer meant tracing it across several agents instead of one.
- Speed. One agent calls the tools directly and answers right away.
What made one agent workable was the system prompt. It carries all four use cases, and it's strict about where answers can come from. Here's its structure with the domain details stripped out:
# Role
Assistant for [domain]. Answer only from data the tools return.
# Rules
1. Data sources: tool output only. If the tools don't return it, say so.
2. Output format: one link format and one response structure for every answer.
3. Consistency: the same question gets the same answer.
4. Interaction limits: a fixed number of clarifying questions per conversation.
5. Quality checks: specific validation rules for each use case.
# Use cases
## Discovery
Triggers: the user describes a need or problem.
Process: match to categories, then rank vendors by relevance.
Output: structured results with the reasoning shown.
## Lookup, Recommendation, Comparison
Same trigger / process / output shape for each.On the n8n side the setup is small: custom session keys so each conversation keeps its own context, a 30-message memory window to keep token use in check, and the MCP tools connected straight to the AI Agent node, so there's no custom HTTP plumbing between the agent and the server.
When splitting agents does pay off
There's one case where I'd still reach for more than one agent: cost. n8n's AI Agent tool lets you nest one agent inside another as a tool, all within a single workflow. That makes it easy to run a premium-tier model on the parent agent for analysis and synthesis, and hand routine sub-tasks (pulling data, calling APIs, simple summaries) to a fast-tier model that costs a fraction as much per token.
That trade works when the sub-tasks are independent and their output is easy to check, and you should check what the cheaper model returns before anything acts on it. For this assistant it didn't pay. Every use case needed the full conversation context, and delegation overhead was one of the reasons the multi-agent version cost more.
Performance
ElasticSearch requests ask only for the fields a tool needs and apply filters in the query instead of afterward wherever possible:
var searchRequest = new VendorSearchRequest
{
Index = elasticConfiguration.Value.VendorsIndex,
SearchPhrase = searchTerm,
Skip = offset,
Take = limit,
// Only the fields the tool returns
SourceIncludes = new[] { "name", "description", "status", "categories", "imageUrl" },
// Filter at query time, not after
Filters = new VendorFilters
{
Statuses = statusList,
CategoryIds = categoryGuids
}
};Category lookups go through a 15-minute in-memory cache:
public async Task<CategoryModel?> GetCategoryByKeyAsync(Guid categoryKey)
{
var cacheKey = $"category_{categoryKey}";
if (_cache.TryGetValue(cacheKey, out CategoryModel? cachedCategory))
return cachedCategory;
var category = await FetchCategoryFromElasticSearch(categoryKey);
if (category != null)
_cache.Set(cacheKey, category, TimeSpan.FromMinutes(15));
return category;
}Complex questions came back in about 2 to 3 seconds end to end.
What I'd Keep, and What I'd Change
The patterns held up:
- Start with one tool and one agent, and add more only when a real question needs it.
- Keep each tool to one job. The search tool searches, and the comparison tool compares.
- Return structured data with consistent field names, plus the metadata (confidence scores, match reasons) the agent needs to explain its answer.
- Let search and templates do the work that has to be repeatable, and let the model do the talking.
The strict data rules did their job. Answers stayed grounded in the catalog instead of the model's general knowledge, and the same question got the same answer. Building on the existing search index also made it faster to ship than a custom RAG pipeline, with one set of logs to read when something went wrong.
What I'd change is the protocol layer. The MCP wiring in this post targets the spec from before July 2026, so a new build should start from the current spec and SDK. The tool design should carry over as is.
