The problem
The retailer sells specialty pet nutrition, where the right product depends on species, age, allergies and health conditions. Customers arrive knowing their animal's problem but not which product solves it. The existing site search matched keywords, so a shopper describing a symptom in everyday language got poor results, and staff answered the same product questions repeatedly through the support desk.
Keyword search could not connect a described symptom to the right product when the product listing used different terminology.
Customers described needs in everyday language, such as a dog with a sensitive stomach and a chicken allergy, which the search engine could not interpret.
Basic product-selection questions were reaching the customer support desk instead of being answered at the point of browsing.
Shoppers had no guided way to narrow a large catalogue down to the few products that genuinely fitted their animal.
Product data changed continuously, so any recommendation layer risked going stale against live prices and stock.
What we designed
We designed a retrieval-augmented system rather than letting a language model answer from its own knowledge. The system searches the live product catalogue first, retrieves the genuine matches, and only then asks the model to write the recommendation. This grounds every answer in real catalogue data and makes it structurally impossible for the assistant to invent a product that does not exist. We presented three architecture options tied to where the client's product data actually lives, because that decision, not the AI, determined cost and complexity.
Shopping assistant
Shopper
Assistant
- Limited-ingredient salmon recipe
- Single-protein lamb and rice
- Sensitive digestion turkey-free blend
What the system did
- 3 products retrieved from the live catalogue
- Answer grounded in retrieved items only
Semantic product index
Understands that a described symptom relates to product attributes even when the wording differs
Retrieval layer
Finds the top matching products from the live catalogue before the model is called
Response generator
Writes a conversational recommendation explaining why each product fits
Product card renderer
Returns images, prices and direct purchase links alongside the answer
Catalogue sync
Webhook-driven updates within seconds, with a scheduled full reconciliation as a safety net
Guardrail layer
Role enforcement, input validation, output filtering, rate limiting and spend caps
How it works
From trigger to result.
How the system works
7 stages
01Shopper question
TriggerTyped into the storefront chat widget.
02Backend
LogicThe widget holds no keys. The backend converts the question to a semantic query.
03Product index
DataAn embedded representation of every catalogue item.
04Retrieval
LogicTop five to ten genuine matches with real prices and attributes.
05Grounded generation
AIThe model only describes products it was handed.
06Product cards
ResultText plus images, prices and purchase links.
07Catalogue sync
LogicWebhooks re-index changes within seconds. Full reconciliation every six hours.
The model never queries the catalogue itself. It only ever sees products the retrieval layer has already verified as real, which is what prevents fabricated recommendations.
Our approach
How we worked.
Setup and data ingestion
Connect to the catalogue, process the full product set and build the semantic index.
Core engine
Build question understanding, retrieval and grounded response generation.
Widget and integration
Build the chat interface, style it to brand and embed it on the storefront.
Testing and tuning
Run 50+ real-world queries, test edge cases, tune tone and recommendation quality, harden security.
Launch and handover
Go live, monitor, set up alerting and transfer documentation.
The outcome
What the client received.
We produced three architecture options, so the client could choose based on where their product data lives and their data-residency policy.
We addressed prompt injection directly, with a layered defence and an honest statement that no system is fully immune, referencing documented incidents where companies were held liable for chatbot output.
We mapped the design against the OWASP Top 10 for LLM applications and documented our exposure on each risk.
We confirmed US-only data residency was achievable across every component, and verified that neither AI provider trains on customer data.
We defined the boundary explicitly: the assistant handles product discovery and does not touch customer accounts, payment systems or order history.
Questions
What people ask.
How do you stop an AI shopping assistant recommending products you do not sell?
By never letting the model search the catalogue itself. The system retrieves genuine matching products first, then passes only those verified items to the model to write the recommendation. The model can only describe products that were handed to it, so fabricated recommendations are structurally prevented rather than merely discouraged.
How quickly does the assistant know about price changes or products going out of stock?
Within seconds. Catalogue webhooks fire on any product change and re-index that item immediately. A full reconciliation runs every six hours as a safety net, so even a missed notification self-corrects the same day.
Is prompt injection a real risk for a retail chatbot?
Yes, and there are documented cases of companies being held liable for what their chatbot said. No system is fully immune. We use layered defences: role enforcement, input validation, output filtering, grounding every answer in real data, and session limits. If one layer is bypassed, the others still apply.
Does the chatbot replace existing customer support software?
No. It handles product discovery and recommendation. Support queries about orders, returns and accounts stay with the existing support platform, and the assistant directs customers there when a question falls outside its scope.

