The problem
Technical support, sales support and operational decisions inside the business depended on a handful of individuals. Knowledge lived across a documentation platform, years of Slack history, a support ticketing system, product PDFs, and — critically — in people's heads with nothing written down. A previous attempt at an AI assistant had already failed because it returned incorrect answers, which made accuracy the defining constraint of the engagement rather than a nice-to-have.
Answers to technical and warranty questions depended on a few key individuals, creating a bottleneck and a single point of failure.
Knowledge was scattered across a documentation platform, Slack history, support tickets and product PDFs with no unified way to search it.
Decision guidance including warranty rules, return logic and escalation thresholds existed only as tribal knowledge, with no written standard operating procedures.
A previous AI assistant attempt had failed by returning incorrect answers, damaging internal trust in the approach.
Product documentation contained diagrams and images carrying information that text-only search could not reach.
What we designed
We built the architecture around the failure mode rather than the feature list. Because a previous attempt had failed on accuracy, we designed five layered safeguards so that an unanswerable question routes to a human instead of producing a confident guess. A correct handover counts as a good outcome. We also treated the undocumented expertise as a first-class data source, extracting it through structured interviews and publishing it as standard operating procedures, which leaves the business with a durable asset independent of the AI.
#support-help
Team member
Assistant
[1] Returns policy, section 4[2] Support ticket archive
Team member
Assistant
Escalated to a person
What the system did
- Answer 1: confidence above threshold, returned with citations
- Answer 2: below threshold, escalated to a human
Ingestion pipeline
Connectors for the documentation platform, Slack history, support tickets and product PDFs with semantic chunking and rich metadata
Vision extraction
Converts diagrams and images in product documentation into searchable text while preserving links to the originals
Hybrid retrieval
Combines keyword and semantic search so exact part numbers and described symptoms both resolve correctly
Reranking layer
Reorders retrieval candidates by true relevance before the model sees them
Citation enforcement
Every claim must cite a source passage, and the assistant declines to answer when it cannot
Confidence gate
Low-confidence responses escalate to a human expert instead of being returned
Slack bot
Threads, direct messages and reaction-based feedback capture
Evaluation framework
Expert-validated question set with automated accuracy testing after every change
How it works
From trigger to result.
How the system works
8 stages
01Knowledge sources
TriggerDocumentation platform, Slack history, support tickets, product PDFs, expert interviews.
02Ingestion
LogicSemantic chunking with source, category, date and confidence metadata.
03Vision extraction
AIDiagrams become searchable text, linked to the original images.
04Hybrid retrieval
LogicKeyword and semantic search run in parallel.
05Reranking
LogicCandidates reordered by true relevance.
06Cited answer
AIEvery claim must cite a source chunk.
07Confidence gate
LogicAbove the threshold: answer in Slack. Below it: escalate to a named expert.
08Feedback loop
PersonReactions queue low-rated answers for expert review.
Expert-verified content carries higher retrieval priority than automatically ingested content, so captured institutional knowledge outranks a stale document when both match a query.
Our approach
How we worked.
Discovery and knowledge engineering
Structured expert interviews to extract undocumented rules, a data source audit, taxonomy design, and a post-mortem of the previous failed attempt.
Ingestion pipeline
Build source connectors, chunking with metadata, vision extraction for diagrams, and initial index population.
Retrieval engine
Hybrid search, reranking, citation enforcement, confidence scoring and model integration with guardrails.
Slack bot and escalation
Bot deployment with thread and direct message handling, feedback capture, and routing to human experts.
Evaluation and tuning
Build an expert-validated question set, run automated accuracy testing, and tune iteratively.
Internal pilot and hardening
Deploy to a limited support-team group on real queries, then add monitoring, documentation and training.
The outcome
What the client received.
We designed around the previous failure rather than ignoring it, making a post-mortem of the earlier attempt a formal phase of the work.
A question the assistant cannot answer reliably goes to a named expert instead of producing a guess.
The extraction of undocumented expertise produces written standard operating procedures, an asset that holds value independently of the AI system.
We required a measurable accuracy benchmark to be agreed before any commitment.
Questions
What people ask.
Why do internal AI assistants usually fail?
They answer confidently when they should not answer at all. A single wrong response on a warranty or technical question destroys trust faster than a hundred correct ones build it. The fix is architectural: enforce citations, score confidence, and escalate to a human below a threshold rather than generating a plausible guess.
How do you get knowledge that only exists in people's heads into an AI system?
Through structured interviews with the individuals who hold it, converted into written standard operating procedures and ingested as a first-class source with higher priority than automatically gathered content. This is real effort, and it produces documentation the business keeps regardless of the AI.
Why use hybrid search instead of semantic search alone?
Semantic search understands meaning but misses exact strings. In technical support, part numbers and product codes must match precisely. Combining keyword and semantic retrieval catches both the exact identifier and the described symptom, which pure semantic search would miss.
How is accuracy actually measured rather than asserted?
With an evaluation set of real questions that have known correct answers, validated by your experts and run automatically after every change to retrieval or prompts. The benchmark is agreed with the client up front, and a correct handover to a person counts as success.

