Skip to content

Internal AI assistant built for accuracy, not confidence

We designed an enterprise knowledge assistant that answers technical support questions in Slack from scattered documentation and undocumented expertise, with citations and escalation instead of guesses.

  • 5

    Layers of accuracy safeguards

  • 5

    Knowledge sources brought into one place

The problem

Technical support, sales support and operational decisions inside the business depended on a handful of individuals. Knowledge lived across a documentation platform, years of Slack history, a support ticketing system, product PDFs, and — critically — in people's heads with nothing written down. A previous attempt at an AI assistant had already failed because it returned incorrect answers, which made accuracy the defining constraint of the engagement rather than a nice-to-have.

  • Answers to technical and warranty questions depended on a few key individuals, creating a bottleneck and a single point of failure.

  • Knowledge was scattered across a documentation platform, Slack history, support tickets and product PDFs with no unified way to search it.

  • Decision guidance including warranty rules, return logic and escalation thresholds existed only as tribal knowledge, with no written standard operating procedures.

  • A previous AI assistant attempt had failed by returning incorrect answers, damaging internal trust in the approach.

  • Product documentation contained diagrams and images carrying information that text-only search could not reach.

What we designed

We built the architecture around the failure mode rather than the feature list. Because a previous attempt had failed on accuracy, we designed five layered safeguards so that an unanswerable question routes to a human instead of producing a confident guess. A correct handover counts as a good outcome. We also treated the undocumented expertise as a first-class data source, extracting it through structured interviews and publishing it as standard operating procedures, which leaves the business with a durable asset independent of the AI.

#support-help

Team member

What is the return window for part AX-2210 with a cracked lens?

Assistant

Returns for physical damage are accepted within the window set out in the returns policy, provided the part was not installed.

[1] Returns policy, section 4[2] Support ticket archive

Team member

Does that change for fleet customers?

Assistant

I do not have a verified source for fleet terms. I have sent this to the warranty specialist.

Escalated to a person

What the system did

  • Answer 1: confidence above threshold, returned with citations
  • Answer 2: below threshold, escalated to a human
  • Ingestion pipeline

    Connectors for the documentation platform, Slack history, support tickets and product PDFs with semantic chunking and rich metadata

  • Vision extraction

    Converts diagrams and images in product documentation into searchable text while preserving links to the originals

  • Hybrid retrieval

    Combines keyword and semantic search so exact part numbers and described symptoms both resolve correctly

  • Reranking layer

    Reorders retrieval candidates by true relevance before the model sees them

  • Citation enforcement

    Every claim must cite a source passage, and the assistant declines to answer when it cannot

  • Confidence gate

    Low-confidence responses escalate to a human expert instead of being returned

  • Slack bot

    Threads, direct messages and reaction-based feedback capture

  • Evaluation framework

    Expert-validated question set with automated accuracy testing after every change

How it works

From trigger to result.

How the system works

8 stages

  1. 01Knowledge sources

    Trigger

    Documentation platform, Slack history, support tickets, product PDFs, expert interviews.

  2. 02Ingestion

    Logic

    Semantic chunking with source, category, date and confidence metadata.

  3. 03Vision extraction

    AI

    Diagrams become searchable text, linked to the original images.

  4. 04Hybrid retrieval

    Logic

    Keyword and semantic search run in parallel.

  5. 05Reranking

    Logic

    Candidates reordered by true relevance.

  6. 06Cited answer

    AI

    Every claim must cite a source chunk.

  7. 07Confidence gate

    Logic

    Above the threshold: answer in Slack. Below it: escalate to a named expert.

  8. 08Feedback loop

    Person

    Reactions queue low-rated answers for expert review.

Expert-verified content carries higher retrieval priority than automatically ingested content, so captured institutional knowledge outranks a stale document when both match a query.

Our approach

How we worked.

  1. Discovery and knowledge engineering

    Structured expert interviews to extract undocumented rules, a data source audit, taxonomy design, and a post-mortem of the previous failed attempt.

  2. Ingestion pipeline

    Build source connectors, chunking with metadata, vision extraction for diagrams, and initial index population.

  3. Retrieval engine

    Hybrid search, reranking, citation enforcement, confidence scoring and model integration with guardrails.

  4. Slack bot and escalation

    Bot deployment with thread and direct message handling, feedback capture, and routing to human experts.

  5. Evaluation and tuning

    Build an expert-validated question set, run automated accuracy testing, and tune iteratively.

  6. Internal pilot and hardening

    Deploy to a limited support-team group on real queries, then add monitoring, documentation and training.

The outcome

What the client received.

  • We designed around the previous failure rather than ignoring it, making a post-mortem of the earlier attempt a formal phase of the work.

  • A question the assistant cannot answer reliably goes to a named expert instead of producing a guess.

  • The extraction of undocumented expertise produces written standard operating procedures, an asset that holds value independently of the AI system.

  • We required a measurable accuracy benchmark to be agreed before any commitment.

Questions

What people ask.

Why do internal AI assistants usually fail?

They answer confidently when they should not answer at all. A single wrong response on a warranty or technical question destroys trust faster than a hundred correct ones build it. The fix is architectural: enforce citations, score confidence, and escalate to a human below a threshold rather than generating a plausible guess.

How do you get knowledge that only exists in people's heads into an AI system?

Through structured interviews with the individuals who hold it, converted into written standard operating procedures and ingested as a first-class source with higher priority than automatically gathered content. This is real effort, and it produces documentation the business keeps regardless of the AI.

Why use hybrid search instead of semantic search alone?

Semantic search understands meaning but misses exact strings. In technical support, part numbers and product codes must match precisely. Combining keyword and semantic retrieval catches both the exact identifier and the described symptom, which pure semantic search would miss.

How is accuracy actually measured rather than asserted?

With an evaluation set of real questions that have known correct answers, validated by your experts and run automatically after every change to retrieval or prompts. The benchmark is agreed with the client up front, and a correct handover to a person counts as success.

Related work

See all work

Have a problem like this one?