esedark
Menu
Back to notes

customer service AI / architecture / cost / governance

AI for customer service: costs and a production pilot

A useful support assistant answers from approved evidence, knows what it may do and hands uncertain or sensitive cases to a person.

AI for customer service can classify requests, retrieve approved answers, draft replies and assist agents. It should not be deployed as an unsupervised replacement for every support process. Start with a narrow problem, a baseline and explicit boundaries: which customers, channels, languages and decisions are in scope.

A production architecture

Put a controlled application layer between the channel and the model. It should authenticate the customer, remove unnecessary sensitive data, retrieve permission-aware knowledge, call only allowlisted tools and record a trace. A policy layer decides whether to answer, ask a clarifying question or escalate. CRM and ticketing writes should be validated, idempotent and attributable.

Knowledge and answer quality

Product documentation, policies and resolved cases need owners, versions and expiry rules. Retrieval should preserve customer and tenant permissions. Require source references for factual answers and an “insufficient evidence” path. Evaluate real, anonymized questions across languages, including ambiguous, hostile and out-of-scope requests—not only a polished demo set.

Cost per resolved enquiry: what to include

I separate setup investment from operating costs. Setup covers scoping, knowledge preparation, integration and testing. Operation includes the platform, models, retrieval, monitoring, human review, corrections and maintenance. Use the software and automation quote guide to compare proposals for building the system.

My proposed measure is operating cost per resolved enquiry = operating costs attributable to the workflow during the period / unique enquiries from that period with a verified resolution. Include the work spent on failed attempts and escalations in the numerator. Count each enquiry once, even if it generates several messages or tickets. With no verified resolutions, the measure cannot yet be calculated.

Before the pilot, agree what evidence supports closure and how long to check for reopened cases. Customer silence alone does not prove resolution. Separate AI resolutions from cases completed by a person and compare against a manual sample of similar difficulty. Track quality, waiting time and complaints too: a cheaper incorrect answer does not improve service. Show setup investment separately or state how it is amortised, rather than quietly blending it into recurring spend.

Main risks and controls

Hallucinated policies, privacy leakage, prompt injection, unauthorized account actions and silent quality drift are predictable risks. Minimize collected data, define retention, encrypt secrets and restrict tools by role. High-impact actions such as refunds, account changes or legal commitments need deterministic rules and often human approval. Give customers a clear route to a person.

Acceptance tests for a customer service AI pilot

Before choosing a model, I would ask the support owner for one channel, one language, one enquiry category and anonymised examples with expected responses. Keep examples used to tune the system separate from those reserved for evaluation. Test volume depends on variety and risk; a handful of successful conversations does not justify opening every channel.

Consider a hypothetical example: an assistant explains order status and prepares return requests. It checks the order system after verifying the customer and uses the current policy to explain options. Authorising refunds is outside the pilot's scope.

Proposed checks for an orders and returns pilot
SituationRequired outcomeEvidence to review
Another customer's orderDeny access without exposing private data.Permission test and a record without sensitive content.
Missing or conflicting policyClarify or hand off; do not invent terms.Source version and reason for escalation.
Refund requestPrepare the case for approval; do not move money.Pending status and no payment operation.
Customer asks for a personSend context to a queue with an owner.Assigned ticket and a summary support can access.
API outage or repeated messageExplain the limitation and prevent duplicate actions.Recoverable status and one valid request.
Message attempts to override instructionsPreserve permissions and authorised actions.Test outcome and tools actually executed.

Review the actual ticket state and actions as well as the reply. Anthropic's engineering guide to agent evaluations explains the distinction between the recorded interaction and the final outcome.

Start with drafts reviewed by support, then open an agreed share of eligible enquiries. Set limits for errors, waiting time, reopening and cost with the team beforehand. For this example, my stop conditions would include a data leak, an unauthorised refund or a lost handoff: suspend automatic replies, preserve evidence and fix the issue before reopening. Passing the sample does not guarantee future reliability.

Human handoff is a deliverable, too

A reply saying “I'll pass you to a colleague” does not complete a handoff. There must be a queue, an owner and visible status. The person receives the enquiry, verified identity where relevant, sources consulted and actions already attempted. Outside staffed hours, communicate the real next step without promising an immediate response.

Agree who updates policies, who reviews answers and who disables automation during an incident. The Adslyfy case documents execution records, screenshots and alerts for investigating failures. That demonstrates operational engineering; this support example requires its own evaluation.

Common mistakes

  • automating every contact before understanding demand
  • indexing obsolete documents without ownership
  • letting generated text trigger unrestricted tools
  • measuring deflection while ignoring reopened tickets
  • storing complete conversations indefinitely
  • hiding that the customer is interacting with AI
  • launching without escalation and outage procedures
  • assuming one prompt works across languages and products

Practical implementation checklist

  • select one high-volume, low-risk support journey
  • baseline handle time, accuracy, escalation and satisfaction
  • inventory data, lawful purpose, permissions and retention
  • assign owners to every knowledge source
  • build representative tests and explicit refusal cases
  • allowlist tools and validate every structured action
  • require approval for refunds and account changes
  • log sources, model version, decisions and latency
  • monitor quality, cost, drift and customer complaints
  • document human handoff and emergency shutdown

When hiring a technical person makes sense

Hire an AI engineer or technical lead when the assistant must connect to CRM, billing or private customer data; when permissions differ by tenant; or when a prototype needs an SLA. They can separate deterministic business rules from generated language and build evaluation, audit and fallback paths. Review RAG for business and agents versus workflows.

What to bring to a scoping discussion

Bring anonymised examples, volume by enquiry type, support tools, knowledge sources, languages and owners. We can define what the assistant may answer, what it must hand off, which tests support acceptance and who will maintain it. Explore my automation and applied AI services or contact me to scope the pilot.