Search theflght

Find a useful decision.

Join theflght

Get practical guides, straight to your inbox.

Pricing, hiring, positioning — the decisions that come after the idea. No spam, no fluff.

Strategy And Future

Should I Let AI Produce Work That Customers Will See?

Decide where AI can support customer-facing work. Assess error consequences, confidentiality and human review before letting generated output reach customers.

Should I Let AI Produce Work That Customers Will See?

Decide which customer-facing AI uses are safe enough to test, set human review thresholds, protect sensitive information and measure corrections.

Short answer: Yes, for narrow, reversible work such as first drafts of routine messages, provided a named person checks every factual claim, customer detail and promise before sending. Do not let AI independently give safety, legal, financial or diagnostic guidance, decide a consequential customer outcome, or process confidential information until you have verified the legal, contractual and technical controls. Pilot 50 low-consequence outputs with 100% human review before expanding the use.

The useful distinction is not whether customers see the words. It is what happens if one word is wrong. A poor caption can be corrected. An invented price, unsafe instruction or false deadline can create loss before anyone notices.

AI output can sound finished while still requiring verification. Fluency is not evidence.

> Jurisdiction note: AI use can engage data protection, consumer, intellectual-property, equality, confidentiality and sector rules. Requirements vary by country and use case, and guidance changes. Check current regulator guidance, customer contracts and provider terms, then use qualified legal or data-protection advice for material decisions.

Use the Customer Exposure Review

The Customer Exposure Review is a five-factor framework. Assess the specific output and process, not “AI” as one category.

  • Consequence: Low exposure: Awkward wording or easy correction; High exposure: Safety, rights, money or access affected; Control question: What is the worst plausible result of one wrong output?
  • Detectability: Low exposure: Reviewer can compare with a trusted source in seconds; High exposure: Error requires specialist knowledge or appears later; Control question: Can a competent reviewer reliably spot the mistake?
  • Reversibility: Low exposure: Message can be corrected before action; High exposure: Customer acts or loses an opportunity immediately; Control question: Can you restore the original position?
  • Information sensitivity: Low exposure: Public product facts; High exposure: Personal, confidential or restricted data; Control question: Are you permitted to provide these inputs to this system?
  • Accountability: Low exposure: Named person owns approval and complaints; High exposure: Nobody can explain or override the result; Control question: Who signs off and who can stop the process?

My position is that a small business should use AI as a drafting assistant before using it as a customer representative. A human must own the final communication. Removing that review is a separate decision that needs much stronger evidence than saving a few minutes.

Sort customer-facing work by consequence

Begin with actual tasks. “Customer service” is too broad.

  • Draft a routine appointment reminder from confirmed details: Starting position: Test with human review; Reason: Facts are limited and easy to compare
  • Rephrase a published explanation in plainer language: Starting position: Test with source comparison; Reason: Meaning can be checked against an approved source
  • Summarise a long customer message for an employee: Starting position: Test without replacing the original; Reason: Useful for triage, but nuance can disappear
  • Diagnose a fault or recommend a safety action: Starting position: Keep with a competent human; Reason: Error consequence and specialist judgement are high
  • Decide refund eligibility, credit, pricing or access: Starting position: Keep human decision authority; Reason: Outcome materially affects the customer
  • Handle a vulnerable customer's unusual complaint: Starting position: Human-led; Reason: Context, fairness and empathy require accountable judgement

A reminder containing medical appointment details is not low exposure merely because it is a reminder.

Specify what the system may do and what it must not do. “Draft a polite confirmation using only the five fields supplied, without adding advice, guarantees or new dates” is more testable than “answer the customer”.

Build the review before the prompt

A human review is meaningful only if the reviewer has the source information, competence, time and authority to reject the output. Reading it quickly for tone is not enough.

For each use case, create an approval check:

  1. Match names, dates, amounts, locations and reference numbers against the source record.
  2. Confirm that every promise is authorised and operationally achievable.
  3. Remove unsupported claims, advice and implied certainty.
  4. Check that sensitive information was handled through an approved route.
  5. Read from the customer's position for clarity, fairness and next action.
  6. Record material corrections during the pilot.

The ICO's current AI and data-protection guidance is under review following changes made by the Data (Use and Access) Act. It emphasises context-specific risk, accountability and data-protection principles. The ICO's human-review guidance says reviewers should have the knowledge, authority and independence to challenge decisions. Check the latest version for your use case.

Do not ask a junior reviewer to validate specialist advice they cannot independently assess. That is ceremony, not control.

Pilot against defined failure classes

Select one low-consequence use and 50 real outputs. Fifty is a working operational threshold, not a statistical guarantee. It is enough to expose recurring defects without inviting an open-ended experiment.

Before starting, define three failure classes:

  • Critical: Example: Wrong price, unsafe instruction, false contractual promise, prohibited data use; Response: Stop the pilot and investigate before any further output
  • Material: Example: Wrong date, missing customer constraint, misleading certainty; Response: Correct, record and revise the process
  • Editorial: Example: Tone, repetition or wording that does not change meaning; Response: Correct if useful and count separately

Review every output before the customer sees it. Record the correction, likely cause and whether the reviewer could detect it from the available source.

Do not approve unsupervised use merely because no critical failure appears in 50 cases. Consider volume, changing inputs, new customer groups and the cost of one missed error. Keep ongoing samples and an immediate stop route.

Worked example: North Mere Appliance Repair

North Mere sends 180 appointment confirmations and routine preparation messages a month. Writing each manually takes an observed six minutes, so monthly effort is: 180 × 6 minutes = 1,080 minutes = 18 hours.

At an internal admin cost of £28 per hour, manual cost is 18 × £28 = £504.

The business tests AI-assisted drafts for 60 routine messages. A coordinator compares every draft with the booking record and approved preparation text. Drafting and review take two minutes per message, which would be six hours across 180 messages.

  • Coordinator review: Calculation: 6 hours × £28; Amount: £168
  • AI service charge: Calculation: Quoted monthly amount; Amount: £60
  • Ongoing total: Calculation: £168 + £60; Amount: £228
  • Capacity value released: Calculation: £504 - £228; Amount: £276

Initial setup and pilot design take another three hours at £28, or £84. First-month saving is therefore £276 - £84 = £192.

During the 60-message pilot, the reviewer records seven material corrections, including dates and preparation steps. None reaches a customer because review precedes sending. North Mere keeps the process assisted rather than automatic. It excludes fault diagnosis, electrical-safety advice, prices and compensation decisions, where a plausible error would carry a higher consequence.

The result is not “AI is accurate”. It is that the bounded workflow releases £276 of monthly capacity while retaining a review that caught real defects.

Decide when to tell the customer

Practitioners disagree about disclosure. One view is that every AI-assisted output should be labelled. Another is that customers care about accuracy and data use, not whether a draft received automated help alongside spelling or formatting support.

My view is to disclose when AI use is material to what the customer is buying, how their information is processed, or how a consequential result is reached. Also disclose when the contract, regulator or customer's stated condition requires it.

If a client hires you for personal expert judgement, quietly replacing that judgement with automated output changes the service. If AI drafts a routine collection reminder that a person verifies, a label on every sentence may add little. Legal transparency duties and reasonable customer expectations still need specific assessment.

Make it easy for a customer to reach a competent person, challenge a result and correct their information.

Keep ownership with the business

Name one owner for the use case. That person controls approved inputs, review criteria, correction records, customer complaints and suspension. A provider update, staff change or new data source triggers another test.

Retain the source used to approve important facts. If a customer challenges a date, you need more than the generated message. Record who approved consequential communications and when.

Do not upload customer data until you understand what data is sent, where, for what purpose, under which terms, who can access it and how long it remains. Use the minimum information required. If the business cannot answer those questions, use synthetic examples for process testing and keep live data out.

Related guides

Run a controlled test over ten working days

Today, choose one low-consequence customer communication and define prohibited content, approved sources and the five exposure factors. Within two days, confirm contractual and data-handling requirements and write the human approval check.

Over the next week, process 50 outputs with 100% review, recording every critical, material and editorial correction. Stop immediately after a critical failure. On day ten, compare time, corrections and customer outcomes with the manual route. Expand only one boundary at a time and retain a named person who can suspend the use without waiting for permission.

Frequently asked questions

Should I tell every customer whenever AI helped write a message?

Not necessarily, but disclose when the use materially affects the service, customer expectation, personal-data processing or a consequential decision, and whenever law, contract or regulator requires it. Reasonable practitioners disagree about labelling minor drafting assistance. My view is that honesty should describe the real operating change, not add a generic badge to ordinary proofreading.

If the customer bought your personal judgement, substituting automated judgement is material. Check current legal duties for your jurisdiction and make sure customers can reach a responsible human, ask how their information was used and challenge an outcome.

Can I publish an AI draft after reading it once?

Only if that reading is an evidence-based review appropriate to the consequence. Compare facts with trusted sources, verify customer details and promises, remove unsupported advice, and confirm the tone and next action. A quick read for grammar will not reveal a plausible but false date or technical claim.

For specialist work, the reviewer must be competent in the subject. Keep high-consequence outputs human-led even when the draft appears strong. If you cannot state what the reviewer checks and which source they use, the review is too vague to protect the customer.

What customer information should I keep out of an AI system?

Keep out any personal, confidential, contract-restricted or security-sensitive information until you have confirmed a lawful purpose, necessary minimum, provider terms, security controls, retention, access and customer transparency. The exact answer depends on the system, deployment and jurisdiction. Replace real details with synthetic ones during early testing. Do not assume removing a customer's name makes the remaining content harmless; an address, unusual incident or account detail may still identify them. Ask a data-protection professional to review material processing, particularly sensitive information or decisions affecting individuals.

Who is responsible when an AI-assisted message is wrong?

Your business remains accountable to the customer for the communication it approves and sends, subject to the applicable contract and law. Blaming the system does not correct the customer's loss or restore trust. Name an internal owner, retain the source and approval record, and give staff authority to stop the workflow.

Review provider remedies separately, but do not promise customers that a third party will fix the issue. For a material error, correct the facts promptly, explain the practical consequence, offer the appropriate remedy and investigate why the review failed before restarting.

Which customer-facing tasks are safest to test first?

Start with frequent, low-consequence drafts whose facts come from a small, reliable source and can be checked in under two minutes. Appointment confirmations, plain-language versions of approved text and internal summaries for a human agent can fit, depending on the information involved.

Avoid diagnosis, safety instructions, eligibility, credit, pricing exceptions, legal positions and vulnerable-customer complaints. Select one task and one customer group. A broad chatbot handling anything is not a first test. Measure corrections and time against the existing manual process before considering a wider scope.

How do I know whether human review is actually working?

Seed the test with known edge cases, record what reviewers change and check whether they can explain why an output was accepted. A meaningful reviewer has the original source, relevant competence, enough time and authority to reject or rewrite.

If approval rates approach 100% while known defects pass through, the process has become a rubber stamp. Sample completed communications against source records and review customer corrections or complaints. Rotate a second competent reviewer through a small sample. Do not measure review quality by speed alone, because pressure can erase the control you intended to retain.

When should I stop using AI for a customer-facing task?

Stop immediately after a critical failure, prohibited data use, unexplained change in output, inability to conduct meaningful review or customer harm that your controls did not anticipate. Pause after repeated material corrections exceed the limit you set for the pilot.

Investigate whether the cause is the input, process, provider, reviewer or unsuitable use case. Resume only after the correction is tested on representative cases. Also stop if the time spent reviewing and repairing exceeds the capacity released. Continuing because setup took effort is sunk-cost thinking, not an operating decision.

BUSINESS ADVISER — Editor at theflght

Practical guides for founders making the decisions after the idea.

Comments (0)