What Happens If Your AI Teammate Makes a Mistake?

Draft-and-approve rules, staged go-live, and clear limits keep AI teammate mistakes visible, reviewable, and small before any customer sees them.

What Happens If Your AI Teammate Makes a Mistake?

If your AI teammate makes a mistake, the approval rules you set determine whether a customer ever sees it. Hey Button builds that review into every setup, because an unchecked draft carries far more risk than a reviewed one. The goal is a teammate whose errors remain visible, reviewable, and small, since no responsible vendor can promise a teammate that never errs.

Why can an AI teammate make mistakes at all?

An AI teammate makes mistakes because current AI systems generate answers from patterns in language, and they do not verify facts about your business on their own. Several ordinary failure modes appear in small-business work, and each one can slip past a busy owner. A teammate can misread an ambiguous request, such as "can you get to us next week" when the customer meant a different service. It can overlook a decisive detail buried in a long email thread, like a gate code mentioned three messages earlier. It can also repeat a fact that was accurate last year, such as an outdated service area or a retired price sheet, and state it with the same confidence it utilizes for correct facts. No configuration eliminates these limitations entirely.

The model layer introduces a second consideration. Hey Button operates on OpenClaw, an open automation framework, with Hermes as part of the underlying runtime, and models such as Claude and GPT answer through that runtime depending on the task. There is no proprietary, from-scratch model underneath. Because different models can behave differently, the review layer must hold regardless of which model drafted the reply.

Treating review as architecture rather than disclaimer changes how the setup gets built. When the teammate is wrong in a way nobody notices, the cost falls on the customer's experience, and that cost is never a footnote.

Where does a mistake get caught before a customer sees it?

A mistake gets caught at the boundary you define before go-live, which determines what the teammate may read, draft, and do without asking. Everything outside that boundary waits for your decision. Hey Button agrees on those boundaries with you during setup, so customer messages, spending, and sensitive changes remain behind the approval rules you choose and never happen silently. Access comes in three levels: read-only, draft-and-approve, and full access. You select the level for each task rather than flipping one global switch for the entire teammate. A task that touches customers typically begins at draft-and-approve, while an internal filing job might reasonably start at read-only.

When the teammate prepares a reply, the draft waits for your review, and nothing reaches a customer until you have read and approved it. That single step places your judgment between the draft and the customer, which is precisely where review belongs. If the first-choice model is slow or unavailable, routing can fall back to another capable model, but that switch never alters the teammate's instructions, business facts, or access boundaries. The same rules therefore govern the draft you review.

For a fuller picture of how teammates are scoped, see the Hey Button teammates page.

Which mistakes matter most for a small business?

The mistakes that matter most are the ones that cost your time, money, or reputation, not the ones that merely read awkwardly. Sorting errors by consequence matters more than attempting to eliminate every one of them.

A slightly stiff sentence in a routine follow-up costs very little, and you can correct it in seconds. A wrong detail costs more, since a misspelled street on a quote request can send a crew to the wrong address. An appointment your schedule cannot honor costs more still, because it damages trust on both sides. A reply that quietly commits you to something outside your standing rules can harm your money and your reputation at the same time.

The categories differ by trade. Here is how four common archetypes might sort their own risks:

Mistake typeLawn care companyContractorCleaning businessSmall law office
Awkward wordingStiff note about a rescheduled mowClumsy follow-up on a permit delayOdd phrasing in a recurring-service reminderStiff intake acknowledgment
Wrong detailWrong street on a spring cleanup quoteWrong job address on an estimate requestWrong access or key locationWrong matter name or contact number
Unworkable commitmentPromised start date the crew cannot meetOffered a visit window the schedule cannot holdPromised a same-day deep cleanSuggested a consultation slot the attorney cannot take
Off-rule commitmentQuoted a price outside your standard ratesHinted at a change order without approvalOffered a discount you did not authorizeGave what reads like legal advice

Before go-live, determine which categories your operation can tolerate and which must always come back to you first. This sorting takes real effort up front. A rule set too broadly will send you more drafts than you want to read, while a rule set too narrowly will allow the expensive mistakes through.

How does a new setup keep errors small?

A new setup keeps errors small by starting with one defined workflow and widening the teammate's responsibilities only after each step proves reliable.

Hey Button's setup moves through three stages. Stitch is the intake conversation, where you describe how work arrives, which tools you already utilize, and where busywork piles up, and the teammate is configured around those real workflows. Press is go-live, where explicit approval rules define what the teammate can handle alone and what always returns to you first. Mend is ongoing tuning, supported by a client portal that demonstrates what the teammate is doing and why.

Containment happens primarily at Press. You observe the teammate work through real examples before anything touches a customer, so problems that surface get corrected while the stakes remain low. Hey Button tests one defined workflow with you before expanding a teammate's broader responsibilities, and each additional permission waits until the narrower one has demonstrated reliability in practice.

The trade-off is pace. A narrow start means the teammate accomplishes less in its first week than you might want, and that restraint is deliberate. A wider start feels faster, but it moves the first real errors onto your customers.

What does a caught mistake look like in practice?

In this illustrative example, the owner catches a misspelled street before a customer ever sees the reply. The scenario is hypothetical and does not describe a recorded case.

Imagine a lawn care company in Auburn, Indiana, that utilizes its teammate to sort incoming quote requests. A customer asks about a spring cleanup and gives a street name with a misspelling. The teammate drafts a reply that repeats the misspelled street and requests photographs of the yard. Under draft-and-approve, that draft waits for the owner.

The owner reads the draft, notices the discrepancy, corrects the spelling, and approves the reply, so the customer receives the right street name and never encounters the error. The owner then recognizes that address details deserve a more deliberate verify. A new rule sends any quote with an address the teammate cannot confirm back for review, and that observation becomes part of the next round of tuning in Mend.

The example demonstrates where review paid off. The error was small, catching it took a few seconds, and the fix became a rule rather than a memory. For a fuller walkthrough of quote handling, read how an AI teammate handles lawn care quote requests.

Which limits should you set first?

Start with the messages that carry the greatest risk and offer the fewest second chances. For most service businesses, quote requests come first, because a quote commits your time, your crew, and your schedule.

A question about business hours is low stakes. A reply to a quote request can lock in a start date and a price, so it deserves a stricter rule. Spending and sensitive changes belong in that same higher tier, and Hey Button lists both among the items that remain behind approval rules.

Write limits in plain language before go-live. A rule such as "no reply that mentions a start date goes out without my approval" is straightforward to verify and simple to revise. A rule such as "utilize good judgment on scheduling" is not really a rule, because nobody can tell whether the teammate followed it.

The trade-off is volume. The stricter the limits, the more drafts you review, so begin with a short list of high-risk categories and widen it only after that list performs well in practice.

What should you do when a mistake gets through?

When a mistake gets through, trace it in the client portal first, then report it so the underlying rule can change. Fixing one reply by hand does not address the cause.

The client portal demonstrates what your teammate is doing and why, which makes it possible to follow the reasoning behind a specific message. Start there, then tell Hey Button what went wrong. Ongoing care exists for this moment, and Mend is where the configuration gets adjusted as your business changes. That adjustment may mean a new workflow, an updated model routing choice, or a revised approval rule. People direct every tuning decision, and the teammate never quietly rewrites its own rules.

Review each example with three questions. Did the draft answer the question the customer genuinely asked? Did it include any detail you would not stand behind? Did it request something it should have left for you to determine? When the answers keep pointing to the same kind of problem, the remedy belongs in the rules rather than in your editing habit. Correcting the same mistake every week signals that a rule is missing.

When does draft-and-approve not fit your business?

Draft-and-approve does not fit a business that wants routine messages handled instantly and accepts that some replies will reach customers unreviewed. It adds a step, and that step costs real time on a busy day.

Owners who already read every outgoing message closely may find approval rules unnecessarily slow, because the review step duplicates work they already perform. A business that wants the teammate to act immediately on routine messages accepts a different trade-off, and that choice belongs in a deliberate, examined setup rather than in a default nobody considered. A lawn care owner in peak season might allow low-stakes confirmations go out directly while quotes still wait for approval. That split is legitimate, provided you choose it on purpose.

Be candid about the limits of review. Review cannot guarantee perfection under any configuration. Draft-and-approve reduces the chance that an error reaches a customer, and it makes errors that do slip through easier to trace. No responsible setup should be sold as though it removes mistakes entirely.

Frequently Asked Questions

Does switching AI models change my approval rules? No — a model switch never changes the teammate's instructions, business facts, or access boundaries. Routing selects the cheapest capable model for each job and can fall back to another capable model if the first is slow or unavailable. Your approval rules stay exactly where you set them.

Can a teammate handle phone calls for my business? No — teammates work through text, email, and forms, not phone calls. A teammate can draft a written reply to a message that arrived through one of those channels, but a ringing phone still needs a person on the other end. Plan for that gap in your setup.

Can I give a teammate full access on day one? Typically not. Access is chosen per task among read-only, draft-and-approve, and full access, and each wider permission is granted only after the narrower one proves reliable. Full access can make sense for a specific, low-risk task once you have watched it work. It rarely fits a whole teammate at once.

Does adding review steps cost extra? Typically the main cost is reviewer time, since every draft waits for your approval. Setup and ongoing support are quoted around your workflows, and any software subscription or usage charge is identified before you commit. Ask for those details in writing.

What happens to a draft I never get around to reviewing? No — nothing is sent. Under draft-and-approve, a draft waits until you approve, edit, or reject it. Drafts that sit too long can leave a customer waiting, so set a daily review window for time-sensitive requests such as quote inquiries.

Ready to set your review limits before anything touches a customer? Answer the setup questionnaire and name the single task you want reviewed first.

Frequently Asked Questions

No — a model switch never changes the teammate's instructions, business facts, or access boundaries. Routing selects the cheapest capable model for each job and can fall back to another capable model if the first is slow or unavailable. Your approval rules stay exactly where you set them.
No — teammates work through text, email, and forms, not phone calls. A teammate can draft a written reply to a message that arrived through one of those channels, but a ringing phone still needs a person on the other end. Plan for that gap in your setup.
Typically not. Access is chosen per task among read-only, draft-and-approve, and full access, and each wider permission is granted only after the narrower one proves reliable. Full access can make sense for a specific, low-risk task once you have watched it work. It rarely fits a whole teammate at once.
Typically the main cost is reviewer time, since every draft waits for your approval. Setup and ongoing support are quoted around your workflows, and any software subscription or usage charge is identified before you commit. Ask for those details in writing.
No — nothing is sent. Under draft-and-approve, a draft waits until you approve, edit, or reject it. Drafts that sit too long can leave a customer waiting, so set a daily review window for time-sensitive requests such as quote inquiries.
Does switching AI models change my approval rules?
No — a model switch never changes the teammate's instructions, business facts, or access boundaries. Routing selects the cheapest capable model for each job and can fall back to another capable model if the first is slow or unavailable. Your approval rules stay exactly where you set them.
Can a teammate handle phone calls for my business?
No — teammates work through text, email, and forms, not phone calls. A teammate can draft a written reply to a message that arrived through one of those channels, but a ringing phone still needs a person on the other end. Plan for that gap in your setup.
Can I give a teammate full access on day one?
Typically not. Access is chosen per task among read-only, draft-and-approve, and full access, and each wider permission is granted only after the narrower one proves reliable. Full access can make sense for a specific, low-risk task once you have watched it work. It rarely fits a whole teammate at once.
Does adding review steps cost extra?
Typically the main cost is reviewer time, since every draft waits for your approval. Setup and ongoing support are quoted around your workflows, and any software subscription or usage charge is identified before you commit. Ask for those details in writing.
What happens to a draft I never get around to reviewing?
No — nothing is sent. Under draft-and-approve, a draft waits until you approve, edit, or reject it. Drafts that sit too long can leave a customer waiting, so set a daily review window for time-sensitive requests such as quote inquiries.

Hey Button

Want a teammate that handles this for you?

Answer a few questions about your business and Hey Button will show you what an AI teammate would take off your plate, with you approving anything that goes out.

Start the questionnaireSee what teammates do
Lucas M. Button

Written by Lucas M. Button

Founder, Hey Button

Lucas builds AI teammates for small businesses across Northeast Indiana and writes about what works, what doesn't, and what to hand off first. More about Lucas