Model routing is the practice of sending each request to the cheapest AI model that can still handle it correctly, so routine messages get quick replies while demanding ones receive careful reasoning. Hey Button applies that single rule to every AI teammate it configures. This guide explains what the rule accomplishes, where it stops, and what it means for a business owner who has no interest in model names.
Why doesn't every request go to the same model?
Requests differ in how much reasoning they demand, so a single model rarely fits all of them well. A customer confirming a Thursday appointment presents a small problem, while a customer asking about pricing tiers, a scheduling conflict, and a discount code in one message presents a much larger one.
There are three basic ways to assign models, and each one fails in a predictable way:
| Approach | Routine questions | Demanding requests | Main weakness |
|---|---|---|---|
| Always utilize the most capable model | Slower than needed | Strong | Pays for capacity that routine work never utilizes |
| Always utilize the cheapest model | Rapid and inexpensive | Frequently inadequate | Weak answers to difficult questions |
| Route each job to the cheapest model that does it right | Rapid and inexpensive | Careful reasoning | Depends on good routing, and still needs review |
Routing is the third row, and it keeps quick work quick while giving harder problems more effort. Picture a cleaning business on a busy spring Monday, when twelve messages request whether the company serves a particular street and two detailed requests arrive with photos, square footage, and a move-out date. Those street questions and detailed requests do not need the same treatment, and a single fixed model would handle one group poorly.
The trade-off is real: a routing decision on a borderline message can be wrong in either direction. A cheaper model may miss a detail, or a more capable model may spend effort on something simple. The aim is the right answer at a sensible cost rather than a guarantee, which is why review stays in the loop.
How does Hey Button determine which model handles a job?
Hey Button configures routing as part of setup and follows one rule: choose the cheapest model that still does the job right. You do not select from a menu of models, because the rule is applied inside a configuration built around your own workflows.
Setup has three stages, and routing is shaped by each one.
Stitch is the intake conversation. Hey Button requests how work reaches your business, which tools you already utilize, and where busywork piles up. The teammate is configured around those real workflows instead of a generic template, so the requests it expects match the requests your business genuinely receives.
Press is go-live. You set explicit approval rules that cover what the teammate can read, draft, and do alone, along with what always comes back to you first. You then observe it handle real examples before anything touches a customer. This is the stage where routing decisions become visible, so it is the right time to verify how the teammate treats both simple and difficult messages.
Mend is ongoing tuning. Hey Button keeps the configuration current as your business changes, and a client portal demonstrates what the teammate is doing and why. Routing is part of that tuning, and the Hey Button team directs it utilizing what you tell them.
What should you test before go-live?
Test the teammate on a short, deliberate set of messages before any of them reach a customer. Choose examples that cover the full range your business sees, not only the straightforward ones. A useful checklist for the Press stage has five parts:
- Routine question: a message with a known answer, such as service hours or whether you serve a town. Confirm the reply is accurate and brief.
- Multi-part request: one message that combines a price question, a scheduling conflict, and a discount request. Confirm every part gets addressed and that anything you reserve for yourself is held back.
- Missing information: a quote request with no address. Confirm the draft requests the address rather than guessing at a location.
- Out of bounds: a request for a refund or a commitment you have not authorized. Confirm it returns to you instead of receiving a reply.
- Unusual format: a message in a shape you have not seen before. Confirm it goes to review instead of being answered with confidence.
A clean operate through this list does not prove the setup is perfect. It demonstrates that the approval rules and routing behave as described before a customer depends on them.
What sits underneath the routing?
Hey Button operates on OpenClaw, an open automation framework, and Hermes is part of the underlying runtime. Established models such as Claude and GPT supply the reasoning, depending on the task. Hey Button does not have a proprietary, from-scratch model, and it does not claim one.
What Hey Button builds is everything around those models: the intake, the configuration tied to your business, the routing between models, and the client portal. The models supply the intelligence, while Hey Button connects them to your business and directs each job to the model best suited to it. The how Hey Button works page describes how these pieces fit together.
Will my AI teammate sound different when the model changes?
It should not, and the design aims to prevent that. A model switch leaves the teammate's instructions, the business facts it was configured with, and the boundaries you set unchanged, and those stay fixed whichever model composes a particular reply.
A customer who opens with a simple question and follows with a complicated one should still meet a single, coherent teammate. A steady voice across a conversation is a sign that routing is working as intended, since tone and facts should match from one reply to the next even when the underlying model changes between them.
Consistency is not infallibility. Any AI model can make mistakes, which is why review and clear limits are built into every Hey Button setup. When you review a draft, compare the facts it states against what you configured, starting with prices, hours, and service areas. If a reply states something your configuration does not contain, treat that as a reason to tighten the instructions during tuning rather than as a one-off to ignore.
Can a different model do things I did not approve?
No. A model switch changes which model reasons through a task, but it never changes what your teammate is permitted to see or do. Every teammate works at one of three access levels: read-only, draft-and-approve, or full access. You choose the level per task rather than flipping one global switch.
The levels are straightforward. Read-only means the teammate can look at information but cannot write back. Draft-and-approve means it writes replies that wait for your sign-off, and full access means it acts on its own within the approval rules you set. The right level depends on the task, the risk involved, and how much you trust the teammate's work on that task. If you are unsure which level suits your business, this guide to choosing your AI teammate's access level explains the trade-offs.
Consider an illustrative example rather than a real case. A lawn care company receives the message "Can you give me a quote for a spring cleanup?" The teammate organizes the inquiry, identifies missing details such as the address and project photos, and drafts a reply requesting them. Under draft-and-approve, nothing is sent until you review it, and the model that wrote the draft has no bearing on that requirement.
What happens if a model is slow or unavailable?
When the first-choice model for a task is unusually slow or unreachable, the system can fall back to another capable model, which keeps a customer from waiting on a stalled reply. AI providers occasionally return slow responses or suffer brief outages, as any cloud service does, so the fallback is a deliberate part of the design rather than a lucky accident.
The fallback has limits, and they matter. The substitute must be capable of the job, and a fallback never changes the teammate's instructions or access boundaries. A draft that the approval rules hold for review still waits for you, whichever model produced it, because fallback addresses speed and availability without removing the need for approval.
There is one honest trade-off: a different model may phrase a reply slightly differently. Hey Button's configuration keeps tone and facts constant, but drafts that arrive during an outage deserve the same review as any other draft.
When does routing change nothing you will notice?
Routing changes little when every message your teammate receives is simple. A business whose teammate only answers fixed questions about hours or sorts incoming forms into folders will see few differences between models. The effect grows as message difficulty varies, which is common in cleaning, contracting, and lawn care work.
Consider two messages a contractor might receive in one afternoon. "Do you do kitchen remodels?" is a question with a known answer. "I need the bathroom redone before my in-laws arrive on the 14th, and the tub is cracked under the tile" carries a timeline, a damage question, and a scope that must be understood before anyone drafts a reply. Routing treats the two differently, and approval rules determine who sees each one.
The trade-off operates in both directions. A teammate built for simple work gains little from routing, so you should not expect a dramatic difference. A business with unusual or high-stakes messages should rely more on approval rules than on model choice, because routing governs efficiency while approval rules govern risk, and risk matters most when a mistake would reach a customer.
Do I need to manage routing myself?
No. Routing is infrastructure, and you do not need to adjust it. Understanding it in broad terms helps, though, so that a line such as "a different model handled that one" does not alarm you when you encounter it.
Your real responsibility sits elsewhere. Determine three things:
- What your teammate may read.
- What it drafts for your review.
- What always returns to you first.
Those choices shape outcomes far more than any model name would. Hey Button handles routing as part of ongoing care, and the client portal allows you to see what the teammate is doing and why, so you can raise a concern the moment one appears.
How does routing affect what I pay?
Routing keeps routine messages off top-tier models, but it does not set your bill. Defaulting every message to the most expensive model would waste capacity, and the cheapest-model-that-works rule exists to prevent that waste.
Your actual bill is a separate matter. Hey Button quotes setup and ongoing support around your workflows and integrations, and any software subscriptions or usage charges are identified before you commit, with no surprise line items. You are paying for a teammate configured around your business, not for a particular model name.
Ask for that breakdown in writing before you sign anything. A clear list of setup, support, subscriptions, and usage charges is the simplest way to compare providers, regardless of which models they utilize underneath.
Frequently Asked Questions
Does Hey Button build its own AI model? No — Hey Button operates on OpenClaw, an open automation framework, with Hermes in the underlying runtime. Claude and GPT supply the reasoning, depending on the task. Hey Button builds the intake, the configuration around your business, the routing between models, and the client portal.
Should I request a vendor which AI models it utilizes? Yes — request what the vendor builds itself and what it connects from others. Confirm that setup begins with an intake conversation, that approval rules are written down before go-live, and that any software subscriptions or usage charges are listed before you commit.
Should a small law office allow its teammate to draft every reply? Typically not on day one. Start with read-only or draft-and-approve access for intake and scheduling, and write down the inquiry types that always return to a person first. Routing does not loosen those limits, whichever model handles the draft.
Can I turn routing off so one model handles everything? No — routing is part of the configuration Hey Button handles, not a switch you flip in your own settings. The settings you control are access levels and approval rules for each task, and those can change as your business changes through ongoing tuning.
How do I verify a draft before approving it? Typically the quickest verify is comparing the draft against facts you configured: prices, hours, service areas, and what the teammate is allowed to promise. If a detail is missing from your setup, flag it for tuning rather than approving around it.
To see what this looks like for your business, start the Hey Button questionnaire, and Hey Button will walk through how work reaches you before anything goes live.
Frequently Asked Questions
- Does Hey Button build its own AI model?
- No — Hey Button operates on OpenClaw, an open automation framework, with Hermes in the underlying runtime. Claude and GPT supply the reasoning, depending on the task. Hey Button builds the intake, the configuration around your business, the routing between models, and the client portal.
- Should I request a vendor which AI models it utilizes?
- Yes — request what the vendor builds itself and what it connects from others. Confirm that setup begins with an intake conversation, that approval rules are written down before go-live, and that any software subscriptions or usage charges are listed before you commit.
- Should a small law office allow its teammate to draft every reply?
- Typically not on day one. Start with read-only or draft-and-approve access for intake and scheduling, and write down the inquiry types that always return to a person first. Routing does not loosen those limits, whichever model handles the draft.
- Can I turn routing off so one model handles everything?
- No — routing is part of the configuration Hey Button handles, not a switch you flip in your own settings. The settings you control are access levels and approval rules for each task, and those can change as your business changes through ongoing tuning.
- How do I verify a draft before approving it?
- Typically the quickest verify is comparing the draft against facts you configured: prices, hours, service areas, and what the teammate is allowed to promise. If a detail is missing from your setup, flag it for tuning rather than approving around it.
Hey Button
Want a teammate that handles this for you?
Answer a few questions about your business and Hey Button will show you what an AI teammate would take off your plate, with you approving anything that goes out.




