Who Checks AI's Work When It Outpaces Your Team?

A new math release shows AI output can arrive faster than people can review it. Here is how a small business should decide who checks AI work.

Who Checks AI's Work When It Outpaces Your Team?

A named person on your team checks AI work before it reaches a customer, and that check has to be designed in, not added after the fact. The lesson originates in a story about mathematics, yet it applies with equal force to quotes, estimates, and customer replies at a plumbing company, a clinic, or a retail shop in Northeast Indiana.

Key Takeaways

  • Producing output with AI has become inexpensive, while understanding and approving that output has not, and that widening gap is where small businesses get hurt.
  • OpenAI's release of 372 novel mathematical results shows how quickly machine output can outrun human review, even among trained experts.
  • Mathematicians argue that a result only counts after someone understands, connects, and tests it, which makes review capacity the true constraint.
  • Most solved problems will not change daily work, though the matrix multiplication result hints at cheaper AI inference over time.
  • Decide in advance who reads which outputs, how much gets sampled, and what triggers a full review, before the volume arrives.
  • An AI teammate should draft and then wait for approval by default, and the access level you choose sets its limits.

This post draws on The AI Daily Brief episode of October 9, 2026, which covered the release in considerable detail. The episode focused mostly on the future of research, but its central observation translates cleanly to a company with five to fifty people. This post extracts those parts and leaves the deeper mathematics to the mathematicians who can evaluate it properly.

Why does a math release matter to a small business?

It matters because it demonstrates what happens when output arrives faster than people can absorb it. The same pattern can appear in quotes, estimates, and customer replies long before any formal review process has a chance to catch up with the rising volume.

The host framed the episode as a possible first field-level disruption from AI, which is a more specific claim than the usual sweeping predictions. Coding has changed more than any other profession, and demand for skilled coders has nonetheless continued to climb. Mathematics appears different in kind, because a large share of its hardest outstanding problems fell in a single release, faster than human experts could read the underlying work.

For a business owner, the useful question is not whether AI can accomplish something difficult. The useful question is whether your process can distinguish a correct result from a merely plausible one once the volume rises. A shop that adds AI to its estimating may produce several times more quotes each week. The owner, however, still has the same number of hours available to read them. That widening gap between production and review is precisely where mistakes slip through unnoticed.

What did OpenAI actually release, and how fast did it happen?

OpenAI released 372 novel mathematical results, accompanied by 722 supporting papers, in a public repository on GitHub. The results originated from an internal frontier model that the public cannot currently access. The company also reported that the average result consumed roughly three hours of ChatGPT Pro-level thinking, a figure the host described as the most striking detail in the release.

That number carries weight because of how rapidly it fell. Only a month earlier, the Navier-Stokes result had required roughly 10,000 parallel AI runs and 88 hours of computation. The efficiency gain arrived within a few weeks rather than over several years. Whatever one thinks of the mathematics itself, a cost curve that steep alters what a small business can afford to attempt and how often it can afford to attempt it.

Scale matters as well. Measured against an AI-generated list of the field's top 500 open problems, compiled by Proof Atlas, OpenAI fully solved 90 of them. Dozens of additional partial proofs would qualify as significant contributions on their own. Among those partial results were advances on the Riemann hypothesis and the Hodge conjecture, two of the seven Millennium Prize problems.

OpenAI also coordinated the release with a newly formed mathematics advisory committee. At the committee's suggestion, the company published reasoning traces for a sample of ten problems and disclosed the average compute spent per result. That habit deserves imitation in your own vendor relationships. When a vendor shows its work, you can verify it independently. When a vendor offers only a headline, you are trusting that headline without any means of checking it.

Why do mathematicians say checking is the real bottleneck?

A result only counts once someone understands it, connects it to existing knowledge, and tests it against what is already established. Producing correct statements has stopped being the scarce ingredient in the process. Reading, judging, and reusing those statements is now the constraint that determines how much genuine progress actually occurs.

Mathematician Francesco Maggi made this point most directly. He argued that mathematics advances through understanding, connection, and reuse, rather than through an ever-growing stockpile of correct statements. Results that nobody picks up, studies, and links to the wider body of knowledge remain inert. The Navier-Stokes proof from the previous month was still under verification, and hundreds of additional results now waited in the queue behind it.

Physicist Steve Hsu pushed the logic considerably further. He suggested that formal systems might eventually verify the proofs, yet humans could still lack the context required to comprehend the web of machine-invented concepts beneath them. In that scenario, people would receive selected explanations from a larger system rather than comprehend the work themselves. Hsu's scenario is speculative, but the bottleneck he describes already appears in smaller forms today.

Translating that observation to a small business is straightforward. Picture an Auburn landscaping crew whose AI teammate drafts forty quotes in a single week. The scarce resource is no longer the drafting itself. The scarce resource is the owner's attention on the handful of quotes that carry unusual risk, such as a new commercial client, an unfamiliar material, or a price far outside the usual range. Before you enable any AI output, decide three things in advance:

  1. Who reads what. Name the person who approves each category of outgoing message, and keep that list short enough that the person can realistically keep pace with it.
  2. How much gets sampled. Select a fixed share of routine output for weekly review, and raise that share whenever errors begin appearing in the sample.
  3. What triggers a full review. Define clear conditions in advance, such as a price above a chosen threshold, a customer complaint, or any request that falls outside the usual scope of work.

None of this argues against using AI output. It argues for deciding, before the volume arrives, who has both the time and the judgment to check it. Without that decision, the review step quietly disappears, and the business ends up learning about its errors from its customers.

Which results could change everyday work?

Very few of them, at least for the foreseeable future, will alter daily work. Most of the solved problems are impressive but impractical for daily work. The one with clear near-term reach is a theoretical improvement in matrix multiplication, the arithmetic that sits underneath AI inference. Even that effect must travel through substantial engineering before anyone notices it in a product.

Most of the other results will leave how most people work essentially unchanged. Navier-Stokes illustrates the point well. Engineers designing aircraft and other fluid systems already rely on simplified versions of the equations, which means a proof of the full problem does not alter their everyday instruments. The Erdős problems mentioned in the episode are similarly abstract, with no obvious application in the physical world.

Matrix multiplication differs in one important respect. The episode described a new bound that lowers the efficiency exponent from roughly 2.37 to no more than 2.25. Cornell's Steven Strogatz compared the leap to Bob Beamon's famous long jump, a performance that stood far apart from everything preceding it.

The catch is that a theoretical bound is not a shipped product, and it does not guarantee that any software runs faster tomorrow. It means researchers have identified a more efficient method in principle, and engineers must still translate that method into reliable, working software.

For a business owner, the practical lesson is to judge tools by measured results on your own work. A headline about a theoretical record should prompt a question about evidence rather than trigger a purchase. Ask any vendor to demonstrate how its speed or accuracy changed on tasks resembling yours, and remain cautious about any claim that cannot be independently checked.

What else was in the October 9 episode?

The episode also covered a correction to OpenAI's reported revenue, a dispute over how Anthropic counts its revenue, and new survey data about how organizations are spending on AI. Each of those points affects how a business should interpret vendor claims, so each deserves a brief summary here.

  • Revenue. The Financial Times reported that OpenAI told investors it reached about $50 billion in annualized revenue at the end of September. A widely circulated $68 billion figure had originated from an investor's estimate rather than from OpenAI itself. OpenAI also reported 77% run-rate growth and 107% enterprise growth, yet semiconductor stocks still fell sharply, with Nvidia down 3% and Oracle down 5.5%.
  • Accounting. The host noted that Anthropic reports revenue before revenue sharing with cloud partners, whereas OpenAI reports net figures. He expects Anthropic's audited financial statements to look considerably lower than the numbers private investors have seen so far.
  • Survey data. In The Information's subscriber survey, 35% of respondents said their organization returns multiples on AI spend, while only 8% described AI spend as a net negative. Just 22% said they were hiring less because of AI. The host cautioned that two-thirds of those readers work at organizations that build AI applications, so the sample leans heavily toward enthusiasts.
  • Open-weight models. About 72% of respondents use open-weight models with some regularity, and 20% said those models carry the majority of their organization's workload. Only 21% use Chinese open-weight models, which suggests many organizations remain cautious about relying on foreign-built systems.
  • Terms of use. Anthropic's new terms prohibit sustained and needless abusive behavior directed at its models. The company states that the rule is intended for extreme circumstances rather than for ordinary frustration or disagreement.

For a Northeast Indiana shop choosing a vendor, the open-weight figure matters mainly as a question: which model runs the work, and where does it run? The practical takeaway is that a vendor's figure remains a claim until you understand its underlying basis. Ask what the number counts, over what period it was measured, and whether it is reported gross or net before comparing it with anything else.

Which fields could AI reach next, and in what order?

A Google distinguished scientist offered an ordering that places structured fields first and fields built on tacit judgment last. In his framing, the speed of automation runs inversely to the entropy of a domain, meaning how noisy, unstable, and ambiguous its information happens to be. Mathematics and coding come first, followed by the hard sciences, financial markets, medicine, and law.

Economists are already making comparable predictions. Arpit Gupta of NYU Stern argued that the kinds of advances now visible in mathematics are headed toward economics and the other social sciences. Ethan Mollick of Wharton described two concurrent effects: published work will be reread and rejudged in ways its authors never anticipated, and novel discoveries will begin arriving far more rapidly than before.

Applied to a small business, that ordering sorts routine work fairly quickly. For a Fort Wayne or Auburn contractor, booking crews and drafting invoices sit nearer the structured end, where automation tends to arrive first and prove most dependable.

Diagnosing a failing system or negotiating a difficult job sits nearer the messy end, where human review should remain heavy the longest. The ordering works best as a diagnostic lens rather than a forecast, because it tells you where AI is likely to earn trust first.

Does this kind of AI work carry over to messy business tasks?

Not automatically, and certainly not without deliberate verification. The math results came from problems in which a correct answer can be verified mechanically, and that kind of structure fits reinforcement learning methods unusually well. Pedro Domingos of the University of Washington summarized the pattern as Moravec's paradox, meaning tasks that are hard for people can prove easy for machines when the rules are explicit and the answers can be checked.

Most small business work is considerably less tidy than a theorem. A quote depends on unspoken constraints, such as whether a crew can reach the backyard with equipment or how old the roof truly is. A reply to an upset customer depends on tone, history, and whatever was promised last spring. Those tasks are harder to verify than a proof, so the review step has to carry more of the weight rather than less.

Some cases still fit the pattern well. Routine work with a clear right answer, such as confirming an appointment, collecting missing details, or checking a price against a published rate sheet, makes a sensible starting point. Work built on relationships, unusual judgment, or sensitive commitments makes a poor starting point, at least until your review process has demonstrated its reliability on easier tasks first.

What does an AI teammate do with this kind of output?

A Hey Button AI teammate drafts, and a person approves. Under the draft-and-approve default, nothing goes out without your sign-off, and the access level you select establishes what the teammate may do on its own. The three access levels are read-only, draft-and-approve, and full access. Most businesses should begin with draft-and-approve and widen the scope only after the work has earned that trust.

Consider an illustrative workflow, not a customer result. Suppose a homeowner asks for a price on a spring cleanup. The teammate organizes the request, identifies the missing details such as the address and photos, and drafts a reply asking for them. The owner reviews that draft, then sends it or edits it first. Customer messages, spending, and sensitive changes remain behind the approval rules you establish during setup.

Setup begins with an intake conversation about how work arrives, which tools you already use, and where busywork accumulates. Go-live means writing clear rules for what the teammate may read, draft, and do independently, and what always returns to you first. You watch it handle real examples before anything touches a customer. The teammate continues to be tuned over time, and a client portal shows what it is doing and why.

For more on how errors get handled, read what happens when an AI teammate makes a mistake. Setup and ongoing support are quoted around your particular workflows and integrations, and any software subscriptions or usage charges are identified before you commit to anything.

If you want to begin with one narrow workflow and a review step that you control, the next move is brief. Start the Hey Button questionnaire and tell us which task consumes the most of your team's time each week. If you would rather understand the role before taking any step, read what is an AI teammate first.

Sources

  • The AI Daily Brief, "What Happens When AI Solves Your Life's Work," October 9, 2026 (episode, show notes, and transcript): https://aidailybrief.ai/e/2026-10-09
  • OpenAI mathematics release of 372 novel results and 722 supporting papers, published to a public repository and described in the episode. Numbers used: 372 results, 722 papers, about three hours of compute per result, and 90 of the top 500 open problems fully solved.
  • Financial Times report on OpenAI's annualized revenue, as described in the episode. Numbers used: about $50 billion at the end of September, the $68 billion investor-estimate figure, 77% run-rate growth, and 107% enterprise growth.
  • The Information subscriber survey, as described in the episode. Numbers used: 35% reporting multiples on AI spend, 8% reporting net negative returns, 22% hiring less because of AI, 72% using open-weight models, and 21% using Chinese open-weight models.

Frequently Asked Questions

Sample often at the start, then taper only after errors stay rare for several weeks. Pull a fixed share of routine messages each week, and raise that share the same day a sampled draft contains a wrong price or a factual error. Write the schedule down so every reviewer follows the same rule.
Your approval rules should say so explicitly. A sound practice is to flag the item and hold it for a person rather than guess at an answer. Write that instruction into the rules, so nobody has to remember it during a busy afternoon.
The steps stay the same, but the owner usually does more of the sampling. In a shop of five people, one person can often cover approvals for the whole team if the trigger list is short. Write that list down before the volume grows, not after the first complaint.
Under draft-and-approve, a person sees each message before it goes out, so a mistake is caught at that step. Errors still happen during review, which is why sampling sent messages against the approved drafts each week matters. Compare the two regularly to catch patterns early.
Ask who approves outgoing messages, what the AI may do without approval, how changes are recorded, and what happens when a model is slow or unavailable. Vague answers tell you more about the vendor than the name of the model does.
How often should I sample AI output for review?
Sample often at the start, then taper only after errors stay rare for several weeks. Pull a fixed share of routine messages each week, and raise that share the same day a sampled draft contains a wrong price or a factual error. Write the schedule down so every reviewer follows the same rule.
What should the AI do when it is unsure?
Your approval rules should say so explicitly. A sound practice is to flag the item and hold it for a person rather than guess at an answer. Write that instruction into the rules, so nobody has to remember it during a busy afternoon.
Does a five-person shop need a different review process than a larger firm?
The steps stay the same, but the owner usually does more of the sampling. In a shop of five people, one person can often cover approvals for the whole team if the trigger list is short. Write that list down before the volume grows, not after the first complaint.
Can an AI teammate send a mistake to a customer?
Under draft-and-approve, a person sees each message before it goes out, so a mistake is caught at that step. Errors still happen during review, which is why sampling sent messages against the approved drafts each week matters. Compare the two regularly to catch patterns early.
What should I ask a vendor about its review process?
Ask who approves outgoing messages, what the AI may do without approval, how changes are recorded, and what happens when a model is slow or unavailable. Vague answers tell you more about the vendor than the name of the model does.

Hey Button

Want a teammate that handles this for you?

Answer a few questions about your business and Hey Button will show you what an AI teammate would take off your plate, with you approving anything that goes out.

Start the questionnaireSee what teammates do
Lucas M. Button

Written by Lucas M. Button

Founder, Hey Button

Lucas builds AI teammates for small businesses across Northeast Indiana and writes about what works, what doesn't, and what to hand off first. More about Lucas