What AI Should and Should Not Do in Your Business
Every year or two, business software goes through a phase where one new capability gets promised as the answer to nearly everything. Right now that capability is AI, and the promises range from reasonable to absurd, often in the same sentence. Somewhere underneath the noise is a genuinely useful set of tools — and a genuinely important set of limits. Sorting one from the other is worth doing carefully, because getting it wrong in either direction has a real cost: over-trusting automation hands a judgment call to something that can't make it responsibly, and under-trusting it means paying a person to do tedious work a system could do just as well, and probably more consistently.
What "AI" means in a business context
It helps to drop the word "AI" for a moment and ask what the software is actually doing. Most of what gets sold as AI in day-to-day business tools falls into a few concrete jobs: drafting text (an email, a summary, a first pass at a reply), classifying or routing something (is this support ticket urgent, which list does this lead belong on), extracting structure from unstructured input (turning a rambling voicemail transcript into three bullet points), and surfacing a pattern a person would otherwise have to notice by hand (this customer hasn't ordered in four months, unusual for them). None of that is exotic once it's named plainly, and none of it requires believing the system understands the business the way a person does. It's pattern-matching at a scale and speed a person can't match, applied to specific, bounded tasks.
That framing matters because it draws the line in the right place. A tool that drafts a reply is doing something fundamentally different from a tool that decides whether to issue a refund, terminate a contract, or tell a customer their claim is denied. The first produces a draft a person can accept, edit, or discard in five seconds. The second makes a decision that changes someone's outcome, often with no easy way to reverse it once it's been acted on. Confusing the two — treating a drafting tool as a decision-maker, or a decision as if it were just a draft that got shipped without review — is where most of the real damage happens.
Where the case for automation is strong
There's a solid, boring, entirely defensible case for AI in the categories where it's doing pattern-matching on well-defined, repetitive, low-stakes work:
- Summarizing a long call or thread so a person doesn't have to re-read all of it before responding.
- Drafting a first version of a routine reply, which a person then reads before it goes out.
- Flagging anomalies for a human to look at — a spike in refund requests, a customer who suddenly stopped opening emails — rather than acting on them directly.
- Sorting and routing incoming requests to the right queue or the right person, where being wrong just means a slightly slower handoff, not a wrong outcome.
- Filling in a first draft of a report from data that already exists, so a person edits and interprets rather than starting from a blank page.
The volume argument here is real. Stanford's 2025 AI Index reports that business use of AI accelerated sharply through 2024, with a large majority of organizations now using it in at least one function [1], and Deloitte's research into enterprise AI adoption describes the more advanced organizations as ones that have learned to let AI handle end-to-end execution on narrow, well-understood tasks while reserving judgment, exception-handling and oversight for people [2]. That's a genuinely useful division of labor, and it's the one worth aiming for: AI does the parts of the job that are repetitive and reversible, a person does the parts that require judgment and carry consequences.
Where it should not be making the call
The harder, more important half of this is naming where a human needs to stay the one actually deciding, not just the one who could theoretically catch a mistake after the fact. The U.S. National Institute of Standards and Technology's AI Risk Management Framework is direct about this: human oversight is a core expectation, and organizations are expected to define who is responsible for monitoring and guiding an AI system, understand its limitations, and retain the ability to override or disengage it when something looks wrong [3]. That's not a compliance footnote — it's a description of where the actual risk in AI deployment lives, and it lives specifically in the high-stakes, hard-to-reverse decisions, not in the low-stakes drafting work.
Concretely, that means keeping a person in the loop, and not a nominal one, for anything that:
- Determines whether someone gets paid, refunded, billed, or charged a fee.
- Ends a relationship — cancels an account, terminates a contract, fires a warning shot at a customer.
- Makes a claim about facts that could be wrong in a way that damages trust if it is — a legal statement, a medical suggestion, a promise about what a product can do.
- Can't be easily undone once it's acted on.
It's worth adding one warning of our own on top of the research: a human in the loop only helps if the human is actually looking. MIT Sloan Management Review's account of a well-known 1983 near-miss makes the underlying point plainly — a Soviet early-warning system wrongly reported incoming missiles, and the officer on duty chose to question the machine rather than act on its recommendation immediately; the piece's conclusion is that a human can add real value by scrutinizing a system's results before acting on them [4]. Note what that requires: actual scrutiny, not just a person's name attached to a decision. An approval step that exists on paper, that everyone clicks through without reading, is not meaningfully different from having no approval step at all — and it's worth being honest with yourself about which one your business actually has.
The test that actually works
A useful, simple test for any specific task: if the AI gets this wrong, what happens next? If the answer is "someone reads a slightly worse draft and fixes it before it goes out," that's a low-stakes task and automation is probably a good fit. If the answer is "a customer gets an outcome they can't easily appeal, or the business makes a promise it can't keep," that's a high-stakes decision, and the AI's job there is to prepare information for a person, not to make the call.
This is the same distinction we described from the customer-record side in our piece on what a CRM is actually for: a CRM can hold every fact about a customer, and an AI on top of it can draft the follow-up email or flag that the account looks at risk, but deciding whether that customer gets a discount, a warning, or a phone call from a manager is still a judgment a person should make, informed by facts the system surfaced rather than replaced. The same logic runs through the point we made about why your tools do not talk to each other — an AI system is only as good as the data it can see, and even a well-designed automation step making a low-stakes decision on an incomplete picture can produce a confidently wrong answer, which is arguably worse than an honest "I don't know."
What this looks like in practice
None of this requires treating AI with suspicion in general. It requires being specific, task by task, about which side of the line a given piece of work falls on, and building the review step to match: light-touch for drafts nobody will act on unread, and a real, accountable human decision for anything that changes an outcome for a customer or the business. Getting that split right is less exciting than either "AI changes everything" or "AI can't be trusted with anything," and it's also the only version of the answer that actually holds up once you're the one who has to live with the result.
Sources
- [1] The 2025 AI Index Report — Stanford HAI
- [2] The State of AI in the Enterprise — Deloitte
- [3] AI Risk Management Framework — NIST
- [4] Justifying Human Involvement in the AI Decision-Making Loop — MIT Sloan Management Review