What AI agents can actually do for your business today

· 6 min read
AI-generated image: What AI agents can actually do for your business today
AI-generated image

Most of what productivity-software makers published today circled a single question: not whether AI agents can act, but how much of the work they can be trusted to finish without a person watching. The answers that arrived were more measured than the slogans of a year ago, and for anyone choosing tools that restraint is worth more than any promise.

An agent, in the sense everyone now uses the word, is software that can take several steps on its own to reach a goal — read a ticket, look something up, decide what to do, and do it. The interesting part of every announcement below is the same: where the line sits between what the machine finishes alone and what it hands back to a person. That line is the thing you are actually buying. We have written before about decisions automation should never make, and today's news is a useful test of whether the industry is drawing that line in roughly the same place.

How much support an AI agent really closes

The most useful number of the day came from a review of real resolution rates across thirteen vendors. "Resolution rate" sounds precise but hides a choice: does it mean the customer's problem was solved, or only that the conversation ended without a human stepping in? When you compare tools, insist on the second, stricter reading, because the first is easy to inflate.

The headline finding is that the best-performing AI support agents resolve about 70% of issues, and the median across vendors is 48% [1]. Read those two figures together before you read either one alone. A median near half means that, for a typical business adopting a typical tool today, a person still handles the other half — and the half the machine keeps tends to be the repetitive, well-documented half, while the hard, angry, or unusual cases flow to people. That is not a shortcoming to hide; it is the shape of the work. The reviewer frames the whole exercise plainly: "It's a good test of just exactly where the latest models and agents are" [1].

For a buyer, the practical lesson is to treat any single vendor's resolution claim as a range, not a guarantee, and to ask what counts as resolved. A tool that quietly closes tickets to raise its score is solving the vendor's metric, not your customer's problem.

Letting an agent act without letting it improvise

If the first story is about how much an agent can do, the second is about how to let it do anything safely. Salesforce published a piece on giving AI autonomy without, in its words, letting it run wild. The mechanism it describes is worth understanding because it is becoming the standard pattern: pair the part of the system that is good at open-ended reasoning with a separate part that enforces fixed rules.

The article explains that "pairing a probabilistic model with deterministic logic gives agents the ability to take action — without giving up the guardrails" [2]. Those two words carry the whole idea. A probabilistic model is the language model: flexible, creative, and occasionally wrong in confident ways. Deterministic logic is ordinary code: it does the same thing every time and can refuse. The model proposes; the rules dispose. A refund over a set amount, a change to a signed contract, a message to a regulated customer — these are the places where you want a rule that the model cannot talk its way past.

When you evaluate an agentic feature, ask the vendor which decisions are governed by fixed rules and which are left to the model's judgement. If the answer is "the model handles everything," you are being sold flexibility where you wanted a guardrail. This is the same principle we covered in what AI should and should not do in your business: the value is not in how much the model is allowed to decide, but in how clearly the boundaries are drawn.

Folding agents back into the workflow

Zapier made a change today that reflects where thinking about agents has moved. It is retiring Zapier Agents as a separate product and folding its capabilities into the ordinary automation editor. In the company's description, "Every capability that made Zapier Agents powerful — tool calling, reasoning, autonomous action — now lives inside a single AI step in the Zap editor, using AI by Zapier" [3].

The trade-off here is worth naming. A standalone agent product invites you to think of the agent as a separate thing that works on its own. A single step inside a workflow invites you to think of it as one more action in a chain you already control — it runs when the steps before it run, hands off to the steps after it, and sits where you can see it. For a business, the second framing is usually easier to reason about, because the agent is bounded by a process you designed rather than roaming free. The cost is that very open-ended, multi-step autonomy may fit less naturally inside a fixed chain. Which matters more depends on how much you want the machine to improvise versus follow a path you set.

Building for when the author is a machine

GitHub described rebuilding the infrastructure underneath its version-control system to prepare for a future where far more code is written by agents than by people. In its framing, the company is "rebuilding GitHub's Git infrastructure while GitHub keeps running, creating a foundation for agent-scale software development" [4].

Most readers will never touch this layer directly, but the signal is useful. The phrase "agent-scale" assumes a volume of automated activity large enough that the plumbing built for human-paced work needs replacing. When a company rebuilds its foundations on that assumption — and does it while the service stays live, which is the harder way — it is telling you where it expects the load to come from. If you build software, or buy from people who do, the practical consequence is that the tools around code are being reshaped for machine authors, and the review and approval steps that catch mistakes matter more, not less, as the volume rises.

When writing code by hand becomes the exception

The sharpest statement of the day came from Basecamp's David Heinemeier Hansson, who used a talk at a developer conference to argue that hand-writing code is on its way to being the exception rather than the rule. He recounts: "Two weeks ago at Rails World, I told my fellow programmers that it's time to put down the pencils" [5]. His reasoning is economic, not sentimental — "We're not going to write the vast majority of code by hand any longer" [5].

Take this as one experienced practitioner's forecast rather than a settled fact, and weigh it against the day's more measured numbers. A median resolution rate near half [1] and a deliberate emphasis on guardrails [2] both suggest the handover to machines is partial and supervised, not total. The two pictures are not in conflict. The volume of work a machine produces can rise sharply while the share it finishes entirely on its own stays well short of everything. The job that remains for people shifts from doing the work to directing and checking it.

What ties the day together

Read across all five, and the thread is clear: the industry is settling into a view where agents do a great deal but finish less than the marketing implied, and where the safe designs are the ones that keep a human and a set of fixed rules in the loop. That is a healthier place to be choosing tools than a year ago, because it lets you ask better questions. What share of the work does this finish alone, measured honestly? Which decisions are governed by rules the model cannot override? Where does the handover to a person happen, and can I see it?

If you are comparing options this week, carry those three questions into every demo. For more in this series, see yesterday's briefing. The vendors making real progress will answer plainly; the ones selling a slogan will change the subject.

Sources

  1. [1] How Much of Your Customer Support Can AI Really Resolve? The Best Get About 70%. The Median Is 48%. Here's the Real Data From 13 Vendors — SaaStr
  2. [2] Trustworthy, Explainable, and Accountable: How to Give AI Autonomy Without Letting It Run Wild — Salesforce
  3. [3] Zapier Agents is now AI by Zapier — Zapier
  4. [4] Building Git infrastructure for agent-scale development — The GitHub Blog
  5. [5] Over my dead pencil — David Heinemeier Hansson (HEY World)

The 360REV newsletter

What is actually changing across productivity software, written for operators and cited to sources. No more than one email a day.

Double opt-in — we send one confirmation link and nothing else until you click it. Unsubscribe from any edition. We never sell or share your address.