AI you can test, see, and undo
A pattern runs through today's announcements: the software industry has stopped presenting artificial intelligence as a feature you bolt on and started presenting it as a decision-maker you have to supervise. The tools shipped today are less about what the AI can do and more about whether you can test it before it faces a customer, watch it while it works, and undo it when it gets something wrong.
That shift matters to anyone choosing tools right now. For most of the past two years the question in a sales demo was "does it have AI". The better question, and the one these releases are built to answer, is "can I govern it". A system that acts on your behalf — answering a customer, matching a payment, writing a record — is only safe to adopt if you can see what it did and reverse it. Below are the five items from today that speak most directly to that question, in order of how much they should affect a buying decision.
Testing an AI agent before it ever answers a customer
If you are going to let software reply to customers without a person in the loop, you need a way to know how it behaves before it goes live, not after. That is what evaluation means in practice: you take the changes you are about to ship, run them against real or realistic conversations, and measure the outcome while the stakes are zero. Then you keep measuring once it is live, because a system that was right last month can drift as your product and your customers change.
Intercom put this exact loop into a product today, built around its support agent Fin. The company describes "a complete evaluation system for Fin" where "you can now test changes before they go live, roll them out with control, evaluate every live conversation, and have confidence in the experience Fin delivers" [1]. Notice the three stages in that sentence — before, during, after. The lesson for a buyer is not that one vendor shipped one feature. It is that automated customer contact should never be a switch you flip and forget. If a tool you are evaluating cannot show you how its AI performed on a sample before you commit, and cannot show you every live conversation afterwards, you are accepting a risk you cannot see. This is the practical side of a point we have made before about what AI should and should not do in your business: the boundary is not the model's cleverness, it is your ability to check its work.
Keeping an automated decision visible and reversible
The second item is smaller in scope and larger in principle. Xero updated its Automatic Bank Reconciliation, the task of matching the lines on a bank statement to the records in your books. The company says the feature "is designed to take the repetitive work out of bank reconciliation, making it faster while keeping every decision visible and reversible" [2]. Those last three words are the whole point.
Reconciliation is exactly the kind of work people want to hand to software: repetitive, rules-based, and tedious. But money matching is also where an unseen mistake compounds quietly for months. The right design is not "the AI did it, trust it". It is "the AI proposed a match, here is why, and you can undo it". That is the difference between automation you can audit and automation you have to take on faith. We have written separately about the audit trail nobody thinks about and about decisions automation should never make — both come down to the same test a buyer should apply here. When a tool automates a financial action, ask to see the record of what it changed and the single click that reverses it. If either is missing, the speed it offers is not worth the exposure.
Turning a spreadsheet into something your team can use
Google launched Sheets canvas, which it describes as "a new Gemini-powered capability that transforms your spreadsheets into custom, interactive, read-write applications using simple natural language prompts" [3]. Strip away the phrasing and the idea is familiar: a lot of businesses run on spreadsheets that have quietly outgrown themselves. The sheet started as a list and became the system of record for a process nobody formally built.
What is interesting is the direction of travel. Instead of asking a non-technical person to learn formulas and scripting, the tool lets them describe the interface they want in plain words and generates something interactive on top of the data. That lowers the cost of making a messy sheet usable. It does not change the underlying question, which is one we have covered in when a spreadsheet stops being enough: a nicer front end on a spreadsheet is still a spreadsheet underneath. For light, self-contained work it may be exactly enough. For anything that several people depend on, or that needs to connect to the rest of your tools, the read-write layer is a convenience, not a substitute for a real system. Treat it as a fast way to prototype, not as a reason to delay the decision you were already putting off.
Capturing a meeting without leaving it
Google also made its "Take Notes for me" feature available for in-person meetings. The company frames the problem plainly: "We know that balancing active participation with note taking during face-to-face discussions can be a challenge", and argues that "you shouldn't have to choose between staying present in the moment and capturing critical context for later" [4].
The tension is real and it is not only about meetings. Any record that depends on someone typing it up afterwards tends to be thin, late, or missing, because the person was busy doing the thing the record is about. Automated capture closes that gap — but only if the capture lands somewhere it will actually be used. A note that an assistant writes and then abandons in a document nobody opens is no better than no note at all. The value appears when the captured context flows into the place your team already works, which is the same discipline behind keeping customer data current: a record is only worth capturing if it stays where decisions get made.
How enterprise AI actually arrives — through people
The last item is a reminder that adoption is a human problem before it is a technical one. IBM announced a partnership with OpenAI, and the headline detail is about training rather than technology: IBM "plans to train and certify tens of thousands of consultants on OpenAI's technologies as part of this deal" [5].
That number tells you something about where the bottleneck sits. Capability is no longer the scarce resource; the scarce resource is people who can fit the capability to a specific business. For a company choosing tools, the practical takeaway is to weigh the help as heavily as the software. A powerful system that nobody on your side knows how to configure, govern, or connect will underperform a modest one that comes with people who understand your situation. This is part of what we mean by learning to choose software worth using — the product is only half of what you are buying.
Seen together, the day's releases are the same message from five directions. The industry is past arguing about whether to use AI and into the harder, more useful work of making it accountable: tested before it ships, visible while it runs, reversible when it errs, and backed by people who can make it fit. Those are the things to ask a vendor about now, because they are the things that decide whether the tool helps you or quietly costs you.
Sources
- [1] Announcing Evals and Releases: Evaluate Fin before, during, and after you go live — Intercom
- [2] What's new in Auto Bank Reconciliation — Xero
- [3] Use Sheets canvas to visualize data in custom, interactive mini-apps — Google Workspace
- [4] Take Notes for me for in-person meetings is now available — Google Workspace
- [5] IBM partners with OpenAI to bolster enterprise AI push — TechCrunch