Testing before you trust, and pricing before you assume
Three announcements today share a single idea: a sound decision about software rests on evidence you can see, not on a price tag or a default setting. One is about testing an AI model before you trust it with real work, one is about learning what your own product is worth, and one is about keeping unwanted items out of a calendar.
Test an AI model on your own data, not on its reputation
There is a quiet assumption in a great deal of software buying: the dearest option must be the safest one. For ordinary tools that assumption is often wrong, and for AI models it can be costly. A model's price reflects how it was built and how much the vendor has chosen to charge, not how well it will do the particular job you have in mind. The only honest way to know how a model performs on your work is to run your work through it and score the results.
That is what Rippling did. The company ran 2,100 scored agent runs per model against real payroll data, and reported that the cheapest model tied the most expensive one [1]. The figure in the headline is not the part that should change your behaviour. The method is. They used their own data rather than a public benchmark. They scored every run, so the comparison could be repeated instead of argued over. And they ran enough trials that a single lucky answer could not tip the result.
For a business choosing an AI tool, the lesson carries over without much translation. Before you commit to a model, or to a product built around one, decide what a good answer looks like for your task, collect a sample of your real inputs, and run them through. Settle in advance on what you are optimising for. The Rippling result came with a condition attached: the cheaper model matched the dearer one when speed did not matter. If response time matters to you, the ranking may well move. A number only changes a decision once you know what the number measures, a habit we have written about in numbers that change a decision.
This is also why the freedom to swap models, or to bring your own, is worth looking for when you choose software worth using. A tool that ties you to one model quietly takes the test out of your hands, and a test you cannot run is a claim you have to take on faith.
Price is a hypothesis, and you test it on the next deal
Most prices are set once and then left alone, because changing them feels risky. The risk people picture is upsetting the customers they already have. The advice from SaaStr today separates those two worries cleanly. If you want to move upmarket, the suggestion is to double your pricing on your next deal, and pointedly not on your existing customers [2].
The reasoning is worth drawing out, because it is a general principle wearing the clothes of a tactic. Your current customers bought at a price they agreed to, and raising it on them breaks a bargain they entered in good faith. A new customer has made no such bargain. The next deal is therefore a low-cost experiment. If the higher figure holds, you have learned that your product is worth more than you were charging for it. If it does not hold, you have lost one negotiation rather than a base of loyal accounts. Either way, you come out of it knowing more than you did going in.
The caution is that a price is a signal as much as a figure. Charging twice as much for a larger customer only works if the value on offer has grown to match it — more scope, more support, more certainty that the thing will keep working. This connects to a failure we have described before in how pricing pages fail: a price the buyer cannot tie to anything concrete reads as arbitrary, and an arbitrary price gets argued down. Test the number on new deals by all means, but make sure the thing standing behind the number has moved as well.
A calendar setting that is really a governance decision
The third item looks small and is not. Google is adding the ability to block the senders of unsolicited event invitations in Google Calendar [3]. On the surface it removes a nuisance. Underneath, it is about who gets to put something on your team's time without being asked.
A calendar is an open door by default. Anyone who knows an address can place an entry on that person's day, and in most organisations that entry arrives looking exactly like a real meeting. Unsolicited invitations are not only an annoyance. They are a way to crowd a schedule, to probe whether an address is live, and to slip a link in front of someone during a busy morning. A block control turns the calendar from something anyone can write to into something the owner can curb.
The wider point for anyone choosing tools is that the small settings are where day-to-day control actually lives. When you assess a product, it is tempting to judge it by its headline features. The features that decide how much noise your team absorbs, how much unwanted contact gets through, and how much tidying each person does by hand are usually the quiet ones — a block list, a default that favours the recipient, a way to filter what reaches an inbox or a diary. A feature that keeps unwanted things out deserves as much attention as a feature that lets new things in, because the first one is what protects the hours your team has already committed.
The thread, and what to do with it
Put the three together and a single working rule emerges. Decide on what you can see. Test the AI model against your own inputs rather than its billing tier. Test your price on a new deal rather than guessing at your worth. And give each person real control over what lands on their calendar, instead of leaving the door open because it came that way. None of these is a grand strategy. Each is a small, checkable step that replaces an assumption with a fact, and the value of a fact is that it holds up when someone pushes back on it.
This is the posture we try to build into 360REV: let people see the numbers behind a decision, keep the controls in the owner's hands, and avoid defaults that quietly decide things on their behalf. The tools on offer change every week. The discipline of deciding on evidence does not.
Sources
- [1] Rippling Ran 2,100 Scored Agent Runs Per Model on Real Payroll Data. The Cheapest Model Tied the Most Expensive One. — SaaStr
- [2] On Your Next Big Deal? Double Your Pricing. — SaaStr
- [3] Managing unsolicited event invitations with user blocking in Google Calendar — Google Workspace Updates