Weco got a top AI model to improve its own AI. The same review habit can make the AI your business already runs cheaper and better each month.
A research lab called Weco just published something that sounds like science fiction. They built an AI system that improved its own AI, with no person in the loop. Over eight days it produced seven better versions of itself, each one doing the same work for less. They’re calling it the first real evidence of recursive self-improvement.
If you run a business, the headline isn’t the part that matters to you. You’re not trying to build an AI that rewrites itself. The useful bit is quieter, and you can copy it this quarter. Weco pointed one very capable AI model at the everyday agents doing the work, told it to find cheaper and better ways to do the same jobs, and kept only the changes that genuinely helped. That habit, not the science, is what your business can borrow.
Most teams set up an AI tool once and never look at it again. It keeps doing the job the way it did on day one, and quietly costs more as you use it more. A regular review, run by a top model against a record of what your AI actually did, is how you fix that.
What Weco actually did
Weco’s system runs on two loops. The inner loop is the workhorse. It’s an ordinary AI agent doing a task, and it runs on a cheaper, faster model because it has to run many times over. The outer loop is the reviewer. It uses the most capable model they had, a top Claude model, and its job is to look at how the workhorse performed and propose a better version of it.
The clever part is the budget. Every proposed change had to work under a fixed cost, so the system couldn’t win by simply spending more. A change only counted if it did the same work for less, or better work for the same money. About nine in ten proposals were rejected. The ones that survived were real gains.
One of those gains is worth understanding, because it’s the kind of thing a review of your own AI would find. The system worked out that its agents were carrying far too much text in every prompt. It rebuilt how they handled that and cut the prompt size by about 16 times against the naive approach of stuffing in the full history. The tokens it saved were put straight back into doing more work. The reviewer’s own running cost, by the way, was a small fraction of the total. A top model doesn’t need to run often to earn its keep.
The part that applies to your business
You already have AI doing repetitive work somewhere. It might be drafting replies, sorting documents, or answering the same customer questions. Left alone, it does that job the same way forever. Weco’s method turns into two plain habits for you.
First, log everything. Every time your AI does a job, record what went in, what came out, and what it cost. Most tools can do this already. Claude, and the coding and agent tools built on it, can copy their full history into one central place in a structured way. The point is to have an honest record of what your AI is actually doing, not what you assume it’s doing.
Second, review it with a top model. Once a month, point one of the most capable AI models at that record and ask a simple question: where are we doing this the slow or expensive way, and what would do the same job for less? You’re renting the frontier for an afternoon to improve work that runs all month.
Picture a growing online homewares retailer. A year ago they set up an AI agent to draft replies to order and returns questions, the sort of everyday work AI agents for retailers handle well. It works. It also runs thousands of times a day, and nobody has looked at it since.
Where the savings come from
When that retailer finally logged a month of replies and had a top model read them, the review found something dull and valuable. Every reply was carrying the shop’s full returns policy in the prompt, even for questions that had nothing to do with returns. Trimming that cut the tokens used on each reply. On its own, a smaller prompt saves a fraction of a cent. Run it across 3,000 replies a day, every day, and the fraction turns into real money. Those figures are only an example to show how the saving adds up, not a measured result.
The review found more than cost. It spotted a common question the agent kept passing to a person when it could safely handle it, and a place where its wording confused customers. So the same afternoon that made the agent cheaper also made it a little better.
None of this pays off in a dramatic first month. A couple of hundred dollars of frontier-model time might find one or two savings to start with. But you do it again the next month, and the next. The lessons build on each other, and as you put AI agents into more corners of the business, the same habit spreads the gains across all of them.
Two honest limits keep this grounded. A review can only find savings in what you record, so the boring discipline of logging is the real work here, not the model. And a model reading your logs won’t run the business for you. It suggests, you decide what to change, and you carry the risk if a change goes wrong. Keep a person on the final call.
What this means for you
The technology in Weco’s report is years ahead of what most businesses need, and that’s fine. You don’t need an AI that rewrites itself. You need the habit underneath it. Keep an honest record of what your AI does, and every month spend a little on a top model to find the same work done cheaper and better. Small savings, repeated across every AI task you run, add up to a lot.
If you’d like help setting that up, from the logging through to the monthly review, take our AI Roadmap Interview. You’ll talk through where your team already uses AI and where it’s costing more than it should, then get a plan for where to start.
