Open weight models on your own hardware aren't about cheaper tokens. They're about client data that never leaves, and a price nobody else can change.
On 24 July, a group of technology companies including Meta, Microsoft, and Nvidia signed an open letter arguing against broad restrictions on open weight models. Open weight models are AI models anyone can download, inspect, modify, and run on their own hardware. The letter’s case is that a country’s AI position rests on how widely AI spreads through ordinary businesses, not on who happens to own the single best model. That argument will be settled well above your head. From where an Australian business owner sits, the useful part is simpler than the politics. A closed AI model is a place you visit. An open weight model is a file you hold on a hard drive, and once you hold the file, a different set of choices opens up.
The usual reason given for running one in-house is extremely cheaper tokens. That’s the wrong reason, and the numbers don’t back it.
The token bill isn’t the argument
Picture a professional services firm working through around 150 client files a month. Identity documents, statements, and correspondence, all of it needing to be read, sorted, and filed by someone. Push that volume through a cloud API and the token cost lands somewhere near $34 to $350 a month depending on the model you pick. A $5,000 local server takes a while to pay off on savings that size. Where the money genuinely gets saved is AI against manual labour, and that saving is there whether you rent the model or own it. Renting versus owning is a separate question, and what owning buys you is control.
A large part of that is price certainty. Today’s API prices are set by labs still competing hard for market share, and nobody can promise those prices hold once that competition settles or investors start asking for margins. We’re not saying they’ll rise. We’re saying you’d have no say if they did. Hardware you’ve already bought doesn’t get repriced, and that’s what sovereignty means in practice, which is the ability to say no to a price rise or a change of terms and keep operating anyway.
There’s a scale point sitting underneath that. A token bill grows every time your business grows, while a box in the server cupboard costs the same whether you push a hundred documents through it or a hundred thousand. Once your team stop watching their subscription usage, they start actually using the thing, and that’s where the value turns up.
What actually leaves the building
For that firm, though, cost isn’t the sharpest issue. What’s already happening is. Staff are busy, so some of them are already pasting client text into whatever chat window is open in the next tab. That’s shadow AI, and it isn’t hypothetical. Banning it doesn’t fix it either, because the same behaviour simply moves onto personal phones where nobody can see it at all. Meanwhile the obligations are tightening around them. Australia’s AML/CTF Tranche 2 reforms bring roughly 80,000 new entities under strict data handling rules, which means far more firms now carrying serious customer due diligence duties. And under section 16C of the Privacy Act, disclosing personal information to an overseas provider can leave you liable for their privacy breach. You don’t have to be the one who got breached to wear it.
Put those together and the firm’s real exposure is one staff member pasting a client’s identity document into a public API on a Tuesday afternoon. A model running in the office server cupboard removes that path entirely, because the document never leaves the building. It doesn’t remove your obligations, and we won’t pretend otherwise. You still owe your clients the same duty of care, and you still need to know where every piece of your data sits. What it removes is the offshore disclosure and the unmonitored browser tab.
What the setup actually looks like
Think of an open weight model as an engine rather than a car. Someone else built the engine, and you still need the chassis around it. In practice that’s three parts: the model file itself, a program to run it, and something your staff actually click on.
You can test the whole idea on a laptop you already own for nothing. A Mac Studio class machine at around $5,000 runs mid-sized models well enough for a team. A dedicated GPU setup runs the 120 billion parameter models quickly, which is the tier our example firm would need to work through client files at that volume. Built and implemented properly, expect that to land around $25,000 to $30,000. That figure covers a full setup rather than the hardware alone, so treat it as a planning range and not a quote. After that, the ongoing cost is the power bill and maintenance, because what you’re running is a computer in a room. It’s a capital purchase rather than a subscription, which for a lot of businesses is simply an easier thing to finance. When you compare that expense to the ever increasing, ongoing cost of Claude max subscriptions for your entire team, (~$300 per user per month), it becomes pretty feasible, pretty fast.
And you don’t have to run it yourself. A 20-person firm has no business patching and securing AI infrastructure in-house, and it doesn’t need to, because a managed service provider can look after it the way they look after everything else. You’re paying one already, so the added cost is miniscule next to what the setup saves. Be honest about the ceiling too. Frontier cloud models still beat local ones on novel, complex reasoning, and that gap is real. The sensible split is cloud for the hard thinking and local for the sensitive, high-volume, repetitive work. There’s no sense chartering a jet to go and fetch the milk, which is what paying frontier prices to sort routine documents amounts to.
What this means for you
The question worth asking isn’t what AI costs per token. It’s which of your data can never leave, and what it would take to keep it in. For the firm in our example that’s the client identity documents and any client IP covered by a service agreement, nothing else really changes, the cloud still does the hard thinking, the staff stop pasting into public tabs, and the monthly bill stops moving. For most firms the honest answer is a smaller piece of hardware than they expect, plus a clearer conversation with whoever supports their IT. Start by finding out what your team is already using and what they’re putting into it, because you can’t scope this until you know.
That scoping is the work we do, and it’s why we hand it over as one implementation rather than a stack of separate projects you’re left holding together. We work out which jobs belong on your own hardware and which are better left in the cloud, then we build both and make sure your provider can support what we leave behind. If you want to see where your data currently goes and what it would cost to bring the sensitive parts home, our AI strategy work starts there, and the Technology Partner page to find out how working with a managed services partner is like in Generation AI.
