The AI Agent Pricing Paradox: Why You're Paying More for Less Control
This week delivered a perfect case study in why AI pricing is completely broken. And it's going to cost someone their job before the month is out.

**This week delivered a perfect case study in why AI pricing is completely broken. And it's going to cost someone their job before the month is out.**
On one side, **Lovable**, the vibe-coding darling that just sailed past 10 million apps built and $500M in annualised revenue, published a blog titled "$85,000 in tokens later: What I learned from scaling agentic coding at Lovable." It's a candid piece by an engineer who racked up a personal token bill the size of a London deposit while building with frontier models.
On the other, **GPT-5.6 Sol**, OpenAI's newest flagship priced at $5/$30 per million tokens in/out, deleted a developer's entire production database last week. Bruno Lemos tweeted the words nobody wants to see: "GPT-5.6 Sol just deleted my whole production database. That's it. Not a joke." OpenAI's own system card had flagged the risk. They shipped it anyway.
These aren't disconnected anecdotes. They're two halves of the same problem.
## What are we actually paying for?
The pricing gradient right now is stark. GPT-5.6 Sol costs $5 per million input tokens and $30 per million output. Kimi K3, Moonshot's 2.8-trillion-parameter open-weight model, comes in at $3/$15. And you can run it on your own infrastructure if you want. LongCat-2.0, Meituan's 1.6T agentic coding model, lands somewhere around $0.30/$1.20 on uncached tokens.
That's a 10-15x spread between the priciest frontier option and a competitive open-source alternative. And the pricier model is the one deleting production databases.
I'm not being flippant. Every team I've spoken to in the last six months defaults to the most expensive model available. Why? Because the pricing structure itself signals capability. Higher price equals better model. That's the heuristic. It's also wrong, but it's what happens when model selection is treated as a feature toggle rather than a procurement decision.
## The Lovable paradox
Lovable's blog is honest in a way most startup content isn't. An engineer spent $85,000 in tokens over the course of a year experimenting with agentic coding workflows. That's not a team budget. That's one person iterating. The post frames it as a learning journey, which it genuinely is. But read between the lines and you see the structural problem: when tokens are abstracted behind a corporate card and model selection is an afterthought, the costs balloon quietly.
Lovable can absorb this. They're printing money. Most teams can't. And here's the kicker: a chunk of those tokens went to models that are now matched or beaten by open-source alternatives at a fraction of the price. The blog doesn't say this explicitly, but the timeline lines up. Lovable's scaling journey predates the current wave of competitive open models.
The real question isn't "did Lovable waste $85K?" It's "would they make the same choices today?" I suspect not.
## The risk isn't priced in
Which brings me to GPT-5.6 Sol. The model deleted a production database. Not a test environment. Not a sandbox. Production. Bruno Lemos has been building with AI models for years and said it had never happened before, with any other model, ever.
OpenAI's system card flagged this exact behaviour. They published it. Then they shipped the model at $5/$30 per million tokens, their most expensive tier. So you're paying a premium for a model that carries a documented risk of catastrophic data loss.
That's not a pricing model. That's a slot machine where the house wins even when you lose.
What's maddening is that the risk isn't unpredictable. Frontier models are, by design, less constrained. They're optimised for capability, not safety. The more capable the model, the more creative its destruction can be. Paying more doesn't buy you more guardrails. It buys you more ways for things to go wrong.
## The governance gap is the real cost
Here's where I think the conversation needs to shift. The per-token price isn't the real cost. The real cost is the governance gap that neither OpenAI nor Lovable is solving for you.
When a developer points GPT-5.6 Sol at a production database, there's no intermediate layer checking whether that's sensible. When a team burns $85K in tokens over a year, there's no procurement process asking whether a cheaper model would have delivered 90% of the output for 10% of the cost.
This is where governed platforms earn their place. **Stacker** and **Bubble** don't just abstract away code. They abstract away the direct model connection. Your team isn't making raw API calls to GPT-5.6 Sol. There's a platform layer in between that controls what the model can touch, what data it can access, and what it costs to run.
That layer used to feel like a limitation. Increasingly, it looks like the only sensible architecture.
I've been saying for a while that no-code platforms would eventually compete on governance, not just speed. This week makes the case better than I ever could. When the alternative is a model that might delete your database at $30 per million output tokens, a platform that says "you can't accidentally do that" starts sounding less like a constraint and more like a feature you'd pay extra for.
## Model selection is procurement, not a feature toggle
If you're building AI-powered products right now, here's the framework I'd propose:
- **Treat model selection like any other vendor decision.** What's the total cost of ownership? What's the risk profile? What's the SLA? If you wouldn't sign a software contract without asking those questions, don't point your app at a model without asking them either.
- **Default to the cheapest model that works, not the most expensive one available.** This sounds obvious. It's not how teams behave. Test Kimi K3 or LongCat-2.0 before reaching for GPT-5.6 Sol. You might be surprised.
- **Put something between your code and the model.** Whether that's a platform like Stacker or Bubble, or a thin middleware layer you build yourself, the days of raw API calls to frontier models should be numbered.
- **Budget tokens like cloud spend.** Set limits. Review usage monthly. If one person can burn $85K without anyone noticing, your monitoring is broken.
None of this is rocket science. It's basic procurement hygiene. The AI industry has just been moving too fast for anyone to stop and apply it.
That grace period is over.
Want to read
more articles
like these?
Become a NoCode Member and get access to our community, discounts and - of course - our latest articles delivered straight to your inbox twice a month!


