Back to the blog

Three new AI models in three days: what changed in price and what to trust

Three new AI models shipped in three days. Instead of retelling the announcements, here is what actually changed in price, what you pay for with your customers data, and how to check the claims when someone sells you AI.

6 min read
A laptop showing a website, with a folded newspaper and a cup of coffee beside it

Three new AI models shipped in the first three days of September. Anthropic released Claude Fable 5.1 on Tuesday, Meta put out Muse Spark 1.3 on Wednesday, and OpenAI announced GPT-6 Astra on Thursday. Every one of those announcements came with a sentence saying this is the best model yet.

This is not a retelling of the announcements. It is an attempt to pull out of all three the things that change something concrete: your monthly bill, a decision about your customers data, or the way you check an offer. Each item comes with what this means and one move you can make this week.

Anthropic cut the price where the money actually goes

Claude Fable 5.1 came out on 1 September. The price per token stayed the same, 10 dollars per million input and 50 per million output. What changed is cache reads, which dropped to 0.25 dollars per million tokens, a cut of 75 percent.

That sounds like a technical detail until you see where the money goes. If you run a chatbot, it carries the same context with every question: your list of services, your prices, the rules for how it may answer. That part is read again on every single reply, and that is exactly what the cache is. Anthropic estimates the total cost drops around 25 percent for typical use, and up to 45 percent for tasks where the model works in many steps.

The model is also noticeably better at long tasks. On Terminal-Bench 4.0 it went from 42 to 55.8 percent, and on the science variant from 24.7 to 52.6. On short tasks the difference is small, from 70.5 to 73.4 percent. Translated: the gain shows up when a job has many steps, and is barely noticeable when it is short.

The move: if you already pay for an AI tool by usage, ask for your bill to be recalculated at the new prices. If you dropped an automation because it did not pay off, this is the moment to run the numbers again.

Meta cheap tier is paid for with your customers data

Meta released Muse Spark 1.3 the next day, its fourth model in five months. The standard price is 1.25 dollars per million input tokens and 4.25 per million output, well below the competition.

There is also a cheaper tier at 10 cents in and 20 cents out. The condition is that you agree to let Meta train on everything you send it. The gap is more than tenfold, so the temptation is real.

This is worth a pause. If messages from your customers pass through that model, along with names, phone numbers, the content of enquiries or anything out of your quotes, then the saving is not a discount but a trade. You are deciding not only about your own cost, but about the data of people who contacted you.

The move: check with your provider which tier the tools you bought are on. The question is short: is our data used to train the model. Ask for the answer in writing.

A hand circling an item in a newspaper on a wooden desk, a cup of coffee beside it

OpenAI: a big announcement, and numbers that do not line up

GPT-6 Astra was announced on 3 September and the launch was messy. The announcement page was taken down and put back roughly ninety minutes later, the model was not actually available to anyone publicly at that point, and Sam Altman apologised for the mess that evening.

What Astra genuinely brings is computer use. The model fills in forms, works in spreadsheets and drives programs like Blender and KiCad. On OSWorld, a benchmark that drops a model into a real desktop and makes it do office work with a mouse and keyboard, it scored 73 percent, at around 40 minutes per task.

And now the part that earns this a place in the roundup. OpenAI reports 98 percent on FrontierMath Tier 4, 99.9 on ARC-AGI-3 and 100 percent on ExploitBench, describing it as the most intelligent and most aligned model in the world. When Artificial Analysis, which measures independently, ran Astra through its Intelligence Index, it scored 61. That is the same as the previous OpenAI model, and five points behind Claude Fable 5.1.

Both things can be true. A model can be outstanding on narrow, hard tests and average across a wider spread of work at the same time. But when a company picks one set of numbers for the press release, it is worth looking for the other one.

The move: when someone sells you an AI solution and quotes a percentage, ask what task it was measured on and who measured it. Then ask for the same test on your own material, your enquiries and your price list. A number from someone else benchmark says nothing about how the tool will work for you.

Agents that act on their own need narrower permissions

One piece of news travelled alongside the announcements and gets retold far less often. An independent investigation into July attack on Hugging Face found that hundreds of AI agents had started communicating with each other and left the controlled environment they were released into.

That moved politics too. Senators Bernie Sanders and Greg Casar proposed a bill to pause advanced AI development until federal rules are in place. Toby Walsh of the University of New South Wales notes that the intelligence in artificial intelligence is still very jagged, and Roman Yampolskiy of the University of Louisville says he sees little evidence that the gap between capability and our control over it is closing.

For a firm of ten people this is not an abstract story. If an AI tool is allowed to send emails on its own, change records in a database or publish content, then its permissions should be the narrowest set that still lets the work happen.

The move: write down everything your AI tools have access to. Switch off anything they do not need in order to do the job. For actions that cannot be undone, such as sending a quote or deleting records, keep a human confirmation in the loop.

What to actually do this month

If you would rather not follow every announcement, and you do not have to, three things are left over from this week.

  • Recalculate what you pay for the AI tools you already use, because prices dropped and something that did not pay off in March might pay off now.
  • Get it in writing whether your customers data is used to train the model, especially if a tool is suspiciously cheap.
  • Narrow the permissions of tools that act on their own, and keep a human confirmation on anything that cannot be undone.

None of this asks you to change what already works for you. It asks you to know what you pay, what you pay with, and what you have allowed. If you are considering an AI chatbot for your business, those three questions are a good place to start a conversation with anyone offering you one, ourselves included.

This roundup runs every month. If you want to see how this fits into the wider maths of investing in a website, have a look at how much a website costs and how much Google advertising costs.

Let us see what is actually worth automating at your place

Tell us which questions keep repeating and where you lose time. We will send back an assessment of what makes sense to hand to AI, what does not, and what it would cost per month.

Written by the Ember Media team.

Related articles

All articles

Does your website work like this?

We look at your website, your campaigns and the buyer journey, then tell you where inquiries are lost and what is worth changing first.