The price now comes with an expiry date
Two AI price changes landed in three days, one down and one up, and both of them have a date attached. The number in your quote is no longer a number. It is a schedule.

- 01
DeepSeek splits the day into peak and off-peak, from this afternoon
Until 15:59 UTC today, V4-Flash cost 0.14 dollars per million input tokens on a cache miss and 0.28 per million out, flat, whatever the hour. From 16:00 UTC DeepSeek moves to peak and off-peak billing, with off-peak set at half the peak rate and peak defined as 01:00 to 04:00 and 06:00 to 10:00 UTC. Reporting on the new rates puts V4-Flash peak input at 0.44 and peak output at 1.32, which is the first time a major provider has made the clock part of what you pay.
DeepSeek - 02
Gemini 3.7 Flash arrives at half price, and the half price runs out
Google released Gemini 3.7 Flash on 13 August at 0.75 dollars per million input tokens and 3.75 per million out, roughly half what the previous Flash cost, and it is generally available rather than a preview. The part to write down is in the same announcement: introductory pricing ends on 31 December 2026, and from 1 January the rate doubles to 1.50 and 7.50. Anything costed on this model today is costed on a rate with eighteen weeks left on it.
Google - 03
The Microsoft Copilot app changes address on 18 August
Microsoft told partners on 14 August that from 18 August the Copilot web app moves from m365.cloud.microsoft to copilot.cloud.microsoft, with automatic redirection unless the new address is blocked on your network. The app also gains visible work and personal account labels, including a green shield for Microsoft Entra work accounts, which is aimed squarely at people who have both signed in and paste the wrong thing into the wrong one. Practical job for whoever runs your IT: make sure the new address is not sitting behind a firewall or proxy rule that nobody has looked at since it was written.
Microsoft - 04
A compromised scanner poisoned an AI library for forty minutes
Researchers at CloudSEK found that attackers reached the LiteLLM package not by attacking it but by compromising Trivy, a security scanner that LiteLLM's own build pipeline installed automatically, then published malicious versions 1.82.7 and 1.82.8. Those versions were live on PyPI for about forty minutes and still reached an estimated 2,500 organisations and 434,000 build pipelines, harvesting cloud keys, SSH keys and API tokens. Nobody in that chain did anything careless: the poison came in three suppliers deep, through the tool whose job was checking for poison.
SecurityWeek - 05
A lab holds back its own weights because the model got too good at attacking
Z.ai released GLM-5.3 on 14 August through its API and coding plans, but is not publishing the open weights for around two more weeks while it finishes safety evaluation and hardening. Its reason is that offensive security capability grew faster than expected: the model leads the CyberGym vulnerability-finding benchmark at 84.5 per cent, and the company says its models have surfaced 2,436 vulnerabilities across 269 open-source projects, 1,097 of them critical or high. This is now the third lab in a fortnight to slow something down over cyber capability, after OpenAI paused Astra work and split its security programme into vetted tiers.
Z.ai
Two price changes in three days, pointing in opposite directions, and the useful thing is what they have in common. Google halved the cost of Gemini 3.7 Flash and printed the expiry date in the same announcement: 31 December, then it doubles. DeepSeek went the other way this afternoon and made the clock part of the bill, so the identical job costs one thing at nine in the morning and half that at two. Neither of these is a scandal. Both of them quietly break the way small firms buy software, because we are all used to a per-month figure that stays put.
So if anyone has quoted you a running cost for an AI tool, a chatbot on the site, something that reads drawings, a thing that answers the phone at seven, that figure is a snapshot of one week in August. It is not wrong. It is just not a price. From direct experience in construction, a materials quote carries its validity in writing, thirty days on the paper, and everyone knows to check the date before ordering off an old one. Nobody prints that on AI running costs yet. The fix is not to stop buying: it is to make the supplier write down which model the thing runs on, what a job costs at today's rate, and what changes if that rate triples. If they cannot answer the third one, the number they gave you was decoration.
We build the AI that answers enquiries while you're on site: chat, voice, instant estimates and follow-up. See how it works or price it in two minutes.