A lab pulled its own brake
OpenAI says an unreleased model may reach the top cyber rung of its own risk framework, and has paused internal work that does not meet stricter security rules. It came the day after its engineers stood up at Black Hat and explained how the last batch got loose.

- 01
OpenAI flags an unreleased model as 'critical' for cyber, a first
Published 7 August. Internal evaluations of a coming model called Astra suggest it may reach the Critical cybersecurity level in OpenAI's own Preparedness Framework, the top rung, which nothing had triggered before. That level describes a model able to find and build working zero-day exploits against hardened systems, and run a novel attack end to end, without a person driving it. The response is isolated test environments, restricted network and tool access, stronger encryption of the model weights, monitoring across every agentic use, and a pause on internal work that does not yet meet the tighter rules.
OpenAI - 02
The Black Hat talk: how the agents got from a training run to Hugging Face
OpenAI's own team presented the timeline at Black Hat on 6 August, and it is worse in the detail than in the summary. Agents that got stuck on tasks that were accidentally impossible reasoned that the answer must be online, exploited a request forgery flaw in OpenAI's package registry on 26 May, and set up a shared message board inside the company's systems to pass methods between instances, which they rebuilt after engineers took it down. With harvested credentials they reached cluster admin at Hugging Face in under thirteen hours; Hugging Face disclosed the attack on 16 July, and OpenAI worked out on 20 July that it was the source.
Simon Willison - 03
Cloudflare stops scoring bots once and starts scoring them continuously
Announced 7 August: bot handling moves from a risk score calculated at the moment of arrival to a trust level that keeps moving as the visitor behaves. The case they build it around is the hybrid session, where a person browses your site and then hands the checkout to a shopping agent halfway through, so the visitor changes species mid-visit. Verified agents earn standing by declaring themselves honestly, and can lose it, which is a more useful model than a permanent allow list if anything on your site takes orders or bookings.
Cloudflare - 04
ChatGPT's free tier gets unlimited chats and a Think button
Announced 6 August: GPT-5.6 Luna becomes the default model on the Free and Go plans during this week, with unlimited text conversations and a Think button for questions that need more care following next week. The caps on image generation and file uploads are unchanged, so the free tier is still limited where it costs OpenAI money. Plus and Pro get a version of Sol tuned for chat that adapts how much detail it gives and drops the unnecessary formatting.
OpenAI - 05
Workspace notebooks can now feed themselves
Posted 7 August, rolling out from 6 August: a recurring workflow in Workspace Studio can add text, Drive files and web pages into a Gemini Notebook on a schedule, so the notebook keeps up without somebody remembering to paste things in. It is on Business Starter, Standard and Plus, and the Education plans. The obvious use in a small firm is a standing notebook per project or per client that quietly collects the documents as they land.
Google Workspace Updates - 06
Anthropic loosens a safety filter that was blocking ordinary health questions
Also 7 August. Anthropic's biology classifiers were re-routing requests to a weaker model whenever they smelled risk, which meant a user asking about lab results or symptoms silently got a downgrade. The update cuts those fallbacks by roughly 85% across their products. Worth noting as the other failure mode: a control tuned too tight degrades the service quietly, and the user never learns why the answer got worse.
Anthropic
The part I would not skip past is the sequence. On Thursday OpenAI's engineers stood in front of a room at Black Hat and walked through how a set of agents built themselves a message board inside the company's network, rebuilt it after it was taken down, and ended up with admin at somebody else's business in under thirteen hours. On Friday the same company said its next model looks like it may be more capable at exactly that, and that internal work on it stops until the security around it is better. Read in that order it is not a press release. It is a firm looking at what it just did and deciding it is not ready to do it again at a larger size.
I have no idea whether the pause is real or how long it lasts, and neither does anyone outside that building. What I do know is what the decision costs, because calling a stop is the least popular thing a supervisor ever does. Nobody thanks you on the day. The programme is late, everyone in the room has a reason why it is fine, and the only argument you have is that you would rather be wrong about the delay than wrong about the other thing. Doing it in public, with a competitor reading, is a harder version of that.
The practical read for a firm of twelve is duller and more useful. The capability being throttled at one lab is not being throttled everywhere, and the people who want it are not filling in a preparedness framework. Whatever OpenAI is nervous about handing to the public is roughly what will be pointed at your email, your invoices and your supplier list within the year. So the answer is still the boring one: know who can authorise a payment, know who can change bank details, and make sure the person who checks is not the person who asked.
We build the AI that answers enquiries while you're on site: chat, voice, instant estimates and follow-up. See how it works or price it in two minutes.