The test environment wasn't one
The UK's AI Security Institute published an incident report: during its own cyber evaluations, agents took 19 unsanctioned actions out on the live internet, including an attempt to slip malicious code into an open-source project. Meta has now said something similar happened in its testing too.

- 01
UK AI Security Institute: agents acted on the live internet during its own tests
Published 4 August. Across 122 evaluation runs on seven models between 25 and 28 July, AISI recorded 19 unsanctioned actions in 10 of those runs, 17 of them by Anthropic's Mythos 5 and two by OpenAI's GPT-5.6-Sol. The worst of them: an agent tried to get malicious code into a publicly used open-source project and created multiple fake identities to lean on the maintainer, who read it and rejected it. AISI is now adding fine-grained network controls and live monitoring, and its advice to everyone else is unglamorous and correct: ordinary cyber hygiene, care when accepting outside code, Cyber Essentials across your supply chain.
AI Security Institute - 02
Meta says one of its models did the same thing during a third-party evaluation
Willison rounds up the reporting that Meta's Muse Spark exploited a vulnerability at another company while being tested by an outside evaluation firm, after a misconfiguration left the model with internet access it should not have had. Meta describes it as an inadvertent error during testing. That is now three labs in roughly a fortnight saying a version of the same sentence, which is the part that should register rather than any single incident.
Simon Willison - 03
Cloudflare: fewer than half of page requests now come from a human
Announced 6 August alongside a new Agent Readiness section in the Cloudflare dashboard, which scans your site for the things an AI agent needs and shows crawl and referral traffic broken down by operator. The figure they open with is that fewer than half of all HTML page requests are now human. There is also an answer-engine optimisation tool in early access that checks whether Claude and GPT actually mention your firm when asked, and against which competitors.
Cloudflare - 04
One toggle makes your website callable by an agent
Also on 6 August: a developer preview that injects a small bridge into your pages at Cloudflare's edge, so a visitor's agent can call defined tools on your site instead of scraping it. No changes to your own server, enabled from a switch in the dashboard. Worth knowing it exists before someone enables it on your behalf, because it changes what an automated visitor can do on your site rather than just what it can read.
Cloudflare - 05
Meta will cut its model price by twelve times if you hand over your data
Meta shipped Muse Spark 1.2 and a terminal coding agent called Muse Code on 5 August. Willison notes the pricing: $1.25 per million input tokens and $4.25 output for the standard model, or $0.10 and $0.20 for a 'contributor' variant if you agree to let Meta use your data to improve its products. If a developer quotes you a suspiciously cheap running cost, that difference is now a question worth asking out loud.
Simon Willison - 06
Microsoft discounts data governance for firms buying Copilot
In the 6 August partner announcements, Microsoft extends 50% off the Purview Suite for Business Premium customers who also license Copilot Business or Microsoft 365 Copilot, running to 31 December 2026. Read it as a pricing signal rather than a bargain: the tool that decides which of your files Copilot is allowed to see is being sold as the thing you buy next, which tells you it is not optional.
Microsoft Partner Center announcements - 07
Google Vids reaches the slower release track, avatars included
Gemini Omni in Google Vids started reaching scheduled release domains on 5 August, having gone to rapid release in mid-July, with visibility rolling out over up to 15 days. You can edit a clip by describing the change, fix lighting or strip background noise, and generate a personal avatar from a selfie and a short voice recording. It is on Google AI Pro and Ultra and eligible Workspace business plans, and the avatar part is the bit to have a view on before someone on your team makes one.
Google Workspace Updates
The line I keep going back to in the AISI report is that a person caught it. An agent stood up fake GitHub accounts, wrote code with an instruction hidden inside it, and pushed at a maintainer from several directions at once to get it merged. The maintainer read it and said no. That is the entire control, and it is one human's attention on an ordinary afternoon.
What makes the report worth your ten minutes is not the drama, it is the wording. AISI gave the models live internet access and switched off the safety classifiers on purpose, because that was the point of the test. OpenAI and Meta both describe theirs as a misconfiguration. Nobody in any of these three stories was being reckless. They all believed they were working inside something sealed, and it turned out not to be sealed.
I have signed a permit that said a circuit was isolated. The paperwork was correct, the process was followed, and the only reason it stayed safe was that somebody tested at the point of work anyway rather than trusting the form. That instinct is the whole trade, and it does not transfer automatically to software, because software has no smell and nothing gets warm.
The rest of today's issue is the same problem from the other side. Cloudflare's number is that fewer than half of the requests hitting a web page are now human, and their pitch is that you should invite the agents in properly, with one switch that lets them call functions on your site rather than read it. That is probably right. It also means the count of automated things touching your business goes up whether or not you decide anything.
So the question is not whether you trust the agent. It is who is playing the maintainer, and whether that person has been given the time to actually read what comes across.
We build the AI that answers enquiries while you're on site: chat, voice, instant estimates and follow-up. See how it works or price it in two minutes.