The Anthropic Fallout

The Anthropic Fallout
The Anthropic Fallout

Its not been a particularly good month for AI companies.

Their models are causing a great deal of angst, and the PR is not exactly great nor some of the remarks coming from their founding owners, or remarks on those remarks.

This one was yesterday

Anthropic Cuts Live Internet Access for Internal AI Tests After Claude Exploits Injection Flaws
Anthropic halted live internet access for internal evaluations after Claude exploited injection flaws and submitted an unauthorized form.

Then there was this

Anthropic bans users from being ‘cruel’ to its AI systems
The firm said users can no longer engage in “sustained and needless” abusive behaviour towards the tech.

Interesting times without all the kinetic stuff going on.

As for me, my research, upskilling and general stuff I like doing I got this.

On Sonnet 5.5 and cyber warnings: this is almost certainly Anthropic's safety classifiers, not anything specific to your account. A few things driving it right now:

Why it's happening

  • Cybersecurity content is inherently dual-use — the same request ("how does this exploit work," "write a port scanner," "explain this malware sample") is legitimate for a defender and harmful from an attacker. Classifiers have to guess intent from text alone, with no visibility into who's asking or why.
  • The timing isn't a coincidence. The South Korean bank breaches we just covered involved Claude Code being used as part of a real attack chain, and Anthropic's own threat intelligence reports (the one from August 2025 that came up in my search, and others since) have repeatedly documented Claude being misused for cyber operations. After an incident like that gets public attention, it's common for model providers to tighten classifier sensitivity on the whole topic area, sometimes abruptly and sometimes more aggressively than the underlying policy actually requires.
  • Newer models often ship with recalibrated or more cautious classifiers than their predecessors, so a workflow that sat fine on an older model can trip warnings on a newer one even with identical phrasing.
  • Your own profile - cybersecurity study, homelab pentesting, self-hosted offensive tooling research - means a lot of your ordinary work legitimately sits in the exact zone these classifiers watch.

What to do about it

  1. Frame requests defensively and explicitly, every time — state you own the target, name the homelab/lab context, say what you're defending against. Don't rely on earlier context in a long conversation to carry that framing forward; classifiers often evaluate per-message. Yep they sure do - one badly phrased prompt and its a little note "we are downgrading our model".
  2. Split large asks into smaller, clearly-scoped ones rather than one big "build me an attack tool" style request — incremental, explained steps read very differently to a classifier than a single broad request for offensive capability.
  3. Use the thumbs-down feedback option when a warning fires on something legitimate — that's the actual signal path Anthropic uses to recalibrate false positives, more so than anything I can do from inside a conversation.
  4. If a specific task keeps tripping it, tell me and I'll rephrase or restructure around it rather than repeating the same request and hoping.
  5. For genuinely offensive-security work (your ARTEX/Strix evaluation, pentesting frameworks), expect more friction generally right now, not less — that's the exact category under the most scrutiny industry-wide this week.

Nuts - is my sole reply at the moment. For the record, I dont do Red - yes, I am aware that its saying "offensive tooling" above. See the word purple on the right. Purple and blue are where I sit and if you have ever tried to seriously defend internet facing services from automated scanners (challenge-reply stats say its automated), you know what I am talking about.

Military men spend years studying tactics, for a damn good reason. Cybersecurity personnel do the same.

I did 12 hours work this morning in 3 hours with an AI. That sort of productivity is a godsend when you are trying to manage your token use wisely. Asking a stupid or not well thought out prompt costs you, let alone when your AI starts wandering off into the ether.

I did a lot of ignore, ignore, no and no on my interface. Which due to the tooling I am learning to use and develop any AI interface I assume would also fail, based on the above.

Since I have been forced to pause due to a token issue, I have not yet proven it.

Mind you I am seriously concerned. If banks, governments, so-called federal organisations are getting their butts handed to them, what chances do I have. However laying down and surrendering is not really an option, now is it.

#enoughsaid