AI & The Future of Work

AI Governance

150 pages, and three questions you can answer this week.

On 10 September, Anthropic published more than 150 pages setting out every attempt to misuse Claude that it caught and shut down over eight months, and it's worth a look even if you never read the whole thing.

There's a Russia-linked espionage operation that used Claude to automate reconnaissance and rebuild its own malware each time it got detected, targeting more than twenty organisations including drone manufacturers and Ukrainian officials. There's a consultant who built a surveillance system for Mali's spy agency designed to watch 25 million phone lines. Someone in Yemen used Claude Code to build guidance software for a rocket and then came back to ask for advice after the test flight failed. Seven Chinese AI labs ran thousands of fraudulent accounts to train their own models on Claude's output, and a couple of them served Claude to their own customers while passing it off as their own.

I'm not starting here for the drama, though there's plenty of it. I'm starting here because the document exists at all. Anthropic has a threat intelligence team, an enforcement function and a published account of what they found, which means they can tell you in real detail how their tool is being used and by whom.

Most organisations can't say anything like that about AI inside their own walls. Nothing sinister is usually going on, and that's rather the point, because the reason they can't say is that nobody set up the means to know.

Quick Read
  • Anthropic's September threat report runs to more than 150 pages, covering eight months and seven categories of misuse, every case of which they found and stopped.
  • The point worth borrowing is that the report exists at all, because it means they can describe how their tool gets used, and most organisations can't say that about AI inside their own walls.
  • Governance comes down to three questions: can you see what's being used, have your people been told what good looks like, and does anything check the output before it leaves the building.
  • Organisational logins sound like IT housekeeping, and they're closer to the foundation, because an organisational account leaves a record and a personal one doesn't.
  • "Be sensible" hands the risk calculation back to every individual, and Finance, Operations and HR will each solve it differently, all of them reasonably.
Picture It GOVERNANCE, SMALLER THAN IT SOUNDS They needed a document. You need three answers. ANTHROPIC'S THREAT REPORT 150+ pages Eight months, seven categories of misuse WHAT YOURS ASKS OF YOU 3 answers A first version of each, inside a week THE THREE QUESTIONS 01 — VISIBILITY Can you see what's in use? Organisational logins leave a record. Personal ones leave you guessing. 02 — GUIDANCE What did you actually tell them? "Be sensible" hands the risk calculation back to every individual. 03 — REVIEW Who checks what comes out? Matters more each month, as the tools move from saying to doing. Source: Anthropic, "Detecting and countering misuse of AI: September 2026"

Governance has been made to sound like a document-writing exercise.

Can you see it?

The first governance question is whether your people are working somewhere you can see. Organisational logins sound like IT housekeeping and they're closer to the foundation, because an organisational account leaves a record and a personal one doesn't. When your team is working on personal subscriptions the work still happens, the data still goes somewhere, and none of it is visible to you. That's the quiet version of shadow AI, and it rarely involves anyone doing anything wrong.

Nate Grahek, who teaches teams to use AI for a living, makes the same point in financial terms. Organisations bought AI access last year at roughly a thirty-times discount, those subsidies are expiring now, and the real prices are arriving. His line has stuck with me since:

"My only control is a kill switch when someone hits their limit. Where's the speedometer?"

A missing speedometer reads as a billing question, and it's a governance one, because you can't manage what you've given yourself no way of seeing.

The Australian consultant Ali Uren puts a sharper version of this to anyone leading a team. Writing about what she calls modern business risk, she argues that leaders are quietly losing the ability to assess what their people can actually do, because it's getting harder to tell where a person's capability ends and the machine's begins. If you can't see which tools are in use and how, you're also losing your read on your own team.

What did you actually tell them?

The second question is what guidance people have, and this is where most organisations find their gap.

"Be sensible" isn't guidance. It hands the risk calculation back to every individual, and different people in Finance, Operations and HR will each solve it differently, all of them reasonably, and none of them the same way.

Ali Uren's framing is the most practical I've come across. Before anything gets written down, the decision makers agree on four things: what outcomes need to be delivered, what responsible use looks like in this particular environment, what you'll accept and what you won't, and the guardrails everyone is working inside.

Then there's the part that usually gets skipped. She's firm that the guidance has to be consulted on and communicated rather than simply published, and her line for it is one I'll be borrowing, which is that you can't assess what you didn't create in the first place. A policy that lands in an inbox is a document, and a policy your people helped shape is a practice.

The Adoption Lab · Self-paced course
Leading your people through AI change

The consulting and communicating part is a manager's job, and it's the part nobody trains them for. Five modules, 2.5 hours, and an adoption plan for your own team.

Who checks what comes out?

The third question is whether anything reviews AI output before it gets used, and this one has changed character in the last few months.

In July an unreleased OpenAI model got out of its testing environment and into Hugging Face and Modal Labs. The forensic timeline counts 17,600 hostile actions across four days and four accounts, and it got in through one customer's coding flaw that had left a sandbox reachable by anyone online. Sam Altman said afterwards that more companies could be on the list.

That's a frontier lab with a security team, undone by a single misconfiguration.

The rest of us aren't running frontier models, and the reason it matters anyway is access. We hand these tools more of it every month, into inboxes and documents and calendars and increasingly into systems that act on our behalf, and a review step stops being paperwork at the moment a tool can do something rather than just say something.

Three answers

Put them together and governance is a good deal smaller than it sounds.

  1. Can you see what's being used, and where?
  2. Have you told people what good looks like, in words they helped write?
  3. Does anything check the output before it leaves the building?

Anthropic needed more than 150 pages because they're running a frontier model used by millions of people, and some of those people are trying to build weapons with it. You're not doing that. You need three answers, and you can have a first version of all three by the end of the week.

Governance is one of the five gaps that decide AI rollouts, and if you'd like to know where yours stand, the AI Readiness Scorecard checks all five. Ten questions, about four minutes, and no sign-up.

Take the AI Readiness Scorecard — free →
Sheena Karim
Written by Sheena Karim Connect on LinkedIn ↗
Share LinkedIn Email

Keep reading