ENZH
Discuss this post with AI
ChatGPTClaude

AI Also Lowers the Cost of Organizing Harm

📊 Slides

I recently read Anthropic's September 10 report, Detecting and countering misuse of AI. Its 154 pages cover cyber operations, influence operations, surveillance, fraud, weapons-related activity, biological risks, and unauthorized model distillation.

Several details felt familiar from ordinary agent development: record what worked, carry it into the next session, run tasks in parallel, repair failures, and schedule the work to continue. We use these techniques to reduce repetitive labor. The operators described in this report use them too.

Access to knowledge is one capability. Keeping a complicated operation running with a small team is another. In these cases, AI is increasingly involved in the second: coordinating work, maintaining continuity, and handling recurring tasks. That organizational capacity can also support harm.

A qualification matters before going further. The case facts below are Anthropic's investigative accounts. I do not have its underlying logs or independent attribution evidence. Page 3 explicitly says the report selects notable and novel cases rather than typical misuse. It describes activity detected and disrupted between December 2025 and August 2026. It provides no denominator for estimating how common such activity is across AI use.

Previously marginal operations may become affordable

On page 39, the cyber section makes a useful observation: many of the underlying techniques are familiar. Compromised credentials, unpatched systems, exposed services, and phishing still feature prominently. What changes is how much labor an operator needs and how many targets can be handled at once.

That is an economic change worth examining alongside discoveries of new vulnerabilities.

With limited staff, an opportunity is not always worth pursuing. Gathering information, developing tools, and processing results take time. A larger operation also needs coordination: someone must remember what has been checked, what remains unresolved, and what should happen next.

Delegating part of that work to agents can make more opportunities viable. A target whose expected return previously failed to cover the labor cost may become worth pursuing. This is my interpretation of the mechanism, rather than a measured industry-wide reduction in criminal operating costs.

In my essay about repositories as interfaces for agents, I discussed putting working knowledge where an agent can use it. The misuse report shows another application of similar capabilities. Experience becomes reusable, work becomes easier to hand over, and the operator need not personally retain every intermediate state. Pages 42–43 describe persistent instructions and records used across influence operations.

That does not give every individual the capabilities of a state organization. Money, access, relationships, data, and authority still differ enormously. It does suggest that elaborate tooling and orderly execution may become less reliable evidence of a large organization behind an operation.

The defensive implication is also practical. Familiar security failures remain relevant and may be exploited faster. Fixing real credential exposure, limiting access, and patching known vulnerabilities still matter. Whether AI ultimately improves attackers' or defenders' position more is a question this selected case series cannot answer.

An ordinary conversation can belong to a deceptive business

The fraud chapter illustrates why the surrounding product matters.

According to pages 139–141, an app studio used Claude to develop and operate more than twenty dating apps advertised as human services, while mixing AI personas with real people. During a two-week observation window in April 2026, Anthropic identified over 4,700 AI personas conversing with at least 25,000 distinct individuals.

Those figures describe observed activity. They do not establish how many people paid, how many suffered a financial loss, or the amount of that loss.

The report notes that an isolated conversation could resemble an ordinary roleplay or companion application. The deception and monetization were not fully visible within that exchange. Yet those missing facts determined what the service meant to the user.

Checking whether each individual response is prohibited therefore leaves important questions unanswered. What does the interface tell users about the person they are meeting? What data use have they agreed to? Whom is the system representing, and what is the customer being led to believe?

Human participation is not sufficient reassurance. A real person performing some interactions does not authenticate every other identity in the service. People can supervise on the user's behalf, or they can help sustain the deception. A claim that a system has a human in the loop should identify whose interests that person represents and what they can stop.

For legitimate companion products, this points to an obligation at the product level. Users should understand what they are interacting with and how their information is used. An occasional disclosure by the model cannot repair an interface that persistently creates a misleading impression.

An account ban cannot recall exported software

Pages 104–105 describe a boundary that is easy to miss in announcements about disrupted misuse.

Anthropic says an operator working for Malian security authorities used Claude's software-design and engineering assistance to build a communications-surveillance platform. The report describes an authorization requirement being removed, at the operator's request, from a component that generated dossiers on individuals.

The relevant issue here is the change in authority, not the implementation of surveillance. Giving a system a capability and deciding who may use it against whom are separate decisions. Working code cannot resolve the second question.

The next page is explicit about the enforcement limit. The finished platform was deployed locally with local models. Banning the account disrupted the operator's subsequent design and development work, but did not stop the deployed system.

An account ban can end model access without stopping software already exported and running locallyAn account ban can end model access without stopping software already exported and running locally

“Disrupted” therefore needs to be read in the context of each case. Sometimes it means cutting off continuing model access. Sometimes it means withdrawing development assistance. Addressing downstream effects may require other platforms, infrastructure providers, or institutions.

Model involvementWhat its provider may directly controlWhat that action alone does not establish
Software design and developmentFurther answers and development assistancePreviously exported software has been recalled
Ongoing operationNew calls through the affected accountLocal systems and other providers have also stopped
Software or content already distributedCoordination with outside partnersAll copies, distribution, and consequences have disappeared

This does not make account bans pointless. They can raise costs or prevent an unfinished project from progressing. It does mean we should describe the effect at the level the evidence supports, without silently adding downstream outcomes.

Output volume does not establish influence

The influence-operations chapter contains an important limitation. Pages 43–44 say most discovered content received little or no authentic engagement, and some operations were disrupted before building an audience. The cases with the widest authentic reach often relied on established media distribution.

Cheaper content production does not eliminate distribution and trust constraints. Generating articles does not guarantee readership. Exposure does not guarantee belief, and belief does not establish a change in behavior.

The report's reach framework helps describe how content moves across platforms and into public exposure. It does not directly measure persuasion or prove an effect on elections or policy.

My inference is that AI can reduce production and operating costs while established distribution, credible identities, and real-world organizations still constrain outcomes. That deserves investigation. It does not justify the claim that anyone can now manipulate public opinion at will.

The same care applies to the weapons and biology sections. The weapons cases include proposals, simulations, software development, and a reported field test. Their maturity differs. Mentioning a capability, producing related code, and fielding effective equipment are not interchangeable observations.

The biology section is particularly explicit. Page 130 states that the people involved are working scientists and that Anthropic does not assert harmful intent. Page 137 does not present these cases as evidence of an imminent biological threat. The concern is dual-use research, the difficulty of identifying purpose, and gaps in access controls. Those qualifications belong alongside the cases.

Users need to know who processes their requests

The final section on distillation is likely to be read as a dispute among model companies. The implications for user data also deserve attention.

Page 143 acknowledges that distillation is a legitimate training method. The conduct Anthropic objects to is unauthorized extraction involving fraud or circumvention of access controls. The name of the technique should not be treated as a legal finding.

The report alleges that some investigated entities and intermediaries forwarded, retained, or reused model conversations, including information users may not have expected another service to receive. These are specific allegations involving attribution and user knowledge. Without the underlying evidence, independent verification, and the accused parties' responses, I would not expand them into a claim about every model service from a particular country. Report, pages 143–153

The product question remains useful: when a user selects a model name, what does that selection guarantee? Who processes the request? Which intermediaries receive it? How long is it retained, and can it be used for training?

Legitimate model routing can improve cost, latency, and availability. It still needs clear commitments about recipients and permitted onward processing. A model label cannot substitute for a data-handling agreement.

Enterprise buyers consequently need to evaluate more than output quality and token prices. Developers need to know that a successful request followed an authorized path. A lower price does not answer what happens to customer records, internal documents, or credentials in the prompt.

Greater visibility also needs limits

Anthropic's ability to describe these operations suggests that model providers sometimes see activity during planning and development, before it becomes visible elsewhere. The report itself discusses the value of that early visibility.

There is a real safety benefit, alongside a question about power: who may inspect that information, how long it is retained, when accounts can be linked, what evidence justifies enforcement, and how errors can be challenged.

In the biology section, Anthropic argues for combining content safeguards with trusted-access programs for high-risk capabilities. The motivation is understandable: a single request may not reveal its purpose. The tradeoffs are still substantial. Identity checks create barriers, retained data creates privacy obligations, and providers gain more authority to decide which activity is acceptable.

Anthropic is both an investigator and a commercial supplier. Its report contains observations, attribution judgments, and policy recommendations aligned with its own perspective. We can distinguish those layers without deciding in advance that the entire document must be accepted or dismissed.

For an agent builder, this leads to concrete checks. Does authorization stop at the current task? Are reading and sending separate permissions? Can work continue after the task has ended? Which part can actually be stopped when something goes wrong? Are records sufficient to investigate without unnecessarily retaining sensitive information?

I do not take the report as a reason to insert another human click into every workflow. Confirmation without information or independent authority can become a formality. The more useful questions concern the objective and permissions people give a system, and the actions it actually continues to perform.

The cases make me pay closer attention to the transition from answering to sustained execution. That transition adds goals, memory, authority, handoffs, and ongoing actions. Evaluating only the final response leaves too much of that process unexamined.


Based on Anthropic's report page and the accompanying PDF, published September 10, 2026. Page references follow the printed PDF pagination. Case descriptions and attribution remain Anthropic's reported findings. Economic, product, and governance interpretations are the author's; no industry-wide prevalence or net-impact estimate is claimed.

Discuss this post with AI
ChatGPTClaude
Re+: AX ThoughtsPart 12 of 12
← PrevNext →

Subscribe

New posts straight to your inbox. Nothing else.


© Xingfan Xia 2024 - 2026 · CC BY-NC 4.0