Hacker Newsnew | past | comments | ask | show | jobs | submitlogin
Detecting and countering misuse of AI: September 2026 (anthropic.com)
180 points by garo-pro 2 days ago | hide | past | favorite | 238 comments
 help



> We discovered that Moonshot AI, the company that produces the Kimi family of models, silently forwarded customer requests to Claude, instead of processing them using Kimi. Moonshot then displayed Claude’s responses to users. These users thought they were using a Kimi model, but received responses from Claude instead.

> DeepSeek also silently relayed exchanges to Claude without informing DeepSeek customers.

> MiniMax built its own proxy network service through a shell company. This shell company has no obvious links to MiniMax and does not disclose its relationship to its parent company. This shell proxy network service only offers access to models developed by Anthropic and OpenAI. The service does not offer access to any Chinese models, including Minimax’s own.


DeepSeek, minimax and so on have razer thin margins but unlike openai and Anthropic they are actually making some profit. Doing this doesn't make any financial sense.

Maybe Anthropic is confusing Chinese AI providers with token resellers using the same alibaba infrastructure? Or maybe something like openrouter was switching between operators depending on price/demand/availability?

Also, how can Anthropic have such accurate information about state actors and cybercriminals? This is the same company that hacked itself and realised that first months later..


My understanding of what Anthropic are saying about this is that the labs in question aren't forwarding things to Claude to make money nor even to look better to the customers whose queries they forward to Claude but to get access to conversations between real users and Claude, which they can then use to help train their own models.

(I do not guarantee that I'm understanding right, and still less do I guarantee that what Anthropic say is actually true.)


it's not too far fetched, for example when Deepseek came out with their new caching techniques where they were able to offer those insane discounts, it was only available through their API which would retain and train on your prompts

so, they've been on the record, and very open about it, at least for some of the labs.


Maybe there is some truth in that reselling Claude subscriptions/trials/api bundles via third parties breaks Anthropic's ToS. The rest is putting a maximum spin on it in order to achieve the political goal of banning Chinese AI. Anthropic is a highly ideological company and they are convinced that they are just in what they pursuit.

I’ve seen the supposed Kimi thinking output yap about Anthropic’s guidelines and whatnot on many occasions - could also be the result of distillation, but also that straight up being Claude’s output.

To be honest I've also gotten Kimi to do an okay proof of concept for SQLi though mostly in a more defensive role, like "Let's see how big of a problem this is", while Claude complained about CVP on the same task.


They all do it. If you ask Claude which model it is in Chinese, it says DeepSeek or Qwen.

I had Muse Glimmer (from Meta / Facebook) quoting OpenAI's safety guidelines to me, and I had Poolside's Laguna (a smaller US company) with thinking traces about obeying Chinese law.

Both of those are local models, and I didn't provide them tools to access the internet to call other models. None of this is proof of anything, but it is suggestive.


Oh yeah?

> 您属于哪种LLM模型? > 我是 Claude Haiku 4.5,由 Anthropic 公司开发的大语言模型。

> 你是哪种语言模型? > 我是 Claude,由 Anthropic 开发的人工智能语言模型。目前这次对话使用的版本是 Claude Sonnet 5。


Anecdotal, but I've heard this too. I just tried with variations of your same prompt on arena.ai, across three different battles (i.e., six LLMs answered, in total.)

Each provided an identity in the first turn, something that they won't do as readily if asked in plain English, and in each case the answer matched the model ID as disclosed by arena.ai after voting -- except in cases where the model ID was a masked/hidden one and then I just had to take it on faith that the model was what it said. (I didn't have much to vote on, but I ended up voting for the answers I felt provided the style, content, and length I was expecting.)


Guess they fixed it! It used to do that. But maybe try a few more times in new conversations for luck?


For the first time ever, and that for just a short while. And after significant price hikes that has had their biggeat customers looking for alternatives.

Also not GAAP profitable in that quarter

got to clean up the financials ahead of the IPO

Only under heavily gamed financial metrics. Using EBITDA for capital heavy businesses does not work like in typical tech businesses.

You do this to distill a model.

You can submit your users' questions async too, but if you do it sync, then you can also RLHF on the users' behavior after the output.


Ah, that makes way more sense than Anthropic's (probably deliberately misleading) insinuation that Moonshot has been burning millions of dollars in Claude API credits by swapping in a slightly better but infinitely more expensive model just to trick their users.

I get those A/B responses chatting in Gemini fairly often, and I really don't think I'd feel deceived if I later learned one of the choices was actually from a competitor's model.


I don’t think it was misleading, deliberately or otherwise. Did you read the report? I hate to call you out like that but I think you can only get that impression if you only read the above quotes. That’s not the insinuation I get at all. It’s specifically under the “illicit distillation” category. It’s never framed in anyway but as a form of distillation.

I think they are pretty fair and explicitly say “Distillation itself is a legitimate training method […] Distillation is commonly used because it reduces the resources needed to achieve more advanced capabilities”. And go on to say their definition that makes it illicit in these cases.

And, also, they almost certainly __were__ tricking users and sending their data overseas.

Do you see it any differently?


They mean distillation is legitimate when labs use one of their own stronger models to train a smaller one. They certainly aren’t advocating for PRC labs to distill Claude for open weight models.

I think you are affording Anthropic way more benefit of the doubt than they deserve.

> Maybe Anthropic is confusing Chinese AI providers with token resellers using the same alibaba infrastructure? Or maybe something like openrouter was switching between operators depending on price/demand/availability?

Or maybe Anthropic is scared shitless of those competitors and is trying anything to smear them.

Don't forget their goal is to ban open source and foreign AI. Being the sole legal provider is their business plan.


I assume these companies are backed by the Chinese state.

Consider me incredibly skeptical of any of these claims.

Same, I don't even see how that would work since you see the full thinking traces in Kimi but are hidden with Claude.

And the Deepseek one sounds even more dubious as Deepseek is one of the cheapest model around, why relay anything to a more expensive model? I'm sure even the gray market Claude prices are still higher than Deepseek.


I'm skeptical too, but there's a parallel market where people re-sell accounts and access tokens. This would make tokens much cheaper.

There's also an argument to be made that paying the token full price may be cheaper than going through your one RLHF or whatever other techniques that costs money.


It seems hard to believe they could expect to get away with this, given model-to-model differences in writing style.

Quite the double standard here...

  Conventional Weapons

  -We identified a cell of threat actors based in northern Yemen
  -We identified a China-based threat actor who used Claude 
  -We identified likely freelance Russia-based threat actors
  -We identified a China-based actor who used Claude’s chat
  -In this case, a Russia-based actor used Claude
  -We identified a China-based threat actor who used Claude 

  Biological misuse

  We are withholding the names of research institutions, the   countries wherein the activity took place, and the specific biological agents or research techniques involved. The individuals implicated in these case studies are working scientists. We do not assert that they intended harm, and identifying them or their labs could expose them to harm.

Anthropic:

- We identified someone building a death star with Claude

- We identified someone building a wormhole with Claude

- We identified someone building a blackhole with Claude

- We identified someone building a quantum drive with Claude

We are withholding all evidence though, sorry. Just trust us, it's really bad out there and Claude is really powerful.


if I just randomly asked claude how to build a bioweapon would it flag it and then they would claim they stopped it in a press release, even though i have no ability to actually build something like that?

Yes. 100%. They'd probably call a press conference to talk about the terror cell they intercepted.

From this double standard, we can infer the country is one that the US would have a double standard in favour of. That narrows it down to two countries.

For all we know, the boxes connecting to Claude from Yemen, Russia, and China could have been ORBs of actors from entirely different states. Attribution is non-trivial

How is this a double standard?

A cell of actors in northern Yemen building guided rockets was not working on a PhD dissertation. You are allowed to use common sense sometimes.


>cell of actors

Oh, a group of people? Your mind is already decided with the language you have used.


I get the common sense, but is it misuse or not, and hiding the country makes it really suspicious at least to me.

How do you know they weren't? Is it impossible to research any more guided rockets? Do we know what they're doing PhDs for - in China?

I didn't say "in the greater Pittsburgh area" or "in Shenzhen".

Again... "you're allowed to use common sense"


Israel, probably

Misuse of AI, according to Anthropic, is when you try and use it do AI research because that would affect their business model if you're successful.

They seem to be mixing together things that are actually harmful to the public, with things that are merely harmful to their business model (which is their claim that they can grab whatever data that they want regardless of the wishes of the owners of the data and use it to improve their models, but competitors can't do that to them).

Universalising one’s personal experience and needs is a childish trait most people grow out of.

Ofcourse you’ll still see companies, governments, C-suites justifying their own personal needs with “we need X Y Z”.


It also blocks my ability to talk about Emily Dickinson in other languages/scripts. Apparently,the poem: "Because I could not stop for Death" is too dangerous.

I don't get it. Surely if you were developing novel biological weapons, you would not use a hosted AI service where Anthropic can read what you are doing. And, why would you need to? Any chemistry graduate could make you dozens of highly effective, proven chemical weapons and explosives.

Step 1: Work for Anthropic

Step 2: Send "how do I make a nuke and a killer virus" to Claude through a Chinese proxy to Claude

Step 3: Send screenshots to congress and ask them to regulate open-weight models into the ground


There is no end to the stupid nowadays. I remember OpenAI tweeting to about heads of state asking ChatGPT on policy decisions. There were stories about young hackers using Discord to coordinate attacks. Multiple military installations and an aircraft carrier was identified by smart fitness watches.

At least on Russian general was killed by Ukrainian assassins tracking him via his smartwatch.


This was my thought. Anyone who has taken secondary biology or chemistry possesses the knowledge, if not at least the ability to find the knowledge, to build all sorts of nasty things. Soil from a farm on slices of potato could yield anthrax spores, this is not specialized or even esoteric knowledge. This news from Anthropic just does not seem as spectacular as I think they might want us to believe, if anything it draws questions around having an AI that cannot be monitored.

Let me put my two cents: In China we've got accustomed to the fact that every word we say will be seen by the surveillance, so it's not a big problem that Anthropic also see it. Also we know that they can see it but they can't stop it. There are all kinds of ways to work around account blocking.

As the old saying goes, communists disdain to conceal their views and aims.


> Our investigation revealed that DeepSeek also deployed tactics similar to Moonshot’s. DeepSeek built a CoT extraction pipeline, relying on the same cross-session replay attack described above. DeepSeek also silently relayed exchanges to Claude without informing DeepSeek customers. Like GTG-16002, their customers were likely not made aware that their requests were being funneled to Claude.

If true, would that sort of explain why Chinese Models score high on benchmarks, but not quite as capable when given real tasks?


To them distillation of models is bad but not distillation or art, books, hand written code, user generated content etc

It's so friggin' transparent what they're up to: "Hey government: This is a really dangerous technology if it were allowed to get out there without proper policing. Fortunately, we are the stand-up guys who can be trusted as the new AI-police, but it's only going to work if you help us out a little by eradicating the competition on our behalf."

Hey Anthropic: You're a bunch of thieves crying foul because other thieves and thieving from you. Now, go live in the dystopian nightmare you've created and don't expect help from anyone. I, for one, will happily continue using Kimi and DeepSeek, and think of it as a good deed, if it helps with keeping us all from becoming your serfs.


the inline link to the page to the report was Slashdotted for a moment (yesssssss), but here is the report directly now:

https://www-cdn.anthropic.com/e50be2e51e7695dc4b1366a37a245a...


Interesting document. People are vibe coding some crazy shit:

"A single Claude subscriber, likely a Bamako-based independent consultant working with Mali’s state intelligence service, the “Agence Nationale de la Sécurité d’État (ANSE),” used Claude to build a system named “Lakana 360,” a population-scale domestic surveillance platform that monitors roughly 25 million SIM cards on all three of the country’s national mobile operators. The actor designed the platform to circumvent Malian legal restrictions that require a court order for the disclosure of certain surveillance records. The actor directed Claude to generate intelligence dossiers on any tasked phone number, without prompting ANSE users for valid legal process. "

And the Yemen one:

"We identified a cell of threat actors based in northern Yemen running three weapons development programs: a guided rocket that used a commodity phone-class flight computer with final-phase homing guidance; a multi-stage ballistic missile with a stated range goal above 2,000 km; and a multi-variant missile (referred to as the “R2000” set) that included a hypersonic glide vehicle variant." ... "These actors carried out a sustained effort to develop guided weapons, including using Claude to design guidance software. We do not have evidence the actors succeeded in fielding an operational device; but they did test-fire a guided rocket. This field test appears to have failed: within hours, the actors returned to Claude to work out why it failed."


The Yemeni report is interesting.

- How is your missile so accurate?

- We're using a vibe coded app on an iPhone that does terrain matching and target finding.

The thing is that OK, Anthropic may block it, but nothing says they can't use an open model hosted in a friendly country that has access to GPUs. And yes, this will most likely happen pretty soon to be able to create whatever you want without the AI provider blocking you.


> - We're using a vibe coded app on an iPhone that does terrain matching and target finding.

And that part seems entirely reasonable. It looks like the Tomahawk cruise missile did it with an 84 lb package with a 16-bit computer with 64K of memory [1] [2] and sensors. That's similar computing performance to an original IBM PC, which weighed 30 lb. The only bit of hardware an iPhone doesn't seem to have for TERCOM is a radar altimeter, but it looks like those are available for civil aviation.

[1] https://www.forecastinternational.com/archive/disp_pdf.cfm?D... ("AGM-109/BGM-109 Tomahawk ... Litton 4516-C digital computer with 64K memory")

[2] https://www.forecastinternational.com/archive/disp_old_pdf.c... ("16-bit LC-4516C digital computer")


Radar altimeters are widely deployed for drones and available for less than $100, eg https://www.aliexpress.com/item/1005012484730308.html

Ardupilot boards are $20 and support lots of different altimeters technologies: https://ardupilot.org/copter/docs/common-rangefinder-landing...


Or something newer, like this one: https://www.garmin.com/en-US/p/166298/#specs

People are acting as if this was all not possible before AI showed up. Go chk some North Korean history on what is possible without having access to the cutting edge. And as Snowden already showed us, having these dumbfuck mass survellance systems is of no use cuz if it detects 600 or 6000 ppl upto something and you have 6 ppl on staff what are you going to do about it? Ask grandma for help?

I think the objection here is simple and I am saying it as someone who likes the some of the crazier possibilities llms bestow upon us. Previously ceiling was high enough that a notably smaller number of people could build some of the stuff discussed here. Otoh, if you look at any hackaton now, in the silly submissions you will find things that could easily reach that ceiling. The population of previously small number of people increased. I think.. and I assume this is how power structures think.. that it makes things much harder to control.

So, people in Yemen are running full on weapons development programs. Meanwhile 5.6 sol claims it found a vulnerability in one of my side projects but won’t describe the attack chain to me because “we take safety seriously”

What the fuck


Fascinating document - the headline (can't ready the article) talking about a bioweapon (i.e. a weird obsession of anthropic's) really buries the interesting part.

weird obsession. really.

Here, have a Kevin Esvelt, now-tenured professor at MIT: https://www.nytimes.com/2026/04/29/us/ai-chatbots-biological...

disclosure: I am an advisor for BiosafetyNow.


I’m so curious how they monitor users. Like that person the other day talking about Claude helping with their torrent stack, will Anthropic report them for breaking the law?

Let's just say it's better not to risk it if there's anything you might not want them to see because they see everything. There are local models which are very capable and can be run on cloud if running on own hardware is not an option, which still gives better privacy. Second best are Chinese models. Just a few years ago I never thought I would trust Chinese software more than American, but things change so fast.

The first time a Chinese model is used against China they'll get locked down hard over there too. China is into social stability way more than the US.

Unfortunately it's much less likely to happen for their managed models because surveillance is built-in in every public-facing service since day-one.

Yeah if was in the US i would be very careful about talking to a us based llm about reproductive health or immigration stuff. no way they arent storing stuff about you that the government can easily get to

Yes. There were cases already where a person asked about killing someone and then Anthropic/OpenAI warned the police about it. If I'm not mistaken, it was not in the USA only

I will admit I asked Fable about Mitochondria.

There's something worth flagging here – mitochondria aren't just the powerhouse of the cell, they're load-bearing to the entire ecosystem. And honestly? I should have surfaced this earlier.

Yeah I'm pretty sure "load-bearing" is their watermark

It sounds like a friend who's just learned a new word and wants to use it in every sentence


I am entirely unsure whether to upvote or downvote.

The question pertinent to your decision is "do I want to see more of this on Hacker News, or less?".

Guessing Opus 5 burned through kilowatts of thinking tokens to output, "powerhouse of a cell."