Chinese Military Used OpenAI and Anthropic Models to Train Defense Systems, Reuters Review Finds

Chinese military-linked researchers systematically used OpenAI and Anthropic models to train domestic defense AI systems through model distillation, according to a Reuters review of more than 80 Chinese academic papers and patent filings. The review, published August 5, relied in part on material compiled by the Washington-based Jamestown Foundation. Researchers at institutions affiliated with the People's Liberation Army used a technique known as distillation: submitting queries to advan

2 min
Chinese Military Used OpenAI and Anthropic Models to Train Defense Systems, Reuters Review Finds

Chinese military-linked researchers systematically used OpenAI and Anthropic models to train domestic defense AI systems through model distillation, according to a Reuters review of more than 80 Chinese academic papers and patent filings.

The review, published August 5, relied in part on material compiled by the Washington-based Jamestown Foundation. Researchers at institutions affiliated with the People's Liberation Army used a technique known as distillation: submitting queries to advanced Western models and using the outputs as synthetic training data for smaller, locally controlled systems they could deploy on tactical hardware.

The applications span battlefield hardware, naval warfare, cyber operations, and social monitoring.

A 2024 paper from the PLA's National University of Defense Technology describes reducing an image-processing model so unmanned aerial vehicles can analyze live video and make navigation decisions in real time, even during communications blackouts. Researchers at China's Academy of Military Sciences used distillation to run target-recognition models on tactical hardware during simulated maritime operations involving ships, drones, and unmanned submarines.

Researchers in PLA Unit 96941, a Beijing-based cyber-warfare unit, used OpenAI's GPT-3.5 to summarize military source code and then trained a local model to operate within classified military networks. At the North University of China, researchers used Anthropic's Claude 3 Haiku to generate synthetic data for text classification and content monitoring.

"Teaching a model the right answer is one thing but teaching it the reasoning behind the answer is much harder," said Sunny Cheung, a Jamestown Foundation fellow. "These papers show Chinese military-linked researchers are trying to transfer that expensive, proprietary reasoning from Western models into smaller systems they can control and deploy locally."

The findings expose a strategic challenge that existing export controls cannot fully address. While Washington can restrict the export of high-end GPU hardware, limiting access to public API outputs or leaked model responses is substantially harder. The result is an asymmetric situation in which Western labs bear the cost of frontier model development while rival militaries extract targeted reasoning to deploy on satellites, drones, and field radios.

However, the report notes that distilled systems have inherent limitations. They can only reproduce reasoning patterns already present in the teacher model and lack the ability to extend beyond their training distribution.

Sources

Written by

More to read

  • Speech-to-Text Serving in Production: Comparing Faster-Whisper, Moonshine, SenseVoice, and NeMo Canary Architecture, Streaming Latency, and GPU Economics

    In conversational voice AI and real-time agentic workflows, the speech-to-text (STT) layer sets the hard lower bound on system responsiveness. Human conversational cadence expects turn-taking latencies between 200ms and 500ms. When an AI pipeline must accommodate downstream large language model (LLM) time-to-first-token generation (100ms to 250ms) and text-to-speech (TTS) audio synthesis (100ms to 200ms), the automatic speech recognition (ASR) stage cannot exceed 100ms to 150ms of processing ove

    1 min
  • Writer Releases Palmyra X6 Flagship Agentic Model with Rebuilt Enterprise Agent Harness

    Enterprise generative AI platform Writer has launched Palmyra X6, its new flagship agentic foundation model, alongside a rebuilt runtime harness engineered for multi-step workflow execution and governance. The model release introduces substantial latency and efficiency improvements over previous Palmyra iterations, cutting inference costs by 52% while accelerating output generation by 48%. Writer reported average generation speeds of 82 tokens per second and a mean task completion time of 26 se

    1 min
  • River AI Secures .1B Led by General Catalyst to Build Open-Weight Model Infrastructure

    River AI, an artificial intelligence startup founded by former xAI co-founder Igor Babuschkin, has secured $1.1 billion in early-stage funding to build an open-weight model stack and decentralized AI infrastructure platform. The financing round was led jointly by General Catalyst and public benefit corporation AMP PBC, with strategic participation from NVIDIA, AMD Ventures, Y Combinator, and Singapore sovereign fund Temasek. Babuschkin, whose prior engineering background spans OpenAI, Google De

    1 min