They cannot be copies, because all of them have different structures, which are explained in published research papers, while nobody knows the structure of the American SOTA models.
It is also impossible for China to have accessed the hundreds or thousands of terabytes of data that must have been hoarded by OpenAI, Anthropic or Google, so the Chinese must have gathered training data sets by themselves, by crawling the Internet and using pirate libraries, the same like Anthropic and OpenAI.
If the American allegations were true, then some Chinese companies might have extracted some blackbox data about the American models, by API querying, in an amount small enough that it could be used only for the post-training of a model created independently.
Such an information extraction is a breach of contract, but it cannot be called "copying", because it can be only an infinitesimal fraction of the information stored in a SOTA LLM.
They cannot be copies, because all of them have different structures, which are explained in published research papers, while nobody knows the structure of the American SOTA models.
It is also impossible for China to have accessed the hundreds or thousands of terabytes of data that must have been hoarded by OpenAI, Anthropic or Google, so the Chinese must have gathered training data sets by themselves, by crawling the Internet and using pirate libraries, the same like Anthropic and OpenAI.
If the American allegations were true, then some Chinese companies might have extracted some blackbox data about the American models, by API querying, in an amount small enough that it could be used only for the post-training of a model created independently.
Such an information extraction is a breach of contract, but it cannot be called "copying", because it can be only an infinitesimal fraction of the information stored in a SOTA LLM.