Perspective | Preliminary Legal Analysis of DEEPSEEK Distillation Behavior
Published:
2025-02-06
DeepSeek is like a flash of lightning, illuminating the path for China's AI large models to overtake others, and has caused a huge shock in the international artificial intelligence field. This shock is not only due to its energy efficiency, high performance, and open-source nature, but also because of its innovative algorithms. It is speculated that its algorithms are comparable to the latest internal testing version of ChatGPT, which has led to accusations from leading foreign AI large models, spearheaded by ChatGPT, claiming that DeepSeek has illegally used their model's training data through distillation technology. The Spring Festival holiday has become a joyful ocean for tech enthusiasts. As a lawyer who is likely to be among the first to be replaced by artificial intelligence, I couldn't help but attempt to analyze the accusations against DeepSeek from a legal perspective, regarding the alleged illegal use of their data through distillation technology. Additionally, I will briefly analyze the issues surrounding the protection of AI large model algorithms. This analysis is based on hypothetical scenarios and is limited to academic research. If the feedback is positive, I will delve deeper into this topic.
DeepSeek is like a lightning bolt, illuminating the path for China's AI large model to overtake, and causing a huge shock in the international artificial intelligence field. Its shock is not only due to energy saving, efficiency, and open source, but also because of its algorithm's unique approach, which is speculated to be related toChatGPTThe latest internal test version is equally powerful, and has also led to accusations from leading foreign AI large models, led by ChatGPT, claiming that it uses the formed data of its large model through distillation technology.
The Spring Festival holiday has become a joyful ocean for tech men. As a lawyer who is likely to be the first to be replaced by artificial intelligence, I can't help but try to analyze from a legal perspective the accusations against DeepSeek for breaching (illegally) using the data results of ChatGPT and other companies through distillation technology, and briefly analyze the issue of AI large model algorithm protection. This analysis is based on assumptions and is limited to academic research. If feedback is good, I will explore it in depth.
1. About Big Data Distillation Technology
The origin of big data distillation technology can be traced back to the needs of machine learning and data science, especially in the training of deep learning models. With the explosive growth of data scale, processing massive amounts of data has become a challenge. The core idea of distillation technology is to extract the core features or knowledge of the data, reduce the amount of data while retaining key information, thereby lowering computational costs and improving efficiency. This technology was first proposed by Hinton et al. in 2015, mainly for model compression. By training a small model (student model) to mimic the behavior of a large model (teacher model), the knowledge of the large model is "distilled" into the small model. It was later expanded to the data level, aiming to extract a small but precise subset of data from large-scale datasets that can approximately represent the features and distribution of the original dataset.
Thus, big data distillation technology refers to the technique of extracting core information or knowledge from large-scale datasets to generate a smaller but more information-dense dataset or model. Its goal is to reduce the amount of data or model complexity while maintaining or approaching the performance of the original data or model. It is akin to the brewing process in winemaking, where impurities are continuously distilled away to retain the essence. Through this technology, the knowledge of complex models is transferred to simpler models; or a small but representative subset is extracted from large-scale datasets. Distillation technology has been widely researched and applied in academia and industry in recent years, especially in the fields of deep learning, natural language processing (NLP), and computer vision (CV). Related papers on knowledge distillation and data distillation are frequently published at top conferences (such as NeurIPS, ICML, CVPR, etc.), covering topics such as optimization of distillation algorithms and expansion of application scenarios. Many tech companies (such as Google, Facebook, OpenAI, etc.) apply distillation technology to model compression and deployment to improve model efficiency and scalability. Some implementations of distillation technology have been open-sourced, with tools and examples for knowledge distillation provided in frameworks like TensorFlow and PyTorch. The distillation technology and the data it generates have significant commercial and technical value, leading major AI companies or platforms to adopt technical protection measures such as encryption, watermarking (anti-counterfeiting measures), access restrictions, and model protection.
In addition to the above technical protection measures, industry-leading AI companies also apply for patents on distillation technology; protect the code, models, and data generated during the distillation process through copyright; and regulate and constrain through technical agreements. Additionally, they sign data usage agreements with partners or users to clarify the scope, limitations, and responsibilities of data use; and sign confidentiality agreements with employees and partners to prevent technical leaks.
2. Preliminary Legal Analysis
1. From the perspective of copyright
According to the Berne Convention andthe WIPO Copyright Treaty, original works (including software and data) are protected by copyright. However, works or data generated entirely by artificial intelligence, without the involvement of natural persons, cannot be protected under relevant US or EU laws and case law, as these works lack authors. This is also the mainstream view in current Chinese judicial practice. Assuming that the data results of companies like ChatGPT have been calibrated or participated in by humans, they are protected works as stipulated by the aforementioned international conventions. If DeepSeek uses the data results of companies like ChatGPT without authorization, it may constitute copyright infringement.
2. From the perspective of patent rights:
If the technology or data processing methods of companies like ChatGPT have been patented, assuming that DeepSeek uses the patented distillation technology without permission, it may involve patent infringement.
3. From the perspective of contract law and anti-competitive law
If there is a contractual relationship between DeepSeek and companies like ChatGPT, and the contract explicitly prohibits the use or reproduction of data results, then assuming that DeepSeek distills the data results for its own large model may constitute a breach of contract. At the same time, unauthorized use of others' data or technology in the field of artificial intelligence, or using distilled data from competing companies for its own products, especially if that product competes with the products of the data source company, may constitute unfair competition.
4. From the perspective of data protection law:
If it involves EU user data, assuming that DeepSeek's distillation behavior may violate the General Data Protection Regulation (GDPR), which has strict regulations on the collection, storage, and use of data. Developed countries like the US also have similar regulations or cases regarding data protection, which I will discuss one by one after collecting and organizing.
5. The Agreement on Trade-Related Aspects of Intellectual Property Rights (TRIPS)
This agreement requires member countries to protect intellectual property rights, including copyright and patent rights. Although the WTO is currently being abandoned by countries led by the US, assuming that DeepSeek's actions may violate the TRIPS agreement, it may also be submitted to the WTO for dispute resolution. If companies like ChatGPT have bilateral or multilateral agreements with the country where DeepSeek is located, relevant agreements may have specific provisions on intellectual property protection and data use, which may also be cited.
6. From the perspective of industry practices:
In the tech industry, open source and sharing technology are common practices, but they usually need to comply with specific licenses (such as GPL, MIT, etc.). If DeepSeek does not comply with relevant licenses, it may violate industry practices or user agreements.
Therefore, if DeepSeek is accused of breaching (illegally) using data from companies like ChatGPT through distillation technology, it is likely to lead to the largest legal dispute case in the AI field in 2025. Due to the complex legal, industry, and technical issues involved, and the delicate state of US-China relations, this issue presents the greatest opportunities and challenges for both the legal and tech communities. This article is merely a legal analysis based on assumptions, and hopes to achieve a win-win situation from the perspective of preventing the abuse of intellectual property and promoting technological progress.
Key words:
Related News
Zhongcheng Qingtai Jinan Region
Address: Floor 55-57, Jinan China Resources Center, 11111 Jingshi Road, Lixia District, Jinan City, Shandong Province