Japan is moving toward greater transparency in the artificial intelligence industry, with the government preparing new guidelines that would encourage generative AI companies to disclose information about the data used to train their models.
The move comes as concerns continue to grow over whether AI companies are using copyrighted books, articles, artwork, photographs, music, films and other creative works without permission. Japan’s proposed framework aims to give creators and rights holders greater visibility into how their work may be collected and used by AI developers.
The proposed rules are not designed as a strict legal mandate. Instead, Japan plans to use a “comply or explain” approach, under which AI companies would either follow the guidelines or publicly explain why they have chosen not to do so.
The framework is expected to apply not only to Japanese AI companies but also to overseas developers providing AI services in Japan.

Japan Targets AI Training Transparency
Generative AI systems are built by training models on enormous quantities of information. Depending on the system, those datasets can contain text, images, videos, websites, books, software code and other forms of digital content.
The training process allows AI models to recognize patterns and generate new material based on what they have learned.
However, the scale of modern AI training has made it difficult for creators to determine whether their work has been included in a particular dataset.
Japan’s proposed guidelines seek to address that problem by encouraging AI developers to publish information about how they collect training data and the general scope of the information used.
The objective is not necessarily to force companies to publish every individual piece of data used to train a model. Instead, the focus is on giving the public and rights holders a clearer picture of where training material comes from and how it is collected.
Copyright Concerns Drive New Rules
Copyright is one of the biggest issues behind Japan’s proposed framework.
Authors, artists, publishers, musicians and other creators have increasingly questioned whether AI companies should be allowed to use their work to train commercial models without permission or compensation.
The debate has become especially important as generative AI systems have become capable of producing sophisticated text, images, music and other creative content.
For creators, one of the biggest challenges is determining whether their work was included in an AI model’s training dataset in the first place.
Without access to information about training data, it can be difficult to investigate potential copyright violations or determine whether a particular AI-generated output may have been influenced by copyrighted material.
Japan’s proposed disclosure system could provide rights holders with a new way to seek that information.
Rights Holders Could Request Information
Under the proposed framework, AI businesses would be encouraged to respond to certain requests from copyright holders.
For example, if a creator believes that a particular webpage or copyrighted work was used to train an AI model, the company could be expected to provide information about whether that material was included, subject to the conditions established under the guidelines.
The approach could make it easier for creators to investigate suspected misuse of their intellectual property.
At the same time, the government is expected to recognize that AI companies cannot reveal every detail of their training processes. Training datasets can contain enormous amounts of information, while data-collection methods and model-development techniques may also represent commercially sensitive information.
Japan’s challenge will therefore be finding a balance between transparency and protecting legitimate business interests.
Pirate Websites Also Under Scrutiny
Japan is also urging AI companies to avoid collecting training material from websites that distribute pirated content.
The issue is particularly significant for Japan because the country has some of the world’s largest manga, anime, publishing, music and entertainment industries.
Creators and rights holders in these sectors have expressed concerns about unauthorized use of their work in AI systems.
The proposed guidelines would encourage developers to take steps to prevent their training processes from relying on obviously illegal sources.
That does not necessarily mean that every copyrighted work found online would be prohibited from being used in AI training. Instead, the government is attempting to establish clearer expectations around responsible data collection and respect for intellectual property rights.
“Comply or Explain” Model
One of the most notable aspects of Japan’s approach is that the proposed code would not operate like a conventional law with mandatory penalties for every violation.
Instead, companies would be encouraged to comply with the principles or explain why they have decided not to follow them.
This “comply or explain” model reflects Japan’s broader approach to AI governance, which seeks to encourage innovation while introducing safeguards around emerging risks.
The approach could also make it easier for Japan to regulate a rapidly changing industry without imposing rules that become outdated as AI technology evolves.
However, the effectiveness of the system could depend heavily on how companies respond.
Critics may argue that companies could simply provide explanations for refusing to disclose information, potentially reducing the practical impact of the framework.
Supporters could counter that public explanations would still create greater accountability and allow users, creators and policymakers to compare how different companies handle training-data transparency.
Global AI Industry Faces More Scrutiny
Japan’s proposal is part of a broader international movement toward greater transparency in AI development.
Governments are increasingly examining not only what AI systems produce but also how they are developed and what information is used to build them.
That shift could create new compliance challenges for global AI companies.
A developer operating in multiple markets may eventually have to provide different levels of information about its training data depending on where its services are offered.
For major AI companies, maintaining detailed records about data sources could therefore become increasingly important.
Companies may need to know whether information was publicly available, licensed, collected through automated systems or obtained from other sources. They may also need clearer procedures for responding to copyright complaints and information requests.
A New Era of AI Accountability
Japan’s proposal comes at a critical moment for the AI industry.
Generative AI is rapidly becoming part of everyday life, with businesses, students, creators and consumers using AI systems for writing, research, programming, design and entertainment.
As adoption grows, questions about the origins of the information powering these systems are becoming harder to ignore.
Japan’s proposed framework could establish a new expectation that AI developers should be more transparent about the data behind their models.
The rules are still being developed, and their nonbinding nature means their ultimate impact will depend on how companies respond.
Nevertheless, the proposal sends a clear message: AI innovation cannot be separated from questions about intellectual property and the rights of creators.
For AI companies, greater transparency could mean additional compliance work and potential exposure to copyright disputes. For creators, it could provide a valuable tool for understanding how their work is being used.
If Japan’s approach proves effective, it could also influence AI policy beyond the country, adding momentum to a global push for greater accountability over the data used to train increasingly powerful artificial intelligence systems.




