Do scientists know what is written in the intellectual property of Generative AI models? It’s a challenge for institutions, publishers and academia to understand the impact of the Industrial Revolution
New laws will ultimately establish more robust expectations around ownership and transparency of the data used to train generative AI (genAI) models. There are some precautions that researchers can take to protect their intellectual property.
Ed said opt-outs were unfair and that some may not even know when they are offered. “It’s particularly good to see the DPA calling for opt-ins,” he says.
Some authors are not sure if their work will be fed into an artificial intelligence system. Edward Ballister is a cancer Biologist at Columbia University in New York City and he says that he doesn’t feel confident in knowing how Artificial Intelligence might affect him or his work. To be open and transparent about their plans is a responsibility that institutions and publishers have.
The burst of artificial intelligence technology and the fact that there are clear answers to questions such as where it falls under existing copyright legislation, who owns it, and what you have to consider when you feed data into your models are only starting to catch up with international policy. There is fast technological developments but legislation is lagging, according to a legal scholar in Rome. “The challenge is how we establish a legal framework that will not disincentivize progress, but still take care of our human rights.”
Tudorache sees the act as an acknowledgement of a new reality, in which artificial intelligence will stay. The industrial revolutions of previous times have profoundly affected different sectors of the economy and society, but none of them have had the powerful impact on society that he thinks is going to be caused by artificial intelligence.
Academics often sign their IP over to institutions or publishers, giving them less leverage in deciding how their data are used. But Christopher Cornelison, the director of IP development at Kennesaw State University in Georgia, says it’s worth starting a conversation with your institution or publisher if you have concerns. These entities should pursue litigation when there is likely to be an violation of the license agreement with the company. He says that they don’t want an acerbity with their faculty and that they’re working towards a common goal.
Scientists can now detect whether visual products, such as images or graphics, have been included in a training set, and have developed tools that can ‘poison’ data such that AI models trained on them break in unpredictable ways. “We basically teach the models that a cow is something with four wheels and a nice fender,” says Ben Zhao, a computer-security researcher at the University of Chicago in Illinois. Nightshade is a tool that works on the basis of an artificial intelligence model associate a corrupted pattern with a different kind of image, such as a dog. Unfortunately, there are not yet similar tools for poisoning writing.
Specialists broadly agree that it’s nearly impossible to completely shield your data from web scrapers, tools that extract data from the Internet. There are steps that can add an extra layer of oversight, for example making resources open and available only by request, or hosting data on a private server. Several companies, including IBM, allow their customers to create their own bot with their own data, that can be isolated in this way.
It may feel like it’s missing out on a golden opportunity if you don’t use Genai. But for certain disciplines — particularly those that involve sensitive data, such as medical diagnoses — giving it a miss could be the more ethical option. “Right now, we don’t really have a good way of making AI forget, so there are still a lot of constraints on using these models in health-care settings,” says Uri Gal, an informatician at the University of Sydney in Australia, who studies the ethics of digital technologies.
Other publishers, such as Wiley and Oxford University Press, have brokered deals with AI companies. Taylor & Francis, for example, has a US$10-million agreement with Microsoft. The Cambridge University Press is developing policies that will give authors an ‘opt in’ agreement, in which they will receive remuneration. The managing director of academic publishing at the CUP, who is based in Oxford, UK, told The Bookseller that the company will be looking at more than 300 research journals and over 24,000 e-books.
Springer Nature and the American Association for the advancement of science, which publish the Science family of journals, have not entered such agreements, according to representatives. Nature is editorially independent of its publisher, which is Springer Nature.
Two studies this year have found evidence of widespread genAI use to write both published scientific manuscripts3 and peer-review comments4, even as publishers attempt to place guardrails around the use of AI by either banning it or asking writers to disclose whether and when AI is used. Legal scholars and researchers who spoke to Nature made it clear that, when academics use chatbots in this way, they open themselves up to risks that they might not fully anticipate or understand. “People who are using these models have no idea what they’re really capable of, and I wish they’d take protecting themselves and their data more seriously,” says Ben Zhao, a computer-security researcher at the University of Chicago in Illinois who develops tools to shield creative work, such as art and photography, from being scraped or mimicked by AI.
When contacted for comment, an OpenAI spokesperson said the company was looking into ways to improve the opt-out process. “As a research company, we believe that AI offers huge benefits for academia and the progress of science,” the spokesperson says. We understand that some content owners don’t want their publicly available works used to help teach our Artificial Intelligence, which is why we offer ways for them to opt out. We’re also exploring what other tools may be useful.”
The technology underlying genAI, which was first developed at public institutions in the 1960s, has now been taken over by private companies, which usually have no incentive to prioritize transparency or open access. As a result, the inner mechanics of genAI chatbots are almost always a black box — a series of algorithms that aren’t fully understood, even by their creators — and attribution of sources is often scrubbed from the output. This makes it difficult to know what went into a model’s answer to a prompt. Users have been asked to make sure outputs used in other work do not violate intellectual-property and copyright regulations, or divulge sensitive information, such as a person’s location, age, ethnicity or contact information. Studies show that genAI tools are capable of doing both.
“There’s an expectation that the research and synthesis is being done transparently, but if we start outsourcing those processes to an AI, there’s no way to know who did what and where the information is coming from and who should be credited,” he says.
Timothée Poisot, a computational ecologist at the University of Montreal in Canada, has made a successful career out of studying the world’s biodiversity. Poisot is hoping that his research will be one of the others considered later this year when the 16th Conference of the Parties (COP16) to the United Nations Convention on Biological Diversity is held. “Every piece of science we produce that is looked at by policymakers and stakeholders is both exciting and a little terrifying, since there are real stakes to it,” he says.



