Rebuilding the Digital Commons for the AI Era

How Creative Commons is reimagining the infrastructure behind the world’s information, creativity, and knowledge 

Creative Commons is an international nonprofit dedicated to defending and nurturing a commons of shared knowledge and culture. Their pioneering copyright licenses and public domain tools are attached to billions of pieces of digital works, granting permission to reuse a work based on pro-social conditions, like giving credit to the original creator. The central premise of all of these legal tools is that creators need to have choices about the conditions under which they share their work. The rise of generative AI and the rising scale of machine reuse of data has not changed that fundamental premise, but it has changed how creators think about sharing and the tools needed to nourish the digital commons.

We sat down with Anna Tumadóttir, CEO of Creative Commons (CC), to discuss the current landscape of the digital commons; how CC is responding to these changes by developing new infrastructures and tools; what it will take to strengthen the commons and encourage knowledge sharing; and the very high stakes for getting this right. 

Before we talk about Creative Commons, let’s start with a basic question. What exactly is the “commons” in the context of the internet and why is it so important to nourish and protect the commons?

To take a step back, “the commons” is a general term for shared resources in which each stakeholder has an equal interest. Think: shared water sources, public parks, or natural resources. The commons in the context of the internet is where digital works and knowledge are shared freely and intended for reuse. 

CC was founded in 2001, and the first set of licenses was released in 2002. Before then, the digital commons was largely anything NOT protected by copyright that could be used freely — that is, public domain materials or materials that weren’t copyrightable in the first place. 

CC licenses allow for a more robust and voluntary human-powered commons, where people choose when and how to share knowledge online, across geographies and cultures. 

Our licenses are designed to sit in this space between all rights reserved copyright on one hand, and public domain materials on the other hand, hence the term “some rights reserved” that was coined by Glenn Otis Brown, who served as executive director of the organization early on and now is vice chair of our board of directors. Instead of an individual needing to give special permission every time someone wants to reuse their work, they are able to share reuse rights simply and easily with everyone in the world. 

CC and the licenses that it developed to allow that sharing of knowledge really created the commons that we’ve had for the last 25 years. As we’ll discuss, things are changing on the commons front, but that notion of balancing the sharing of rights with creative expression remains.

Licenses have been core to Creative Commons’ work since its inception. How do the CC licenses seek to balance copyright with openness and freedom of expression?

The CC licenses and public domain tools that we have today build on four primary conditions: attribution, sharealike, noncommercial, and no derivatives.

Attribution is also known as CC BY. That’s the most commonly used license. What that means is that anyone, anywhere in the world can use this thing that I created, as long as they attribute it to me. For example, if I write an article in Icelandic and license it CC BY, you would be free to translate it into English for a broader audience, as long as you provide attribution. 

The second condition is sharealike. The shorthand for that is SA. The equivalent in the software world is a copyleft or a viral license. It means that you can reuse this work, but you must share it under the same conditions. It’s about continued sharing and ensures a knock-on effect for everybody to continue to benefit from what you put out there. For example, if I take a Wikipedia entry, make modifications, and publish it on my own website, I need to share it under the same sharealike conditions as the original Wikipedia article, because Wikipedia is licensed under CC BY-SA.

There are two other conditions: noncommercial and no derivatives. Both of these conditions, when included in a license, allow people to share their works but under stricter terms. The NC clause restricts commercial use — that is, making money from the remix.  The ND clause means only exact copies can be shared by others.

Those choosing to share can select from six combinations of these four conditions, which makes up the core CC license suite. 

In addition, we have two public domain tools. Most commonly known is CC0 (or CC Zero), which allows anyone marking their work to formally dedicate it to the public domain. This means absolutely no restrictions on reuse, or no rights reserved. The second is PDM (Public Domain Mark), which is used to mark works that are already in the public domain, so that reuse rights are perfectly clear. This is commonly used by cultural heritage institutions.

Because the terms of these licenses and legal tools are completely standardized, there’s no custom negotiation that happens. It means that as long as you have works under compatible licenses, you can translate, modify, remix, reuse, or compile, without having to go through individual legal clearance. That is incredibly liberating for creators, researchers, educators or anyone who believes in open sharing and can benefit from the commons.

How has Creative Commons’ approach evolved in response to the rise of generative AI?

There are two distinct challenges here. First, the infrastructures that power commons are not adequately resourced. This was a problem before generative AI became mainstream. In the case of CC, the whole point of these licenses was to provide a self-service tool, so there is no fee included when using a license. Anyone, anywhere can use the licenses. In addition, business models around open sharing are a challenge at the best of times, and often the infrastructures that support the sharing of open materials — like Wikipedia, or open access repositories, or cultural heritage collections — are sustained through donations or the support of their home institutions.

The critical infrastructure of the commons is hidden and a lot of people don’t think about the level of stewardship required. Again, using our legal tools as an example, in a perfect world, we would have more translations, FAQs that are more in-depth and communicate using different formats, more free training materials, better tech readiness, and more grassroots efforts to get license selection integrated into curricula, publishing workflows, public policies, and beyond. We would have better broadcasted the benefit of the commons. 

When generative AI enters the mix, you have a second challenge, which is that the two origin stories of CC — creator agency and open sharing — are now in tension. 

With generative AI, the folks that associated CC with agency and creator choice say, ‘I need different choices today. I don’t want my work being used without me getting credit. I don’t want to be replaced. I don’t want my work being taken out of context.’ A lot of these concerns come back to wanting acknowledgement for the thing that you put in, and having some sort of reciprocity back in the ecosystem. 

At the same time, you have a separate group of people who say the whole point was always openness. They say, ‘This is great. We can learn from all the world’s knowledge now. End goal reached.’

Of course, I am exaggerating for storytelling purposes. Most people sit somewhere between these two poles, but CC aims to serve both. The organization was founded to build in nuanced gradients where everything was previously binary. 

We need to rethink the gradients. They’ll look a little bit different, just as the digital environment has changed, but the goal is ultimately still to encourage people to choose to share. Because if people do not choose to share, then there will not be a thriving commons and that is, quite simply, bad for everyone. 

AI systems are built on publicly available data. So the commons enabled this technological change to occur quite rapidly. Yet the governance of data — the choices available to those sharing data, the understanding of what has been used where and for what purposes — has not kept pace with rapid development. As a result, some people feel that this is happening to them and not with them. People don’t feel they have a say in what is happening. 

What needs to happen in order to change that dynamic, to recreate the commons for the generative AI era? 

We need mechanisms for attribution, transparency, and honoring the intentions of creators. 

If those had been established norms at the onset of this tech wave, that would have been the best outcome. If we can build those mechanisms into licenses and tools, it would go a long way towards a more reciprocal commons and toward closing the large schism that I described. 

The application of copyright law is very frequently unclear when it comes to whether or not materials can be used for AI systems at-large, whether training or retrieval augmented generation (RAG) or anything in-between. The licenses certainly speak to machine reuse, but I don’t think anyone ever imagined the scale that we have now. The licenses were built with human reuse in mind and formed the basis of a social contract. 

The licenses are built on top of copyright, and for humans all over the world they work the same, but copyright limitations and exceptions vary widely across the globe when it comes to AI. This means that if you are going to use material for machine learning purposes, you have to know what jurisdiction you’re in, and what you can and can’t legally do in that jurisdiction.

Our challenge today is this: how do we take the same components and spirit of the licenses and make that the standard norm or expectation? Attribution is probably the most salient and well-known example, but there are others.

You can imagine how complicated it can be. In a world where you have an agent interacting with another agent because they want to retrieve some information, it’s really important to have clearly documented what is accessible, where it came from, or who the original source was, so that these agents know what they can and can’t do, but also so they can surface the provenance of the information to the eventual human user. 

If an agent is assigned to research XYZ subject and approaches different repositories, the agents of those repositories need to know whether they have the information, whether they can give it to the first agent, and under what conditions. For example, you can only use it if you indicate where it came from, or you can only use it if you are an independent researcher, you’re in this geography, or whatever the future rules of the road are. We’d like to simplify knowledge sharing once again.

What are the practical applications of developing such rules and infrastructure for our current moment?

Here’s one example around attribution. It does not make sense to me that if you query a chatbot today and you ask for information that you wouldn’t also get sources cited. Currently there’s no legal obligation to cite sources or explain how an interpretation was reached. I see some of the mainstream products developing in that direction, but it really needs to become standard. The chatbot might explain that the answer comes from community-created knowledge like Wikipedia or from publicly-funded research, for example.

Knowing where things come from, giving credit, acknowledging effort, should be the norm. The stakes are very high. The fundamental trust in information spreads to societal trust, and then spreads to democratic trust. That all goes away if you don’t know where a thing came from. We broke down this topic in detail recently on our blog, in collaboration with the fine folks from OpenMined.

CC is exploring approaches that extend beyond copyright, focusing on infrastructure and power dynamics. Why have you made this shift?

I don’t think that we should limit our imagination just to an existing legislative framework, because those take a long time to catch up to technological and societal change. I think it is our job in this moment to say, ‘What is the world that we want to see ten years from now or twenty years from now? How do we want our grandchildren to research and understand knowledge and understand one another and collaborate with people? What is the tooling of the future that enables that?’ 

The tooling today builds on copyright frameworks, but until those are globally harmonized and caught up to speed and developed in the direction that I think we’d want, we can’t afford to wait. We need to go out and make things like attribution the norm. An interesting analogy from the early internet is the robots.txt protocol, which signalled to web crawlers whether and how they could crawl a website. It’s the handshake of the internet and is not legally binding. Attempts are being made to leverage that structure but are proving challenging and won’t suffice for all types of asset sharing.

We have to be creative when developing these new infrastructures. There are going to be technical components to this, and it is really important that those are built with public interest protections in mind. 

I’ll give you a concrete example. It is possible to put a website online and block all bot and crawling traffic. If you do that, it becomes harder for anything but a human clicking around at a human speed to interact with your information. There are systems administrators that are doing this for sites because the barrage of bot traffic is so high, the costs have gone up so much, or perhaps they’re just philosophically opposed to their work being crawled. 

But what does that mean for someone who’s doing legitimate research? What does that mean for someone who has accessibility needs and maybe needs that information read out loud to them, or transposed in some way? 

In one case, a language research institute had to put some measures in place because they were getting hit by too much crawler or scraping traffic. Inadvertently, the settings were so rigorous that it blocked the local school district from being able to access the various datasets the research institute hosted, because the school district had shared IP addresses with a thousand students logging on to the same location. That was identified as a suspicious traffic pattern and the students weren’t able to access these educational materials. 

There are fundamental uses that you should never be able to disallow based on what’s in the public interest. The bottom line is that we need to figure out what the non-negotiables are when it comes to sharing information and develop ways to protect those. You really can’t protect them through legislation when technology quietly overrides it. It would be a game of whack-a-mole.

What tooling is CC developing in order to address these issues? 

All content is data. We need to figure out how to prepare that data so that it is ready to actually be understood by a crawler or by an agent, and for them to know how that data can be reused. Licensing information or preference signals would go along with the reading of the data. There’s a future world where there are tools specifically designed to address this interaction. 

The ideal is that people continue to share, but they won’t if they don’t feel like the conditions are in their favor. We’re already seeing a retrenchment happening. We’re seeing an increasing use of our more restrictive licenses, which is no good for anyone. We need to close this gap between feeling like anything you put online is fair game for anyone to use in any way or feeling that the other choice is that nobody ever sees this thing that I created. Filling in the choice gaps is a problem the organization wants to solve. 

Ideally, a year from now we will have an experimental legal tool in the hands of a set of pilot adopters, to facilitate the relationship between the datasets that these stewards hold and AI use for training, inference, or other uses. That testing will allow us to see what the potholes are as we develop a tool to restore balance to the commons ecosystem. We’re going to do that work as much in the open as possible. 

What’s your roadmap for the future of CC? What are some of the concerns that you have in realizing that vision?

We need to make sure that we have sufficient financial backing for this work. It’s not a small task and we need to make sure that we have the resources to pull it off. We are in our 25th anniversary year and a big focus is growing the number of organizations who are part of CC’s Open Infrastructure Circle, especially engaging with those who use our legal tools as part of their baseline operations, such as academic publishers with open access program, philanthropic foundations with open access policies for grantees, and platforms that host independent, openly licensed creative works.

After that, we need to make sure that machines engage with the commons in ways that are transparent and aligned with the public good. Reciprocity is the key foundational value in allowing the commons to continue to thrive and be accessible. 

I am excited when we get to imagine creating this next generation of infrastructure of sharing. It’s such a mouthful — the legal infrastructure of sharing — but that’s what the licenses are. It’s magical knowing that you could put something out into the world, and without any friction, without any interaction, someone clean on the other side of the world could find it and say, ‘This is awesome, I’m going to build on it.’ It has an impact. It gives them some inspiration. It relieves some resource constraints. 

But I am kept up at night by the flipside of that possibility: What if it doesn’t work? What if in a decade we don’t have a commons, or it doesn’t look the way that it used to look? If people don’t share, there’s less knowledge, and if there’s less knowledge, then education fails, and if education fails, we don’t trust one another as much because we forget how to interact with one another. At that point we’ve lost the glue that holds society together. This possible future really illustrates why this work is so important and why we need to devote resources to getting it right.

—-

Anna Tumadóttir is CEO of Creative Commons where she previously served as COO and Director of Product. Before joining CC, Anna spent a decade building out all aspects of the operations of three distributed performance marketing start-ups. Anna grew up in Iceland, Scotland, and Malawi, leaving home to study in Norway, the United States, South Africa, and back to Iceland. She now lives in the United States.