The crisis of authorship and the right to know the origin of text in the age of artificial intelligence

🇵🇱 Polski
The crisis of authorship and the right to know the origin of text in the age of artificial intelligence

📚 Based on

Who Wrote This
ISBN: 9781503643574

👤 About the Author

Naomi S Baron

American University

Naomi S. Baron (born September 27, 1946) is a prominent linguist and Professor Emerita of Linguistics at American University in Washington, D.C. She earned her B.A. from Brandeis University and her Ph.D. from Stanford University. Throughout her distinguished academic career, she has held positions at institutions including Brown University, the Rhode Island School of Design, Emory University, and Southwestern University. A former Guggenheim Fellow and Fulbright Fellow, Baron is internationally recognized for her research on the intersection of language, technology, and literacy. Her work extensively explores how digital media, mobile communication, and artificial intelligence reshape human reading, writing, and cognitive processes. She has authored numerous books, including "Always On," "Words Onscreen," "How We Read Now," and "Who Wrote This?," contributing significantly to the understanding of linguistic evolution in the digital age.

Introduction

The rise of generative AI is triggering a crisis in the traditional understanding of authorship. The boundary between human creativity and algorithmic output is becoming fluid, raising critical questions about the right to know the origin of a text.

In this article, we analyze the shift from debates over machine creativity toward the concept of responsible agency. You will learn how legal systems define an author in the age of AI and why content transparency is essential for protecting the public sphere and human cognitive abilities.

Legal Authorship as a Human Normative Construct

Currently, AI systems cannot be recognized as authors under the law. Authorship is a normative construct that assigns rights and responsibilities to specific entities, typically human beings.

In the United States, the Stephen Thaler case reaffirmed the requirement for human authorship. The court ruled that a work must be the product of a human being to be eligible for protection; AI is merely a tool, not a legal subject.

Different jurisdictions adopt different approaches. Polish tradition links a work to the individual character of creativity. Conversely, the United Kingdom allows authorship to be attributed to the person who organizes the generation process of a computer-generated work.

Protection of Expression vs. Prompting Effort

The effort invested in prompting and the subsequent selection of results are not sufficient on their own to grant authorship. The law protects specific expression, not an idea or the time spent operating a tool.

An AI user often lacks control over the details of execution, which distinguishes them from a traditional creator. Selecting a finished result from an AI is a form of curation rather than authorship in a legal sense.

However, protection applies where there is meaningful human control. This occurs when a human makes creative modifications, arrangements, or edits to material generated by the system.

Authorship as a Legislative Choice and Human Creative Freedom

The recipient has a right to know whether content was created by AI. This is the foundation of informational symmetry, allowing one to calibrate their trust in a message based on its source.

The European Union's AI Act introduces transparency obligations. It requires synthetic content to be labeled in a machine-readable format, preventing the act of counterfeiting humanity.

Precise labeling protects against model collapse—the degradation of knowledge that occurs when AI is trained on synthetic data. Transparency allows for the separation of primary data from machine simplifications, which is essential for preserving cultural diversity.

Summary

In a world saturated with synthetic abundance, it is easy to mistake the possession of a text for the thought process itself. However, true authorship is not merely the final result, but primarily the labor and resistance encountered in arriving at a specific sentence.

Preserving writing as praxis is crucial for human autonomy. The value of creativity lies in the process of shaping agency, which cannot be fully delegated to algorithms without risking cognitive degradation.

📖 Glossary

Author-in-fact vs Author-in-law
Rozróżnienie między podmiotem faktycznie generującym znaki (np. AI) a podmiotem, któremu system prawny przypisuje prawa autorskie.
Human authorship
Wymóg, aby utwór był dziełem człowieka, co stanowi obecnie główną barierę w przyznawaniu ochrony prawnej treściom w pełni autonomicznym.
Model collapse
Zjawisko degradacji modeli AI wynikające z trenowania ich na danych syntetycznych, co prowadzi do utraty kontaktu z pierwotnym ludzkim doświadczeniem.
Epistemiczna sprawiedliwość
Koncepcja dotycząca prawa osób do bycia uznanymi za wiarygodne źródła wiedzy i świadectw, zagrożona przez masową podaż treści syntetycznych.
Capability Approach
Podejście Amartyi Sena oceniające rozwój nie przez ilość dóbr, lecz przez realne zdolności człowieka do prowadzenia życia zgodnego z jego wartościami.
Asymetryczna delegacja
Strategia współpracy z AI polegająca na przekazywaniu maszynie zadań rutynowych, przy zachowaniu silnej ludzkiej kontroli nad decyzjami kluczowymi i etycznymi.

Frequently Asked Questions

Can AI systems be recognized as authors of works under current law?
AI systems cannot be recognized as authors of works because the law (using US regulations as an example) requires that a work be primarily the creation of a human. Although a computational system may actually produce content, the status of author is a normative construct that machines are denied.
Is the effort put into creating prompts and selecting AI results alone enough to be recognized as the author of a work?
The effort spent on prompting and selection may not be sufficient for authorship, as law protects the creative expression rather than the labor itself. Protection may be granted if the user makes constitutive decisions, such as specific changes to fragments, determining the final form, or creating an original structure of the work.
Who is the legal author of AI-generated content in light of different legal systems?
In the UK, authorship of a work generated without human intervention is attributed to the person who took the steps necessary for its creation, e.g., the user entering the prompt. In Polish law and the EU system, authorship is linked to human creative activity and free intellectual choices, which precludes the AI system itself from being a subject of copyright.
Does choosing a ready-made result from an AI make the user its author under the law?
No, simply selecting a satisfactory result from an AI does not make the user its author, as the law distinguishes curation from creation. A user may hold rights to a compilation or the arrangement of elements, but they do not automatically gain a monopoly over every element of the generated output.
Does the recipient of a text have the right to know whether the content was written by a human or generated by AI?
Yes, this has become a right for recipients in the European Union, where transparency requirements resulting from Art. 50 of the AI Act will apply from August 2, 2026. This demand stems from the need to counteract information asymmetry and prevent the 'faking of humanity' in synthetic content.
How do legal regulations regarding the labeling of AI content affect the understanding of authorship and responsibility for a text?
AI system providers must ensure machine-readable labels for synthetic content, moving the issue of text origin to the level of technical infrastructure. In the case of content concerning public affairs, disclosure of AI involvement is required unless the material has undergone editorial control and a person or entity bears responsibility for it.
How should the obligation to label AI content be introduced to avoid information noise and actually help the recipient?
Standard proofreading and activities that do not significantly change the content and semantics of the data should be exempt from the labeling obligation. Instead of a binary division, it is worth introducing a common semantics that distinguishes the degree of AI involvement (e.g., from technical correction to full automation) and the possibility of certifying fully human-made content.
Why is information about whether a text was created by AI important for the recipient, and what risks does it entail?
Information about the origin of the content allows the recipient to properly calibrate their level of trust and interpret the message depending on the identity of the sender. The main risk is the possibility of making decisions based on the false belief that a human is behind the communication, which constitutes a form of exploitation of social trust.
What should AI content labeling look like in practice, and is it technically possible to enforce?
AI content labeling is intended to take the form of a layered system, including model-side markings, metadata, provenance registries, and publisher declarations. Full enforcement is difficult due to the possibility of removing tags and paraphrasing text; therefore, legal requirements are adapted to the current state of the art and the scope of feasibility.
Does technical labeling of content as AI-generated solve the problem of authorship and responsibility?
No, technical labeling of content does not solve the broader problem of authorship or moral responsibility for words. It only informs about the participation of AI in text generation, but it does not indicate who conceived the arguments, verified the sources, or bears agency for the content.
How are recommendation algorithms and generative AI changing the functioning of the public sphere?
Recommendation algorithms act as automatic gatekeepers that promote content maximizing engagement instead of using substantive criteria, which can lead to the creation of information environments based on emotion and hostility. Meanwhile, generative AI drastically lowers the cost of producing mass, contextually diverse content, changing the dynamics of the public sphere through the ability to simulate individual presence on a massive scale.
How do bots influence public discourse, if not by directly persuading people of their arguments?
Bots influence discourse by altering the topology of information flow and creating synthetic social proof, which distorts the image of collective beliefs. Instead of directly persuading people, they change users' meta-beliefs about what others think, thereby affecting the willingness to speak and mobilize.
How does the mass production of content by AI affect our ability to conduct authentic public debate?
The mass production of content by AI lowers the cost of flooding the public space with information, which, given the limited attention spans of audiences, can lead to the dominance of statistically visible topics rather than substantive ones. This phenomenon promotes the creation of a hyperactive but epistemically poor public sphere, where an increased volume of utterances does not signify a deepening of debate or an increase in its quality.
How can AI simulate a pluralism of opinions and influence our perception of truth in the public sphere?
AI can simulate pluralism by creating so-called synthetic audiences—an environment of many seemingly independent voices and interactions. By generating mutually reinforcing messages, these systems can create the impression of social consensus, which influences the perception of truth, as people often perceive information repeated by multiple sources as stronger evidence of its credibility.
How do AI infrastructure and platforms change the nature of public communication and the reliability of societal data?
AI infrastructure and platforms change public communication by taking on the role of co-author and distributor of content, creating a system of synthetic feedback loops. This leads to the pollution of the evidentiary environment, where digital data ceases to reflect actual human behavior and instead becomes a product of algorithms.
Why is labeling content as AI-generated important not only for humans but also for information systems themselves?
Labeling content as AI-generated serves as input for other systems, enabling them to distinguish primary content from synthetic content. This prevents the phenomenon of model collapse, in which successive generations of models learn from the errors and averaged outputs of their predecessors, leading to a statistical erosion of knowledge.
What is the phenomenon of model collapse and why is training AI on synthetic data dangerous?
Model collapse is a process where AI models gradually drift away from the real data distribution as a result of training subsequent generations on synthetic data. This is dangerous because it leads to a loss of diversity and the disappearance of rare information (the so-called tails of the distribution), which can result in a lack of representation for marginalized groups and niche aspects of reality.
How does model collapse differ from a standard AI hallucination, and what are its consequences for knowledge?
A hallucination is an error at the level of a single output, whereas model collapse is a generational phenomenon involving the inheritance of errors as training data. This leads to the contamination of the knowledge reproduction apparatus and the creation of a loop in which models learn from representations that are increasingly detached from reality.
Does every use of synthetic data inevitably lead to the degradation of knowledge (model collapse)?
No, model collapse is not inevitable in every system using synthetic data. Degradation can be avoided if new content is cumulatively added to a dataset where original real-world data remains, rather than replacing it.
Can the degradation of AI models resulting from training on synthetic data be prevented, and how will this change the value of human-created content?
Model degradation can be countered through methods such as ForTIFAI, which reduces the weight of tokens typical of generative content, as well as by applying provenance standards and filtering synthetic data. As a result, the value of primary (human-origin) data will increase, as certainty regarding the human origin of information will become a valuable asset in a world saturated with AI content.
Why is model collapse dangerous not only for AI technology, but for all of human knowledge and culture?
Model collapse leads to the loss of rare and atypical information, which is crucial for science, medicine, and cultural development. This phenomenon creates a loop of epistemic recursion, where systems consume their own simplifications and multiply mediocrity instead of relying on real observations and diversity of knowledge.
When does human authorship matter and why should we not delegate every writing process to AI?
Human authorship matters where the process of creating a text constitutes part of its value, builds accountability, testifies to experience, or shapes the capacity for judgment. Not every writing process should be delegated to AI because total outsourcing can weaken human agency and deprive us of the internal goods of practice, such as the development of thinking and cognitive competencies.
When is the use of AI a support for development, and when does it become a dangerous replacement for human abilities?
AI supports development when it is used to delegate routine tasks that are not key mechanisms of professional or personal growth. It becomes a dangerous replacement when it takes over functions that are prerequisites for human autonomy and skills that a person should acquire during the educational process.
How can one use AI to increase productivity without losing the ability to think independently and make decisions?
Functions should be consciously delegated to AI so that cognitive offloading frees up resources for higher-order tasks, rather than replacing key decisions and the practice necessary to maintain competence. The key is using "supertools" systems that combine high automation with a high level of human control and clear decision points.
How can one determine in practice which elements of the writing process can be entrusted to AI and which must remain under strict human control?
AI can be entrusted with routine tasks and the creation of instrumental information, such as summarizing documents. Deep forms of expression and elements with high epistemic, relational, and educational stakes—where an error is difficult to reverse or undermines the nature of testimony—must remain under strict human control.
Why is human authorship important if AI can generate convincingly sounding content?
Human authorship ensures a flow of new, primary data and experiences into culture, protecting it from merely replicating existing representations. It is crucial for building trust in communication and protecting the epistemic representation of individuals whose testimonies could be drowned out by a mass supply of synthetic content.
How can we distinguish between the use of AI that enhances human capabilities and that which leads to the degradation of our cognitive abilities?
This distinction depends on the direction of dependency: AI enhances capabilities when it helps humans see and think more and expands their actual capabilities, whereas it leads to degradation when it allows for the avoidance of thinking and replaces cognitive processes essential for developing judgment. The key is maintaining human agency in areas requiring competence, responsibility, and independent decision-making.
Why do we still need people capable of authorship in a world dominated by AI, even if machines can generate texts?
Society needs people capable of authorship to be able to recognize new facts, formulate problems, and challenge dominant patterns. They are essential for taking real responsibility for content and for describing rare experiences that machines cannot undergo.

🧠 Thematic Groups

Tags: crisis of authorship the right to know the origin of a text generative AI human authorship author-in-fact vs author-in-law creative expression prompting AI Act Art. 50 model collapse faking humanity asymmetric delegation epistemic justice computer-generated works Capability Approach concept