Not Plausible, But Verifiable: Scholarly Metadata in the Age of AI

Not Plausible, But Verifiable: Scholarly Metadata in the Age of AI

Toby Steiner  & Vincent W.J. van Gerven Oei

Preprint; to appear in tendencia editorial UR – Nº 41 Especial · Universidad del Rosario, Bogotá, 2026

Introduction

Across scholarly communications, artificial intelligence is being touted as both a diagnosis of and a miracle cure for the wide variety of issues plaguing the sector. The diagnosis is familiar: too much information, too little time; too many records, too few cataloguers; too many systems, too little interoperability. The proposed cure is automation at scale—automation facilitated through AI.

The massive increase in higher education institutions’ (HEI) spending on the technology (UNESCO, 2025) points to a rather unique persuasiveness of this framing by commercial vendors over the last few years, particularly among upper management executives without direct access to technical expertise: to them, AI is being sold as the one “magic bullet” to fix the many underlying problems they had already been persuaded to invest in over the previous years. Here, the pivot to AI can be understood as just the next step in a larger move to fully outsource HEI infrastructure to commercial vendors (see, e.g., Bowie, 2025).

Meanwhile, the underlying plethora of problems becomes even more severe: metadata work remains under-resourced. Small and scholar-led publishers commonly lack the technical capacity of large commercial firms. Library and publishing workflows are fragmented across systems, standards, and platforms. Books and chapters remain particularly difficult to describe, distribute, discover, and preserve consistently.

Yet the promised all-purpose solution of AI conceals a more consequential set of questions. What kind of automation is being normalised? Who owns and governs the systems that enable automation? What costs—financial, environmental, and social—disappear behind the shiny interface of an AI agent? Whose actual labour is displaced, subordinated, or rendered invisible? Who looks after and is ultimately responsible for code churned out as AI slop (Weedon et al., 2026), which some have declared our societies’ modern asbestos, being shoveled into the walls of our collective digital ecosystems (Doctorow, 2025)? What happens when public and community-produced knowledge becomes an input to proprietary systems that cannot be inspected or governed by the communities from which that knowledge originated?

For Thoth Open Metadata, the non-profit we work for, these are not abstract concerns or mere philosophical ponderings. The underlying issues, one might argue, are very much about the overall purpose, ownership, and future of scholarly infrastructure.

To provide a bit of context on what we do: Thoth is a non-profit, community-led open platform for managing and disseminating fully open, reusable, high-quality metadata for books and chapters. It was developed by and for publishers, with a particular focus on open-access titles and an overarching aim to support small, scholar-led, university, and library presses in tackling a variety of structural barriers to participation in the book supply chain. Thoth’s legal form, a UK Community Interest Company limited by guarantee, reflects our explicit commitment to community benefit rather than shareholder value. Thoth’s technical architecture is open source; its Application Programming Interfaces (APIs) are open; and its metadata is released into the public domain (via a CC0 dedication made explicit within the metadata records) to facilitate open reuse across libraries and other stakeholders active in the book supply chain.

This positionality also shapes Thoth’s approach to AI. For us, the issue is not whether automation can assist with metadata work. It already does. The more important question is whether it can strengthen communities and professional agency or reorganise them around the profit-making priorities of extractive proprietary systems. We are strongly opposed to reducing the future of metadata and cataloguing to yet another layer in an extractive corporate AI stack. Instead, we argue, AI tools should form part of a broader public knowledge infrastructure that is transparent, bibliodiverse, multilingual, interoperable, reusable, and accountable to the communities that sustain it.

Hence, for us, imagining an alternative approach to the current commercial AI regime would be guided by the following set of rules:

  • Not a single universal platform, but a trusted ecosystem.
  • Not labour subordinated to machines, but technology governed to support professional and community judgement.
  • Not platform dependency and expansion through capture, but interoperability through open APIs and collaborative relationships.
  • Not proprietary data assets and extraction, but open, reusable, and verifiable public-domain metadata in reciprocal exchange.

In the following sections, we will expand on each of these points by discussing four major issues that can be identified under the current regime of commercial AI provision. Following this discussion, we will outline how the work we do within the broader publishing community might provide an alternative, bibliodiverse and equitable model that serves the communities it is governed by, rather than stakeholders’ pockets.

Issue 1: AI, Politics of Scale, and Questions of Sustainability

Recent commercial AI is organised around an expansive and consolidating model of scale. Larger models require larger datasets, greater computational capacity, more data centre infrastructure, and deeper concentrations of capital. Scale here functions simultaneously as a technical strategy and a competitive weapon. In this environment, scale does not simply mean serving more users. It often means consolidating control over the infrastructures through which people work, publish, search, and communicate. Once functions are bundled together, institutions become progressively dependent on a small number of providers.

The concentration of AI provision therefore can well be understood as the logical continuation of a longer process of platformisation (Poell et al., 2019) in scholarly communication. The already-familiar issues of vendor lock-in, opaque analytics, restricted metadata reuse, and extractive work linked to maintaining metadata quality in the context of closed walled gardens have not disappeared. On the contrary, AI is very likely intensifying them by adding another proprietary front-house interface and smoke screen between users and these complex issues (see also Van Gerven Oei 2026).

A practice tied to the neoliberal notion of “scaling up,” which we see rampantly at play with the major commercial AI providers, is the provision of services that can be sustained even while remaining heavily indebted, environmentally damaging, socially extractive, or politically vulnerable. A commercial infrastructure may appear stable because it has substantial financial backing propped up by circular funding (Down & Milmo, 2026), but it can still expose communities to acquisition, strategic redirection, service withdrawal, sudden price changes, or revised data policies.

For open infrastructure, sustainability must include governance, technical openness, succession, community accountability, and the ability of users to continue their work if an organisation changes or disappears.

This has direct implications for the use of AI. Uptake of a heavily cross-subsidised AI service may appear inexpensive today, but it will also create deep, long-term dependencies. Its prices, environmental practices, access terms, or data policies may change once institutions have reorganised their workflows around new profitability targets. Commercial AI provision also externalises costs. Users absorb the labour of checking and correcting outputs; open infrastructures absorb the bandwidth and server costs generated by automated harvesting through bots and agents; communities whose work informs models may receive neither recognition nor material support.

Tied to the above, the overall rhetoric around AI encourages users to experience it as frictionless. A prompt is entered; an answer appears. The underlying infrastructure becomes opaque, yet AI remains a materially resource-intensive technology. Its provision depends on massive data centres, global semiconductor supply chains, public electricity grids, environmentally damaging cooling systems, increasingly scarce water resources, physical land, and continuous hardware replacement. As recent research by Aczel et al. (2026) argues, AI’s environmental impact must be assessed across carbon, water, and land footprints, rather than through electricity use alone. The report emphasises that inference, i.e., the routine use of deployed systems, can become a major source of resource consumption at scale.

Investigations into data centre expansion have also drawn attention to facilities operating or under development in water-stressed regions (e.g., Barratt et al., 2025). This complicates the notion that AI necessarily produces efficiency. A process may appear efficient to the user because environmental burdens have been displaced elsewhere. The same applies to labour. The time required to verify, correct, and contextualise AI output is often omitted from productivity claims. So too is the cost of maintaining defensive systems against aggressive automated scraping. A serious assessment of AI therefore needs to include the infrastructural footprint of its entire stack, from geology to the user interface (see Bratton, 2016). Universities, libraries, publishers, and funders should demand environmental transparency and determine whether a system’s computational intensity is proportionate to the task.

One might argue that, particularly in the field of metadata and cataloguing, the problem space at hand does not require a general-purpose large-scale language model. As we show below, the issues can be addressed more reliably through structured data, validation rules, authority files, controlled vocabularies, and narrowly targeted, purposeful automation.

Issue 2: AI Internalises Profits While Externalising Operating Costs

The consequences of commercial AI’s extractive behaviour are already visible at the infrastructural level. In the books space, aggregators such as OAPEN and the Directory of Open Access Books have reported aggressive bot traffic that has affected platform stability and required traffic-management and bot-protection systems (Davison et al., 2026).

The Public Knowledge Project faces a similar structural challenge across the distributed ecosystem of Open Journal Systems and Open Monograph Press installations. Because these systems are often hosted by universities, libraries, societies, and small publishers, the costs of automated scraping are distributed across institutions with very different technical capacities.

Project Gutenberg provides another example. Its public-domain corpus is openly available and widely useful, but the infrastructure that serves it is not cost-free. Repeated large-scale crawling consumes bandwidth and processing capacity and relies on the labour of volunteers who digitise, proofread, structure, and maintain the collection.

Wikimedia faces the same conflict on an even larger scale. Open knowledge infrastructures pay specialists to curate, store, serve, and protect information, while automated systems create additional demand and may simultaneously redirect users away from the original source. This creates both operational and wider-ranging social imbalances. Operationally, community infrastructures bear the costs of access. Socially, they lose the recognition, participation, and support that accompany direct engagement.

As Van Gerven Oei and Fradenburg Joy (2025) point out, the problem is not automated access per se. Preservation, indexing, and metadata exchange have always relied on machine-to-machine communication. The problem is uncoordinated extraction without restraint, attribution, provenance, or reciprocity. Responsible (re)users should prefer APIs, Open Archives Initiative Protocol for Metadata Harvesting (OAI-PMH) endpoints, bulk datasets, mirrors, and rate-limited feeds where available. They should properly identify themselves, avoid repeated retrieval of unchanged resources, and contribute materially when their use creates high costs, in line with fair-use principles. Open access cannot mean that public-interest infrastructures are expected to subsidise limitless private accumulation.

Issue 3: Extractive Access and a Dissolution of Provenance

Current commercial AI provision challenges basic assumptions about openness. The promise of open access was to enable knowledge to be read, shared, reused, preserved, and built upon. Yet open availability is increasingly interpreted by commercial AI companies as permission for unrestricted extraction.

The controversy surrounding Taylor & Francis’s agreement with Microsoft clearly illustrates the underlying problem. Authors reported that they had not been meaningfully consulted about the use of their work for AI-related purposes (Battersby, 2024). Recent investigations into the use of pirated book collections for model training reveal even more serious failures in consent and accountability (Reisner, 2025). These developments point to the limitations inherent in current implementations of an open knowledge paradigm—openness that focuses solely on a legal definition can be easily co-opted and exploited by actors whose practices are neither reciprocal nor transparent.

The Wikimedia Foundation’s recent discussion of the “social contract of free knowledge” offers a valuable framework for understanding these tensions (Rogers, 2026). Free knowledge depends on more than mere availability; it relies on relationships among contributors, publishers, infrastructures, and (re)users. Attribution, provenance, and the possibility of contributing back to the wider ecosystem are part of the social arrangement that sustains a commons. Commercial AI systems can disrupt this arrangement in several ways. They ingest material without meaningful disclosure, or any disclosure at all. They provide answers without adequate citation. Their system design may keep users inside proprietary interfaces, preventing them from encountering original sources or the communities that maintain them. They may return little to the commons from which they derive value. All of this points not simply to a failure of etiquette; it signifies an illegal transfer of control and power over content away from its originators.

When users receive information through one of the large commercial AI interfaces—be they Gemini, Claude, ChatGPT, or any other recent “silver bullet” offerings available, often integrated wholesale into browser interfaces—without seeing its sources, the provider effectively co-opts the relationship between reader and knowledge producer. The originating author, contributor, publisher, repository, catalogue, or community is rendered invisible (for more on this, see, e.g., Strauss et al., 2025). The commercial AI system appears to be the source, and all the labour involved in creating the information is devalued and rendered void. For us, this is a clear violation of the spirit inherent in the participatory commons that open knowledge represents.

Related to the above, the question of provenance becomes a central issue. In web search, proper attribution is often treated as a secondary concern, added after an answer is generated. In scholarly communication, however, it is a crucial element ingrained in good academic practice. Citation allows readers to evaluate evidence, trace claims, understand context, and identify the people and organisations responsible for a work. It creates pathways between users and knowledge communities.

The same principle applies to metadata. A bibliographic record may combine information from title pages, publisher feeds, authority files, DOI registrations, institutional identifiers, catalogues, and controlled vocabularies, among many other sources. If an AI system synthesises these sources without explaining which assertions came from where, errors become difficult to diagnose and accountability becomes diffuse.

One might argue that a meaningful commitment to tracking provenance should therefore extend beyond providing a generic list of links. For example, it might be advisable for systems to distinguish publisher-supplied metadata from inferred data or from input generated by aggregators and libraries once metadata have been released into the wider supply chain. Ideally, metadata could indicate when records have been mapped, transformed, or enriched, and could also document the vocabularies and authority sources used.

Third-party providers should also disclose which catalogues, repositories, websites, and datasets have contributed to their systems, when the material was collected, and how it was processed. As we at Thoth have repeatedly argued, in line with open data experts (see, e.g., ICOLC, 2023; Courtney et al., 2024), any metadata record should be placed under a public-domain dedication (e.g., CC0) to ensure easy, frictionless downstream reuse, with the relevant licensing statement explicitly included in the metadata record.

Without such measures, commercial actors, and through them, AI, can not only reuse open knowledge but also erase the chain of provenance that makes that knowledge trustworthy.

Issue 4: “No Metadata” Is Better Than “Wrong Metadata”

The use of generative AI can be particularly problematic when plausibility is not sufficient. Metadata and cataloguing depend on verifiable outcomes. A record should accurately identify a specific work, edition, contributor, licence, subject, or persistent identifier. It must be consistent enough to move across systems and precise enough to support discovery, dissemination, and preservation.

Generative AI systems, by contrast, are designed to produce plausible language. That design difference creates a fundamental mismatch with metadata work—a mismatch that has been identified at both functional and ethical levels (Frenzel, 2025). A recent study of ChatGPT even classifies the usability of models such as ChatGPT as “bullshit” (Hicks et al., 2026).

The underlying problem is not merely that large language models (LLMs) sometimes make mistakes; it is that, with AI, producing an authoritative-sounding answer does not equate to generating a verifiable output. In cataloguing, this limitation is consequential. Fabricated ISBNs, licences, contributor identities, or subject terms can quickly propagate into catalogues, aggregator databases, repositories, and knowledge graphs. As records move further from their source, errors become increasingly difficult to identify and correct. One might argue that, in the context of cataloguing and metadata work, having no metadata at all would be preferable to having wrong metadata.

Research on AI-assisted cataloguing has thus called for extreme caution rather than a quick replacement of existing metadata and cataloguing workflows. Experiments at the U.S. Library of Congress found greater potential for AI in relatively straightforward descriptive fields than in areas such as subject and genre assignment, where substantial human review remained necessary (Saccucci & Potter, 2025). An earlier study of AI-assisted MARC21 record creation reached a similarly qualified conclusion (Taniguchi, 2024). In a similar vein, AI has been shown to produce fabricated titles, inaccurate identifiers, missing fields, malformed structures, and invented content (Aycock, 2025). As a consequence, the Program for Cooperative Cataloging (2025) stresses the need for human review to ensure provenance, accountability, and explicit responsibility for machine-generated metadata.

It also seems important to highlight the need to honestly address and assess the verification labour that is becoming necessary to control and oversee output generated by AI. When every generated field must be checked against an authoritative source, AI-driven automation quickly redistributes labour rather than reducing it. Correcting a plausible but misleading record can require more labour and expertise than creating an accurate record directly.

Scaling Small: Seeding a Trusted Ecosystem

Building on the issues above, it may be worth considering a number of central questions governing the use of AI: Who governs the system? What specific task does it perform? Can its outputs and assumptions be inspected? Can users leave the system? Does it strengthen or diminish professional agency? What environmental and social costs does it create?

The concept of scaling small (Barnes & Gatti, 2019; Adema & Moore, 2021; Adema et al., 2024) that informs Thoth’s activities within a wider network of open infrastructures may serve us here and offer a useful counter-positioning. The concept’s core argument does not romanticise smallness or reject the need for coordination. It challenges the assumption that infrastructures must grow through concentration, standardisation, and market capture to be successful (see, e.g., Sutton, 1991).

In contrast, scaling small entails a relational model of collaborative development. Capacity can grow through the federation of interoperable systems that communicate via open data and open APIs. A diverse plurality of organisations supports one another through the mutual sharing of resources and knowledge, all grounded in a shared set of values and participatory governance models. The overall objective is not to construct a single platform capable of absorbing every function and content type, but to cultivate an ecosystem in which specialist infrastructures co-operate while remaining accountable to their communities.

Applied to AI, this distinction is fundamental. A system that helps publishers translate structured metadata into all major industry formats (e.g., ONIX, KBART, MARC21, etc.) in an openly licensed way is not equivalent to a proprietary model that ingests vast amounts of scholarly content under questionable legal frameworks and then generates outputs of uncertain provenance and reliability. Both of these systems may involve automation, but they embody very different assumptions about ownership, control, accountability, and value. One expands distributed capacity. The other effectively concentrates and flattens it.

Non-profit, community-led open infrastructures such as Thoth Open Metadata and the Open Book Collective offer a practical articulation of this alternative. Its call to “seed” a not-for-profit, community-led ecosystem for open access books (Steiner et al., 2025) deliberately rejects the aspiration to build a single dominant platform. “Seeding” here suggests cultivation rather than capture: a landscape of organisations rooted in particular communities and areas of expertise, connected through common values, open standards, and shared infrastructures. Within this ecosystem, publishing, metadata, preservation, collective funding, discovery, and governance need not be controlled by a single organisation. They can be distributed across a network of community-led services that cooperate without merging. This is a non-competitive philosophy in a substantive sense. The success of one infrastructure does not require the disappearance or absorption of another. Organisations can exchange data, coordinate development, share documentation, and direct users toward one another. Their relationships need not be organised primarily through market conquest.

The distinction matters because commercial providers frequently treat integration as a rationale for consolidation. Community-led infrastructures can instead treat interoperability as a condition for plurality. When systems communicate via open APIs, when data can be exported in open formats, and when records can move across systems without proprietary restrictions, cooperation becomes possible without institutional assimilation. Drawing on Ostrom (2009), Thoth’s role in this ecosystem is deliberately “bounded.” It does not seek to become the one-and-only infrastructure for open access books. Its purpose is to provide an open, reusable translational metadata layer through which books and chapters can move among publishing systems, repositories, catalogues, aggregators, knowledge bases, preservation services, and discovery platforms. In this context, boundedness is a strength. It makes interdependence possible without making any one organisation indispensable.

Responsible Community-Led Automation in Practice

Current discourse on the promises of greater automation through AI often collapses all automation into a single category, obscuring important differences among tools, purposes, and governance models. In some contexts, AI-augmented automation might actually prove useful. One example is the development of Annif, an open-source toolkit for automated subject indexing developed by the National Library of Finland that works with controlled vocabularies and selected training collections. Its scope is bounded, its outputs can be evaluated, and its limitations can be studied in relation to particular languages, vocabularies, and domains. The German National Library’s work on automated subject cataloguing (Poley et al., 2025) offers another example of institutionally governed machine assistance.

Such systems operate within professional contexts, are subject to explicit evaluation, and align with established bibliographic responsibilities. These tools differ substantially from general-purpose proprietary models presented as universal solutions. They are narrower, more transparent, and more accountable. Their usefulness does not depend on replacing professional human judgement. This distinction is also central to Thoth’s position. The relevant choice is not between technological enthusiasm and refusal; it is between technologies organised around different institutional values and modes of operation. When AI is used as a tool overseen by a capable specialist and helps them automate tedious tasks, it might actually prove useful.

Thoth’s development work between 2023 and 2026 illustrates how automation can be organised around community needs rather than labour substitution (Steiner, Arias, Gatti, et al., 2026). A central objective has been to reduce duplication of work imposed on publishers by fragmented supply chains. Instead of repeatedly entering and reformatting the same information for different recipients, publishers can maintain rich metadata in a single open record and then generate outputs in multiple standards as needed, including ONIX, MARC, crossref XML, KBART, JSON, CSV, and BibTeX.

By implementing the good metadata practice recommendations outlined in a recent report (Steiner, Arias, Bennett, et al., 2026), persistent identifiers such as ORCID and ROR can be incorporated into these records. Controlled vocabularies such as Thema and BISAC can be applied in structured form. Open APIs enable other services to retrieve and reuse the open data available through Thoth. Providing an OAI-PMH endpoint enables harvesting via an established repository protocol. Bulk-ingest workflows reduce repetitive entry and make it easier to bring existing records into the system. These forms of automation remove avoidable friction while leaving the publisher in control of the underlying metadata.

Commercial AI systems frequently treat linguistic variety as a problem to be overcome by using larger training corpora. To meet the need for ever more data to feed these training corpora, commercial AI companies have recently started to include the wholesale ingestion and subsequent destruction of pre-2022 printed books (see, e.g., Landymore, 2026). If metadata production, translation, discovery, and evaluation are mediated by the same commercial systems, diversity at the level of content risks becoming flattened into uniformity at the level of infrastructure.

Current modes of commercial AI provision in the area of metadata and discoverability intensify this risk. Systems trained predominantly on Western languages, publishing models, and disciplinary conventions are prone to reproducing those priorities, especially when marketed as “universal.” Materials that fall outside the patterns on which AI has been trained may be mistranslated, misclassified, and/or made less visible. Local categories and diverse contributor roles may be forced into schemas that do not adequately represent them.

For books, these risks are especially acute. Books and chapters support long-form, multilingual, edited, experimental, and locally situated scholarship that is particularly relevant to the Humanities and Social Sciences. Their metadata may need to describe multiple contributor roles, multiple relational connections among, for example, a series, title, and its chapters, rights statements, accessibility features, translations, and locally meaningful subject classifications.

It is worth noting that the complexities inherent in long-form publishing are not a defect to be flattened or removed. Quite to the contrary, they are part and parcel of the cultural and bibliographic diversity that defines academic research cultures. Scholarly infrastructure that cannot represent such complexity does more than simply produce subpar, incomplete records. It limits which publications become discoverable.

By contrast, open metadata infrastructure approaches it as a design obligation. Languages, scripts, and regional conventions should be represented explicitly and structurally, not reconstructed probabilistically from the perspective of dominant linguistic patterns. Thoth’s commitment to multilingual metadata and structured bibliographic relationships, informed by its commitment to bibliodiversity (Adema et al., 2024), reflects an effort to preserve underlying diversity at the infrastructural level. The intention is not to treat every publication as an instance of the most common model, but to develop systems sufficiently flexible to represent the beautiful variety that makes up the scholarly record across disciplines.

Thoth’s service model follows the same logic. Thoth Oasis provides free metadata management and open-format exports. Thoth Obelisk adds dissemination, DOI registration, hosting, and archiving; Thoth Sphinx provides privacy-conscious usage statistics; and Thoth Pyramid enables publishers to host websites and integrated catalogues under their own web domains. All of these services integrate and are interoperable, with a rich array of industry standards maintained by multiple communities (e.g., ONIX, MARC) and third-party services often maintained by non-profit organisations, such as Crossref, DataCite, ORCID, and ROR, as well as aggregators and repositories, such as the OAPEN Foundation, the Directory of Open Access Books (DOAB), the Internet Archive, Zenodo, JSTOR, Project MUSE, and the OPERAS Metrics service (Arias et al., 2025). In addition, data and content are submitted to commercial platforms to maximise visibility of books and chapters while still maintaining the open nature of the content and its metadata upstream.

The underlying objective here is not to create yet another expanding proprietary stack. Publishers can use modular services while always retaining control over their identity, records, and workflows. We would argue this is automation without enclosure: the reduction of unnecessary repetition without transferring authority from communities to opaque algorithmic black boxes.

Interoperability Without Assimilation

Discussions about academic books at last year’s Guadalajara Book Fair (Ramalho & Steiner, 2025) underscore why bibliodiversity, metadata autonomy, and interoperability need to be considered together. Participants in a metadata seminar at the Fair identified recurring problems in Ibero-American academic publishing: fragmented records, missing identifiers, incomplete rights information, inconsistent metadata, and weak interoperability among publishers, libraries, indexes, and evaluation systems. As a result, many academic books from Latin America remain poorly represented in discovery environments even when their content is openly accessible (see also the preliminary results of the CSIC-led Cartografía de la edición académica iberoamericana project).

We do not want to propose a solution that assimilates diverse publishing communities into a single global platform. On the contrary, the focus ought to be on developing infrastructures that express local and linguistic specificity while enabling information exchange through shared standards. The respective approaches of SciELO Livros and Thoth illustrate this possibility. SciELO Livros has a long track record of public infrastructure development grounded in Latin American scholarly communication, including book and chapter records, ePub production, DOI management, preservation, and metadata exchange. Thoth addresses related challenges through open metadata management and dissemination.

Here, we see the value of community-led collaboration precisely in the fact that neither system needs to absorb the other. Through interoperable open protocols, infrastructures can connect while preserving regional, institutional, and epistemic autonomy. This is scaling small in practice: difference becomes the basis of cooperation rather than an obstacle to it.

Open APIs are therefore not merely technical conveniences. They are governance instruments. They allow organisations to collaborate without consolidating ownership. In this context, open, reusable public-domain metadata serves a similar purpose and acts as an antidote to the extractive practices outlined above by enabling circulation, subsequent enrichment, and redistribution without tying records to a single proprietary service. Interoperability must, however, be based on trust. Systems need to exchange information in transparent, inspectable, and reversible ways. Participants should be able to identify where data originated, how it has been transformed over time, and under which conditions reuse is permitted. These conditions should then be enforced across the ecosystem so that participants can trust their applicability.

In Need of a Feedback Loop: A Commons-Based Approach to AI

A commons survives because people can contribute to it. This is one of the strongest arguments in the Wikimedia report. AI interfaces often provide no meaningful route for users to correct errors, improve source materials, or support the communities from which the reproduced information originates. In short, much knowledge flows out of the commons; very little flows back. Generative systems can intensify this asymmetry by producing large quantities of plausible but unverified material. Creating a summary, subject description, or catalogue record may take seconds. Verifying it may require specialist expertise and substantial time. The cost of generation is borne by the model provider. The cost of correction is handed over to the communities without prior consultation or any protocol in place to facilitate those mediations.

For metadata, this is particularly damaging. Shared records improve when publishers, libraries, and researchers collaborate to add identifiers, correct affiliations, refine subject classifications, and update rights information. Those improvements should circulate through accountable workflows rather than remain trapped within proprietary products and walled-garden ecosystems. Thoth’s open APIs and publisher-controlled records are intended to support this form of exchange. Fully open metadata released under a CC0 dedication can move freely between systems, be enriched, and return to the wider ecosystem without any single company becoming the permanent owner and thus the controller of the process.

Cory Doctorow’s account of the recent AI bubble further sharpens the distinction between useful technical capabilities and the commercial ideology surrounding them (2026). Doctorow argues that AI is being marketed not merely as software but as a vehicle for replacing labour and capturing part of the resulting wage savings. In Doctorow’s reading of automation theory, “centaurs” can be understood as information professionals who use AI as a productive tool to improve their work practices. His concept of the “reverse centaur” then describes a worker subordinated to cater to the needs of the machine: required to monitor automated output, to correct its failures, and absorb responsibility when AI is wrong (Ouellette, 2026).

The “reverse centaur” metaphor seems a particularly pertinent warning for cataloguing and metadata work. A genuinely assistive system might identify candidate subjects, reconcile identifiers, or detect inconsistencies. A reverse-centaur workflow would increase production targets, reduce specialist staffing, and leave a substantially smaller number of professionals responsible for checking ever-increasing volumes of machine-generated records. The difference lies in who controls the workflow and for whose benefit it is organised.

As a silver lining, Doctorow (2026) suggests that a number of useful tools might be salvaged from the eventual collapse of the current, highly speculative AI bubble. Indeed, open models, commodity hardware, and specialised systems—such as Annif or other specialist workbenches that use AI, e.g., for automatically extracting references—may prove useful once detached from claims of universal automation and staff replacement.

This possibility aligns well with Thoth’s positionality. The aim should be not to reject computational tools but to situate them within open infrastructures governed by community-led values. A useful system will be transparent, inspectable, proportionate, and subordinate to human and community purposes. It will strengthen professional agency rather than reduce people to monitors of machine output. It will operate on open standards rather than deepen proprietary dependence.

Change of Ambitions

The debate about AI is often framed as a choice between embracing and resisting technological progress. A more meaningful choice concerns the technological and institutional future that scholarly communities are prepared to build.

As described above, one path leads to further concentration: a small number of commercial providers control models, interfaces, data, and access (Hao 2026). Knowledge is extracted from open sources, captured and processed by opaque systems, and delivered through proprietary environments. Professional labour is reorganised around machine output. Substantial environmental costs remain largely hidden. Communities become dependent on systems they cannot govern.

The other path is more demanding because it requires network-building rather than adopting a single silver-bullet product. It involves creating open, interoperable, community-led infrastructures that support shared workflows without centralising authority. Thoth’s contribution to this second path is necessarily partial. We provide open-source tools, open APIs, and openly reusable public-domain metadata for books and chapters. Our contribution seeks to strengthen publisher agency, reduce duplicate labour, and enable collaboration among infrastructures rather than replace them.

This approach is grounded in trust—but not blind trust. It is trust supported by transparency: inspectable code, visible provenance, open standards, good metadata practices, documented interfaces, accountable governance—and the belief that publishers should retain control over their metadata and operational infrastructure.

It is also grounded in a different understanding of efficiency. For us, efficiency should not mean shifting costs onto workers, communities, or the environment. It should mean reducing unnecessary duplication while preserving quality, agency, and participation that contribute to the enhancement of shared knowledge across infrastructures.

The emerging, broader ecosystem envisioned through active collaborations among open infrastructures such as Thoth Open Metadata, SciELO Livros, the Public Knowledge Project, Open Book Collective, OAPEN, DOAB, Project Gutenberg, Wikimedia, Pressbooks, the many small- to medium-sized publishers, and an ever-increasing number of academic libraries points toward such a future.

These infrastructures differ in purpose, scale, and history. Their value lies not in becoming uniform, but in strengthening bibliodiversity.

We believe the future of metadata and cataloguing should not be reduced to yet another layer in an extractive corporate AI stack. It should form part of a broader public infrastructure for knowledge: transparent, bibliodiverse, multilingual, interoperable, reusable, and accountable to the communities that sustain it. For open-access books, for free knowledge, and for the infrastructures on which they depend, this is the more ambitious alternative—we feel it is also the one worth fighting for.

References

The bibliography of all sources referenced in this article are also available via an open Zotero collection: https://www.zotero.org/groups/6626908/ai_and_metadata

Aczel, M., Chamanara, S., Matin, M., Farsi, A., Marwala, T., & Madani, K. (2026). The Environmental Cost of AI’s Energy Use: Carbon, Water and Land Footprints. United Nations University Institute for Water, Environment and Health (UNU INWEH). https://doi.org/10.53328/INR26RMA002

Adema, J., & Moore, S. A. (2021). Scaling Small; Or How to Envision New Relationalities for Knowledge Production. Westminster Papers in Communication and Culture, 16(1), Article 1. https://doi.org/10.16997/wpcc.918

Adema, J., Barnes, L., Deville, J., Fathallah, J., Gatti, R., Grady, T., Hopkins, K., Hughes, A., McGann, C., Sanders, K., & Steiner, T. (2024, December 19). The COPIM Perspective on Bibliodiversity. COPIM. https://doi.org/10.21428/785a6451.86a892a7

Arias, J., Van Gerven Oei, V. W. J., Gatti, R., Higman, R., Hillen, H., O’Connell, B, Ramalho, A., & Steiner, T. (2025, October 20). Open Access, Open Data, Open Archiving: Liberating Metadata Flows across the OA Books Landscape. Zenodo. https://doi.org/10.5281/zenodo.17340814

Aycock, M. (2025). Prompting Generative AI to Catalog: The Promise and the Reality. College & Research Library News, 86(09). https://doi.org/10.5860/crln.86.10.423

Barnes, L., & Gatti, R. (2019). Bibliodiversity in Practice: Developing Community-Owned, Open Infrastructures to Unleash Open Access Publishing. ELPUB 2019 23rd edition of the International Conference on Electronic Publishing, June 2019. Marseille, France. https://doi.org/10.4000/proceedings.elpub.2019.21

Barratt, L, Gambarini, C., Witherspoon, A., & Uteuova, A. (2025, April 9). Revealed: Big Tech’s New Datacentres Will Take Water from the World’s Driest Areas. The Guardian. https://www.theguardian.com/environment/2025/apr/09/big-tech-datacentres-water

Battersby, M. (2024, July 19). Academic Authors “shocked” After Taylor & Francis Sells Access to Their Research to Microsoft AI. The Bookseller. https://www.thebookseller.com/news/academic-authors-shocked-after-taylor–francis-sells-access-to-their-research-to-microsoft-ai

Bowie, S. (2025). Small Changes: Taking Back Control of Research through Open Software. Edinburgh Open Research, 4: Conference Proceedings. https://doi.org/10.2218/eor.2025.10953

Bratton, B. (2016). The Stack: On Software and Sovereignty. MIT Press.

Courtney, K., DeLaurenti, K., Kopel, M., & Zimmerman, K. (2024). No One “Owns” That Metadata, Copyright, and the Problems with [Library] Vendor Agreements. DASH. https://nrs.harvard.edu/URN-3:HUL.INSTREPOS:37379715

Davison, S., Wałek, A., & Varachkina, H. (June 16, 2026). Traffic Management and Bot Protection for the OAPEN Library and DOAB: Implementing Cloudflare and Anubis. OAPEN: A World of Scholarly Books: Open to All, Built to Last. https://doi.org/10.58079/16es1

Doctorow, C. (2025, September 27). Pluralistic: The Real (Economic) AI Apocalypse Is Nigh (27 Sep 2025). Pluralistic: Daily Links from Cory Doctorow. https://pluralistic.net/2025/09/27/econopocalypse/

Doctorow, C. (2026, January 18). AI Companies Will Fail. We Can Salvage Something From the Wreckage. The Guardian. https://www.theguardian.com/us-news/ng-interactive/2026/jan/18/tech-ai-bubble-burst-reverse-centaur

Down, A., & Milmo, D. (2026, February 5). What Does the Disappearance of a $100bn Deal Mean for the AI Economy? The Guardian. https://www.theguardian.com/technology/2026/feb/05/disapperance-100bn-deal-ai-circular-economy-funding-nvidia-openai

Frenzel, F. (2025). Smart Enough to Mislead: The Functional Shortcomings and Ethical Dilemmas of Generative AI Use in Metadata Work. Catalogue and Index, 211, 39–57.

Hao, Karen. 2026. Empire of AI: Dreams and Nightmares in Sam Altman’s OpenAI. Penguin Books.

Hicks, M. T., Humphries, J., & Joe Slater, J. (2024). ChatGPT Is Bullshit. Ethics and Information Technology, 26(2), 38. https://doi.org/10.1007/s10676-024-09775-5

ICOLC (International Coalition of Library Consortia). (2023). Joint Statement on the Metadata Rights of Libraries. ICOLC. https://icolc.net/statements/joint-statement-metadata-rights-libraries

Landymore, F. (2026, July 25). AI Companies Are Buying Antique Books, Ingesting Their Contents to Train Models, and Then Destroying Them at Incredible Scale, Even If Almost No Copies Remain. Futurism. https://futurism.com/artificial-intelligence/ai-companies-destroying-rare-books

Ostrom, E. (2009). Building Trust to Solve Commons Dilemmas: Taking Small Steps to Test an Evolving Theory of Collective Action. In S. A. Levin (Ed.), Games, Groups, and the Global Good (pp. 207–28). Springer. https://doi.org/10.1007/978-3-540-85436-4_13

Ouellette, J. (2026, June 23). How to Burst the AI Bubble: Strike at Its Roots. Sci-Fi Author/Tech Journalist Cory Doctorow on His New Book, The Reverse Centaur’s Guide to Life After AI. ars technica. https://arstechnica.com/gadgets/2026/06/how-to-burst-the-ai-bubble-strike-at-its-roots/

Poell, T., Nieborg, D., & Van Dijck, J. (2019). Platformisation. Internet Policy Review, 8(4). https://doi.org/10.14763/2019.4.1425

Poley, C., Uhlmann, S., Busse, F., Jacobs, J.-H., Kähler, M., Nagelschmidt, M., & Schumacher, M. (2025). Automatic Subject Cataloguing at the German National Library. LIBER Quarterly: The Journal of the Association of European Research Libraries, 35(1), 1–29. https://doi.org/10.53377/lq.19422

Program for Cooperative Cataloging. (2025). PCC Task Group on AI and Machine Learning in Cataloging and Metadata: Final Report. https://www.loc.gov/aba/pcc/taskgroup/AI-and-Machine-Learning-TG-final-report.pdf

Ramalho, A, & Steiner, T. (2026, January 20). What the FIL Guadalajara Debates Reveal about Metadata for Academic Books and How Thoth Open Metadata and SciELO Books Respond to This Challenge. Thoth Open Metadata. https://doi.org/10.70950/sfsv8305

Reisner, A. (2025, March 20). The Unbelievable Scale of AI’s Pirated-Books Problem. Technology. The Atlantic. https://www.theatlantic.com/technology/archive/2025/03/libgen-meta-openai/682093/

Rogers, J. (2026, July 17). How AI Threatens the Social Contract of Free Knowledge. Diff. https://diff.wikimedia.org/2026/07/17/how-ai-threatens-the-social-contract-of-free-knowledge/

Saccucci, C., & Potter, A. (2025). Results of AI Experimentation for Cataloging at the Library of Congress. 89th IFLA World Library and Information Congress (WLIC), Satellite Meeting: Artificial Intelligence, Bibliographic Control and Legal Matters: Navigating New Horizons, Astana, Kazakhstan, August 14-15, 2025. https://repository.ifla.org/items/4e25dfe1-336e-49a5-be2a-912e337e7624/full

Steiner, T., Barnes, L., Fathallah, J., Findanis, J., Gatti, R., Deville, J., Sanders, K., Higman, R., Stern, N., & Stone, G. (2025, April 16). Seeding for a Not-for-Profit Community-Led OA Books Ecosystem. COPIM. https://doi.org/10.21428/785a6451.d35141ca

Steiner, T, Arias, J., Bennett, M., Booth, E., Edmunds, J., Gatti, R., Higman, R., Hillen, H., Laakso, M., Nason, M., O’Connell, B., Pogačnik, A., Rabar, U., Ramalho, A., Stone, G., Van Gerven Oei, V. W. J., & Wake Hye, Z. (2026, January 27). International Metadata Recommendations, and Platform-Specific Requirements for Open Access Books and Chapters. Thoth Open Metadata. https://doi.org/10.5281/zenodo.18173982

Steiner, T., Arias, J., Gatti, R., Higman, R., Hillen, H., Ramalho, A., & Van Gerven Oei, V. W. J. (2026, May 5). Building the Conditions for Open Access Books to Thrive: A Reflection on Thoth Open Metadata’s Work Across Infrastructures, Research, and Communities between 2023-26. Thoth Open Metadata. https://doi.org/10.70950/jlke1310

Strauss, I., Yang, J., O’Reilly, T., Rosenblat, S., & Moure, I. (2025). The Attribution Crisis in LLM Search Results: Estimating Ecosystem Exploitation. SSRC AI Disclosures Project Working Paper Series (SSRC AI WP 2025-06), Social Science Research Council. https://doi.org/10.35650/AIDP.4114.d.2025

Sutton, J. (1991). Sunk Costs and Market Structure: Price Competition, Advertising, and the Evolution of Concentration. The MIT Press.

Taniguchi, S. (2024). Creating and Evaluating MARC 21 Bibliographic Records Using ChatGPT. Cataloging & Classification Quarterly, 62(5), 527–46. https://doi.org/10.1080/01639374.2024.2394513

UNESCO. “UNESCO Survey: Two-Thirds of Higher Education Institutions Have or Are.” n.d. Accessed July 26, 2026. https://www.unesco.org/en/articles/unesco-survey-two-thirds-higher-education-institutions-have-or-are-developing-guidance-ai-use

Van Gerven Oei, V. W. J. (2025, April 3). An Update on Punctum Books Usage Data: Interoperability, Interpretability, and AI. punctum books. https://doi.org/10.21428/ae6a44a6.52f495bc

Van Gerven Oei, Vincent W. J. (2026, January 9). Let Our Brains Rot but Let Them Be Ours. Ibraaz. https://ibraaz.org/ibraaz-publishing/read/let-our-brains-rot-but-let-them-be-ours

Van Gerven Oei, V. W. J., & Fradenburg Joy, E. A. (2025, September 10). Open or Proprietary? AI Scraping of OA Content Warrants a Collective Response. punctum books. https://doi.org/10.21428/ae6a44a6.6b5058ea

Weedon, J., François, C., & Ponak, J. (2026). AI Slop and the Information Ecosystem. Institute of Global Politics. https://igp.sipa.columbia.edu/sites/igp/files/2026-06/AI%20Slop%20and%20the%20Information%20Ecosystem_IGP%20Report.pdf

 

CC BY 4.0 Not Plausible, But Verifiable: Scholarly Metadata in the Age of AI by Toby Steiner and Vincent W.J. van Gerven Oei is licensed under a Creative Commons Attribution 4.0 International License.