The Smithsonian’s (In)Visible Women Project, based at the National Museum of the American Indian, is developing pan-institutional data practices within Wikidata to support the disambiguation of artists, designers, scientists, etc. documented across fourteen Smithsonian repositories. This work seeks to connect siloed internal collections information systems (ArchivesSpace, TMS, EMu, etc.), an issue that has compounded with gaps in existing authority identifier sources (LCNAF, ULAN, etc.) and historical Smithsonian information practices to obscure the achievements of women.
Our lightning talk will focus on our work to date and reflect on our successes, but also lessons learned from managing a collaborative, pan-institutional, linked data project. Specifically, we will discuss: 1) the process of developing our collaborative, Smithsonian-specific Wikidata data model; 2) the challenges we have encountered in existing Wikidata structures; 3) our approach to managing the privacy of those represented in Smithsonian datasets; and 4) the uptake of Wikidata editing by Smithsonian staff and our strategies to build Wikidata editing skillsets.
The transition toward linked data environments, entity-based models, and interoperable metadata ecosystems is reshaping cataloguing into a dynamic, collaborative, and distributed practice based on shared cataloguing. To respond to this change, the Share Family community has developed JCricket, the Linked Data Entity Editor designed for collaboratively curating entities in a multi-institutional, shared space. Drawing on over ten years of developments within the Share Family Ecosystem, this presentation will show the concrete implementation of JCricket and how it streamlines libraries’ operational workflows and real-world cataloguing practices. Through granular cataloguing, Provenance-based data management, and integrations with FOLIO and Alma library services platforms (and others to come), data entered in libraries’ ILS/LSP is reflected in the central Cluster Knowledge Base underlying JCricket and vice-versa. These new copy cataloguing features allow for a full roundtrip of data transitioning from record transfer workflows to collaborative, entity-based data management. In JCricket, bibliographic entities are no longer isolated local records or static imported copies, but interconnected data objects continuously enriched, synchronised, corrected, and reused across institutions. Since JCricket operates on a multi-Provenance dataset, its inherent complexity requires guidelines: a set of best practices is being elaborated to consolidate shared cataloguing workflows and ensure consistent data integration with external systems.
Part of the task force for the second revision cycle of the Art and Rare Materials (ARM) BIBFRAME ontology extension, the Manuscripts Subteam has been working to define the role of unique items, such as manuscripts, within the conceptual framework BIBFRAME is built from. Describing manuscripts poses challenges as it requires sets of definitions that can be adapted to objects that, by definition, do not always fit in standardized categories. These unique works and the relationships they may have to other works must be described in such a way that their metadata remains unique but also serves discovery to other resources. Our goal is to extend the ontology to create a system that is flexible enough to be collapsed in on itself when the boundaries between work-instance-item are less distinct but also supports these distinctions where valuable. This presentation will explore the meaning and ontological implications of "uniqueness," the overlapping definitions of manuscripts, and how some of the distinctive characteristics of manuscripts can be expressed within the BIBFRAME ontology.
Libraries around the world are at various stages in their transition to BIBFRAME. Some are learning the impact and benefits of leveraging BIBFRAME data; many are experimenting with BIBFRAME data in their cataloging and discovery workflows; and others are ready to fully productionize BIBFRAME data across their metadata management ecosystems.
When it comes to widespread adoption and implementation, the need for scalable, consistent, and quality data becomes the cornerstone. And while the data formats and workflows evolve, the global infrastructure to support such an important transition already exists. For nearly 60 years, WorldCat has provided the foundation upon which the global library community creates, manages, and shares knowledge—evolving every step of the way as library collections and needs of the communities they serve change.
Building on OCLC’s history of collaboration with the library community on metadata management and linked data innovation, we’re excited to share the next phase of WorldCat’s evolution. In this presentation, we’ll share updates on WorldCat BIBFRAME interoperability, workflows for managing data in both MARC and BIBFRAME, and feedback from our latest library partnership on cataloging in BIBFRAME.
The Blue Core project successfully completed its prototype development in 2025, establishing a community-operated and owned BIBFRAME data store. This prototype has served as a critical testing environment for integration with two open-source linked data editors: Sinopia and Marva. Following comprehensive scalability testing, metadata operations validation, and integration assessments with local Library Services Platforms, the project has advanced to its second MVP (Minimum Viable Product) work cycle. The current MVP incorporates agentic AI capabilities designed to enhance cataloger workflows. The project team is now preparing for a phased BIBFRAME implementation scheduled to begin in early 2027. This session will provide an overview of the project's background, current development status, and planned future activities.
The Program for Cooperative Cataloging (PCC) aims to reduce costs associated with producing bibliographic and authority records to support discovery within public catalogs. The PCC meets this aim by developing metadata application profiles and implementing an apprenticeship-style training scheme within its global membership. The PCC launched a new program in spring 2025, the Entity Management Cooperative (EMCO). The EMCO program aims to empower library staff to manage linked data entities from any source for use in descriptive metadata practices, thereby significantly lowering barriers for participating in the PCC, while fostering the development of the technical expertise needed to forge new paths for linked data adoption within the PCC as a whole. Working in communities of practice, participants have learned to create and describe linked data entities in ISNI, WorldCat Entities, and Wikidata. This panel brings together EMCO participants and organizers to explain why and how the EMCO program was created, explore its work and impact, and discuss what we have learned in our attempts during EMCO’s inaugural year to corral over 250 data enthusiasts into an organized program.
Metadata and Access Librarian, University of Pittsburgh
Hello! This will be my first conference and I'm extremely excited that I have the opportunity to attend. I'm interested in learning as much about Linked Data and BIBFRAME as possible in transitioning from MARC. Electronic resources are also another area of interest. And, I'm in the... Read More →
Metadata Librarian, University of Toronto Libraries
Kyla Jemison is a metadata librarian at the University of Toronto working with special formats. She is interested in music and audiovisual cataloguing, and how metadata affects discovery.
As part of two larger projects in 2025 I identified four articles about bibliometrics published in journals identified as sources of racial hereditarian research. These four articles are relatively insignificant in the larger bibliometrics literature, but they form a natural experiment for both testing the scope of what is available in the bibliometrics data sources and the impact of racial hereditarian research on bibliometrics as a field.
An outlier paper in the study is the most highly cited paper for all of the authors in the paper. While the total number of citations is small, the impact on these particular authors is large. This unique linked data set identifies a new problem with the naive use of citation counts as a basis of quality in scientific communication.
Finding where to contribute on Wikidata isn't always straightforward, especially with structured and lexicographical data spread across multiple WikiProjects and languages. Identifying a gap, and knowing how to act on it, can be harder than making the edit itself.
This session introduces Dustpan and Broomstick, two tools developed by Wikicollabs that help contributors discover opportunities to improve Wikidata. Dustpan helps contributors find Items with missing information by selecting a WikiProject, a category of missing information, and, where available, an additional scope. For example, contributors working with WikiProject Sum of All Paintings can use Dustpan to find paintings that are missing information, such as the creator or material used. Broomstick focuses on lexicographical data, helping contributors uncover Lexemes that can be improved across different languages, including issues involving Senses, Forms, and misplaced statements.
Through live demonstrations, participants will see how finding a gap leads to a specific edit they can make on Wikidata. By the end of the session, participants will be able to use Dustpan and Broomstick to find where to contribute on Wikidata. We also welcome suggestions for which WikiProjects and languages the tools should support next.
This session will provide an overview of Wiki Loves Names and its contribution to Wiktionary. It will highlight the importance of documenting personal names, their meanings, origins, pronunciations, linguistic forms, and cultural contexts, particularly names from underrepresented languages and communities.
The session will explore how Wiki Loves Names supports the preservation and visibility of naming traditions by encouraging contributors to add well-sourced information about names to Wiktionary. It will also highlight the value of Wiktionary as a free, multilingual dictionary and how name-related contributions can help improve knowledge about languages and cultures that are often underrepresented online.
The session will share the objectives, approach, progress, and potential impact of the initiative, while showcasing examples of contributions and opportunities for community participation. It will be an introductory overview rather than a hands-on training session.
Designed for digital humanities research, Semantic Kompakkt is a free, open-source platform for exploring and semantically annotating visual media in an interconnected open-data environment. Semantic Kompakkt, provides accessible features through graphical interfaces while Wikibase manages metadata storage as linked open data, thus enabling the publication and annotation of 2D, 3D, and audiovisual media.
The proposed talk shares lessons from an evaluation of the metadata presented on Semantic Kompakkt field names as reviewed by digital archaeologists. The review was done as part of a project to reduce fuzziness within the platform. Fuzziness, also referred to as wobbliness, serves as an overarching concept encompassing indistinguishable combinations of uncertainty and vagueness within statements or datasets, as well as broader informational ambiguities such as unreproducible literature claims or imprecise ontological mappings.
As a result some key field names were updated, which include; the “persons and institutions” which became “actors”; and descriptions on object types, provenance, material, dimensions, collection and other such, which were re-directed to the original object data from the existing collection. This resulted in having to re-upload the items on Wikibase with all the new statements so as to better serve the objects on display in Semantic Kompakkt.
Agentic AI coding tools such as Claude Code, OpenAI Codex, and the open source OpenCode enable users to create sophisticated software applications from natural-language feature descriptions. This approach, often called “vibe coding,” allows people without software development experience to build applications. It also changes how experienced developers can iteratively orchestrate AI agents to engineer software within highly compressed time frames.
This presentation describes an experiment conducted over a single weekend: developing a static-site RDF editor that uses a graph-based editing approach rather than the form-based interfaces used by tools such as Sinopia, Marva, and Share-VDE’s JCricket. The editor is available at https://ld4p.github.io/graph-editor/, and its source code is hosted at https://github.com/LD4P/graph-editor/.
Software Engineer, Stanford University Libraries, Software Developer
Jeremy Nelson is a software engineer at Stanford University Libraries. He is on the team that migrated Stanford to FOLIO and is active developer in the Sinopia Linked Data Editor and Blue Core projects. His research interests include applying AI to improving library workflows.
Tuesday September 29, 2026 3:20pm - 3:50pm EDT TBA
In this presentation, the PCC Sinopia Cataloging Affinity Group will share updates to the Sinopia editor made for Blue Core, including updates to the interface, functionality, and seed data. We will provide an overview of Sinopia, compare Sinopia Classic to Sinopia for Blue Core, and highlight plans for production cataloging in BIBFRAME. We will review recent developments to operational metadata such as data associated with the BIBFRAME AdminMetadata class, and outline ways to get started with BIBFRAME cataloging in Sinopia.
The National Library Board’s (NLB) Entity Data Service (EDS) is a linked data initiative built on NLB’s knowledge graph, comprising some 2.6 million transformed records from the NLB and National Archives of Singapore (NAS) collections, with entities representing people, organisations and places described across these collections.
This presentation will introduce the origins of the EDS and how entities are managed within the knowledge graph. Using examples of Person entities, it will demonstrate how entity resolution brings together different descriptions and references to the same person across disparate collections, enabling related resources to be connected.
The presentation will conclude with lessons from developing the service and considerations for its future growth, including the importance of authority data in supporting entity reconciliation and enabling entity data to be shared with others.
The presentation will explore how diverse initiatives can complement one another to strengthen the Malagasy language in digital spaces. Key strategies include contributing linguistic data to Unicode and CLDR to enhance technical representation, participating in initiatives supported by organizations such as AFLC and SILICON at Stanford University, expanding open knowledge through Wikipedia and Wikidata, and improving language access through offline community initiatives such as Ainteny.
Rather than presenting language preservation as a single project, the session will show it as an evolving ecosystem in which technical standards, linked open data, open knowledge, media, technology, and community participation all play a role. The Malagasy experience illustrates how contributions at different levels from language data and digital infrastructure to structured knowledge, content creation, and cultural expression can work together to ensure that an underrepresented language remains visible, usable, and sustainable in the digital age.
As linked data increasingly moves from experimentation toward implementation and production, GLAMs face a practical dilemma: how do we build scalable workflows while also ensuring that this data represents the collections and communities we seek to describe? Archival collections especially make this challenge visible. Their records often contain historically underrepresented creators that are inconsistently described or absent altogether from catalogs and authority systems. Linked data is aptly poised to be a lever for richer and more inclusive information access.
This presentation shares a developing linked data production workflow at the University of New Mexico using OCLC Meridian to create, connect, and enrich entities from archival collections. Beginning with MARC records and archival finding aids, the workflow identifies candidates for entity work, gathers disparate pieces of evidence, connects existing authority data, and models relationships. Experiments with OpenRefine, APIs, and Python explore how entity identification and reconciliation can scale while preserving human review.
This case study will share lessons learned about where automation helps, where it breaks down, and how gaps in existing descriptive infrastructure shape who, what, and where can become visible through linked data. The project argues that sustainable linked data workflows should consider not only technical scalability, but representation: using metadata expertise and local knowledge to build richer connections and ensure that underrepresented creators are present in emerging knowledge systems.
Part of BIBFRAME’s promise is that bibliographic descriptions move between systems on the web. The Library of Congress (LC) and the Blue Core project have been using the Concise Bounded Description (CBD) as their unit of description for BIBFRAME data. However, there is no shared, machine checkable definition of what a valid one looks like. In this talk we will describe what LC’s CBD is, and how it came to be. After a brief overview of how the CBD is used at the Library of Congress and in the Blue Core project we will describe some of the challenges involved with moving CBDs around between systems. We will close with a discussion of the significance of ongoing work to validate BIBFRAME using DC Tabular Application Profiles, SHACL and JSON Schema.
The records of a single archival collection are inherently linked via a complex web of interrelationships between topics, people, places, and historical events. Current archival practices and metadata standards, however, do not easily facilitate the discovery of these relationships. In January 2025, Archivists at Dana Library on the campus of Rutgers University - Newark launched a pilot project utilizing Wikidata to make the complex data-sets about a jazz collection and a collection about an historic Rutgers-Newark student protest available to researchers. This presentation will explore how the archivists repurposed existing metadata to create new records in Wikidata, as well as provide an overview of the challenges, surprises, and successes.
The Art and Rare Materials BIBFRAME Extension (ARM), first published in 2018 as part of the Andrew W. Mellon Foundation-funded Linked Data for Production (LD4P) project (2016-2018), is currently undergoing its second revision cycle. It has grown in scope to cover four distinct and interrelated domains; in its current form it seeks to extend the BIBFRAME ontology with additional classes and properties to describe Art, Rare Books, Archives, and Manuscripts. The current revision task force, under the joint auspices of the Society of American Archivists (SAA) and the Rare Books and Manuscripts Section (RBMS), is now in the second year of its two-year work cycle. This presentation will provide an update on the work to date, with brief reports from each of the four subject domains as well as current considerations for the overall direction of the extension, technical upgrades, and integrated approaches to ontology development.
This presentation shares the experience of Cartonera Digital do Cariri, a university extension project at the Federal University of Cariri (UFCA), Brazil, that combines cartonera publishing, cultural memory, sustainability, accessibility, open access, and collaborative knowledge production. Inspired by the Latin American cartonera movement and the cultural traditions of the Cariri region, the project brings students, community participants, and external collaborators together to create physical and digital cultural resources using reused materials, digital publishing, podcasts, QR Codes, audio, and Brazilian Sign Language (Libras).
The presentation focuses particularly on the development of an Interactive Digital Library and its potential to evolve from a collection of community-generated publications into an interconnected cultural knowledge environment. It introduces a gradual Linked Data roadmap involving metadata development, identification of people, works, places, cultural traditions, institutions, and digital manifestations, followed by the use of shared vocabularies, persistent identifiers, machine-readable data, and connections to broader GLAM knowledge ecosystems.
Because the project is still under development, the presentation discusses partial results, challenges, and lessons learned rather than presenting a completed Linked Data implementation. The session highlights how small community-based projects can begin building the foundations for interoperable cultural heritage data while preserving local context, authorship, accessibility, and community participation. Full article : https://docs.google.com/document/d/1UHH1g-DkVqXNF2QH4YSEm4SJWaDVZEJkY4CRkEsrEwo/edit?usp=sharing