The Persistent Identifiers for Projects Community Dialogue, hosted by DataCite and Metadata Game Changers, brought together a diverse group of experts to explore how PIDs can transform the identification and documentation of projects and related resources of many kinds. With over 30 attendees from life sciences, astronomy, and other domains, the discussion highlighted the shared need for standard approaches to connecting diverse resources and for clarity on roles and responsibilities for building and maintaining capacity. This blog summarizes the discussions and focuses on the key “How Might We” themes from the presentations and breakout groups to help focus future efforts. These themes point to practical steps the research community can take to address the challenges of creating and managing connected project metadata.
Landscape of Project Metadata: Presentation Summaries
The dialog opened with presentations from experts across several fields, setting the stage for the following discussions. Each presentation brought unique insights into managing metadata for projects in fields ranging from environmental sampling to citizen science to Indigenous knowledge, illustrating the need for flexible, interoperable metadata standards to enhance project discoverability, impact tracking, and user control over data:
1. Project Framing Use Case – Erin Robinson (Metadata Game Changers) presented a case study focused on marine field stations, such as the Gump South Pacific Research Station, and their role in the Moorea Biocode Project. She discussed the challenges in adequately connecting field stations to specimen collection events and related works. By leveraging infrastructure like DataCite, the project successfully created persistent links between field locations and specimens over time. This case illustrated the importance of detailed, place-based metadata in tracking project impacts and making connections between specimens, plans, research results, and field stations more discoverable.
2. Current Project Metadata in DataCite – Ted Habermann (Metadata Game Changers) discussed how projects are currently represented within DataCite, noting that mandatory fields are consistently used across repositories. In contrast, recommended and optional fields that capture project-specific details vary widely in use. He highlighted how project metadata can document connections among samples, events, people, and funding. This presentation underscored the diversity in metadata application across 42 DataCite repositories already documenting projects and the potential for metadata fields to capture relationships and enhance project discoverability.
3. RAiD: Extending Project Metadata – Shawn Ross (Australian Research Data Commons, ARDC) introduced RAiD, a new identifier service, explaining its customized metadata schema, which includes “Core,” “Extended,” and “Local” blocks to handle diverse project needs. RAiD incorporates an ISO standard for registration authorities and a standardized landing page. Designed for flexibility, RAiD’s metadata can support multiparty administration and track project history. He noted that RAiD is developing automated change management for DOIs and enhanced crosswalks with DataCite, to enable better interoperability across platforms.
4. CitSci: Science Data and Projects – Greg Newman, CitSci & Colorado State University presented CitSci, a platform designed to support the diverse needs of citizen science projects, which has enabled integrated management of over two million observations by public participants. CitSci allows users to create self-described projects, emphasizing storytelling and ease of use, to encourage broad engagement. Newman highlighted the need for standardized metadata to document projects effectively without overwhelming users. The platform enables connections between projects and various scientific observations, including unique use cases like monitoring perennial grains and tracking mountain goats, and aims to make project impacts traceable.
5. Local Contexts: Connecting Indigenous Rights to Projects – Jane Anderson (NYU) & Stephany RunningHawk Johnson (Local Contexts) described Local Contexts, a platform that enables indigenous communities to manage rights and access to their cultural heritage data. Local Contexts empowers communities to create and control projects, allowing institutions and researchers to collaborate within a structure that respects community governance. Projects on Local Contexts can describe a range of events and resources, bridging tangible and intangible heritage elements. The presenters stressed the importance of metadata that respects community authority and tracks data use outside of the community, supporting cultural continuity.
Recordings of these presentations are available on the DataCite YouTube channel. Slides are also available.
Cross-Cutting Themes and “How Might We” Questions
Following the presentations, breakout groups discussed several themes and formulated “How Might We” questions to guide future efforts in improving instrument metadata. These questions reflect common concerns across domains and provide a framework for moving forward.
1. Project Discovery and Linkage
Participants highlighted the need for metadata that enables clear, consistent linkage between projects, datasets, and related instruments.
- “How might we…”
- Create and find references to projects in academic papers?
- Link datasets and instruments to project metadata?
- Connect non-traditional project outcomes, like policy impacts, back to originating research projects?
2. Metadata Consistency and Interoperability
Interoperability across open scholarly infrastructures like DataCite, Crossref, RAiD, ROR, and ORCID requires a shared vocabulary and standard relation types. With each system operating unique schemas and relation types, cross-platform compatibility is challenging, and inconsistent metadata hinders seamless integration. Participants noted that tools like RAiD are developing crosswalks to improve this consistency.
- “How might we…”
- Use relation types consistently across metadata platforms?
- Share a common vocabulary across platforms to enhance metadata consistency?
- Balance between human-readable and machine-readable identifiers, especially in citizen science or public-facing research?
3. Project Versioning and Evolution
As projects evolve, metadata must be able to track version changes without losing prior records. RAiD’s metadata structure offers a model for managing project versions and supporting multi-party contributions. Participants from platforms like OSF shared insights on the need for time-based updates to track project evolution and cumulative impacts over time.
- “How might we…”
- Manage project (DOI) versions and support time-sensitive updates?
- Track project evolution with PIDs that can accommodate contributions from multiple collaborators?
- Include start and end dates to clarify project phases within metadata?
4. Permissions, Ownership, and Accountability
CitSci’s permissions model, where project creators manage co-managers and participant access, inspired discussion on metadata governance and accountability. Participants emphasized the need for transparent permissions and ownership transfer options, especially for projects with multiple stakeholders. Provenance—tracking who updated metadata and when—was highlighted as essential for multi-institutional projects.
- “How might we…”
- Control who creates and updates project metadata?
- Establish accountability for metadata and PID creation and maintenance?
- Transfer project ownership between individuals or institutions while retaining metadata continuity?
5. Reducing Redundant Metadata Entry
Participants frequently cited the challenge of duplicative data entry across platforms, which increases inconsistencies and inefficiencies. RAiD’s proposal for a “single source of truth” for metadata resonated with attendees, as did the suggestion to integrate local systems with persistent identifiers to streamline updates.
- “How might we…”
- Reduce repetitive metadata entry across systems?
- Establish a central source of truth for project metadata to ensure consistency?
6. Tracking and Measuring Project Impacts
Many projects, particularly conservation and community-driven research, produce impacts beyond traditional academic outputs. Tracking intangible or non-academic outcomes, such as policy changes or community benefits, is essential in capturing the societal impact of research. These are difficult to represent in current metadata structures.
- “How might we…”
- Track societal impacts, such as policy changes, that lack DOIs?
- Manage metadata granularity to connect project-wide and sample-level impacts?
- Link research outputs that serve as inputs for subsequent projects, creating a web of interlinked research?
Next Steps
The insights and questions raised during these dialogues underscore a strong commitment to collaborative action for metadata interoperability and enhanced project documentation. Here are some next steps that emerged from the discussions:
- Develop Guidance on Temporal Updates and Accountability: There is a clear need for guidance on managing temporal updates within metadata systems to maintain accountability and consistent project tracking over time. Addressing this will support robust metadata provenance and help align with the dual view of metadata as both time-stamped assertions and records of project elements throughout the life cycle.
- Continue the Project Metadata Conversation: A diverse set of organizations are focused on project PIDs and metadata standardization. The participants recommended a forum that would further increase collaborative efforts and interoperability that already exist across platforms like RAiD, DataCite, ORCID, and other key stakeholders.
Conclusion
This community dialogue clarified the importance of collective action in developing adaptable, interoperable metadata conventions. Participants highlighted the demand for a metadata framework that supports traditional research outputs and broader societal impacts, especially in fields like citizen science, conservation, and community-led projects. By continuing to engage the research community, we aim to create a metadata infrastructure that makes research more accessible, reproducible, and impactful, setting a new standard for persistent identifiers across the Life Sciences, Astronomy, and other disciplinary communities. Stay involved in these ongoing conversations by contributing to the discussions on the DataCite Suggestions GitHub forum or participating in the PID Forum. For more details about the dialogue and the recording, visit the Persistent Identifiers for Projects Community Dialogue page.