Repository Guide

These tools answer three questions about your repository's metadata: what's in it, what it connects to, and how it got that way — measured live from DataCite, in your browser.

Nothing to install, no account, no API key. Nothing you look at is uploaded or stored: each tool reads the public DataCite API directly and does the work in the page. If you can open a web page, you can use them. This guide walks you through a first session end to end.

You only need one thing to start: your repository's name (e.g. "Metadata Game Changers") or its DataCite client ID (e.g. sjyq.oozvia). Every tool takes the same input, and they hand your selection to each other as you go.

Your first session, step by step

Work through these in order the first time. It takes about twenty minutes and ends with a concrete list of records to fix.

1
Start with Metadata Completeness — what's in your metadata?

Type your repository name or client ID and press Explore. The tool samples your records and scores them against the FAIR use cases.

What you'll see: a FAIR Total and four use-case chips — Text (can people find it?), Identifiers (can people connect to it?), Connections (who and what is connected?), and Contacts (Who can contact if I have questions?). Click any chip to see the concepts behind it, each with a completeness bar.
If you are interested, use cases are also available for Project Metadata, Relations, or Contributors, click on Projects or Extras above the completeness chips to focus on these.

Low numbers here are normal. They tell you which parts of your records might be candidates for improvement.

Try it on an example repository →
2
Drill into a weak concept — what's actually there?

Click any concept name (e.g. Project Funder). You'll get the real values from your own records.

What you'll see: each distinct value with its Occurrences (how many times it appears) and DOIs (how many records it appears in) — plus how many DOIs carry the concept at all. This is where "we do record funders" meets "…in 14 of 4,000 records."

This is one way to spot common values and some inconsistencies: five spellings of one funder, or a concept that scores well but holds incorrect values.

3
Move to Metadata Connectivity — does your metadata connect?

Completeness asks whether a name is present. Connectivity asks whether that name carries the identifier that links it to the rest of the world — an ORCID for a person, a ROR for an organization, funder or publisher.

Use the Share with Metadata Connectivity link in the Completeness Top Bar and both tools measure the same records — otherwise a new random sample is drawn and the numbers may differ.

What you'll see: one bar per question — "How many identifiers do creators have?", "How many affiliations do creators have?", and the same for contributors, funders, publishers, and rights. Green = identified in every occurrence, yellow = some, red = never.
4
Click a bar — get a list of records to fix

Every bar opens a list of the distinct people, organizations, funders, or rights behind it.

What you'll see: for each one — Occurrences, Identified, and To fix. All three are links: they open the exact DOIs, so "we should add ORCIDs" becomes a specific list of records.

Start with the quick wins (shown first, tagged in yellow). A quick win is an entity that already has an identifier on some records but not all — you don't have to look anything up, just copy what you already have.

5
Find the identifiers you're missing — ROR and ORCID Retrievers

In the drill-downs, click Find RORs or Find ORCIDs. Every name that still needs an identifier is handed to the matching retriever automatically.

What you'll see: candidate matches for each name — for people, each ORCID iD with their most recent employment and the journal of their most recent paper to help tell two "J. Smith"s apart; for organizations, ROR matches for each affiliation, funder, or publisher string. The results are color-coded with the best being green. Export the results as CSV/TSV and take them back to your repository.

Organizations and people are routed correctly on their own: an organizational creator goes to ROR, a person goes to ORCID.

6
Check for duplicates — the same entity with several spellings

Scroll to Potential duplicates in Connectivity. It clusters names that look like the same person, organization, or license.

What you'll see: a canonical name with its spellings, each with Occurrences and Identified (both link to the records). A ⚠ shared-id conflict means two spellings carry the same identifier — sometimes they are the same name, but sometimes one needs to be fixed. Sometimes you see two people with the same ORCID or multiple ORCIDs for one person!

This section is advisory: it suggests, you decide.

7
Understand how you got here — the history tools
  • Re-Curation Watch — one DOI at a time: changes since 2019 with focus on recent changes, with before/after values and a timeline showing which parts of the record changed and when.
  • Repository Activity — your whole repository over time: what's being curated, and by whom (you, DataCite, or someone else). Changes are shown as record timelines and specific changes.
  • Repository History — completeness for every year you've been registering DOIs, so you can see how completeness changes. You can share data back to Completeness or Connectivity to see details for any year.

These answer the question a completeness score can't: is this improving, and who is doing the invisible work?

8
Take it with you

Every tool exports. Reports as JSON, HTML, or PDF; the numbers as CSV; charts and timelines as PNG (save or copy straight to the clipboard for a slide). Every view is also a plain URL — bookmark it or paste it to a colleague and they'll see exactly what you saw.

Which tool answers which question?

Your questionTool
Are the elements I care about present?Metadata Completeness
What values exist in a given element?Completeness → click a concept
Do my people and organizations have ORCIDs and RORs?Metadata Connectivity
Which records should I fix first?Connectivity → quick wins / To fix
What ORCID/ROR should this name have?ORCID / ROR Retriever
Is the same person/organization in the metadata under three spellings?Connectivity → Potential duplicates
Who changed this record, and what did they change?Re-Curation Watch
Is my repository being curated, and by whom?Repository Activity
Is our metadata getting more complete over the years?Repository History

Reading the numbers honestly

Completeness is presence, not quality. A 100% score means the element is there in every record — not that it's correct, rich, or useful. A title of "Untitled" still counts. Use completeness to find gaps, then look at the values.
Sampling. For large repositories the tools score a random sample (Max sets the size) rather than every record — a few hundred records gives a representative picture in seconds. Two random runs give slightly different numbers; that's the sample, not your metadata. Turn off Random sample or raise Max for a deeper look, and use Share with… so two tools measure the identical records.
Connectivity percentages are occurrence-based. The bar's % is identified occurrences ÷ all occurrences. The bar's colored segments count distinct entities — one person appearing in 40 records is one entity. Both views matter: one tells you how much of the metadata is connected, the other how many people you'd have to look up.
Repository History is a cohort view. Each year is scored using the records as they are today. A strong early year can mean the metadata was good then — or that it was re-curated since. Pair it with Re-Curation Watch or Repository Activity to tell those apart.

Focusing on repository subsets for specific re-curation projects

Three controls, on both Completeness and Connectivity:

You can also paste a single DOI into the repository box to look at exactly one record — handy for checking a fix, or for sharing what a well-formed record looks like with colleagues.

Words you'll see

TermWhat it means here
Client IDYour repository's DataCite identifier, like sjyq.oozvia. Every tool accepts it (or your repository's name).
Use caseA group of metadata concepts that together support a purpose — the four FAIR use cases are Text, Identifiers, Connections, Contacts. The Project use cases are Teams, Items, and Relations.
ConceptOne thing a use case needs, e.g. Project Funder or Resource Author Affiliation. Concepts map to DataCite fields and can be also used in other metadata dialects.
OccurrenceOne appearance of a value. A funder named on 30 awards across 14 DOIs is 30 occurrences in 14 DOIs.
EntityOne distinct person, organization, funder, or rights statement — however many records it appears in.
Quick winAn entity that already has an identifier on some records but not others. Copy it across — no lookup needed.
Re-curationMetadata improved after the DOI was first registered — the sign of an actively maintained repository.
ORCID / RORPersistent identifiers for people (ORCID) and organizations (ROR). They're what turn a name into a connection.

Where to go next

This site uses Google Analytics for anonymous usage statistics. Google Privacy Policy ↗