Skip to content
JCDL 2004
JCDL.2004
Digital Libraries Summit
The Persistent Identifiers Holding Digital Scholarship Together
← All posts

The Persistent Identifiers Holding Digital Scholarship Together

The modern scholarly record has a problem it rarely advertises: the web it is built on will not stay still. Pages move, servers are decommissioned, journals change platforms, institutions reorganise, and the links that connected one piece of knowledge to another quietly break. A citation that pointed somewhere real in 2015 leads, a decade later, to a "not found" page — and with it, a thread in the fabric of human knowledge goes slack. Against this constant erosion stands an unglamorous piece of infrastructure that most readers never notice and could not name: the persistent identifier. It is one of the quiet foundations on which the reliability of digital scholarship depends.

The problem persistent identifiers solve

To understand why persistent identifiers matter, you have to appreciate how fragile a plain web address is. A URL is really just the current location of a thing, and locations change. When a publisher migrates to a new system or a repository restructures its folders, every URL pointing to the old location dies, even though the underlying article or dataset still exists somewhere. This is the reference rot that steadily degrades the scholarly record, a close cousin of the storage-level decay we examined in bit rot, the quiet catastrophe threatening everything we've digitised. The information has not been destroyed; the pointer to it has simply gone stale.

A persistent identifier is designed to break this dependence on location. Instead of citing where something currently lives, you cite a stable, permanent identifier for the thing itself. That identifier is registered in a system that keeps track of the object's current location, so that when the object moves, the identifier is updated to point to its new home. The citation never changes, but it always resolves to the right place. In effect, persistent identifiers add a layer of indirection between a name and a location — and that layer is what keeps scholarship findable as the web shifts beneath it.

How they actually work

The mechanism is elegant precisely because it is simple. A persistent identifier is a unique, permanent string assigned to an object — an article, a dataset, a person, a book. That string is entered into a resolution system: a service whose job is to translate the identifier into the object's current location and send the user there. When you follow a persistent identifier, you are not going directly to a file; you are asking the resolution system, "where does this identifier point right now?" and being forwarded accordingly.

The crucial responsibility in this arrangement is maintenance. A persistent identifier is only as good as the commitment to keep its resolution up to date. If an object moves and no one updates the record, the identifier breaks like any other link — so the system depends not only on technology but on institutions agreeing to steward these records over time. This is why persistent identifiers are as much a social and organisational commitment as a technical one. The string is trivial; the promise to keep it resolving, for decades, is the hard part.

The identifiers you have already met

Several persistent identifier systems quietly underpin modern research, whether or not readers recognise them. The most familiar is the DOI, the Digital Object Identifier, which is assigned to journal articles, datasets and other research outputs. When you see a DOI on a paper, you are looking at a permanent handle that will continue to resolve to the article even if it changes journals or platforms, which is why scholarly citation has increasingly standardised around it rather than around raw URLs.

Just as important, and often overlooked, is the identifier for people rather than documents. Researchers share names, change institutions, and publish under different forms of their name over a career, which makes attributing work correctly surprisingly difficult. Persistent identifiers for researchers, such as ORCID, solve this by giving each person a unique, stable identifier that ties their body of work together unambiguously, regardless of name changes or institutional moves. Alongside these sit other systems — handles, and the long-established identifiers for books and journals — all doing the same fundamental job: attaching a stable name to something that needs to remain findable over time.

Why this matters for the scholarly record

The stakes here are larger than convenience. The integrity of scholarship depends on the ability to cite, locate and build upon previous work, and that ability collapses if the links between works cannot be trusted to endure. A body of knowledge is not just a collection of documents; it is a web of references, each pointing to the evidence and the arguments it rests upon. When those pointers rot, the structure weakens — claims lose their sources, data becomes uncitable, and the record of what was known, and on what basis, grows patchy and unreliable.

Persistent identifiers are the infrastructure that holds this web together against the entropy of the moving internet. They are what allow a citation written today to remain usable in twenty years, a dataset to be found long after the project that produced it has ended, and a researcher's contributions to be reliably connected to their name across a lifetime. This reliable findability is also the foundation on which newer discovery tools are built; the value of semantic search or machine-assisted retrieval, discussed in what semantic search actually does to a digital library, depends entirely on the objects it surfaces still being locatable.

The quiet work of keeping knowledge findable

There is something fitting about the invisibility of persistent identifiers. They do their most important work precisely when nothing appears to happen — when a link that should have broken instead resolves cleanly, and a reader reaches the source they were promised without ever knowing how close it came to being lost. This is the ordinary miracle of good infrastructure: it is noticed only in its absence.

For the communities that steward the scholarly record — libraries, archives, publishers and the researchers who depend on them — persistent identifiers are not a technical footnote but a core commitment. Building and maintaining these systems is the unglamorous, essential work of ensuring that the knowledge we produce remains connected and findable, rather than dissolving into a sea of broken links. The web will keep moving; locations will keep changing. Persistent identifiers are how the scholarly record intends to survive the motion — a quiet promise, kept one resolution at a time, that what we know today can still be found tomorrow.

Keep reading

More from Web Innovations