From 29e841cebcb58b5422fce7958e5c4e374e5046b0 Mon Sep 17 00:00:00 2001 From: damienriehl Date: Sun, 26 Jul 2026 13:10:50 -0500 Subject: [PATCH 1/6] docs: add position paper on identifier durability and opaque canonical IRIs Draft 1 for committee discussion, with the two CDCF drafts it responds to included verbatim under standards/drafts/ as citable companions (both CC-BY). Co-Authored-By: Claude Fable 5 Claude-Session: https://claude.ai/code/session_01MoZuH8vKiwL2rsX3a1zSqc --- ...tifier-durability-opaque-canonical-iris.md | 458 +++++++++ .../drafts/cdcf-catholic-uri-scheme-03.md | 952 ++++++++++++++++++ .../drafts/cdcf-identifier-rationale-00.md | 165 +++ 3 files changed, 1575 insertions(+) create mode 100644 research/identifier-durability-opaque-canonical-iris.md create mode 100644 standards/drafts/cdcf-catholic-uri-scheme-03.md create mode 100644 standards/drafts/cdcf-identifier-rationale-00.md diff --git a/research/identifier-durability-opaque-canonical-iris.md b/research/identifier-durability-opaque-canonical-iris.md new file mode 100644 index 0000000..56dbec3 --- /dev/null +++ b/research/identifier-durability-opaque-canonical-iris.md @@ -0,0 +1,458 @@ +# Identifier Durability and Opaque Canonical IRIs + +| | | +| :---------------- | :---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | +| **Document type** | Position paper — public comment | +| **Status** | Draft 1 — prepared for submission during the `draft-cdcf-catholic-uri-scheme-03` (v0.4.0) 60-day comment window | +| **Relationship** | Responds to [draft-cdcf-identifier-rationale-00](../standards/drafts/cdcf-identifier-rationale-00.md) and [draft-cdcf-catholic-uri-scheme-03](../standards/drafts/cdcf-catholic-uri-scheme-03.md); informs the [CDCF Standards program](../standards/overview.md) | +| **License** | CC BY 4.0 | + +--- + +## Table of Contents + +1. [Executive Summary](#executive-summary) +2. [The Shared Premise](#the-shared-premise) +3. [What We Agree On](#what-we-agree-on) +4. [The Evidence: Transparency Produced Divergence, Not Convergence](#the-evidence-transparency-produced-divergence-not-convergence) +5. [The Steelmen — and Where Each Fails](#the-steelmen--and-where-each-fails) +6. [Where Does the Churn Land](#where-does-the-churn-land) +7. [Worked Example: Canonicalizing Levels of Authority](#worked-example-canonicalizing-levels-of-authority) +8. [The Proposal](#the-proposal) +9. [Universal Slugs Across Standards](#universal-slugs-across-standards) +10. [Costs We Accept — and Their Remedies](#costs-we-accept--and-their-remedies) +11. [Recommended Amendments to Draft 0.4.0](#recommended-amendments-to-draft-040) +12. [Bibliography](#bibliography) + +--- + +## Executive Summary + +**Goals — the ground we already share.** The ambition behind CDCF is not a Catholic filing system. It is one common taxonomy and ontology every Catholic can use; then one every +Christian can use; then one the Abrahamic traditions can share wherever they genuinely refer to the same thing; and ultimately a machine-readable articulation of what those +traditions hold to be true, usable by AI systems generally as alignment-with-human-values infrastructure. That ladder is why the identifier question matters at all: a scheme that +cannot survive the second rung will never reach the fourth. + +**Problems — names and structures are exactly where traditions diverge.** The Church comprises 24 churches _sui iuris_, one Latin and 23 Eastern, and draft 0.4.0 already requires +that each "MUST be representable without privileging the Latin Church as the default" (§4.7.1). An identifier that hard-codes a Latin-rite taxonomy segment, or an English slug, or +an Italian one, quietly violates that requirement in every string it mints — and the problem sharpens as the ladder extends. To an Eastern Orthodox adopter, an identifier reading +`institution/circumscription/…` in Latin-derived English is not neutral infrastructure; it reads as _the Latin thing, not ours_. This is not speculative. Transparency has already +produced measurable divergence inside CDCF's and CatholicOS's own repositories: one verse now carries three committee-minted transparent identifiers in eight months, shipped +registry IDs have been renamed in place, and a calendar-date anchor baked into an identifier moved because the Latin and Italian editions of one book disagree about the date. + +**Solutions — one opaque spine, one guaranteed affordance layer.** We propose that every canonical identifier CatholicOS mints — ontology and data registries alike — be an opaque +base62-encoded UUID of the shape `R…`, and that _all_ human readability move into a layer the standard guarantees rather than one the identifier improvises: every existing slug and +registry key preserved as a **permanent resolvable alias**, never deprecated and never reused; **multilingual labels** on every entity, so naming disputes are settled by _adding_ a +label rather than _changing_ an identifier; and a **rendering rule** requiring production surfaces to show a label beside every canonical ID. This is deliberately a _small delta_ +to draft 0.4.0, not a rival document. Draft 0.4.0 already built the machinery — the two-artifact model (§3.6), the `notations` array (§5.9), the never-reassign stability guarantee +(§3.4), the `exact-match`/`close-match` relations (§4.8.2). We keep all of it and invert which artifact is primary. + +--- + +## The Shared Premise + +We begin where the committee itself began. Asked whether CDCF is meant to become the meta-level disambiguation layer for Catholic data — "kinda like what DOI does for published +URLs" — the answer was yes, with one refinement: `cdcf:` identifiers dereference to structured JSON-LD carrying typed relationships, so the model is "more like what Wikidata does +as an entity graph with typed statements."[^1] We accept that self-description completely; it is the strongest available framing of what CDCF is for. Three consequences follow. + +**First, both named models mint opaquely and carry readability in metadata.** A DOI name is, in the DOI Handbook's own words, "an opaque string" or "dumb number" — "nothing at all +can or should be inferred from the number," and "the only secure way of knowing anything about the entity that a particular DOI name identifies is by looking at the metadata that +the Registrant of the DOI name declares."[^2] Wikidata's Q-numbers work the same way: `Q7186` is one identifier labelled _Marie Curie_ in English and French and _Maria +Skłodowska-Curie_ in Polish.[^3] Neither system is illegible in practice; both are illegible in the _string_ and legible in the _payload_. + +**Second, the richer the resolution payload, the less semantic work the identifier string must do.** DOI resolves to a bare target URI and still succeeds. Draft 0.4.0 resolves to +JSON-LD with `notations`, `crossReferences`, `licenses`, `doctrinalHistory`, and typed authority metadata (§5.4, §5.9). CDCF has built a resolution layer far richer than DOI's, +which means the marginal legibility a transparent string buys is far smaller here than in the systems transparency's advocates usually cite. The payload has absorbed the job. + +**Third, a disambiguation layer must be neutral among the names it arbitrates.** The purpose of such a layer is to adjudicate between competing names, spellings, languages, and +structural placements for one referent. A transparent identifier pre-commits to one side of exactly the disputes the layer exists to resolve — silently, in every citation, forever. +It is a strange arbiter that writes its verdict into its own name. + +--- + +## What We Agree On + +We want to be precise about how much of `draft-cdcf-identifier-rationale-00` we accept, because it is more than the disagreement. + +**The four-axes framework is right, and we adopt it.** The rationale doc separates axis A (grammar), axis B (transparency), axis C (structure), and axis D (consumption), and +observes that fixing one does not fix the others (§2). That is correct and clarifying, and this paper argues inside that vocabulary. **We concede axis A entirely** — "you want a +grammar either way" is simply true, and under this proposal the ABNF work is not discarded but moves to the layer where hand-authored strings actually live, governing the notation +and alias vocabulary plus one trivial production for the canonical shape. **We concede axis D entirely** — a reasoner must treat an IRI as a rigid designator, and the doc is right +that opacity-to-reasoners does not entail opaque minting. Our case for opaque minting is independent of D; it rests on durability and neutrality, not reasoner correctness. + +**Draft 0.4.0's two-artifact model is convergence, not conflict.** The most important thing in the 0.4.0 revision is §3.6: the recognition that graph identity and public citation +are two jobs, and that both can be carried on one entity. The `notations` array (§5.9), the scheme URNs, the never-reassign stability guarantee (§3.4), and the +`exact-match`/`close-match` relations that correctly refuse blanket `owl:sameAs` (§4.8.2, §3.6.1) are precisely the infrastructure an opaque-primary architecture needs. **This +proposal reuses all of it.** We ask the committee to build nothing it has not already specified — only to decide which artifact carries the stability guarantee. + +**The slug schemes are the right vocabulary, in the wrong slot.** The registry slugs are careful, well-researched, and genuinely useful; they are exactly the alias vocabulary the +standard needs — and CatholicOS has already built the mechanism we propose to generalize, since CRMEDR ships `data/deprecated_ids.json` alongside `i18n/la.json`, `i18n/it.json`, +and `i18n/en.json`.[^4] Deprecation records plus multilingual labels beside an identifier is not an architecture we are importing; it is a pattern this organization built once +already, and our proposal is that it be applied universally rather than per-repo. **CSC.rdf's modeling instincts are right too, and we endorse them by name:** the Catholic Semantic +Canon ontology attaches the edition by property (`hasEdition`, with `John_1_14` pointing at `NovaVulgata`) rather than baking it into the base text unit's IRI, and models +vernacular renderings as first-class `Translation` artifacts linked by `hasTranslation`/`translationOf`.[^5] Both are what this paper proposes to make universal: volatile and +language-specific facts belong in properties, not identifiers. + +The dispute, then, is narrow. It lives on axes B and C: whether the _primary_ minted string carries meaning, and whether it encodes hierarchy. + +--- + +## The Evidence: Transparency Produced Divergence, Not Convergence + +The rationale doc's strongest empirical claim is that transparency is the field's answer for hand-authored citation strings. Ours is narrower and closer to home: **inside this +committee's own work, transparency has produced divergence rather than convergence — and quickly.** + +### Two entities, six spellings + +| Entity | CSC.rdf fragment IRI | CSC.rdf `identifier` | Third live spelling | +| :-------------------------- | :------------------- | :--------------------------------- | :---------------------------------------------- | +| **John 1:14** | `csc:John_1_14` | `urn:catholic:scripture:john:1:14` | `cdcf:verse/jn/1/14` (per 0.4.0 §4.1.2 grammar) | +| **Trent, Sess. XIII ch. 4** | `csc:Trent_S13_Ch4` | `Trent-Session13-Ch4` | `urn:catholic:council:trent:session13:chapter4` | + +The first two spellings of each pair sit in one file, on adjacent entities, under `purl.org/cdcf/ontology/catholic-semantic-canon#`;[^5] the third Trent form is the +reference-linking example in the Rome working-session materials.[^6] That same file also carries `CCC-1376`, `ST-III-75-4`, and `CIC1983-915` — four shape conventions inside one +`identifier` property. None of this is carelessness; each spelling is locally reasonable. That is the point: transparent identifiers are locally reasonable in incompatible ways, +and no mechanical test detects the divergence. **One opaque canonical ID would have carried all six spellings as notations, and the divergence would have been visible as what it +is: six citation forms for two entities.** + +### The recorded in-repo record + +**CRMEDR** (Roman Martyrology eulogies) has already corrected shipped identifiers. A slug that captured an entry's introductory words rather than its subject was renamed +`mr:0323-itemcoronae-sanctonim-martyrum` → `mr:0323-domitius-et-socii`; a later commit transliterated the Polish `ł` across ten identifiers (`mr:0308-vincentius-kad-ubek` → +`mr:0308-vincentius-kadlubek` and nine siblings), stating the policy: "IDs are drafts pending committee review, so renamed in place with no deprecated-alias." Most instructive is +the moved anchor — `mr:1210-marcus-antonius-durando` became `mr:0610-marcus-antonius-durando`, because the Latin _editio altera_ 2004 places the blessed on June 10 while the +Italian (CEI) edition of the same book places him on December 10.[^4] The identifier hard-codes a fact that two editions of one work disagree about, so it must move whenever the +anchor edition is reconsidered. + +**CLEDR**'s crosswalk records the same phenomenon across projects. On 26 January 2021 the Congregation for Divine Worship decreed that 29 July be designated the Memorial of Saints +Martha, Mary and Lazarus, replacing the celebration of Martha alone.[^7] CLEDR's row carries the Latin title _Sanctorum Marthæ, Mariæ et Lazari_ — and three irreconcilable keys: +litcal froze `StMartha` (now factually wrong), romcal renamed to the 56-character `martha_of_bethany_mary_of_bethany_and_lazarus_of_bethany`, and eprex kept `martha`.[^8] Three +projects, one decree, three divergent responses — and the row's source column still credits `missale_romanum_1970`, so the decree that caused the divergence appears nowhere in the +identifier layer. + +**CECDR** (ecclesiastical circumscriptions) states a strip rule: generic type words such as "Diocese of" and "Diocesi di" are stripped, because "the _type_ is an attribute, not +part of the identity." Against 2,935 live IDs, 58 still contain `arcidiocesi-di-` and 31 military and personal ordinariates carry a type word slugged across at least ten languages +(`ordinariato militare`, `obispado castrense`, `diocese aux armees`, `ordynariat polowy`, `vojensky ordinariat`, and more) — 89 live identifiers at odds with the repository's own +rule. Separately, `circ:it-opus-dei` was renamed `circ:int-opus-dei` because "the personal prelature of the Holy Cross and Opus Dei is supranational, so tying it to Italy … was +wrong," and two homonymous Chinese sees are disambiguated by bare ordinals, `circ:cn-xinjiang-1` and `circ:cn-xinjiang-2`, both noted "qualifier pending committee review."[^9] + +Three sibling registries add three more shapes of the same problem. **CLBDR** contradicts itself between front page and schema: `README.md` documents +`__` with the example `martyrologium_romanum_cei_2004`, while `docs/schema.md` specifies `__` — and the data follows the +schema (`martyrologium_romanum_2004_it_IT`), so the README's example identifier exists nowhere in the registry. **CICLSALDR** records that its own `icl:` prefix "reads as the +repository's shorthand, not as a canonical classification of every entry," because its scope includes societies of apostolic life, canonically distinct from consecrated life; +`icl:cm` (_Congregatio Missionis_, typed `society_of_apostolic_life`) is live proof. **CMDDR** renamed its own acronym — "Fix typo in project name from CLDDR to CMDDR" — before +minting a single identifier.[^10] + +**Draft 0.4.0 itself** supplies two more. It retired the `order` supertype before reaching 1.0, leaving `cdcf:institution/order/{slug}` resolvable only via successor pointers +(§4.7). And §4.6.1 requires an ordinal on every papal identifier, producing `cdcf:person/pope-francis-i` — an ordinal the Church did not use: the Vatican listed the name simply as +_Francis_ in its first bulletin, and the press office's own gloss was that "it will become Francis I after we have a Francis II."[^11] The identifier asserts a fact about the +Church that the Church declined to assert. Meanwhile Open Issue 1 (Psalm numbering) and Open Issue 5 (multilingual slugs) have survived four revisions unresolved (§10) — both +disputes that exist _only because_ the identifier string must choose. + +Every one of these is a governance-quality problem a diligent committee will keep solving. A committee member may fairly object that draft registries are _supposed_ to churn — that +is what a pre-1.0 registry is for. Two answers. First, litcal, romcal, and ePrex are not drafts; they are production systems, and the St Martha divergence happened in their +_shipped_ keys — a 2021 decree left `StMartha` frozen factually wrong in a production API while a sibling project renamed to a 56-character key — so the mechanism operates after +normativity, not only before it. Second, the churn's causes — multilingual naming, movable anchors, editorial judgment about which name is _the_ name — do not end at 1.0: popes +keep being elected, dioceses keep merging, decrees keep expanding memorials. Normativity freezes the identifiers, not the world. That is precisely the cost we ask the committee to +weigh: transparency's maintenance burden is not a one-time migration but a standing obligation, recurring whenever the world, an edition, a decree, or a translation policy changes. + +--- + +## The Steelmen — and Where Each Fails + +We are not neutral, but the transparent case deserves its full strength, because a proposal that defeats only weak versions of the opposition deserves to lose. Each steelman below +is followed by the failure mode we think it develops over time, a concrete mitigation from the affordance layer, and an invitation. If the transparent camp can name a remedy for +one of these failure modes that costs less than ours, that should decide the question. + +### Steelman 1 — The BCP 47 / Unicode layered model + +**At full strength.** The rationale doc's §5 is its best section. BCP 47 composes stable atomic registry codes by ABNF into `zh-Hant-TW`: a tag simultaneously transparent to a +human, grammar-validated, built from stable atoms, and opaque to a matcher. Unicode pairs an opaque code point `U+0041` with a name — `LATIN CAPITAL LETTER A` — frozen forever by +its stability policy. The i18n stack runs on these, in CDCF's exact use case: hand-authored, quoted-in-the-wild identifiers, for two decades. Nobody writes `lang="Q1860"`. + +**Failure mode.** BCP 47's registry survives change only through alias machinery — which is our architecture, not theirs. RFC 5646 §3.1.6 and §3.1.7 define `Deprecated` and +`Preferred-Value` precisely so the registry can carry `iw` forever while canonicalizing it to `he`; the §3.4 stability guarantee is that `Subtag`, `Type`, and `Added` "MUST NOT be +changed" — the subtag is never removed, only redirected.[^12] That is an opaque-spine architecture wearing legible clothes. More decisively, language tags have a property no +Catholic entity has: they are written in the one alphabet definitionally neutral for their domain, and their referents are the naming authorities themselves. There is no +Latin-versus-Italian dispute about how to spell `en`. There is exactly such a dispute about `pope-{name}-{roman}`: Leo, or Leone, or León? CECDR's ten-language ordinariate slugs +are that dispute already lost, at scale. + +**Mitigation and invitation.** Adopt the BCP 47 architecture in full, including the part that does the work. Under this proposal `Preferred-Value` is not approximated; it _is_ the +alias layer, and every competing spelling (`pope-leo-xiv`, `papa-leone-xiv`, `papa-leon-xiv`) becomes a permanent resolvable notation on one canonical ID, none privileged and none +wrong. We would welcome a proposal for how a transparent primary IRI carries three co-equal language forms without electing one. + +### Steelman 2 — Code ergonomics + +**At full strength.** `if (key == "immaculate-conception")` is readable in a diff, greppable in a log, and self-documenting in a stack trace. `if (key == "Rk8f3vQ2…")` is none of +those. Most humans who ever touch these identifiers are application developers, not ontologists, and their productivity is a real cost that ontological purity does not pay. + +**Failure mode.** The legibility is real but bound in the wrong place: in a literal, at every call site. When the referent's boundary or preferred name changes, every call site is +stale in a way no compiler catches. The string was never checked against the record; it was trusted because it looked right. + +**Mitigation and invitation.** Named constants bound to canonical IDs — `const IMMACULATE_CONCEPTION = "R…"` — which is how every codebase already handles hex colors, port numbers, +and country codes. That restores full call-site legibility while leaving exactly one authoritative binding to audit, and the rendering rule (§8, item 4) extends it to registry +source files, serializations, and generated code so a label always sits beside the ID. If the constant-binding overhead is too high for a particular consumer, we would like to see +that workflow, so the affordance layer can be shaped around it. + +### Steelman 3 — Church-oversight verifiability + +**At full strength.** Ecclesiastical review is a genuine requirement, and reviewers are theologians and canonists, not engineers. A bishop's delegate can read +`cdcf:magisterium/pope-leo-xiii/rerum-novarum` and confirm it is right. Nobody can review a page of base62. + +**Failure mode.** Transparency converts review from _verification_ into _recognition_. The reviewer confirms the string looks correct; the string is not thereby checked against the +record. When a slug is subtly wrong — a garbled Latin incipit captured as a subject name, a date drawn from the wrong edition, an ordinal the Church never used — it passes review +precisely because it reads plausibly. Every recorded CRMEDR correction above was a plausible-reading slug that shipped. + +**Mitigation and invitation.** The first half needs no tooling at all: under commitment 4 of the Proposal, the registry source file a canonist reviews _remains a plain text file_, +with the Latin label in the column adjacent to the ID. Nothing is taken away from text-file reviewers — the opaque ID is an added column, not a substitute — so recognition-style +review continues exactly as it does today, while verification-style review becomes possible on top of it. Verification against labels rendered _from_ canonical IDs is strictly +safer, because it checks the record rather than trusting the key. Under the rendering rule a review surface shows the canonical ID with its `skos:prefLabel` in the reviewer's own +language — Latin, Italian, English — resolved live from the registry, so a reviewer sees what the system actually believes rather than what a past minting decision asserted. We +would welcome the committee's review-workflow requirements as design input for that rule, which is the natural place to encode them. + +### Steelman 4 — Things versus concepts + +**At full strength.** The rationale doc's §6 is its most original contribution and genuinely explanatory: OBO Foundry and the Gene Ontology went opaque because biological +categories are reclassified as science advances, while BCP 47 and OSIS went transparent because languages and scriptural books are fixed.[^13] CDCF spans both, so it should apply +both rules per dataset. The residual fringe among things — antipopes, the Stephen II/III ambiguity, the skipped John XX — is finite, enumerable, and already adjudicated by the +Church's own historical record. + +**Failure mode.** The criterion cross-cuts the doc's own evidence. Its §4.1 table lists **VIAF, Getty (TGN/ULAN/AAT), and GeoNames** as opaque — and those are authority files of +_things_: persons, places, named artifacts. It lists **schema.org, FOAF, Dublin Core, and SKOS** as transparent — and those are _concept_ vocabularies of classes and properties. If +things→transparent and concepts→opaque were the rule, those rows sit on the wrong side of it; and the row the doc itself flags as drift-prone is **DBpedia**, transparent +identifiers for things, derived from article titles, which break on rename (§4.1). The line the field actually draws is simpler, and it is the doc's own summary sentence in §4.1: +**opacity clusters where a resolver always mediates.** That describes CDCF exactly — draft 0.4.0 specifies a resolution server, mirror discovery, content negotiation, a change +feed, and one-year immutable caching (§5.1–§5.9). CDCF is not choosing whether to be resolver-mediated; it has already chosen. + +**Mitigation and invitation.** Apply the resolver criterion, which CDCF satisfies, rather than the things/concepts criterion, which the doc's own evidence table does not support — +and keep the things/concepts insight where it is undeniably right: as the rule for which _notation scheme_ to feature. Scripture keeps OSIS, canons keep their numbers, CCC keeps +its paragraph numbers; this proposal never asks anyone to stop writing them, only that they be notations on a durable spine rather than the spine itself. If the committee believes +there is a closed class of entities whose canonical naming is genuinely finished, we would like to see it enumerated — our reading of the six registries is that every candidate +class has already recorded a rename. + +### Steelman 5 — "Just take a vote on the language" + +**At full strength.** Standards bodies decide contested questions by deliberation and vote all the time. Latin is the Church's own language and an obvious Schelling point. +Committees exist to make exactly these calls; declaring the question unanswerable is an abdication. + +**Failure mode.** A vote produces a winner and a resentful minority — per identifier, permanently, and visibly in every citation string. That is a governance tax recurring with +every new entity and compounding with every tradition the standard hopes to serve. It is also empirically what has happened: Open Issue 5 has been open across four revisions, +CRMEDR mints Latin lemmas, CLEDR mints English snake_case, and CECDR's ordinariates ended up in ten languages without anyone ever deciding they should. + +**Mitigation and invitation.** Multivalued labels produce no losers. `skos:prefLabel` is language-tagged; a Polish reader gets Polish and a Latin reader gets Latin, from one +record, with no election held. The precedent is trivially familiar: _honor_ and _honour_ are both correct, and no standard had to choose, because they are labels rather than keys. +If a vote is nonetheless preferred for a given domain, note that under this proposal it decides which notation carries `prefLabel: true` — a reversible, low-stakes call — rather +than which identifier the world cites for a century. + +--- + +## Where Does the Churn Land + +Strip away the vocabulary and one question remains. **Every identifier architecture must absorb world-change somewhere.** Editions get revised, dioceses are erected and merged, a +decree expands a memorial from one saint to three, a prelature turns out to be supranational, a committee decides `order` should have been `institute`. The choice is not whether to +absorb change but which layer takes the hit. + +**Transparent-primary architectures absorb it in the canonical layer.** Every change becomes a deprecation, a successor pointer, a permanently stale citation in a document nobody +will revise. Draft 0.4.0 handles this correctly — §3.4 requires that both original and successor stay resolvable indefinitely — but "handles correctly" means "accumulates +forever," and the discipline it requires is perpetual institutional rigor, which even ISO failed at once: `CS` was used for Czechoslovakia and then reused for Serbia and +Montenegro, and ISO's own archival code for the latter had to be changed from `CSHH` to `CSXX` to stop the collision. ISO's success story is the other half of the same standard — +Burma → Myanmar archived as `BUMM`, withdrawn alpha-2 codes transitionally reserved for at least fifty years before any possible reuse.[^14] The rationale doc reads this correctly +(§7): reassignment is the danger, and it is a governance failure available to both camps. + +**Opaque-primary architectures absorb it in the alias layer, which is built to age.** When a slug turns out to be wrong it is not corrected — it is _joined_. +`mr:1210-marcus-antonius-durando` and `mr:0610-marcus-antonius-durando` both resolve, forever, to one entity; the Latin edition and the CEI edition are each right about their own +book; nothing published ever breaks. `StMartha`, `martha`, and the 56-character romcal key all resolve to one celebration, and the 2021 decree becomes a property on the record +rather than a naming crisis in three projects. Aliases are supposed to pile up. Canonical identifiers are not. + +And the cost asymmetry has inverted since this trade-off was last argued seriously. The rationale doc calls opacity's cost "permanent and unsolvable by design" — a mandatory lookup +for every human who reads the identifier (§7). That was true in 2005. Hover labels, IDE inlays, resolver-backed link previews, and agents that never hand-type an identifier have +collapsed the lookup cost toward zero, and CDCF's own resolution protocol is precisely the substrate those affordances run on. CatholicOS need not take this on faith: ontokit +already renders resolver-backed labels beside opaque `osc:` R-IDs today — its source-view hover resolves a full IRI to its label, and its `useIriLabels` hook label-joins API +responses.[^15] Drift costs move the other way: they rise with every integration, every downstream project, every new edition, and every tradition the standard reaches for. +**Legibility costs are falling toward zero; drift costs rise with adoption.** The transparent trade-off was right for 2005. We do not think it is right for a standard minted in +2026 to last a century. + +--- + +## Worked Example: Canonicalizing Levels of Authority + +The committee has an open question about how to canonicalize levels of authority.[^1] It is an ideal test of the architecture, because three formulations already coexist and they +are not 1:1 mappable. + +| Source | Count | Values | +| :----------------------------- | :---- | :----------------------------------------------------------------------------------------------------------------------------------------------- | +| Rome working-session materials | 6 | Primary Revelation · Universal Ordinary Magisterium · Papal Magisterium · Universal Theological Commentary · Local Magisterium · Private Opinion | +| CSC.rdf individuals | 6 | `PrimaryRevelation` · `UniversalOrdinary` · `PapalMagisterium` · `CommentaryLevel` · `LocalOrdinary` · `NonMagisterial` | +| Draft 0.4.0 §6.2 | 5 | `solemn-definition` · `ordinary-universal` · `definitive-doctrine` · `authentic-doctrine` · `pastoral-guidance` | + +Two things are worth noticing. First, the Rome materials disagree with _themselves_ about ordering: the pyramid slide runs Primary Revelation, Universal Magisterium, Papal +Magisterium, Local Magisterium, Commentary, Private opinion — placing Local Magisterium above Commentary — while the authority-metadata table on a later slide places Universal +Theological Commentary above Local Magisterium.[^6] Second, the two six-value lists and the five-value list do not measure the same thing: the first two answer _who teaches_, and +draft 0.4.0's answers _what assent is owed_ (§6.1). Those axes correlate but do not align; the same document can sit high on one and lower on the other. + +Harmonizing this is real theological work, and not this paper's to do — draft 0.4.0 §7.3 rightly requires theologian and canonist review for exactly these classifications. What we +can say is what the harmonization will do to the identifiers. It will rename levels, reorder them, split at least one, and possibly separate the two axes into two vocabularies. If +level-IDs are transparent strings, every one of those moves breaks every text already tagged: a corpus tagged `"UniversalOrdinary"` is stranded the moment the level is renamed or +its boundary redrawn, and the migration is a rewrite of the annotation layer rather than of a lookup table. If level-IDs are opaque, with today's names carried as labels and all +three existing vocabularies carried as notations, the theology can develop and the data survives — the record's label changes, the tagged corpus does not move, and the crosswalk +between the who-teaches and what-assent axes becomes a property rather than a renaming. That is the whole argument in one case. **The identifier architecture should let the +theology be revised. It should not require the theology to be finished first.** + +--- + +## The Proposal + +Five commitments. Nothing here replaces draft 0.4.0's machinery; each item names the 0.4.0 mechanism it rides on. + +1. **Canonical identifiers are opaque, `R`-shaped, and org-wide.** Every canonical ID CatholicOS mints — ontology and every data registry — is a base62-encoded 128-bit UUID + prefixed `R` (23 characters in practice; 122 random bits, so decentralized minting needs no counter and no central allocator, and the leading letter keeps it QName-safe for + RDF/XML). This is not new for CDCF: it is the shape already shipping in `ontology-semantic-canon`, where `osc:RChKPk9K152BirrIYgAREsY` is _Clergy_, and the shape ontokit already + mints.[^15] It is also FOLIO's shape — `folio.openlegalstandard.org/R7Ttdyo4FsvaupPKT35Qry0` is _Murder_.[^16] +2. **Every existing slug, key, and registry prefix becomes a permanent resolvable alias.** `mr:`, `circ:`, `icl:`, CLEDR keys, CLBDR edition IDs, litcal/romcal/eprex keys, + `cdcf:verse/jn/1/14`, `cdcf:concept/C0000418` — all carried as 0.4.0 `notations` (§5.9) under declared scheme URNs, resolvable per §3.4, **never deprecated and never reused**. + This is the IP/DNS model, and the analogy is worth stating plainly: nobody argues that an IP address should be human-readable, and nobody has to, because the domain name + resolves to it. No adopter loses a working key. That is the adoption story. +3. **Every entity carries multilingual labels.** `skos:prefLabel` and `skos:altLabel`, language-tagged, on every entity in every registry. Naming disputes are resolved by **adding + a label**, never by changing an identifier. CRMEDR's `i18n/{la,it,en}.json` is the existing precedent; this generalizes it. +4. **Production surfaces MUST render a label beside every canonical ID.** Registry source files carry a label column or comment beside each ID; serializations carry the label + inline; UIs and generated code render it. This commitment is what makes opacity livable, and it belongs in the standard rather than in each implementer's good intentions. +5. **Structural and volatile facts live in properties, never in canonical identifiers.** Dates and calendar position (CRMEDR's `MMDD` anchor), country codes (CECDR's ISO 3166 + prefix), taxonomy supertypes (`circumscription`, `institute`, the retired `order`), edition years (CLBDR), chapter/verse hierarchy, ownership, and language are all already + modeled as fields in 0.4.0 responses. Where a fact lives both in a field and in the identifier, the identifier is the copy that goes stale. + +--- + +## Universal Slugs Across Standards + +Draft 0.4.0 §4.9 already states the principle we want to generalize: a sibling registry's slug "MUST be reused verbatim as the final path segment of the corresponding `cdcf:` IRI … +This is a MUST, not a convention: it is what makes the pairing machine-verifiable rather than merely coincidental." That is exactly right, and it is the seed of something larger. +Lift it one level, from sibling registries to peer standards: when CatholicOS and another standard identify the same referent, they reuse the **same opaque local name** under their +own namespaces. + +```text +https://ontology.catholicos.catholic/R7Ttdyo4FsvaupPKT35Qry0 +https://folio.openlegalstandard.org/R7Ttdyo4FsvaupPKT35Qry0 +``` + +Cross-standard identity becomes machine-verifiable by local-name equality — no mapping table, no crosswalk to maintain, no `sameAs` hazard. Each standard keeps its own namespace, +governance, labels, and resolution payload; only the local name is shared. **This is possible only because the shared name asserts nothing in anyone's language.** No tradition will +agree to share `god-the-son` across Jewish, Catholic, and Protestant standards — the string itself is a theological claim, and for two of the three it is the wrong one. But every +tradition can share `R7Ttdyo4FsvaupPKT35Qry0`, because it says nothing at all, and each standard attaches its own label, definition, and typed statements to it. Opacity is not +merely tolerable at the inter-tradition boundary; **it is the only thing that crosses it.** That is the interfaith-alignment benefit named in the goals ladder, and it is concrete +rather than aspirational: shape-compatibility with FOLIO today, and a mechanism that scales to machine-readable identifier work across the Abrahamic traditions tomorrow. The +alternative — a mapping table between every pair of standards, maintained by both parties in perpetuity — is the cost transparency imposes at exactly the boundary where the goals +ladder can least afford it. + +--- + +## Costs We Accept — and Their Remedies + +Opacity has real costs. We would rather name them than let them be discovered later. + +| Cost | The problem, stated plainly | Remedy | +| :----------------------- | :-------------------------------------------------------------------------------------------------------------------------- | :--------------------------------------------------------------------------------------------------------------------------------------- | +| **Diff reviewability** | A slug in a pull request is self-checking; a reviewer sees a wrong one. `Rk8f3vQ2…` is not self-checking and never will be. | Mandatory label columns in registry source files, so every diff line carries an ID **and** a human-readable label that reviews itself. | +| **Silent wrong-paste** | Paste the wrong opaque ID and nothing looks wrong. Paste the wrong slug and something usually does. | Lint rules validating every ID against the registry and asserting label/ID agreement in CI; label comments beside IDs in serializations. | +| **Debugging ergonomics** | A log line, stack trace, or SPARQL result full of base62 is harder to read than one full of slugs. | IDE inlays and resolver-backed hover labels; named constants bound to canonical IDs at call sites; label-joining helpers in tooling. | +| **Onboarding friction** | A newcomer reading raw data cannot orient without a lookup. | The §5.9 `notations` array ships every canonical record with all its familiar slugs, so a newcomer's existing vocabulary still works. | + +None of these remedies is speculative — each exists in shipped systems, and two exist in CatholicOS repositories today. But we do not claim the list is complete. **We invite the +committee, and especially the transparent camp, to name costs we have missed and mitigations we have not thought of.** If a cost turns out to have no adequate remedy, that is a +finding worth having before adoption rather than after. + +--- + +## Recommended Amendments to Draft 0.4.0 + +We offer these as amendments the committee can adopt into the existing text, not as a rival document. Draft 0.4.0's non-identifier machinery — the `licenses` object, authority +metadata, content negotiation, caching, resilience, and the governance process — is endorsed as written. + +| § | Amendment | +| :------------------- | :----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | +| **§3.6** | Make the ontology IRI opaque for **every** domain, not only `concept/`. Each domain's current transparent path form becomes a **guaranteed** notation rather than an optional one. The two-artifact model is unchanged; only which artifact is primary. | +| **§3.3, Appendix D** | Re-scope the ABNF grammars to govern the notation layer, where hand-authored strings live. Add one production for canonical IRIs: `canonical-id = "R" 20*24(ALPHA / DIGIT)` (base62, case-sensitive — this relaxes §3.3's lower-case rule for canonical IDs only). | +| **§4.5** | Extend the opaque option from `concept/` to all domains, and recommend the `R` shape over sequential `C0000418` for new mints. Existing `C…` IDs are kept as permanent notations; **no forced migration** (§1.3 preserved). | +| **§4.9** | Extend the verbatim-slug MUST to **cross-standard** opaque local names: peer standards identifying the same referent SHOULD reuse the same local name under their own namespaces. | +| **§3.4** | State explicitly that the stability guarantee applies to notations as permanent aliases: a published notation MUST remain resolvable and MUST NOT be reassigned, exactly as an IRI. | +| **§5.9** | Make `notations` load-bearing in every domain (withdrawing §3.6's permission to omit it for "thing" domains), and add `skos:prefLabel`/`altLabel` as a REQUIRED language-tagged field on every resolution response. | +| **New §5.10** | Add a production rendering rule: registry source files, serializations, UIs, and generated code MUST render a human-readable label adjacent to every canonical ID. | +| **§10, new issue** | Mint namespace: we recommend **one org-wide namespace** shared by the ontology and all registries. The host choice — `id.catholiccommons.org` (§3.1) versus `ontology.catholicos.catholic` (already live for `osc:`) — is a committee decision we do not presume. | + +Two notes on scope. This proposal executes no data migration; beyond the permanent-alias commitment, migration mechanics are committee work. And it proposes no theological +harmonization — §7.3 review remains the right gate for anything touching doctrinal classification. + +--- + +## Bibliography + +[^1]: + Committee discussion, July 2026 (data-and-standards group), paraphrased without attribution. The exchange characterized CDCF as the meta-level disambiguation layer, "kinda like + what DOI does for published URLs," refined as "more like what Wikidata does as an entity graph with typed statements," and raised the canonicalization of levels of authority. + +[^2]: + International DOI Foundation, _DOI Handbook_, §2.2 (Syntax of a DOI name) and §2.2.1 (General characteristics), https://www.doi.org/doi_handbook/2_Numbering.html. The quoted + phrases "opaque string," "dumb number," and the metadata sentence are from those sections. + +[^3]: + Wikidata, "Help:Items," https://www.wikidata.org/wiki/Help:Items. Item `Q7186` is labelled _Marie Curie_ in English and French and _Maria Skłodowska-Curie_ in Polish; labels, + descriptions, and aliases are per-language while the identifier is not. + +[^4]: + CatholicOS, _CRMEDR_, github.com/CatholicOS/crmedr: commit `afb5a70` ("Fix garbled/spurious registry IDs (rename/remove, not deprecate)"); commit `d6439bd` ("Fix ł-drop in + canonical slugs: transliterate ł→l (10 Polish IDs)"), whose message states the quoted policy; `docs/canonicalization-report.md` (moved date anchor); and the alias/label + precedent in `data/deprecated_ids.json` with `i18n/la.json`, `i18n/it.json`, `i18n/en.json`. + +[^5]: + _Catholic Semantic Canon_ ontology (CSC.rdf, v1.0.0), namespace `http://purl.org/cdcf/ontology/catholic-semantic-canon#`: individuals `John_1_14` and `Trent_S13_Ch4`; + `identifier` values `urn:catholic:scripture:john:1:14`, `Trent-Session13-Ch4`, `CCC-1376`, `ST-III-75-4`, `CIC1983-915`; `hasEdition`/`Edition` and + `hasTranslation`/`translationOf`/`Translation`. The base text unit's IRI is edition-neutral, though the vernacular `Translation` individual (`John_1_14_NABRE`) does concatenate + its edition into its IRI. + +[^6]: + Rome working-session materials, "Why AI for the Church Needs a Hierarchy of Authority" (presentation deck, 2026). The pyramid slide and the authority-metadata table slide order + Commentary and Local Magisterium differently; the reference-linking slide gives `urn:catholic:council:trent:session13:chapter4` as its example target. + +[^7]: + Congregation for Divine Worship and the Discipline of the Sacraments, _Decree on the Celebration of Saints Martha, Mary and Lazarus in the General Roman Calendar_, 26 January + 2021, https://www.vatican.va/roman_curia/congregations/ccdds/documents/rc_con_ccdds_doc_20210126_decreto-santi_en.html. + +[^8]: + CatholicOS, _CLEDR_, github.com/CatholicOS/cledr, `liturgical_events.md` — crosswalk row for _Sanctorum Marthæ, Mariæ et Lazari_, carrying `litcal_key: StMartha`, + `romcal_key: martha_of_bethany_mary_of_bethany_and_lazarus_of_bethany`, `eprex_key: martha`, and source `missale_romanum_1970`. + +[^9]: + CatholicOS, _CECDR_, github.com/CatholicOS/cecdr. Strip rule and Xinjiang note: `docs/schema-proposal.md`. Counts measured against `data/circumscriptions.json` (2,935 IDs; 58 + containing `arcidiocesi-di-`; 31 ordinariates carrying a type word). Opus Dei rename and its quoted rationale: commit `5067fb8`. + +[^10]: + CatholicOS sibling registries, github.com/CatholicOS: _CLBDR_ (`README.md` line 21 against `docs/schema.md` and `data/editions.json`); _CICLSALDR_ (`docs/schema-proposal.md` + prefix note; `data/institutes.json`, `icl:cm` typed `society_of_apostolic_life`); _CMDDR_ (commit `13d3de9`, "Fix typo in project name from CLDDR to CMDDR"). + +[^11]: + Associated Press, "Just Francis," _Philippine Daily Inquirer_, March 2013, https://newsinfo.inquirer.net/373397/just-francis. The Vatican's first bulletin listed the name + without a Roman numeral; the press office's gloss was "It will become Francis I after we have a Francis II." + +[^12]: + A. Phillips and M. Davis, eds., _Tags for Identifying Languages_, BCP 47 / RFC 5646 (IETF, September 2009), §3.1.6 (`Deprecated`), §3.1.7 (`Preferred-Value`, including the + `iw`/`he` example), and §3.4 (stability of `Type`, `Subtag`, `Tag`, and `Added`), https://www.rfc-editor.org/rfc/rfc5646.txt. + +[^13]: + Gene Ontology Consortium, "Ontology Documentation," https://geneontology.org/docs/ontology-documentation/ — "Every term has a GO ID, a unique seven digit identifier prefixed by + GO:"; obsoleted terms keep "the term and ID … in the ontology," tagged obsolete. See also OBO Foundry, Principle 3: URI/Identifier Space, + https://obofoundry.org/principles/fp-003-uris.html. + +[^14]: + ISO 3166-3, _Codes for the representation of names of countries and their subdivisions — Part 3: Code for formerly used names of countries_. Burma → Myanmar is archived as + `BUMM`; `CSHH` was first assigned to Serbia and Montenegro despite its prior use for Czechoslovakia and was subsequently changed to `CSXX`; withdrawn alpha-2 codes are + transitionally reserved for at least fifty years. See https://en.wikipedia.org/wiki/ISO_3166-3 and ISO Online Browsing Platform, https://www.iso.org/obp/ui/#iso:code:3166:CS. + +[^15]: + CatholicOS, _ontology-semantic-canon_, github.com/CatholicOS/ontology-semantic-canon, `queries/jena/01-church-hierarchy.rq` (`osc:RChKPk9K152BirrIYgAREsY` = Clergy) and + `sources/ontology-semantic-canon.ttl`; ontokit, github.com/CatholicOS/ontokit-web, `lib/ontology/iriGeneration.ts` (`uuidToBase62`, prefixed `"R"` "to ensure RDF/XML QName + safety"). The same repository already ships the affordance layer beside those opaque IDs: `lib/hooks/useIriLabels.ts` resolves a set of IRIs to `rdfs:label`s and label-joins + them into API responses (cached per project and branch), and the Turtle editor's Monaco hover provider — `registerHoverProvider` in `components/editor/TurtleEditor.tsx` — + renders `Label: ` beside the resolved full IRI on hover. + +[^16]: + FOLIO (Federated Open Legal Information Ontology), https://folio.openlegalstandard.org. The concept _Murder_ resolves at + `https://folio.openlegalstandard.org/R7Ttdyo4FsvaupPKT35Qry0`; every FOLIO concept IRI carries the same `R`-prefixed base62 local name. diff --git a/standards/drafts/cdcf-catholic-uri-scheme-03.md b/standards/drafts/cdcf-catholic-uri-scheme-03.md new file mode 100644 index 0000000..c0d9200 --- /dev/null +++ b/standards/drafts/cdcf-catholic-uri-scheme-03.md @@ -0,0 +1,952 @@ +# CDCF DRAFT PROPOSAL — REVISION 03 + +## Catholic Digital Commons Foundation URI Scheme + +`draft-cdcf-catholic-uri-scheme-03` + +| | | +|---|---| +| **Document ID:** | `draft-cdcf-catholic-uri-scheme-03` | +| **Supersedes:** | `draft-cdcf-catholic-uri-scheme-02` | +| **Status:** | Draft Proposal — Public Comment | +| **Date:** | July 2026 | +| **Version:** | 0.4.0 | +| **Issuing Body:** | Catholic Digital Commons Foundation (CDCF) | +| **Comment Period:** | 60 days from date of publication | +| **License:** | CC-BY 4.0 International | +| **Repository:** | github.com/CatholicOS/cdcf-uri-scheme | +| **Relates to:** | `draft-cdcf-identifier-rationale-00`; `crmedr`, `cecdr`, `ciclsaldr` (CatholicOS sibling registries) | + +--- + +## Revision History + +This document supersedes `draft-cdcf-catholic-uri-scheme-02` (version 0.3.0). This is a **breaking-change release**: it introduces new required vocabulary and a new field shape on every resolution response. All changes introduced in this revision are marked **[REV-03]** and documented in **Appendix F (Change Rationale v0.4.0)**. + +| Version | Date | Summary of Changes | +|---|---|---| +| 0.1.0 | June 2026 | Initial draft published for public comment. | +| 0.2.0 | June 2026 | Theological/canonical review: authority taxonomy corrected; relationship types added; licenseStatus field; authentic interpretations; documentType field; doctrinal development note; Eastern churches; papal disambiguation; 3 new open issues. | +| 0.3.0 | June 2026 | Software architecture review: ABNF grammar (Appendix D); cdcf:rel/ resolution defined; cdcf:book/ domain formalised; licenseStatus replaced with licenses object; response depth levels added (§5.6); @context specification added (§5.3); error response schema (§5.5); caching and resilience spec (§5.7–5.8); person path grammar fixed; issuer-slug unified to pope-{name}-{roman}; bulk resolution noted; change notification feed noted; duplicate documentType row removed; sub-spec versioning defined (§7.5). | +| **0.4.0** | **July 2026** | **Identifier-architecture review** (`draft-cdcf-identifier-rationale-00`): adopts the two-artifact model (ontology IRI + canonical notation) via a new `notations` field (§3.6, §5.9); opens `cdcf:concept/` to an opaque-IRI option (§4.5); replaces the closed `inst-type` enumeration with a small set of supertypes plus a `circumscriptionType` field (§4.7); formalizes cross-references to sibling CatholicOS registries `circ:`, `icl:`, `mr:` (§4.9); adds `cdcf:rel/exact-match` and `cdcf:rel/close-match` in place of blanket `owl:sameAs` (§4.8.2); formalizes `mr:` as a tenth `cdcf:` domain (Appendix D). | + +> **Reviewer Note** +> +> Version 0.4.0 addresses four gaps identified in the identifier-architecture review: (1) the specification minted only one string per entity, conflating graph identity and public citation, where the recommendation is that CDCF carry both; (2) `cdcf:concept/` had no opacity option despite being the one domain where opacity is defensible; (3) `inst-type` was a closed enumeration that could not accommodate CECDR's or CICLSALDR's actual scope; (4) sibling CatholicOS registries (`crmedr`, `cecdr`, `ciclsaldr`) had minted their own prefixes with no declared relationship to `cdcf:`, in tension with the §7.5 sub-specification rule. Theological and canonical content, and all software-architecture content from 0.3.0 not listed above, are unchanged. + +--- + +# Abstract + +This document defines the cdcf: URI scheme — a system of stable, globally unique, dereferenceable identifiers for canonical Catholic entities including Scripture, Magisterial documents, Catechism paragraphs, Canon Law, theological concepts, saints, popes, ecclesiastical institutions, and — new in this revision — Roman Martyrology eulogies referenced from sibling registries. + +The scheme is designed to serve as the foundational reference layer for interoperable Catholic digital infrastructure. It draws on established internet standards (URI syntax per RFC 3986, content negotiation per RFC 7231, Linked Data principles per W3C, ABNF grammar per RFC 5234, and SKOS/ADMS vocabulary for citation-layer identifiers per W3C) while respecting the theological authority structures of the Catholic Church. + +This is a Draft Proposal published for public comment. The comment period closes 60 days from the date of publication. + +# 1. Introduction + +## 1.1 The Problem of Catholic Data Fragmentation + +Catholic digital projects — Bible apps, catechetical platforms, canonical law databases, theological reasoning engines — currently lack any shared reference system. A magisterial document cited in one application cannot be reliably matched to the same document in another. + +- Siloed digitization: each project independently digitizes the same texts, introducing independent errors that cannot be corrected centrally. +- Vendor lock-in: data held in proprietary formats cannot be consumed by other applications, placing governance of Church data outside Church institutions. + +## 1.2 Purpose of This Specification + +1. A canonical identifier for every major category of Catholic entity +2. A resolution protocol returning structured, machine-readable data on dereferencing +3. A formal path grammar ensuring machine-parseable, unambiguous identifiers +4. A governance model for identifier assignment and long-term stability + +## 1.3 Design Principles + +- Stability: once assigned, an identifier must never be reassigned or broken +- Readability: identifiers should be interpretable by humans without a lookup +- Dereferenceability: every identifier resolves to a structured data response +- Openness: the scheme is an open standard, not proprietary to any vendor +- Theological fidelity: metadata respects Catholic authority hierarchies as defined by the Magisterium +- Backwards compatibility: no published identifier may be invalidated by future revisions +- Machine-parseability: all identifier paths conform to the ABNF grammar in Appendix D + +# 2. Terminology + +The key words MUST, MUST NOT, REQUIRED, SHALL, SHOULD, RECOMMENDED, and OPTIONAL are to be interpreted per RFC 2119. ABNF notation is per RFC 5234. + +| Term | Definition | +|---|---| +| Canonical Entity | A Catholic object of theological, juridical, or historical significance with a defined authoritative source | +| Canonical Identifier | A globally unique, stable URI assigned to a canonical entity under this specification | +| Ontology IRI | **[REV-03]** The `cdcf:` URI's role as a reasoner-facing, opaque-to-reasoners atom used for graph identity. See §3.6. | +| Canonical Notation | **[REV-03]** A typed citation string carried in the `notations` array, distinct from the ontology IRI, intended for human citation and interchange with sibling registries. See §3.6. | +| Dereferencing | The act of resolving a URI to its data resource via an HTTP request | +| Response Depth | A request parameter controlling how much related data is inlined in a resolution response (summary vs. full) | +| licenses object | A structured metadata object mapping each content field in a resolution response to its applicable licenseStatus value | +| Assent of Faith | The unconditional assent (assensus fidei) required for truths proposed as divinely revealed | +| Definitive Assent | The firm assent (assensus definitivus) required for truths definitively proposed but not explicitly revealed | +| Religious Submission | Obsequium religiosum — submission of intellect and will owed to authentic but non-definitive magisterial teaching (LG §25) | +| DS Number | Denzinger-Schönmetzer reference number — established scholarly ID for magisterial texts | +| OSIS Code | Open Scripture Information Standard book abbreviation, used for Scripture book references | +| Sui iuris Church | One of the 24 autonomous Catholic churches (1 Latin, 23 Eastern) in full communion with the Bishop of Rome | + +# 3. URI Syntax and Structure + +## 3.1 Base URI + +The authoritative base URI for all cdcf: identifiers is: + +> `https://id.catholiccommons.org/` + +The prefix `cdcf:` is a shorthand notation. Example: + +> `cdcf:verse/jn/3/16` → `https://id.catholiccommons.org/verse/jn/3/16` + +## 3.2 General Structure + +> `cdcf:{domain}/{path}` + +Where `{domain}` is one of the **ten** defined domains (Section 4; nine as of 0.3.0, plus `mr` formalized in this revision — see §4.9 and Appendix D) and `{path}` conforms to the domain-specific ABNF production rules in Appendix D. + +## 3.3 Syntax Rules + +1. All characters MUST be lower-case ASCII +2. Words within a path segment MUST be separated by hyphens (-) +3. Path segments MUST be separated by forward slashes (/) +4. No trailing slashes +5. No query strings in canonical identifiers (query parameters are permitted on API endpoints, which are distinct from canonical identifiers) +6. No version numbers or dates in identifier paths, except where the date is intrinsic to the entity identity (e.g., authentic interpretations of canons are dated acts; martyrology eulogies are anchored to a calendar date per §4.9) +7. Identifiers MUST NOT exceed 256 characters +8. All identifier paths MUST conform to the ABNF grammar in Appendix D + +## 3.4 Stability Guarantee + +Any identifier published in a CDCF stable release MUST remain resolvable and MUST NOT be reassigned. Deprecated identifiers MUST return HTTP 301 with a Location header pointing to the successor. The successor MUST be a new URI. Both the original and successor identifiers MUST remain resolvable indefinitely. + +## 3.5 Licenses Object + +Every resolution response MUST include a licenses object. This replaces the scalar licenseStatus field present in version 0.2.0. A scalar field was insufficient for composite resources whose constituent parts carry different copyright statuses (e.g., a Latin text that is public domain alongside an English translation that is proprietary). + +The licenses object maps each content-bearing field in the response to its applicable licenseStatus value. Fields not listed inherit the value of the root key ("*"). + +```json +"licenses": { + "*": "public-domain", + "text.en-RSVCE": "proprietary-licensed", + "text.en-NABRE": "proprietary-licensed" +} +``` + +Permitted licenseStatus values: + +| licenseStatus value | Meaning | +|---|---| +| public-domain | No copyright restrictions; free use and redistribution | +| cc-by | Creative Commons Attribution; attribution required | +| cc-by-nc | Creative Commons Attribution Non-Commercial | +| proprietary-licensed | CDCF holds a distribution licence; contact CDCF for terms | +| rights-holder-contact-required | No CDCF licence; contact the rights holder directly before redistribution | + +CDCF does not claim to own content it does not hold rights to. The licenses object is informational, not a grant of rights. + +## 3.6 The Two-Artifact Model [NEW — REV-03] + +Every canonical entity under this specification has exactly **two** identifying artifacts, not one: + +1. **The ontology IRI** — the `cdcf:` URI defined by the domain grammars in Appendix D. Its governing requirement is stability and opacity-to-reasoners. A reasoner or triple store MUST treat it as an atomic identifier and MUST NOT parse it to recover meaning, regardless of whether its characters happen to be human-legible. +2. **The canonical notation** — a typed string, carried in the resolution response's `notations` array (§5.9), intended for citation, hand-authoring, and interchange. A notation MAY equal the final path segment of the ontology IRI (the common case for "thing" domains — Scripture, Magisterium, Canon Law, Persons) or MAY be an entirely separate registry-native identifier (the case for sibling registries, §4.9). + +This is not redundancy for its own sake. For domains where the two already coincide (verse, book, ccc, canon, magisterium, person), implementations MAY omit an explicit `notations` entry and treat the IRI's final segment as the default notation. The array becomes load-bearing only where the two genuinely diverge — `concept/` (§4.5) and sibling-registry entities (§4.9). + +### 3.6.1 Non-Normative Note: Why Not `owl:sameAs` Everywhere + +Mature identifier systems that separate graph identity from citation form (BCP 47's registry-plus-grammar model, Unicode's stability policy, Getty/VIAF's authority-record-plus-notation pattern) do not assert blanket identity between the two layers. A notation is not the IRI; a close external match is not an exact one. §4.8.2 and §4.9 define two relationship types — `cdcf:rel/exact-match` and `cdcf:rel/close-match` — that carry this distinction into the cross-reference graph, rather than relying on `owl:sameAs`, whose reasoner-merge semantics are inappropriate wherever two linked resources are similar but not strictly identical. + +# 4. Identifier Domains + +The cdcf: namespace is divided into ten domains as of this revision (nine as of 0.3.0, plus `mr`, formalized in §4.9). Domains revised or added in version 0.4.0 are marked **[REV-03]**; domains revised in 0.3.0 remain marked **[REV-02]** for historical continuity. + +## 4.1 Scripture (cdcf:verse/) and Books (cdcf:book/) + +Version 0.2.0 referenced cdcf:book/ in the resolution example but did not define it. Version 0.3.0 formally introduced cdcf:book/ as a sibling of cdcf:verse/ within the Scripture domain. A book identifier is the parent entity of all verse identifiers within that book. + +### 4.1.1 Book Identifiers (cdcf:book/) + +> `cdcf:book/{osis-book}` + +| URI | Refers To / Description | +|---|---| +| cdcf:book/jn | The Gospel of John | +| cdcf:book/gen | Genesis | +| cdcf:book/ps | Psalms (Vulgate numbering as canonical) | +| cdcf:book/1macc | 1 Maccabees (deuterocanonical) | + +### 4.1.2 Verse Identifiers (cdcf:verse/) + +> `cdcf:verse/{osis-book}/{chapter}/{verse}` +> `cdcf:verse/{osis-book}/{chapter}/{start-verse}-{end-verse}` — verse range +> `cdcf:verse/{osis-book}/{chapter}` — whole chapter + +| URI | Refers To / Description | +|---|---| +| cdcf:verse/jn/3/16 | John 3:16 | +| cdcf:verse/rom/8/28-30 | Romans 8:28–30 | +| cdcf:verse/ps/22 | Psalm 22 (Vulgate numbering, canonical) | +| cdcf:verse/mt/26/26-28 | Matthew 26:26–28 | + +Note: Identifiers MUST use Vulgate Psalm numbering as canonical. The resolution response MUST include a psalmNumberingMap field with the Hebrew/Protestant equivalent where they differ. See Open Issue 1 (Section 10). + +## 4.2 Magisterial Documents (cdcf:magisterium/) + +> `cdcf:magisterium/{issuer-slug}/{document-slug}` +> `cdcf:magisterium/{issuer-slug}/{document-slug}/{section-type}/{section-id}` + +Issuer slug convention: Version 0.2.0 used inconsistent conventions (e.g., leo13 in the magisterium domain vs. pope-leo-xiii in the person domain for the same pontiff). Version 0.3.0 unifies all papal slugs to the form pope-{name}-{roman-numeral} across both domains. A magisterialIssuedBy field in the resolution response carries the canonical cdcf:person/ identifier of the issuing authority, making the link machine-derivable. + +| Domain | Issuer Slug Form | Example | +|---|---|---| +| cdcf:magisterium/ | pope-{name}-{roman} | cdcf:magisterium/pope-leo-xiii/rerum-novarum | +| cdcf:magisterium/ | pope-{name}-{roman} | cdcf:magisterium/pope-pius-ix/ineffabilis-deus | +| cdcf:magisterium/ | {council-slug} | cdcf:magisterium/trent/decree-justification | +| cdcf:magisterium/ | {dicastery-slug} | cdcf:magisterium/ddf/fiducia-supplicans | +| cdcf:person/ | pope-{name}-{roman} | cdcf:person/pope-leo-xiii | + +### 4.2.1 Document Type Field + +Every magisterial document resolution response MUST include a documentType field. + +| documentType value | Latin Term | Description / Binding Weight | +|---|---|---| +| dogmatic-constitution | Constitutio dogmatica | Ecumenical council; defines doctrine or condemns error; assent of faith required | +| apostolic-constitution | Constitutio apostolica | Highest papal legislative act; used for dogma definition and major law | +| solemn-definition | Definitio sollemnis | Ex cathedra papal definition; highest authority; assent of faith required | +| encyclical | Litterae encyclicae | Circular letter on faith, morals, discipline; authentic magisterial teaching | +| apostolic-exhortation | Adhortatio apostolica | Post-synodal or devotional; pastoral; not definitional | +| declaration | Declaratio | Formal declaration clarifying doctrine or discipline | +| instruction | Instructio | Practical guidance for implementing magisterial teaching; dicastery-issued | +| council-decree | Decretum conciliare | Conciliar decree; weight depends on papal confirmation | +| response-authentic | Responsio authentica | Authentic interpretation by Dicastery for Legislative Texts; force of law | + +### 4.2.2 Papal Approval Field + +The resolution response MUST include a papalApproval field: + +- ex-cathedra — solemn definition by the pope personally +- in-forma-specifica — papal approval incorporating the document as the pope's own act +- in-forma-communi — papal approval of the dicastery's competence but not the document's specific content +- dicastery-only — issued by a dicastery without explicit papal approval +- conciliar — promulgated by ecumenical council confirmed by the pope + +> **Example: Fiducia Supplicans (2023)** +> `cdcf:magisterium/ddf/fiducia-supplicans`: documentType=declaration, papalApproval=in-forma-communi, authorityLevel=authentic-doctrine. + +## 4.3 Catechism (cdcf:ccc/) + +Resolution responses for CCC text MUST carry `licenses: { "text.*": "proprietary-licensed" }` as copyright in all modern translations is held by Libreria Editrice Vaticana and/or the relevant bishops' conference. + +> `cdcf:ccc/{paragraph}` +> `cdcf:ccc/{start}-{end}` + +| URI | Refers To / Description | +|---|---| +| cdcf:ccc/1324 | CCC §1324 — The Eucharist as source and summit | +| cdcf:ccc/891 | CCC §891 — Papal infallibility | +| cdcf:ccc/1730-1748 | CCC §§1730–1748 — Human freedom section | + +## 4.4 Canon Law (cdcf:canon/) + +> `cdcf:canon/{code}/{canon}` +> `cdcf:canon/{code}/{canon}/{para}` +> `cdcf:canon/{code}/{canon}/authentic-interpretation/{date}` + +| URI | Refers To / Description | +|---|---| +| cdcf:canon/cic1983/1024 | CIC 1983, Canon 1024 | +| cdcf:canon/cic1983/844/1 | CIC 1983, Canon 844 §1 | +| cdcf:canon/cic1983/230/authentic-interpretation/1994-11-11 | Auth. interp. of Canon 230, 11 Nov 1994 | +| cdcf:canon/cceo/7 | CCEO, Canon 7 | +| cdcf:canon/cic1917/1023 | CIC 1917, Canon 1023 (historical) | + +The ABNF grammar in Appendix D makes the three path forms unambiguous: a canon paragraph is always numeric; authentic-interpretation is a fixed keyword; dates follow ISO 8601. No heuristic parsing is required. + +## 4.5 Theological Concepts (cdcf:concept/) [REVISED — REV-03] + +> **Limitation: Doctrinal Development** +> This domain treats each concept as a single stable entity. Doctrinal concepts develop over time (cf. Newman, *Essay on the Development of Christian Doctrine*; Dei Verbum §8). The doctrinalHistory field (an ordered array of cdcf:magisterium/ references) surfaces this development without resolving the formal question of how to represent it. A dedicated doctrinal-development specification is planned. + +Version 0.3.0 and earlier defined this domain as a plain mnemonic slug with no opacity option, despite the callout above acknowledging that concept boundaries are a live theological question rather than a settled fact — precisely the case in which opacity is defensible (cf. Gene Ontology/OBO Foundry precedent, and §3.6). Version 0.4.0 opens the grammar to accept **either** form: + +> `cdcf:concept/{slug}` — mnemonic form (unchanged from 0.3.0) +> `cdcf:concept/{opaque-id}` — **[NEW]** opaque form, e.g. `cdcf:concept/C0000418` + +New concept entries SHOULD mint an opaque primary IRI and carry the existing mnemonic form as a `skos:notation`-typed entry in `notations` (§5.9). Concepts already assigned a mnemonic IRI under 0.3.0 are **not** required to migrate (§1.3, backwards compatibility); the opaque form is available for new entries and RECOMMENDED wherever the committee expects the concept's boundary to be revised. A `supersededNotation` field MAY record a notation retired after a boundary revision. + +| URI | Refers To / Description | Notation (§5.9) | +|---|---|---| +| cdcf:concept/C0000418 | Real change of substance in the Eucharist | transubstantiation | +| cdcf:concept/hypostatic-union | Union of divine and human natures in Christ (legacy mnemonic form) | — | +| cdcf:concept/purgatory | State of purification after death | — | +| cdcf:concept/natural-law | Moral law knowable by unaided reason | — | +| cdcf:concept/filioque | Procession of the Holy Spirit from Father and Son | — | +| cdcf:concept/original-sin | Sin of Adam and its transmission to all humanity | — | + +Resolution response MUST include: canonical definition (from CCC or proximate magisterium), doctrinalHistory array, authorityLevel, notations array (§5.9), and suiIurisChurch where the concept has distinct elaboration in an Eastern tradition. + +## 4.6 Persons (cdcf:person/) + +Version 0.3.0 corrects the path grammar for works. The four-segment pattern in version 0.2.0 was broken by its own Summa Theologiae example, which required five segments. The path now uses a variable-depth segment for works. + +> `cdcf:person/{person-slug}` +> `cdcf:person/{person-slug}/{work-slug}` +> `cdcf:person/{person-slug}/{work-slug}/{path-segment}+` — one or more sub-divisions + +The `{path-segment}+` production allows arbitrary depth to accommodate works with three or more levels of internal structure (e.g., Summa Theologiae: part / question / article; Scriptum super Sententiis: book / distinction / question / article). See Appendix D for the ABNF production. + +### 4.6.1 Papal Name Disambiguation + +All papal identifiers MUST include an ordinal suffix, even for names held by only one pope, to guarantee forward stability. + +| URI | Refers To / Description | +|---|---| +| cdcf:person/pope-francis-i | Pope Francis (Jorge Mario Bergoglio) | +| cdcf:person/pope-john-paul-ii | Pope John Paul II (Karol Wojtyła) | +| cdcf:person/pope-benedict-xvi | Pope Benedict XVI (Joseph Ratzinger) | +| cdcf:person/pope-leo-xiii | Pope Leo XIII (Vincenzo Gioacchino Pecci) | +| cdcf:person/thomas-aquinas | Thomas Aquinas | +| cdcf:person/thomas-aquinas/summa-theologiae | Summa Theologiae | +| cdcf:person/thomas-aquinas/summa-theologiae/iii/q75/a1 | ST III, Q.75, A.1 (variable-depth path) | + +## 4.7 Institutions (cdcf:institution/) [REVISED — REV-03] + +> `cdcf:institution/{supertype}/{slug}` + +Version 0.3.0's `inst-type` was a closed seven-item enumeration (`diocese`, `conference`, `council`, `order`, `dicastery`, `seminary`, `eastern-church`) that could not represent CECDR's actual scope (eparchies, exarchates, territorial prelatures, apostolic vicariates, military and personal ordinariates, personal prelatures, missions sui iuris) or CICLSALDR's institute/family distinction. Version 0.4.0 replaces the closed enum with a small, stable set of **supertypes**, moving canon-law-specific detail into a response field rather than the URI path: + +| Supertype | Covers | Subtype field | +|---|---|---| +| circumscription | Dioceses, archdioceses, eparchies, archeparchies, exarchates, territorial prelatures/abbacies, apostolic vicariates/prefectures/administrations, military and personal ordinariates, personal prelatures, missions sui iuris | `circumscriptionType` (open string, e.g. "diocese", "eparchy", "territorial-prelature") | +| institute | Individual institutes of consecrated life and societies of apostolic life | — | +| family | Religious-family groupings under Praenotanda n. 38 (e.g. the Franciscan family) | — | +| conference | Bishops' conferences | *(unchanged)* | +| council | Ecumenical and particular councils | *(unchanged)* | +| dicastery | Roman Curia dicasteries | *(unchanged)* | +| seminary | Seminaries | *(unchanged)* | +| eastern-church | Whole sui iuris Churches (distinct from individual eparchies, which fall under circumscription) | — | + +The `order` supertype from 0.3.0 is retired in favor of the `institute`/`family` split. Existing `cdcf:institution/order/{slug}` identifiers continue to resolve per §1.3 (backwards compatibility); their responses carry a `successor` field pointing to the equivalent `cdcf:institution/institute/{slug}`. + +Examples: + +``` +cdcf:institution/circumscription/us-boston (circumscriptionType: "diocese") +cdcf:institution/circumscription/al-shkodre-pult (circumscriptionType: "eparchy") +cdcf:institution/institute/ofmcap +cdcf:institution/family/franciscan +cdcf:institution/eastern-church/maronite (unchanged — whole Church, not an eparchy) +``` + +Per §4.9, `us-boston`, `ofmcap`, and `franciscan` are the same slugs CECDR and CICLSALDR use in `circ:us-boston`, `icl:ofmcap`, and `icl:familia-franciscana` respectively (with the `familia-` prefix dropped, since the supertype segment already disambiguates institute from family). + +### 4.7.1 Eastern Catholic Churches + +The Catholic Church comprises 24 sui iuris churches. Each MUST be representable without privileging the Latin Church as the default. Eastern Church identifiers use the eastern-church supertype. + +> `cdcf:institution/eastern-church/{slug}` + +| URI | Refers To / Description | +|---|---| +| cdcf:institution/eastern-church/maronite | Maronite Catholic Church (Antiochene rite) | +| cdcf:institution/eastern-church/coptic-catholic | Coptic Catholic Church (Alexandrian rite) | +| cdcf:institution/eastern-church/melkite | Melkite Greek Catholic Church (Byzantine/Antiochene) | +| cdcf:institution/eastern-church/chaldean | Chaldean Catholic Church (East Syriac rite) | +| cdcf:institution/eastern-church/ukrainian-greek-catholic | Ukrainian Greek Catholic Church (Byzantine rite) | +| cdcf:institution/council/trent | Council of Trent (1545–1563) | +| cdcf:institution/dicastery/ddf | Dicastery for the Doctrine of the Faith | +| cdcf:institution/conference/usccb | United States Conference of Catholic Bishops | + +## 4.8 Relationship Types (cdcf:rel/) + +The cdcf:rel/ domain defines a controlled vocabulary of relationship types for cross-reference graph edges. This was introduced in version 0.2.0 but left the question of what dereferencing a cdcf:rel/ identifier returns unanswered. Version 0.3.0 resolved this; version 0.4.0 adds two relationship types for cross-registry alignment (§4.8.2, §4.9). + +### 4.8.1 Resolution of cdcf:rel/ Identifiers + +cdcf:rel/ identifiers are dereferenceable. When resolved, they return an OWL ObjectProperty declaration. The resolution response for `Accept: application/ld+json` is a JSON-LD `@type: owl:ObjectProperty` document. The response for `Accept: application/rdf+xml` returns the OWL declaration for import into an ontology. The response for `Accept: text/html` returns a human-readable description. + +Example: `GET https://id.catholiccommons.org/rel/magisterial-proof-text` with `Accept: application/ld+json` returns: + +```json +{ + "@context": { + "owl": "http://www.w3.org/2002/07/owl#", + "cdcf": "https://id.catholiccommons.org/" + }, + "@id": "cdcf:rel/magisterial-proof-text", + "@type": "owl:ObjectProperty", + "rdfs:label": "magisterial proof text", + "rdfs:comment": "The subject is cited by the Magisterium as a scriptural foundation for the object doctrine or concept.", + "rdfs:domain": "cdcf:BibleVerse", + "rdfs:range": "cdcf:Concept" +} +``` + +### 4.8.2 Relationship Type Vocabulary + +| cdcf:rel/ value | OWL Domain | OWL Range | Description | +|---|---|---|---| +| magisterial-proof-text | BibleVerse | Concept | Cited by the Magisterium as scriptural foundation | +| patristic-interpretation | BibleVerse | Concept | Interpreted in this doctrinal sense by a Church Father | +| liturgical-use | BibleVerse | Concept | Used in liturgy expressing this doctrine | +| theological-argument | BibleVerse | Concept | Used in theological argument; not magisterial | +| typological-sense | BibleVerse | BibleVerse | OT type fulfilled in NT antitype | +| anagogical-sense | BibleVerse | Concept | Points toward eschatological realities | +| defines-concept | MagisterialDoc | Concept | Document formally defines this concept | +| supersedes | MagisterialDoc | MagisterialDoc | Later document formally abrogates earlier | +| interprets | MagisterialDoc | Canon | Authentic interpretation of a canon | +| contradicted-by | Any | Any | Scholarly annotation; not a doctrinal assertion | +| **exact-match** **[NEW — REV-03]** | Any | Any | Maps to `skos:exactMatch`. Strict identity between a `cdcf:` entity and an entity in a sibling or external registry. Use sparingly — see §3.6.1. | +| **close-match** **[NEW — REV-03]** | Any | Any | Maps to `skos:closeMatch`. Overlapping but non-identical referent (e.g. a GeoNames territorial point standing in for a diocese-as-institution). | + +## 4.9 Sibling and External Registries [NEW — REV-03] + +CatholicOS maintains sibling data repositories — CRMEDR (Roman Martyrology eulogies, `mr:`), CECDR (ecclesiastical circumscriptions, `circ:`), CICLSALDR (institutes of consecrated life, `icl:`) — that mint their own citation-layer prefixes for entities that also have `cdcf:` ontology IRIs. Version 0.4.0 formalizes their relationship to this specification: + +- A sibling registry's identifier is carried as an entry in `notations` (§5.9), with `scheme` set to the registry's declared scheme URN (`cdcf:scheme/circ`, `cdcf:scheme/icl`, `cdcf:scheme/mr`). +- The sibling registry's own slug MUST be reused verbatim as the final path segment of the corresponding `cdcf:` IRI wherever one exists (e.g. `circ:us-boston` ↔ `cdcf:institution/circumscription/us-boston`). This is a MUST, not a convention: it is what makes the pairing machine-verifiable rather than merely coincidental. +- Cross-registry links that assert **strict identity** use `crossReferences` with `@type: "cdcf:rel/exact-match"` (§4.8.2). Links to **external, non-identical authority records** (Wikidata, GeoNames, VIAF) that model an overlapping but non-identical referent MUST use `@type: "cdcf:rel/close-match"` rather than an identity-asserting relation, to avoid the reasoner-merge hazard `owl:sameAs` creates between non-identical resources (§3.6.1). `cdcf:rel/exact-match` maps to `skos:exactMatch` in the JSON-LD `@context`; `cdcf:rel/close-match` maps to `skos:closeMatch`. Neither maps to `owl:sameAs`; a future sub-specification MAY define narrower conditions under which `owl:sameAs` is warranted. +- Sibling registries publishing under this section MUST declare a `compatibleWith` field in their own schema documents (§7.5, unchanged) naming the `cdcf-uri-scheme` version they target — `compatibleWith: "0.4.0"`. +- `mr:` is additionally formalized as a `cdcf:` domain in its own right (Appendix D §D.3, §D.5), including ABNF for the owner-segment composite form (`mr:cecdr/us-boston:0705-{slug}`, `mr:ciclsaldr/ofm:0824-{slug}`), previously prose-only in CRMEDR's documentation. + +# 5. Resolution Protocol + +## 5.1 HTTP Resolution + +All cdcf: identifiers MUST be resolvable via HTTPS GET to the base URI. The resolution server MUST support content negotiation via the HTTP Accept header. + +## 5.2 Content Negotiation + +| Accept Header | Format | Use Case | +|---|---|---| +| text/html | HTML | Human-readable browser view | +| application/json | JSON | Application integration | +| application/ld+json | JSON-LD | Linked Data / semantic web | +| application/rdf+xml | RDF/XML | OWL ontology systems | +| text/turtle | Turtle RDF | SPARQL / triplestore import | + +## 5.3 JSON-LD Context Specification + +Version 0.2.0 referenced `https://id.catholiccommons.org/context/v1` in example responses without specifying the context document. Version 0.3.0 corrected this. Version 0.4.0 requires the context document to additionally declare the `skos:` and `adms:` prefixes used by the `notations` field (§5.9) and the two new relationship types (§4.8.2). + +### 5.3.1 Context Document Requirements + +- The context document at `https://id.catholiccommons.org/context/v1` MUST map all field names used in cdcf: resolution responses to their fully qualified IRIs +- The context document MUST be served with `Cache-Control: public, max-age=31536000, immutable` (one year; it is versioned and therefore stable) +- The context document MUST be embedded inline in resolution responses (the `@context` field MUST contain the expanded context object, not only the URL) when the request includes `Prefer: return=representation` or when the response is for a cdcf:rel/ identifier +- Context versions are immutable: v1 will never change. If a new field requires a different mapping, a v2 context MUST be published. Responses using v2 fields (including `notations`, `skos:exactMatch`, `skos:closeMatch`) MUST reference `context/v2` **[REV-03]** +- The context MUST declare the cdcf: prefix, all @type values, and all relationship predicates used in §4.8, and, as of v2, the `skos:` and `adms:` prefixes used in §5.9 **[REV-03]** + +### 5.3.2 Context Availability + +Consuming applications MUST NOT assume the context URL is always reachable at request time. Implementations SHOULD cache the context document locally. CDCF MUST publish the context document as a static file on its GitHub repository so that implementations can bundle it directly, eliminating the remote dependency. + +## 5.4 Example Resolution Response + +The following example shows the full response for `cdcf:verse/jn/6/53`, including the licenses object, inline `@context`, and typed relationship edges: + +```json +{ + "@context": { + "cdcf": "https://id.catholiccommons.org/", + "owl": "http://www.w3.org/2002/07/owl#", + "rdfs": "http://www.w3.org/2000/01/rdf-schema#" + }, + "@id": "cdcf:verse/jn/6/53", + "@type": "cdcf:BibleVerse", + "book": "cdcf:book/jn", + "chapter": 6, "verse": 53, + "osisRef": "John.6.53", + "licenses": { + "*": "public-domain", + "text.en-RSVCE": "proprietary-licensed", + "text.en-NABRE": "proprietary-licensed" + }, + "text": { + "la-Vulgata": "Dixit ergo eis Iesus: Amen amen dico...", + "en-DRC": "Then Jesus said to them: Amen, amen..." + }, + "crossReferences": [ + { "@type": "cdcf:rel/magisterial-proof-text", + "target": "cdcf:concept/real-presence", + "citedIn": "cdcf:magisterium/trent/decree-eucharist/canon/1" }, + { "@type": "cdcf:rel/patristic-interpretation", + "target": "cdcf:concept/real-presence", + "citedIn": "cdcf:person/cyril-of-alexandria/commentary-on-john" } + ] +} +``` + +See §5.9 for the `notations` field, illustrated there with a sibling-registry example. + +## 5.5 HTTP Status Codes and Error Response Schema + +Version 0.2.0 defined status codes only. Version 0.3.0 added a mandatory error response body schema for all 4xx responses. + +| HTTP Code | Condition | Error Response Body Required | +|---|---|---| +| 200 | Success | Full resource representation | +| 301 | Deprecated; use Location header | `{ "error": "deprecated", "successor": "cdcf:...", "since": "YYYY-MM-DD" }` | +| 404 | Identifier not yet assigned | `{ "error": "not-found", "identifier": "...", "message": "...", "suggestion": null }` | +| 406 | Accept type not supported | `{ "error": "not-acceptable", "supported": [...] }` | +| 410 | Retired; no successor | `{ "error": "gone", "identifier": "...", "retiredOn": "YYYY-MM-DD" }` | + +All error responses MUST be served with `Content-Type: application/json`. The message field in 404 responses SHOULD include human-readable detail (e.g., "Chapter 99 does not exist in the Gospel of John"). The suggestion field MAY carry an alternative identifier if a close match exists in the registry. + +## 5.6 Response Depth + +Version 0.2.0 did not define how much related data is inlined in responses. Response depth is controlled by the `Depth` request header (per the pattern of RFC 7240 Prefer header). Two depth levels are defined: + +- `Depth: 0` (default) — summary: returns core metadata and identifier arrays only; no nested objects are inlined. `crossReferences` is an array of cdcf: URI strings. +- `Depth: 1` — full: returns core metadata with one level of related objects inlined. `crossReferences` is an array of objects with `@type`, `target`, and `citedIn` fields as shown in §5.4. + +Depth values above 1 are reserved for future use. Implementations MUST treat unrecognised Depth values as `Depth: 0`. The default behaviour (no Depth header) is equivalent to `Depth: 0`. + +> **Design Note: N+1 and Bulk Resolution** +> `Depth: 1` mitigates the most common N+1 request pattern. A bulk resolution endpoint (`POST /resolve` with an array of identifiers) is noted as a near-term implementation requirement and is planned for version 1.0.0 of the API specification. It is outside the scope of this identifier scheme document but is acknowledged here to prevent incompatible implementations. + +## 5.7 Caching + +The resolution server MUST set the following headers on all 200 responses for stable identifiers (those in a published CDCF release): + +- `Cache-Control: public, max-age=31536000, immutable` — one year; stable identifiers do not change +- `ETag` — a hash of the response body, supporting conditional GET requests (If-None-Match) +- `Last-Modified` — the date of the last substantive change to the resource + +Draft identifiers (identifiers in a draft release, not yet stable) MUST be served with `Cache-Control: public, max-age=86400` (one day). + +## 5.8 Operational Resilience + +The resolution server at `id.catholiccommons.org` is the single canonical hostname and MUST have a 99.9% monthly availability SLA. To meet this SLA and to provide geographic distribution, the following are REQUIRED by version 1.0.0: + +1. At least two geographically distributed mirror servers hosting identical read-only replicas +2. A mirror discovery endpoint at `https://id.catholiccommons.org/.well-known/cdcf-mirrors` returning a JSON array of mirror base URIs +3. A change notification feed at `https://id.catholiccommons.org/changes` (Atom format, RFC 4287) publishing an entry for every identifier deprecation, addition, and correction +4. DNS records for `id.catholiccommons.org` MUST be protected with DNSSEC + +Consuming applications SHOULD implement fallback logic: if the canonical hostname is unreachable, the application SHOULD attempt the first mirror from the most recently cached mirror list before failing. Because stable identifiers carry `max-age=31536000`, most production applications will serve responses from local cache without contacting the resolution server for most requests. + +## 5.9 Notations Field [NEW — REV-03] + +Every resolution response MUST include a `notations` array, per the two-artifact model of §3.6 (empty or containing only the IRI's own final segment where no divergent notation exists): + +```json +"notations": [ + { "value": "transubstantiation", "scheme": "cdcf:mnemonic", "prefLabel": true } +] +``` + +- `value` — the notation string. +- `scheme` — a declared CDCF or sibling-registry scheme identifier (`cdcf:mnemonic` for a domain's own legacy slug form, or `cdcf:scheme/circ` / `cdcf:scheme/icl` / `cdcf:scheme/mr` for sibling registries, per §4.9). +- `prefLabel` — OPTIONAL boolean; true marks the notation as the preferred display form where more than one exists. + +Example — `cdcf:institution/circumscription/us-boston` carrying its CECDR sibling notation: + +```json +"notations": [ + { "value": "us-boston", "scheme": "cdcf:scheme/circ", "prefLabel": true } +] +``` + +# 6. Theological Authority Metadata + +Unchanged from version 0.2.0. The three-level taxonomy follows the 1989 Profession of Faith and the 1998 CDF commentary. + +> **Normative Sources** +> 1989 Profession of Faith (AAS 81, 1989, pp. 104-106); Ad Tuendam Fidem, John Paul II (1998); Commentary on Professio fidei, CDF/Card. Ratzinger (AAS 90, 1998, pp. 544-551); Donum Veritatis, CDF (1990); Lumen Gentium §25. + +## 6.1 Three Levels of Assent + +### Level 1 — Assent of Faith (credenda) + +Truths contained in the Word of God and proposed by the Magisterium as divinely revealed. Require assent of theological faith. Binding under penalty of heresy if formally denied. Include both solemn definitions and truths proposed by the ordinary universal magisterium (LG §25). Both sub-types require the same assent of faith. + +### Level 2 — Definitive Assent (tenenda) + +Truths proposed definitively but not explicitly contained in divine revelation. Require definitive assent — firm and irrevocable but not an act of theological faith. Denial is error but not formal heresy. + +### Level 3 — Religious Submission (sequenda) + +Teachings of the authentic ordinary magisterium not proposed definitively. Require religious submission of intellect and will (obsequium religiosum). Largest category; includes most encyclicals and dicastery documents. + +## 6.2 Authority Level Taxonomy + +| cdcf:authority/ value | Assent Type | Notes | +|---|---|---| +| solemn-definition | Assent of faith | Ex cathedra papal definition; de fide definita | +| ordinary-universal | Assent of faith | Ordinary universal magisterium; de fide catholica; same assent as solemn-definition | +| definitive-doctrine | Definitive assent | Definitively tenenda; denial is error but not heresy | +| authentic-doctrine | Religious submission | Authentic ordinary magisterium; obsequium religiosum | +| pastoral-guidance | Respectful reception | Pastoral or prudential; no doctrinal definition | + +# 7. Governance and Change Process + +## 7.1 Identifier Assignment + +New identifiers require: entity to be identified, proposed URI slug, justification for slug form, reference to at least one authoritative source, confirmation the identifier conforms to the ABNF grammar in Appendix D. + +## 7.2 Change Process + +1. Draft: proposal published on CDCF GitHub for public comment (minimum 30 days) +2. Last Call: final comment period (14 days) after revision +3. Published: ratified and versioned release + +No change may invalidate a previously published stable identifier. + +## 7.3 Theological and Canonical Review + +Proposals touching doctrinal classifications, authority-level metadata, or concept definitions MUST receive review from at least one qualified theologian and one qualified canonist before entering Last Call. + +## 7.4 Versioning + +This specification follows semantic versioning (MAJOR.MINOR.PATCH). Version 1.0.0 will be declared when the specification has been implemented in at least two independent systems and has passed a 90-day stability review. + +## 7.5 Sub-Specification Versioning + +Future sub-specifications (particular law, Eastern liturgy, patristics, liturgical domain) will be published as separate documents under the naming convention: + +> `draft-cdcf-{topic}-{nn}` + +Examples: draft-cdcf-particular-law-00, draft-cdcf-eastern-liturgy-00, draft-cdcf-patristics-00. Sub-specifications MUST reference the main cdcf-uri-scheme version they are compatible with using a `compatibleWith` field in their header table. A sub-specification that introduces a new domain or modifies an existing domain's ABNF MUST be accompanied by a patch release of the main specification incorporating the grammar change. **[REV-03: CRMEDR, CECDR, and CICLSALDR are the first sibling registries to be brought into compliance with this section — see §4.9.]** + +# 8. Security Considerations + +- The resolution server MUST be served exclusively over HTTPS (TLS 1.2 minimum) +- The registry MUST be maintained under version control with a public audit log +- The resolution server MUST NOT accept write operations via the public endpoint +- Implementations MUST validate that responses originate from `id.catholiccommons.org` or a listed mirror +- DNS records MUST be protected with DNSSEC +- The JSON-LD context document MUST be served with Subresource Integrity (SRI) hashes published on the CDCF GitHub repository, so that implementations bundling the context locally can verify its integrity + +# 9. Open Standards Alignment + +- RFC 2119 — Requirement Levels +- RFC 3986 — URI Generic Syntax +- RFC 4287 — Atom Syndication Format (change notification feed) +- RFC 5234 — ABNF for Syntax Specifications +- RFC 7231 — HTTP/1.1 Semantics (content negotiation) +- RFC 7240 — Prefer Header for HTTP (response depth pattern) +- W3C JSON-LD 1.1 — resolution response format +- W3C OWL 2 — cdcf:rel/ property declarations +- **W3C SKOS — `skos:exactMatch` / `skos:closeMatch` cross-registry relationships (§4.8.2, §4.9) [NEW — REV-03]** +- ISO/IEC 21838-2 — Basic Formal Ontology (BFO) +- OSIS 2.1.1 — Scripture book codes + +# 10. Open Issues for Public Comment + +## Issues from Version 0.1.0 (unresolved) + +1. Psalm numbering: Vulgate vs. Hebrew as primary? Should both be first-class? +2. Deuterocanonical books: Catholic canon as normative; Protestant versification as secondary mapping? +3. Amended magisterial documents: sub-identifiers or replacement documents for post-promulgation corrections? +4. Liturgical domain: dependencies requiring inclusion in version 1.0.0? +5. Multilingual slugs: should non-English document title forms be permitted? +6. Patristic works: is a dedicated patristics sub-specification needed? + +## Issues from Version 0.2.0 (unresolved) + +7. Dicastery approval: should a disputedApproval flag be included for papalApproval fields where the approval mode is contested? +8. Eastern theological concepts: separate identifiers for Eastern equivalents (theosis, epiclesis) or an easternEquivalents field on the Latin-rite identifier? +9. Particular law: cdcf:particular-law/ domain in version 1.0.0 or a standalone sub-specification? + +## Issues from Version 0.3.0 (unresolved) + +10. Bulk resolution API: the POST /resolve endpoint is noted in §5.6. Should its request/response schema be specified in this document or in a separate API specification? +11. Change notification scope: the Atom feed in §5.8 covers identifier changes. Should it also include changes to the cdcf:rel/ OWL declarations and the JSON-LD context? What is the feed retention policy? +12. ABNF stability: the grammar in Appendix D is normative. How should future domains that require new path patterns trigger a revision to the grammar? Should the ABNF be maintained in a separate machine-readable file in the repository? + +## New Issues — Version 0.4.0 [REV-03] + +13. `owl:sameAs` narrowing: §4.9 declines to map anything to `owl:sameAs` pending a future sub-specification defining narrower identity conditions. Should that sub-specification be drafted now, or should `owl:sameAs` simply remain unused indefinitely? +14. Notation scheme registry: `scheme` values in §5.9 (`cdcf:mnemonic`, `cdcf:scheme/circ`, etc.) are introduced ad hoc in this revision. Should CDCF maintain a formal, dereferenceable registry of notation scheme URNs, analogous to the IANA Language Subtag Registry? +15. `circumscriptionType` vocabulary: §4.7 leaves `circumscriptionType` an open string rather than a controlled vocabulary. Should it be closed to an enumerated list (diocese, archdiocese, eparchy, archeparchy, exarchate, territorial-prelature, territorial-abbacy, apostolic-vicariate, apostolic-prefecture, apostolic-administration, military-ordinariate, personal-ordinariate, personal-prelature, mission-sui-iuris), and if so, who maintains that list as new forms arise? +16. Retroactive concept migration: §4.5 does not require existing mnemonic concept IRIs to migrate to the opaque form. Should there be a threshold (e.g., a doctrinal-development sub-specification reaching Last Call) that triggers mandatory migration for a specific concept? + +# 11. References + +## 11.1 Normative References + +- RFC 2119: Bradner, S. (1997) +- RFC 3986: URI Generic Syntax. Berners-Lee, T. et al. (2005) +- RFC 4287: Atom Syndication Format. Nottingham, M. & Sayre, R. (2005) +- RFC 5234: ABNF for Syntax Specifications. Crocker, D. & Overell, P. (2008) +- RFC 7231: HTTP/1.1 Semantics. Fielding, R. & Reschke, J. (2014) +- RFC 7240: Prefer Header for HTTP. Snell, J. (2014) +- W3C SKOS Reference. Miles, A. & Bechhofer, S. W3C Recommendation, 2009. **[NEW — REV-03]** +- Catechism of the Catholic Church, Second Edition (1997). Libreria Editrice Vaticana. +- Code of Canon Law (CIC 1983). Libreria Editrice Vaticana. +- Code of Canons of the Eastern Churches (CCEO 1990). Libreria Editrice Vaticana. +- ISO/IEC 21838-2:2021. Basic Formal Ontology (BFO). +- Ad Tuendam Fidem, John Paul II. AAS 90 (1998), pp. 457-461. +- Commentary on the Professio fidei, CDF/Card. Ratzinger. AAS 90 (1998), pp. 544-551. +- Donum Veritatis, CDF (1990). AAS 82, pp. 1550-1570. + +## 11.2 Informative References + +- Denzinger-Schönmetzer. Enchiridion Symbolorum. Freiburg: Herder, 1965. +- Newman, J.H. Essay on the Development of Christian Doctrine. London: Toovey, 1845. +- OSIS Specification v2.1.1. American Bible Society, 2014. +- W3C JSON-LD 1.1. Sporny, M. et al. W3C Recommendation, 2020. +- W3C OWL 2 Primer. Hitzler, P. et al. W3C Recommendation, 2012. +- BibleGet I/O API. Grasso, J. github.com/BibleGet-I-O. +- CatholicOS ontology-semantic-canon. D'Orazio, J. github.com/CatholicOS. +- CatholicOS crmedr, cecdr, ciclsaldr. github.com/CatholicOS. **[NEW — REV-03]** +- draft-cdcf-identifier-rationale-00. CDCF Identifier Architecture Committee. **[NEW — REV-03]** + +# Appendix A. OSIS Book Code Reference (Selected) + +| OSIS | Book | OSIS | Book | Canon | +|---|---|---|---|---| +| Gen | Genesis | Matt | Matthew | OT/NT | +| Exod | Exodus | Mark | Mark | OT/NT | +| Ps | Psalms | Luke | Luke | OT/NT | +| Prov | Proverbs | Jn | John | OT/NT | +| Isa | Isaiah | Rom | Romans | OT/NT | +| Dan | Daniel | 1Cor | 1 Corinthians | OT/NT | +| Tob | Tobit | Eph | Ephesians | Deuterocanon/NT | +| Sir | Sirach | Rev | Revelation | Deuterocanon/NT | +| 1Macc | 1 Maccabees | 2Macc | 2 Maccabees | Deuterocanon | + +# Appendix B. Eastern Catholic Sui Iuris Churches (Complete List) + +| cdcf: slug | Church | Rite / Tradition | +|---|---|---| +| eastern-church/maronite | Maronite Catholic Church | West Syriac (Antiochene) | +| eastern-church/melkite | Melkite Greek Catholic Church | Byzantine (Antiochene) | +| eastern-church/ukrainian-greek-catholic | Ukrainian Greek Catholic Church | Byzantine | +| eastern-church/chaldean | Chaldean Catholic Church | East Syriac | +| eastern-church/syro-malabar | Syro-Malabar Catholic Church | East Syriac | +| eastern-church/syriac-catholic | Syriac Catholic Church | West Syriac | +| eastern-church/coptic-catholic | Coptic Catholic Church | Alexandrian | +| eastern-church/ethiopian-catholic | Ethiopian Catholic Church | Alexandrian | +| eastern-church/armenian-catholic | Armenian Catholic Church | Armenian | +| eastern-church/ruthenian | Ruthenian Catholic Church | Byzantine | +| eastern-church/romanian-greek-catholic | Romanian Greek Catholic Church | Byzantine | +| eastern-church/melkite-jerusalem | Melkite (Jerusalem Patriarchate) | Byzantine | +| eastern-church/syro-malankara | Syro-Malankara Catholic Church | West Syriac | +| eastern-church/bulgarian-greek-catholic | Bulgarian Greek Catholic Church | Byzantine | + +*Note: this list enumerates whole sui iuris Churches under the unchanged `eastern-church` supertype (§4.7). Individual eparchies and exarchates of these Churches are represented separately under `cdcf:institution/circumscription/` as of 0.4.0 — see §4.7.* + +# Appendix C. Change Rationale (Version 0.2.0) + +Carried forward from version 0.2.0 for continuity. See that document for full rationale on authority taxonomy, relationship types, licenseStatus, authentic interpretations, documentType, doctrinal development, Eastern churches, and papal disambiguation. + +# Appendix D. Formal ABNF Grammar [REVISED — REV-03] + +This appendix defines the normative ABNF grammar for all cdcf: identifier paths. All implementations MUST validate identifiers against these productions. The grammar follows RFC 5234 conventions. Productions changed or added in version 0.4.0 are marked **[REV-03]**. + +> **Normative Status** +> This ABNF grammar is normative. Any cdcf: identifier that does not conform to these productions is invalid. Implementations MUST reject non-conforming identifiers. Future domains introduced by sub-specifications MUST extend this grammar via a patch release of this document. + +## D.1 Core Productions + +``` +; Base character classes +ALPHA = %x61-7A ; a-z (lower-case only) +DIGIT = %x30-39 ; 0-9 +HYPHEN = %x2D ; - +SLASH = %x2F ; / +COLON = %x3A ; : [NEW — REV-03] + +slug = 1*(ALPHA / DIGIT / HYPHEN) ; general-purpose slug +roman = 1*(%x69 / %x76 / %x78 / %x6C / %x63 / %x64 / %x6D) + ; i v x l c d m (Roman numerals) + +iso-date = 4DIGIT HYPHEN 2DIGIT HYPHEN 2DIGIT ; YYYY-MM-DD +``` + +## D.2 Top-Level Production [REVISED — REV-03] + +``` +; A cdcf: identifier is a domain followed by a slash and a domain path +cdcf-id = domain SLASH domain-path + +domain = "verse" / "book" / "magisterium" / "ccc" / "canon" + / "concept" / "person" / "institution" / "rel" + / "mr" ; [NEW — REV-03] +``` + +## D.3 Domain Path Productions + +``` +; ── Scripture ──────────────────────────────────────────────────── +domain-path =/ verse-path / book-path +book-path = osis-book +verse-path = osis-book SLASH chapter [SLASH verse-ref] +osis-book = slug ; per OSIS 2.1.1 abbreviations +chapter = 1*DIGIT +verse-ref = 1*DIGIT [HYPHEN 1*DIGIT] ; single verse or range + +; ── Magisterium ────────────────────────────────────────────────── +domain-path =/ magisterium-path +magisterium-path = issuer-slug SLASH doc-slug [SLASH section-type SLASH section-id] +issuer-slug = slug ; council, dicastery, or pope-{name}-{roman} +doc-slug = slug +section-type = "chapter" / "canon" / "paragraph" / "article" +section-id = 1*DIGIT + +; ── Catechism ──────────────────────────────────────────────────── +domain-path =/ ccc-path +ccc-path = para-num [HYPHEN para-num] +para-num = 1*DIGIT + +; ── Canon Law ──────────────────────────────────────────────────── +domain-path =/ canon-path +canon-path = code SLASH canon-num [SLASH canon-qualifier] +code = "cic1983" / "cic1917" / "cceo" +canon-num = 1*DIGIT +canon-qualifier = para-num ; paragraph number — always numeric + / "authentic-interpretation" SLASH iso-date + ; fixed keyword + ISO date — unambiguous + +; ── Theological Concepts ──────────────────────────────────────── [REVISED — REV-03] +domain-path =/ concept-path +concept-path = slug / opaque-id +opaque-id = "C" 1*7DIGIT + ; e.g. C0000418. Slug form (legacy) and opaque form (new, RECOMMENDED + ; for new entries) are both valid; see §4.5 and §3.6 for minting guidance. + +; ── Persons ────────────────────────────────────────────────────── +domain-path =/ person-path +person-path = person-slug [SLASH work-slug [SLASH 1*path-segment]] +person-slug = slug +work-slug = slug +path-segment = slug ; one or more sub-division segments +; NOTE: 1*path-segment allows arbitrary depth for works with 3+ structural levels +; e.g., summa-theologiae/iii/q75/a1 = work / part / question / article + +; ── Institutions ───────────────────────────────────────────────── [REVISED — REV-03] +domain-path =/ institution-path +institution-path = supertype SLASH slug + +supertype = "circumscription" / "institute" / "family" + / "conference" / "council" / "dicastery" + / "seminary" / "eastern-church" + / "order" ; DEPRECATED — resolvable for backwards compatibility only; + ; MUST NOT be used for new identifiers (see §4.7) + +; ── Relationship Types ─────────────────────────────────────────── [REVISED — REV-03] +domain-path =/ rel-path +rel-path = rel-slug +rel-slug = "magisterial-proof-text" / "patristic-interpretation" + / "liturgical-use" / "theological-argument" + / "typological-sense" / "anagogical-sense" + / "defines-concept" / "supersedes" / "interprets" + / "contradicted-by" + / "exact-match" / "close-match" ; [NEW — REV-03] + ; cdcf:rel/exact-match → skos:exactMatch (strict identity; use sparingly, §3.6.1) + ; cdcf:rel/close-match → skos:closeMatch (overlapping, non-identical referent) +``` + +## D.4 Grammar Validation Notes + +- The canon-qualifier production makes `cdcf:canon/cic1983/844/1` (paragraph) and `cdcf:canon/cic1983/230/authentic-interpretation/1994-11-11` unambiguous: a paragraph qualifier is 1*DIGIT; an authentic interpretation begins with the fixed keyword "authentic-interpretation". No heuristic parsing is required. +- The person-path `1*path-segment` production allows Thomas Aquinas's Summa (three sub-levels: part/question/article) and other scholastic works of arbitrary structural depth, fixing the broken four-segment grammar of version 0.2.0. +- All slug productions are restricted to lower-case ASCII (ALPHA = %x61-7A). Upper-case characters are syntactically invalid in cdcf: identifiers. +- The iso-date production in canon-qualifier enforces ISO 8601 date format for authentic interpretations, ensuring consistent machine parsing across implementations. +- **[NEW — REV-03]** The concept-path opaque-id production and the institution-path supertype production are both additive: every identifier valid under 0.3.0's grammar remains valid under 0.4.0's. No published 0.3.0 identifier is invalidated by this revision, satisfying the backwards-compatibility principle of §1.3. +- **[NEW — REV-03]** The mr-path production (§D.5) is the first ABNF production for a sibling-registry domain under §7.5, and the reference for how future sibling registries should formalize their own grammars via patch release rather than prose description. + +## D.5 Martyrology (cdcf:mr/) [NEW — REV-03, formalizes the `mr:` domain per §7.5] + +``` +domain-path =/ mr-path + +mr-path = mr-owner-path / mr-universal-path + +mr-universal-path = mmdd-date HYPHEN slug + ; e.g. mr:0731-ignatius-de-loyola + +mr-owner-path = owner-registry SLASH owner-key COLON mmdd-date HYPHEN slug + ; e.g. mr:cecdr/us-boston:0705-{slug} + ; e.g. mr:ciclsaldr/ofm:0824-{slug} + +owner-registry = "cecdr" / "ciclsaldr" / iso3166-alpha2 + ; a bare ISO 3166-1 alpha-2 code denotes a national proprium + ; (no owner-key) + +owner-key = slug ; MUST equal the referenced entity's own registry slug + ; (the circ: or icl: slug, without that registry's prefix) + +mmdd-date = 2DIGIT 2DIGIT + ; month, day; "0229" permitted (leap-day identity decisions are + ; per crmedr/docs/canonicalization-report.md) + +iso3166-alpha2 = 2ALPHA ; per ISO 3166-1 alpha-2 +``` + +This production is normative for `mr:` as a `cdcf:`-recognized domain as of 0.4.0. CRMEDR (once it declares `compatibleWith: "0.4.0"`) is the reference implementation. + +# Appendix E. Change Rationale (Version 0.3.0) + +## E.1 ABNF Grammar Added (Appendix D) + +Version 0.2.0 described path patterns in prose only. Prose descriptions are insufficient for implementation: different developers reading the same prose will produce different parsers, leading to identifier strings that are valid according to one implementation and invalid according to another. The ABNF grammar in Appendix D is the definitive machine-readable specification. It also resolves the canon-qualifier ambiguity (paragraph number vs. authentic-interpretation keyword) and the person path depth problem, both of which were defects in version 0.2.0. + +## E.2 cdcf:rel/ Resolution Defined (§4.8.1) + +Version 0.2.0 introduced cdcf:rel/ identifiers but did not specify what dereferencing them returns. This made them inconsistent with the dereferenceability principle in §1.3. Version 0.3.0 defines them as OWL ObjectProperty declarations, with domain and range typed to the appropriate cdcf: entity classes. This allows ontology systems to import the relationship vocabulary directly. + +## E.3 Scalar licenseStatus Replaced with licenses Object (§3.5) + +A single scalar field cannot represent the copyright status of composite resources, where a Latin text may be public domain but its English translation is proprietary. The licenses object maps each content field independently, with a wildcard (*) for fields not explicitly listed. This is backward-compatible in principle: a resource where all fields share the same status uses only the wildcard key. + +## E.4 Response Depth Added (§5.6) + +Without a defined response depth mechanism, implementations will diverge on how much related data to inline. This creates incompatibility between clients and servers built from the same specification. The Depth header follows the established pattern of RFC 7240 and provides the minimum necessary control (summary vs. full) without requiring a full query language at this stage. + +## E.5 @context Specification Added (§5.3) + +The JSON-LD @context URL is a load-bearing dependency: without a reachable and stable context document, JSON-LD responses become uninterpretable. Version 0.3.0 specifies the context document's requirements, its immutability guarantee, its caching headers, and the requirement that it be publishable as a bundleable static file. This eliminates the runtime dependency for implementations that bundle the context locally. + +## E.6 Error Response Schema Added (§5.5) + +Status codes alone do not give consuming applications enough information to handle errors gracefully. A 404 for an unassigned identifier is different from a 404 for a malformed path; a 301 with no body requires an additional GET to discover the successor. The error schema defined in §5.5 is minimal but sufficient for all defined error conditions. + +## E.7 Caching and Resilience Specification Added (§5.7, §5.8) + +The resolution server is the single point of failure for the entire infrastructure. Without defined caching headers, consuming applications will generate unnecessary traffic and be vulnerable to outages. The one-year max-age for stable identifiers means the vast majority of production traffic will be served from cache. The mirror discovery endpoint and Atom change feed are designed to be implementable with static file hosting and a simple feed generator, not requiring complex infrastructure. + +## E.8 Issuer Slug Unified (§4.2) + +Version 0.2.0 used different slug conventions for the same pope in the magisterium domain (leo13) versus the person domain (pope-leo-xiii). A consuming application could not link a document to its author without an external mapping table. Version 0.3.0 unifies to pope-{name}-{roman} across both domains. The magisterialIssuedBy field in magisterial document responses carries the canonical cdcf:person/ identifier, making the link machine-derivable. + +## E.9 Duplicate documentType Row Removed (§4.2.1) + +The version 0.2.0 documentType table contained two rows with the value dogmatic-constitution. A controlled vocabulary with duplicate keys is invalid for use in schema validators. The duplicate was a copy-paste error; it has been removed. The correct entry (Constitutio dogmatica — ecumenical council constitution) is retained. + +## E.10 Sub-Specification Versioning Defined (§7.5) + +Version 0.2.0 referenced future sub-specifications without defining how they would be named, versioned, or related to the main specification. This risked fragmentation. Section 7.5 defines the naming convention (draft-cdcf-{topic}-{nn}), the compatibleWith field, and the requirement that grammar-modifying sub-specifications trigger a patch release of the main document. + +# Appendix F. Change Rationale (Version 0.4.0) [NEW — REV-03] + +## F.1 Two-Artifact Model Adopted (§3.6, §5.9) + +Version 0.3.0 minted a single string per entity, requiring it to serve simultaneously as reasoner-facing graph identity and human-facing citation. The identifier-architecture review (`draft-cdcf-identifier-rationale-00`) demonstrated these are independent requirements that mature standards (BCP 47, Unicode, Getty/VIAF) satisfy with two coordinated artifacts rather than one. Version 0.4.0 adds the `notations` field so both jobs can be served without forcing every domain to choose one style globally. + +## F.2 Concept Opacity Option Added (§4.5) + +`cdcf:concept/` was the one domain 0.3.0 left fully transparent despite its own "Limitation: Doctrinal Development" note acknowledging that concept boundaries are a live theological question, unlike the settled "things" domains. This is precisely the case the identifier rationale identifies as where opacity is defensible (cf. Gene Ontology/OBO Foundry precedent). The grammar now accepts an opaque form without breaking any existing mnemonic identifier. + +## F.3 Institution Supertypes Replace Closed Enum (§4.7, Appendix D §D.4) + +The 0.3.0 `inst-type` enumeration could not represent CECDR's actual scope (eparchies, exarchates, territorial prelatures, apostolic vicariates, ordinariates, personal prelatures, missions sui iuris) or CICLSALDR's institute/family distinction, both of which are load-bearing for Praenotanda n. 38 proprium ownership. Rather than re-enumerating an ever-growing list of canonical forms in ABNF, the path grammar now carries a small set of stable supertypes, with canon-law-specific detail moved to a response field (`circumscriptionType`) that can grow without a further grammar patch. + +## F.4 Sibling Registry Cross-References Formalized (§4.9, Appendix D §D.5) + +CRMEDR, CECDR, and CICLSALDR had each minted an independent top-level prefix (`mr:`, `circ:`, `icl:`) with no declared relationship to `cdcf:`, in tension with §7.5's requirement that any sub-specification introducing a new domain be accompanied by a patch release of the main specification. Version 0.4.0 brings all three into compliance: `circ:` and `icl:` are formalized as notation schemes attached to `cdcf:institution/circumscription/` and `cdcf:institution/institute|family/` respectively (§4.9), and `mr:` is formalized as a `cdcf:` domain in its own right (Appendix D §D.5), including ABNF for the previously prose-only owner-segment composite form. + +## F.5 `cdcf:rel/exact-match` and `close-match` Added, in Place of Blanket `owl:sameAs` (§4.8.2, §4.9) + +The identifier rationale's own recommendation used `owl:sameAs` and `skos:exactMatch` somewhat interchangeably. `owl:sameAs` asserts unqualified logical identity and causes reasoners to merge all properties of both resources — a known hazard when the linked resources are close but not identical (e.g. a GeoNames territorial point standing in for a diocese-as-institution). Version 0.4.0 introduces two relationship types mapped respectively to `skos:exactMatch` and `skos:closeMatch`, and does not map anything to `owl:sameAs`; a narrower future sub-specification may reintroduce it under conditions where full identity is actually warranted. + +*— End of Draft 0.4.0 —* diff --git a/standards/drafts/cdcf-identifier-rationale-00.md b/standards/drafts/cdcf-identifier-rationale-00.md new file mode 100644 index 0000000..4e900f4 --- /dev/null +++ b/standards/drafts/cdcf-identifier-rationale-00.md @@ -0,0 +1,165 @@ +# Ontology IRIs and Canonical Identifiers +## A Design Rationale for CDCF Datasets + +| Field | Value | +|---|---| +| Document ID | `draft-cdcf-identifier-rationale-00` | +| Status | Committee discussion paper | +| Purpose | Align the committee on *why* CDCF mints both ontology IRIs and canonical IDs, and *when* each should be opaque or transparent | +| Relates to | `draft-cdcf-catholic-uri-scheme-02`; `draft-cdcf-liturgical-events-00` | +| License | CC-BY 4.0 | + +--- + +## 1. Purpose + +A live disagreement in the committee runs roughly: *opaque, flat, randomly-generated IRIs are sufficient and would end all argument about how to form identifiers.* A competing position holds that mnemonic identifiers (`en-US`, not `uni:/x/Q7f2a9`) are the right middle ground between human and machine consumption. + +This paper argues that the disagreement is real but has been conducted on the wrong axis. "Opaque vs mnemonic" is not one decision; it is several independent decisions that have been collapsed into one. Once separated, most of the apparent conflict dissolves, and the residue resolves cleanly along the distinction between **things** (Popes, published Missals, canonically-erected dioceses) and **concepts** (theological notions whose boundaries develop). The recommendation is not to choose opaque *or* transparent globally, but to mint **both an ontology IRI and a canonical ID per entity**, and to decide transparency **per dataset** using criteria this paper sets out. + +--- + +## 2. The Debate Conflates Four Orthogonal Axes + +An identifier design involves at least four independent choices. They are genuinely independent: fixing one does not fix the others. + +| Axis | Question | Poles | Governs | +|---|---|---|---| +| **A — Grammar** | Is the string's shape formally specified? | ABNF-validated ↔ ad hoc | Well-formedness checking | +| **B — Transparency** | Does the string carry human-legible meaning? | Transparent/mnemonic ↔ opaque | Human & tooling ergonomics | +| **C — Structure** | Does the string encode hierarchy? | Hierarchical/path ↔ flat | Namespacing, enumeration | +| **D — Consumption** | Do machines parse the string to recover meaning? | Parsed ↔ opaque-to-reasoners | Coupling, correctness | + +The proposal that identifiers be flat, opaque, and randomly generated is a specific point in this space: opaque on B, flat on C, and — implicitly — it assumes that opacity on B is *required* to get correctness on D. That last inference is the crux, and it is mistaken. The opaque-identifier analysis raised in the discussion correctly observes that a UUID is *both* flat (C) and opaque (B), and that the two are independent. It is right about that independence, and the same independence extends to all four axes: + +- **A is orthogonal to everything.** `id = "Q" 1*DIGIT` is a perfectly good ABNF grammar for opaque Wikidata-style IDs. Adopting ABNF does not commit you to transparency, and choosing opacity does not free you from needing a grammar. **You want a grammar either way.** So the ABNF work in the URI-scheme draft is not in tension with the opaque-ID proposal at all — it applies to it. +- **B is orthogonal to D.** This is the decisive point. In RDF/OWL an IRI is a *rigid designator*: a reasoner must treat it as an opaque atom and must not parse it to recover meaning. That is a real and correct discipline. But "reasoners treat IRIs as opaque" (a rule about D, *consumption*) does not entail "IRIs must be randomly generated" (a rule about B, *minting*). `en-US` is a rigid designator that a reasoner treats as opaque *and* a human reads at a glance. The two facts coexist without contradiction. Transparency is an affordance for the humans and tooling that *handle* the string; opacity-to-reasoners is a discipline for the software that *reasons over* it. **You get both at once.** + +Those who hold that the identifiers should be opaque are therefore right about D — and are using that correctness to argue for B, where it does not reach. + +--- + +## 3. Two Artifacts, Two Jobs — and Why the "Opaque Is Best Practice" Advice Answers a Different Question + +CDCF is not minting one identifier per entity. It is minting two artifacts that do two different jobs: + +1. **The ontology IRI** — the entity's identity *inside the graph*, consumed by reasoners. Its governing requirement is stability and opacity-to-reasoners (axis D). Whether its characters are legible (axis B) is, to the reasoner, irrelevant. +2. **The canonical ID** — the string that appears *in the wild*: in a footnote, a citation, a URL, an MCP tool argument, a content author's markup, a config file. Its governing requirement is that a human and ordinary tooling can read, write, quote, and verify it without a lookup round-trip. + +The circulated case for opacity is sound — **for the job it is describing, which is application-internal database keys.** Its three rationales are worth taking at face value and then locating precisely: + +- *"Security & privacy: users cannot enumerate or guess other IDs."* This is a virtue for private application data (customer records, session tokens). For a **public reference standard it is inverted**: CDCF *wants* its identifiers to be discoverable, guessable, and enumerable. `cdcf:verse/jn/3/16` being predictable from `cdcf:verse/jn/3/17` is the entire point of a citation scheme. Un-guessability is an anti-feature here. +- *"Decoupling: an ID with `NY` baked in breaks when the customer moves to `CA`."* This is the drift argument, and it is real (§7). But it is a claim about volatile *business state*, not about identity. A Pope who has reigned does not "move to CA." +- *"Database flexibility: flat opaque keys let you reorganize storage without migrating IDs."* A storage-layer concern. CDCF's canonical IDs are not storage keys; they are public citations whose whole value is that they *don't* change under reorganization. + +So that case is not wrong; it is answering "how should I key rows in my application's database?" — where flat + opaque is indeed best practice. CDCF's question is "how should the world cite a Catholic entity for the next century?" These have different, in places opposite, force profiles. Conflating them is the single most common error in identifier debates. + +Crucially, the two artifacts are **not a dilemma**. The Semantic Web already provides the vocabulary to carry both on one entity (§8): an opaque-if-you-like IRI for identity, plus a transparent canonical ID attached as a typed `skos:notation`, plus `owl:sameAs` links outward. Choosing one does not cost you the other. + +--- + +## 4. What the Field Actually Does + +The abstract argument can run forever. The empirical record is clearer: mature ontologies and standards have already made these choices at scale, and they did **not** converge on one answer. They converged on a *rule* for choosing. + +### 4.1 Real ontologies + +| System | IRI style | Why | +|---|---|---| +| **Wikidata** | Opaque (`wd:Q42` = Douglas Adams; `wdt:P31` = "instance of") | A queried database with a search UI; IDs are supplied by tooling, never hand-typed. Its own users pay for this daily: raw SPARQL/dumps are unreadable without constant label lookups. | +| **Gene Ontology / OBO Foundry** | Opaque numeric (`GO:0008150` = biological_process) | **Deliberate policy.** Biological categories are reclassified, merged, and split as science advances; semantics-free IDs immunize the identifier against that churn. | +| **Getty (TGN/ULAN/AAT), VIAF, GeoNames** | Opaque numeric with rich labels | Authority files behind resolvers; used for reconciliation, not citation. | +| **Schema.org** | Transparent (`schema:Person`, `schema:birthDate`) | An *authoring* vocabulary embedded by hand in web-page markup — the same surface CDCF's citations live on. | +| **FOAF, Dublin Core, SKOS** | Transparent (`foaf:knows`, `dc:creator`, `skos:broader`) | Human-written interchange vocabularies. | +| **DBpedia** | Transparent, derived (`dbr:Pope_Francis`, from the article title) | Legible, but *drift-prone*: article renames break the mnemonic — a live illustration of the risk (§7). | + +The pattern is not "the sophisticated people chose opacity." It is: **opacity clusters where a resolver always mediates and where referents are fluid; transparency clusters where the string is authored and quoted by hand and referents are fixed.** + +### 4.2 Real canonical-ID standards + +| Standard | Style | Notes | +|---|---|---| +| **BCP 47 language tags** | Transparent (`en-US`, `zh-Hant-TW`) | Composed by ABNF from ISO 639/15924/3166 + UN M.49. The i18n stack — HTML `lang`, HTTP `Accept-Language`, CLDR — runs on these *because they are authored by hand*. Nobody writes `lang="Q1860"`. | +| **Unicode / ISO 10646** | *Both, layered* | The code point `U+0041` is opaque; the character name `LATIN CAPITAL LETTER A` is transparent **and frozen forever** by Unicode's name-stability policy. Even Unicode does not put semantics *in the number* — it maps the number to an immutable descriptive name in a registry. | +| **OSIS book codes** | Transparent (`Gen`, `Jn`, `Rev`) | Already reused by the URI-scheme draft — the right instinct. | +| **Canon numbers, CCC paragraphs, Denzinger (DS)** | Transparent-numeric | The number *is* the authoritative, centuries-stable citation. Discarding it for a UUID would be perverse. | +| **DOI, ORCID** | Opaque behind a resolver | Work precisely because a resolver is *always* interposed; the raw DOI is not meant to be read. | + +The observation that settles the `uni:/x/Q7f2a9` vs `en-US` example: **the entire internationalization ecosystem already chose `en-US`, in exactly CDCF's use case (hand-authored, quoted-in-the-wild identifiers), and has run on it for two decades.** The opaque-ID proposal is not "align with best practice"; it is "diverge from the closest and most successful precedent CDCF has." + +--- + +## 5. The Unicode/BCP-47 Model, Precisely — CDCF's Template + +The appeal to the Unicode Consortium points at the right model; it is worth stating its architecture precisely, because the precision *is* the design CDCF should copy. It is not one thing — it is a stack, and each layer answers one of our axes: + +1. **Code lists (vocabulary).** ISO 639 (language), ISO 15924 (script), ISO 3166 (region) — flat registries of atomic, descriptive-yet-stable codes. *Axis B: transparent. Axis C: flat.* +2. **A grammar (syntax).** BCP 47 (RFC 5646) uses **ABNF** to compose those atoms into a tag. *Axis A.* +3. **A stability keeper (governance).** The IANA Language Subtag Registry snapshots the ISO code lists and adds `Added`, `Deprecated`, and `Preferred-Value` fields — the machinery that makes canonicalization declarative. +4. **A matching layer.** RFC 4647 defines fallback (`zh-Hant-TW` → `zh-Hant` → `zh`). + +Note what this buys: the tag `zh-Hant-TW` is simultaneously **transparent** (a human reads it), **grammar-validated** (ABNF), **composed of stable registry atoms**, and **opaque to a matcher** (which treats it as a token to truncate, not a sentence to parse). All four axes, resolved independently, in one identifier. This is the existence proof that the committee's dilemma is false. CDCF's URI-scheme draft and CLEDR strawman already replicate layers 1–4; this paper's only addition is to name the model explicitly so the committee stops treating "transparent" and "machine-consumable" as opposed. + +--- + +## 6. The Real Decision Criterion: Things vs Concepts + +The strongest contribution to this debate has been the distinction between **things** and **concepts**, and it turns out to be *the* criterion the field is implicitly using. + +- **Things** — entities with fixed extension and existing authoritative naming: Popes who have reigned, published editions of the Roman Missal, canonically-erected Latin-rite dioceses, the books of Scripture, promulgated magisterial documents. The Church has already done the canonical-naming work, often over centuries. Here **transparent identifiers are correct**, and the real-world evidence agrees: BCP 47, OSIS, ISO 3166, canon/CCC/DS numbering all chose transparent for well-defined things. +- **Concepts** — entities whose boundaries are genuinely contestable or develop over time: some theological notions, categories whose articulation shifts under doctrinal development. Here the drift and individuation risks are highest, and **opacity is defensible**. The evidence again agrees: the Gene Ontology and OBO Foundry — enormous, mature, sophisticated — chose *opaque numeric IDs precisely because their domain is nothing but concepts whose boundaries move.* + +This is the reconciliation. The biomedical ontologists and the BCP 47 editors did not disagree; **they were identifying different kinds of referent.** Biology is almost entirely concepts, so it went opaque. Language, geography, and Scripture are things, so they went transparent. CDCF spans both kinds, so it should apply both rules — per dataset. + +A necessary honesty check on "things": even things have a finite fringe of genuine edge cases. Papal numbering carries historical wrinkles (antipopes; the Stephen II/III ambiguity; the skipped John XX). Dioceses are erected, renamed, merged, and suppressed. But these are **finite, enumerable, and already adjudicated by the Church's own historical record** — a bounded list of known cases with authoritative answers — unlike a concept whose boundary is a *live* theological question. A small, closed set of documented exceptions is exactly what a registry with a deprecation policy is built to hold. It is not an argument for opacity. + +--- + +## 7. The Drift Objection, and How Standards Actually Answer It + +The one genuinely strong argument for opacity is **semantic drift**: a transparent identifier is a small promise about the world, and history can falsify it. Rename a diocese and `cdcf:institution/diocese/old-name` mildly lies. Opaque IDs never lie because they never asserted anything. This must be conceded plainly. + +But it is answered — no standardized dataset is fossilized; they are revised — and the standards show *how* to revise without opacity. The instructive case is ISO 3166, because it contains both the failure and the fix. + +ISO 3166's discipline: when a country's code is withdrawn, the two-letter code is **held in transitional reservation** (at least five years, often longer for the three-letter form), the retired code is **archived in ISO 3166-3** with a four-letter successor code recording what it became (Burma `BU` → Myanmar `MM`, archived as `BUMM`), and the numeric code is not reassigned casually. This is precisely the URI-scheme draft's §3.4 contract — never reassign, deprecate rather than delete, record the successor — and it lets descriptive codes survive geopolitical upheaval. + +The cautionary half is just as useful. `CS` was used for Czechoslovakia, then **reused** for Serbia and Montenegro after Yugoslavia was renamed. That single violation of the never-reassign rule caused lasting confusion — even ISO's own archival code for Serbia and Montenegro had to be changed from `CSHH` to `CSXX` to avoid colliding with Czechoslovakia. The lesson is not "descriptive codes are dangerous." It is **"reassignment is dangerous"** — and reassignment is a governance failure, available to opaque and transparent schemes alike. Drift is not solved by making identifiers meaningless; it is solved by never reusing them and always recording their succession. CDCF's spec already mandates exactly this. + +In short: the drift objection is real, it is a maintenance burden transparency carries and opacity avoids, and it is nonetheless the *lesser* cost for a citation layer — because the standards prove it is a **solved** problem, while opacity's cost (a mandatory lookup for every human who ever reads the identifier) is permanent and unsolvable by design. + +--- + +## 8. Recommendation: Coexistence via SKOS, Decided Per Dataset + +CDCF should stop framing this as a choice and mint, for each entity, the following — the same shape the Getty and Library of Congress vocabularies use to satisfy both machine and human consumers at once: + +- an **ontology IRI** — the rigid designator for the graph (may be opaque *or* transparent, per §6); +- a `skos:prefLabel` — the human display name, language-tagged; +- a `skos:notation` **carrying the ABNF-governed canonical ID**, datatyped to a declared CDCF scheme (this is the transparent citation string — and note that in RDF the canonical ID most naturally lives *as a typed notation on the concept*, not necessarily as the raw IRI); +- `owl:sameAs` / `skos:exactMatch` links to **Wikidata, VIAF, GeoNames** for cross-walk stability. + +This gives the opaque-ID position what it correctly asks for, exactly where it earns its keep (an ontology-internal identity that no reasoner parses, plus external cross-references that never lie) and gives adopters legibility exactly where it earns its keep (the citation and interchange layer), with ABNF validating that layer — and it is the standard, boring, widely-deployed pattern, not a novel compromise. + +### Per-dataset recommendation + +| Dataset | Kind | Canonical ID (notation) | Notes | +|---|---|---|---| +| Bible books / verses | Thing | **Transparent** (OSIS) | Already decided; authoritative external scheme reused. | +| Roman Missal editions (CRMETDR) | Thing | **Transparent** | Published artifacts with fixed identity. | +| Popes (pontiffs DR) | Thing | **Transparent** (`pope-{name}-{roman}`) | Fringe cases (antipopes, numbering) are finite and adjudicated → registry entries, not a reason for opacity. | +| Latin-rite dioceses | Thing | **Transparent**, with strict deprecation | The dataset that changes most; ISO 3166's never-reassign discipline is the model. | +| Canon law / CCC | Thing | **Transparent-numeric** | The number is the authoritative citation. | +| Liturgical celebrations (CLEDR) | Thing (mostly) | **Transparent + compositional** | See CLEDR strawman; commemoration atoms are things. | +| Magisterial documents | Thing | **Transparent** | Promulgated artifacts. | +| Theological concepts (`concept/`) | **Concept** | **Opaque IRI + transparent notation** *defensible* | The one domain where the case for opacity has real force; individuation is a live question. Consider opaque primary IRI with a mnemonic `skos:notation` and a `Preferred-Value` discipline for boundary revisions. | + +--- + +## 9. Bottom Line for the Committee + +- **ABNF is not the thing under dispute.** It applies whether identifiers are opaque or transparent, and CDCF needs it either way. +- **"Reasoners treat IRIs as opaque" is true and does not imply "mint random IRIs."** Those are different axes. +- **Opacity does not "prevent disagreement"; it relocates it** — from visible slug choices to invisible individuation and minting-policy choices, which are the ones that actually matter for authoritative Catholic data and which governance, not string format, must settle. +- **The opaque-ID best practices commonly cited are correct for application databases and mis-aimed at a public citation standard,** whose requirements (discoverability, quotability, legibility, century-scale stability) are in places the opposite. +- **The field already chose the answer for our use case:** the entire i18n ecosystem cites `en-US`, not an opaque code, because those identifiers are authored and quoted by hand — exactly like CDCF's. +- **The right rule is the things/concepts rule,** empirically confirmed by OBO (concepts → opaque) versus BCP 47/OSIS (things → transparent). CDCF spans both, so it applies both — per dataset, per the table in §8 — and carries both an IRI and a canonical ID on every entity via SKOS, rather than choosing between them. From 348a4e100aed8b2a9c3d7130ff94e409f117bea3 Mon Sep 17 00:00:00 2001 From: damienriehl Date: Sun, 26 Jul 2026 13:10:50 -0500 Subject: [PATCH 2/6] docs: register identifier-durability paper in build surfaces; exclude verbatim drafts from lint Co-Authored-By: Claude Fable 5 Claude-Session: https://claude.ai/code/session_01MoZuH8vKiwL2rsX3a1zSqc --- .github/workflows/deploy-docs.yml | 2 ++ README.md | 11 ++++++----- RELEASING.md | 15 ++++++++++++--- package.json | 6 +++--- scripts/build-combined-pdf.sh | 1 + scripts/build-standalone-html.sh | 1 + 6 files changed, 25 insertions(+), 11 deletions(-) diff --git a/.github/workflows/deploy-docs.yml b/.github/workflows/deploy-docs.yml index 68d2a48..70d34c9 100644 --- a/.github/workflows/deploy-docs.yml +++ b/.github/workflows/deploy-docs.yml @@ -28,6 +28,7 @@ jobs: "research/fragmented-catholic-digital-governance" "research/governance-as-code-catholic-technology" "research/trusted-data-infrastructure-catholic-ministry" + "research/identifier-durability-opaque-canonical-iris" "project-governance/definitions" "project-governance/project-types" "project-governance/lifecycle" @@ -106,6 +107,7 @@ jobs: ["research/fragmented-catholic-digital-governance"]="Fragmented Catholic Digital Governance at Scale" ["research/governance-as-code-catholic-technology"]="Governance-as-Code for Catholic Technology Deployment" ["research/trusted-data-infrastructure-catholic-ministry"]="Trusted Data Infrastructure for Catholic Ministry" + ["research/identifier-durability-opaque-canonical-iris"]="Identifier Durability and Opaque Canonical IRIs" ["project-governance/definitions"]="CDCF Governance Definitions" ["project-governance/project-types"]="CDCF Project Types: Foundation Projects and Community Projects" ["project-governance/lifecycle"]="CDCF Project Lifecycle" diff --git a/README.md b/README.md index 3352ac1..10024c4 100644 --- a/README.md +++ b/README.md @@ -49,11 +49,12 @@ The core frameworks for any project seeking CDCF endorsement. Supplementary research memos informing the design of the vetting criteria. -| Document | Type | Description | -| :-------------------------------------------------------------------------------------------------------------- | :------------ | :------------------------------------------------------ | -| [fragmented-catholic-digital-governance.md](./research/fragmented-catholic-digital-governance.md) | Research memo | The urgency of shared digital governance standards. | -| [governance-as-code-catholic-technology.md](./research/governance-as-code-catholic-technology.md) | Research memo | Machine-enforceable deployment governance architecture. | -| [trusted-data-infrastructure-catholic-ministry.md](./research/trusted-data-infrastructure-catholic-ministry.md) | Research memo | Trusted data infrastructure for Catholic ministry. | +| Document | Type | Description | +| :-------------------------------------------------------------------------------------------------------------- | :------------ | :------------------------------------------------------------------------------------------------------------------- | +| [fragmented-catholic-digital-governance.md](./research/fragmented-catholic-digital-governance.md) | Research memo | The urgency of shared digital governance standards. | +| [governance-as-code-catholic-technology.md](./research/governance-as-code-catholic-technology.md) | Research memo | Machine-enforceable deployment governance architecture. | +| [trusted-data-infrastructure-catholic-ministry.md](./research/trusted-data-infrastructure-catholic-ministry.md) | Research memo | Trusted data infrastructure for Catholic ministry. | +| [identifier-durability-opaque-canonical-iris.md](./research/identifier-durability-opaque-canonical-iris.md) | Research memo | Position paper: opaque canonical IRIs with human-readable affordances — public comment on the CDCF URI-scheme draft. | --- diff --git a/RELEASING.md b/RELEASING.md index 86e0437..55582b1 100644 --- a/RELEASING.md +++ b/RELEASING.md @@ -137,11 +137,20 @@ The following documents are included in builds and deployments: | Research | `research/fragmented-catholic-digital-governance.md` | | Research | `research/governance-as-code-catholic-technology.md` | | Research | `research/trusted-data-infrastructure-catholic-ministry.md` | +| Research | `research/identifier-durability-opaque-canonical-iris.md` | | Standards | `standards/overview.md` | | Standards | `standards/committees.md` | To add a new document, update the `DOCS` array in all three places: -1. `.github/workflows/deploy-docs.yml` (lines 27-38 and 87-98) -2. `scripts/build-standalone-html.sh` (lines 10-21) -3. `scripts/build-combined-pdf.sh` (lines 12-24) +1. `.github/workflows/deploy-docs.yml` (lines 27-39 and 106-118) +2. `scripts/build-standalone-html.sh` (lines 10-22) +3. `scripts/build-combined-pdf.sh` (lines 10-25) + +### Excluded: `standards/drafts/` + +`standards/drafts/` holds external committee documents reproduced **verbatim** +(e.g. CDCF Catholic URI Scheme drafts). Their text must not be edited or +reformatted, so they are excluded from `npm run lint:md`, `npm run lint:md:fix`, +and `npm run format:md`, and they are deliberately absent from the `DOCS` arrays +above -- they are neither deployed to WordPress nor included in build artifacts. diff --git a/package.json b/package.json index 0bf5e63..7b98a31 100644 --- a/package.json +++ b/package.json @@ -3,9 +3,9 @@ "private": true, "description": "Governance documentation for the Catholic Digital Commons Foundation", "scripts": { - "lint:md": "markdownlint-cli2 \"**/*.md\" \"#**/node_modules\" \"#.yarn\"", - "lint:md:fix": "markdownlint-cli2 --fix \"**/*.md\" \"#**/node_modules\" \"#.yarn\"", - "format:md": "prettier --write \"**/*.md\" \"!**/node_modules/**\"", + "lint:md": "markdownlint-cli2 \"**/*.md\" \"#**/node_modules\" \"#.yarn\" \"#standards/drafts\"", + "lint:md:fix": "markdownlint-cli2 --fix \"**/*.md\" \"#**/node_modules\" \"#.yarn\" \"#standards/drafts\"", + "format:md": "prettier --write \"**/*.md\" \"!**/node_modules/**\" \"!standards/drafts/**\"", "build:html": "bash scripts/build-standalone-html.sh", "build:pdf": "bash scripts/build-combined-pdf.sh", "prepare": "husky" diff --git a/scripts/build-combined-pdf.sh b/scripts/build-combined-pdf.sh index f344e0f..b339639 100755 --- a/scripts/build-combined-pdf.sh +++ b/scripts/build-combined-pdf.sh @@ -18,6 +18,7 @@ DOCS=( "research/fragmented-catholic-digital-governance.md" "research/governance-as-code-catholic-technology.md" "research/trusted-data-infrastructure-catholic-ministry.md" + "research/identifier-durability-opaque-canonical-iris.md" # Standards "standards/overview.md" "standards/committees.md" diff --git a/scripts/build-standalone-html.sh b/scripts/build-standalone-html.sh index cb12163..5d8c0e5 100755 --- a/scripts/build-standalone-html.sh +++ b/scripts/build-standalone-html.sh @@ -16,6 +16,7 @@ DOCS=( "research/fragmented-catholic-digital-governance.md|fragmented-catholic-digital-governance|Fragmented Catholic Digital Governance" "research/governance-as-code-catholic-technology.md|governance-as-code-catholic-technology|Governance-as-Code for Catholic Technology" "research/trusted-data-infrastructure-catholic-ministry.md|trusted-data-infrastructure-catholic-ministry|Trusted Data Infrastructure for Catholic Ministry" + "research/identifier-durability-opaque-canonical-iris.md|identifier-durability-opaque-canonical-iris|Identifier Durability and Opaque Canonical IRIs" "standards/overview.md|standards-overview|CDCF Standards Overview" "standards/committees.md|standards-committees|CDCF Standards Committees" ) From 35c8374ed3673893db43f34598945c55f370db9c Mon Sep 17 00:00:00 2001 From: damienriehl Date: Sun, 26 Jul 2026 14:24:32 -0500 Subject: [PATCH 3/6] =?UTF-8?q?docs:=20Draft=202=20=E2=80=94=20machine-rea?= =?UTF-8?q?dable/human-readable=20terminology,=20actorless=20voice,=20exem?= =?UTF-8?q?plar-rich=20executive=20summary,=20et-et=20framing,=20objection?= =?UTF-8?q?s=20moved=20to=20disputatio=20position?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Co-Authored-By: Claude Fable 5 Claude-Session: https://claude.ai/code/session_01MoZuH8vKiwL2rsX3a1zSqc --- .github/workflows/deploy-docs.yml | 2 +- README.md | 12 +- ...tifier-durability-opaque-canonical-iris.md | 484 +++++++++++------- scripts/build-standalone-html.sh | 2 +- 4 files changed, 297 insertions(+), 203 deletions(-) diff --git a/.github/workflows/deploy-docs.yml b/.github/workflows/deploy-docs.yml index 70d34c9..0d11f23 100644 --- a/.github/workflows/deploy-docs.yml +++ b/.github/workflows/deploy-docs.yml @@ -107,7 +107,7 @@ jobs: ["research/fragmented-catholic-digital-governance"]="Fragmented Catholic Digital Governance at Scale" ["research/governance-as-code-catholic-technology"]="Governance-as-Code for Catholic Technology Deployment" ["research/trusted-data-infrastructure-catholic-ministry"]="Trusted Data Infrastructure for Catholic Ministry" - ["research/identifier-durability-opaque-canonical-iris"]="Identifier Durability and Opaque Canonical IRIs" + ["research/identifier-durability-opaque-canonical-iris"]="Identifier Durability: Machine-Readable Canonical IRIs" ["project-governance/definitions"]="CDCF Governance Definitions" ["project-governance/project-types"]="CDCF Project Types: Foundation Projects and Community Projects" ["project-governance/lifecycle"]="CDCF Project Lifecycle" diff --git a/README.md b/README.md index 10024c4..8c57328 100644 --- a/README.md +++ b/README.md @@ -49,12 +49,12 @@ The core frameworks for any project seeking CDCF endorsement. Supplementary research memos informing the design of the vetting criteria. -| Document | Type | Description | -| :-------------------------------------------------------------------------------------------------------------- | :------------ | :------------------------------------------------------------------------------------------------------------------- | -| [fragmented-catholic-digital-governance.md](./research/fragmented-catholic-digital-governance.md) | Research memo | The urgency of shared digital governance standards. | -| [governance-as-code-catholic-technology.md](./research/governance-as-code-catholic-technology.md) | Research memo | Machine-enforceable deployment governance architecture. | -| [trusted-data-infrastructure-catholic-ministry.md](./research/trusted-data-infrastructure-catholic-ministry.md) | Research memo | Trusted data infrastructure for Catholic ministry. | -| [identifier-durability-opaque-canonical-iris.md](./research/identifier-durability-opaque-canonical-iris.md) | Research memo | Position paper: opaque canonical IRIs with human-readable affordances — public comment on the CDCF URI-scheme draft. | +| Document | Type | Description | +| :-------------------------------------------------------------------------------------------------------------- | :------------ | :---------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | +| [fragmented-catholic-digital-governance.md](./research/fragmented-catholic-digital-governance.md) | Research memo | The urgency of shared digital governance standards. | +| [governance-as-code-catholic-technology.md](./research/governance-as-code-catholic-technology.md) | Research memo | Machine-enforceable deployment governance architecture. | +| [trusted-data-infrastructure-catholic-ministry.md](./research/trusted-data-infrastructure-catholic-ministry.md) | Research memo | Trusted data infrastructure for Catholic ministry. | +| [identifier-durability-opaque-canonical-iris.md](./research/identifier-durability-opaque-canonical-iris.md) | Research memo | Identifier Durability: Machine-Readable Canonical IRIs — machine-readable canonical IRIs with a guaranteed human-readable layer; public comment on the CDCF URI-scheme draft. | --- diff --git a/research/identifier-durability-opaque-canonical-iris.md b/research/identifier-durability-opaque-canonical-iris.md index 56dbec3..455928e 100644 --- a/research/identifier-durability-opaque-canonical-iris.md +++ b/research/identifier-durability-opaque-canonical-iris.md @@ -1,9 +1,9 @@ -# Identifier Durability and Opaque Canonical IRIs +# Identifier Durability: Machine-Readable Canonical IRIs | | | | :---------------- | :---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | **Document type** | Position paper — public comment | -| **Status** | Draft 1 — prepared for submission during the `draft-cdcf-catholic-uri-scheme-03` (v0.4.0) 60-day comment window | +| **Status** | Draft 2 — prepared for submission during the `draft-cdcf-catholic-uri-scheme-03` (v0.4.0) 60-day comment window | | **Relationship** | Responds to [draft-cdcf-identifier-rationale-00](../standards/drafts/cdcf-identifier-rationale-00.md) and [draft-cdcf-catholic-uri-scheme-03](../standards/drafts/cdcf-catholic-uri-scheme-03.md); informs the [CDCF Standards program](../standards/overview.md) | | **License** | CC BY 4.0 | @@ -13,14 +13,14 @@ 1. [Executive Summary](#executive-summary) 2. [The Shared Premise](#the-shared-premise) -3. [What We Agree On](#what-we-agree-on) +3. [Where the Proposal Agrees with the Rationale Doc](#where-the-proposal-agrees-with-the-rationale-doc) 4. [The Evidence: Transparency Produced Divergence, Not Convergence](#the-evidence-transparency-produced-divergence-not-convergence) -5. [The Steelmen — and Where Each Fails](#the-steelmen--and-where-each-fails) -6. [Where Does the Churn Land](#where-does-the-churn-land) -7. [Worked Example: Canonicalizing Levels of Authority](#worked-example-canonicalizing-levels-of-authority) -8. [The Proposal](#the-proposal) -9. [Universal Slugs Across Standards](#universal-slugs-across-standards) -10. [Costs We Accept — and Their Remedies](#costs-we-accept--and-their-remedies) +5. [Where Does the Churn Land](#where-does-the-churn-land) +6. [Worked Example: Canonicalizing Levels of Authority](#worked-example-canonicalizing-levels-of-authority) +7. [The Proposal](#the-proposal) +8. [Universal Slugs Across Standards](#universal-slugs-across-standards) +9. [Costs the Proposal Accepts — and Their Remedies](#costs-the-proposal-accepts--and-their-remedies) +10. [Objections at Full Strength — and Their Remedies](#objections-at-full-strength--and-their-remedies) 11. [Recommended Amendments to Draft 0.4.0](#recommended-amendments-to-draft-040) 12. [Bibliography](#bibliography) @@ -28,70 +28,153 @@ ## Executive Summary -**Goals — the ground we already share.** The ambition behind CDCF is not a Catholic filing system. It is one common taxonomy and ontology every Catholic can use; then one every -Christian can use; then one the Abrahamic traditions can share wherever they genuinely refer to the same thing; and ultimately a machine-readable articulation of what those -traditions hold to be true, usable by AI systems generally as alignment-with-human-values infrastructure. That ladder is why the identifier question matters at all: a scheme that -cannot survive the second rung will never reach the fourth. - -**Problems — names and structures are exactly where traditions diverge.** The Church comprises 24 churches _sui iuris_, one Latin and 23 Eastern, and draft 0.4.0 already requires -that each "MUST be representable without privileging the Latin Church as the default" (§4.7.1). An identifier that hard-codes a Latin-rite taxonomy segment, or an English slug, or -an Italian one, quietly violates that requirement in every string it mints — and the problem sharpens as the ladder extends. To an Eastern Orthodox adopter, an identifier reading -`institution/circumscription/…` in Latin-derived English is not neutral infrastructure; it reads as _the Latin thing, not ours_. This is not speculative. Transparency has already -produced measurable divergence inside CDCF's and CatholicOS's own repositories: one verse now carries three committee-minted transparent identifiers in eight months, shipped -registry IDs have been renamed in place, and a calendar-date anchor baked into an identifier moved because the Latin and Italian editions of one book disagree about the date. - -**Solutions — one opaque spine, one guaranteed affordance layer.** We propose that every canonical identifier CatholicOS mints — ontology and data registries alike — be an opaque -base62-encoded UUID of the shape `R…`, and that _all_ human readability move into a layer the standard guarantees rather than one the identifier improvises: every existing slug and -registry key preserved as a **permanent resolvable alias**, never deprecated and never reused; **multilingual labels** on every entity, so naming disputes are settled by _adding_ a -label rather than _changing_ an identifier; and a **rendering rule** requiring production surfaces to show a label beside every canonical ID. This is deliberately a _small delta_ -to draft 0.4.0, not a rival document. Draft 0.4.0 already built the machinery — the two-artifact model (§3.6), the `notations` array (§5.9), the never-reassign stability guarantee -(§3.4), the `exact-match`/`close-match` relations (§4.8.2). We keep all of it and invert which artifact is primary. +**Goals — the ground already shared.** The ambition behind CDCF is not a Catholic filing system. + +- **A ladder, not a registry.** One common taxonomy and ontology every Catholic can use; then one every Christian can use; then one the Abrahamic traditions can share wherever + they genuinely refer to the same thing; and ultimately a machine-readable articulation of what those traditions hold to be true, usable by AI systems generally as + alignment-with-human-values infrastructure. +- **The ladder is the test.** A scheme that cannot survive the second rung will never reach the fourth, which is why the identifier question matters at all. +- **Two jobs, both required.** The standard owes its adopters a name that never moves _and_ a record every reader can interpret in their own language. Nothing below trades one + away for the other. + +**Problems — names and structures are exactly where traditions diverge.** + +- **The Church is 24 churches _sui iuris_,** one Latin and 23 Eastern, and draft 0.4.0 already requires that each "MUST be representable without privileging the Latin Church as + the default" (§4.7.1). An identifier that hard-codes a Latin-rite taxonomy segment, or an English slug, or an Italian one, quietly violates that requirement in every string it + mints — and the problem sharpens as the ladder extends. To an Eastern Orthodox adopter, an identifier reading `institution/circumscription/…` in Latin-derived English is not + neutral infrastructure; it reads as _the Latin thing, not ours_. +- **The divergence is already recorded,** inside CDCF's and CatholicOS's own repositories: one verse now carries three committee-minted transparent identifiers in eight months, + shipped registry IDs have been renamed in place, and a calendar-date anchor baked into an identifier moved because the Latin and Italian editions of one book disagree about the + date. +- **The pressure never ends.** Popes keep being elected, dioceses keep merging, decrees keep expanding memorials, editions keep being revised. Every one of those events asks a + transparent identifier to change something it promised never to change. + +**Solution — a machine-readable spine _and_ a guaranteed human-readable layer.** + +- **Canonical identifiers become machine-readable** — in the literature's term, _opaque_ — durable precisely because they assert nothing: a base62-encoded UUID of the shape `R…`, + minted once for every entity CatholicOS names, ontology and data registries alike. +- **All human readability moves into a layer the standard guarantees** rather than one the identifier improvises: every existing slug and registry key preserved as a **permanent + resolvable alias**, never deprecated and never reused; **multilingual labels** on every entity, so naming disputes are settled by _adding_ a label rather than _changing_ an + identifier; and a **rendering rule** requiring production surfaces to show a label beside every canonical ID. +- **The frame is Catholic: _et-et_, not _aut-aut_.** The question is not identifier stability _or_ human interpretability. It is both — each carried in the layer built for it, + the stable identifier canonical and the human-readable layer guaranteed rather than optional. +- **A small delta, not a rival document.** Draft 0.4.0 already provides the machinery — the two-artifact model (§3.6), the `notations` array (§5.9), the never-reassign stability + guarantee (§3.4), the `exact-match`/`close-match` relations (§4.8.2). The proposal keeps all of it and inverts which artifact is primary. + +**What it looks like in practice — four surfaces, one canonical ID.** In every case the canonical machine-readable IRI _is_ the code, and every human-readable rendering rides +beside it as a comment or a quoted label value. + +**Ontology source.** The Turtle a modeller edits — the identifier carries no language, the labels carry every language. + +```turtle +osc:Rh895FOUgLMfrylaO713w1 # "God the Son"@en · "Fílius Dei"@la · "Dio Figlio"@it + skos:prefLabel "God the Son"@en , "Fílius Dei"@la , "Dio Figlio"@it ; + skos:altLabel "the Word"@en , "Verbum"@la , "Lógos"@grc . +``` + +**Registry source file.** The row a canonist reviews — plain text, ID and label side by side, with the familiar slug preserved forever. + +```text +id: R7kQp2mXf4LdTbz9Ns3Hc1 # canonical — minted once, never re-minted +alias: mr:0629-petri-et-pauli # permanent — resolves forever, never reused +labels: "Sanctorum Petri et Pauli, Apostolorum"@la · + "Santi Pietro e Paolo, apostoli"@it · + "Sts. Peter & Paul, Apostles"@en +``` + +**Production resolution.** The response an application receives — one request, the reader's own language, every legacy key still on the record. + +```http +GET /R7kQp2mXf4LdTbz9Ns3Hc1 HTTP/1.1 +Host: id.catholiccommons.org +Accept: application/ld+json +Accept-Language: it + +HTTP/1.1 200 OK +Content-Type: application/ld+json + +{ + "@id": "R7kQp2mXf4LdTbz9Ns3Hc1", + "prefLabel": "Santi Pietro e Paolo, apostoli", + "altLabel": ["Sanctorum Petri et Pauli, Apostolorum", "Sts. Peter & Paul, Apostles"], + "notations": [ + { "value": "mr:0629-petri-et-pauli", "scheme": "cdcf:scheme/mr", "prefLabel": true }, + { "value": "StsPeterPaulAp", "scheme": "cdcf:scheme/litcal" }, + { "value": "saints_peter_and_paul_apostles", "scheme": "cdcf:scheme/romcal" } + ] +} +``` + +`notations` is 0.4.0's field as written (§5.9), carrying the litcal and romcal production keys unchanged; the language-tagged `prefLabel`/`altLabel` are the one addition the +proposal asks for (commitment 3, amendment to §5.9). + +**Editor, IDE, and AI agent.** The call site a developer or an agent actually reads — a named constant, with the label supplied by the tooling. + +```text +if (celebration === IDs.PETER_AND_PAUL) { // = "R7kQp2…" · hover: Peter and Paul, Apostles +``` + +That last surface is not aspirational: ontokit renders exactly this today, resolving a full IRI to its label in a Monaco hover and label-joining API responses beside machine-readable +`osc:` R-IDs.[^15] + +**Proposal preview — five commitments (§7).** + +- **1 — Canonical identifiers are machine-readable, `R`-shaped, and org-wide,** across the ontology and every data registry. +- **2 — Every existing slug, key, and registry prefix becomes a permanent resolvable alias,** carried as a 0.4.0 `notation`, never deprecated and never reused. +- **3 — Every entity carries multilingual labels,** `skos:prefLabel` and `skos:altLabel`, so a naming dispute is settled by adding a label. +- **4 — Production surfaces MUST render a label beside every canonical ID** — source files, serializations, UIs, and generated code alike. +- **5 — Structural and volatile facts live in properties, never in canonical identifiers** — dates, country codes, supertypes, edition years, hierarchy, and language. --- ## The Shared Premise -We begin where the committee itself began. Asked whether CDCF is meant to become the meta-level disambiguation layer for Catholic data — "kinda like what DOI does for published -URLs" — the answer was yes, with one refinement: `cdcf:` identifiers dereference to structured JSON-LD carrying typed relationships, so the model is "more like what Wikidata does -as an entity graph with typed statements."[^1] We accept that self-description completely; it is the strongest available framing of what CDCF is for. Three consequences follow. +This paper begins where the committee itself began. Asked whether CDCF is meant to become the meta-level disambiguation layer for Catholic data — "kinda like what DOI does for +published URLs" — the answer was yes, with one refinement: `cdcf:` identifiers dereference to structured JSON-LD carrying typed relationships, so the model is "more like what +Wikidata does as an entity graph with typed statements."[^1] That self-description is accepted here completely; it is the strongest available framing of what CDCF is for. Three +consequences follow. -**First, both named models mint opaquely and carry readability in metadata.** A DOI name is, in the DOI Handbook's own words, "an opaque string" or "dumb number" — "nothing at all -can or should be inferred from the number," and "the only secure way of knowing anything about the entity that a particular DOI name identifies is by looking at the metadata that -the Registrant of the DOI name declares."[^2] Wikidata's Q-numbers work the same way: `Q7186` is one identifier labelled _Marie Curie_ in English and French and _Maria -Skłodowska-Curie_ in Polish.[^3] Neither system is illegible in practice; both are illegible in the _string_ and legible in the _payload_. +**First, both named models mint machine-readable identifiers and carry readability in metadata.** A DOI name is, in the DOI Handbook's own words, "an opaque string" or "dumb +number" — "nothing at all can or should be inferred from the number," and "the only secure way of knowing anything about the entity that a particular DOI name identifies is by +looking at the metadata that the Registrant of the DOI name declares."[^2] Wikidata's Q-numbers work the same way: `Q7186` is one identifier labelled _Marie Curie_ in English and +French and _Maria Skłodowska-Curie_ in Polish.[^3] Neither system is illegible in practice; both are illegible in the _string_ and legible in the _payload_. Neither asks its +readers to give up legibility — each relocates it. **Second, the richer the resolution payload, the less semantic work the identifier string must do.** DOI resolves to a bare target URI and still succeeds. Draft 0.4.0 resolves to JSON-LD with `notations`, `crossReferences`, `licenses`, `doctrinalHistory`, and typed authority metadata (§5.4, §5.9). CDCF has built a resolution layer far richer than DOI's, -which means the marginal legibility a transparent string buys is far smaller here than in the systems transparency's advocates usually cite. The payload has absorbed the job. +which means the marginal legibility a transparent string buys is far smaller here than in the systems transparency's advocates usually cite. The payload has absorbed the job — +and it does the job better, in every language at once. **Third, a disambiguation layer must be neutral among the names it arbitrates.** The purpose of such a layer is to adjudicate between competing names, spellings, languages, and -structural placements for one referent. A transparent identifier pre-commits to one side of exactly the disputes the layer exists to resolve — silently, in every citation, forever. -It is a strange arbiter that writes its verdict into its own name. +structural placements for one referent. A transparent identifier pre-commits to one side of exactly the disputes the layer exists to resolve — silently, in every citation, +forever. It is a strange arbiter that writes its verdict into its own name. --- -## What We Agree On +## Where the Proposal Agrees with the Rationale Doc -We want to be precise about how much of `draft-cdcf-identifier-rationale-00` we accept, because it is more than the disagreement. +Precision about how much of `draft-cdcf-identifier-rationale-00` this paper accepts is worth the space, because the agreement is larger than the disagreement. -**The four-axes framework is right, and we adopt it.** The rationale doc separates axis A (grammar), axis B (transparency), axis C (structure), and axis D (consumption), and -observes that fixing one does not fix the others (§2). That is correct and clarifying, and this paper argues inside that vocabulary. **We concede axis A entirely** — "you want a -grammar either way" is simply true, and under this proposal the ABNF work is not discarded but moves to the layer where hand-authored strings actually live, governing the notation -and alias vocabulary plus one trivial production for the canonical shape. **We concede axis D entirely** — a reasoner must treat an IRI as a rigid designator, and the doc is right -that opacity-to-reasoners does not entail opaque minting. Our case for opaque minting is independent of D; it rests on durability and neutrality, not reasoner correctness. +**The four-axes framework is right, and this paper adopts it.** The rationale doc separates axis A (grammar), axis B (transparency), axis C (structure), and axis D (consumption), +and observes that fixing one does not fix the others (§2). That is correct and clarifying, and the argument below runs inside that vocabulary. **Axis A is conceded entirely** — +"you want a grammar either way" is simply true, and under this proposal the ABNF work is not discarded but moves to the layer where hand-authored strings actually live, governing +the notation and alias vocabulary plus one trivial production for the canonical shape. **Axis D is conceded entirely** — a reasoner must treat an IRI as a rigid designator, and +the doc is right that opacity-to-reasoners does not entail machine-readable minting. The case for machine-readable minting made here is independent of D; it rests on durability +and neutrality, not reasoner correctness. **Draft 0.4.0's two-artifact model is convergence, not conflict.** The most important thing in the 0.4.0 revision is §3.6: the recognition that graph identity and public citation -are two jobs, and that both can be carried on one entity. The `notations` array (§5.9), the scheme URNs, the never-reassign stability guarantee (§3.4), and the -`exact-match`/`close-match` relations that correctly refuse blanket `owl:sameAs` (§4.8.2, §3.6.1) are precisely the infrastructure an opaque-primary architecture needs. **This -proposal reuses all of it.** We ask the committee to build nothing it has not already specified — only to decide which artifact carries the stability guarantee. +are two jobs, and that both can be carried on one entity. That recognition is the same _et-et_ instinct this paper builds on. The `notations` array (§5.9), the scheme URNs, the +never-reassign stability guarantee (§3.4), and the `exact-match`/`close-match` relations that correctly refuse blanket `owl:sameAs` (§4.8.2, §3.6.1) are precisely the +infrastructure a machine-readable-primary architecture needs. **This proposal reuses all of it.** The committee is invited to build nothing it has not already specified — only to +decide which artifact carries the stability guarantee. **The slug schemes are the right vocabulary, in the wrong slot.** The registry slugs are careful, well-researched, and genuinely useful; they are exactly the alias vocabulary the -standard needs — and CatholicOS has already built the mechanism we propose to generalize, since CRMEDR ships `data/deprecated_ids.json` alongside `i18n/la.json`, `i18n/it.json`, -and `i18n/en.json`.[^4] Deprecation records plus multilingual labels beside an identifier is not an architecture we are importing; it is a pattern this organization built once -already, and our proposal is that it be applied universally rather than per-repo. **CSC.rdf's modeling instincts are right too, and we endorse them by name:** the Catholic Semantic -Canon ontology attaches the edition by property (`hasEdition`, with `John_1_14` pointing at `NovaVulgata`) rather than baking it into the base text unit's IRI, and models -vernacular renderings as first-class `Translation` artifacts linked by `hasTranslation`/`translationOf`.[^5] Both are what this paper proposes to make universal: volatile and -language-specific facts belong in properties, not identifiers. +standard needs — and CatholicOS has already built the mechanism this paper proposes to generalize, since CRMEDR ships `data/deprecated_ids.json` alongside `i18n/la.json`, +`i18n/it.json`, and `i18n/en.json`.[^4] Deprecation records plus multilingual labels beside an identifier is not an imported architecture; it is a pattern this organization built +once already, and the proposal is that it be applied universally rather than per-repo. **CSC.rdf's modeling instincts are right too, and this paper endorses them by name:** the +Catholic Semantic Canon ontology attaches the edition by property (`hasEdition`, with `John_1_14` pointing at `NovaVulgata`) rather than baking it into the base text unit's IRI, +and models vernacular renderings as first-class `Translation` artifacts linked by `hasTranslation`/`translationOf`.[^5] Both are what this paper proposes to make universal: +volatile and language-specific facts belong in properties, not identifiers — and, once there, they can all be true at once. The dispute, then, is narrow. It lives on axes B and C: whether the _primary_ minted string carries meaning, and whether it encodes hierarchy. @@ -99,8 +182,8 @@ The dispute, then, is narrow. It lives on axes B and C: whether the _primary_ mi ## The Evidence: Transparency Produced Divergence, Not Convergence -The rationale doc's strongest empirical claim is that transparency is the field's answer for hand-authored citation strings. Ours is narrower and closer to home: **inside this -committee's own work, transparency has produced divergence rather than convergence — and quickly.** +The rationale doc's strongest empirical claim is that transparency is the field's answer for hand-authored citation strings. The claim made here is narrower and closer to home: +**inside this committee's own work, transparency has produced divergence rather than convergence — and quickly.** ### Two entities, six spellings @@ -112,8 +195,8 @@ committee's own work, transparency has produced divergence rather than convergen The first two spellings of each pair sit in one file, on adjacent entities, under `purl.org/cdcf/ontology/catholic-semantic-canon#`;[^5] the third Trent form is the reference-linking example in the Rome working-session materials.[^6] That same file also carries `CCC-1376`, `ST-III-75-4`, and `CIC1983-915` — four shape conventions inside one `identifier` property. None of this is carelessness; each spelling is locally reasonable. That is the point: transparent identifiers are locally reasonable in incompatible ways, -and no mechanical test detects the divergence. **One opaque canonical ID would have carried all six spellings as notations, and the divergence would have been visible as what it -is: six citation forms for two entities.** +and no mechanical test detects the divergence. **One machine-readable canonical ID would have carried all six spellings as notations — none of them lost, none of them +privileged — and the divergence would have been visible as what it is: six citation forms for two entities.** ### The recorded in-repo record @@ -153,96 +236,9 @@ Every one of these is a governance-quality problem a diligent committee will kee is what a pre-1.0 registry is for. Two answers. First, litcal, romcal, and ePrex are not drafts; they are production systems, and the St Martha divergence happened in their _shipped_ keys — a 2021 decree left `StMartha` frozen factually wrong in a production API while a sibling project renamed to a 56-character key — so the mechanism operates after normativity, not only before it. Second, the churn's causes — multilingual naming, movable anchors, editorial judgment about which name is _the_ name — do not end at 1.0: popes -keep being elected, dioceses keep merging, decrees keep expanding memorials. Normativity freezes the identifiers, not the world. That is precisely the cost we ask the committee to -weigh: transparency's maintenance burden is not a one-time migration but a standing obligation, recurring whenever the world, an edition, a decree, or a translation policy changes. - ---- - -## The Steelmen — and Where Each Fails - -We are not neutral, but the transparent case deserves its full strength, because a proposal that defeats only weak versions of the opposition deserves to lose. Each steelman below -is followed by the failure mode we think it develops over time, a concrete mitigation from the affordance layer, and an invitation. If the transparent camp can name a remedy for -one of these failure modes that costs less than ours, that should decide the question. - -### Steelman 1 — The BCP 47 / Unicode layered model - -**At full strength.** The rationale doc's §5 is its best section. BCP 47 composes stable atomic registry codes by ABNF into `zh-Hant-TW`: a tag simultaneously transparent to a -human, grammar-validated, built from stable atoms, and opaque to a matcher. Unicode pairs an opaque code point `U+0041` with a name — `LATIN CAPITAL LETTER A` — frozen forever by -its stability policy. The i18n stack runs on these, in CDCF's exact use case: hand-authored, quoted-in-the-wild identifiers, for two decades. Nobody writes `lang="Q1860"`. - -**Failure mode.** BCP 47's registry survives change only through alias machinery — which is our architecture, not theirs. RFC 5646 §3.1.6 and §3.1.7 define `Deprecated` and -`Preferred-Value` precisely so the registry can carry `iw` forever while canonicalizing it to `he`; the §3.4 stability guarantee is that `Subtag`, `Type`, and `Added` "MUST NOT be -changed" — the subtag is never removed, only redirected.[^12] That is an opaque-spine architecture wearing legible clothes. More decisively, language tags have a property no -Catholic entity has: they are written in the one alphabet definitionally neutral for their domain, and their referents are the naming authorities themselves. There is no -Latin-versus-Italian dispute about how to spell `en`. There is exactly such a dispute about `pope-{name}-{roman}`: Leo, or Leone, or León? CECDR's ten-language ordinariate slugs -are that dispute already lost, at scale. - -**Mitigation and invitation.** Adopt the BCP 47 architecture in full, including the part that does the work. Under this proposal `Preferred-Value` is not approximated; it _is_ the -alias layer, and every competing spelling (`pope-leo-xiv`, `papa-leone-xiv`, `papa-leon-xiv`) becomes a permanent resolvable notation on one canonical ID, none privileged and none -wrong. We would welcome a proposal for how a transparent primary IRI carries three co-equal language forms without electing one. - -### Steelman 2 — Code ergonomics - -**At full strength.** `if (key == "immaculate-conception")` is readable in a diff, greppable in a log, and self-documenting in a stack trace. `if (key == "Rk8f3vQ2…")` is none of -those. Most humans who ever touch these identifiers are application developers, not ontologists, and their productivity is a real cost that ontological purity does not pay. - -**Failure mode.** The legibility is real but bound in the wrong place: in a literal, at every call site. When the referent's boundary or preferred name changes, every call site is -stale in a way no compiler catches. The string was never checked against the record; it was trusted because it looked right. - -**Mitigation and invitation.** Named constants bound to canonical IDs — `const IMMACULATE_CONCEPTION = "R…"` — which is how every codebase already handles hex colors, port numbers, -and country codes. That restores full call-site legibility while leaving exactly one authoritative binding to audit, and the rendering rule (§8, item 4) extends it to registry -source files, serializations, and generated code so a label always sits beside the ID. If the constant-binding overhead is too high for a particular consumer, we would like to see -that workflow, so the affordance layer can be shaped around it. - -### Steelman 3 — Church-oversight verifiability - -**At full strength.** Ecclesiastical review is a genuine requirement, and reviewers are theologians and canonists, not engineers. A bishop's delegate can read -`cdcf:magisterium/pope-leo-xiii/rerum-novarum` and confirm it is right. Nobody can review a page of base62. - -**Failure mode.** Transparency converts review from _verification_ into _recognition_. The reviewer confirms the string looks correct; the string is not thereby checked against the -record. When a slug is subtly wrong — a garbled Latin incipit captured as a subject name, a date drawn from the wrong edition, an ordinal the Church never used — it passes review -precisely because it reads plausibly. Every recorded CRMEDR correction above was a plausible-reading slug that shipped. - -**Mitigation and invitation.** The first half needs no tooling at all: under commitment 4 of the Proposal, the registry source file a canonist reviews _remains a plain text file_, -with the Latin label in the column adjacent to the ID. Nothing is taken away from text-file reviewers — the opaque ID is an added column, not a substitute — so recognition-style -review continues exactly as it does today, while verification-style review becomes possible on top of it. Verification against labels rendered _from_ canonical IDs is strictly -safer, because it checks the record rather than trusting the key. Under the rendering rule a review surface shows the canonical ID with its `skos:prefLabel` in the reviewer's own -language — Latin, Italian, English — resolved live from the registry, so a reviewer sees what the system actually believes rather than what a past minting decision asserted. We -would welcome the committee's review-workflow requirements as design input for that rule, which is the natural place to encode them. - -### Steelman 4 — Things versus concepts - -**At full strength.** The rationale doc's §6 is its most original contribution and genuinely explanatory: OBO Foundry and the Gene Ontology went opaque because biological -categories are reclassified as science advances, while BCP 47 and OSIS went transparent because languages and scriptural books are fixed.[^13] CDCF spans both, so it should apply -both rules per dataset. The residual fringe among things — antipopes, the Stephen II/III ambiguity, the skipped John XX — is finite, enumerable, and already adjudicated by the -Church's own historical record. - -**Failure mode.** The criterion cross-cuts the doc's own evidence. Its §4.1 table lists **VIAF, Getty (TGN/ULAN/AAT), and GeoNames** as opaque — and those are authority files of -_things_: persons, places, named artifacts. It lists **schema.org, FOAF, Dublin Core, and SKOS** as transparent — and those are _concept_ vocabularies of classes and properties. If -things→transparent and concepts→opaque were the rule, those rows sit on the wrong side of it; and the row the doc itself flags as drift-prone is **DBpedia**, transparent -identifiers for things, derived from article titles, which break on rename (§4.1). The line the field actually draws is simpler, and it is the doc's own summary sentence in §4.1: -**opacity clusters where a resolver always mediates.** That describes CDCF exactly — draft 0.4.0 specifies a resolution server, mirror discovery, content negotiation, a change -feed, and one-year immutable caching (§5.1–§5.9). CDCF is not choosing whether to be resolver-mediated; it has already chosen. - -**Mitigation and invitation.** Apply the resolver criterion, which CDCF satisfies, rather than the things/concepts criterion, which the doc's own evidence table does not support — -and keep the things/concepts insight where it is undeniably right: as the rule for which _notation scheme_ to feature. Scripture keeps OSIS, canons keep their numbers, CCC keeps -its paragraph numbers; this proposal never asks anyone to stop writing them, only that they be notations on a durable spine rather than the spine itself. If the committee believes -there is a closed class of entities whose canonical naming is genuinely finished, we would like to see it enumerated — our reading of the six registries is that every candidate -class has already recorded a rename. - -### Steelman 5 — "Just take a vote on the language" - -**At full strength.** Standards bodies decide contested questions by deliberation and vote all the time. Latin is the Church's own language and an obvious Schelling point. -Committees exist to make exactly these calls; declaring the question unanswerable is an abdication. - -**Failure mode.** A vote produces a winner and a resentful minority — per identifier, permanently, and visibly in every citation string. That is a governance tax recurring with -every new entity and compounding with every tradition the standard hopes to serve. It is also empirically what has happened: Open Issue 5 has been open across four revisions, -CRMEDR mints Latin lemmas, CLEDR mints English snake_case, and CECDR's ordinariates ended up in ten languages without anyone ever deciding they should. - -**Mitigation and invitation.** Multivalued labels produce no losers. `skos:prefLabel` is language-tagged; a Polish reader gets Polish and a Latin reader gets Latin, from one -record, with no election held. The precedent is trivially familiar: _honor_ and _honour_ are both correct, and no standard had to choose, because they are labels rather than keys. -If a vote is nonetheless preferred for a given domain, note that under this proposal it decides which notation carries `prefLabel: true` — a reversible, low-stakes call — rather -than which identifier the world cites for a century. +keep being elected, dioceses keep merging, decrees keep expanding memorials. Normativity freezes the identifiers, not the world. That is precisely the cost the committee is asked +to weigh: transparency's maintenance burden is not a one-time migration but a standing obligation, recurring whenever the world, an edition, a decree, or a translation policy +changes. Under this proposal none of those events costs an identifier — each of them adds a label or a notation, which is what those layers are for. --- @@ -259,18 +255,19 @@ Montenegro, and ISO's own archival code for the latter had to be changed from `C Burma → Myanmar archived as `BUMM`, withdrawn alpha-2 codes transitionally reserved for at least fifty years before any possible reuse.[^14] The rationale doc reads this correctly (§7): reassignment is the danger, and it is a governance failure available to both camps. -**Opaque-primary architectures absorb it in the alias layer, which is built to age.** When a slug turns out to be wrong it is not corrected — it is _joined_. +**Machine-readable-primary architectures absorb it in the alias layer, which is built to age.** When a slug turns out to be wrong it is not corrected — it is _joined_. `mr:1210-marcus-antonius-durando` and `mr:0610-marcus-antonius-durando` both resolve, forever, to one entity; the Latin edition and the CEI edition are each right about their own book; nothing published ever breaks. `StMartha`, `martha`, and the 56-character romcal key all resolve to one celebration, and the 2021 decree becomes a property on the record -rather than a naming crisis in three projects. Aliases are supposed to pile up. Canonical identifiers are not. +rather than a naming crisis in three projects. Aliases are supposed to pile up. Canonical identifiers are not. Note what this preserves: every legible string anyone has ever +written keeps working, in the layer where legibility belongs. -And the cost asymmetry has inverted since this trade-off was last argued seriously. The rationale doc calls opacity's cost "permanent and unsolvable by design" — a mandatory lookup -for every human who reads the identifier (§7). That was true in 2005. Hover labels, IDE inlays, resolver-backed link previews, and agents that never hand-type an identifier have -collapsed the lookup cost toward zero, and CDCF's own resolution protocol is precisely the substrate those affordances run on. CatholicOS need not take this on faith: ontokit -already renders resolver-backed labels beside opaque `osc:` R-IDs today — its source-view hover resolves a full IRI to its label, and its `useIriLabels` hook label-joins API -responses.[^15] Drift costs move the other way: they rise with every integration, every downstream project, every new edition, and every tradition the standard reaches for. -**Legibility costs are falling toward zero; drift costs rise with adoption.** The transparent trade-off was right for 2005. We do not think it is right for a standard minted in -2026 to last a century. +And the cost asymmetry has inverted since this trade-off was last argued seriously. The rationale doc calls the cost of opaque (durable) identifiers "permanent and unsolvable +by design" — a mandatory lookup for every human who reads the identifier (§7). That was true in 2005. Hover labels, IDE inlays, resolver-backed link previews, and agents that +never hand-type an identifier have collapsed the lookup cost toward zero, and CDCF's own resolution protocol is precisely the substrate those human-readable surfaces run on. +CatholicOS need not take this on faith: ontokit already renders resolver-backed labels beside machine-readable `osc:` R-IDs today — its source-view hover resolves a full IRI to +its label, and its `useIriLabels` hook label-joins API responses.[^15] Drift costs move the other way: they rise with every integration, every downstream project, every new +edition, and every tradition the standard reaches for. **Legibility costs are falling toward zero; drift costs rise with adoption.** The transparent trade-off was right for 2005. +It is not right for a standard minted in 2026 to last a century. --- @@ -290,32 +287,36 @@ Magisterium, Local Magisterium, Commentary, Private opinion — placing Local Ma Theological Commentary above Local Magisterium.[^6] Second, the two six-value lists and the five-value list do not measure the same thing: the first two answer _who teaches_, and draft 0.4.0's answers _what assent is owed_ (§6.1). Those axes correlate but do not align; the same document can sit high on one and lower on the other. -Harmonizing this is real theological work, and not this paper's to do — draft 0.4.0 §7.3 rightly requires theologian and canonist review for exactly these classifications. What we -can say is what the harmonization will do to the identifiers. It will rename levels, reorder them, split at least one, and possibly separate the two axes into two vocabularies. If -level-IDs are transparent strings, every one of those moves breaks every text already tagged: a corpus tagged `"UniversalOrdinary"` is stranded the moment the level is renamed or -its boundary redrawn, and the migration is a rewrite of the annotation layer rather than of a lookup table. If level-IDs are opaque, with today's names carried as labels and all -three existing vocabularies carried as notations, the theology can develop and the data survives — the record's label changes, the tagged corpus does not move, and the crosswalk -between the who-teaches and what-assent axes becomes a property rather than a renaming. That is the whole argument in one case. **The identifier architecture should let the -theology be revised. It should not require the theology to be finished first.** +Harmonizing this is real theological work, and not this paper's to do — draft 0.4.0 §7.3 rightly requires theologian and canonist review for exactly these classifications. What can +be said now is what the harmonization will do to the identifiers. It will rename levels, reorder them, split at least one, and possibly separate the two axes into two vocabularies. +If level-IDs are transparent strings, every one of those moves breaks every text already tagged: a corpus tagged `"UniversalOrdinary"` is stranded the moment the level is renamed or +its boundary redrawn, and the migration is a rewrite of the annotation layer rather than of a lookup table. If level-IDs are machine-readable, with today's names carried as labels +and all three existing vocabularies carried as notations, the theology can develop _and_ the data survives — every name in the table above stays readable and citable, the record's +label changes, the tagged corpus does not move, and the crosswalk between the who-teaches and what-assent axes becomes a property rather than a renaming. That is the whole argument +in one case. **The identifier architecture should let the theology be revised. It should not require the theology to be finished first.** --- ## The Proposal -Five commitments. Nothing here replaces draft 0.4.0's machinery; each item names the 0.4.0 mechanism it rides on. +Five commitments. Nothing here replaces draft 0.4.0's machinery; each item names the 0.4.0 mechanism it rides on. The shape of the whole is _et-et_: not _aut-aut_ — a durable +identifier _or_ an interpretable record — but both, the machine-readable ID canonical and the human-readable layer guaranteed by the standard rather than left to each +implementer's good intentions. -1. **Canonical identifiers are opaque, `R`-shaped, and org-wide.** Every canonical ID CatholicOS mints — ontology and every data registry — is a base62-encoded 128-bit UUID - prefixed `R` (23 characters in practice; 122 random bits, so decentralized minting needs no counter and no central allocator, and the leading letter keeps it QName-safe for - RDF/XML). This is not new for CDCF: it is the shape already shipping in `ontology-semantic-canon`, where `osc:RChKPk9K152BirrIYgAREsY` is _Clergy_, and the shape ontokit already - mints.[^15] It is also FOLIO's shape — `folio.openlegalstandard.org/R7Ttdyo4FsvaupPKT35Qry0` is _Murder_.[^16] +1. **Canonical identifiers are machine-readable, `R`-shaped, and org-wide.** Every canonical ID CatholicOS mints — ontology and every data registry — is a base62-encoded 128-bit + UUID prefixed `R` (23 characters in practice; 122 random bits, so decentralized minting needs no counter and no central allocator, and the leading letter keeps it QName-safe + for RDF/XML). This is not new for CDCF: it is the shape already shipping in `ontology-semantic-canon`, where `osc:RChKPk9K152BirrIYgAREsY` is _Clergy_, and the shape ontokit + already mints.[^15] It is also FOLIO's shape — `folio.openlegalstandard.org/R7Ttdyo4FsvaupPKT35Qry0` is _Murder_.[^16] 2. **Every existing slug, key, and registry prefix becomes a permanent resolvable alias.** `mr:`, `circ:`, `icl:`, CLEDR keys, CLBDR edition IDs, litcal/romcal/eprex keys, `cdcf:verse/jn/1/14`, `cdcf:concept/C0000418` — all carried as 0.4.0 `notations` (§5.9) under declared scheme URNs, resolvable per §3.4, **never deprecated and never reused**. This is the IP/DNS model, and the analogy is worth stating plainly: nobody argues that an IP address should be human-readable, and nobody has to, because the domain name resolves to it. No adopter loses a working key. That is the adoption story. -3. **Every entity carries multilingual labels.** `skos:prefLabel` and `skos:altLabel`, language-tagged, on every entity in every registry. Naming disputes are resolved by **adding - a label**, never by changing an identifier. CRMEDR's `i18n/{la,it,en}.json` is the existing precedent; this generalizes it. +3. **Every entity carries multilingual labels.** `skos:prefLabel` and `skos:altLabel`, language-tagged, on every entity in every registry — `"Sanctorum Petri et Pauli, +Apostolorum"@la`, `"Santi Pietro e Paolo, apostoli"@it`, `"Sts. Peter & Paul, Apostles"@en` on one record. Naming disputes are resolved by **adding a label**, never by changing + an identifier. CRMEDR's `i18n/{la,it,en}.json` is the existing precedent; this generalizes it. 4. **Production surfaces MUST render a label beside every canonical ID.** Registry source files carry a label column or comment beside each ID; serializations carry the label - inline; UIs and generated code render it. This commitment is what makes opacity livable, and it belongs in the standard rather than in each implementer's good intentions. + inline; UIs and generated code render it. This commitment is the _et_ that makes the other _et_ livable, and it belongs in the standard rather than in each implementer's good + intentions. 5. **Structural and volatile facts live in properties, never in canonical identifiers.** Dates and calendar position (CRMEDR's `MMDD` anchor), country codes (CECDR's ISO 3166 prefix), taxonomy supertypes (`circumscription`, `institute`, the retired `order`), edition years (CLBDR), chapter/verse hierarchy, ownership, and language are all already modeled as fields in 0.4.0 responses. Where a fact lives both in a field and in the identifier, the identifier is the copy that goes stale. @@ -324,10 +325,10 @@ Five commitments. Nothing here replaces draft 0.4.0's machinery; each item names ## Universal Slugs Across Standards -Draft 0.4.0 §4.9 already states the principle we want to generalize: a sibling registry's slug "MUST be reused verbatim as the final path segment of the corresponding `cdcf:` IRI … +Draft 0.4.0 §4.9 already states the principle worth generalizing: a sibling registry's slug "MUST be reused verbatim as the final path segment of the corresponding `cdcf:` IRI … This is a MUST, not a convention: it is what makes the pairing machine-verifiable rather than merely coincidental." That is exactly right, and it is the seed of something larger. -Lift it one level, from sibling registries to peer standards: when CatholicOS and another standard identify the same referent, they reuse the **same opaque local name** under their -own namespaces. +Lift it one level, from sibling registries to peer standards: when CatholicOS and another standard identify the same referent, they reuse the **same machine-readable local name** +under their own namespaces. ```text https://ontology.catholicos.catholic/R7Ttdyo4FsvaupPKT35Qry0 @@ -335,48 +336,141 @@ https://folio.openlegalstandard.org/R7Ttdyo4FsvaupPKT35Qry0 ``` Cross-standard identity becomes machine-verifiable by local-name equality — no mapping table, no crosswalk to maintain, no `sameAs` hazard. Each standard keeps its own namespace, -governance, labels, and resolution payload; only the local name is shared. **This is possible only because the shared name asserts nothing in anyone's language.** No tradition will -agree to share `god-the-son` across Jewish, Catholic, and Protestant standards — the string itself is a theological claim, and for two of the three it is the wrong one. But every -tradition can share `R7Ttdyo4FsvaupPKT35Qry0`, because it says nothing at all, and each standard attaches its own label, definition, and typed statements to it. Opacity is not -merely tolerable at the inter-tradition boundary; **it is the only thing that crosses it.** That is the interfaith-alignment benefit named in the goals ladder, and it is concrete -rather than aspirational: shape-compatibility with FOLIO today, and a mechanism that scales to machine-readable identifier work across the Abrahamic traditions tomorrow. The -alternative — a mapping table between every pair of standards, maintained by both parties in perpetuity — is the cost transparency imposes at exactly the boundary where the goals -ladder can least afford it. +governance, labels, and resolution payload; only the local name is shared, and each side's readers still see their own tradition's words. **This is possible only because the shared +name asserts nothing in anyone's language.** No tradition will agree to share `god-the-son` across Jewish, Catholic, and Protestant standards — the string itself is a theological +claim, and for two of the three it is the wrong one. But every tradition can share `Rh895FOUgLMfrylaO713w1`, because it says nothing at all, and each standard attaches its own +label, definition, and typed statements to it — the very record shown in the first exemplar above. A machine-readable name is not merely tolerable at the inter-tradition boundary; +**it is the only thing that crosses it, and it is what lets each tradition keep its own language on its own side.** That is the interfaith-alignment benefit named in the goals +ladder, and it is concrete rather than aspirational: shape-compatibility with FOLIO today, and a mechanism that scales to machine-readable identifier work across the Abrahamic +traditions tomorrow. The alternative — a mapping table between every pair of standards, maintained by both parties in perpetuity — is the cost transparency imposes at exactly the +boundary where the goals ladder can least afford it. --- -## Costs We Accept — and Their Remedies +## Costs the Proposal Accepts — and Their Remedies -Opacity has real costs. We would rather name them than let them be discovered later. +Machine-readable identifiers have real costs. Naming them here is better than letting them be discovered later. Each remedy is a piece of the human-readable layer, which is why +the layer is a commitment of the proposal rather than an aspiration. | Cost | The problem, stated plainly | Remedy | | :----------------------- | :-------------------------------------------------------------------------------------------------------------------------- | :--------------------------------------------------------------------------------------------------------------------------------------- | | **Diff reviewability** | A slug in a pull request is self-checking; a reviewer sees a wrong one. `Rk8f3vQ2…` is not self-checking and never will be. | Mandatory label columns in registry source files, so every diff line carries an ID **and** a human-readable label that reviews itself. | -| **Silent wrong-paste** | Paste the wrong opaque ID and nothing looks wrong. Paste the wrong slug and something usually does. | Lint rules validating every ID against the registry and asserting label/ID agreement in CI; label comments beside IDs in serializations. | +| **Silent wrong-paste** | Paste the wrong machine-readable ID and nothing looks wrong. Paste the wrong slug and something usually does. | Lint rules validating every ID against the registry and asserting label/ID agreement in CI; label comments beside IDs in serializations. | | **Debugging ergonomics** | A log line, stack trace, or SPARQL result full of base62 is harder to read than one full of slugs. | IDE inlays and resolver-backed hover labels; named constants bound to canonical IDs at call sites; label-joining helpers in tooling. | | **Onboarding friction** | A newcomer reading raw data cannot orient without a lookup. | The §5.9 `notations` array ships every canonical record with all its familiar slugs, so a newcomer's existing vocabulary still works. | -None of these remedies is speculative — each exists in shipped systems, and two exist in CatholicOS repositories today. But we do not claim the list is complete. **We invite the -committee, and especially the transparent camp, to name costs we have missed and mitigations we have not thought of.** If a cost turns out to have no adequate remedy, that is a -finding worth having before adoption rather than after. +None of these remedies is speculative — each exists in shipped systems, and two exist in CatholicOS repositories today. But the list is not claimed to be complete. **The +committee, and especially the transparent camp, is invited to name costs this paper has missed and mitigations it has not thought of.** If a cost turns out to have no adequate +remedy, that is a finding worth having before adoption rather than after. + +--- + +## Objections at Full Strength — and Their Remedies + +This paper is not neutral, but the transparent case deserves its full strength, because a proposal that defeats only weak versions of the opposition deserves to lose. Each +objection below is stated at its best, then followed by the failure mode it develops over time, a concrete mitigation from the human-readable layer, and an invitation. If the +transparent camp can name a remedy for one of these failure modes that costs less than the one offered here, that should decide the question. + +### Objection 1 — The BCP 47 / Unicode layered model + +**At full strength.** The rationale doc's §5 is its best section. BCP 47 composes stable atomic registry codes by ABNF into `zh-Hant-TW`: a tag simultaneously transparent to a +human, grammar-validated, built from stable atoms, and opaque to a matcher. Unicode pairs an opaque code point `U+0041` with a name — `LATIN CAPITAL LETTER A` — frozen forever by +its stability policy. The i18n stack runs on these, in CDCF's exact use case: hand-authored, quoted-in-the-wild identifiers, for two decades. Nobody writes `lang="Q1860"`. + +**Failure mode.** BCP 47's registry survives change only through alias machinery — which is this proposal's architecture, not transparency's. RFC 5646 §3.1.6 and §3.1.7 define +`Deprecated` and `Preferred-Value` precisely so the registry can carry `iw` forever while canonicalizing it to `he`; the §3.4 stability guarantee is that `Subtag`, `Type`, and +`Added` "MUST NOT be changed" — the subtag is never removed, only redirected.[^12] That is a machine-readable spine wearing legible clothes: the durable layer and the readable +layer, already both present. More decisively, language tags have a property no Catholic entity has: they are written in the one alphabet definitionally neutral for their domain, +and their referents are the naming authorities themselves. There is no Latin-versus-Italian dispute about how to spell `en`. There is exactly such a dispute about +`pope-{name}-{roman}`: Leo, or Leone, or León? CECDR's ten-language ordinariate slugs are that dispute already lost, at scale. + +**Mitigation and invitation.** Adopt the BCP 47 architecture in full, including the part that does the work. Under this proposal `Preferred-Value` is not approximated; it _is_ the +alias layer, and every competing spelling of the reigning pope — `pope-leo-xiv`, `papa-leone-xiv`, `papa-león-xiv` — becomes a permanent resolvable notation on one canonical ID, +none privileged, none wrong, and none dropped. An Anglophone developer keeps writing `pope-leo-xiv`; an Italian diocesan office keeps writing `papa-leone-xiv`; both resolve to the +same record and neither had to win. A proposal for how a transparent primary IRI carries three co-equal language forms without electing one would be welcome. + +### Objection 2 — Code ergonomics + +**At full strength.** `if (key == "immaculate-conception")` is readable in a diff, greppable in a log, and self-documenting in a stack trace. `if (key == "Rk8f3vQ2…")` is none of +those. Most humans who ever touch these identifiers are application developers, not ontologists, and their productivity is a real cost that ontological purity does not pay. + +**Failure mode.** The legibility is real but bound in the wrong place: in a literal, at every call site. When the referent's boundary or preferred name changes, every call site is +stale in a way no compiler catches. The string was never checked against the record; it was trusted because it looked right. + +**Mitigation and invitation.** Named constants bound to canonical IDs — `const PETER_AND_PAUL = "R7kQp2mXf4LdTbz9Ns3Hc1"` — which is how every codebase already handles hex colors, +port numbers, and country codes. That restores full call-site legibility while leaving exactly one authoritative binding to audit, and the rendering rule (§7, item 4) extends it to +registry source files, serializations, and generated code so a label always sits beside the ID — the fourth exemplar in the executive summary is the shipped version of this, in +ontokit today. If the constant-binding overhead is too high for a particular consumer, that workflow would be worth seeing, so the human-readable layer can be shaped around it. + +### Objection 3 — Church-oversight verifiability + +**At full strength.** Ecclesiastical review is a genuine requirement, and reviewers are theologians and canonists, not engineers. A bishop's delegate can read +`cdcf:magisterium/pope-leo-xiii/rerum-novarum` and confirm it is right. Nobody can review a page of base62. + +**Failure mode.** Transparency converts review from _verification_ into _recognition_. The reviewer confirms the string looks correct; the string is not thereby checked against the +record. When a slug is subtly wrong — a garbled Latin incipit captured as a subject name, a date drawn from the wrong edition, an ordinal the Church never used — it passes review +precisely because it reads plausibly. Every recorded CRMEDR correction above was a plausible-reading slug that shipped. + +**Mitigation and invitation.** The first half needs no tooling at all: under commitment 4 of the Proposal, the registry source file a canonist reviews _remains a plain text file_, +with the Latin label in the column adjacent to the ID — the second exemplar in the executive summary is that file. Nothing is taken away from text-file reviewers — the +machine-readable ID is an added column, not a substitute — so recognition-style review continues exactly as it does today, while verification-style review becomes possible on top +of it. Verification against labels rendered _from_ canonical IDs is strictly safer, because it checks the record rather than trusting the key. Under the rendering rule a review +surface shows the canonical ID with its `skos:prefLabel` in the reviewer's own language — Latin, Italian, English — resolved live from the registry, so a reviewer sees what the +system actually believes rather than what a past minting decision asserted. The committee's review-workflow requirements would be welcome as design input for that rule, which is +the natural place to encode them. + +### Objection 4 — Things versus concepts + +**At full strength.** The rationale doc's §6 is its most original contribution and genuinely explanatory: OBO Foundry and the Gene Ontology went opaque because biological +categories are reclassified as science advances, while BCP 47 and OSIS went transparent because languages and scriptural books are fixed.[^13] CDCF spans both, so it should apply +both rules per dataset. The residual fringe among things — antipopes, the Stephen II/III ambiguity, the skipped John XX — is finite, enumerable, and already adjudicated by the +Church's own historical record. + +**Failure mode.** The criterion cross-cuts the doc's own evidence. Its §4.1 table lists **VIAF, Getty (TGN/ULAN/AAT), and GeoNames** as opaque — and those are authority files of +_things_: persons, places, named artifacts. It lists **schema.org, FOAF, Dublin Core, and SKOS** as transparent — and those are _concept_ vocabularies of classes and properties. If +things→transparent and concepts→opaque were the rule, those rows sit on the wrong side of it; and the row the doc itself flags as drift-prone is **DBpedia**, transparent +identifiers for things, derived from article titles, which break on rename (§4.1). The line the field actually draws is simpler, and it is the doc's own summary sentence in §4.1: +**opacity clusters where a resolver always mediates.** That describes CDCF exactly — draft 0.4.0 specifies a resolution server, mirror discovery, content negotiation, a change +feed, and one-year immutable caching (§5.1–§5.9). CDCF is not choosing whether to be resolver-mediated; it has already chosen. + +**Mitigation and invitation.** Apply the resolver criterion, which CDCF satisfies, rather than the things/concepts criterion, which the doc's own evidence table does not support — +and keep the things/concepts insight where it is undeniably right: as the rule for which _notation scheme_ to feature. Scripture keeps OSIS, canons keep their numbers, CCC keeps +its paragraph numbers; this proposal never asks anyone to stop writing them, only that they be notations on a durable spine rather than the spine itself. `John 1:14` stays +`John 1:14` — in OSIS as `John.1.14`, in CSC.rdf as `csc:John_1_14`, in the 0.4.0 grammar as `cdcf:verse/jn/1/14` — all three on one record, none of them the thing that must never +change. If the committee believes there is a closed class of entities whose canonical naming is genuinely finished, an enumeration of it would be welcome — a reading of the six +registries suggests every candidate class has already recorded a rename. + +### Objection 5 — "Just take a vote on the language" + +**At full strength.** Standards bodies decide contested questions by deliberation and vote all the time. Latin is the Church's own language and an obvious Schelling point. +Committees exist to make exactly these calls; declaring the question unanswerable is an abdication. + +**Failure mode.** A vote produces a winner and a resentful minority — per identifier, permanently, and visibly in every citation string. That is a governance tax recurring with +every new entity and compounding with every tradition the standard hopes to serve. It is also empirically what has happened: Open Issue 5 has been open across four revisions, +CRMEDR mints Latin lemmas, CLEDR mints English snake_case, and CECDR's ordinariates ended up in ten languages without anyone ever deciding they should. + +**Mitigation and invitation.** Multivalued labels produce no losers. `skos:prefLabel` is language-tagged; a Polish reader gets Polish and a Latin reader gets Latin, from one +record, with no election held. The precedent is trivially familiar: _honor_ and _honour_ are both correct, and no standard had to choose, because they are labels rather than keys. +If a vote is nonetheless preferred for a given domain, note that under this proposal it decides which notation carries `prefLabel: true` — a reversible, low-stakes call — rather +than which identifier the world cites for a century. --- ## Recommended Amendments to Draft 0.4.0 -We offer these as amendments the committee can adopt into the existing text, not as a rival document. Draft 0.4.0's non-identifier machinery — the `licenses` object, authority +These are offered as amendments the committee can adopt into the existing text, not as a rival document. Draft 0.4.0's non-identifier machinery — the `licenses` object, authority metadata, content negotiation, caching, resilience, and the governance process — is endorsed as written. | § | Amendment | | :------------------- | :----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -| **§3.6** | Make the ontology IRI opaque for **every** domain, not only `concept/`. Each domain's current transparent path form becomes a **guaranteed** notation rather than an optional one. The two-artifact model is unchanged; only which artifact is primary. | +| **§3.6** | Make the ontology IRI machine-readable for **every** domain, not only `concept/`. Each domain's current transparent path form becomes a **guaranteed** notation rather than an optional one. The two-artifact model is unchanged; only which artifact is primary. | | **§3.3, Appendix D** | Re-scope the ABNF grammars to govern the notation layer, where hand-authored strings live. Add one production for canonical IRIs: `canonical-id = "R" 20*24(ALPHA / DIGIT)` (base62, case-sensitive — this relaxes §3.3's lower-case rule for canonical IDs only). | -| **§4.5** | Extend the opaque option from `concept/` to all domains, and recommend the `R` shape over sequential `C0000418` for new mints. Existing `C…` IDs are kept as permanent notations; **no forced migration** (§1.3 preserved). | -| **§4.9** | Extend the verbatim-slug MUST to **cross-standard** opaque local names: peer standards identifying the same referent SHOULD reuse the same local name under their own namespaces. | +| **§4.5** | Extend the opaque (durable) option from `concept/` to all domains, and recommend the `R` shape over sequential `C0000418` for new mints. Existing `C…` IDs are kept as permanent notations; **no forced migration** (§1.3 preserved). | +| **§4.9** | Extend the verbatim-slug MUST to **cross-standard** machine-readable local names: peer standards identifying the same referent SHOULD reuse the same local name under their own namespaces. | | **§3.4** | State explicitly that the stability guarantee applies to notations as permanent aliases: a published notation MUST remain resolvable and MUST NOT be reassigned, exactly as an IRI. | | **§5.9** | Make `notations` load-bearing in every domain (withdrawing §3.6's permission to omit it for "thing" domains), and add `skos:prefLabel`/`altLabel` as a REQUIRED language-tagged field on every resolution response. | | **New §5.10** | Add a production rendering rule: registry source files, serializations, UIs, and generated code MUST render a human-readable label adjacent to every canonical ID. | -| **§10, new issue** | Mint namespace: we recommend **one org-wide namespace** shared by the ontology and all registries. The host choice — `id.catholiccommons.org` (§3.1) versus `ontology.catholicos.catholic` (already live for `osc:`) — is a committee decision we do not presume. | +| **§10, new issue** | Mint namespace: **one org-wide namespace** shared by the ontology and all registries is recommended. The host choice — `id.catholiccommons.org` (§3.1) versus `ontology.catholicos.catholic` (already live for `osc:`) — is a committee decision. | Two notes on scope. This proposal executes no data migration; beyond the permanent-alias commitment, migration mechanics are committee work. And it proposes no theological harmonization — §7.3 review remains the right gate for anything touching doctrinal classification. @@ -449,9 +543,9 @@ harmonization — §7.3 review remains the right gate for anything touching doct [^15]: CatholicOS, _ontology-semantic-canon_, github.com/CatholicOS/ontology-semantic-canon, `queries/jena/01-church-hierarchy.rq` (`osc:RChKPk9K152BirrIYgAREsY` = Clergy) and `sources/ontology-semantic-canon.ttl`; ontokit, github.com/CatholicOS/ontokit-web, `lib/ontology/iriGeneration.ts` (`uuidToBase62`, prefixed `"R"` "to ensure RDF/XML QName - safety"). The same repository already ships the affordance layer beside those opaque IDs: `lib/hooks/useIriLabels.ts` resolves a set of IRIs to `rdfs:label`s and label-joins - them into API responses (cached per project and branch), and the Turtle editor's Monaco hover provider — `registerHoverProvider` in `components/editor/TurtleEditor.tsx` — - renders `Label: ` beside the resolved full IRI on hover. + safety"). The same repository already ships the human-readable layer beside those machine-readable IDs: `lib/hooks/useIriLabels.ts` resolves a set of IRIs to `rdfs:label`s and + label-joins them into API responses (cached per project and branch), and the Turtle editor's Monaco hover provider — `registerHoverProvider` in + `components/editor/TurtleEditor.tsx` — renders `Label: ` beside the resolved full IRI on hover. [^16]: FOLIO (Federated Open Legal Information Ontology), https://folio.openlegalstandard.org. The concept _Murder_ resolves at diff --git a/scripts/build-standalone-html.sh b/scripts/build-standalone-html.sh index 5d8c0e5..fafd952 100755 --- a/scripts/build-standalone-html.sh +++ b/scripts/build-standalone-html.sh @@ -16,7 +16,7 @@ DOCS=( "research/fragmented-catholic-digital-governance.md|fragmented-catholic-digital-governance|Fragmented Catholic Digital Governance" "research/governance-as-code-catholic-technology.md|governance-as-code-catholic-technology|Governance-as-Code for Catholic Technology" "research/trusted-data-infrastructure-catholic-ministry.md|trusted-data-infrastructure-catholic-ministry|Trusted Data Infrastructure for Catholic Ministry" - "research/identifier-durability-opaque-canonical-iris.md|identifier-durability-opaque-canonical-iris|Identifier Durability and Opaque Canonical IRIs" + "research/identifier-durability-opaque-canonical-iris.md|identifier-durability-opaque-canonical-iris|Identifier Durability: Machine-Readable Canonical IRIs" "standards/overview.md|standards-overview|CDCF Standards Overview" "standards/committees.md|standards-committees|CDCF Standards Committees" ) From 7075ae690ee927c7a7f51b58d0ebe9f8f7d6bb15 Mon Sep 17 00:00:00 2001 From: damienriehl Date: Sun, 26 Jul 2026 16:40:04 -0500 Subject: [PATCH 4/6] =?UTF-8?q?docs:=20Draft=203=20=E2=80=94=20plain-Engli?= =?UTF-8?q?sh=20and-not-or,=20God=20the=20Son=20exemplar,=20dispute-powder?= =?UTF-8?q?=20argument,=20Prevost=20timeline,=20authorship=20trendline,=20?= =?UTF-8?q?Agile=20framing,=20all-parties=20scaling,=20costs-to-zero?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Co-Authored-By: Claude Fable 5 Claude-Session: https://claude.ai/code/session_01MoZuH8vKiwL2rsX3a1zSqc --- ...tifier-durability-opaque-canonical-iris.md | 133 +++++++++++++----- 1 file changed, 95 insertions(+), 38 deletions(-) diff --git a/research/identifier-durability-opaque-canonical-iris.md b/research/identifier-durability-opaque-canonical-iris.md index 455928e..cb350c6 100644 --- a/research/identifier-durability-opaque-canonical-iris.md +++ b/research/identifier-durability-opaque-canonical-iris.md @@ -3,7 +3,7 @@ | | | | :---------------- | :---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | **Document type** | Position paper — public comment | -| **Status** | Draft 2 — prepared for submission during the `draft-cdcf-catholic-uri-scheme-03` (v0.4.0) 60-day comment window | +| **Status** | Draft 3 — prepared for submission during the `draft-cdcf-catholic-uri-scheme-03` (v0.4.0) 60-day comment window | | **Relationship** | Responds to [draft-cdcf-identifier-rationale-00](../standards/drafts/cdcf-identifier-rationale-00.md) and [draft-cdcf-catholic-uri-scheme-03](../standards/drafts/cdcf-catholic-uri-scheme-03.md); informs the [CDCF Standards program](../standards/overview.md) | | **License** | CC BY 4.0 | @@ -56,8 +56,8 @@ - **All human readability moves into a layer the standard guarantees** rather than one the identifier improvises: every existing slug and registry key preserved as a **permanent resolvable alias**, never deprecated and never reused; **multilingual labels** on every entity, so naming disputes are settled by _adding_ a label rather than _changing_ an identifier; and a **rendering rule** requiring production surfaces to show a label beside every canonical ID. -- **The frame is Catholic: _et-et_, not _aut-aut_.** The question is not identifier stability _or_ human interpretability. It is both — each carried in the layer built for it, - the stable identifier canonical and the human-readable layer guaranteed rather than optional. +- **The frame is "and," not "or"** (the classical Catholic _et-et_: both together, not one at the other's expense). The question is not identifier stability _or_ human + interpretability. It is both — each carried in the layer built for it, the stable identifier canonical and the human-readable layer guaranteed rather than optional. - **A small delta, not a rival document.** Draft 0.4.0 already provides the machinery — the two-artifact model (§3.6), the `notations` array (§5.9), the never-reassign stability guarantee (§3.4), the `exact-match`/`close-match` relations (§4.8.2). The proposal keeps all of it and inverts which artifact is primary. @@ -75,17 +75,18 @@ osc:Rh895FOUgLMfrylaO713w1 # "God the Son"@en · "Fílius Dei"@la · "Dio Figli **Registry source file.** The row a canonist reviews — plain text, ID and label side by side, with the familiar slug preserved forever. ```text -id: R7kQp2mXf4LdTbz9Ns3Hc1 # canonical — minted once, never re-minted -alias: mr:0629-petri-et-pauli # permanent — resolves forever, never reused -labels: "Sanctorum Petri et Pauli, Apostolorum"@la · - "Santi Pietro e Paolo, apostoli"@it · - "Sts. Peter & Paul, Apostles"@en +id: Rh895FOUgLMfrylaO713w1 # canonical — minted once, never re-minted +alias: god-the-son # permanent — resolves forever, never reused +labels: "Fílius Dei"@la · + "Dio Figlio"@it · + "God the Son"@en +altLabels: "Lógos"@grc ``` **Production resolution.** The response an application receives — one request, the reader's own language, every legacy key still on the record. ```http -GET /R7kQp2mXf4LdTbz9Ns3Hc1 HTTP/1.1 +GET /Rh895FOUgLMfrylaO713w1 HTTP/1.1 Host: id.catholiccommons.org Accept: application/ld+json Accept-Language: it @@ -94,24 +95,24 @@ HTTP/1.1 200 OK Content-Type: application/ld+json { - "@id": "R7kQp2mXf4LdTbz9Ns3Hc1", - "prefLabel": "Santi Pietro e Paolo, apostoli", - "altLabel": ["Sanctorum Petri et Pauli, Apostolorum", "Sts. Peter & Paul, Apostles"], + "@id": "Rh895FOUgLMfrylaO713w1", + "prefLabel": "Dio Figlio", + "altLabel": ["Fílius Dei", "God the Son", "Lógos"], "notations": [ - { "value": "mr:0629-petri-et-pauli", "scheme": "cdcf:scheme/mr", "prefLabel": true }, - { "value": "StsPeterPaulAp", "scheme": "cdcf:scheme/litcal" }, - { "value": "saints_peter_and_paul_apostles", "scheme": "cdcf:scheme/romcal" } + { "value": "god-the-son", "scheme": "cdcf:scheme/osc", "prefLabel": true }, + { "value": "filius-dei", "scheme": "cdcf:scheme/osc-la" }, + { "value": "logos", "scheme": "cdcf:scheme/osc-grc" } ] } ``` -`notations` is 0.4.0's field as written (§5.9), carrying the litcal and romcal production keys unchanged; the language-tagged `prefLabel`/`altLabel` are the one addition the -proposal asks for (commitment 3, amendment to §5.9). +`notations` is 0.4.0's field as written (§5.9), carrying every hand-authored key the entity has ever been given — here the English, Latin, and Greek slugs, none of them privileged +in the identifier and none of them lost; the language-tagged `prefLabel`/`altLabel` are the one addition the proposal asks for (commitment 3, amendment to §5.9). **Editor, IDE, and AI agent.** The call site a developer or an agent actually reads — a named constant, with the label supplied by the tooling. ```text -if (celebration === IDs.PETER_AND_PAUL) { // = "R7kQp2…" · hover: Peter and Paul, Apostles +if (concept === IDs.GOD_THE_SON) { // = "Rh895F…" · hover: God the Son ``` That last surface is not aspirational: ontokit renders exactly this today, resolving a full IRI to its label in a Monaco hover and label-joining API responses beside machine-readable @@ -163,8 +164,8 @@ the doc is right that opacity-to-reasoners does not entail machine-readable mint and neutrality, not reasoner correctness. **Draft 0.4.0's two-artifact model is convergence, not conflict.** The most important thing in the 0.4.0 revision is §3.6: the recognition that graph identity and public citation -are two jobs, and that both can be carried on one entity. That recognition is the same _et-et_ instinct this paper builds on. The `notations` array (§5.9), the scheme URNs, the -never-reassign stability guarantee (§3.4), and the `exact-match`/`close-match` relations that correctly refuse blanket `owl:sameAs` (§4.8.2, §3.6.1) are precisely the +are two jobs, and that both can be carried on one entity. That recognition is the same "and," not "or" instinct this paper builds on. The `notations` array (§5.9), the scheme +URNs, the never-reassign stability guarantee (§3.4), and the `exact-match`/`close-match` relations that correctly refuse blanket `owl:sameAs` (§4.8.2, §3.6.1) are precisely the infrastructure a machine-readable-primary architecture needs. **This proposal reuses all of it.** The committee is invited to build nothing it has not already specified — only to decide which artifact carries the stability guarantee. @@ -232,13 +233,32 @@ _Francis_ in its first bulletin, and the press office's own gloss was that "it w Church that the Church declined to assert. Meanwhile Open Issue 1 (Psalm numbering) and Open Issue 5 (multilingual slugs) have survived four revisions unresolved (§10) — both disputes that exist _only because_ the identifier string must choose. -Every one of these is a governance-quality problem a diligent committee will keep solving. A committee member may fairly object that draft registries are _supposed_ to churn — that -is what a pre-1.0 registry is for. Two answers. First, litcal, romcal, and ePrex are not drafts; they are production systems, and the St Martha divergence happened in their -_shipped_ keys — a 2021 decree left `StMartha` frozen factually wrong in a production API while a sibling project renamed to a 56-character key — so the mechanism operates after -normativity, not only before it. Second, the churn's causes — multilingual naming, movable anchors, editorial judgment about which name is _the_ name — do not end at 1.0: popes -keep being elected, dioceses keep merging, decrees keep expanding memorials. Normativity freezes the identifiers, not the world. That is precisely the cost the committee is asked -to weigh: transparency's maintenance burden is not a one-time migration but a standing obligation, recurring whenever the world, an edition, a decree, or a translation policy -changes. Under this proposal none of those events costs an identifier — each of them adds a label or a notation, which is what those layers are for. +A committee member may fairly object that draft registries are _supposed_ to churn — that is what a pre-1.0 registry is for. Two answers. First, litcal, romcal, and ePrex are not +drafts; they are production systems, and the St Martha divergence happened in their _shipped_ keys — a 2021 decree left `StMartha` frozen factually wrong in a production API while +a sibling project renamed to a 56-character key — so the mechanism operates after normativity, not only before it. Second, the churn's causes — multilingual naming, movable +anchors, editorial judgment about which name is _the_ name — do not end at 1.0: popes keep being elected, dioceses keep merging, decrees keep expanding memorials. **Normativity +freezes the identifiers, not the world.** Transparency's maintenance burden is therefore not a one-time migration but a standing obligation, recurring whenever the world, an +edition, a decree, or a translation policy changes. + +That standing obligation is usually costed as engineering time, and that is the smaller half of the bill. Every one of the disputes above is a governance-quality problem a +diligent committee will keep solving — and solving each one costs something scarcer than engineering time. Identifier disputes of this kind will not number in the dozens. Across +six registries, twenty-four churches _sui iuris_, and a goals ladder that reaches past Catholicism entirely, they will number in the thousands. Each one produces a winner and at +least one loser, and each loser's grievance is individually small, individually reasonable, and cumulative. Accumulated far enough, the grievance stops being about any particular +identifier and becomes a sentence a stakeholder says on the way out the door: _I'll leave the standard — it's too Latin-centric._ Or too English-centric, or too Roman, or too +Western. A standard has a finite supply of dispute powder, and it should be spent on the disputes that actually matter — the levels of authority, the doctrinal classifications, +the scope of the ontology, the terms of ecclesiastical review — not on what to call things. + +The design conclusion follows directly: make the identifiers **maximally neutral**, and make each identifier's concept **objective rather than subjective** — a referent +stakeholders can agree exists, rather than a name they must agree to use. Two results follow. Disputes resolve faster, because the question shrinks from _which name owns the +citation string for the next century_ to _which label carries `prefLabel` in which language_, which is a reversible call nobody has to win. And adoption widens, because the +standard's answer to "you have spelled our name wrong" stops being a change-control negotiation and becomes **"yes — added as a synonym."** That is what an adaptive, durable +standard sounds like from the outside. + +And there is a kicker, which is that when the identifier itself must be argued out with every stakeholder, even winning the argument produces the drift the argument was meant to +prevent. Suppose one reading prevails and the string is minted; a later revision, a new stakeholder, or a decree concedes the other form. The identifier is now deprecated, and the +concept wears a word that differs from the word in the string every earlier document cites — precisely the mismatch a transparent identifier exists to avoid. The minority loses +immediately and the majority loses on the next revision, so everyone loses. Under this proposal none of those events costs an identifier at all — each of them adds a label or a +notation, which is what those layers are for, and nobody has to lose anything. --- @@ -261,6 +281,16 @@ book; nothing published ever breaks. `StMartha`, `martha`, and the 56-character rather than a naming crisis in three projects. Aliases are supposed to pile up. Canonical identifiers are not. Note what this preserves: every legible string anyone has ever written keeps working, in the layer where legibility belongs. +**One human being, five names, one identifier.** The clearest case is a person now living. The reigning pope was born Robert Francis Prevost; he was _Bob_ to his brother, then Fr. +Prevost, then Bishop and later Cardinal Prevost, and is now Pope Leo XIV. Nothing about the human being changed at any of those moments except what people call him — and the +standard needs exactly one canonical identifier for him across the whole span. A human-readable identifier has to be minted at one point on that timeline, and it then has two +options, both bad. It can drift with each new name — `robert-prevost`, then `cardinal-prevost`, then `pope-leo-xiv` — breaking every citation minted before each change, three +times over one lifetime. Or it can stay frozen where it was minted, which leaves an identifier reading `robert-prevost` as the canonical string for the reigning pope: defensible +as history, plainly inaccurate in the present, and drawing exactly the ire the transparent string was supposed to avoid — the same failure mode `cdcf:person/pope-francis-i` +exhibits from the other direction, asserting an ordinal the Church declined to use. A machine-readable ID never faces that choice. It carries every name and every era at once, as +labels and notations, and every citation ever written keeps resolving. Objection 1 below shows the same mechanism absorbing the _other_ axis of the papal-name problem — Leo, +Leone, León — by the same move. + And the cost asymmetry has inverted since this trade-off was last argued seriously. The rationale doc calls the cost of opaque (durable) identifiers "permanent and unsolvable by design" — a mandatory lookup for every human who reads the identifier (§7). That was true in 2005. Hover labels, IDE inlays, resolver-backed link previews, and agents that never hand-type an identifier have collapsed the lookup cost toward zero, and CDCF's own resolution protocol is precisely the substrate those human-readable surfaces run on. @@ -269,6 +299,15 @@ its label, and its `useIriLabels` hook label-joins API responses.[^15] Drift cos edition, and every tradition the standard reaches for. **Legibility costs are falling toward zero; drift costs rise with adoption.** The transparent trade-off was right for 2005. It is not right for a standard minted in 2026 to last a century. +**The trajectory is visible in who writes the code.** Ask the question in rough terms, since the exact figures are not the point. Of the code written in 2025, how much was drafted +and reviewed by a human being — nearly all of it? By early 2026, perhaps two-thirds. By mid-2026, perhaps a quarter or less. What will the human-drafted, human-reviewed share be +in July 2027? In 2037? Those numbers are illustrative estimates of a visible trajectory rather than measurements, and nothing here depends on any one of them being right; the +argument depends only on the direction, which is not seriously in dispute. It has a design consequence. A standard minted in 2026 and intended to last a century should optimize +for what stays hard, not for what is already being solved. What is already being solved is interpretability: IDEs, resolvers, and agents render a label in the reader's own +language and rite, on demand, at the moment of reading, and they will do it more thoroughly every year. What stays hard is durability across a century of world-change, and +flexibility among contentious stakeholders who disagree with each other and are not finished arriving. Nothing but architecture has ever solved that second problem, and +optimizing the canonical layer for a reading cost that is falling spends the architecture on the wrong half. + --- ## Worked Example: Canonicalizing Levels of Authority @@ -295,28 +334,39 @@ and all three existing vocabularies carried as notations, the theology can devel label changes, the tagged corpus does not move, and the crosswalk between the who-teaches and what-assent axes becomes a property rather than a renaming. That is the whole argument in one case. **The identifier architecture should let the theology be revised. It should not require the theology to be finished first.** +That is a claim about method as much as about identifiers. Under a waterfall model — the Latin Church, all twenty-three Eastern churches _sui iuris_, and every episcopal +conference reviewing and signing off on the vocabulary before anything is built — human-readable identifiers would carry considerably less risk, because the names would be settled +before the first string was minted. Freeze the requirements, then mint. That is not how this standard is being built. It is being built iteratively: artifacts ship, stakeholders +read them, and the next revision absorbs what they said — which means stakeholders weigh in _after_ the artifacts exist, not before, and the set of stakeholders is itself still +growing as the goals ladder extends. Iterative building requires exactly two properties of its identifiers. **Durability:** nothing already published breaks when a later +stakeholder arrives with a correction. **Flexibility:** the new stakeholder's name for a thing arrives as a synonym rather than as a rename. A machine-readable canonical ID with a +guaranteed label layer supplies both. A transparent string, minted before the stakeholders who will argue about it have arrived, supplies neither. + --- ## The Proposal -Five commitments. Nothing here replaces draft 0.4.0's machinery; each item names the 0.4.0 mechanism it rides on. The shape of the whole is _et-et_: not _aut-aut_ — a durable -identifier _or_ an interpretable record — but both, the machine-readable ID canonical and the human-readable layer guaranteed by the standard rather than left to each +Five commitments. Nothing here replaces draft 0.4.0's machinery; each item names the 0.4.0 mechanism it rides on. The shape of the whole is "and," not "or" — not a durable +identifier _or_ an interpretable record, but both, the machine-readable ID canonical and the human-readable layer guaranteed by the standard rather than left to each implementer's good intentions. 1. **Canonical identifiers are machine-readable, `R`-shaped, and org-wide.** Every canonical ID CatholicOS mints — ontology and every data registry — is a base62-encoded 128-bit UUID prefixed `R` (23 characters in practice; 122 random bits, so decentralized minting needs no counter and no central allocator, and the leading letter keeps it QName-safe for RDF/XML). This is not new for CDCF: it is the shape already shipping in `ontology-semantic-canon`, where `osc:RChKPk9K152BirrIYgAREsY` is _Clergy_, and the shape ontokit already mints.[^15] It is also FOLIO's shape — `folio.openlegalstandard.org/R7Ttdyo4FsvaupPKT35Qry0` is _Murder_.[^16] -2. **Every existing slug, key, and registry prefix becomes a permanent resolvable alias.** `mr:`, `circ:`, `icl:`, CLEDR keys, CLBDR edition IDs, litcal/romcal/eprex keys, +2. **Every existing slug, key, and registry prefix becomes a permanent resolvable alias.** `mr:0629-petri-et-pauli`, `circ:`, `icl:`, CLEDR keys, CLBDR edition IDs, + litcal/romcal/eprex keys, `cdcf:verse/jn/1/14`, `cdcf:concept/C0000418` — all carried as 0.4.0 `notations` (§5.9) under declared scheme URNs, resolvable per §3.4, **never deprecated and never reused**. This is the IP/DNS model, and the analogy is worth stating plainly: nobody argues that an IP address should be human-readable, and nobody has to, because the domain name resolves to it. No adopter loses a working key. That is the adoption story. -3. **Every entity carries multilingual labels.** `skos:prefLabel` and `skos:altLabel`, language-tagged, on every entity in every registry — `"Sanctorum Petri et Pauli, -Apostolorum"@la`, `"Santi Pietro e Paolo, apostoli"@it`, `"Sts. Peter & Paul, Apostles"@en` on one record. Naming disputes are resolved by **adding a label**, never by changing - an identifier. CRMEDR's `i18n/{la,it,en}.json` is the existing precedent; this generalizes it. +3. **Every entity carries multilingual labels.** `skos:prefLabel` and `skos:altLabel`, language-tagged, on every entity in every registry — `"Fílius Dei"@la`, `"Dio Figlio"@it`, + `"God the Son"@en`, `"Lógos"@grc` on one record. Naming disputes are resolved by **adding a label**, never by changing an identifier. CRMEDR's `i18n/{la,it,en}.json` is the + existing precedent; this generalizes it. 4. **Production surfaces MUST render a label beside every canonical ID.** Registry source files carry a label column or comment beside each ID; serializations carry the label - inline; UIs and generated code render it. This commitment is the _et_ that makes the other _et_ livable, and it belongs in the standard rather than in each implementer's good - intentions. + inline; UIs and generated code render it. The rule is baked into the standard's documentation and into its coding guidelines — the conventions agentic coders read and follow — + so compliance does not depend on each implementer remembering it; and as agentic coding becomes prevalent, every production surface will carry its human-readable labels, in the + user's preferred language and rite, as a matter of course. This commitment is the half of the "and" that makes the other half livable, and it belongs in the standard rather + than in each implementer's good intentions. 5. **Structural and volatile facts live in properties, never in canonical identifiers.** Dates and calendar position (CRMEDR's `MMDD` anchor), country codes (CECDR's ISO 3166 prefix), taxonomy supertypes (`circumscription`, `institute`, the retired `order`), edition years (CLBDR), chapter/verse hierarchy, ownership, and language are all already modeled as fields in 0.4.0 responses. Where a fact lives both in a field and in the identifier, the identifier is the copy that goes stale. @@ -342,8 +392,10 @@ claim, and for two of the three it is the wrong one. But every tradition can sha label, definition, and typed statements to it — the very record shown in the first exemplar above. A machine-readable name is not merely tolerable at the inter-tradition boundary; **it is the only thing that crosses it, and it is what lets each tradition keep its own language on its own side.** That is the interfaith-alignment benefit named in the goals ladder, and it is concrete rather than aspirational: shape-compatibility with FOLIO today, and a mechanism that scales to machine-readable identifier work across the Abrahamic -traditions tomorrow. The alternative — a mapping table between every pair of standards, maintained by both parties in perpetuity — is the cost transparency imposes at exactly the -boundary where the goals ladder can least afford it. +traditions tomorrow. The alternative — a mapping table between every pair of standards, maintained by all parties in perpetuity — is the cost transparency imposes at exactly the +boundary where the goals ladder can least afford it. And the peer-standard set is not a pair. It is many ontologies at once — Catholic, Protestant, and Jewish standards, and legal +standards such as FOLIO and SALI, among others, with more arriving as the ladder extends — so pairwise mapping tables scale quadratically in the number of standards, while a +shared machine-readable local name scales not at all. --- @@ -363,6 +415,11 @@ None of these remedies is speculative — each exists in shipped systems, and tw committee, and especially the transparent camp, is invited to name costs this paper has missed and mitigations it has not thought of.** If a cost turns out to have no adequate remedy, that is a finding worth having before adoption rather than after. +One note on all four rows together. Every cost above is a cost paid by a human being handling raw identifiers by hand — pasting them, scanning them in a diff, reading them in a +log — and that is the population shrinking fastest. As agentic coding becomes prevalent, an agent neither hand-pastes an identifier nor reads a diff unaided: it resolves the ID +and renders the label, in the reader's language, as an ordinary part of doing the work. Each row therefore trends toward zero on the same trajectory described in §5. The costs are +real today and should be weighed as real today — but they are shrinking, which is the exact opposite of the drift costs a transparent canonical layer accumulates. + --- ## Objections at Full Strength — and Their Remedies @@ -397,7 +454,7 @@ those. Most humans who ever touch these identifiers are application developers, **Failure mode.** The legibility is real but bound in the wrong place: in a literal, at every call site. When the referent's boundary or preferred name changes, every call site is stale in a way no compiler catches. The string was never checked against the record; it was trusted because it looked right. -**Mitigation and invitation.** Named constants bound to canonical IDs — `const PETER_AND_PAUL = "R7kQp2mXf4LdTbz9Ns3Hc1"` — which is how every codebase already handles hex colors, +**Mitigation and invitation.** Named constants bound to canonical IDs — `const GOD_THE_SON = "Rh895FOUgLMfrylaO713w1"` — which is how every codebase already handles hex colors, port numbers, and country codes. That restores full call-site legibility while leaving exactly one authoritative binding to audit, and the rendering rule (§7, item 4) extends it to registry source files, serializations, and generated code so a label always sits beside the ID — the fourth exemplar in the executive summary is the shipped version of this, in ontokit today. If the constant-binding overhead is too high for a particular consumer, that workflow would be worth seeing, so the human-readable layer can be shaped around it. From fb7d20eaabcf7fd0fb340754a2556c5675c5ceb3 Mon Sep 17 00:00:00 2001 From: damienriehl Date: Sun, 26 Jul 2026 17:09:37 -0500 Subject: [PATCH 5/6] =?UTF-8?q?docs:=20Draft=204=20=E2=80=94=20two-directi?= =?UTF-8?q?onal=20IDE=20exemplar;=20production=20break-or-drift=20conseque?= =?UTF-8?q?nce=20in=20executive=20summary?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Co-Authored-By: Claude Fable 5 Claude-Session: https://claude.ai/code/session_01MoZuH8vKiwL2rsX3a1zSqc --- ...entifier-durability-opaque-canonical-iris.md | 17 +++++++++++++---- 1 file changed, 13 insertions(+), 4 deletions(-) diff --git a/research/identifier-durability-opaque-canonical-iris.md b/research/identifier-durability-opaque-canonical-iris.md index cb350c6..9e0c4cd 100644 --- a/research/identifier-durability-opaque-canonical-iris.md +++ b/research/identifier-durability-opaque-canonical-iris.md @@ -3,7 +3,7 @@ | | | | :---------------- | :---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | **Document type** | Position paper — public comment | -| **Status** | Draft 3 — prepared for submission during the `draft-cdcf-catholic-uri-scheme-03` (v0.4.0) 60-day comment window | +| **Status** | Draft 4 — prepared for submission during the `draft-cdcf-catholic-uri-scheme-03` (v0.4.0) 60-day comment window | | **Relationship** | Responds to [draft-cdcf-identifier-rationale-00](../standards/drafts/cdcf-identifier-rationale-00.md) and [draft-cdcf-catholic-uri-scheme-03](../standards/drafts/cdcf-catholic-uri-scheme-03.md); informs the [CDCF Standards program](../standards/overview.md) | | **License** | CC BY 4.0 | @@ -46,6 +46,13 @@ - **The divergence is already recorded,** inside CDCF's and CatholicOS's own repositories: one verse now carries three committee-minted transparent identifiers in eight months, shipped registry IDs have been renamed in place, and a calendar-date anchor baked into an identifier moved because the Latin and Italian editions of one book disagree about the date. + - **The consequence lands downstream.** Had any Catholic system already put those identifiers into production — a diocesan database, a liturgical app, a publisher's citation + layer — one of exactly two things happens. + - **The system breaks.** The identifier it stored changed underneath it, so the citation, the query, or the join simply stops resolving. + - **Or the system drifts.** It keeps running against an identifier the registry has abandoned, quietly serving a date, a name, or a spelling that is no longer the standard's + answer — the worse failure, because nothing reports it. + - **And that is fatal for a nascent standard.** Across six registries and a corpus this size it is bound to happen thousands of times, and each occurrence buys a sentence an + adopter says on the way out: _I tried the CDCF standard, but the IDs kept changing_ — or _I tried the CDCF standard, but it favored one language and phrasing over mine._ - **The pressure never ends.** Popes keep being elected, dioceses keep merging, decrees keep expanding memorials, editions keep being revised. Every one of those events asks a transparent identifier to change something it promised never to change. @@ -109,10 +116,12 @@ Content-Type: application/ld+json `notations` is 0.4.0's field as written (§5.9), carrying every hand-authored key the entity has ever been given — here the English, Latin, and Greek slugs, none of them privileged in the identifier and none of them lost; the language-tagged `prefLabel`/`altLabel` are the one addition the proposal asks for (commitment 3, amendment to §5.9). -**Editor, IDE, and AI agent.** The call site a developer or an agent actually reads — a named constant, with the label supplied by the tooling. +**Editor, IDE, and AI agent.** The call site a developer or an agent actually reads — keyed either way, by the human-readable constant or by the canonical ID, because the +tooling supplies whichever half the code does not carry; the reader sees both at once, the human-explainable name _and_ the durable machine name. ```text -if (concept === IDs.GOD_THE_SON) { // = "Rh895F…" · hover: God the Son +if (concept === IDs.GOD_THE_SON) { // = "Rh895FOUgLMfrylaO713w1" · hover: God the Son +if (concept === IDs.Rh895FOUgLMfrylaO713w1) { // = GOD_THE_SON · hover: God the Son ``` That last surface is not aspirational: ontokit renders exactly this today, resolving a full IRI to its label in a Monaco hover and label-joining API responses beside machine-readable @@ -206,7 +215,7 @@ privileged — and the divergence would have been visible as what it is: six cit `mr:0308-vincentius-kadlubek` and nine siblings), stating the policy: "IDs are drafts pending committee review, so renamed in place with no deprecated-alias." Most instructive is the moved anchor — `mr:1210-marcus-antonius-durando` became `mr:0610-marcus-antonius-durando`, because the Latin _editio altera_ 2004 places the blessed on June 10 while the Italian (CEI) edition of the same book places him on December 10.[^4] The identifier hard-codes a fact that two editions of one work disagree about, so it must move whenever the -anchor edition is reconsidered. +anchor edition is reconsidered — and any downstream system already holding the old anchor breaks or drifts, as the executive summary notes. **CLEDR**'s crosswalk records the same phenomenon across projects. On 26 January 2021 the Congregation for Divine Worship decreed that 29 July be designated the Memorial of Saints Martha, Mary and Lazarus, replacing the celebration of Martha alone.[^7] CLEDR's row carries the Latin title _Sanctorum Marthæ, Mariæ et Lazari_ — and three irreconcilable keys: From 3a14fc5dd13079f2a2869fc1f7e2a110a4b3fb9a Mon Sep 17 00:00:00 2001 From: damienriehl Date: Sun, 26 Jul 2026 17:21:42 -0500 Subject: [PATCH 6/6] docs: nest break/drift sub-bullets under the downstream-consequence bullet Co-Authored-By: Claude Fable 5 Claude-Session: https://claude.ai/code/session_01MoZuH8vKiwL2rsX3a1zSqc --- research/identifier-durability-opaque-canonical-iris.md | 6 +++--- 1 file changed, 3 insertions(+), 3 deletions(-) diff --git a/research/identifier-durability-opaque-canonical-iris.md b/research/identifier-durability-opaque-canonical-iris.md index 9e0c4cd..a0c0480 100644 --- a/research/identifier-durability-opaque-canonical-iris.md +++ b/research/identifier-durability-opaque-canonical-iris.md @@ -48,9 +48,9 @@ date. - **The consequence lands downstream.** Had any Catholic system already put those identifiers into production — a diocesan database, a liturgical app, a publisher's citation layer — one of exactly two things happens. - - **The system breaks.** The identifier it stored changed underneath it, so the citation, the query, or the join simply stops resolving. - - **Or the system drifts.** It keeps running against an identifier the registry has abandoned, quietly serving a date, a name, or a spelling that is no longer the standard's - answer — the worse failure, because nothing reports it. + - **The system breaks.** The identifier it stored changed underneath it, so the citation, the query, or the join simply stops resolving. + - **Or the system drifts.** It keeps running against an identifier the registry has abandoned, quietly serving a date, a name, or a spelling that is no longer the standard's + answer — the worse failure, because nothing reports it. - **And that is fatal for a nascent standard.** Across six registries and a corpus this size it is bound to happen thousands of times, and each occurrence buys a sentence an adopter says on the way out: _I tried the CDCF standard, but the IDs kept changing_ — or _I tried the CDCF standard, but it favored one language and phrasing over mine._ - **The pressure never ends.** Popes keep being elected, dioceses keep merging, decrees keep expanding memorials, editions keep being revised. Every one of those events asks a