IMPORTANT: Developer documentation for the current development branch. This content is unreleased, may change without notice, and must not be treated as Buildish release documentation.
Flexible component publication model
The central design choice is to separate component identity from public routing:
slugis a stable internal identifier,- component repositories own content and lifecycle metadata,
- the consumer-owned catalog owns publication layout, and
- the pipeline resolves final publication paths and exposes them to renderers.
Why change the model
The current contract assumes a shared publication shape centered on
/components/<slug>/....
That is simple, but it couples:
- component identity,
- grouping,
- public URL structure, and
- stage layout assumptions.
If the site needs:
- different groups under different path prefixes,
- some components under custom paths that do not match their group, or
- each component under an explicitly chosen path,
then publication should be modeled explicitly rather than inferred from slug.
Design goals
- let each component publish under an explicit path,
- support multi-host publication, including per-component hostnames,
- support groups as a reusable defaults layer,
- keep component repositories renderer-neutral,
- keep public path design under consumer control,
- expose resolved paths to renderers as authoritative metadata, and
- validate collisions and ambiguous routing strictly.
Recommended contract split
Component repository contract
Component-owned metadata should describe only component identity, content roots, and lifecycle hints.
Preferred structure for site/component.yaml:
1schemaVersion: 1
2component:
3 slug: spark
4 displayName: Apache Spark
5content:
6 pagesRoot: site/pages
7 docsRoot: site/docs
8 assetsRoot: site/assets
9lifecycle:
10 latestStable: v4.0.0
The component repository should not own its final public site path.
It should, however, be able to author repository-local content roots safely. In
practice that means a component with content.docsRoot can publish a moving
development docs surface even before any consumer models explicit artifacts. If
the consumer resolves that component to /components/site-pipeline/development/, the
component-owned docsRoot should populate that /development/ tree without first
inventing an artifact in site/catalog.yaml.
Why the contract is split across two files
New maintainers often expect one authored file to own everything. The pipeline intentionally does not work that way.
site/component.yaml answers “what is true about this repository no matter who
consumes it?” It owns stable component identity, repository-local content roots,
and optional lifecycle hints that remain meaningful across consuming sites.
site/catalog.yaml answers “what does this specific site want to publish from
the available repositories and sources?” It owns workspace source bindings,
publication layout, release or ref selection, artifact decomposition, and other
policy that can legitimately differ from one consumer site to another.
That split matters because one repository can be published in more than one way.
One site might publish Apache Spark docs under /components/spark/, while
another site might group the same repository under /analytics/spark/ and only
publish two selected release lines. Neither site should require the repository
itself to rewrite its metadata just to fit one consumer’s layout or selection
policy.
The same rule applies to source bindings. A local checkout path, a workspace overlay, or a consumer-specific source alias is not repository truth. Those are properties of the current build workspace, so they belong in the consumer catalog rather than in component-owned metadata.
Artifacts also stay in the consumer catalog on purpose. A component can begin
with only repository-owned docsRoot content, which lets one consumer publish a
single development surface immediately. Later, if a consumer needs independent
version selection, release visibility, or publication routing for multiple docs
surfaces, that consumer can model artifacts in site/catalog.yaml without
forcing every other consumer of the repository to adopt the same artifact split.
Consumer catalog contract
The consumer catalog should define both inventory and publication.
This is also where artifact decomposition belongs. That is initially surprising, but it follows the boundary above: artifacts affect source bindings, publication selection, release visibility, and final routing, so they are not just repository facts. They are part of what a specific consumer site chooses to publish from a repository without letting the repository silently redefine consumer-owned policy.
Recommended top-level fields in site/catalog.yaml:
schemaVersiondefaultssiteoriginssourcesgroupscomponents
Recommended site section fields:
pagesRootassetsRootvendorAssets
Recommended defaults additions:
publication.originpublication.developmentSegmentpublication.docsSegmentpublication.assetsSegment
Recommended origins model:
baseUrl- optional
canonical - optional
labels
Recommended groups model:
displayNamepathPrefixnavigationSectionweightpublication
Recommended sources model:
localDir- optional
repository - optional
defaultBranch - optional
metadataFile
Recommended per-component additions:
weightgroupcontent.sourcepublication.originpublication.pathSegmentpublication.mountPathpublication.componentPathpublication.developmentPathpublication.docsPathpublication.assetsPathartifacts[]
Example catalog
1schemaVersion: 2
2defaults:
3 metadataFile: site/component.yaml
4 pagesRoot: site/pages
5 docsRoot: site/docs
6 assetsRoot: site/assets
7 publication:
8 origin: main
9 developmentSegment: development
10 assetsSegment: assets
11site:
12 pagesRoot: site/root-pages
13 assetsRoot: site/root-assets
14 vendorAssets:
15 - source: vendor/project-brand
16 mountPath: /assets/vendor/project-brand/
17 kind: vendorStatic
18origins:
19 main:
20 baseUrl: https://www.example.org
21 products:
22 baseUrl: https://products.example.org
23 spark:
24 baseUrl: https://spark.example.org
25groups:
26 libraries:
27 displayName: Libraries
28 pathPrefix: /libraries/
29 navigationSection: Libraries
30 platforms:
31 displayName: Platforms
32 pathPrefix: /platform/
33 publication:
34 origin: products
35components:
36 - slug: spark
37 localDir: ../spark
38 group: libraries
39 publication:
40 origin: spark
41 mountPath: /
42 - slug: kafka
43 localDir: ../kafka
44 group: platforms
45 publication:
46 mountPath: /streaming/kafka/
47 - slug: camel
48 localDir: ../camel
49 publication:
50 mountPath: /integration/
In this model:
sparkpublishes on its own hostname athttps://spark.example.org/,kafkainherits theproductsorigin fromplatformsand publishes athttps://products.example.org/streaming/kafka/, andcamelpublishes on the default hostname athttps://www.example.org/integration/.
Hostname-aware publication
Once multiple hostnames are allowed, the publication target is no longer just a path. It becomes a route made of:
- an
origin, such ashttps://www.example.org, and - a public path, such as
/libraries/spark/.
That means the unique publication key is effectively (origin, path).
This model handles several consumer needs cleanly:
- many components on one hostname,
- one hostname per group,
- one hostname per component, and
- the same path reused on different hostnames without collision.
The pipeline should therefore model reusable origins explicitly rather than copying raw hostnames onto every component.
Independent release artifacts
Some components do not have one release cadence. They have several independently versioned deliverables, each with its own tags, support window, and sometimes its own repository.
That means lifecycle and versioning should not be modeled only at the component level. The contract needs one more layer:
- a
componentis the user-facing product or documentation space, - an
artifactis an independently versioned deliverable within that component, - a
sourceis the repository or checkout from which content and tags are read.
In this model, a component may:
- use one source for shared landing pages,
- use several artifact sources for versioned docs,
- have several
tagPatternvalues, one per artifact, and - mix monorepo and multi-repo release inputs.
Recommended artifact shape:
keydisplayNamesource- optional
docsRoot - optional
assetsRoot versioning.developmentRefversioning.tagPattern- optional
versioning.namedRefs[] - optional
publicationSelection - optional
lifecycle.latestStable - optional
lifecycle.releaseLines - optional
lifecycle.releases[] - optional
lifecycle.supportStatusVocabulary
Example:
1sources:
2 spark-repo:
3 localDir: ../spark
4 operator-repo:
5 localDir: ../spark-k8s-operator
6components:
7 - slug: spark
8 content:
9 source: spark-repo
10 artifacts:
11 - key: runtime
12 displayName: Spark Runtime
13 source: spark-repo
14 docsRoot: docs/runtime
15 versioning:
16 developmentRef: main
17 tagPattern: ^v[0-9]+\.[0-9]+\.[0-9]+$
18 namedRefs:
19 - key: preview
20 ref: preview/docs
21 displayName: Preview
22 maturity: preview
23 publicationSelection:
24 namedRefs: [preview]
25 releases:
26 mode: latestPerLine
27 - key: kubernetes-operator
28 displayName: Spark Kubernetes Operator
29 source: operator-repo
30 docsRoot: docs
31 versioning:
32 developmentRef: main
33 tagPattern: ^operator-v[0-9]+\.[0-9]+\.[0-9]+$
With this structure, lifecycle metadata belongs primarily to artifacts. A component-level lifecycle can also be emitted by the pipeline as a summary, but it should be treated as derived convenience metadata rather than authored truth.
Publication visibility should stay separate from lifecycle meaning.
That means an artifact may also carry a publicationSelection policy used during
planning and staging. When both component and artifact policy are present, the
artifact policy should win.
Effective authored configuration should resolve predictably:
defaults<group<component<artifact- scalar values use the nearest defined value
- maps merge by key, with the nearer level winning per key
- arrays replace rather than concatenate
- absent values inherit; explicitly empty arrays or maps clear inherited values
To avoid terminology confusion, artifact in this document should mean an
independently versioned release unit or documentation stream, not every
downloadable file belonging to a release. Individual tarballs, signatures,
checksums, SBOMs, and attestations should usually hang off release records as
optional asset metadata rather than become top-level pipeline identities.
Compatibility relationships are first-class metadata
Release identity and support status are not enough for ecosystems such as Quarkus + Quarkiverse, Camel-family docs, or tool-to-server documentation.
The model should therefore allow explicit compatibility assertions between:
- components,
- artifacts,
- release lines,
- exact releases, or
- named API or protocol levels.
Those relationships should be emitted as aggregate metadata rather than inferred
from matching version strings. That keeps documentation statements such as
“extension line 3.15.x is compatible with platform line 3.15” separate from
release identity and separate from support posture.
Generated and imported documentation mounts
Many projects publish a mix of authored content and generated or imported reference trees under one version root.
That means the publication model should treat mounted subtrees as a first-class input, not as an awkward side channel.
Examples include:
- Javadoc or Scaladoc under
/api/java/ - generated Python reference content under
/api/python/ - imported OpenAPI or CLI reference bundles
The important distinction is that mounted subtrees share the same publication
system as authored pages without pretending they are identical in provenance or
indexing behavior. They should therefore carry an explicit trustClass that
distinguishes passive mounts from imported active browser content.
Release lines and support phases
Adding release lines such as 1.x and 1.1.x does add complexity, but it is the
useful kind of complexity. Without them, renderers can show exact versions, but
they cannot easily build support tables, grouped version selectors, or listings
such as “latest in 1.x” and “latest in 1.1.x”.
The clean model is:
- exact releases are immutable versions such as
1.1.7, - release lines are named groupings such as
1.xor1.1.x, - lines belong to artifacts, not components, and
- moving labels such as
stableorlateststay separate from release lines.
Hierarchical lines should be allowed. A release can belong to a narrower line and that line can point to a broader parent line.
Recommended release line shape:
keydisplayName- optional
parent latest- optional
supportStatus - optional
aliases
Example:
1artifacts:
2 - key: runtime
3 displayName: Spark Runtime
4 versioning:
5 tagPattern: ^v[0-9]+\.[0-9]+\.[0-9]+$
6 lifecycle:
7 supportStatusVocabulary:
8 active:
9 displayName: Active
10 order: 10
11 stable:
12 displayName: Stable
13 order: 20
14 bugfix-only:
15 displayName: Bugfix only
16 order: 30
17 security-fix-only:
18 displayName: Security fix only
19 order: 40
20 eol:
21 displayName: End of life
22 order: 90
23 releaseLines:
24 - key: 1.x
25 displayName: 1.x
26 latest: 1.4.3
27 supportStatus: security-fix-only
28 - key: 1.1.x
29 displayName: 1.1.x
30 parent: 1.x
31 latest: 1.1.9
32 supportStatus: eol
On naming: status versus something more specific
I would avoid a generic field name if this is meant to express maintenance or support posture. A better name is something like:
supportStatus, orsupportPhase
I would slightly prefer supportStatus because it reads naturally on a release
line and is easy for renderers to consume.
The important part is not the exact field name, though. The important part is that the pipeline should not hardcode a global enum for all projects.
Instead, the contract should allow a project-defined vocabulary. For example:
- a component may define a default
supportStatusVocabulary, - an artifact may override or narrow that vocabulary,
- each release line may optionally reference one vocabulary key, and
- if no support status is authored, the line simply has no status.
That gives projects room for values like active, stable, lts,
bugfix-only, security-fix-only, community-supported, or eol without
forcing the pipeline to pretend those terms are universal.
The pipeline’s job should be to validate references and preserve the vocabulary metadata, not to impose semantics beyond basic structure.
Authored named refs and publication selection
Intentional preview-style publications should be authored explicitly.
Recommended namedRefs[] fields under versioning:
keyref- optional
displayName - optional
maturity - optional
description
Those authored named refs become the stable identities used by planning, staging, and aggregate metadata. Provider data may enrich them, but provider records should not be required to create them.
Publication visibility should be controlled by a separate
publicationSelection policy rather than by supportStatus or release-line
semantics.
Recommended policy dimensions:
developmentlineHeadsreleasesnamedRefscandidates
Recommended built-in default:
- include development docs
- include all authored line heads
- include the latest stable release per release line
- include only explicitly selected authored named refs
- exclude release candidates unless explicitly requested
This keeps public visibility intentional without making route and lifecycle metadata carry planning semantics.
Example:
1 artifacts:
2 - key: runtime
3 versioning:
4 developmentRef: main
5 tagPattern: ^v[0-9]+\.[0-9]+\.[0-9]+$
6 namedRefs:
7 - key: preview
8 ref: preview/docs
9 displayName: Preview
10 maturity: preview
11 publicationSelection:
12 development: true
13 lineHeads:
14 mode: allAuthored
15 releases:
16 mode: latestPerLine
17 namedRefs: [preview]
18 candidates:
19 mode: none
Support windows should be structured, but optional
Some projects need more than a support-status key. They need dates and policy links that explain how long a line or release remains supported.
That is best modeled as a small structured support-window object attached to release lines and, when needed, exact releases.
Typical fields include:
releaseDatemaintenancePhaseendOfActiveSupportDateendOfSupportDateendOfLifeDatesupportPolicyUrl
This is mostly additive lifecycle metadata. It should enrich the current model, not replace the support-status vocabulary and not absorb compatibility matrices.
Exact-release publication state
Exact releases sometimes need publication behavior that differs from their support posture.
That should be modeled separately through exact-release entries under
lifecycle.releases[].
Recommended fields:
version- optional
releaseLine - optional
supportStatus - optional
supportWindow - optional
publicationState - optional
withdrawalBehavior - optional
redirectTarget - optional
reason
Recommended publicationState vocabulary:
publishedhiddenwithdrawntombstoned
Recommended withdrawalBehavior vocabulary:
noticeredirectomit
This keeps maintenance meaning and publication behavior separate:
supportStatussays how a release is maintainedpublicationStatesays whether and how it is publicly surfaced
For withdrawn or tombstoned releases, the model should support both:
- a notice or tombstone page at the preserved route, and
- an explicit redirect policy to another route or URL
Example:
1 artifacts:
2 - key: runtime
3 lifecycle:
4 releases:
5 - version: 4.1.0
6 publicationState: withdrawn
7 withdrawalBehavior: notice
8 reason: Recalled pending security fix.
9 - version: 3.2.0
10 publicationState: tombstoned
11 withdrawalBehavior: redirect
12 redirectTarget: /security/runtime/3.2.0/
Routing rules
Core rules
slugis never used implicitly to derive the public URL.- groups are optional and are not the source of truth for the final URL.
- the final resolved route is always explicit, even when derived from defaults.
- a route is the combination of resolved
originand resolved public path.
Resolution order
The pipeline should resolve the publication origin in this order:
- per-component
publication.origin - group
publication.origin defaults.publication.origin- validation failure if no origin can be resolved
The pipeline should resolve publication paths in this order:
- explicit per-component paths:
componentPath,developmentPath,docsPath,assetsPath - per-component
mountPath group.pathPrefixplus per-componentpublication.pathSegment- validation failure if no component mount can be resolved
Derived defaults
When only origin and mountPath are provided, the pipeline should derive:
componentPath = mountPathdevelopmentPath = componentPath + <developmentSegment>/docsPath = developmentPathunless an explicitdocsPathor extradocsSegmentoverride is configuredassetsPath = componentPath + <assetsSegment>/
It should also derive fully qualified URLs by joining each path to the resolved
origin baseUrl.
This keeps the common case simple while also allowing a fully explicit routing map when needed.
Canonical routes, aliases, and redirects
The route model should distinguish between:
- the published target,
- the canonical route for that target,
- alias routes that also resolve to the same target, and
- redirect routes that forward to another route or URL.
This matters for:
- moving labels such as
latestorstable - legacy path preservation after a site reorganization
- host migrations
- canonical URL generation for search and feeds
Redirects should be modeled as pipeline-owned route metadata, not primarily as page markup conventions in Markdown or AsciiDoc.
The pipeline should therefore emit:
- a complete route inventory, and
- a derived redirect inventory that downstream tools can use to generate Apache
httpd, Nginx, CDN, or other deployment-specific config.
Locale and translation as route dimensions
Locale should be an optional publication dimension that is orthogonal to component, artifact, and version context.
The model should support:
- no locale path transform (
none) - locale-prefixed paths such as
/fr/docs/(prefixAll) - default-locale behavior
- translation linkage between equivalent pages
- partial translation coverage when some pages exist in only one locale
Host-based locale publication is a later extension rather than part of the initial implementation contract.
Renderers should receive locale and translation data directly rather than trying to infer them from path prefixes alone.
Only one locale routing mode should be active for one effective publication. In
the initial implementation that means either none or prefixAll.
Translation linkage itself should be page-authored rather than catalog-authored.
Pages that belong to the same translation set should carry a shared
translationKey in authored page metadata. The pipeline should validate those
keys, derive locale sibling relationships, and emit data/translations.json
from the staged page set.
That keeps translation equivalence close to the pages that actually vary by locale and avoids a brittle central registry in the catalog model.
How groups should behave
Groups should be treated as a convenience layer rather than as a hard routing layer.
They are useful for:
- shared hostnames or origins,
- shared path prefixes,
- shared navigation defaults,
- shared weighting defaults, and
- coarse organization in the catalog.
They should not prevent a component from publishing under an arbitrary path.
External release providers
Site Pipeline should be able to consume release metadata from external systems, with Apache Trusted Releases (ATR) as one possible provider rather than a hard coded special case.
That argues for a provider boundary with three properties:
- the pipeline consumes a versioned, normalized snapshot,
- providers may attach extension data without reshaping the core contract, and
- downstream renderers consume staged metadata rather than querying providers directly.
This keeps the architecture flexible enough for ATR, Git hosting release APIs, package registries, or future foundation-wide tooling without turning the core pipeline into a provider-specific orchestration engine.
Integration model
The recommended architecture is:
- an external sync tool or provider adapter pulls provider data,
- it writes a normalized, versioned snapshot into a pipeline input location,
build()loads that snapshot during the resolve-and-plan phase,- the pipeline merges provider data with authored metadata,
- staged front matter and aggregate metadata are written from the merged view, and
- renderers consume the staged outputs only.
In this model, webhooks are useful as triggers for refreshing provider snapshots or starting builds, but they are not the authoritative source of release state.
That fits cleanly into the broader build architecture: provider data is just another resolved input during planning, not a special renderer-time dependency.
Normalized provider snapshot
The provider snapshot should have a small stable core and an explicit extension area.
Recommended top-level shape:
schemaVersionproviders[]records[]
Recommended providers[] fields:
keytype- optional
displayName - optional public
baseUrl fetchedAt
Recommended normalized records[] core fields:
providerkinddevelopmentnamedReflineHeadcandidatereleased
componentSlugartifactKey- optional
sourceKey - optional
externalId - optional
externalUrl - optional
version - optional
displayVersion - optional
tag - optional
ref - optional
namedRefKey - optional
commitSha - optional
releaseLine - optional
releaseLineAncestors - optional
supportStatus - optional
publicationState - optional
maturity - optional
candidateSequence - optional
voteStatus - optional
createdAt - optional
publishedAt - optional
updatedAt - optional
urls - optional
assets
Minimal normalized records[] schema
To keep provider integrations predictable without freezing the model too early, the pipeline should define a small minimum contract for every record.
Required for every record:
providerkindcomponentSlugartifactKey
Additionally, each record should provide at least one stable locator from this set:
externalId,version,tag, orref
Kind-specific expectations should remain simple:
released: should normally provideversionand usuallytagcandidate: should normally provideversion;candidateSequenceis recommended when the provider has onenamedRef: should providereflineHead: should providereleaseLineand usuallyrefdevelopment: should usually provideref
Recommended normalization rules:
kindshould come from the small shared vocabulary abovecomponentSlugandartifactKeyshould reference known pipeline identitiesversionanddisplayVersionmay differ when the provider exposes a friendly label such as4.1.0-rc2maturity,supportStatus, andvoteStatusshould be optional metadata, not required classification keysnamedRefKeyshould be present when a provider record enriches an authored named ref rather than introducing an ad hoc discovered ref- provider-specific details outside the normalized contract should stay out of the v1 public schema for now
This keeps the snapshot useful for indexing and rendering while also allowing a provider such as ATR to expose richer state over time.
Example minimum shape:
1records:
2 - provider: atr
3 kind: released
4 componentSlug: spark
5 artifactKey: runtime
6 externalId: atr:release:spark-runtime:4.0.0
7 version: 4.0.0
8 tag: v4.0.0
This is intentionally a normalized core rather than an exhaustive universal release schema. Providers may have richer source data, but renderers and pipeline logic should rely first on the stable core fields and keep unmatched detail out of the v1 public contract for now.
Example:
1schemaVersion: 1
2providers:
3 - key: atr
4 type: atr
5 displayName: Apache Trusted Releases
6 baseUrl: https://release-test.apache.org
7 fetchedAt: 2026-04-02T12:00:00Z
8records:
9 - provider: atr
10 kind: candidate
11 componentSlug: spark
12 artifactKey: runtime
13 externalId: atr:candidate:spark-runtime:4.1.0:2
14 externalUrl: https://release-test.apache.org/candidates/spark-runtime/4.1.0/2
15 version: 4.1.0
16 displayVersion: 4.1.0-rc2
17 maturity: rc
18 candidateSequence: 2
19 voteStatus: open
20 releaseLine: 4.x
21 - provider: atr
22 kind: namedRef
23 componentSlug: spark
24 artifactKey: runtime
25 ref: feature/docs-reorg
26 displayVersion: docs-reorg preview
27 maturity: preview
Merge rules
The merge boundary should be explicit.
Authored catalog and component metadata should remain authoritative for:
- publication routing,
- content roots,
- component and artifact identity,
- local grouping and navigation defaults, and
- project-defined support-status vocabularies.
Provider snapshots should be authoritative for externally observed release state, such as:
- discovered releases,
- release candidates and vote status,
- state for authored named refs or preview refs,
- timestamps,
- external URLs, and
- downloadable asset inventories.
Consumer-authored metadata should remain authoritative for:
- which named refs are intentionally publishable,
- publication-selection policy,
- exact-release publication behavior such as notice vs redirect, and
- final public route ownership.
The pipeline should not let provider data silently redefine consumer-owned URL layout or component identity. Conversely, authored metadata should not need to copy volatile provider state such as candidate numbers or vote windows.
When projects need to correct or enrich provider data locally, that should happen through an explicit override layer rather than by mutating the normalized provider snapshot in place.
How much release-file detail to model
This is where the contract can become messy very quickly.
My recommendation is to stop the core model at the level of:
- component,
- artifact,
- release line,
- named ref,
- release candidate, and
- exact release.
The core model should not make every downloadable file a first-class identity.
Instead, exact releases and candidates may optionally carry a generic assets[]
list for commonly useful file-level detail. A minimal generic asset shape could
include:
name- optional
kind url- optional
mediaType - optional
size - optional
checksums - optional
signatureUrl - optional
sbomUrl - optional
provenanceUrl
Minimal optional assets[] schema
assets[] should stay intentionally lightweight. It exists to power download
tables and release detail pages, not to model every package-management concept in
the core contract.
Required for every asset entry:
nameurl
Recommended optional fields:
kindmediaTypesizechecksumssignatureUrlsbomUrlprovenanceUrl
kind should remain advisory rather than exhaustive. Useful values might be:
archivesignaturechecksumsbomprovenancecontainer-imagepackage
For checksums, the normalized model should prefer a simple map keyed by algorithm, for example:
sha512sha256
Example:
1assets:
2 - name: spark-4.0.0-src.tgz
3 kind: archive
4 url: https://downloads.example.org/spark-4.0.0-src.tgz
5 mediaType: application/gzip
6 size: 123456789
7 checksums:
8 sha512: abcdef...
9 signatureUrl: https://downloads.example.org/spark-4.0.0-src.tgz.asc
10 sbomUrl: https://downloads.example.org/spark-4.0.0-src.spdx.json
That is enough for download listings and release detail pages without forcing the pipeline to understand every archive, signature, checksum, attestation, or registry object as its own top-level domain entity.
If a provider exposes richer package or distribution metadata, the pipeline should leave it out of the v1 public contract rather than inflate the core schema prematurely.
Renderer contract
The pipeline should expose resolved paths to renderers in staged metadata and front matter.
For multi-host publication, exposing only paths is no longer enough. A renderer or template often needs:
- the current component’s origin,
- the current page path relative to that origin,
- the fully qualified canonical URL,
- the component home URL,
- the docs root URL,
- the assets base URL, and
- enough structure to generate cross-links without recomputing routing rules.
The front matter should therefore expose both route parts and resolved URLs.
Pipeline-owned staged front matter should live under a reserved top-level
pipeline namespace. Authored page metadata remains outside that namespace, and
authored pages should fail validation if they attempt to define pipeline.
Front matter should also expose artifact context when a page belongs to one specific release stream.
Recommended component-level front matter shape:
1pipeline:
2 component:
3 slug: spark
4 displayName: Apache Spark
5 publication:
6 origin: { key: spark, baseUrl: https://spark.example.org, hostname: spark.example.org }
7 paths: { component: /, development: /development/, docs: /development/, assets: /assets/ }
8 urls: { component: https://spark.example.org/, development: https://spark.example.org/development/, docs: https://spark.example.org/development/, assets: https://spark.example.org/assets/ }
9 artifacts:
10 - { key: runtime, displayName: Spark Runtime, latestStable: 4.0.0, releaseLines: [{ key: 4.x, latest: 4.0.0, supportStatus: active }] }
11 - { key: kubernetes-operator, displayName: Spark Kubernetes Operator, latestStable: 1.3.0 }
Recommended page-level front matter shape:
1pipeline:
2 page:
3 kind: docsPage
4 section: docs
5 artifactKey: runtime
6 path: /releases/4.0.0/sql/
7 url: https://spark.example.org/releases/4.0.0/sql/
8 canonicalUrl: https://spark.example.org/releases/4.0.0/sql/
9 locale: en
10 defaultLocale: true
11 translationKey: runtime-sql-overview
12 componentPath: /
13 componentUrl: https://spark.example.org/
14 version:
15 kind: released
16 label: 4.0.0
17 path: /releases/4.0.0/
18 url: https://spark.example.org/releases/4.0.0/
19 docsPath: /releases/4.0.0/
20 docsUrl: https://spark.example.org/releases/4.0.0/
21 publicationState: published
22 releaseLine: { key: 4.x, supportStatus: active, ancestors: [], supportWindow: { maintenancePhase: active, endOfSupportDate: 2027-06-30T00:00:00Z } }
23 translations: [{ locale: fr, url: https://spark.example.org/fr/releases/4.0.0/sql/ }]
Front matter attribute guidance
The necessary attributes break down into a few categories:
- identity:
slug,displayName - artifact context:
artifactKey,artifacts[] - origin selection:
origin.key,origin.baseUrl,origin.hostname - public paths:
component,development,docs,assets - full URLs:
component,development,docs,assets - current page route: page
pathand pageurl - route semantics:
canonicalUrl, optional aliases, and redirect-aware routing - locale and translation context:
locale, default-locale state, and translated siblings derived from page-authoredtranslationKey - version context: development or released version label and URLs
- release-line context: current line, ancestor lines, optional support status, and optional support-window metadata
- release provenance:
provider,externalId,externalUrl - preview and candidate context:
ref,namedRefKey,maturity,candidateSequence, and optionalvoteStatus
For reliability, the pipeline should expose both:
- normalized public paths for renderer-relative logic, and
- absolute URLs for canonical links, feeds, sitemaps, redirects, and cross-host navigation.
Renderers should consume those resolved values directly and should not reconstruct
URLs from slug, group name, internal stage layout, or hostname conventions.
Front matter versus aggregate metadata
Front matter is necessary, but it is not sufficient for maximum renderer freedom.
Front matter is best for page-local context:
- what component and artifact the page belongs to,
- what the current resolved URLs are,
- what version context the page is in, and
- what links are immediately relevant to that page.
For global navigation, listings, landing pages, sitemaps, and cross-component UI, the pipeline should also emit normalized aggregate metadata files.
Recommended staged metadata outputs:
data/components.jsonfor component identity, summaries, groups, origins, and publication rootsdata/artifacts.jsonfor artifact-level lifecycle, tag patterns, source refs, latest stable versions, release lines, support-status vocabularies, and docs rootsdata/releases.jsonfor exact release records from authored and provider inputs, including publication state for hidden, withdrawn, or tombstoned releasesdata/candidates.jsonfor in-flight release candidates and vote-related statedata/refs.jsonfor development, maintenance, feature, and preview refs when they are intentionally exposed, including authored named ref keys when presentdata/routes.jsonfor resolved origin/path/url mappings and canonical route inventorydata/redirects.jsonfor resolved redirect inventory derived from route metadatadata/translations.jsonfor translation-set relationships and locale-specific sibling routesdata/compatibility.jsonfor cross-component, cross-artifact, or cross-line compatibility assertionsdata/mounts.jsonfor mounted generated or imported documentation subtreesdata/providers.jsonfor loaded provider snapshot provenance and fetch statedata/content-index.jsonfor one normalized entry per staged pagedata/diagnostics.jsonfor non-fatal warnings, skipped optional inputs, and stale-provider notices when present
Recommended content-index entry fields:
- stable page id
- component slug
- optional artifact key
- page kind and section
- origin key
- public path and absolute URL
- title, link title, description, summary
- weight and ordering hints
- parent id or ancestor ids derived from the staged content tree
- source file path within the staged tree
- version context
- release-line and support-status context
- support-window metadata when available
- locale and translation-set context when available
- provider provenance and optional maturity/candidate context
That gives a renderer enough information to build:
- arbitrary sidebars,
- drop-down menus from content structure,
- component and artifact listings,
- release selectors,
- release-candidate or preview listings,
- host-aware navigation, and
- locale switchers,
- compatibility tables,
- generated-reference mount sections, and
- alternate views such as cards, tables, or landing-page groupings.
The important distinction is:
- front matter powers rendering of the current page, and
- aggregate metadata powers global queries and information architecture.
Validation expectations
If publication becomes explicit, validation should also become explicit.
These are natural site-pipeline check failures and should be aggregated in one
validation pass where practical.
The pipeline should reject:
- undefined publication origins,
- non-absolute or non-canonical public paths,
- duplicate
(origin, path)publication routes, - duplicate canonical routes for the same published target,
- redirect loops or redirects to unknown internal targets,
- redirect targets with unsupported URL schemes,
- overlapping component mounts within the same origin,
- overlapping generated/imported mounts with conflicting ownership,
- duplicate artifact keys within a component,
- ambiguous artifact-to-route mappings,
- duplicate locale entries within one translation set,
- incompatible locale route-mode and origin configuration,
- duplicate authored named ref keys within one artifact,
- authored pages that define the reserved
pipelinefront matter namespace, - publication-selection policies that reference unknown release lines, named ref keys, or exact versions,
- compatibility assertions that reference unknown subjects or targets,
- conflicting tag patterns within one artifact definition,
- provider records that reference unknown components or artifact keys,
- duplicate provider records for the same
(provider, externalId)pair, - unknown support-status keys on release lines,
- withdrawn or tombstoned exact releases that declare
withdrawalBehavior: redirectwithout a redirect target, - malformed support-window date ordering,
- cycles or broken references in release-line parent chains,
- path-bearing authored or mounted inputs that escape declared roots after normalization and symlink resolution,
- collisions with consumer-authored site content,
- collisions between page, docs, and asset mounts for the same component when the result would be ambiguous, and
- any configuration that leaves a component without a resolvable public mount.
Detailed path-safety, XSS-defense, redirect-safety, and mounted-content trust rules are defined in security-and-trust-model.md.
Opinionated recommendation
If the project wants maximum flexibility without legacy constraints, the best contract is:
- keep
slugas identity only, - model publication as
origin + path, - support groups as defaults rather than as routing truth, and
- treat artifacts as the unit of release and lifecycle truth,
- ingest external release state through normalized provider snapshots, and
- keep file-level distribution details as optional metadata rather than core pipeline identity.
That model scales cleanly from a simple grouped site to a fully custom information architecture.