IMPORTANT: Developer documentation for the current development branch. This content is unreleased, may change without notice, and must not be treated as Buildish release documentation.
Provider snapshot to staged metadata mapping
provider-snapshot-schema.md by answering not just what a provider may supply, but what the pipeline should emit for renderers.Mapping principles
- authored metadata remains authoritative for identity, routing, content roots, publication selection, and exact-release publication behavior
- provider metadata enriches lifecycle, release, candidate, ref, and asset state
- page front matter should contain only page-local and immediately relevant context
- aggregate files in
data/should contain cross-page and cross-component query data - provider-specific details outside the normalized contract should stay out of the public staged outputs for now rather than being copied through wholesale
Staged outputs overview
Recommended staged outputs that may receive provider-derived data:
- page front matter
data/providers.jsondata/components.jsondata/artifacts.jsondata/releases.jsondata/candidates.jsondata/refs.jsondata/content-index.json
Serialization format recommendation
Page front matter commonly uses YAML, but aggregate metadata needs a more portable baseline for renderer and tooling integration.
In practice:
- Markdown front matter is commonly YAML across mainstream documentation tools
- aggregate data-file support varies more across renderers and site generators
- JSON is the safest baseline format for machine-consumed staged metadata
Recommended rule:
- keep page front matter YAML
- emit aggregate staged metadata in JSON as the contract format
- treat optional alternate serializations as non-authoritative mirrors
What goes into page front matter
Page front matter should expose only what is needed to render the current page and closely related navigation affordances.
Pipeline-owned staged values should live under the reserved top-level
pipeline.page namespace so authored page metadata and pipeline metadata do not
collide silently.
Recommended provider-derived fields in page front matter:
providerexternalIdexternalUrlversion.kindversion.label- optional
version.tag - optional
version.ref - optional
version.maturity - optional
version.releaseLine - optional
version.releaseLineAncestors - optional
version.supportStatus - optional
version.candidateSequence - optional
version.voteStatus
These values should be present only when the page belongs to a provider-backed release, candidate, or named ref context.
Example page-level shape:
1pipeline:
2 page:
3 artifactKey: runtime
4 version:
5 kind: candidate
6 label: 4.1.0-rc2
7 maturity: rc
8 releaseLine: 4.x
9 candidateSequence: 2
10 voteStatus: open
11 provider:
12 key: atr
13 externalId: atr:candidate:spark-runtime:4.1.0:2
14 externalUrl: https://release-test.apache.org/candidates/spark-runtime/4.1.0/2
What belongs in aggregate metadata
Anything that must be queried across pages, components, or artifacts should be
written to data/ files instead of repeated into front matter.
data/providers.json
Purpose:
- inventory of loaded providers
- fetch timestamps and provenance
- diagnostics for stale or missing provider data
Recommended fields:
keytype- optional
displayName - optional public
baseUrlwhen safe to disclose fetchedAt
data/components.json
Provider data should influence component summaries only when it is useful to show derived lifecycle state at component level.
Recommended provider-derived additions:
- latest published release per artifact
- latest candidate per artifact
- available provider keys for the component
The full release inventory should not be duplicated here.
data/artifacts.json
This should be the main aggregate view for artifact-level lifecycle and version state.
Recommended provider-derived fields:
providerKeyslatestReleaselatestCandidatereleaseLines[]namedRefs[]
Each releaseLines[] entry may include:
key- optional
parent - optional
latest - optional
supportStatus - optional
headRef
data/releases.json
Purpose:
- one normalized entry per exact released version
- source for release listings, version selectors, and download pages
Recommended fields:
providerexternalIdexternalUrlcomponentSlugartifactKeyversion- optional
displayVersion - optional
tag - optional
releaseLine - optional
releaseLineAncestors - optional
supportStatus - optional
publicationState - optional
withdrawalBehavior - optional
redirectTarget - optional
maturity - optional
publishedAt - optional
assets - optional
urls
data/candidates.json
Purpose:
- one normalized entry per in-flight or historical release candidate
- source for vote pages, candidate listings, and release-vote UIs
Recommended fields:
providerexternalIdexternalUrlcomponentSlugartifactKeyversion- optional
displayVersion - optional
candidateSequence - optional
releaseLine - optional
maturity - optional
voteStatus - optional
createdAt - optional
publishedAt - optional
assets
data/refs.json
Purpose:
- moving refs intentionally exposed to renderers
- development, maintenance, feature, and preview refs
Recommended fields:
providercomponentSlugartifactKeykind- optional
namedRefKey ref- optional
displayVersion - optional
releaseLine - optional
maturity - optional
externalUrl
data/content-index.json
The content index should denormalize the most useful provider context for each page so renderers can query page collections without joining multiple files.
Recommended provider-derived additions:
provider- optional
externalId - optional
versionKind - optional
versionLabel - optional
releaseLine - optional
supportStatus - optional
maturity - optional
candidateSequence - optional
voteStatus
Mapping by record kind
Recommended mapping behavior:
released- emit entry in
data/releases.json - merge authored exact-release publication state and withdrawal behavior when present
- enrich
data/artifacts.jsonlatest release and line summaries - copy page-local release context into front matter for pages staged from that release
- emit entry in
candidate- emit entry in
data/candidates.json - enrich
data/artifacts.jsonlatest candidate summary - copy page-local candidate context into front matter for candidate pages
- emit entry in
namedRef- emit entry in
data/refs.json - carry the authored
namedRefKeywhen the ref matches one - copy ref context into front matter for pages staged from that ref
- emit entry in
lineHead- emit entry in
data/refs.json - enrich matching release-line summaries in
data/artifacts.json - copy
lineHeadcontext into front matter when applicable
- emit entry in
development- emit entry in
data/refs.json - mark artifact development context in
data/artifacts.json - copy development ref context into front matter for development pages
- emit entry in
What should not be copied into front matter
To keep front matter compact, the pipeline should avoid embedding:
- full release inventories
- full candidate histories
- every downloadable asset for unrelated versions
- complete provider payloads
- large provider-specific passthrough payloads
Only the normalized subset of those inventories belongs in aggregate files. Raw provider payloads and passthrough-only fields should stay out of the public staged contract.
Provider-specific passthrough data
When provider data includes extra fields that do not fit the normalized staged contract, the pipeline should not copy them into public staged metadata in v1.
Recommended rule:
- normalize only the documented shared fields into
data/*.json - do not copy provider-specific passthrough payloads into page front matter
- do not expose private or API-only provider endpoints via public staged outputs
Validation expectations
The pipeline should reject or warn on at least:
- provider-backed pages that refer to no matching normalized record
- multiple normalized records competing for the same staged page context
- provider records mapped to the wrong artifact or component
- provider-specific passthrough fields copied into public staged outputs
- large provider payloads copied wholesale into front matter
Worked example
Given this provider record:
1- provider: atr
2 kind: candidate
3 componentSlug: spark
4 artifactKey: runtime
5 version: 4.1.0
6 displayVersion: 4.1.0-rc2
7 releaseLine: 4.x
8 candidateSequence: 2
9 voteStatus: open
The pipeline should typically emit:
- one entry in
data/candidates.json - an updated candidate summary in
data/artifacts.json - candidate/version context in front matter for pages staged from that candidate
- denormalized candidate fields in
data/content-index.jsonfor those pages