SMC Consumer Contract — v1.23
Status: v1.23 proposed — remote MCP projection, audit, and deployment hardening, shipped on 2026-09-03; serving endpoint and direct-protocol subset verified. v1.22 proposed — read-only remote MCP access for Codex, Claude Code, and agentic BIMei consumers. v1.21 proposed — the recorded rights grant IS the featured-preview nomination, and a correction to reuse_decision, which reported a granted-but-unrendered material as a rights refusal. Accepted by Build Digital on 2026-09-02. v1.20 proposed — GET /api/smc/vocabularies/{vocabulary}, closing the gap v1.19 left open: the open material_type set can now be read rather than guessed. v1.19 proposed — material_type is an OPEN vocabulary, governed in the database and extensible between releases; consumers must not fail closed on an unrecognised value. v1.18 proposed — one additive material_type value (academic_article_or_chapter); no shape, field or status-code change. v1.17 staged — inactive release-pinned classification path. The exact receipt route, active-pin filter, release-derived compatibility projection, P3B.2 storage/guards/private evidence tooling and guarded cutover/rollback substrate exist. The shipped bootstrap admitted the exact taxonomy v1.1 P3B.3 trust anchors idempotently in production behind an explicit write-intent flag. v1.0 is a predecessor/disposable regression fixture and is not a production import target. No active pin, runtime/cutover row, capability, consumer acknowledgement or reader/writer switch exists. v1.16 page previews remain shipped, with worksheet-range activation approved and in release. v1.15 proposed (country codes are ISO 3166-1 alpha-2, normalised on intake); v1.14 proposed (archived organisations and buckets stop serving); v1.13 proposed; v1.12 proposed; v1.11 proposed; v1.10 proposed; v1.9 proposed; v1.8 proposed; v1.7 proposed; v1.6 proposed; v1.5 remains the signed-off baseline.
Author: SMC (Bilal Succar)
Signed-off by (history): MAD, RPF (v1 on 2026-05-15; v1.1 amendment on 2026-05-15; v1.2 amendment on 2026-05-16; v1.3 amendment on 2026-05-26; v1.4 amendment on 2026-05-26; v1.5 amendment on 2026-05-27)
Onboarded, non-approving: BIMd (integration guides in docs/consumers/; calls production today), Build Digital KB (v1.13, scoped credential)
Future consumers: BIMei KB (approves at onboarding via amendment)
Last updated: 2026-09-03 (v1.23 shipped; direct-protocol subset verified)
What changed in v1.23
Behaviour-narrowing privacy, audit, and deployment correction for the remote MCP only. The five tool names and input field names, read-only authority, REST routes, REST credentials, and durable citation identities are unchanged. Validation bounds and contradictory-selector rejection are tighter. This amendment was deployed on 2026-09-03; the serving endpoint and direct-protocol subset were verified. Exact evidence and the remaining client/audit/monitor gates are recorded in the operator runbook.
Full passage results are an explicit consumer projection
get_chunks full mode and get_material(include_chunks: true) no longer pass
the database-backed chunk row through to an agent client. Each full passage is
projected deliberately to the established consumer fields:
conversionId,chunkId,granularity,anchor,text,hashSha256,chunkRecordId, andpassageVersionId;headingPath,pageStart,pageEnd,charStart,charEnd, andtokenEstimate; andlanguage,language_declared, andlanguage_detected.
The implementation-only id, materialId, chunkIndex, metadata, and
createdAt fields are no longer returned by the MCP. They were documented in
v1.22 as leakage rather than stable contract fields. Manifest mode remains the
same minimal chunkId/hashSha256 projection selected by fields.
Schemas and result size are bounded
Every tool advertises a strict output schema. Input objects reject unknown
properties; material/source query text is capped at 500 characters, material
keys are 2–120 lowercase ASCII letters, digits, or hyphens, other locators have
explicit bounds, and citation selectors reject contradictory
passage_id plus anchor combinations. The serialized MCP result is capped at
131,072 bytes. A result that would exceed the cap is replaced by an opaque
isError: true response with error_code: "result_too_large" and
max_bytes: 131072; callers narrow the query or use the text-free manifest
projection.
Wrong-material conversion selectors and ambiguous logical passage locators use
the safe, caller-correctable codes conversion_not_for_material and
ambiguous_passage_id. Internal exception detail remains only in SMC logs.
Search counts apply eligibility before they cross the boundary
search_sources.material_summary keeps its envelope, but its total now counts
only materials that satisfy the complete consumer-eligibility rule at query
time. Its byStatus array is empty when the count is zero and otherwise contains
only { status: "accepted", count }. Candidate, rejected, incomplete,
admin_only, restricted, and otherwise withheld material counts are not exposed.
For MCP reads, the conversion named by latest_conversion_id must exist for
that material and have status completed or completed_with_warnings. A later
failed conversion attempt does not replace the named current conversion, and a
missing or non-serving current conversion fails closed.
Authority class is applied within the source query rather than after its page is
read. Material search likewise applies the consumer predicate before pagination,
so candidate_total—retained under its compatibility name—and continuation
describe the eligible result set rather than a wider candidate set with
post-read gaps. Every selected material is still rechecked immediately before
projection to fail closed on a concurrent state change.
Governed reads require a durable audit record
A schema-valid MCP tool call must persist its mcp_tool_call event before the
domain read. If that required insert fails, the tool returns the existing opaque
execution error and does not return evidence. The row records the declared
consumer identity with actor_type=agent, target_type=mcp_tool, target key
equal to the tool name, and an allowlisted payload: tool,
identity_assurance: "declared_shared_key", optional credential_slot, and an
optional governed material_key. Queries, filters, excerpts, anchors, passage
selectors, and credentials are not recorded. Citation helpers attempt a
secondary, best-effort citation_requested event with the same identity instead
of defaulting nested MCP citations to smc_admin; the required pre-read
mcp_tool_call remains the durable audit gate.
The operator-facing consumer monitor and its export include the codex and
claude_code identities and count REST api_call, successful MCP
mcp_tool_call, and authenticated mcp_tool_rejected activity. An
authenticated known-tool schema rejection, unknown tool name, or JSON-RPC batch
body attempts a best-effort mcp_tool_rejected event containing only a
sanitized tool label, invalid_arguments, unknown_tool, or
batch_not_supported reason, bounded argument-name paths, credential slot, and
identity-assurance label. It never stores argument values and never performs a
domain read. An authenticated batch is rejected as a whole with HTTP 400 /
JSON-RPC -32600 before any member reaches the SDK. Malformed JSON and HTTP
authentication failures are sanitized boundary logs rather than domain-audit
rows. No MCP-specific retention period is introduced by this amendment.
Credential overlap and the browser boundary are explicit
SMC_MCP_API_KEY remains the required current credential. The runtime can also
accept one optional SMC_MCP_API_KEY_PREVIOUS during a controlled rotation
overlap and identifies which credential slot matched for audit metadata. A
previous key never makes a missing current key valid. The hardening deployment
does not bind this optional variable or rotate the credential; using the overlap
is a later reviewed operational change.
Native clients that omit Origin remain supported. When an Origin header is
present, it must exactly match the canonical origin derived from SMC_BASE_URL.
The earlier arbitrary-origin allowlist is no longer honoured. The route still
has no CORS preflight or CORS response headers, so this does not add
browser-hosted MCP support; OPTIONS /api/mcp returns 405.
Repeated bearer-authentication failures share one in-memory token bucket per Cloud Run instance, with capacity 10 and refill 5 per minute. Exhaustion returns 429 with retry headers. The bucket is shared across rejected bearer values on that instance, but not across the service; this is a brute-force speed bump, not a distributed rate limit or general tool-call quota.
The deployment binds the existing MCP key by numeric version
The deployed workflow binds SMC_MCP_API_KEY to
smc-mcp-api-key:2, replacing the mutable :latest alias. This changes only
the Cloud Run current-key reference; it neither binds
SMC_MCP_API_KEY_PREVIOUS nor rotates or reveals the credential. The numeric
binding was observed on serving revision smc-app-00349-l6c after deploy run
33748624376 completed for merge commit
7a033b0bb7647a43caa19c1ff9bf2780f8a6f5e7.
What changed in v1.22
Additive agent-consumer surface. Existing
/api/smc/*routes, response shapes, credentials and consumers are unchanged.
POST /api/mcp — stateless Streamable HTTP
SMC now exposes its governed evidence reads as a remote MCP server. The endpoint uses stateless Streamable HTTP with JSON responses, so it works across Cloud Run instances without sticky sessions or in-memory session state.
Five read-only tools are advertised:
| Tool | Purpose |
|---|---|
search_sources | Find approved source registers. |
search_materials | Discover accepted, consumer-shareable materials and their stable keys. |
get_material | Read a material metadata projection and citation bundle. |
get_chunks | Read current or explicitly versioned passages. |
get_citation_bundle | Resolve material or passage provenance before citing it. |
There are no write, ingestion, acceptance, classification, preview-grant or administration tools. Any future mutation is a separate amendment and must be proposal-first with an explicit human approval boundary.
Authentication is separate from the REST super-key
Every request requires both:
Authorization: Bearer <SMC_MCP_API_KEY>
X-SMC-Consumer-App: codex | claude_code | rpf | mad | bimei_kg | bimd | build_digital_kb | smc_admin
SMC_MCP_API_KEY is accepted only by /api/mcp. It is deliberately distinct
from the existing SMC_API_KEY, which can write through REST routes. A leaked MCP
key therefore cannot approve, upload, deprecate or otherwise mutate SMC through
the consumer API. The consumer header is an audit tag inside the shared internal
MCP trust boundary, not independent authentication; per-client credentials or
OAuth remain a later hardening option.
In v1.22, a browser Origin was admitted when it matched the SMC origin or an
entry in SMC_MCP_ALLOWED_ORIGINS. V1.23 supersedes that behaviour: only the
canonical origin derived from SMC_BASE_URL is accepted. CLI clients normally
send no Origin. This remains an origin guard, not CORS support: v1 has no
preflight handler or CORS response headers, so browser-hosted clients are not
supported.
The SMC eligibility and citation rules still apply
Material search and every material, passage and citation read re-apply the shared
consumer-eligibility predicate at request time: material accepted, source
approved, current conversion present, visibility not admin_only, and reuse
policy not restricted. A refusal is returned as not found; callers must not
infer existence.
For durable evidence references, consumers store passage_version_id and
conversion_id. chunk_id remains a current logical locator and is not a
permanent passage identity. The MCP server advertises this rule in its initialize
instructions as well as its tool descriptions.
Historical v1.22 compatibility notes, superseded by v1.23
The following described the deployed v1.22 behaviour before the v1.23 hardening release. They are retained as release history, not current contract behaviour:
- source and material
qfilters searched metadata only; this limitation remains current because passage-text search still requires a known material key andget_chunks; - tool results included
structuredContent, but definitions did not advertise formaloutputSchemavalues; - full passage responses included DB-backed internal fields and arbitrary chunk metadata;
- source
material_summaryaggregated all lifecycle states rather than only consumer-visible materials; and - schema-invalid and HTTP-authentication failures did not create MCP audit rows,
handler audit writes were fail-open, and nested citation events were attributed
to
smc_admin.
What changed in v1.21
One correction and one clarification. No new endpoint, no new field, no status-code change. One existing field,
reuse_decision, starts giving a different answer in one case — it is the case where the old answer was wrong. A consumer that already branches onpreviewsbeing empty is unaffected.
The recorded grant is the featured-preview nomination
v1.16 said SMC does not write alt text, and the worksheet section said the curator picks the range. What neither said is what a consumer should do to find the picked one, because the answer was implicit in machinery already described: renditions of ungranted locators are omitted from the list.
So a page-scoped grant naming one page, or a worksheet grant naming one range,
already narrows previews[] to exactly the locator a curator chose. That is
the nomination. There is no separate nomination field, and SMC does not
maintain a per-consumer editorial selection.
The invariant, as Build Digital asked for it to be stated:
The nomination is the sole distinct locator represented by
previews[]. Multiple distinct locators are ambiguous and must not be automatically resolved by the consumer.
Concretely, a consumer displaying a single preview must:
- Group
previews[]bysource_locator— not bypreview_id, and never takepreviews[0]. One locator is up to three renditions. - Render only when there is exactly one distinct permitted locator.
- Fail closed on zero and on more than one. Zero means nothing is permitted or nothing is rendered yet; more than one means the rights decision is wider than the editorial choice and SMC has not been told which page is the card. Neither is a state a consumer may resolve by picking.
- Only then select the rendition nearest its preferred width.
Read width from the response. A rendition is fitted, never stretched, so
the nominal 1440 commonly arrives as 1439 or narrower — on the Build Digital
corpus, a top width below 1440 on every material measured. A hard-coded
width === 1440 lookup finds nothing.
Why SMC does not add a nomination field. It would be a second copy of a choice the response already carries, and a copy drifts. The same reasoning keeps alt text and captions on the consumer's side, which v1.16 stated and this amendment does not disturb. If a curator ever needs to grant wider than they feature, or a second consumer needs a different image for the same material, that is a new decision and gets a real record; it is not this.
Correction: reuse_decision reported a rights refusal it never received
This is a behaviour change, and it is a fix.
reuse_decision was derived from how many renditions survived the rights
filter. So a material a curator had explicitly allowed reported denied
whenever nothing was rendered — which is not a corner case:
- automatic generation on conversion is off by default, so a material can be granted long before anything is rasterised;
- previews are bound to one conversion, so every re-conversion orphans the previous conversion's renditions and leaves a well-granted material with none current.
denied is what a consumer escalates as a rights question. This is part of why
BD-MAT-3134's four-day outage was diagnosed as a rights problem twice when it
was an eligibility problem.
From v1.21, reuse_decision reports the rights decision and nothing else.
| Situation | Before | Now |
|---|---|---|
| Granted, renditions present | allowed | allowed |
| Granted, nothing rendered yet | denied ❌ | allowed, previews: [], total: 0 |
| Granted, renditions orphaned by a re-conversion | denied ❌ | allowed, previews: [], total: 0 |
| Refused, or no grant and a policy that does not permit a facsimile | denied | denied |
| Page-scoped grant, original since replaced | denied | denied |
allowed with an empty previews array is now a real state, and it means
permitted, not rendered yet — ask again. It is the state the earlier guidance
described in prose as "denied with total: 0", now said properly by the field.
Consumers must branch on
previewsbeing empty, never onreuse_decisionalone. This was already the standing guidance issued to Build Digital on 2026-08-18; v1.21 makes the field agree with it.
denied continues to return an empty previews array rather than a 404, and the
content route continues to return 404. Nothing about which bytes are servable
changes — the per-locator filter and the per-request re-authorisation on the
content route are untouched. This amendment changes one summary field, not a
permission.
What changed in v1.20
Additive — one new read endpoint. No existing response shape, field or status code changes. Consumers on v1.19 keep working unchanged.
GET /api/smc/vocabularies/{vocabulary}
v1.19 made material_type an open set and asked consumers not to fail closed on
an unrecognised value — but left them no way to read the set. A consumer
wanting its own picker had to hardcode a list, which is the exact duplication
v1.19 removed from SMC's own code, or infer one from whatever the corpus
happened to contain. This closes it.
GET /api/smc/vocabularies/material_type
{
"vocabulary": "material_type",
"total": 32,
"terms": [
{ "value": "framework", "label": "Framework",
"description": null, "deprecated": false, "deprecated_reason": null },
{ "value": "case_study", "label": "Case study", "description": null,
"deprecated": true, "deprecated_reason": "Seeded in error from a pre-v1.6 list." }
]
}
Retired terms are included, flagged rather than filtered, because the two audiences need opposite things:
| You are… | Use |
|---|---|
| rendering a picker | terms where deprecated is false |
labelling a material's existing material_types | all terms — a material keeps the tag it was given, and a retired term still needs a label |
A mode parameter would have let a caller pick the wrong one silently. One shape
serves both; deprecated says which is which.
material_type is the only vocabulary served. jurisdiction_level has rows
in the same table and is deliberately excluded: nothing in SMC reads them — the
jurisdiction picker is still a constant in the page — so serving them would
publish a list that merely resembles the one in force, with no way for you to
tell. It becomes available when something starts reading it. Any other name
returns 404 with code: "unknown_vocabulary" naming what is available.
Unpaginated, deliberately: the vocabulary is tens of rows and is only useful entire. A paginated picker list is a picker with a bug in it.
Conditional requests are supported and worth using. The response carries an
ETag over the served payload and Cache-Control: private, max-age=0, must-revalidate. Send If-None-Match and an unchanged vocabulary costs you a
bodyless 304. There is no freshness window on purpose — a term an admin adds
is live in SMC immediately, and a consumer holding a stale picker for even a few
minutes would offer a list SMC has already moved past. Weak validators (W/"…")
are honoured. The tag moves when a term is added, relabelled or retired — a
retirement changes what you should offer, so it must change the validator.
Auth. The shared API key works. For a scoped credential the route
requires the read:vocabulary capability, granted to Build Digital KB on
2026-08-20; any other scoped credential gets 403 until granted the same way.
Unusually, this capability carries no organisation predicate. Every other scoped read proves the thing being read belongs to the credential's organisation; this response contains no material, no bucket and no rights decision — the same rows for every caller, already published in this document — so there is nothing to scope it to. The absence is deliberate, not an oversight, and it is the reason the capability is granted per consumer rather than assumed: "harmless" is a judgement made once and written down.
The read-only rule is unchanged: a scoped credential holding this capability
still cannot use any method but GET.
What changed in v1.19
material_typeis now an OPEN vocabulary. No response shape, field or status code changes, and every value valid under v1.18 is still valid. What changes is the promise about the value set: it is governed in SMC's database rather than in SMC's source, so it can gain a term between releases.
Treat material_type as an open set
Until now the vocabulary was a list in SMC's source, duplicated into two database constraints, a curator's picker and an AI prompt. Keeping four copies in step by hand failed twice — most visibly when the curator's edit form spent weeks offering eighteen values it then refused to save. There is now one definition, a governed table, and everything reads from it.
For consumers this means one thing: do not fail closed on a material_type
you do not recognise. Carry it, display it, ignore it in a filter you have no
rule for — but do not treat it as a malformed response. The values listed in §7
are a snapshot at the time of writing, not a closed enum, and SMC does not
consider adding one a breaking change. Values are still ^[a-z0-9_]+$, still
lowercase, and are still never removed — a term that falls out of use is
retired, which stops it being assigned to anything new while every material
already carrying it keeps its tag and keeps serving it.
document_type is unaffected and remains a closed enum. The distinction is
deliberate: SMC's code branches on document_type to select a converter
pipeline, so a value no deployed code knows about is not a new option, it is a
document nothing can process. material_type is descriptive and carries no such
branch.
POST/PATCH on materials now return 400, not 500, for a bad type
A material_type outside the vocabulary previously escaped as an unhandled
error and surfaced as a 500 — the same response SMC returns when it has itself
fallen over, which told a consumer nothing about what to fix. It now returns:
{ "error": { "code": "invalid_material_type",
"message": "Material type not in the vocabulary: nonsense_type.",
"unknown": ["nonsense_type"], "deprecated": [] } }
This is a correction of an error path, not a new one: no request that succeeded before fails now, and a request that failed before still fails — with a status that says whose fault it is.
Since added, in v1.20: GET /api/smc/vocabularies/material_type enumerates
the terms, so a consumer rendering its own picker no longer has to hardcode a
list or infer one from the corpus.
What changed in v1.18
Additive, and a value only. One new
material_typeterm. No response shape, field, status code or existing value changes, and nothing a consumer sends today starts failing.
academic_article_or_chapter joins the material type vocabulary
A peer-reviewed article, a conference paper or a chapter in an edited volume had
no honest home in the vocabulary. Curators reached for report, which is the
band for point-in-time evidence artefacts generally and says nothing about the
scholarly provenance that makes an academic source worth citing differently.
Adding a value to an existing enum is contract-safe by the v1.6 precedent,
so this ships without blocking on MAD/RPF sign-off. A consumer that does not
know the term treats it the way it already treats every other term it has no
special handling for. A consumer that filters ?material_type=report will stop
matching material reclassified onto the new term — which is the point of the
change, not a regression of it.
Review cadence: 36 months, alongside report, assessment and survey.
An academic article is evidence OF a date rather than a live rule; it does not
go stale, it ages.
Already applied: the 14 materials under source succar-publications carry
material_types: ["academic_article_or_chapter"] as of this amendment. Their
prior types (report, framework, guidance, guide) were replaced, not
appended. No material changed status, visibility or reuse policy, so no
material.published / material.deprecated mutation event was emitted and no
conversionId or chunk hash moved — a consumer polling change anchors sees only
updatedAt.
What changed in v1.17
Additive, expand-only and inactive. P3A established the storage and type contract. P3B adds the pre-activation receipt, filter, verification, shadow migration, exact inactive import and controlled-cutover paths. The active release pin, capability grant and active reader/writer remain unchanged; the historical legacy constant remains a one-release rollback fixture and has no runtime importer.
Exact release cache and scheme pins
SMC stores an immutable imported package release using its authority, release ID, full manifest SHA-256 and exact artefact URI. Each cached scheme and concept is bound to that release ID/hash. Labels, definitions and notation live in this immutable cache; assignment rows never copy them.
Import is not activation. P3B.2 stages exact pin targets in a separate immutable table; a staged pin is never read as active. Per-scheme activation is a separate atomic operation that creates a new active pin and retains the previous pin as history. P3A adds the empty structures only: it imports no package and activates no pin.
One generic assignment identity
The v1.17 substrate applies to exactly two subject types: material and
course. An assignment records:
- subject authority, subject type and stable subject ID;
- scheme IRI, package release ID/full hash and concept IRI;
- state and any stale reason/change reference;
- immutable evidence references;
- proposer, method, model and confidence; and
- reviewer, decision reference, timestamps and audit ID.
Assignment identity includes immutable shadow import, subject, scheme and
concept. Topic and Resource Kind remain single-valued per shadow import through
a database uniqueness guard; Lifecycle Relevance is release-defined
multi-select and may carry multiple concepts for one material. Candidate
targets and evidence are retained in immutable review
rounds; state transitions are append-only review events. A material with no
applicable lifecycle concept requires one immutable SMC-owned
not_applicable decision with evidence. Active N/A evidence and an
approved/stale lifecycle assignment are mutually exclusive.
Administrative state and public receipt boundary
The administrative transition is proposed → reviewed → approved|rejected,
with proposed|reviewed|approved → withdrawn and approved → stale. A stale
assignment remains stale while a new immutable review round moves through
proposed and reviewed. Only the final owner decision moves the assignment to
approved, rejected or withdrawn; re-review never temporarily hides the
stale signal.
The consumer receipt contains only:
approved, withprojection_eligible=true; orstale, withprojection_eligible=false, the last approved decision, invalidating release/change, reason and timestamps.
proposed, reviewed, rejected and withdrawn are administrative/audit
states and never appear in that receipt. The endpoint is
GET /api/smc/classifications?subject_authority=…&subject_id=…; both identifiers
are exact and required once. Scoped credentials require
read:classification_receipts, which no current registration receives.
Out-of-scope and unknown subjects return the same 404.
Material scope is resolved through the existing material-to-organisation registry. P3B.2 adds an empty generic course-scope registry and exact authority-to-organisation guard. Course receipts remain fail-closed until one complete, hash-verified 30-row authority vector is atomically registered for an active organisation. Zero, partial, duplicate, wrong-type, wrong-organisation or inactive matches all return the same 404; an IRI, assignment row or caller claim is never accepted as scope evidence. For an external course, vector validation and receipt projection run in one repeatable-read/read-only database snapshot, so deactivating any target or non-target vector member cannot race a previous scope decision.
GET /api/smc/materials?classification={schemeKey}:{notation} resolves a known
scheme through its one active exact release pin and matches approved
assignments only. A known scheme with unknown notation produces an empty set;
an unknown scheme is 400. Without the optional filter, the existing materials
response is unchanged. Because no release or pin is active yet, this path
cannot project v1 classifications in the current posture. The materials-list
route remains shared-key-only: its handler declares no scoped capability or
organisation predicate, so scoped credentials receive the existing fail-closed
403. A future scoped list capability must define material collection scope
before that route can be documented or served as scoped.
P3B pre-activation and cutover boundary
Release import accepts the exact P1 eight-file approved bundle plus every
declared dependency proof. It verifies filenames, SHA256SUMS, per-file bytes,
schema, RFC 8785 canonical JSON, payload hash/size/media type, UUIDv5 release
identity, approved status, scheme summaries and the ONT decision reference.
Release authority is derived from the manifest's schema-locked BIMe Initiative rights holder; caller-supplied authority is refused, so the same
bytes cannot be rebound to another authority.
The release/scheme/concept cache graph is derived from the verified collection;
callers cannot supply it. One transaction provides concurrent-safe insertion
and compares the complete persisted graph for idempotence. The persistence
interface has no activation or active-pin operation. The guarded bootstrap
repository trust-anchors exact evidence/file/register hashes, re-verifies every
nested evidence projection, and can persist complete immutable shadow evidence
plus five inactive staged pin targets in one serializable, privileged,
inventory-locked transaction. Public release bytes are compared with the trust
anchor before dependency resolution. An exact repeat is a no-op; partial,
extra, active or drifted state is refused. P3B.2 retains the frozen v1.0
disposable regression fixture. P3B.3 binds the exact public v1.1 bytes,
promotion attestations, frozen ONT activation evidence and Build Digital
authority; its production CLI still refuses without explicit write intent and
the exact production inactive import completed idempotently with one shadow
import and five staged pins.
Pre-switch consumer evidence is exported only through a private CLI. Schema
smc-taxonomy-staged-evidence-v2 is emitted only for the exact active, paused
replayed cutover and carries a top-level canonical cutover section/hash even
before an acknowledgement exists. That section proves, without disclosing the
pause token, that runtime still owns the exact cutover pause and that the live
active-pin rows and generation remain byte-semantically identical to the
immutable pin snapshot captured at preparation. It derives
staged pins, scopes, assignments, complete review histories, receipts,
lifecycle N/A decisions and acknowledgements from one privileged
repeatable-read/read-only snapshot, rejects hand-copied publication decisions,
and emits canonical section/state/evidence hashes to a pre-created owner-only
path. An exported acknowledgement carries the exact still-replayed cutover ID
and cutover SHA,
replay time and database update time, and must chronologically follow that replay
while preceding its database-stamped insertion and snapshot capture. There is no
private-export HTTP route.
Shadow migration consumes the exact accepted 109-row CSV at SHA-256
ad88226478f9bb92c422dec499f5f9dd78346b5312bf2538a80680ad749d77b5
(93 defer, 16 map_existing) and a canonical envelope binding those bytes to
the verified release and ONT decision. Legacy proposed remains proposed;
legacy confirmed becomes non-consumer reviewed with an immutable review
round and append-only proposed/reviewed events. Deferred, unknown and duplicate
rows become explicit exceptions. Nothing becomes approved.
Each pause, final-delta replay, switch, resume, rollback or pre-switch abort transition is transactional. Preparation derives and locks the legacy cursor and affected consumers. Replay evidence compares full canonical legacy inventories as well as persisted assignments, rounds, events, exceptions and delta rows; request callers cannot provide its counts, parity flag or hashes. PostgreSQL blocks all classification/material/carrier writes while paused, with no ordinary session variable bypass. Runtime, run and immutable active-pin generations are checked on every transition. The atomic switch obtains ONT activation and only post-import, post-replay consumer acknowledgements from persisted state. Every assignment is bound to one immutable shadow import, and acknowledgement insertion time is database-stamped and checked against replay and observation time. The switch rechecks the still-active exact pin in final receipt/filter statements, and computes AST-derived legacy-runtime import evidence from the executing build. Rollback is permitted only while the matching switched cutover remains paused; once resumed, rollback is refused until a reverse-replay design exists. Public operations expose no caller policy or runtime-import-evidence override. These are technical integrity gates, not another governance approval.
A failed prepared/paused/replayed exercise can be terminally aborted with the exact generation and pause ownership. Abort leaves pins and reader/writer authority unchanged, clears only that pause, and appends database-stamped actor/reference/time evidence; it is unavailable after switch.
The v1.1 adapter preserves two different hashes: Build Digital
packageReleaseSha256=ed5536dd6d1d3803a1b123cc75dcfd48e969237e64fdc40edbc39ee34d4301fe
means taxonomy payload bytes, while SMC package_release_sha256=
ac6d8892169e8b088657fac46bbff52c5c6653dbf0829852698e387c240c5ea0
means the authoritative release manifest. The reviewed Build Digital authority
artifact at main 4ca4aa2b9a79855c6d8e7372ca4ff1a6f360f327 is
evidence/taxonomy/v1.1.0/build-digital-content-authority.json, SHA-256
5f79b6178d41e9bcbac023d319df112d367a28d414ce795792076fbcfae2de10.
Explicit non-effects of the inactive v1.17 path
- No legacy
confirmedassignment is promoted toapproved. - No inactive release import is an activation; no active scheme pin, runtime row, cutover run, approval migration/backfill or seed is created.
- No existing bare materials-list response changes; the new filter is opt-in.
- No capability grant, consumer acknowledgement or active-pointer switch.
- No hosted mutation or deployment follows from the pre-activation code alone.
- The automatic main-branch deployment for PR #241 applied the seven empty v1.17 foundation tables to production on 2026-08-18. Immediate read-only verification found zero rows in every table. This is schema presence only: no release, scheme pin, assignment, review round or review event exists.
- PR #249 merged and deployed the read-only P3B.1 evidence layer and empty cutover-control schema without importing or activating SMC taxonomy state.
- The P3B.2 main deployment applies an expand-only migration for immutable
shadow imports, separate staged pins, replay exceptions/deltas, course scope,
lifecycle N/A decisions and consumer acknowledgements. Schema deployment
alone inserts no release, scope, runtime row, active/staged pin, assignment,
N/A decision, acknowledgement or cutover record. P3B.3 now supplies the exact
v1.1-only inactive bootstrap path; it remains an explicit, separately audited
production operation and has not been executed, as documented in
docs/runbooks/taxonomy-v117-p3b3-execution.md.
What changed in v1.16
Additive, and OFF by default. No existing endpoint, envelope, field or status code changes. Two new routes appear, gated behind a capability that no credential exercises until it is deliberately granted, and behind a rights decision no material carries until a curator records one. A consumer that ignores this amendment is unaffected in every respect.
Page previews — a governed image surface
Until now every byte SMC served a consumer was text: metadata, markdown, chunks, citations. Build Digital's WordPress plugin needs something SMC has never offered — an image of a real page of the source document, to show on a material card and, if an editor chooses it, as a public featured image.
Two routes:
| Endpoint | Returns |
|---|---|
GET /api/smc/materials/{key}/previews?conversion_id= | The previews bound to one conversion. Metadata and locators; no bytes. |
GET /api/smc/materials/{key}/previews/{previewId}/content | The image bytes for one preview. |
{
"material_key": "ie-010-information-manager-bim-role-profiles-2025",
"conversion_id": "d3915df9-29c0-4d2b-9f04-2fbc3b772001",
"is_current_conversion": true,
"reuse_decision": "allowed", // allowed | denied — see "The rights gate"
"granted_pages": [3], // pages the decision covers; null = the whole document
"previews": [
{
"preview_id": "a9334d3b-eb0a-4996-84f9-82f63d6ed53b",
"kind": "page_preview",
"source_locator": { "type": "page", "page": 1, "label": "Page 1" },
"mime_type": "image/webp", // image/webp | image/png
"width": 480,
"height": 679,
"byte_size": 47212,
"sha256": "fc91c6c5ca387c46930baafac30885aa9ef6f868b8fe44ebc6ea34f7fb366797",
"content_path": "/api/smc/materials/{key}/previews/{preview_id}/content",
"reuse_decision": "allowed",
"created_at": "2026-08-16T21:08:00.000Z"
}
],
"total": 63
}
202 when the material exists but has no conversion yet (poll, as with
/markdown). 400 for a malformed conversion_id. 404 for an unknown
material — and for a well-formed conversion_id belonging to a different
material, because saying otherwise would confirm that conversion exists
somewhere.
What SMC will not do here
Four boundaries, stated because each one is a thing a preview surface could plausibly have done and this one does not.
- SMC does not generate, redraw or reconstruct a page. PDFium draws the
publisher's own page. The only transformations are: scale down, composite onto
white, encode. No crop, no rotation, no colour correction, no overlay. The
whole page is always present —
object-fit: containsemantics, enforced in the scaler rather than promised in prose. - SMC does not write alt text.
labelis a factual locator ("Page 3"), not a description. The public alt text and caption are the consumer's editorial act, and manufacturing one from the material title would produce something that reads as a description and is not one. - SMC does not hand out a URL to storage. No signed URL, no redirect, no bucket or object path in any response. Bytes are proxied through the route so that authorisation is re-decided on every single retrieval, rather than delegated to a URL that outlives the permission that created it.
- SMC does not infer image rights from
reuse_policy. See below.
The rights gate, and why it is separate — and per page
A full-page raster reproduces the publisher's typography, layout, and any
photography or licensed diagram on the page. That is a materially larger
reproduction than the markdown SMC already serves, and reuse_policy was
written before any image surface existed.
So the derived rule is the strictest defensible one: only original_allowed —
the single policy that already permits redistributing the original file —
implies a page image is permitted, because a page image is strictly less than
the file it came from. Every other policy denies by default, including
metadata_and_citation, which is what the Build Digital pilot materials carry.
The way a material becomes previewable is therefore a recorded human
decision: a curator records allowed with a stated basis — which licence,
permission or ownership makes it right — attributed and timestamped.
The decision names pages, not documents. The rights position across a single file is rarely uniform: on both Build Digital pilots the cover carries commissioned photography while the interior pages carry only the publisher's own diagrams and typesetting. A whole-document yes/no forces the curator to answer for the riskiest page, which either withholds the safe pages or publishes the risky one — neither being the decision anyone meant. So a grant carries a page set; leaving it empty means the whole document, which is the older behaviour and still available where it is what the curator means.
A page-scoped decision is bound to the file its pages were counted in. Page numbers are positional. If the stored original is replaced — a new edition, a corrected file — page 3 of the new document is not the page anybody approved, so the grant stops applying until it is made again. A re-conversion of unchanged bytes does not invalidate it, because re-converting does not re-paginate anything: the binding is to the document, not to the conversion. This is the same rule Build Digital asked for on the spreadsheet side, applied to the surface that exists.
An explicit denial overrides even original_allowed, and a restricted
material can never be re-opened by a grant.
For a consumer this shows up in two fields:
granted_pages— the pages the decision covers, ornullfor the whole document. It is how a consumer tells "page 4 was not rendered" apart from "page 4 is not permitted".reuse_decision—deniedreturns an emptypreviewsarray rather than a 404, and the content route returns404.
Renditions of ungranted pages are omitted from the list, not itemised as
denied, and the content route refuses them with the same 404 it gives a
preview that does not exist. A consumer cannot map the granted set by probing.
Preview identity is immutable and conversion-bound
preview_id is permanent. Its bytes never change: object paths are
content-addressed, writes are create-only, and the database refuses UPDATE on
the row outright. Re-running generation over an unchanged conversion returns the
same ids and the same hashes.
A new conversion mints new previews with new ids. The old ones are not
deleted and keep resolving, which is deliberate: a WordPress page that has
already published one must not break because SMC re-converted the document.
What changes is that they are no longer current, and the response says so —
is_current_conversion on the list, X-SMC-Preview-Current on the bytes.
Consumer action on re-conversion: SMC's latestConversionId remains the
material-level refresh anchor (v1.9). When it changes, re-read …/previews and
re-select; do not assume a stored preview_id is still the current image.
The narrower race — a preview superseded between selection and display — is
signalled by X-SMC-Preview-Current: false on the bytes. The pattern Build
Digital adopted is the recommended one, and it is recorded here so the next
consumer does not have to derive it: re-read the list and re-select once, then
fail closed if the replacement is also superseded. Re-selecting in a loop would
turn a fast-re-converting material into an unbounded retry, and serving the
superseded bytes anyway would defeat the header.
Caching
ETag is the content sha256, quoted and strong. If-None-Match is honoured
with weak comparison, so W/"…" from a proxy still yields 304. Last-Modified,
Content-Type and Content-Length are accurate.
Cache-Control: private, max-age=0, must-revalidate — not immutable.
The bytes are immutable; the permission to read them is not. Bucket membership,
material eligibility and the rights decision are all re-evaluated on every
request, and a long-lived cache would keep a withdrawn image on a public page
for as long as its max-age. Revalidation is cheap precisely because the
validator is the content hash. private keeps the bytes out of shared caches,
which would otherwise serve a credentialled response to callers with no
credential.
Renditions
Approximately 480, 960 and 1440 px wide per page. Each page is rasterised once at the master size and the smaller renditions are reductions of that same bitmap, so the three are the same picture at three sizes rather than three independent rasterisations that can differ in hinting and anti-aliasing.
Three rather than one immutable master, because the consumer surface is a card grid: a 1440 master costs roughly 190 kB where the 480 costs 47 kB, and resizing in the browser pays that on every view with no shared cache. The extra cost is two encodes, not two rasterisations.
Full sizes exist only for pages the decision covers. Every page of a document is rendered at a 480 px review size so a curator can choose one; the 960 and 1440 renditions are produced only for pages the recorded rights decision permits. This is invisible to a consumer — ungranted pages are not listed at all — but it is why a newly-approved page may briefly offer fewer widths than the rest until SMC finishes adding them, which it does automatically when the decision is recorded.
Read width and height from the response rather than assuming the nominal
value. A rendition is fitted, never stretched: a page whose geometry rounds
down comes back at 1439, and a page too tall to honour the width (the master is
bounded at 5760 px high, so a banner-shaped page is not rasterised into
hundreds of megabytes) comes back narrower than requested and reports it.
WebP is the delivery format. The encoder produces lossy and lossless candidates and keeps lossless whenever it is not materially larger — which, on the Build Digital corpus, is every text and diagram page, so those are byte-exact reproductions. PNG is encoded as a third candidate and kept only when it beats lossless WebP; on this corpus it never does, and it stays for the corpus where it will. Metadata is stripped: no EXIF, XMP, ICC or producer string.
A rendition above 3 MB is refused and reported, not silently truncated, and generation stops after 50 pages.
Fonts do not vary by host, so the hash does not either. PDFium here is a WebAssembly build with no filesystem access, carrying its own font data — verified by rendering a PDF that references Helvetica without embedding it, which draws correctly. A page therefore renders to the same bytes on any machine, which is what lets the content hash serve as the ETag. Confirmed by execution, not by argument: both pilot documents render byte-identically on macOS/arm64 and inside the production linux/amd64 image. Both pilot documents happen to subset-embed every font they use (Roboto and Roboto Slab), so the question does not even arise for them.
The maintenance consequence is contractual: upgrading the PDFium build is a
renderer_version change. A new PDFium may draw the same page differently, and
existing previews are never re-rendered, so a consumer holding a preview_id
keeps exactly the image it fetched.
Capability and audit
read:material_previews — a sixth read capability, granted separately from
read:material. A consumer permitted a material's text has not thereby been
permitted its pictures, and revoking the pictures must not cost it the text.
Both routes are recorded in the existing consumer audit trail on every call.
Formats
PDF page previews and XLSX worksheet-range previews. A material with no stored original file cannot have previews at all, however eligible it otherwise is.
Worksheet-range previews (revision 4, pilot activation)
Open item O-v1.13-1 records that spreadsheets have no positional locator. This extension covers the preview half of that gap. It is implemented behind curator and consumer gates; only individually reviewed ranges may be generated.
Revision 2 folded in Build Digital's review of 2026-08-16, which accepted the baseline and added twelve conditions. All twelve are adopted. Three of them turned out to change the shape of the thing rather than add a rule to it, and those are marked.
Revision 3 settled the two questions revision 2 left open — the range ceiling and the font set — with Build Digital's decisions of the same date. Revision 4 puts the legibility flag on the wire, which was the last thing left undecided. Every part of this design is now specified. The three-workbook pilot is the first activation; it does not authorise collection-wide generation.
{
"type": "worksheet_range",
"sheet_name": "Deliverables",
"sheet_id": 3,
"range_a1": "A1:H32",
"print_area": "A1:H32",
"range_basis": "print_area"
}
Identity and selection
Sheet identity and renaming. sheet_name is what a human reads and is
not identity: renaming a sheet must not silently repoint a published image.
Identity is the workbook's internal sheet id (sheet_id), which survives a
rename and a reorder. A preview records both.
A sheet id is not sufficient on its own (adopted; shape change). Build
Digital's condition — a range selection is reviewed again after every new
conversion, and never silently promoted by matching only sheet_id — is
correct, and it generalises: a workbook can keep a sheet id while replacing
everything in the sheet. A range selection is therefore bound to the source
workbook's content hash, exactly as a page-scoped rights decision is bound to
the PDF it was counted in, and a new conversion of changed bytes requires
the curator to confirm the selection again. Unchanged bytes re-converted keep
it. This is now the same mechanism on both surfaces rather than two rules that
happen to agree.
Hidden sheets are ineligible by default (adopted). A hidden sheet, like a hidden row, was hidden by someone. It can be selected only by an explicit curator override recorded in the same way as any other rights decision.
Range precedence. In order: an explicit curator-selected range_a1; then a
defined name scoped to the sheet; then the sheet's print area; then the used
range, clamped. range_basis reports which rule applied, so a consumer never
has to guess whether a range was chosen or inferred. A used range clamped by the
cap is reported as used_range_clamped — an image of part of a sheet must
announce that it is partial.
Merged cells are rendered as merged. A range that cuts a merged cell in half is expanded to whole cells — never clipped mid-cell, for the same reason pages are never cropped. Where that expansion would take the range outside the selection the curator made, it is not performed silently (adopted): the curator is shown the expansion and confirms it, because a merged cell can reach a long way and the difference between "the range I chose" and "the range plus whatever a merge dragged in" is exactly the kind of thing that gets published by accident.
What is rendered, and what is not
Hidden rows and columns are respected, not revealed. Rendering them would publish something the publisher withheld. They are omitted from the image and their count is recorded, so the omission is legible rather than invisible.
Filtered-out rows follow the same rule (adopted). A filter is a hiding mechanism with a different name, and the fact that it is easier to undo does not make the hidden rows more publishable.
Formulas use cached displayed values only (adopted; already the baseline). The cached displayed value is what the publisher last saw and published. Formulas are never recalculated: recalculation would need a full evaluation engine, would produce values the publisher never approved, and volatile functions would make the render non-deterministic — which would break hashing and therefore caching. A sheet whose cached values are absent yields no preview rather than a grid of zeroes.
Macros are never executed, and external connections and links are never refreshed (adopted). Both follow from the same principle as recalculation, and macros additionally make rendering an untrusted-code-execution surface. A workbook is data to this renderer, never a program.
Comments, notes and external-link tooltips are not rendered (adopted). They are frequently internal working commentary — the reviewer's marginalia, not the publication.
Charts, shapes and embedded images require explicit inclusion rules and their own rights review (adopted; shape change). An embedded chart or logo is a distinct work that may carry distinct rights, so it cannot ride along with a decision made about a cell range. Two consequences: a range containing floating objects is rendered without them unless the curator includes them, and the inclusion is recorded in the rights decision rather than in the range selector.
Locale-sensitive dates and numbers use the workbook's stored display format
(adopted). The publisher's 31/01/2026 must not become 01/31/2026 because a
server's locale differs; the stored number format is the publisher's statement
of how the value is meant to read.
Password-protected or encrypted workbooks fail closed (adopted). No preview, no partial render, no attempt to open. Excel's sheet-level protection flag is an authoring control rather than encryption, but it also fails closed by default. A curator may permit read-only rendering only by recording an explicit confirmation against the exact hash-bound sheet and range. That confirmation does not unlock, edit, recalculate, execute or refresh the workbook; it allows the renderer to read the same cached displayed values Excel exposes while preserving every other worksheet-preview safeguard.
Determinism, bounds and rights
Print scaling, orientation and clipping. The sheet's own print setup is honoured — orientation, fit-to-width, scale — because it is the publisher's statement about how the sheet is meant to appear on a page. Where no print setup exists, the range is rendered at natural column widths with no scaling. Nothing is ever clipped: if a range cannot fit the master bounds, it is scaled down.
Fonts must be deterministic or the hash is a lie. A workbook naming a font the renderer lacks would otherwise substitute differently on different hosts and produce different bytes for the same input.
Note that this is a problem the PAGE surface does not have: PDFium is a WebAssembly build carrying its own font data, so a PDF renders identically everywhere. A spreadsheet renderer draws text itself and therefore has to be given a font set explicitly.
Settled in revision 5. Four pinned families, and nothing else:
| Requested | Rendered | Why |
|---|---|---|
| Roboto | Roboto | Preserved, not substituted. Both PDF pilots subset-embed Roboto and Roboto Slab, so it is the house face across this corpus. |
| Arial | Arimo | Metrically compatible, so line breaks and column fits land where the author put them. |
| Aptos Narrow | Roboto Condensed | Explicit, recorded substitution to a pinned OFL condensed face; never a host-font fallback. |
| (missing glyphs) | Noto Sans Symbols | Pinned glyph fallback only — never a text face. |
An unsupported font with no approved mapping fails closed. No preview, and no quiet fall-through to whatever the platform offers: a silent substitution changes the bytes, and therefore the hash, without changing anything a reader could notice was wrong.
Substitution is visible in two places, because two different people need it:
- The curator UI shows the exact requested → rendered mapping for the range being previewed, so the person approving the image can see that the publisher's Arial became Arimo before they approve it.
- The consumer envelope carries
font_substituted: true|false, and when true a compact list:
{
"font_substituted": true,
"font_substitutions": [
{ "requested": "Arial", "rendered": "Arimo" }
]
}
The boolean is the field a consumer branches on; the list is for audit and for the support conversation that starts "this doesn't look like our spreadsheet". Adding both is cheap now and impossible to retrofit cleanly once anyone is reading the envelope.
Explicit limits (adopted; settled in revision 3). Two numbers, doing two different jobs — a distinction Build Digital drew and which the first draft missed by conflating them.
| Bound | Value | Behaviour when exceeded |
|---|---|---|
| Renderer ceiling — rows × columns | 200 × 40 | Refuse. This is a safety limit on what the renderer will attempt, not a claim about readability. |
| Curation guidance — rows × columns | 40 × 12 | Render, and mark the selection legibility_review_required: true. |
| Rendered pixels (master) | 1440 × 5760, as for pages | Scale down, never crop |
| Bytes per rendition | 3 MB, as for pages | Refuse and report |
| Ranges per material | 10 | Refuse |
The ceiling is not a legibility target, and the design must not imply that it is. A 200 × 40 range fits inside the renderer's bounds and is unreadable on a 960 px card. So a selection above the curation guidance still renders — a curator may have a good reason — but it is flagged for a legibility decision rather than passing silently.
legibility_review_required is on the wire (settled in revision 4)
{ "legibility_review_required": true }
Revision 3 left this open, on the argument that legibility is a property of the card and the card is the consumer's surface. Build Digital's answer is better, and it is the reason the field ships: without it every consumer re-implements A1 range parsing and then has to track which version of SMC's guidance it was comparing against. SMC already knows both, so reporting it is cheaper for everyone and moves no authority — the consumer still decides whether and where to show the rendition.
What the field means, stated narrowly, because a flag like this attracts meanings it cannot support:
truemeans exactly one thing: the selected range exceeds SMC's 40 × 12 curation guidance for the 960 px rendition. It is a deterministic property of the range, computable from the locator alone.falseis not a claim that the text is readable in any particular layout. It says the range is inside the guidance, and nothing more.- It is an advisory selection heuristic, never an accessibility or readability guarantee. Final display size, responsive behaviour and accessibility review belong to the consumer, and this field does not discharge any of them.
- Exceeding the 200 × 40 renderer ceiling remains a refusal, not a flag. The two are different mechanisms and must not be confused: one withholds an image, the other annotates one.
Build Digital's stated use — true requires editorial confirmation before a
rendition becomes a card feature image, while remaining usable in an expanded
detail view — is exactly the shape of decision the field is for.
The ceiling covers the three review workbooks as currently authored, which is what makes it a ceiling rather than a guess:
| Workbook | Largest inspected range | Representative range |
|---|---|---|
| BD-MAT-3121 | Instruction!A1:C52 | Deliverables IDP!A1:V14 |
| BD-MAT-3093 | NamingConvention!A1:Z171 | Deliverables!A1:V14 |
| BD-MAT-3129 | Responsibility!A1:G160 | — |
Note that both representative ranges (A1:V14 — 22 columns) exceed the 12-column
curation guidance while sitting far inside the ceiling. That is the flag doing
its job on the first three workbooks anyone will try, which is a reasonable
signal that the guidance is set about right rather than decoratively.
Conversion binding and hashing are unchanged from page previews: the row is
NOT NULL on conversion_id, the object path is content-addressed, and the
sha256 is the ETag. The locator becomes part of the rendition's unique key, so
a sheet may carry several ranges without collision.
Reuse eligibility uses the same recorded-decision mechanism, with the same default and the same binding to the source file's hash. A worksheet range is if anything a more exposing reproduction than a document page — a rendered range can amount to the data itself rather than a depiction of a document — so the decision names sheet and range, as the page decision names pages.
Curator selection. Unlike pages, ranges are not enumerable: a workbook has no natural candidate set. The curator picks — the contact sheet gains a per-sheet range field defaulting to the precedence rule above, renders a proposed range, and the curator accepts or adjusts it before anything is stored. An automatic "representative range" would be SMC making an editorial judgement about a document it did not write, which is the boundary this whole surface is built to respect.
Implementation. SMC renders cached cell values through a bounded, deterministic worksheet layout engine with hidden-row and hidden-column omission, stored number formats, floating-object suppression and a pinned font stack with fail-closed substitution. The renderer does not execute macros, refresh links or recalculate formulas.
Nothing is open on the contract. The two questions revision 2 raised are answered, every condition is implemented, and activation is limited to the individually reviewed pilot ranges.
Not to be scaled. No collection-wide XLSX preview generation until this contract is agreed and representative pilot workbooks have been reviewed individually. Build Digital's three spreadsheet pilots (BD-MAT-3121, BD-MAT-3093, BD-MAT-3129) are the proposed review set.
What changed in v1.15
Additive and backwards-compatible. Nothing a consumer sends today starts failing. What changes is what SMC stores and returns.
Country codes are ISO 3166-1 alpha-2
Every country_code SMC stores or serves is two uppercase letters — IE,
GB, BR — or null when the country is genuinely unknown. This was always
the intent and was never written down, so both forms accumulated:
| Where | Before v1.15 |
|---|---|
smc_sources.country_code | 20 rows GBR beside 2 rows GB |
smc_curation_requests.country_code | 83 rows, all alpha-3 — IRL ×68, BRA ×7, GBR ×3 |
smc_materials.country_code | alpha-2, plus 7 rows of empty string |
Nothing errored. An exact-match filter on IE simply returned zero rows for a
document stored under IRL, which reads as "SMC has nothing for Ireland".
On intake, SMC normalises rather than rejects. A consumer that sends IRL
gets a 200, and SMC stores IE. Rejecting would have broken RPF's live
curation-request flow to enforce a convention it was never told about; this way
the data is correct immediately and each consumer migrates on its own clock.
A code SMC cannot map becomes null, not a guess. Filing a document under
the wrong country is worse than recording that the country is unknown.
What consumers should do: send alpha-2. If you hold alpha-3 locally, map at
your boundary rather than relying on SMC's normalisation — it is a safety net,
not an interface. Where you previously mapped SMC's output (RPF maps IRL→IE
on read), that mapping is now a no-op and can be retired once you have
confirmed no alpha-3 remains on your side.
Enforced at rest, not only in application code: CHECK (country_code is null or country_code ~ '^[A-Z]{2}$') on smc_sources, smc_materials,
smc_curation_requests and smc_source_requests.suggested_country_code. A
three-letter code can no longer be stored by any path — including a future
script or agent that forgets this rule.
What changed in v1.14
This is a behaviour-NARROWING amendment, not an additive one. It removes reads that a consumer can technically obtain today, so it needs explicit notice rather than the usual "ships as proposed" treatment.
Archived organisations and buckets serve nothing
smc_organisations.status and smc_org_buckets.status are now load-bearing.
Until v1.14 they were written (CID revoking an organisation's SMC grant archived
it) and displayed in the admin UI, but no read path consumed them — so an
organisation whose grant CID had revoked kept its bucket fully listed and served
to every shared-key consumer, and its scoped credential kept reading until
someone rotated the environment variable.
From v1.14:
GET /bucketsomits buckets that are archived, and buckets belonging to an archived organisation.GET /buckets/{id}/materialsreturns an empty page (items: [],total: 0) for such a bucket — not a 404, consistent with the existing unknown-bucket rule that avoids existence disclosure.- A scoped credential bound to an archived organisation is refused
(
404, matching the existing out-of-scope semantics).
Curation is preserved. Archiving is a switch, not a delete: re-granting the organisation in CID, or reactivating the bucket in the admin UI, restores exactly the membership that was there.
Notification status (2026-08-13): consumers have NOT been notified. Verified in
production on the same day: both organisations (build-digital, internal) and both
buckets are active, and 0 bucket members would stop serving under the new rule.
The narrowing therefore has no live blast radius today, and notice was deliberately
deferred rather than skipped. Send the notice before archiving the first organisation
or bucket — at that moment this stops being theoretical for RPF (the heavy caller)
and Build Digital KB (scoped to build-digital, the only org with members).
Consumer action: none required for correct behaviour — a revoked organisation simply stops returning data, which is the intent. If your app caches bucket contents, treat an empty bucket response as authoritative rather than as a transient failure.
Why this is a correction, not a feature: the previous state let a takedown look effective in the admin UI while the data kept flowing. Individual material eligibility still applied, so this was not a confidentiality leak in itself, but it was a control that appeared to do something and did nothing.
Bucket updated_at now moves
Membership changes previously left smc_org_buckets.updated_at at its creation
value, so the one field that looked like a sync hint was created_at in
disguise. It is now touched on every membership add and remove. Bucket
membership still emits no mutation event (unchanged from v1.12) — consumers
re-read on cadence, and updated_at is now a usable cheap check.
What changed in v1.13
Additive for existing consumers, with one behaviour change that can break a malformed caller — see "Unknown consumer tags now fail closed". No endpoint, envelope or field changes.
This amendment does the thing §7 called roadmap: it introduces an enforced per-consumer credential, and onboards the first consumer that requires one. It also corrects two statements in §7 and v1.12 that scoped credentials make untrue.
Two credential types
Shared SMC_API_KEY | Per-consumer scoped key | |
|---|---|---|
| Methods | all | GET only |
| Routes | all /api/smc/* | only routes that explicitly opt in |
| Organisation scope | none | exactly one CID organisation |
| Consumer identity | self-declared tag | bound to the credential |
| Holders | MAD, RPF, admin tooling | Build Digital (build_digital_kb) |
Nothing changes for MAD or RPF. The shared key keeps its existing unrestricted behaviour, and no route they call has been narrowed.
Enforcement is fail-closed
A scoped credential is refused unless the route it calls names the capability it serves. Every route that predates this amendment names nothing, and therefore denies scoped credentials without being edited. A route added later inherits the same default. Opening a route to a scoped consumer is an explicit, reviewable act; leaking one is not something forgetfulness can cause.
Five read capabilities exist, one per consumer-facing GET route:
read:buckets GET /api/smc/buckets
read:bucket_materials GET /api/smc/buckets/{bucket_id}/materials
read:material GET /api/smc/materials/{shared_material_key}
read:material_chunks GET /api/smc/materials/{shared_material_key}/chunks
read:citation GET /api/smc/citations/chunk/{passage_version_id}
Organisation scope, and what it is keyed on
A scoped credential may read a material because that material is a member of one of its organisation's buckets — owned or referenced — not because it holds the key. Shared standards remain reachable exactly when a curator has added them as referenced members, which is the same curation act that already declares them shareable.
On GET /api/smc/buckets the organisation_id filter is overridden, not validated: an omitted or mismatched value narrows to the credential's own organisation. There is no argument a scoped consumer can pass that widens the list, including passing none.
Out-of-scope reads return 404, not 403. A 403 would confirm that some other organisation holds a material under the requested key, which is itself a disclosure. Consumers must not infer existence from the distinction.
Scope and consumer-eligibility are independent gates and both apply. A material inside the credential's organisation is still withheld unless it is accepted, its source approved, a current conversion present, visibility not admin_only and reuse policy not restricted (v1.12). Scoping does not widen eligibility.
Unknown consumer tags now fail closed
Previously a X-SMC-Consumer-App value outside the registered enum was silently rewritten to smc_admin, attributing an unregistered consumer's calls to the administrator app in the audit log. It is now rejected with 403.
- An absent header still defaults to
smc_admin— unchanged, and existing consumers rely on it. - A present but unrecognised value is now an error rather than a downgrade.
- Under a scoped credential, identity comes from the credential; a header contradicting it is refused rather than believed.
Registered values are rpf, mad, bimei_kg, bimd, build_digital_kb, smc_admin. Note the underscores — a hyphenated build-digital-kb is not a registered value and now returns 403 instead of silently logging as smc_admin.
Correction to §7 and to v1.12
Two statements elsewhere in this contract are no longer unconditionally true, and consumers were explicitly told to design against them:
- §7's "SMC v1 uses a single shared super-key" and "
X-SMC-Consumer-Appis a tag, not authentication" now hold only for the shared key. Under a scoped credential the consumer-app identity is authenticated. - v1.12's "Any key holder can read any bucket" is now false for scoped credentials, which can read exactly one organisation's buckets.
MAD and RPF hold the shared key, so the original guidance still describes their situation accurately. No change is required of them.
Credential handling
Scoped credentials live in Secret Manager and reach the service as environment configuration — never Git, never a database row, never browser-reachable code. Both the key and its organisation binding are injected; a credential whose organisation binding is missing or malformed authenticates as nobody rather than falling back to an unscoped read. Rotation and revocation are documented in the consumer's onboarding handoff.
Onboarded in v1.13
build_digital_kb — Build Digital Knowledge Navigator. Read-only projection of the Build Digital organisation bucket into a consumer-owned search index, displaying attributed excerpts. CID organisation 4f6134fd-8d6e-4b0a-a9c3-b36b6e1921e5.
Deferred, not dropped: spreadsheet passage locators
SMC chunks carry page_start / page_end and no sheet or cell-range locator. Spreadsheet-derived passages therefore cannot currently be cited more precisely than the whole document.
Build Digital has accepted this for the v1 connector pilot, which uses PDF-backed materials carrying real page locators, on the basis that it is a deferral rather than a removal. A consumer must not present an unlocatable spreadsheet passage as a citation-safe excerpt. Sheet/cell-range locators are recorded as a post-MVP schema and contract amendment for design — see Open items.
What changed in v1.12
Additive — two new read endpoints. No existing endpoint, envelope or field changes. A consumer can now retrieve the materials an organisation has curated for itself, so org-scoped content can be grounded on that organisation's own evidence rather than only on the shared corpus.
What a bucket is
An organisation bucket is a curated grouping of materials belonging to, or chosen for, one organisation. It carries two kinds of membership:
owned— the organisation's own document (its BEP, EIR, AIR, OIR).referenced— a curated pointer at a shared standard that applies to it.
Both are links to ordinary smc_materials rows. Nothing about ingestion, conversion, chunking or shared_material_key derivation changes, so a bucket material is fetched, cited and bound exactly like any other.
Organisations are identified by their CID organisation id. SMC mirrors the organisation from CID's signed provisioning push rather than minting its own identifier, so the organisation a consumer knows and the organisation SMC knows are provably the same tenant.
Read the buckets for an organisation
GET /api/smc/buckets?organisation_id={cid_org_uuid}
{
"items": [
{
"bucket_id": "…uuid",
"organisation_id": "…uuid", // the CID organisation id
"organisation_slug": "build-digital",
"organisation_name": "Build Digital",
"slug": "default",
"name": "Build Digital",
"status": "active",
"material_count": 25,
"updated_at": "2026-08-07T00:00:00.000Z"
}
],
"total": 1
}
material_count is the raw membership count, not a page total. See the filtering note below: the materials endpoint may legitimately return fewer.
Read the materials in a bucket
GET /api/smc/buckets/{bucket_id}/materials?limit=&offset=
{
"materials": [ { /* the standard material shape, plus: */
"membership_kind": "owned", // owned | referenced
"added_at": "2026-08-07T00:00:00.000Z"
} ],
"total": 25,
"limit": 100, // default 100, clamped 1–500
"offset": 0,
"bucket_id": "…uuid"
}
Items carry the same shape as GET /api/smc/materials, through the same mapMaterialForApi projection: the internal status column is never on the wire and review_status is derived. The v1.3 wire-shape convention still applies to the material fields.
An unknown bucket_id returns an empty page (200), not 404, matching the chunk-manifest projection's behaviour for an unknown material. A consumer polling a bucket that has not been provisioned yet should back off rather than treat it as an error.
There is deliberately no bucket-scoped chunks endpoint. Fetch chunks per material through the existing GET /api/smc/materials/{key}/chunks.
This read is filtered, and the filter is not stable over time
A bucket serves only material currently shareable with a consumer: status = accepted, an approved source, a current conversion, visibility ≠ admin_only, and reuse_policy ≠ restricted. This is the same predicate already applied to web-evidence fulfilment snapshots.
It is re-applied on every read, not only when a curator adds a material, because every column it depends on mutates afterwards. A material can therefore leave a bucket's consumer-visible set with no membership change and no notification. Two consequences for consumers:
totalmay be lower than the bucket'smaterial_count, and may fall between polls.- A consumer mirroring a bucket must treat each read as the current truth and prune what is no longer returned, rather than only adding what is new. A material that disappears has not necessarily been removed from the bucket; it may have been deprecated, restricted or hidden upstream.
Bucket membership changes emit no mutation event in v1.12. Consumers re-read on their existing sync cadence.
Scope, not access control
Bucket membership does not restrict who may read a material. SMC v1 authenticates every consumer with a single shared SMC_API_KEY, and X-SMC-Consumer-App remains a tag rather than authentication (§7). Any key holder can read any bucket.
Amended by v1.13. This paragraph now describes the shared key only. A per-consumer scoped credential reads exactly one organisation's buckets, and its consumer identity is authenticated rather than self-declared. Bucket membership is the access boundary for scoped credentials.
What a bucket guarantees is the other direction: everything inside it is already curated as shareable with a consumer app. Buckets are therefore safe for retrieval scoping, so one organisation's evidence does not ground another organisation's content, and must not be relied on for confidentiality. Per-consumer keys and org-scoped authorisation are a prerequisite before any organisation places non-public material in a bucket.
What changed in v1.11
Additive — existing curation-request bodies and responses continue to work. MAD can now identify a request as web evidence, replay it idempotently, inspect acquisition progress, and receive an immutable snapshot containing only approved evidence.
Request intake
POST /api/smc/curation-requests accepts these optional fields in addition to the v1.7 body:
{
"request_kind": "web_evidence", // default: open_discovery
"consumer_request_id": "education-programmes-BRA",
"preferred_languages": ["pt-BR", "en"],
"intended_use": "Evidence for MAD education-programme map layers",
"inclusion_criteria": ["Official higher-education programme page"],
"exclusion_criteria": ["Training-provider marketing page"],
"seed_urls": ["https://example.edu/programmes"],
"freshness": "current" // current | recent_5_years | any
}
When consumer_request_id is supplied it is unique within the authenticated consumer_app:
- First submission:
201with{ request_id, status, replayed: false }. - Exact replay:
200with the original{ request_id, status, replayed: true }. - Same ID with different content:
409with{ error, request_id }.
Legacy submissions that omit these fields retain their existing behaviour.
Request detail
GET /api/smc/curation-requests/{request_id} returns the authenticated consumer's request, a compact acquisition summary, and—only after fulfilment—the fulfilment payload. It does not expose unapproved candidate URLs or crawler output. A request belonging to another consumer returns 404.
Approved-evidence fulfilment
Web-evidence fulfilment is stricter than the legacy manual attachment flow. A material is eligible only when its source is approved, its material status is accepted, it has a current conversion, it is shareable beyond admin_only, and its reuse policy is not restricted.
The fulfilment payload adds:
request_kindconsumer_request_idevidence_snapshot
evidence_snapshot is created once at fulfilment and contains the eligible material keys, source IDs, conversion IDs, and immutable passage-version UUIDs. Subsequent edits, re-conversions, or attachment changes do not silently rewrite it. material_keys and source_ids in the fulfilment are read from this snapshot. Partial fulfilment requires a curator coverage note.
MAD remains the requester and consumer. SMC exclusively owns Vertex discovery, native retrieval, Firecrawl fallback, review gates, costs, retries, and audit history.
What changed in v1.10
Additive — existing request shapes continue to work, but logical chunkId reads now deterministically select the material's latest conversion. This amendment fixes the reconversion hazard where a positional locator could resolve an arbitrary retained historical row.
| # | Change | Surface | Breaking? |
|---|---|---|---|
| 1 | Added immutable passage identity | Full chunk rows and citations | No |
| 2 | Clarified/fixed default current-version scope | /chunks, material chunks, counts, retrieval summaries, logical citation/locator reads | No |
| 3 | Added validated historical version selection | Authenticated material-scoped chunk/citation reads | No |
Immutable passage identity
Each smc_chunks row is now retained permanently for its conversion. Its existing UUID (id) is the immutable passage version and is exposed additively on full camelCase chunk rows as chunkRecordId and passageVersionId (both equal id). It is exposed in snake_case citation payloads as chunk_record_id and passage_version_id.
chunkId remains a logical positional locator. It is intentionally reused when a material is re-converted, including when output bytes are identical. Therefore consumers that need a permanent citation, audit reference, or cached version must store chunkRecordId/passageVersionId and conversionId, not only chunkId.
Conversion-owned artefacts (v1.10 forward-only)
From SMC-CONTRACT-01 onward, normalized_markdown, review_markdown, english_markdown, and obsidian_bundle are written under deterministic conversion-specific object paths and retain a row with that same conversion_id. Writes use create-only object generation preconditions; a 412 is an idempotent retry only when the retained artefact row has the same conversion and hash. Current markdown and bundle reads resolve through material.latestConversionId; a failed finalisation may leave an unreferenced new-version artefact, but cannot overwrite the prior current version.
This is forward-only. SMC does not relabel, migrate, or infer legacy pre-v1.10 object provenance: older material-level paths may not be a reliable historical artefact reference. Reconciliation reports current passage-row completeness; it does not attempt historical artefact repair.
Every citation chunk object now includes:
{
"chunk_record_id": "uuid", // immutable retained row UUID
"passage_version_id": "uuid", // alias of chunk_record_id
"conversion_id": "uuid",
"chunk_id": "material-0001", // logical positional locator
"granularity": "compound",
"hash_sha256": "…",
"anchor": "…",
"heading_path": ["…"],
"page_start": 1,
"page_end": 1,
"text_excerpt": "…"
}
Current-first reads and historical access
GET /api/smc/materials/{key}/chunksandGET /api/smc/chunksreturn only chunks whoseconversionId = material.latestConversionIdby default. Counts, search results, and retrieval summaries use the same scope.GET /api/smc/chunks/{logicalChunkId}andGET /api/smc/citations/chunk/{logicalChunkId}resolve that positional locator only in the current conversion.GETby a UUIDchunkRecordIdretrieves that exact retained historical passage.- Authenticated material-scoped reads may name
?conversion_id=<UUID>to inspect a retained version. The server validates that it belongs to the namedmaterial_key; corpus-wide historical search is not exposed. This uses the same existing API-key access boundary and does not make private text public. - A byte-identical same-order reconversion creates new immutable passage UUIDs but does not emit a
material.rechunkedevent. Changed text, duplicated-hash differences, or reordered chunks do emit it. Consumers should continue usinglatestConversionIdas the material-level refresh anchor.
Consumer action
BIMd should use chunk_record_id / conversion_id in stored evidence references and use a logical chunk_id only for a current-view lookup. On an SMC latestConversionId change, refresh the material-scoped manifest before updating its own index. See ../consumers/bimd-versioned-passages-integration.md for the coordinated rollout order.
What changed in v1.9
Additive — no existing response shape, field, or status code changes. Consumers on v1.8 keep working unchanged. Origin: RPF built its own vector index over SMC chunks (RPF ADR-0004, rpf-app docs/adr/0004-semantic-layer.md; rpf-app PR #597), exactly as §2 Embedding locus sanctions ("Consumers … own their own vector indices"). Building it surfaced three consumer-side gaps. All three are additive; none touch the chunker or re-chunk any existing material.
| # | Change | Surface | Breaking? | Notes |
|---|---|---|---|---|
| 1 | Added — chunk-manifest projection | GET /api/smc/materials/{key}/chunks?fields=chunk_id,hash_sha256 | No (additive; opt-in query param) | Projects each chunk to just its change-detection anchors — no chunk text. Lets a consumer diff its external index against SMC's current chunk set cheaply. Omit ?fields= for the unchanged full-row shape. |
| 2 | Clarified — material-level change anchors | GET /api/smc/materials, …/{key} | No (doc only) | Names updatedAt + latestConversionId as the stable material-level change anchors, and declares un-named whole-row-spread fields (e.g. currentMarkdownHashSha256) unstable — not change anchors. |
| 3 | Fixed — future embedding-column key sketch | §2 Embedding locus | No (doc only; future capability) | The sketched SMC-side embedding column is now keyed by (hash_sha256, embedding_model), not (chunk_id, embedding_model) — chunk_id is positional and renumbers on re-convert. |
1. Chunk-manifest projection — ?fields=chunk_id,hash_sha256
A consumer keeping an external vector index fresh has to answer two questions: did this material's chunks change, and if so which chunks moved (so it re-embeds only those, not the whole document). Before v1.9 neither was cheap: the "did it change" probe (GET /api/smc/chunks?material_key=…&limit=1 for conversionId + total, what RPF does today) can't see a content shift that keeps the chunk count constant, and getting per-chunk identity meant paging the full chunk text.
v1.9 adds an opt-in projection on the material-scoped chunks endpoint:
| Query | Returns |
|---|---|
GET /api/smc/materials/{key}/chunks?fields=chunk_id,hash_sha256 | The same envelope ({ chunks, total, materialKey, granularity }) plus an echoed fields: ["chunk_id","hash_sha256"], where each chunk is projected to only the requested anchors — { "chunkId": "ie-010-…-0001", "hashSha256": "…" }. No text, headingPath, metadata, or language fields. |
?fields=hash_sha256 (a subset) | Only that field per chunk. |
?fields= omitted / empty | Unchanged — the full v1.0–v1.8 chunk row. |
?fields=text (or any non-anchor token) | 400 Bad Request → { "error": "Invalid fields 'text'. Expected a comma-separated subset of: chunk_id, hash_sha256." } (mirrors the ?granularity= 400). |
Notes on the shape:
- Selector tokens are snake_case (
chunk_id,hash_sha256— the names §2 Idempotency anchors already uses). Emitted keys stay camelCase (chunkId,hashSha256) so a projected row parses identically to the full chunk row a consumer already reads — the projection subsets the shape, it does not rename it. - Composes with
?granularity=and?limit=/?offset=(clamped 1–500) exactly as the full read does.chunk_idis unique per material across granularity stripes (atom carries ana-prefix), so a single?granularity=all&fields=…manifest is unambiguous. - The DB read fetches only those two columns — text never leaves storage — so the manifest is cheap end-to-end, not just on the wire.
Change-detection contract: the manifest changes if and only if the material's logical chunk sequence changes. SMC-CONTRACT-01 always retains a fresh immutable passage row for a re-conversion, but this text-free projection intentionally omits that row UUID. Thus a byte-identical same-order re-conversion leaves its chunk_id and hash_sha256 manifest anchors untouched, while a content, duplicate-count, or order change changes the relevant manifest rows. Recommended consumer flow: keep the cheap conversionId/total probe as the first-line "might have changed" gate, and pull the manifest only when it trips — then diff hash_sha256 (treat chunk_id as opaque; see anchor note in item 3) to decide what to re-embed.
This ships only on GET /api/smc/materials/{key}/chunks. The full-corpus search endpoint (GET /api/smc/chunks) is intentionally unchanged — its q/language search semantics are out of scope for this amendment.
2. Material-level change anchors are updatedAt + latestConversionId (and only those)
The materials endpoints (GET /api/smc/materials, …/{key}, PATCH …/{key}) pass the Drizzle row through verbatim (v1.3 Wire-shape convention), so camelCase columns that are not named in this contract may appear or disappear on the wire without a version bump — exactly what happened to effectiveAt, dropped in v1.6 as "an undocumented whole-row-spread leak — never named in this contract."
The stable, contract-named material-level change anchors are:
updatedAt— bumped on every material mutation (§1 "re-validate againstupdated_at").latestConversionId— changes on every re-conversion.
Fields such as currentMarkdownHashSha256 and currentOriginalHashSha256 currently reach the wire only as un-named spread fields. They are NOT contracted change anchors — consumers MUST NOT key on them. For content-hash change detection use the anchors that are contracted: markdown_hash_sha256 from GET /api/smc/materials/{key}/markdown (v1.4) for the whole document, and the chunk manifest's hash_sha256 (item 1) for per-chunk granularity. (RPF has standardized on conversionId + chunk total, which remains valid; the manifest is the precise upgrade when it needs per-chunk diffs.)
This is a clarification, not a wire change — nothing added or removed. It states the rule that effectiveAt established, so consumers don't accidentally bind to the next accidental spread field.
3. Future embedding-column key → (hash_sha256, embedding_model)
§2 Embedding locus sketches a future SMC-side embedding column "once BIMei KB onboards." That sketch previously keyed it (chunk_id, embedding_model). Corrected to (hash_sha256, embedding_model): chunk_id is positional ({materialId}-{padded-index}, with a per-granularity prefix) and renumbers whenever a re-convert changes the chunk count or order, whereas §2 Idempotency anchors already tells consumers to key re-processing off content hashes. Keying the embedding column on hash_sha256 means an identical chunk keeps its embedding across a re-convert that only shifted other chunks' positions. Still a future capability, still out of scope until BIMei KB onboards — this only fixes the forward-looking sketch so it doesn't bake in the positional-key mistake.
As with prior additive amendments, the live API may ship item 1 ahead of consumer sign-off — v1.8 consumers keep working on the unchanged envelope + full-row chunk shape (the projection is strictly opt-in).
What changed in v1.8
Additive — no existing response shape, field, or status code changes. Consumers on v1.7 keep working unchanged. This closes the loop the v1.7 surface left open: a fulfilled open-discovery curation request now carries the structured list of materials that satisfied it, so the requesting consumer can bind them.
A curator, when fulfilling a request, attaches the resulting materials (and/or sources) to it — a discovery request usually yields several, which the single promoted_source_id of v1.7 can't represent. Those links surface on two consumer reads:
| Endpoint | Method | Query | Response |
|---|---|---|---|
/api/smc/curation-requests | GET | ?status=fulfilled (and any existing filter) | each item in items[] additionally carries material_keys: string[] and source_ids: string[] |
/api/smc/curation-requests/recent-fulfilments | GET | ?since=<ISO8601>&consumer_app=<app> | { "fulfilments": [ … ], "cursor_at": string } (cursored delta — mirror of /source-requests/recent-approvals) |
Each recent-fulfilments entry:
{
"request_id": "CUR-…",
"consumer_app": "rpf",
"topic": "…",
"country_code": "IE", // or null
"status": "fulfilled",
"smc_reference": "…", // or null — free-text the curator may add (v1.7)
"promoted_source_id": "…", // or null — the single source link (v1.7), retained
"material_keys": ["ie-bim-…"], // v1.8 — the shared_material_keys to bind
"source_ids": ["uuid", …], // v1.8 — any whole sources the curator attached
"fulfilled_at": "2026-…Z", // ISO 8601, or null
"fulfilled_by": "curator@…", // or null
"resolution_note": "…" // or null
}
Binding contract: material_keys are shared_material_key values — the same identifiers exposed by GET /api/smc/materials / GET /api/smc/materials/{key}. A consumer resolves each key against the materials endpoints to fetch metadata, markdown, chunks, and citations (subject to the usual review-state vocabulary, §6 — a freshly-attached material may still be candidate). The cursor pattern (since → cursor_at) composes with backfill exactly like /recent-approvals.
Lifecycle (unchanged from v1.7): open → acknowledged → fulfilled | declined, triaged by a curator at /admin/sources/curation-requests. v1.8 only adds what a fulfilment returns, not how it's produced. v1 caveat retained: the curator still gates fulfilment manually; SMC does not auto-fulfil from intake.
As with prior additive amendments, the live API may ship ahead of consumer sign-off — v1.7 consumers keep working on the unchanged item/envelope fields. RPF's follow-up (a poll-and-bind step reading material_keys) is the reciprocal consumer-side change.
What changed in v1.7
Additive — a new endpoint; no existing response shape, field, or status code changes. Consumers on v1.6 keep working unchanged.
New surface: open-discovery curation requests — a consumer asks SMC to find materials for a jurisdiction + topic when it has no source URL yet. Distinct from POST /api/smc/source-requests (v1.1), which is URL-keyed ("curate THIS document").
| Endpoint | Method | Body / Query | Response |
|---|---|---|---|
/api/smc/curation-requests | POST | { "topic": string (3–200), "country_code"?: ISO alpha-2/3, "note"?: string } | 201 → { "request_id": "CUR-…", "status": "open" } |
/api/smc/curation-requests | GET | ?status=&consumer_app=&country_code=&limit=&offset= | { "items": [...], "total", "limit", "offset" } |
Auth: the usual X-SMC-API-Key (+ optional X-SMC-Consumer-App); consumer_app is stamped from the header. Lifecycle: a curator triages each request in the SMC admin UI (/admin/sources/curation-requests), moving it open → acknowledged → fulfilled | declined (a fulfilled request may carry smc_reference + link promoted_source_id). Consumers poll GET …?status=fulfilled for outcomes. v1 caveat: ingestion of the discovered materials is the curator's manual job; this covers request intake + status only.
This document is the load-bearing contract between SMC and consuming applications in the BIMe Initiative (MAD, RPF, BIMei KB, BIMd, and future apps). It is one shape for all consumers — bespoke per-consumer deviations defeat the goal of a shared substrate.
The live API at https://smc.bimexcellence.org is governed by this contract as of 2026-05-15. /api-docs links here as the canonical reference. Amendments follow the Change process at the bottom.
What changed in v1.6
Materials UI/UX refresh (docs/plans/materials-uiux-refinements.md). Four material-payload changes, all on the camelCase materials endpoints (GET /api/smc/materials, GET /api/smc/materials/{key}, PATCH /api/smc/materials/{key}). Consumer-impact audited 2026-06-23 across MAD/RPF/BIMd default branches.
| Change | Field | Type | Breaking? | Notes |
|---|---|---|---|---|
| Added | materialTypes | string[], nullable | No (additive) | The full multi-value semantic-kind set. Curators can now tag a material with more than one type. |
| Retained | materialType | string, nullable | No | Still emitted — now the primary type (first of materialTypes), kept for back-compat. Consumers reading it as a scalar string keep working. New consumers should prefer materialTypes. |
| Changed | documentType | string | No | Value set changed: removed policy_document / guidance / standard / procurement (existing rows remapped to other; the semantic meaning lives on materialType now); added xlsx / pptx / video / form / image / map. Still a plain string. |
| Added | publishedAtPrecision | "day" | "month" | "year", nullable | No (additive) | How much of publishedAt is actually known. The publishedAt date still stores a real first-of-period date. |
| Removed | effectiveAt | — | No* | Dropped. It was an undocumented whole-row-spread leak — never named in this contract. |
Filter change: GET /api/smc/materials?material_type=<value> now matches a material that carries <value> anywhere in its materialTypes set (was scalar equality on the primary). No consumer filters on this param today.
Material type vocabulary — OPEN SET, snapshot at v1.19 (see § What changed in v1.19; governed in SMC's database, may gain values between releases, never loses one): framework, policy, mandate, standard, code, specification, regulation, guidance, guide, manual, protocol, plan, strategy, roadmap, template, workflow, role_profile, classification, schema, contract, agreement, requirements, rfp, proposal, submission, academic_article_or_chapter, report, assessment, survey, certificate, cv_resume, awareness. (Superset of the prior 11 — no existing value has ever been removed. academic_article_or_chapter added in v1.18. Do not treat this list as exhaustive: a curator can add a term without a release, and an unrecognised value is not a malformed response.)
Document type vocabulary (v1.6): pdf, docx, xlsx, pptx, html, web_page, video, form, image, map, other.
Consumer actions:
- RPF —
lib/agent/tools/smc.tsvalidatesmaterial_typeasz.string().nullable(); that still holds (the scalar is retained). To surface multiple types, adoptmaterialTypes: z.array(z.string()). Recommended, not required — see the RPF parallel-session prompt in the plan doc. - MAD — two silent fidelity notes (no breakage): (1) the 6 new
documentTypevalues aren't in MAD'sMAD_DOCUMENT_TYPES, somapDocumentTypecoerces them to"other"— add them if those types matter downstream; (2) MAD'seffective_atcolumn simply stops being populated (dateOnly(undefined) → null). - BIMd — no action; reads none of these fields.
* effectiveAt removal is non-breaking because no consumer reads it (RPF only types it, BIMd has it in a fixture, MAD's writer null-coalesces). It was never a contracted field.
Sign-off: this bumps to "stable v1.6" once RPF acks (the change is non-breaking, so the live API can ship ahead of the ack — consumers keep working on the retained scalar).
What changed in v1.5
Additive — /chunks gains a granularity axis tracking BIMei Ontology v4 § Granularity vocabulary (Atom · Cluster · Compound · Complex · System). Non-breaking — consumers that ignore the new field and don't pass the new query param get exactly the v1.4 behaviour (Compound-only chunks, today's shape). Origin: Phase β v4 alignment (docs/plans/phase-beta-v4-multi-granularity.md). β.1 ships the schema + API surface; β.2 adds the Atom chunker; β.3+ adds Cluster / Complex / System on consumer demand.
Granularity vocabulary
v4 § Granularity vocabulary defines five levels:
| Granularity | Roughly | Status in β.1 | Notes |
|---|---|---|---|
atom | Paragraph-level | Not yet emitted (?granularity=atom → empty array) | β.2 |
cluster | Section / sub-heading group | Not yet emitted | β.3+ on demand |
compound | Full heading section — what v1.0–v1.4 chunks already approximate | Emitted (the only granularity SMC writes today) | All pre-β.1 chunks are tagged compound automatically |
complex | Multi-section topic block | Not yet emitted | β.3+ on demand |
system | Whole-document or major-part scope | Not yet emitted | β.3+ on demand |
/chunks endpoints — ?granularity= filter
Applies to both GET /api/smc/materials/{key}/chunks and GET /api/smc/chunks.
| Query | Returns |
|---|---|
| (omitted) | compound chunks only — back-compat with v1.0–v1.4 consumers |
?granularity=compound | Same as omitted — explicit form |
?granularity=atom|cluster|complex|system | Only that stripe. In β.1 the response is an empty array — documented behaviour, not a 404, because the material exists but SMC hasn't materialised this stripe yet. Consumers can poll or wait for material.rechunked events to surface new stripes. |
?granularity=all | All stripes interleaved, ordered by (granularity, chunk_index). Each chunk in the envelope is tagged with its granularity field. |
?granularity=foo (anything else) | 400 Bad Request with { "error": "Invalid granularity '…'. Expected one of: atom, cluster, compound, complex, system, all." } |
The response envelope gains a granularity field at the top level echoing what was filtered for ("compound" / "atom" / "all" / etc.), and every chunk in the chunks (or items) array carries its own granularity field. Other envelope fields are unchanged from v1.4.
Citations gain granularity
GET /api/smc/citations/chunk/{idOrChunkId} — the chunk sub-object gains a granularity field:
{
// material citation fields (v1.3 shape) ...
"chunk": {
"chunk_id": "ie-010-…-0001",
"granularity": "compound", // v1.5: identifies which granularity this chunk was emitted at
"anchor": "ie-010-…#1-role-scope",
"heading_path": ["Role scope"],
"page_start": 3,
"page_end": 4,
"text_excerpt": "…"
}
}
All chunks created before β.1 are tagged compound automatically by the DB default — existing citations don't need a re-fetch, but consumers polling them will see the new field appear on the next response.
material.rechunked event — per-granularity counts
The event payload extends additively. old_chunk_count / new_chunk_count continue to carry Compound totals for v1.4 back-compat; new optional fields carry the full per-stripe map:
{
"event_type": "material.rechunked",
"target_key": "ie-010-…",
"rechunk": {
"old_chunk_count": 50, // v1.4 — stays as Compound total
"new_chunk_count": 52, // v1.4 — stays as Compound total
"old_chunk_counts": { "compound": 50 }, // v1.5 — per-granularity. β.1: compound only.
"new_chunk_counts": { "compound": 52 } // v1.5 — per-granularity. β.2+ adds keys.
}
}
A given map only carries keys for stripes SMC actually materialises. Consumers reading the v1.5 fields should treat missing keys as "this stripe wasn't recomputed (or doesn't exist) in this event", not as zero.
Embedding locus — clarification, not a change
SMC does not store embeddings. Consumers (RPF in v1.0, MAD in β follow-on, future BIMei KB / BIMd) own their own vector indices and pick their own embedding models. v4 § 1233-1323 anticipates a SMC-side retriever once BIMei KB onboards, at which point an embedding column on smc_chunks (keyed by (hash_sha256, embedding_model) — not chunk_id, which is positional and renumbers on re-convert; see v1.9 amendment item 3) becomes a contract amendment. Not in scope for v1.5. Consumers building their own external vector index today should read the v1.9 chunk-manifest projection (?fields=chunk_id,hash_sha256) for cheap index-freshness diffing.
What changed in v1.4
Additive endpoint. Non-breaking — exposes the full converted markdown of a material in one HTTP call so consumer apps don't have to reassemble paginated chunks or sign GCS URLs. Origin: MAD's MC-cutover migration (docs/plans/mad-migration-and-v4-alignment.md); the same endpoint is useful to any consumer that wants the full text without going through /chunks (BIMd, BIMei KB).
| Endpoint | Method | Returns | Notes |
|---|---|---|---|
GET /api/smc/materials/{key}/markdown | GET | 200 with the markdown envelope (see below) | Streams the markdown body inline from the latest matching artefact in GCS. No signed URLs. |
GET /api/smc/materials/{key}/markdown?profile=machine | GET | Same envelope shape; markdown is the normalized_markdown artefact | Default is ?profile=review (the human-readable review_markdown artefact, drop-in for MC's preview shape). ?profile=machine returns the LLM-ready normalized_markdown artefact. |
Response envelope (snake_case, matches the citation projection convention — and matches MC's /api/preview/{jobId} markdown field name so consumers migrating from MC don't have to remap the field):
{
"material_key": "ie-010-information-manager-bim-role-profiles-2025",
"markdown": "# Title…\n\nBody markdown…",
"profile": "review", // "review" | "machine"
"markdown_hash_sha256": "abc123…", // hash of the artefact bytes; null if missing
"conversion_id": "uuid",
"converter_version": "http-v1", // converter HTTP API version (today); null when the conversion row lacks it
"updated_at": "2026-05-20T12:00:00.000Z" // material.updated_at, ISO8601
}
The current converter_version value is "http-v1" — that's the markdown-converter's HTTP API version. The underlying MC service binary version isn't captured separately today; flagged as a Phase β open item ("conversion-pipeline metadata"). Consumers should treat converter_version as a stable identifier for "which API contract produced this markdown," not a binary version. Combine with markdown_hash_sha256 for change detection — MAD does this; see §2 "Idempotency anchors" below.
Idempotency anchors. Consumers that re-process material content should key off markdown_hash_sha256 from this endpoint (and chunk_id + hashSha256 from /chunks). SMC's own re-chunk detection uses the same chunk-set hash diff (lib/mutations/event-emitter.ts). If SMC ever changes its markdown-normalization step before hashing, the hash flips and consumers see a re-version — which is the intended behaviour, since the underlying bytes really did change.
Non-200 responses (load-bearing):
| Status | When | Body shape |
|---|---|---|
202 Accepted | Material exists but its conversion isn't ready yet (latestConversionId is null) | { "status": "pending", "review_status": "<canonical review_status>", "message": "Material has not been converted yet." } — consumer should poll, not back off to "not found" |
400 Bad Request | ?profile is set to a value other than review or machine | { "error": "Invalid profile: …" } |
404 Not Found | Material with the given key/id doesn't exist | { "error": "Material not found." } |
404 Not Found | Material exists but the requested profile's artefact is absent (rare; suggests a converter regression — different from "not ready yet") | { "error": "No <profile> markdown artefact available for this material." } |
The 202-vs-404 distinction matters for consumers: 202 means "come back later", 404 means "this won't appear by waiting". Don't collapse them.
No new mutation event. Markdown availability already rides on material.published/material.rechunked — see §2. Consumers that want push-notification of new markdown subscribe to those events and call /markdown on receipt.
What changed in v1.3
Additive field on the material payload. Non-breaking — existing consumers ignore unknown fields and stay green; consumers that want to surface it (RPF in v1; MAD optionally) read it from the same endpoints they already poll.
| Endpoint | Field | Where | Type | Notes |
|---|---|---|---|---|
GET /api/smc/materials | copyrightHolder | each item | text, nullable | Structured copyright owner string, e.g. "© BSI Standards Limited, 2024". Distinct from licenseOrUsageNote. Casing note: this endpoint passes the Drizzle row through verbatim, so the key is camelCase — same shape sharedMaterialKey, licenseOrUsageNote, reusePolicy, etc. already ship as. |
GET /api/smc/materials/{key} | copyrightHolder | material object | text, nullable | Same shape and casing as above. |
PATCH /api/smc/materials/{key} | copyright_holder (input) / copyrightHolder (response) | request body / material response | text, nullable, max 220 chars | Input shape is snake_case (Zod uses snake_case for all writeable material fields — shared_material_key, license_or_usage_note, etc.). The response is the same camelCase passthrough as GET. |
GET /api/smc/citations/material/{key} | copyright_holder | response | text, nullable | Citation endpoint has its own explicit snake_case projection (the rest of its keys — material_key, canonical_url, license_or_usage_note — are snake too). Included so consumers building immutable citation snapshots (§1) capture the holder. |
Why a new field instead of overloading license_or_usage_note: license_or_usage_note is free-form prose (usage caveats, attribution sentences, licence strings). Consumers asked for a parseable structured holder slot for citation rendering. Both fields coexist — holder for the structured "©" string, usage note for prose. SMC curators populate the field via /admin/materials/[key]/edit; expect null for most materials until backfill catches up.
Wire-shape convention on materials endpoints (clarification — applies retroactively from v1.1)
The materials list, detail, and PATCH endpoints (GET /api/smc/materials, GET /api/smc/materials/{key}, PATCH /api/smc/materials/{key}) pass the Drizzle row through verbatim, so all response keys are camelCase (sharedMaterialKey, licenseOrUsageNote, copyrightHolder, reusePolicy, updatedAt, latestConversionId, etc.).
One documented exception: review_status is snake_case on these endpoints. It's the only derived field — mapMaterialForApi strips the internal status column and emits the canonical consumer review_status enum per §6. The snake_case naming has been on the wire since the v1.1 amendment (2026-05-15) when the legacy status field was removed. RPF and MAD consume the field at this casing.
Consumer connector types should declare review_status separately from the camelCase fields (typed review_status: ReviewStatus rather than reviewStatus: ReviewStatus) to match the wire. Hand-normalization to camelCase is acceptable consumer-side but not required of the wire — standardising the wire would be a v2 break since both consumers depend on the current shape. v2 (when scoped) is the natural moment to normalise casing across the materials endpoints; until then the asymmetry is documented and stable.
The citation endpoint (/api/smc/citations/material/{key}) is a separate projection that emits all keys in snake_case (material_key, canonical_url, license_or_usage_note, copyright_holder); that convention is unchanged.
No new mutation event type. copyright_holder is editable metadata in the same class as title and publisher; per §1 consumers re-fetch on read against updated_at. If a consumer later needs change-push for this field, raise an Open item and we'll revisit.
What changed in v1.2
This amendment closes the last three open items by either shipping the work or naming the long-term posture. Once signed off, the consumer contract goes from "ongoing" to "stable" — open-items list empty.
Shipped:
- §2 mutation events — B2 + B3 complete.
source.deprecated,material.deprecated,material.superseded,material.rechunked, and thematerial.source_state_changedcascade now emit live. Admin UI for deprecate/supersede lives at/admin/sources/[id]and/admin/materials/[key]. Curator email is captured on B2-emitted events (legacy approve/reject path still emitsactor='system'— small follow-up if either consumer asks).
Closed by amendment (rather than by shipping):
- §3 topic briefs — GitHub-issue path is the long-term answer. SMC weekly triage handles current volume cleanly. The first-class endpoint stays unbuilt; it's a future capability if any consumer crosses >5 briefs/quarter (see §3 below for the named escalation trigger).
- §5 SLA — "best-effort for v1.x line" with an explicit reporting channel (GitHub issue labelled
sla-pain). Real numbers commit once SMC has steady-state load data or a consumer reports operational pain, whichever comes first. - §6 conformance npm runner — skipped for v1.x. The fixture endpoint covers the data half of the §6 commitment. The
@bimei/smc-contract-testrunner is reserved as a BIMei-KB / BIMd onboarding-bundle item for if either of them asks at onboarding time.
Dropped from active roadmap (still future v2):
- §7
Accept: application/vnd.smc.v1+jsonheader pinning — only relevant once v2 is scoped. Not a v1.x gap.
What changed in v1.1
This amendment closes the §6 Phase B harmonisation item: the legacy status field is removed from every consumer-facing material response, and the legacy ?status= query filter is dropped from GET /api/smc/materials. Consumers had migrated to review_status in writing before this amendment shipped (RPF PR #202; MAD b981d053).
Removed / renamed fields (grep your codebase for residual references — per MAD's amendment-shape ask):
| Endpoint | Field / param | Where | Replacement |
|---|---|---|---|
GET /api/smc/materials | status | response item | review_status |
GET /api/smc/materials | ?status= | query param | ?review_status= (canonical enum) |
GET /api/smc/materials/{key} | status | response material object | review_status |
PATCH /api/smc/materials/{key} | status | response material object | review_status |
GET /api/smc/materials/{key}/versions | status | each entry of versions array | review_status |
Source / source-request endpoints already emitted review_status in v1; only materials carried the legacy status name. No other endpoints affected.
Operating principle
SMC is the single curator for the BIMe Initiative. SMC decides whether a source or material is valid and still relevant. Each consuming app parses the curated corpus its own way — MAD for claims and evidence, RPF for competencies and project flows, BIMei KB for the conversational agent and GraphRAG, BIMd for localised term variants. No app duplicates curation work.
1. Consumer posture — copy vs reference
Reference by id, re-fetch on read. SMC owns source/material identity and review state; consumers own their parsed outputs (claims, competencies, terms, graph edges) keyed against shared_source_key and shared_material_key.
Shipping copies of the corpus into consumer DBs as a primary store is discouraged because it forks the source of truth on review state — review_status, visibility, and reuse_policy change continuously and silently invalidate any local copy. Consumers MUST NOT host their own canonical material corpus.
Short-TTL caching for performance is fine — consumer's call, but treat it as cache, not state, and re-validate against updated_at. The contract is: SMC is the only writer; consumers are read-through clients.
Immutable citation snapshots are explicitly allowed
Consumers that publish external artefacts (MAD's CSR reports, future BIMd embeddable widgets, audit trails, legal hold) MAY take point-in-time snapshots of cited materials so the published artefact doesn't silently change when SMC edits the source.
A snapshot:
- MUST sit alongside the reference-by-id binding (not replace it).
- MUST record the SMC version it was taken from (sufficient identifiers from the mutation event payload — see §2 — so the snapshot can be diffed against current state on demand).
- MUST be clearly labelled as a snapshot in any consumer-facing UI.
2. Change notifications & mutation events
Today (polling only)
GET /api/smc/sourcesandGET /api/smc/materialsreturnupdated_atand sort desc. Implement a "fetch since last seenupdated_at" loop on the consumer side.GET /api/smc/source-requests/recent-approvals?since=<ISO8601>&consumer_app=<app>returns a cursored delta of source-request approvals. Originally built for RPF's daily staleness scan; same shape works for any consumer.GET /api/smc/curation-requests/recent-fulfilments?since=<ISO8601>&consumer_app=<app>(v1.8) returns the sibling cursored delta for open-discovery curation-request fulfilments, each carrying thematerial_keysto bind. Samesince→cursor_atpattern. See the v1.8 amendment.
Direction: generic /mutations?since= cursor
SMC commits to extending the cursor pattern into a single generic mutation endpoint for sources+materials (same shape as /recent-approvals) before adding webhooks. Cursors compose with backfill and reconciliation in a way webhooks don't. Webhooks may come after if a consumer hits a latency floor that polling can't meet.
Hard constraint: whatever shape lands has to be one shape — the same endpoint that MAD uses must work for RPF, BIMei KB, and BIMd.
Minimum mutation event shape
When /mutations?since= ships, each event has at least these fields:
{
"event_id": "uuid", // for idempotent consumer processing
"event_type": "source.approved" | "source.deprecated" | "source.superseded"
| "source.reuse_policy_changed" | "source.visibility_changed"
| "material.published" | "material.rechunked" | "material.deprecated",
"target_kind": "source" | "material",
"target_key": "<shared_source_key | shared_material_key>",
"occurred_at": "ISO8601",
"actor": "<curator_id | \"system\">",
"old_state": { "review_status": "...", "visibility": "...", "reuse_policy": "..." } | null,
"new_state": { "review_status": "...", "visibility": "...", "reuse_policy": "..." },
"supersedes": "<shared_material_key | null>", // for supersession events
"rechunk": { "old_chunk_count": N, "new_chunk_count": M } | null,
"reason": "<free-text curator note | null>" // optional; supported when curator provides one
}
The shape gives consumers enough information to:
- process events idempotently (via
event_id); - render a "this source has been updated, pending review" UI banner without a second fetch (
old_state,new_state,reason); - decide which downstream entities to flag (
target_key+event_type); - short-circuit redundant scanners (e.g. RPF's chunk-hash drift check becomes unnecessary once
material.rechunkedevents are emitted).
Cadence floor
Consumers determine their own polling cadence. SMC will give ≥30 days notice before introducing per-consumer rate limits. Pairs with §5's "rate limits hardcoded infinite today" caveat.
Normative: SMC events are advisory
SMC mutation events are advisory. Consumers MUST flag affected downstream entities for human review and MUST NOT auto-modify their own review state in response to SMC events.
SMC publishes facts; consumers decide what to do with them. This is the "two-layer review" of §6 in motion: SMC's state change is information, not a command.
Non-deletion guarantee
SMC MUST NOT hard-delete published sources or materials. State transitions to deprecated or archived are the supported removal mechanism — consumer bindings retain stable references regardless of lifecycle state.
Legal redactions (GDPR, takedown notices, defamation orders) are out-of-band: SMC will notify affected consumers before redacting, with notice period bounded by the legal requirement. This is the only path to a hard delete; routine curator workflows never hard-delete.
3. Discovery requests
Today
POST /api/smc/source-requests with { url, reason, suggested_country_code, suggested_topic_watchlist }. URL-only — the consumer (or its agent) finds the URL, curators in SMC review, promote to candidate source, run the approval workflow.
Discovery channel for topic briefs (v1.2: long-term)
Topic-level briefs ("find me sources on topic X in jurisdiction Y") are the canonical consumer-driven discovery path. SMC — not each consumer — is the right place for that work because the curation-non-duplication principle requires a single curator across BIMei apps.
Open a GitHub issue at changeagents-aec/bimei-smc with the topic-brief label. Include:
- Topic and any controlling terminology.
- Jurisdiction(s) of interest.
- Urgency (rough due date if relevant).
- Consumer context (which app, which downstream need this serves).
SMC triages weekly. The threaded conversation between curators and the consumer (where clarifying questions naturally happen during a brief) makes this materially better than a one-shot POST endpoint at current volume.
Escalation trigger (kept on the books per RPF's v1.2 ask)
If any single consumer crosses >5 briefs per quarter, the GitHub-issue path stops being a good fit and SMC formalises POST /api/smc/topic-briefs as a first-class endpoint. Named here so a future consumer (e.g. a chatbot UI that needs programmatic submission) has a clear, auditable escalation path without re-litigating the decision.
What SMC will not do: build a separate consumer-side discovery channel per app. That re-introduces the duplication this contract is meant to prevent.
4. Per-consumer "I no longer care about this source"
Strictly the consumer's job. SMC stays a clean global corpus; per-consumer dismissal/subscription state lives in the consumer's own DB, keyed by shared_source_key. Two reasons:
- The corpus is small enough that consumer-side filtering is cheap, and consumer relevance criteria differ enough (claims vs competencies vs terms) that SMC can't model them sensibly.
- Per-consumer state on SMC's side expands the contract surface in a direction that doesn't scale across BIMei apps.
If a consumer finds itself wanting SMC to model this, treat it as a signal something else is off and raise an Open item — we'll dig in.
5. SLA on mutations
Best-effort, no formal SLA in v1. Concretely:
- Mutations are visible to consumers the moment the curator action commits (sub-second on the DB; no async pipeline between curator click and consumer read).
- Availability is best-effort — single Cloud Run service in
us-central1, no replication tier yet. - There is no general authenticated REST-read or MCP tool-call quota in v1. The MCP's shared per-instance failed-authentication bucket is a security throttle, not a service-level rate commitment. See §2's cadence floor for the ≥30-day notice commitment before introducing a general consumer quota.
Posture for v1.x (closed in v1.2)
Best-effort for the v1.x line; no committed numbers. Without steady-state production load data from MAD + RPF, any p99/availability number we'd post here is fiction. The contract closes this loop by naming an auditable reporting channel for the only thing that actually matters — operational pain on the consumer side.
Trigger to commit numbers:
- Either consumer reports operational pain via a GitHub issue at
changeagents-aec/bimei-smclabelledsla-pain. The issue triggers a renegotiation cycle: SMC reads operational metrics, the consumer reports observed impact, both parties agree on numbers that go in this section. - Annual review. Whichever comes first.
sla-pain is the SLA equivalent of §3's topic-brief label — a named, auditable channel so a "consumer reports operational pain" verbal commitment becomes a real trail.
Contract-bump trigger (active once numbers are committed)
Once availability or latency numbers are committed, missing a committed target for two consecutive months returns this contract to draft and triggers a renegotiation cycle. This keeps committed numbers operational, not aspirational.
Status page
A BIMei-wide status page documenting live availability and incidents will land alongside the committed numbers. Link to be added here.
6. Review state vocabulary
SMC and each consumer review different things against the same vocabulary. SMC reviews sources and materials; consumers review their own derived entities (RPF: competency bindings; MAD: claims; BIMei KB: graph nodes; BIMd: term variants). Each layer has its own review_status column; the enum values are shared with SMC and changes to the enum are contract-breaking.
Canonical review_status enum (v1)
| Value | Meaning | SMC entities | Consumer entities |
|---|---|---|---|
requested | Submitted by a consumer; awaiting SMC triage | smc_source_requests only | n/a (transient, SMC-side) |
candidate | In SMC's review queue or in a consumer's pre-approval state | sources, materials | yes |
approved | Reviewed and accepted as valid by the owning party | sources, materials | yes |
rejected | Reviewed and not accepted | sources, materials | yes |
deprecated | Was approved; now outdated or no longer current. Bindings remain valid; consumers flag for review. | sources, materials (roadmap — ships with /mutations) | yes |
superseded | Replaced by another approved entity. Successor referenced in the mutation event. | materials (roadmap — ships with /mutations) | yes |
archived | Removed from active workflows but retained for reference | sources, materials | yes |
Notes
- Per-entity applicability:
supersededapplies to materials only — sources have no supersession concept (a deprecated source's replacement is a new source, not an inheriting successor). All other states apply to both sources and materials where indicated in the table. - Mandatory mirroring: consumers MUST use these values for their own
review_statuscolumns. Custom values MAY be added by a consumer for states SMC does not model (e.g. RPF's priorsystem_draft) but MUST NOT collide with the canonical names above. - Internal vs public — Phase A (shipped 2026-05-15, retired by v1.1 cutover): intermediate state where the API returned both
statusandreview_statusadditively. Documented for historical interest; no longer reachable. - Internal vs public — Phase B harmonisation half (shipped 2026-05-15 with the v1.1 amendment): legacy
statusfield removed from every consumer-facing material response;?status=filter dropped. Onlyreview_statusis exposed on the wire. See the "What changed in v1.1" table at the top of this doc for the grep-friendly list of removed fields. The mapping from internalsmc_materials.statusto canonicalreview_statusis:accepted→approved;archivedandduplicate→archived;candidate,processing,ready_for_review,needs_manual_verification→candidate;rejectedunchanged. - Internal vs public — Phase B mutation half (shipped 2026-05-16 with the v1.2 amendment):
deprecatedandsupersededare live internal states for materials;deprecatedis live for sources. All four lifecycle event types now emit (source.deprecated,material.deprecated,material.superseded, plus the existing cascadematerial.source_state_changedthat fires automatically whensource.deprecatedwrites). B3 also lights upmaterial.rechunkedevents from the chunker pipeline. Curator email is captured on these events via thewithActor+ session-var pattern; the legacy approve/reject path still emitsactor='system'(low-priority follow-up if either consumer asks). - Other SMC enums (
visibility,reuse_policy): SMC-internal by default. Consumers MAY mirror them when surfaced to operators or used for filtering; onlyreview_statusis mandatory.
Change policy for the canonical enum
Any addition, rename, split, or removal of a canonical value requires:
- A PR to this contract proposing the change.
- Consumer review by all approving consumers; blocking objections must be resolved before SMC ships.
- For removals: a deprecation period during which the old value is still emitted (with a documented mapping to the replacement) but flagged.
Conformance testing (closed in v1.2)
The fixture endpoint is the canonical conformance surface for v1.x. GET /api/smc/mutations/fixture (public, no auth) returns a deterministic JSON response with one synthetic event of every canonical event_type. Consumers point CI directly at the URL and assert against a committed snapshot — that's what both MAD and RPF do today.
The top-level _fixture: true marker exists so consumer CI can hard-assert the cursor is on the fixture URL (defence against a misconfigured cursor silently consuming synthetic data through production code paths).
@bimei/smc-contract-test npm runner — reserved, not built. The package would centralise assertions in a single installable runner for TS-based CIs. With MAD and RPF the only consumers today and both already conformant via curl-and-assert, the centralising hedge isn't worth the build cost. Reserved as a BIMei-KB / BIMd onboarding-bundle item — if either future consumer asks for it during onboarding, SMC ships the runner. Until then, the fixture endpoint stands.
7. Authentication, identity, and versioning
API versioning
The HTTP API is versioned via media-type header. Consumers pin to v1 by sending:
Accept: application/vnd.smc.v1+json
This contract version tracks the API version. When the API moves to v2, this contract version bumps to v2 in lockstep. The v1 → v2 migration policy (parallel-serve window, deprecation period, breaking-change handling) will be documented in this contract when v2 is scoped — not pre-committed today.
Consumer identity in v1
X-SMC-Consumer-App: <app> is a tag, not authentication. SMC v1 uses a single shared super-key (X-SMC-API-Key); the consumer-app header is for routing, observability, and audit attribution only. Either consumer can technically claim to be the other; SMC does not enforce identity in v1.
This is acceptable in the BIMei v1 trust environment (single org, all consumers operated by the same team). It is not acceptable for multi-tenant or external-partner scenarios — those scenarios require the v1.5+ work below.
Consumers should design their integrations assuming the consumer-app header is not a security boundary in v1. Treat X-SMC-API-Key as a shared operational secret, not a per-consumer credential.
Per-consumer authentication — delivered in v1.13
Superseded. Enforced per-consumer credentials shipped in v1.13; see "What changed in v1.13" for the capability model, organisation scoping and the fail-closed default. The text above describes the shared key, which is unchanged and is still what MAD and RPF hold.
The commercial hooks the original roadmap bundled this with (per-app billing, quota, tiered access — v1 MVP spec §1.2) remain deferred. v1.13 delivers the security boundary only.
Open items
O-v1.13-1 — Positional passage locators are absent for every format. Raised by Build Digital at onboarding (2026-08-10), then found to be broader than reported.
smc_chunks has page_start, page_end, char_start and char_end columns, and they are serialised on the wire — but the chunker never populates them. Verified 2026-08-10 against production across PDF-backed materials in several sources (ISO 19650-1, ISO 19650-2, CEN/TR 17654, and Build Digital's own PDFs): null on every chunk. Spreadsheets additionally have no sheet-name or cell-range column at all.
So the gap is not, as first stated, XLSX-versus-PDF. SMC currently provides no positional locator for any format. The only passage locator it does provide is the heading anchor — anchor plus heading_path — which is stable and section-accurate but cannot express a page or a cell.
Two independent pieces of work, one now closed:
PopulateCLOSED 2026-08-10. The converter already emits a page separator between consecutive pages, so the information was recoverable from the markdown without a converter change. Derivation runs at conversion time, and a backfill populates existing chunk rows in place — updating two integer columns, never re-converting, sopage_start/page_endfor PDF conversions.hash_sha256andpassage_version_idare untouched and existing citations stay bound.- Add spreadsheet sheet/cell-range locators. STILL OPEN. Spreadsheets have no page concept and the office-native conversion path emits no separators, so those chunks remain null by design rather than by omission. Needs a schema change plus a locator-shape amendment. Owner SMC, unscheduled.
Two rules for consumers, both still load-bearing:
- Null means "no locator", never "page 1". Unpaginated documents and unlocatable passages are deliberately left null rather than given a fabricated position. A wrong page in a citation is worse than an absent one, because a consumer cannot tell it is wrong.
- Where a passage has no page, the heading anchor is the locator.
anchorandheading_pathare populated for every chunk and attribute an excerpt to a named section.
Everything raised through the v1 sign-off process remains closed — either by shipping it (B2 + B3) or by naming the long-term posture explicitly in the doc (§3 GitHub-issue path with the >5/quarter escalation trigger; §5 sla-pain channel + "best-effort for v1.x"; §6 fixture endpoint as canonical conformance surface).
Reserved for BIMei KB and BIMd at onboarding time — those consumers will add their own items via PR amendments before signing on as approving consumers. The @bimei/smc-contract-test npm runner is held as a candidate onboarding-bundle item for either of them.
Change process
This contract evolves through PRs in the SMC repo at docs/contracts/smc-consumer-contract-v1.md. Process:
- A consumer team opens a PR amending the relevant section or adding to Open items.
- SMC acks or negotiates in PR review; other approving consumers comment.
- Merge requires sign-off from SMC plus all approving consumers — no outstanding blocking objections.
- Bump the version (
v1→v1.1) when shape changes affect existing consumers.
Approving consumers
| Consumer | Status (v1) |
|---|---|
| MAD | approving |
| RPF | approving |
Build Digital (build_digital_kb) | onboarded in v1.13 — scoped read-only credential |
| BIMei KB | future — approves at onboarding via amendment |
| BIMd | onboarded, non-approving — holds integration guides and calls production today |
Amendment acceptance ledger
Amendments ship as proposed and are implemented immediately — SMC does not block additive changes on sign-off, because a consumer that has not read an amendment is unaffected by an additive one. But "proposed" was accumulating with no way to resolve, so v1.6 through v1.15 all read as outstanding regardless of whether anyone had actually looked.
This ledger closes that. A row moves to acknowledged when a consumer team
confirms they have read it, and to signed when an approving consumer
formally accepts. An amendment that is still proposed months later is not
necessarily a problem — but it should be a visible one.
| Amendment | Shipped | MAD | RPF | BIMd | Build Digital |
|---|---|---|---|---|---|
| v1.6 material types, doc chat, normativity | 2026-06-23 | acknowledged | acknowledged | n/a | n/a |
| v1.7 curation-request intake | 2026-06-30 | proposed | acknowledged | n/a | n/a |
| v1.8 fulfilment material keys | 2026-06-30 | proposed | acknowledged | n/a | n/a |
| v1.9 chunk-manifest projection | 2026-07-09 | proposed | acknowledged | proposed | n/a |
| v1.10–v1.12 | 2026-07 | proposed | proposed | proposed | n/a |
| v1.13 scoped credential | 2026-08 | proposed | proposed | proposed | signed |
| v1.14 | 2026-08 | proposed | proposed | proposed | proposed |
| v1.15 alpha-2 country codes | 2026-08-14 | proposed | proposed | proposed | proposed |
| v1.16 page and worksheet-range previews | 2026-08-16 | n/a | n/a | n/a | accepted; three-workbook pilot approved (2026-08-16) |
| v1.17 release-pinned classification path (inactive) | 2026-08-18 | proposed | proposed | proposed | proposed |
v1.18 academic_article_or_chapter material type | 2026-08 | proposed | proposed | proposed | proposed |
v1.19 material_type is an open vocabulary | 2026-08 | proposed | proposed | proposed | proposed |
| v1.20 vocabulary read endpoint | 2026-08-20 | proposed | proposed | proposed | proposed |
v1.21 grant-is-the-nomination; reuse_decision correction | 2026-09-02 | n/a | n/a | n/a | accepted (2026-09-02, BD-H/BD-J reply) |
| v1.22 read-only remote MCP | 2026-09-03 | proposed | proposed | proposed | proposed |
| v1.23 MCP projection, audit, and deployment hardening | 2026-09-03 | proposed | proposed | proposed | proposed |
n/a means the consumer was not onboarded when the amendment shipped.
The v1.18–v1.20 rows were added on 2026-09-02 from the document's own header, not
from any consumer confirmation — the ledger had stopped at v1.17 while three
further amendments shipped. The v1.22 row was added on 2026-09-03 when its
operating documentation was completed. The v1.23 row received its shipped date
after revision smc-app-00349-l6c was observed serving the merged hardening
release. Authenticated initialization and the direct-protocol subset passed;
fresh in-client calls and direct audit/monitor inspection remain separate gates.
Every proposed cell still means that no consumer-team acknowledgement has been
evidenced.
How to update a row: change the cell, and say in the PR who confirmed and
where. The ledger is only worth having if its cells mean something — a row moved
to acknowledged because it felt stale is worse than leaving it proposed.
Signed-off by MAD and RPF on 2026-05-15 (v1 + v1.1 amendment), 2026-05-16 (v1.2 amendment), 2026-05-26 (v1.3 amendment — sign-off thread: #79), and 2026-05-26 (v1.4 amendment — sign-off thread: #80). The consumer contract is stable. /api-docs links here as the canonical contract.
Resolved-in-v1 log
For reviewers tracking what changed since the first draft:
- Original O1 (Two-layer review = same vocabulary, different fields) — resolved into §6 with the canonical enum pinned.
- Original O2 (Scanner flags, never auto-fixes) — resolved into §2 (generic
/mutations?since=direction, minimum event shape, normative MUST line on advisory events, non-deletion guarantee). - MAD's immutable-citation-snapshots ask — resolved into §1.
- MAD's no-hard-delete ask — resolved into §2 non-deletion guarantee.
- MAD's pin-the-enum ask — resolved into §6 canonical enum table.
- RPF's cadence-floor ask — resolved into §2 cadence floor.
- RPF's name-the-interim-discovery-channel ask — resolved into §3 (GitHub issue with
topic-brieflabel). - RPF's contract-bump-trigger ask — resolved into §5.
- MAD's no-local-material-corpus clarification, status-page placeholder, approving-consumers normative line — all resolved.
- RPF O3 (API versioning) — acked into §7 with Accept-header pinning; v1 → v2 sunset policy deferred to v2 scoping.
- RPF O4 (Consumer identity) — acked into §7 as convention-only for v1; per-consumer keys land with v1.5+ commercial hooks.
- RPF O5 (Contract test fixture corpus) — initial ack into §6 Conformance testing; the fixture endpoint shipped in v1.1 as the data half; the runner closed in v1.2 as "reserved for future-consumer onboarding bundle."
Resolved-in-v1.2 log
For reviewers tracking what changed since v1.1:
- §2 mutation events — B2 + B3 complete (shipped 2026-05-16). Lifecycle states
deprecated(sources + materials) andsuperseded(materials) are live; the corresponding event types emit on transition; the cascadematerial.source_state_changedfires automatically;material.rechunkedemits from the chunker pipeline on re-conversion when the chunk set changes. Admin UI panels at/admin/sources/[id]and/admin/materials/[key]. - §3 topic-brief endpoint (closed in v1.2). GitHub-issue path adopted as long-term answer. >5/quarter trigger language preserved as the named escalation path (RPF's v1.2 ask). The "expected before BIMd onboards" framing is dropped — BIMd is not a forcing function.
- §5 SLA numbers (closed in v1.2). "Best-effort for v1.x line" with an explicit
sla-painGitHub-issue reporting channel (RPF's v1.2 ask). Real numbers commit on operational pain or annual review, whichever comes first. - §6 conformance npm runner (closed in v1.2). Fixture endpoint is the canonical conformance surface for v1.x.
@bimei/smc-contract-testnpm runner reserved as a BIMei-KB / BIMd onboarding-bundle item. - §7 Accept-header API versioning (deferred). Not relevant before v2 is scoped. Dropped from active roadmap.
All raised items now closed. The consumer contract is at v1.2 stable; open-items list empty.