10. Conformance
10.1 Conforming processor
A conforming processor is an implementation that, for every input:
- produces a result without aborting (§2.4);
- applies the §4 normalization and the §5–§7 model; and
- satisfies every
must-level expectation of the conformance suite (§10.3).
A processor need not implement every projection (some have no HTML renderer,
some no serializer); it conforms with respect to the projections it produces.
A processor MUST NOT emit, for a must-level vector, a projection that
contradicts that vector.
10.2 The role of this specification (master)
This specification is designed to be the single source of truth a
processor conforms to. The normative prose, the examples, and the
machine-readable suite are kept in one repository so they cannot drift: every
normative example in §6 corresponds to a vector, and a consuming processor
SHOULD pin a tagged release (vX.Y.Z) of this repository — a git revision
also works — and run the suite in its own CI, failing on any must mismatch.
10.3 Conformance suite
The suite lives at conformance/vectors/. Each case is a vector.json
validating against conformance/schema/vector.schema.json and carrying a
source plus the expected projections (nodes, pairs, serialize,
diagnostics, and optionally html). The runner contract — how a processor
consumes a vector, which projections are compared, and how matches are scored
by level — is conformance/RUNNER.md.
Levels
| Level | Obligation |
|---|---|
| must | Every comparable projection MUST match exactly. A mismatch is non-conformance. |
| should | Projections SHOULD match; a documented mismatch is a warning, not non-conformance. |
| may | Informational; divergence is permitted. |
For a must vector, the nodes, pairs, and diagnostics projections are
mandatory comparisons; serialize is compared when the processor has a
serializer; html is compared when the processor has the reference renderer
(§8) and the vector's level pins it.
10.4 Claiming conformance
An implementation claiming conformance SHOULD state:
- the specification version it targets (a
vX.Y.Zrelease tag; or the git revision); - which projections it produces;
- its results against the suite, including any justified
should/maydivergences.
10.5 Coverage and growth
The §6 catalogue covers the core notation families. The boundary below is
grounded in a full sweep of every work in
P4suta/aozorabunko_text
(17,889 files, re-measured 2026-06-25) — parsed and bucketed by the reference
implementation, not estimated. Counts are occurrences across that sweep. An
earlier draft of this section misreported several of them (noted inline); they
are corrected here from the measured data. The most recent #78 burn-down
promoted the compound-indent modifier set (§6.6) and the structural-marker
leaves (§6.19), dropping the residual Unknown (generic-annotation) total from
3418 to 2848.
Recognised, pending full §6 promotion
These corpus-attested forms are recognised by the reference implementation (typed distinctly, not the generic-annotation fallback). They are listed here because their full normative §6 text and vectors are still being written; until then a conforming processor MAY treat them as generic annotations (§6.14), but the recommended behaviour is the typed form below.
- File-header 凡例 symbols — the de-facto-standard legend that prefixes
nearly every work:
[#](入力者注, ~14k occurrences),[#…](返り点),[#(…)](訓点送り仮名). Empty / placeholder directives — not unrecognised notation. Typed distinctly so they leave the generic-annotation bucket. - Bare-range forms of families whose
ここから…/ここで…終わりblock form is already in §6:[#{N}段階…文字]…[#…文字終わり](font-size, ~15k),[#横組み]…終わり(~3k),[#行右/左小書き]…終わり(~8k),[#キャプション]…終わり(~2.5k). - Block
罫囲み/割り注(theここから…region forms; the bare[#罫囲み]/[#割り注]remain inline),天から{N}字下げ(top-origin indent), and the bare改行天付き、折り返して{N}字下げhanging indent. - Range bouten
「X」~「Y」に<kind>(also〜) — marks the whole preceding run from X to Y. - Input-editor notes:
「X」はママ(sic),「X」は底本では「Y」(source divergence),「X」に「Y」の注記(side annotation — see the correction below), numbered illustrations[#挿絵{N}(…)入る], and caption-before figures「caption」のキャプション付きの(図|挿絵|写真)(file)入る.
Promoted in the #78 burn-down
Two families that this section previously listed as outside the model are now
normative, each with a must vector:
- Structural-marker leaves (§6.19) —
[#本文終わり](body end, 242 occurrences, the #1 unrecognised body before this change) is a block leaf;[#改行](forced break, 165) is an inline leaf. Seebody_endandforced_break. - Compound-indent modifier set (§6.6) — the decorative clauses an indent
opener may stack after
ここからN字下げ、(ゴシック体,横書き/横組み,罫囲み,小さい活字), order-independent on input and re-emitted in a fixed canonical order, are now §6.6-normative (no longer deferred). Page-centring within an indent (ページの左右中央に) is part of the same compound form and is likewise supported. Seeindent_compound_styled. The canonical order, spellings, and the modifier set are fixed by ADR-0004 (extending ADR-0003).
Closed as a boundary (kept lossless-generic)
These forms were reviewed in the #78 burn-down and closed for want of corpus
demand: a conforming processor retains the whole [#…] as a generic annotation
(§6.14), round-tripping byte-exact. The measured occurrence is recorded so the
boundary is auditable.
- Rare bouten in bare-range form —
[#鎖線]…[#鎖線終わり]and the 破線 / 黒三角傍点 variants: 0 occurrences. The forward-reference form (「X」に鎖線の傍点etc.) stays normative; the bare-range deferral is recorded in ADR-0004. - Right-side ruby / annotation — the right side is already the default
|《》ruby, so an explicitの右に…のルビbracket is redundant. Theの右に…のルビandの右に…の注記forms have 0 occurrences (an earlier "≈8" was wrong); the attested side-annotation form「X」に「Y」の注記is recognised (above). - Table cell / row structure —
[#ここから表]marks a region; there is no sanctioned cell or row delimiter (the full-width/is an editor convention).段間に罫appears just once across the sweep and上段/下段never appear as directives (0) — far too sparse to pin a sub-region model. - Column sub-regions —
上段/下段/段組み適用外as in-region directives: 0 occurrences. The段組みregion itself is recognised; it carries no sanctioned sub-region split.
Still deferred
- Block centring with no closer —
ここから中央揃え(0) and the attestedここからページの左右中央(11) appear in the corpus without a matching closer, so they do not fit the paired-container model and remain generic annotations. (Page-centring within an indent is supported — §6.6, above.) - 地寄せ (margined right-align) and embedded-directive compounds —
[#ここからN字下げ、「」は返り点],[#ここから字下げ、ここから数式]and the like: any unrecognised clause triggers the decline-whole rule, so the whole directive is retained losslessly and reportsunrecognised-container-directive. - Left-ruby block form (
[#左にルビ付き]…[#左に「X」のルビ付き終わり]) — the paired-block counterpart of the forward-reference left-side ruby (which is normative). ~2 occurrences, and the reading sits in the closer, so it does not fit the streaming per-run container machinery. Pinned only on demand. ローマ数字、面-区-点gaiji form — needs a node the model does not yet carry; left as a generic annotation until promoted.
The full sweep leaves a residue of generic-annotation occurrences (down from
~194k before successive rounds of corpus grounding). The remainder is
overwhelmingly correct as a generic annotation: open-ended 入力者注 free-text
prose (not a fixed notation) and forward references whose target is genuinely
absent — legend examples like (例)[#「第一章」は中見出し], where the quoted
target never occurs as a preceding run. (A heading whose title merely carries
ruby is not such a case: its target is present, only ruby-split, and is
resolved ruby-stripped per §7.5 — see heading_ruby_hint.)
As a family in the first group gains full normative text it gains a vector, recorded in Annex E. Coverage is tracked openly with full-corpus evidence rather than implied — a processor is never asked to match a behaviour this document has not pinned with a vector.