Skip to content

Compose the Java XML serialization (#691) - #691

Merged
mristin merged 1 commit into
mainfrom
mristin/Compose-Java-XML-serialization
Sep 15, 2026
Merged

mristin merged 1 commit into
mainfrom
mristin/Compose-Java-XML-serialization

Conversation

@mristin

@mristin mristin commented Sep 15, 2026

Copy link
Copy Markdown
Contributor

Write-side dual of #690, as #687 was of #686 in C# and #689 of #688 in Go.

visit{Cls} was a verbatim copy of serializeElement: 38 methods of 18 lines each spelling out the start tag, the namespace, the content and the end tag, around a call which already did exactly that. Six near-identical emitters -- one per property kind -- produced the very same serializeElement call, differing only in the content serializer. And everything was an instance method, because of one mutable topLevel flag.

That flag is what cost the most. It made this::writeStringifiedContent, this::visit and (value, w) -> this.xToSequence(value, w) capturing, and serializeItems(...) and asNamedElementSerializer(...) returned capturing lambdas on top of them, so writing one instance allocated a small pile of closures. It is replaced by a final boolean withNamespace pinned in the constructor of two immutable visitors, ROOT and NESTED: only the outermost element declares the namespace, and an element written from writeClass is nested by construction. Everything else is then static, and a method reference to a static method captures nothing.

The design is otherwise the read side's, mirrored:

  • One shape, not two. ContentWriter<T> writes a value where the writer already is; writeElement frames it in a start and an end tag. There is deliberately no second delegate for a whole element, as there is on the reading side -- an element differs from a content only in what it writes, never in its shape. The reading needs the distinction because a content reader has to be told whether its element was self-closing, and a writer has nothing to be told. Use sites take ContentWriter<? super T>, which is how Java spells the contravariance C# writes as in T, so the single writer of an IClass serves wherever the writer of a more specific interface is expected.

  • visit{Cls} collapses to one call, and the visitor becomes a dispatch shim::

    @Override
    public void visitQualifier(
      IQualifier that,
      XMLStreamWriter writer) {
      writeElement(
        "qualifier",
        that,
        writer,
        withNamespace,
        _VisitorWithWriter::writeQualifierAsSequence);
    }
    
  • The families are dual to the readers and named by the same moniker grammar: write{Cls}AsSequence, writeClass/writeUnion, writeListOf_{M}, writeTupleOf{N}_{M...}, writeAtV_{M} and writeAtV{i}_{M}. The list and the tuple writers inline their loop, which retires serializeItems, asNamedElementSerializer and serializeTuple{N} -- Java needs no partial application here, since the name to bind is fixed per generated function.

  • One property emitter replaces the six. Every property is written by writeProperty, or by writeOptionalProperty, which takes the Optional apart once instead of asking it and then unwrapping it -- the old form called the getter twice, and every call allocates an Optional.

  • _type_moniker, _leaf_moniker, _is_instance_type, _is_dispatched and _collect_needed are reused unchanged: the reading and the writing are composed out of the same shapes, so one gating pass answers for both. Only the two dispatching writers needed a walk of their own, because the position matters on the write side while it does not on the read side: an item of any class is written as its own element, whereas a property of a concrete class without descendants writes its properties straight into the property element.

Two things are genuinely asymmetric, and both save work:

  • There is no writeTextAs_{primitive} family. Once the value is at hand nothing type-specific is left to do with it, so the four toString primitives share writeStringifiedContent and the byte array writeByteArrayContent. An enumeration does have something of its own to write and keeps one function each -- named writeTextAs_{Enum} and not write{Enum}Content, so that it stays in the moniker-keyed name space: were it keyed on the symbol, an enumeration which somebody names Stringified or ByteArray would silently take a shared writer's name.

  • There is no element writer per class. Which element an instance writes follows from its run-time type, and one virtual call answers that for every class at once. The reading needs a dispatcher per interface because it has to decide what to construct before it has read anything.

Two defects of the write side are fixed along the way, both of which the decomposition makes one-line changes rather than sweeping ones.

Serialize.to never flushed, so the serialization was not finished when it returned, and a failure of the underlying stream surfaced at the caller's own flush as a raw XMLStreamException -- outside the one method whose throws clause promises a SerializeException. Probed with a Writer which accepts every character and fails on flush, to used to return normally and now throws. The flush goes in Serialize.to and nowhere else, as in Go's #689, so XMLStreamWriter keeps buffering in between. Mind that this does push the underlying stream: a caller serializing many small documents into one buffered writer now flushes once per document rather than once per buffer. That is the right trade, since the alternative leaves the output silently incomplete whenever the caller forgets.

Every SerializeException carried an empty path, so every failure read ... at: the beginning no matter where it happened. The path is now built as the stack unwinds, each container prepending the one segment it knows, exactly as on the read side where read<Cls>FromSequence prepends the property after its switch and readList prepends the index. Probed with an XMLStreamWriter which refuses to write one chosen value:

a property of the root instance   idShort
nested through two lists          submodels/*[1]/submodelElements/*[1]/value
through a class-typed property    semanticId/keys/*[0]/value

SerializeException renders its message in its constructor, so its path has to be complete by the time it is built and it can not be the thing thrown from inside; a private _SerializeFailure carries a Reporting.Error instead, and Serialize.to renders that with Reporting.generateRelativeXPath and converts. The public exception is untouched. writeElement contributes no segment and writeProperty is writeElement plus the one thing a property knows which nothing below it does -- its own name -- so there is no try around any of the 276 property writes. A list and a tuple put their try outside the loop and advance the index only once an item has been written, so the item named is the one which failed and nothing is paid per item.

Two segments are deliberately absent, and the reasoning is written into the generated Javadoc, since both look like oversights. The outermost element is not named, because Serialize.to takes any IClass and the name would say nothing the caller does not already know; the read side does prepend it, having one facade per class. Neither is the discriminator element of a polymorphic property, which the read side does prepend -- for a failure in an embedded data specification it reports submodel/embeddedDataSpecifications/*[0]/dataSpecificationContent/ dataSpecificationIec61360/preferredName/*[0] where the write side reports embeddedDataSpecifications/*[0]/dataSpecificationContent/ preferredName/*[0]/text. The reading points into a document it is parsing, where that element is a real extra level; the writing points into the instance the caller handed over, where it is not, and value/idShort there is exactly getValue().getIdShort(). Go's #689 made the same choice when it rendered its serialization paths through aasreporting.ToGolangPath rather than as XPaths.

For the AAS meta-model the write region of Xmlization.java goes 3,281 -> 2,873 lines (-12.4%), error paths included. The write region no longer contains a single lambda.

Unlike the read side, that shows up in the measurement. Serializing an Environment of 200 submodels (1.14 MB of XML) into a writer which counts the characters and discards them, 100 runs after 100 warm-up runs,
eight rounds of each variant run alternately: the allocation goes 0.76 ->
0.36 MB per serialization, and the time 7.69 -> 7.49 ms, which is inside the round-to-round noise. XMLStreamWriter writes straight through, so the whole write path allocates ~0.8 MB against the ~49 MB the reading of the same document allocates -- there is no StAX floor here for the closures to hide behind, which is precisely why #690 could not measure its own.

The snippet contract for an implementation-specific class changes: Xmlization/VisitorWithWriter/visit_{cls}.java is no longer read, as the element around the sequence is now written identically for every class, and the remaining {cls}_to_sequence.java has to define a static write{Cls}AsSequence.

Write-side dual of #690, as #687 was of #686 in C# and #689 of #688 in Go.

``visit{Cls}`` was a verbatim copy of ``serializeElement``: 38 methods of
18 lines each spelling out the start tag, the namespace, the content and
the end tag, around a call which already did exactly that. Six
near-identical emitters -- one per property kind -- produced the very same
``serializeElement`` call, differing only in the content serializer. And
everything was an *instance* method, because of one mutable ``topLevel``
flag.

That flag is what cost the most. It made ``this::writeStringifiedContent``,
``this::visit`` and ``(value, w) -> this.xToSequence(value, w)`` capturing,
and ``serializeItems(...)`` and ``asNamedElementSerializer(...)`` returned
capturing lambdas on top of them, so writing one instance allocated
a small pile of closures. It is replaced by a ``final boolean
withNamespace`` pinned in the constructor of two immutable visitors,
``ROOT`` and ``NESTED``: only the outermost element declares the
namespace, and an element written from ``writeClass`` is nested by
construction. Everything else is then ``static``, and a method reference
to a static method captures nothing.

The design is otherwise the read side's, mirrored:

* One shape, not two. ``ContentWriter<T>`` writes a value where the
  writer already is; ``writeElement`` frames it in a start and an end tag.
  There is deliberately no second delegate for a whole element, as there
  is on the reading side -- an element differs from a content only in what
  it writes, never in its shape. The reading needs the distinction because
  a content reader has to be told whether its element was self-closing,
  and a writer has nothing to be told. Use sites take
  ``ContentWriter<? super T>``, which is how Java spells the
  contravariance C# writes as ``in T``, so the single writer of an
  ``IClass`` serves wherever the writer of a more specific interface is
  expected.

* ``visit{Cls}`` collapses to one call, and the visitor becomes
  a dispatch shim::

      @OverRide
      public void visitQualifier(
        IQualifier that,
        XMLStreamWriter writer) {
        writeElement(
          "qualifier",
          that,
          writer,
          withNamespace,
          _VisitorWithWriter::writeQualifierAsSequence);
      }

* The families are dual to the readers and named by the same moniker
  grammar: ``write{Cls}AsSequence``, ``writeClass``/``writeUnion``,
  ``writeListOf_{M}``, ``writeTupleOf{N}_{M...}``, ``writeAtV_{M}`` and
  ``writeAtV{i}_{M}``. The list and the tuple writers inline their loop,
  which retires ``serializeItems``, ``asNamedElementSerializer`` and
  ``serializeTuple{N}`` -- Java needs no partial application here, since
  the name to bind is fixed per generated function.

* One property emitter replaces the six. Every property is written by
  ``writeProperty``, or by ``writeOptionalProperty``, which takes the
  ``Optional`` apart once instead of asking it and then unwrapping it --
  the old form called the getter twice, and every call allocates an
  ``Optional``.

* ``_type_moniker``, ``_leaf_moniker``, ``_is_instance_type``,
  ``_is_dispatched`` and ``_collect_needed`` are reused unchanged: the
  reading and the writing are composed out of the same shapes, so one
  gating pass answers for both. Only the two dispatching writers needed
  a walk of their own, because the *position* matters on the write side
  while it does not on the read side: an item of any class is written as
  its own element, whereas a property of a concrete class without
  descendants writes its properties straight into the property element.

Two things are genuinely asymmetric, and both save work:

* There is no ``writeTextAs_{primitive}`` family. Once the value is at
  hand nothing type-specific is left to do with it, so the four
  ``toString`` primitives share ``writeStringifiedContent`` and the byte
  array ``writeByteArrayContent``. An enumeration does have something of
  its own to write and keeps one function each -- named
  ``writeTextAs_{Enum}`` and not ``write{Enum}Content``, so that it stays
  in the moniker-keyed name space: were it keyed on the symbol, an
  enumeration which somebody names ``Stringified`` or ``ByteArray`` would
  silently take a shared writer's name.

* There is no element writer per class. Which element an instance writes
  follows from its run-time type, and one virtual call answers that for
  every class at once. The reading needs a dispatcher per interface
  because it has to decide what to construct before it has read anything.

Two defects of the write side are fixed along the way, both of which the
decomposition makes one-line changes rather than sweeping ones.

``Serialize.to`` never flushed, so the serialization was not finished when
it returned, and a failure of the underlying stream surfaced at the
caller's own flush as a raw ``XMLStreamException`` -- outside the one
method whose ``throws`` clause promises a ``SerializeException``. Probed
with a ``Writer`` which accepts every character and fails on ``flush``,
``to`` used to return normally and now throws. The flush goes in
``Serialize.to`` and nowhere else, as in Go's #689, so
``XMLStreamWriter`` keeps buffering in between. Mind that this does push
the underlying stream: a caller serializing many small documents into one
buffered writer now flushes once per document rather than once per
buffer. That is the right trade, since the alternative leaves the output
silently incomplete whenever the caller forgets.

Every ``SerializeException`` carried an empty path, so every failure
read ``... at: the beginning`` no matter where it happened. The path is
now built as the stack unwinds, each container prepending the one segment
it knows, exactly as on the read side where ``read<Cls>FromSequence``
prepends the property after its ``switch`` and ``readList`` prepends the
index. Probed with an ``XMLStreamWriter`` which refuses to write one
chosen value:

    a property of the root instance   idShort
    nested through two lists          submodels/*[1]/submodelElements/*[1]/value
    through a class-typed property    semanticId/keys/*[0]/value

``SerializeException`` renders its message in its constructor, so its path
has to be complete by the time it is built and it can not be the thing
thrown from inside; a private ``_SerializeFailure`` carries
a ``Reporting.Error`` instead, and ``Serialize.to`` renders that with
``Reporting.generateRelativeXPath`` and converts. The public exception is
untouched. ``writeElement`` contributes no segment and ``writeProperty``
is ``writeElement`` plus the one thing a property knows which nothing
below it does -- its own name -- so there is no ``try`` around any of the
276 property writes. A list and a tuple put their ``try`` outside the
loop and advance the index only once an item has been written, so the
item named is the one which failed and nothing is paid per item.

Two segments are deliberately absent, and the reasoning is written into
the generated Javadoc, since both look like oversights. The outermost
element is not named, because ``Serialize.to`` takes any ``IClass`` and
the name would say nothing the caller does not already know; the read
side does prepend it, having one facade per class. Neither is the
discriminator element of a polymorphic property, which the read side does
prepend -- for a failure in an embedded data specification it reports
``submodel/embeddedDataSpecifications/*[0]/dataSpecificationContent/``
``dataSpecificationIec61360/preferredName/*[0]`` where the write side
reports ``embeddedDataSpecifications/*[0]/dataSpecificationContent/``
``preferredName/*[0]/text``. The reading points into a document it is
parsing, where that element is a real extra level; the writing points
into the instance the caller handed over, where it is not, and
``value/idShort`` there is exactly ``getValue().getIdShort()``. Go's #689
made the same choice when it rendered its serialization paths through
``aasreporting.ToGolangPath`` rather than as XPaths.

For the AAS meta-model the write region of ``Xmlization.java`` goes 3,281
-> 2,873 lines (-12.4%), error paths included. The write region no longer
contains a single lambda.

Unlike the read side, that shows up in the measurement. Serializing an
``Environment`` of 200 submodels (1.14 MB of XML) into a writer which
counts the characters and discards them, 100 runs after 100 warm-up runs,
eight rounds of each variant run alternately: the allocation goes 0.76 ->
0.36 MB per serialization, and the time 7.69 -> 7.49 ms, which is inside
the round-to-round noise. ``XMLStreamWriter`` writes straight through, so
the whole write path allocates ~0.8 MB against the ~49 MB the *reading*
of the same document allocates -- there is no StAX floor here for the
closures to hide behind, which is precisely why #690 could not measure
its own.

The snippet contract for an implementation-specific class changes:
``Xmlization/VisitorWithWriter/visit_{cls}.java`` is no longer read, as
the element around the sequence is now written identically for every
class, and the remaining ``{cls}_to_sequence.java`` has to define
a ``static`` ``write{Cls}AsSequence``.
@coveralls

Copy link
Copy Markdown

Coverage Report for CI Build 34977000804

Coverage increased (+0.02%) to 85.322%

Details

  • Coverage increased (+0.02%) from the base build.
  • Patch coverage: 6 uncovered changes across 1 file (133 of 139 lines covered, 95.68%).
  • No coverage regressions found.

Uncovered Changes

File Changed Covered %
aas_core_codegen/java/lib/_generate_xmlization.py 139 133 95.68%

Coverage Regressions

No coverage regressions found.


Coverage Stats

Coverage Status
Relevant Lines: 39795
Covered Lines: 33954
Line Coverage: 85.32%
Coverage Strength: 2.56 hits per line

💛 - Coveralls

@mristin
mristin merged commit 8a70656 into main Sep 15, 2026
5 checks passed
@mristin
mristin deleted the mristin/Compose-Java-XML-serialization branch September 15, 2026 14:33
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants