Skip to content

Compose the C# XML serialization out of writers - #687

Merged
mristin merged 1 commit into
mainfrom
mristin/Decompose-csharp-xml-serialization
Sep 13, 2026
Merged

mristin merged 1 commit into
mainfrom
mristin/Decompose-csharp-xml-serialization

Conversation

@mristin

@mristin mristin commented Sep 13, 2026

Copy link
Copy Markdown
Contributor

This is the write-side dual of #686, and finishes the decomposition of csharp/lib/_generate_xmlization.py: every property is now written by one call whose content writer has been composed once into a field, instead of by a lambda built at the call site.

Measured on the AAS meta-model fixture, the writer region of xmlization.cs goes 3,908 -> 2,464 lines (-37%), and the file as a whole 12,471 -> 11,023. Per property 8.1 -> ~5 lines:

if (that.SemanticId != null)
{
    WriteElement(
        "semanticId", that.SemanticId, writer, WriteIReference);
}

The line count is the lesser win. Every content serializer used to be a capturing lambda composed at the call site -- 212 this. references in that region -- which allocates a closure on every single write; the read side measured that at 152 bytes/op and rejected it. All 276 call sites now read a static readonly field instead, so the write path allocates nothing regardless of the consumer's C# version.

One delegate, ContentWriter<in T>, replaces
ElementContentSerializer<T>. in T is the dual of the read side's out T: it is what lets the single static WriteIClass(Aas.IClass, ...) serve as the item writer of a list or a tuple of any interface. Deliberately one delegate, not a content/element pair: on the write side both would have identical signatures, since the content/element split is a read-side concern (isEmpty).

WriteElement and WrapInElement are one body with two entry points. The applied form is what the 276 property sites call; WrapInElement is the partial application C# does not give for free, and is used only for the <v> element of a list item and the positional v1, v2, ... of a tuple item. It carries a doc comment saying so, because its payoff is invisible from its call count: it is what lets WriteList and WriteTupleN take a single item-writer type even though a class item writes its own run-time-named element while a primitive item needs a fixed positional one. The tuples fixture's Tuple[int, Some_item, Abstract_item, Some_item, Positive_int, Result] mixes the two item by item; with two item-writer types it would have needed a combinator per combination.

Names are kept distinct from the read side's on purpose: WrapInElement / WriteEnum / WriteList / WriteTupleN against AtElement / AsEnum / AsList / AsTupleN. There is no dual of AsText (WriteValue cannot fail) and none of AsElement.

The visitor stays, reduced to a dispatch shim: a single _instance plus static WriteIClass, which one virtual call resolves for every abstract class, every concrete class with descendants and -- through WriteIUnion -- every named union. That is why the write side needs neither the read side's 13 per-interface dispatchers nor a combinator to invoke one.

Along the way, two things shared with the read side:

  • The gating pass is now recursive and serves both emitters. _needed_content_readers walked only one level, so a model whose only str sat inside List[List[str]] would have emitted ReadString = AsText<string>(ReadContentAsString, "") while generating neither AsText nor ReadContentAsString -- output that does not compile. Latent only because the per-property gate blocks such models today.
  • _content_types_in_initialization_order replaces the two copies of the registration recursion, so the readers and the writers are de-duplicated and ordered by one walk.

Nesting composes for free in this design -- WriteListOfListOfString = WriteList<List<string>>(WrapInElement(WriteListOfString, "v")) falls out with no new combinator -- but stays gated, as on the read side, since copying, enhancing, jsonization and the test generators bail on it in C# and in all five other backends. The write side now reports that as a proper Error instead of raising NotImplementedError.

The snippet contract for an implementation-specific class changes: Xmlization/VisitorWithWriter/{cls}_to_sequence.cs must now be static, and visit_{cls}.cs must call it without this. -- the same change the read side already imposed on
..._from_sequence.cs.

@coveralls

coveralls commented Sep 13, 2026 •

Copy link
Copy Markdown

Coverage Report for CI Build 34749764234

Coverage increased (+0.004%) to 85.253%

Details

  • Coverage increased (+0.004%) from the base build.
  • Patch coverage: 9 uncovered changes across 1 file (174 of 183 lines covered, 95.08%).
  • 2 coverage regressions across 1 file.

Uncovered Changes

File Changed Covered %
aas_core_codegen/csharp/lib/_generate_xmlization.py 183 174 95.08%

Coverage Regressions

2 previously-covered lines in 1 file lost coverage.

File Lines Losing Coverage Coverage
aas_core_codegen/csharp/lib/_generate_xmlization.py 2 93.41%

Coverage Stats

Coverage Status
Relevant Lines: 39771
Covered Lines: 33906
Line Coverage: 85.25%
Coverage Strength: 2.56 hits per line

💛 - Coveralls

This is the write-side dual of #686, and finishes the decomposition
of ``csharp/lib/_generate_xmlization.py``: every property is now written
by one call whose content writer has been composed once into a field,
instead of by a lambda built at the call site.

Measured on the AAS meta-model fixture, the writer region of
``xmlization.cs`` goes 3,908 -> 2,464 lines (-37%), and the file as
a whole 12,471 -> 11,023. Per property 8.1 -> ~5 lines:

    if (that.SemanticId != null)
    {
        WriteElement(
            "semanticId", that.SemanticId, writer, WriteIReference);
    }

The line count is the lesser win. Every content serializer used to be
a capturing lambda composed at the call site -- 212 ``this.`` references
in that region -- which allocates a closure on every single write; the
read side measured that at 152 bytes/op and rejected it. All 276 call
sites now read a ``static readonly`` field instead, so the write path
allocates nothing regardless of the consumer's C# version.

One delegate, ``ContentWriter<in T>``, replaces
``ElementContentSerializer<T>``. ``in T`` is the dual of the read side's
``out T``: it is what lets the single static ``WriteIClass(Aas.IClass,
...)`` serve as the item writer of a list or a tuple of any interface.
Deliberately *one* delegate, not a content/element pair: on the write
side both would have identical signatures, since the content/element
split is a read-side concern (``isEmpty``).

``WriteElement`` and ``WrapInElement`` are one body with two entry
points. The applied form is what the 276 property sites call;
``WrapInElement`` is the partial application C# does not give for free,
and is used *only* for the ``<v>`` element of a list item and the
positional ``v1``, ``v2``, ... of a tuple item. It carries a doc comment
saying so, because its payoff is invisible from its call count: it is
what lets ``WriteList`` and ``WriteTupleN`` take a *single* item-writer
type even though a class item writes its own run-time-named element
while a primitive item needs a fixed positional one. The ``tuples``
fixture's ``Tuple[int, Some_item, Abstract_item, Some_item,
Positive_int, Result]`` mixes the two item by item; with two item-writer
types it would have needed a combinator per combination.

Names are kept distinct from the read side's on purpose:
``WrapInElement`` / ``WriteEnum`` / ``WriteList`` / ``WriteTupleN``
against ``AtElement`` / ``AsEnum`` / ``AsList`` / ``AsTupleN``. There is
no dual of ``AsText`` (``WriteValue`` cannot fail) and none of
``AsElement``.

The visitor stays, reduced to a dispatch shim: a single ``_instance``
plus static ``WriteIClass``, which one virtual call resolves for every
abstract class, every concrete class with descendants and -- through
``WriteIUnion`` -- every named union. That is why the write side needs
neither the read side's 13 per-interface dispatchers nor a combinator to
invoke one.

Along the way, two things shared with the read side:

* The gating pass is now recursive and serves both emitters.
  ``_needed_content_readers`` walked only one level, so a model whose
  only ``str`` sat inside ``List[List[str]]`` would have emitted
  ``ReadString = AsText<string>(ReadContentAsString, "")`` while
  generating neither ``AsText`` nor ``ReadContentAsString`` -- output
  that does not compile. Latent only because the per-property gate
  blocks such models today.
* ``_content_types_in_initialization_order`` replaces the two copies of
  the registration recursion, so the readers and the writers are
  de-duplicated and ordered by one walk.

Nesting composes for free in this design -- ``WriteListOfListOfString =
WriteList<List<string>>(WrapInElement(WriteListOfString, "v"))`` falls
out with no new combinator -- but stays gated, as on the read side,
since ``copying``, ``enhancing``, ``jsonization`` and the test
generators bail on it in C# and in all five other backends. The write
side now reports that as a proper ``Error`` instead of raising
``NotImplementedError``.

The snippet contract for an implementation-specific class changes:
``Xmlization/VisitorWithWriter/{cls}_to_sequence.cs`` must now be
``static``, and ``visit_{cls}.cs`` must call it without ``this.`` --
the same change the read side already imposed on
``..._from_sequence.cs``.
@mristin
mristin force-pushed the mristin/Decompose-csharp-xml-serialization branch from ab3cad5 to 436da8c Compare September 13, 2026 09:34
@mristin
mristin merged commit 2542c01 into main Sep 13, 2026
5 checks passed
@mristin
mristin deleted the mristin/Decompose-csharp-xml-serialization branch September 13, 2026 09:44
mristin added a commit that referenced this pull request Sep 13, 2026
Write-side dual of #688, as #687 was of #686 in C#.

Measured on the AAS meta-model fixture, the serialization region of
``xmlization.go`` goes 11,080 -> 4,727 lines (-57%), and the file as
a whole 16,692 -> 10,330. ``writeExtensionAsSequence`` and its six
properties go 180 -> 69 lines.

Everything is written through one shape,
``func(encoder, value T) error``. A stringification, a class's own
``write<X>AsSequence`` and the generated item, list and tuple writers
all have it, so ``writeElement`` -- which frames an XML element around
a value written that way -- is the only writer a property needs,
whatever kind of value it holds. What is left per property is that one
call and a decision on presence:

    err = finishProperty(
        "SemanticID()",
        writeOptionalInstance(
            encoder,
            "semanticId",
            that.SemanticID(),
            writeReferenceAsSequence,
        ),
    )
    if err != nil {
        return
    }

``finishProperty`` takes the *result* of the write rather than the write
itself -- the same trick as the read side's ``readOptional``, since Go
evaluates the argument first -- so one function concludes every
property, prepending the getter to the path of the error.

Mind that the getter is spelled out as it is in Go, ``SemanticID()`` and
not ``semanticId``, since a serialization error renders its path through
``aasreporting.ToGolangPath``, while the de-serialization prepends
the XML name, its errors being XPaths. Go has no ``nameof``.

The ``writeOptional*`` family has one member per way Go spells an
optional, each named after that spelling, as the check for the presence
is the only thing which differs between them: a pointer to the value
(``writeOptionalPointer``: a scalar, an enumeration, a tuple), a value
which is nil on its own (``writeOptionalInstance``: an instance, or
a named union, which is a pointer to a struct) and a slice
(``writeOptionalSlice``: a list, the bytes). They can not be collapsed
into one, since Go decides nil-ness by the representation: a type
parameter can not be compared against ``nil``, ``any(that) == nil`` is
false for a nil *pointer* boxed in an ``any`` -- hence the comparison
against the zero value in ``writeOptionalInstance``, which in turn
panics on a slice, which is not comparable. Each of them writes nothing
at all when the value is absent, which is what keeps the absence out of
the generated code.

No closure is allocated on the write path any more: every writer passed
on is a top-level function or an instantiated generic, hence a static
funcval, exactly as on the read side. A list and a tuple get one
generated content writer as well, per item type resp. per tuple
type -- ``writeListOfIReference`` and ``writeTupleOfStringLong`` and
their kin -- since ``writeList`` and ``writeTupleN`` take a writer per
item, and Go gives no partial application to bind those in without
allocating.

Finally, the flush. ``write<X>AsSequence`` used to flush the encoder
after every property, and ``write<X>`` again at its end element, which
defeated the buffering of ``xml.Encoder`` on every write; only the flush
after the last token is needed, as ``EncodeToken`` deliberately leaves
it to the caller. ``Marshal`` is therefore split into ``writeClass``,
which dispatches on the model type, and ``Marshal`` itself, which calls
it and flushes exactly once; ``writeInstance`` and ``writeUnion`` go to
``writeClass``, so a nested instance does not flush either. Serializing
an instance holding 1,000 instances (42 KB of XML) from
the ``list_of_classes`` fixture, the writes which actually reach
the underlying writer go 2,003 -> 11, one per 4 KB of buffer instead of
one per property and per instance; to an ``os.File`` that is 5.29 ->
0.76 ms per serialization, the median of five.
mristin added a commit that referenced this pull request Sep 15, 2026
Write-side dual of #690, as #687 was of #686 in C# and #689 of #688 in Go.

``visit{Cls}`` was a verbatim copy of ``serializeElement``: 38 methods of
18 lines each spelling out the start tag, the namespace, the content and
the end tag, around a call which already did exactly that. Six
near-identical emitters -- one per property kind -- produced the very same
``serializeElement`` call, differing only in the content serializer. And
everything was an *instance* method, because of one mutable ``topLevel``
flag.

That flag is what cost the most. It made ``this::writeStringifiedContent``,
``this::visit`` and ``(value, w) -> this.xToSequence(value, w)`` capturing,
and ``serializeItems(...)`` and ``asNamedElementSerializer(...)`` returned
capturing lambdas on top of them, so writing one instance allocated
a small pile of closures. It is replaced by a ``final boolean
withNamespace`` pinned in the constructor of two immutable visitors,
``ROOT`` and ``NESTED``: only the outermost element declares the
namespace, and an element written from ``writeClass`` is nested by
construction. Everything else is then ``static``, and a method reference
to a static method captures nothing.

The design is otherwise the read side's, mirrored:

* One shape, not two. ``ContentWriter<T>`` writes a value where the
  writer already is; ``writeElement`` frames it in a start and an end tag.
  There is deliberately no second delegate for a whole element, as there
  is on the reading side -- an element differs from a content only in what
  it writes, never in its shape. The reading needs the distinction because
  a content reader has to be told whether its element was self-closing,
  and a writer has nothing to be told. Use sites take
  ``ContentWriter<? super T>``, which is how Java spells the
  contravariance C# writes as ``in T``, so the single writer of an
  ``IClass`` serves wherever the writer of a more specific interface is
  expected.

* ``visit{Cls}`` collapses to one call, and the visitor becomes
  a dispatch shim::

      @OverRide
      public void visitQualifier(
        IQualifier that,
        XMLStreamWriter writer) {
        writeElement(
          "qualifier",
          that,
          writer,
          withNamespace,
          _VisitorWithWriter::writeQualifierAsSequence);
      }

* The families are dual to the readers and named by the same moniker
  grammar: ``write{Cls}AsSequence``, ``writeClass``/``writeUnion``,
  ``writeListOf_{M}``, ``writeTupleOf{N}_{M...}``, ``writeAtV_{M}`` and
  ``writeAtV{i}_{M}``. The list and the tuple writers inline their loop,
  which retires ``serializeItems``, ``asNamedElementSerializer`` and
  ``serializeTuple{N}`` -- Java needs no partial application here, since
  the name to bind is fixed per generated function.

* One property emitter replaces the six. Every property is written by
  ``writeProperty``, or by ``writeOptionalProperty``, which takes the
  ``Optional`` apart once instead of asking it and then unwrapping it --
  the old form called the getter twice, and every call allocates an
  ``Optional``.

* ``_type_moniker``, ``_leaf_moniker``, ``_is_instance_type``,
  ``_is_dispatched`` and ``_collect_needed`` are reused unchanged: the
  reading and the writing are composed out of the same shapes, so one
  gating pass answers for both. Only the two dispatching writers needed
  a walk of their own, because the *position* matters on the write side
  while it does not on the read side: an item of any class is written as
  its own element, whereas a property of a concrete class without
  descendants writes its properties straight into the property element.

Two things are genuinely asymmetric, and both save work:

* There is no ``writeTextAs_{primitive}`` family. Once the value is at
  hand nothing type-specific is left to do with it, so the four
  ``toString`` primitives share ``writeStringifiedContent`` and the byte
  array ``writeByteArrayContent``. An enumeration does have something of
  its own to write and keeps one function each -- named
  ``writeTextAs_{Enum}`` and not ``write{Enum}Content``, so that it stays
  in the moniker-keyed name space: were it keyed on the symbol, an
  enumeration which somebody names ``Stringified`` or ``ByteArray`` would
  silently take a shared writer's name.

* There is no element writer per class. Which element an instance writes
  follows from its run-time type, and one virtual call answers that for
  every class at once. The reading needs a dispatcher per interface
  because it has to decide what to construct before it has read anything.

Two defects of the write side are fixed along the way, both of which the
decomposition makes one-line changes rather than sweeping ones.

``Serialize.to`` never flushed, so the serialization was not finished when
it returned, and a failure of the underlying stream surfaced at the
caller's own flush as a raw ``XMLStreamException`` -- outside the one
method whose ``throws`` clause promises a ``SerializeException``. Probed
with a ``Writer`` which accepts every character and fails on ``flush``,
``to`` used to return normally and now throws. The flush goes in
``Serialize.to`` and nowhere else, as in Go's #689, so
``XMLStreamWriter`` keeps buffering in between. Mind that this does push
the underlying stream: a caller serializing many small documents into one
buffered writer now flushes once per document rather than once per
buffer. That is the right trade, since the alternative leaves the output
silently incomplete whenever the caller forgets.

Every ``SerializeException`` carried an empty path, so every failure
read ``... at: the beginning`` no matter where it happened. The path is
now built as the stack unwinds, each container prepending the one segment
it knows, exactly as on the read side where ``read<Cls>FromSequence``
prepends the property after its ``switch`` and ``readList`` prepends the
index. Probed with an ``XMLStreamWriter`` which refuses to write one
chosen value:

    a property of the root instance   idShort
    nested through two lists          submodels/*[1]/submodelElements/*[1]/value
    through a class-typed property    semanticId/keys/*[0]/value

``SerializeException`` renders its message in its constructor, so its path
has to be complete by the time it is built and it can not be the thing
thrown from inside; a private ``_SerializeFailure`` carries
a ``Reporting.Error`` instead, and ``Serialize.to`` renders that with
``Reporting.generateRelativeXPath`` and converts. The public exception is
untouched. ``writeElement`` contributes no segment and ``writeProperty``
is ``writeElement`` plus the one thing a property knows which nothing
below it does -- its own name -- so there is no ``try`` around any of the
276 property writes. A list and a tuple put their ``try`` outside the
loop and advance the index only once an item has been written, so the
item named is the one which failed and nothing is paid per item.

Two segments are deliberately absent, and the reasoning is written into
the generated Javadoc, since both look like oversights. The outermost
element is not named, because ``Serialize.to`` takes any ``IClass`` and
the name would say nothing the caller does not already know; the read
side does prepend it, having one facade per class. Neither is the
discriminator element of a polymorphic property, which the read side does
prepend -- for a failure in an embedded data specification it reports
``submodel/embeddedDataSpecifications/*[0]/dataSpecificationContent/``
``dataSpecificationIec61360/preferredName/*[0]`` where the write side
reports ``embeddedDataSpecifications/*[0]/dataSpecificationContent/``
``preferredName/*[0]/text``. The reading points into a document it is
parsing, where that element is a real extra level; the writing points
into the instance the caller handed over, where it is not, and
``value/idShort`` there is exactly ``getValue().getIdShort()``. Go's #689
made the same choice when it rendered its serialization paths through
``aasreporting.ToGolangPath`` rather than as XPaths.

For the AAS meta-model the write region of ``Xmlization.java`` goes 3,281
-> 2,873 lines (-12.4%), error paths included. The write region no longer
contains a single lambda.

Unlike the read side, that shows up in the measurement. Serializing an
``Environment`` of 200 submodels (1.14 MB of XML) into a writer which
counts the characters and discards them, 100 runs after 100 warm-up runs,
eight rounds of each variant run alternately: the allocation goes 0.76 ->
0.36 MB per serialization, and the time 7.69 -> 7.49 ms, which is inside
the round-to-round noise. ``XMLStreamWriter`` writes straight through, so
the whole write path allocates ~0.8 MB against the ~49 MB the *reading*
of the same document allocates -- there is no StAX floor here for the
closures to hide behind, which is precisely why #690 could not measure
its own.

The snippet contract for an implementation-specific class changes:
``Xmlization/VisitorWithWriter/visit_{cls}.java`` is no longer read, as
the element around the sequence is now written identically for every
class, and the remaining ``{cls}_to_sequence.java`` has to define
a ``static`` ``write{Cls}AsSequence``.
mristin added a commit that referenced this pull request Sep 15, 2026
Write-side dual of #690, as #687 was of #686 in C# and #689 of #688 in Go.

``visit{Cls}`` was a verbatim copy of ``serializeElement``: 38 methods of
18 lines each spelling out the start tag, the namespace, the content and
the end tag, around a call which already did exactly that. Six
near-identical emitters -- one per property kind -- produced the very same
``serializeElement`` call, differing only in the content serializer. And
everything was an *instance* method, because of one mutable ``topLevel``
flag.

That flag is what cost the most. It made ``this::writeStringifiedContent``,
``this::visit`` and ``(value, w) -> this.xToSequence(value, w)`` capturing,
and ``serializeItems(...)`` and ``asNamedElementSerializer(...)`` returned
capturing lambdas on top of them, so writing one instance allocated
a small pile of closures. It is replaced by a ``final boolean
withNamespace`` pinned in the constructor of two immutable visitors,
``ROOT`` and ``NESTED``: only the outermost element declares the
namespace, and an element written from ``writeClass`` is nested by
construction. Everything else is then ``static``, and a method reference
to a static method captures nothing.

The design is otherwise the read side's, mirrored:

* One shape, not two. ``ContentWriter<T>`` writes a value where the
  writer already is; ``writeElement`` frames it in a start and an end tag.
  There is deliberately no second delegate for a whole element, as there
  is on the reading side -- an element differs from a content only in what
  it writes, never in its shape. The reading needs the distinction because
  a content reader has to be told whether its element was self-closing,
  and a writer has nothing to be told. Use sites take
  ``ContentWriter<? super T>``, which is how Java spells the
  contravariance C# writes as ``in T``, so the single writer of an
  ``IClass`` serves wherever the writer of a more specific interface is
  expected.

* ``visit{Cls}`` collapses to one call, and the visitor becomes
  a dispatch shim::

      @OverRide
      public void visitQualifier(
        IQualifier that,
        XMLStreamWriter writer) {
        writeElement(
          "qualifier",
          that,
          writer,
          withNamespace,
          _VisitorWithWriter::writeQualifierAsSequence);
      }

* The families are dual to the readers and named by the same moniker
  grammar: ``write{Cls}AsSequence``, ``writeClass``/``writeUnion``,
  ``writeListOf_{M}``, ``writeTupleOf{N}_{M...}``, ``writeAtV_{M}`` and
  ``writeAtV{i}_{M}``. The list and the tuple writers inline their loop,
  which retires ``serializeItems``, ``asNamedElementSerializer`` and
  ``serializeTuple{N}`` -- Java needs no partial application here, since
  the name to bind is fixed per generated function.

* One property emitter replaces the six. Every property is written by
  ``writeProperty``, or by ``writeOptionalProperty``, which takes the
  ``Optional`` apart once instead of asking it and then unwrapping it --
  the old form called the getter twice, and every call allocates an
  ``Optional``.

* ``_type_moniker``, ``_leaf_moniker``, ``_is_instance_type``,
  ``_is_dispatched`` and ``_collect_needed`` are reused unchanged: the
  reading and the writing are composed out of the same shapes, so one
  gating pass answers for both. Only the two dispatching writers needed
  a walk of their own, because the *position* matters on the write side
  while it does not on the read side: an item of any class is written as
  its own element, whereas a property of a concrete class without
  descendants writes its properties straight into the property element.

Two things are genuinely asymmetric, and both save work:

* There is no ``writeTextAs_{primitive}`` family. Once the value is at
  hand nothing type-specific is left to do with it, so the four
  ``toString`` primitives share ``writeStringifiedContent`` and the byte
  array ``writeByteArrayContent``. An enumeration does have something of
  its own to write and keeps one function each -- named
  ``writeTextAs_{Enum}`` and not ``write{Enum}Content``, so that it stays
  in the moniker-keyed name space: were it keyed on the symbol, an
  enumeration which somebody names ``Stringified`` or ``ByteArray`` would
  silently take a shared writer's name.

* There is no element writer per class. Which element an instance writes
  follows from its run-time type, and one virtual call answers that for
  every class at once. The reading needs a dispatcher per interface
  because it has to decide what to construct before it has read anything.

Two defects of the write side are fixed along the way, both of which the
decomposition makes one-line changes rather than sweeping ones.

``Serialize.to`` never flushed, so the serialization was not finished when
it returned, and a failure of the underlying stream surfaced at the
caller's own flush as a raw ``XMLStreamException`` -- outside the one
method whose ``throws`` clause promises a ``SerializeException``. Probed
with a ``Writer`` which accepts every character and fails on ``flush``,
``to`` used to return normally and now throws. The flush goes in
``Serialize.to`` and nowhere else, as in Go's #689, so
``XMLStreamWriter`` keeps buffering in between. Mind that this does push
the underlying stream: a caller serializing many small documents into one
buffered writer now flushes once per document rather than once per
buffer. That is the right trade, since the alternative leaves the output
silently incomplete whenever the caller forgets.

Every ``SerializeException`` carried an empty path, so every failure
read ``... at: the beginning`` no matter where it happened. The path is
now built as the stack unwinds, each container prepending the one segment
it knows, exactly as on the read side where ``read<Cls>FromSequence``
prepends the property after its ``switch`` and ``readList`` prepends the
index. Probed with an ``XMLStreamWriter`` which refuses to write one
chosen value:

    a property of the root instance   idShort
    nested through two lists          submodels/*[1]/submodelElements/*[1]/value
    through a class-typed property    semanticId/keys/*[0]/value

``SerializeException`` renders its message in its constructor, so its path
has to be complete by the time it is built and it can not be the thing
thrown from inside; a private ``_SerializeFailure`` carries
a ``Reporting.Error`` instead, and ``Serialize.to`` renders that with
``Reporting.generateRelativeXPath`` and converts. The public exception is
untouched. ``writeElement`` contributes no segment and ``writeProperty``
is ``writeElement`` plus the one thing a property knows which nothing
below it does -- its own name -- so there is no ``try`` around any of the
276 property writes. A list and a tuple put their ``try`` outside the
loop and advance the index only once an item has been written, so the
item named is the one which failed and nothing is paid per item.

Two segments are deliberately absent, and the reasoning is written into
the generated Javadoc, since both look like oversights. The outermost
element is not named, because ``Serialize.to`` takes any ``IClass`` and
the name would say nothing the caller does not already know; the read
side does prepend it, having one facade per class. Neither is the
discriminator element of a polymorphic property, which the read side does
prepend -- for a failure in an embedded data specification it reports
``submodel/embeddedDataSpecifications/*[0]/dataSpecificationContent/``
``dataSpecificationIec61360/preferredName/*[0]`` where the write side
reports ``embeddedDataSpecifications/*[0]/dataSpecificationContent/``
``preferredName/*[0]/text``. The reading points into a document it is
parsing, where that element is a real extra level; the writing points
into the instance the caller handed over, where it is not, and
``value/idShort`` there is exactly ``getValue().getIdShort()``. Go's #689
made the same choice when it rendered its serialization paths through
``aasreporting.ToGolangPath`` rather than as XPaths.

For the AAS meta-model the write region of ``Xmlization.java`` goes 3,281
-> 2,873 lines (-12.4%), error paths included. The write region no longer
contains a single lambda.

Unlike the read side, that shows up in the measurement. Serializing an
``Environment`` of 200 submodels (1.14 MB of XML) into a writer which
counts the characters and discards them, 100 runs after 100 warm-up runs,
eight rounds of each variant run alternately: the allocation goes 0.76 ->
0.36 MB per serialization, and the time 7.69 -> 7.49 ms, which is inside
the round-to-round noise. ``XMLStreamWriter`` writes straight through, so
the whole write path allocates ~0.8 MB against the ~49 MB the *reading*
of the same document allocates -- there is no StAX floor here for the
closures to hide behind, which is precisely why #690 could not measure
its own.

The snippet contract for an implementation-specific class changes:
``Xmlization/VisitorWithWriter/visit_{cls}.java`` is no longer read, as
the element around the sequence is now written identically for every
class, and the remaining ``{cls}_to_sequence.java`` has to define
a ``static`` ``write{Cls}AsSequence``.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants