Skip to content

Compose the C# XML de-serialization out of readers - #686

Merged
mristin merged 1 commit into
mainfrom
mristin/Decompose-csharp-xmlization-property-reading
Sep 13, 2026
Merged

mristin merged 1 commit into
mainfrom
mristin/Decompose-csharp-xmlization-property-reading

Conversation

@mristin

@mristin mristin commented Sep 13, 2026

Copy link
Copy Markdown
Contributor

The generated xmlization.cs repeated itself at three levels. Every property of every concrete class inlined its whole reading procedure into its own case block -- the self-closing check, the end-of-file check, the try/catch around the conversion, an error message naming that property and its class, the PrependSegment marking the path, and, for a list or a tuple, the loop or the positional <v> elements on top. Every concrete class then got its own ...FromElement function, and all of those were the same 43 lines with four tokens substituted in. And every ...FromSequence re-emitted the ~70 lines of framing around its property loop, which said nothing about the class it belonged to.

Everything is now read through one composable shape, and the pieces that do not vary are generated once:

case "category":
    theCategory = ReadString(
        reader, isEmptyProperty, out error);
    break;

The rationale and the design:

  • One shape for every reader. Reading a property's content, a list item's and a tuple item's are the same operation at different nesting depths, so they are one delegate:

    private delegate T ContentReader<T>(
        Xml.XmlReader reader, bool isEmpty, out Reporting.Error? error);
    

    Because the shape is uniform a reader can be an argument to another reader, so a list of tuples -- or anything deeper the meta-model may grow -- falls out of the existing pieces instead of needing a generated helper per combination.

  • A plain T return, not T? and not out T. Failure travels in the error alone and the value is then default!, which the caller never reads. T? would have split the machinery in two, since a nullable return is Nullable<T> for a value type but a nullable reference otherwise and no unconstrained parameter covers both. out T would have compiled, but out is invariant, so a reader of a concrete class could no longer serve as the item reader of an interface-typed list; a return type is covariant in a method-group conversion.

  • Combinators return delegates, because C# has no partial application: AsText, AsEnum, AsList, AsElement and AsTuple1..N each take what varies and return a ContentReader.

  • AtElement is the hinge, and the element name is data. It binds a name to a ContentReader and yields an ElementReader, which is what a list and a tuple take for their items -- and also what a class is. A name cannot be a type argument (v in a list, v1, v2, ... by position in a tuple, its own XML name for a class), so it is passed. With the messages phrased in terms of that name, the combinator needs nothing further to say what it expected, which is what lets one combinator serve both a class and a <v> element.

  • A class's element reader is therefore a binding, not a function. A class's ...FromSequence already is a ContentReader, so its element reader is AtElement(ExtensionFromSequence, "extension") and the 43-line function per class is gone. Call sites are unchanged: a field is invoked exactly like the method was.

  • The composed readers live in static readonly fields. This is not cosmetic. Composing at the call site allocates a closure on every read -- measured at 152 bytes against zero for the code being replaced, a real regression for a de-serialization library. Composing once costs nothing per read and de-duplicates: 276 properties collapse onto 40 readers.

    Two consequences worth knowing. Field initialisers run in declaration order, so a reader is emitted after everything it composes. And ElementReader had to become covariant and internal: as a method group a class's reader widened to its interface for free, but a field is a value, and a field may not be less accessible than its type.

  • Looking ahead an element's name is one helper, PeekElementName, shared by AtElement and by the thirteen readers that dispatch on a discriminator element -- they only ever needed the name, and consuming nothing is what lets the dispatch hand the whole element on.

  • The property loop keeps only its switch. TryNextProperty reads the next property's start tag or reports that the sequence ended, and ConsumeEndElement -- shared with AtElement, which concludes identically -- consumes the end tag. Only the switch stays inline, because its branches assign the locals the constructor is called with.

  • The property is marked on the error path once per class, after the switch, rather than in every case: for a matched case the element name already is the property's XML name.

  • The type in a message comes from typeof(T).Name, so the shared skeletons do not grow with the number of types in the meta-model.

  • Only what a model reaches is emitted, gated along the call graph rather than by property kind -- a list of enumerations needs the enumeration combinator even when no property is one. Without that, a small meta-model would pay for helpers it never calls.

For the AAS meta-model, xmlization.cs goes from 22,572 to 12,471 lines (-44.8%), the per-property blocks from 19.7 to 3.7 lines on average, and the read side loses 38 generated functions along with four helpers.

The generated ``xmlization.cs`` repeated itself at three levels. Every
property of every concrete class inlined its whole reading procedure into
its own ``case`` block -- the self-closing check, the end-of-file check,
the ``try``/``catch`` around the conversion, an error message naming that
property and its class, the ``PrependSegment`` marking the path, and, for
a list or a tuple, the loop or the positional ``<v>`` elements on top.
Every concrete class then got its own ``...FromElement`` function, and all
of those were the same 43 lines with four tokens substituted in. And every
``...FromSequence`` re-emitted the ~70 lines of framing around its property
loop, which said nothing about the class it belonged to.

Everything is now read through one composable shape, and the pieces that
do not vary are generated once:

    case "category":
        theCategory = ReadString(
            reader, isEmptyProperty, out error);
        break;

The rationale and the design:

* **One shape for every reader.** Reading a property's content, a list
  item's and a tuple item's are the same operation at different nesting
  depths, so they are one delegate:

      private delegate T ContentReader<T>(
          Xml.XmlReader reader, bool isEmpty, out Reporting.Error? error);

  Because the shape is uniform a reader can be an argument to another
  reader, so a list of tuples -- or anything deeper the meta-model may
  grow -- falls out of the existing pieces instead of needing a generated
  helper per combination.

* **A plain ``T`` return, not ``T?`` and not ``out T``.** Failure travels
  in the error alone and the value is then ``default!``, which the caller
  never reads. ``T?`` would have split the machinery in two, since
  a nullable return is ``Nullable<T>`` for a value type but a nullable
  reference otherwise and no unconstrained parameter covers both. ``out T``
  would have compiled, but ``out`` is invariant, so a reader of a concrete
  class could no longer serve as the item reader of an interface-typed
  list; a return type is covariant in a method-group conversion.

* **Combinators return delegates**, because C# has no partial application:
  ``AsText``, ``AsEnum``, ``AsList``, ``AsElement`` and ``AsTuple1..N`` each
  take what varies and return a ``ContentReader``.

* **``AtElement`` is the hinge, and the element name is data.** It binds
  a name to a ``ContentReader`` and yields an ``ElementReader``, which is
  what a list and a tuple take for their items -- and also what a class is.
  A name cannot be a type argument (``v`` in a list, ``v1``, ``v2``, ... by
  position in a tuple, its own XML name for a class), so it is passed. With
  the messages phrased in terms of that name, the combinator needs nothing
  further to say what it expected, which is what lets one combinator serve
  both a class and a ``<v>`` element.

* **A class's element reader is therefore a binding, not a function.**
  A class's ``...FromSequence`` already is a ``ContentReader``, so its
  element reader is ``AtElement(ExtensionFromSequence, "extension")`` and
  the 43-line function per class is gone. Call sites are unchanged: a field
  is invoked exactly like the method was.

* **The composed readers live in ``static readonly`` fields.** This is not
  cosmetic. Composing at the call site allocates a closure on every read --
  measured at 152 bytes against zero for the code being replaced, a real
  regression for a de-serialization library. Composing once costs nothing
  per read and de-duplicates: 276 properties collapse onto 40 readers.

  Two consequences worth knowing. Field initialisers run in declaration
  order, so a reader is emitted after everything it composes. And
  ``ElementReader`` had to become covariant and ``internal``: as a method
  group a class's reader widened to its interface for free, but a field is
  a value, and a field may not be less accessible than its type.

* **Looking ahead an element's name is one helper**, ``PeekElementName``,
  shared by ``AtElement`` and by the thirteen readers that dispatch on
  a discriminator element -- they only ever needed the name, and consuming
  nothing is what lets the dispatch hand the whole element on.

* **The property loop keeps only its ``switch``.** ``TryNextProperty``
  reads the next property's start tag or reports that the sequence ended,
  and ``ConsumeEndElement`` -- shared with ``AtElement``, which concludes
  identically -- consumes the end tag. Only the ``switch`` stays inline,
  because its branches assign the locals the constructor is called with.

* **The property is marked on the error path once per class, after the
  ``switch``**, rather than in every case: for a matched case the element
  name already is the property's XML name.

* **The type in a message comes from ``typeof(T).Name``**, so the shared
  skeletons do not grow with the number of types in the meta-model.

* **Only what a model reaches is emitted**, gated along the call graph
  rather than by property kind -- a list of enumerations needs the
  enumeration combinator even when no property is one. Without that,
  a small meta-model would pay for helpers it never calls.

For the AAS meta-model, ``xmlization.cs`` goes from 22,572 to 12,471 lines
(-44.8%), the per-property blocks from 19.7 to 3.7 lines on average, and
the read side loses 38 generated functions along with four helpers.
@coveralls

Copy link
Copy Markdown

Coverage Report for CI Build 34729308442

Coverage increased (+0.007%) to 85.249%

Details

  • Coverage increased (+0.007%) from the base build.
  • Patch coverage: 6 uncovered changes across 1 file (204 of 210 lines covered, 97.14%).
  • 2 coverage regressions across 1 file.

Uncovered Changes

File Changed Covered %
aas_core_codegen/csharp/lib/_generate_xmlization.py 210 204 97.14%

Coverage Regressions

2 previously-covered lines in 1 file lost coverage.

File Lines Losing Coverage Coverage
aas_core_codegen/csharp/lib/_generate_xmlization.py 2 92.84%

Coverage Stats

Coverage Status
Relevant Lines: 39800
Covered Lines: 33929
Line Coverage: 85.25%
Coverage Strength: 2.56 hits per line

💛 - Coveralls

@mristin
mristin merged commit cc1bacf into main Sep 13, 2026
5 checks passed
@mristin
mristin deleted the mristin/Decompose-csharp-xmlization-property-reading branch September 13, 2026 06:08
mristin added a commit that referenced this pull request Sep 13, 2026
This is the write-side dual of #686, and finishes the decomposition
of ``csharp/lib/_generate_xmlization.py``: every property is now written
by one call whose content writer has been composed once into a field,
instead of by a lambda built at the call site.

Measured on the AAS meta-model fixture, the writer region of
``xmlization.cs`` goes 3,908 -> 2,464 lines (-37%), and the file as
a whole 12,471 -> 11,023. Per property 8.1 -> ~5 lines:

    if (that.SemanticId != null)
    {
        WriteElement(
            "semanticId", that.SemanticId, writer, WriteIReference);
    }

The line count is the lesser win. Every content serializer used to be
a capturing lambda composed at the call site -- 212 ``this.`` references
in that region -- which allocates a closure on every single write; the
read side measured that at 152 bytes/op and rejected it. All 276 call
sites now read a ``static readonly`` field instead, so the write path
allocates nothing regardless of the consumer's C# version.

One delegate, ``ContentWriter<in T>``, replaces
``ElementContentSerializer<T>``. ``in T`` is the dual of the read side's
``out T``: it is what lets the single static ``WriteIClass(Aas.IClass,
...)`` serve as the item writer of a list or a tuple of any interface.
Deliberately *one* delegate, not a content/element pair: on the write
side both would have identical signatures, since the content/element
split is a read-side concern (``isEmpty``).

``WriteElement`` and ``WrapInElement`` are one body with two entry
points. The applied form is what the 276 property sites call;
``WrapInElement`` is the partial application C# does not give for free,
and is used *only* for the ``<v>`` element of a list item and the
positional ``v1``, ``v2``, ... of a tuple item. It carries a doc comment
saying so, because its payoff is invisible from its call count: it is
what lets ``WriteList`` and ``WriteTupleN`` take a *single* item-writer
type even though a class item writes its own run-time-named element
while a primitive item needs a fixed positional one. The ``tuples``
fixture's ``Tuple[int, Some_item, Abstract_item, Some_item,
Positive_int, Result]`` mixes the two item by item; with two item-writer
types it would have needed a combinator per combination.

Names are kept distinct from the read side's on purpose:
``WrapInElement`` / ``WriteEnum`` / ``WriteList`` / ``WriteTupleN``
against ``AtElement`` / ``AsEnum`` / ``AsList`` / ``AsTupleN``. There is
no dual of ``AsText`` (``WriteValue`` cannot fail) and none of
``AsElement``.

The visitor stays, reduced to a dispatch shim: a single ``_instance``
plus static ``WriteIClass``, which one virtual call resolves for every
abstract class, every concrete class with descendants and -- through
``WriteIUnion`` -- every named union. That is why the write side needs
neither the read side's 13 per-interface dispatchers nor a combinator to
invoke one.

Along the way, two things shared with the read side:

* The gating pass is now recursive and serves both emitters.
  ``_needed_content_readers`` walked only one level, so a model whose
  only ``str`` sat inside ``List[List[str]]`` would have emitted
  ``ReadString = AsText<string>(ReadContentAsString, "")`` while
  generating neither ``AsText`` nor ``ReadContentAsString`` -- output
  that does not compile. Latent only because the per-property gate
  blocks such models today.
* ``_content_types_in_initialization_order`` replaces the two copies of
  the registration recursion, so the readers and the writers are
  de-duplicated and ordered by one walk.

Nesting composes for free in this design -- ``WriteListOfListOfString =
WriteList<List<string>>(WrapInElement(WriteListOfString, "v"))`` falls
out with no new combinator -- but stays gated, as on the read side,
since ``copying``, ``enhancing``, ``jsonization`` and the test
generators bail on it in C# and in all five other backends. The write
side now reports that as a proper ``Error`` instead of raising
``NotImplementedError``.

The snippet contract for an implementation-specific class changes:
``Xmlization/VisitorWithWriter/{cls}_to_sequence.cs`` must now be
``static``, and ``visit_{cls}.cs`` must call it without ``this.`` --
the same change the read side already imposed on
``..._from_sequence.cs``.
mristin added a commit that referenced this pull request Sep 13, 2026
This is the write-side dual of #686, and finishes the decomposition
of ``csharp/lib/_generate_xmlization.py``: every property is now written
by one call whose content writer has been composed once into a field,
instead of by a lambda built at the call site.

Measured on the AAS meta-model fixture, the writer region of
``xmlization.cs`` goes 3,908 -> 2,464 lines (-37%), and the file as
a whole 12,471 -> 11,023. Per property 8.1 -> ~5 lines:

    if (that.SemanticId != null)
    {
        WriteElement(
            "semanticId", that.SemanticId, writer, WriteIReference);
    }

The line count is the lesser win. Every content serializer used to be
a capturing lambda composed at the call site -- 212 ``this.`` references
in that region -- which allocates a closure on every single write; the
read side measured that at 152 bytes/op and rejected it. All 276 call
sites now read a ``static readonly`` field instead, so the write path
allocates nothing regardless of the consumer's C# version.

One delegate, ``ContentWriter<in T>``, replaces
``ElementContentSerializer<T>``. ``in T`` is the dual of the read side's
``out T``: it is what lets the single static ``WriteIClass(Aas.IClass,
...)`` serve as the item writer of a list or a tuple of any interface.
Deliberately *one* delegate, not a content/element pair: on the write
side both would have identical signatures, since the content/element
split is a read-side concern (``isEmpty``).

``WriteElement`` and ``WrapInElement`` are one body with two entry
points. The applied form is what the 276 property sites call;
``WrapInElement`` is the partial application C# does not give for free,
and is used *only* for the ``<v>`` element of a list item and the
positional ``v1``, ``v2``, ... of a tuple item. It carries a doc comment
saying so, because its payoff is invisible from its call count: it is
what lets ``WriteList`` and ``WriteTupleN`` take a *single* item-writer
type even though a class item writes its own run-time-named element
while a primitive item needs a fixed positional one. The ``tuples``
fixture's ``Tuple[int, Some_item, Abstract_item, Some_item,
Positive_int, Result]`` mixes the two item by item; with two item-writer
types it would have needed a combinator per combination.

Names are kept distinct from the read side's on purpose:
``WrapInElement`` / ``WriteEnum`` / ``WriteList`` / ``WriteTupleN``
against ``AtElement`` / ``AsEnum`` / ``AsList`` / ``AsTupleN``. There is
no dual of ``AsText`` (``WriteValue`` cannot fail) and none of
``AsElement``.

The visitor stays, reduced to a dispatch shim: a single ``_instance``
plus static ``WriteIClass``, which one virtual call resolves for every
abstract class, every concrete class with descendants and -- through
``WriteIUnion`` -- every named union. That is why the write side needs
neither the read side's 13 per-interface dispatchers nor a combinator to
invoke one.

Along the way, two things shared with the read side:

* The gating pass is now recursive and serves both emitters.
  ``_needed_content_readers`` walked only one level, so a model whose
  only ``str`` sat inside ``List[List[str]]`` would have emitted
  ``ReadString = AsText<string>(ReadContentAsString, "")`` while
  generating neither ``AsText`` nor ``ReadContentAsString`` -- output
  that does not compile. Latent only because the per-property gate
  blocks such models today.
* ``_content_types_in_initialization_order`` replaces the two copies of
  the registration recursion, so the readers and the writers are
  de-duplicated and ordered by one walk.

Nesting composes for free in this design -- ``WriteListOfListOfString =
WriteList<List<string>>(WrapInElement(WriteListOfString, "v"))`` falls
out with no new combinator -- but stays gated, as on the read side,
since ``copying``, ``enhancing``, ``jsonization`` and the test
generators bail on it in C# and in all five other backends. The write
side now reports that as a proper ``Error`` instead of raising
``NotImplementedError``.

The snippet contract for an implementation-specific class changes:
``Xmlization/VisitorWithWriter/{cls}_to_sequence.cs`` must now be
``static``, and ``visit_{cls}.cs`` must call it without ``this.`` --
the same change the read side already imposed on
``..._from_sequence.cs``.
mristin added a commit that referenced this pull request Sep 13, 2026
This is the Go counterpart of #686, but the C# design does not carry
over: the shape that works there fails to compile in Go, so the read
side is composed out of generated dispatchers instead of composed
readers.

Measured on the AAS meta-model fixture, the read region of
``xmlization.go`` goes 11,738 -> 5,584 lines (-52%), and the file as
a whole 22,846 -> 16,692. The ``read<X>WithLookahead`` band -- 50
functions, 3,057 lines -- goes to **zero**; 20 ``read<X>Dispatched``
functions stand in its place.

The whole design turns on one asymmetry in the wire format. An instance
element is self-describing: its local name *is* its type. A scalar
element is not: its name is its *position*, ``v`` in a list and ``v1``,
``v2``, ... in a tuple. So: every instance is read by dispatch on its
own name, every scalar by a name its container supplies, and nothing
else in the read path knows the difference.

``readElementDispatched`` is the only place which frames an XML element
around a value. ``readListOf`` and every ``readTupleN`` delegate the
framing to it, which is what deletes ``readListOfScalars``,
``readScalarWithName``, ``asScalarTupleItemReader`` and
``asInstanceTupleItemReader`` outright: one list helper, one item shape,
no adapters. It also returns the token *past* the end element, so the
"``read<X>WithLookahead`` stops at the end element, so we look ahead"
comment and its hand-written ``readNext`` vanish from every call site:

    readListOf(decoder, current, readReferenceDispatched)
    readListOf(decoder, current, readSubmodelElementDispatched)
    readListOf(decoder, current, readLongAtV)

    readTuple6(
        decoder, current,
        readLongAtV1,
        readSomeItemDispatched,
        readAbstractItemDispatched,
        readSomeItemDispatched,
        readLongAtV5,
        readResultAtV6,
    )

Per class, ``read<X>AsSequence`` has for every property case one
assignment of the same ``(value, current, valueErr)`` triple:

    case "supplementalSemanticIds":
        theSupplementalSemanticIDs, current, valueErr = readListOf(
            decoder, current, readReferenceDispatched,
        )
    case "valueType":
        theValueType, current, valueErr = readOptional(
            readTextAsDataTypeDefXSD(decoder, current),
        )

``readOptional`` absorbs the ``var value T; ...; theX = &value`` dance.
It takes the *results* of a read rather than the reader, so Go's
multi-value pass-through lets one helper serve any read however many
arguments that read takes -- including ``readOptional(readTuple2(...))``,
which a reader-taking signature cannot express, since a tuple's item
readers vary in number and in type. That also fixes an optional tuple
property, which is a pointer (a Go struct is not nilable) and used to be
assigned a non-pointer ``readTupleN`` result: output that does not
compile.
mristin added a commit that referenced this pull request Sep 13, 2026
Write-side dual of #688, as #687 was of #686 in C#.

Measured on the AAS meta-model fixture, the serialization region of
``xmlization.go`` goes 11,080 -> 4,727 lines (-57%), and the file as
a whole 16,692 -> 10,330. ``writeExtensionAsSequence`` and its six
properties go 180 -> 69 lines.

Everything is written through one shape,
``func(encoder, value T) error``. A stringification, a class's own
``write<X>AsSequence`` and the generated item, list and tuple writers
all have it, so ``writeElement`` -- which frames an XML element around
a value written that way -- is the only writer a property needs,
whatever kind of value it holds. What is left per property is that one
call and a decision on presence:

    err = finishProperty(
        "SemanticID()",
        writeOptionalInstance(
            encoder,
            "semanticId",
            that.SemanticID(),
            writeReferenceAsSequence,
        ),
    )
    if err != nil {
        return
    }

``finishProperty`` takes the *result* of the write rather than the write
itself -- the same trick as the read side's ``readOptional``, since Go
evaluates the argument first -- so one function concludes every
property, prepending the getter to the path of the error.

Mind that the getter is spelled out as it is in Go, ``SemanticID()`` and
not ``semanticId``, since a serialization error renders its path through
``aasreporting.ToGolangPath``, while the de-serialization prepends
the XML name, its errors being XPaths. Go has no ``nameof``.

The ``writeOptional*`` family has one member per way Go spells an
optional, each named after that spelling, as the check for the presence
is the only thing which differs between them: a pointer to the value
(``writeOptionalPointer``: a scalar, an enumeration, a tuple), a value
which is nil on its own (``writeOptionalInstance``: an instance, or
a named union, which is a pointer to a struct) and a slice
(``writeOptionalSlice``: a list, the bytes). They can not be collapsed
into one, since Go decides nil-ness by the representation: a type
parameter can not be compared against ``nil``, ``any(that) == nil`` is
false for a nil *pointer* boxed in an ``any`` -- hence the comparison
against the zero value in ``writeOptionalInstance``, which in turn
panics on a slice, which is not comparable. Each of them writes nothing
at all when the value is absent, which is what keeps the absence out of
the generated code.

No closure is allocated on the write path any more: every writer passed
on is a top-level function or an instantiated generic, hence a static
funcval, exactly as on the read side. A list and a tuple get one
generated content writer as well, per item type resp. per tuple
type -- ``writeListOfIReference`` and ``writeTupleOfStringLong`` and
their kin -- since ``writeList`` and ``writeTupleN`` take a writer per
item, and Go gives no partial application to bind those in without
allocating.

Finally, the flush. ``write<X>AsSequence`` used to flush the encoder
after every property, and ``write<X>`` again at its end element, which
defeated the buffering of ``xml.Encoder`` on every write; only the flush
after the last token is needed, as ``EncodeToken`` deliberately leaves
it to the caller. ``Marshal`` is therefore split into ``writeClass``,
which dispatches on the model type, and ``Marshal`` itself, which calls
it and flushes exactly once; ``writeInstance`` and ``writeUnion`` go to
``writeClass``, so a nested instance does not flush either. Serializing
an instance holding 1,000 instances (42 KB of XML) from
the ``list_of_classes`` fixture, the writes which actually reach
the underlying writer go 2,003 -> 11, one per 4 KB of buffer instead of
one per property and per instance; to an ``os.File`` that is 5.29 ->
0.76 ms per serialization, the median of five.
mristin added a commit that referenced this pull request Sep 14, 2026
Java is the third backend to have its XML read side decomposed, after C#
(#686) and Go (#688), and it carried the repetition #686 described at four
levels rather than three.

Every property inlined its whole reading procedure into its own ``case``
block. Everything is now read through one composable shape, and
the pieces which do not vary are generated once:

    case "contentType": {
      final Reporting.Result<String> value =
        readTextAs_string(reader, isEmptyProperty);
      if (value.isError()) {
        valueError = value.getError();
      } else {
        theContentType = value.getResult();
      }
      break;
    }

The rationale and the design:

* A method reference to a ``static`` method captures nothing, so its
  ``invokedynamic`` call site is linked once and hands back
  the same instance afterwards; passing
  ``_DeserializeImplementation::readReference`` costs nothing.

  A capturing lambda is not free, and the old code had one in
  every ``try<X>FromElement`` and in every dispatcher, both closing over the
  enclosing ``reader``: an allocation on every single instance read. The
  read side now contains no lambdas at all -- every reader passed
  anywhere is a method reference to a static method -- so nothing can
  accidentally capture.

* Two shapes, and two only. ``ContentReader<T>`` reads the content of an
  element which has already been opened, ``ElementReader<T>`` reads a whole
  element. They are duals, and the wire format forces both, because it is
  asymmetric: an instance element is self-describing -- its local name *is*
  its type -- whereas a scalar element is not, its name being its position
  (the property's name, ``v`` in a list, ``v1``, ``v2``, ... in a tuple).

* One primitive, ``peekElementName``. It consumes nothing, and it is the
  single answer to "we are at an element, and this is its name". The framer
  ``readNamedElement`` checks that name against the one its container
  supplied, the dispatchers switch on it, and every property loop
  selects the property with it. That is also why no third shape is needed to
  carry the name into a dispatcher: a dispatcher peeks, switches, and hands
  the *whole* element on to that type's own ``read<Cls>FromElement``, which
  re-checks the name and consumes it.

* The property loop keeps only its ``switch``, because its branches
  assign the locals the constructor is called with. ``atEndOfSequence`` is
  the loop condition, ``consumeEndElement`` -- shared with
  ``readNamedElement``, which concludes identically -- consumes the end tag,
  and ``unexpectedProperty`` and ``missingRequiredProperty`` are the two
  errors which used to be spelled out per class and per property.

* The property is marked on the error path once per class, after the
  ``switch``, rather than in every case: for a matched case the element
  name already is the property's XML name. That also retires the per-case
  ``castTo(<Cls>.class)`` re-threading, and fixes an inconsistency along the
  way -- a primitive and an enumeration used to be marked with the *Java*
  property name while a list and a class used the XML name.

* The names are unique by construction, so no uniqueness check is
  needed. A reader is keyed either on one symbol or on a type moniker, and
  the two name spaces are kept apart by a single rule: a moniker-keyed name
  contains ``_``, a symbol-keyed name never does. That holds because
  ``naming.capitalized_camel_case``, which every one of our type names goes
  through, emits letters and digits only, with an upper-case initial. The
  monikers are then a Polish notation over ``_``-separated tokens --
  ``ListOf_string``, ``TupleOf2_string_long``,
  ``ListOf_TupleOf2_string_long`` -- with the arity spelled out, so the
  encoding is injective. The primitives are lower-cased for the same reason:
  an enumeration which somebody names ``String`` can then never collide with
  the primitive ``str``.

* ``readNestedElement`` looks redundant and is not. A polymorphic
  property and a list item leave the reader at the very same position, just
  before a self-describing start element, so it is tempting to give
  ``read<X>FromElement`` an ``isEmpty`` parameter and call it from the
  property ``case`` directly. The error path is what forbids it. A property
  wraps its instance in an element of its own, so the discriminator's name
  has to be prepended and the path reads ``value/property/idShort``; a list
  item is not wrapped -- the item element *is* the indexed child -- so the
  same prepend would give ``annotations/*[0]/property/idShort``, which walks
  one level past the element ``*[0]`` already selects and resolves to
  nothing. The reasoning is written into the generated doc comment as well,
  because it is exactly the kind of layer the next reader deletes.

* Only what a model reaches is emitted, gated along the call graph
  rather than by property kind -- a list of enumerations needs the
  enumeration reader even when no property is one, and an enumeration needs
  the text path because it is built on it. The pass recurses into the items
  of a list and of a tuple, without which a model whose only ``str`` sits in
  a ``List[List[str]]`` would emit a reader built from helpers which were
  never generated.

The snippet contract for an implementation-specific class changes:
``Xmlization/DeserializeImplementation/{cls}_from_element.java`` is no
longer read, and the remaining ``{cls}_from_sequence.java`` has to define
``read<Cls>FromSequence`` rather than ``try<Cls>FromSequence``.

For the AAS meta-model, ``Xmlization.java`` goes from 16,647 to 11,455
lines (-31.2%) and its read region from 12,187 to 6,995 (-42.6%).

Finally, the speed, since the read side is where a de-serialization library
is judged: there is no measurable difference, and that is the expected
answer. De-serializing a 1.4 MB AAS environment (200 submodels, 131,806
StAX events; 400 runs after 200 warm-up runs, three measurements of each
variant, run alternately), the medians are 32.2/33.2/33.2 ms before and
32.1/32.7/33.2 ms after -- the same, within the noise of repeated runs of
either one.

Draining that very same document through ``XMLEventReader`` without
building anything at all already costs 27.4 ms and allocates 41.2 MB, so
StAX accounts for ~83% of the time and ~84% of the allocation, and
everything the generated code does is the remaining ~6 ms and ~8.2 MB.
Against that, the lambdas which are gone are worth ~0.15 MB per run at the
median (49.48-49.57 MB before, 49.14-49.44 MB after): consistently in that
direction, but under 1% of the total. The decomposition is
a maintainability change, not a performance one, and it was worth
measuring rather than assuming, since the C# commit's 152 bytes/op is what
a reader would expect to carry over.
mristin added a commit that referenced this pull request Sep 15, 2026
Write-side dual of #690, as #687 was of #686 in C# and #689 of #688 in Go.

``visit{Cls}`` was a verbatim copy of ``serializeElement``: 38 methods of
18 lines each spelling out the start tag, the namespace, the content and
the end tag, around a call which already did exactly that. Six
near-identical emitters -- one per property kind -- produced the very same
``serializeElement`` call, differing only in the content serializer. And
everything was an *instance* method, because of one mutable ``topLevel``
flag.

That flag is what cost the most. It made ``this::writeStringifiedContent``,
``this::visit`` and ``(value, w) -> this.xToSequence(value, w)`` capturing,
and ``serializeItems(...)`` and ``asNamedElementSerializer(...)`` returned
capturing lambdas on top of them, so writing one instance allocated
a small pile of closures. It is replaced by a ``final boolean
withNamespace`` pinned in the constructor of two immutable visitors,
``ROOT`` and ``NESTED``: only the outermost element declares the
namespace, and an element written from ``writeClass`` is nested by
construction. Everything else is then ``static``, and a method reference
to a static method captures nothing.

The design is otherwise the read side's, mirrored:

* One shape, not two. ``ContentWriter<T>`` writes a value where the
  writer already is; ``writeElement`` frames it in a start and an end tag.
  There is deliberately no second delegate for a whole element, as there
  is on the reading side -- an element differs from a content only in what
  it writes, never in its shape. The reading needs the distinction because
  a content reader has to be told whether its element was self-closing,
  and a writer has nothing to be told. Use sites take
  ``ContentWriter<? super T>``, which is how Java spells the
  contravariance C# writes as ``in T``, so the single writer of an
  ``IClass`` serves wherever the writer of a more specific interface is
  expected.

* ``visit{Cls}`` collapses to one call, and the visitor becomes
  a dispatch shim::

      @OverRide
      public void visitQualifier(
        IQualifier that,
        XMLStreamWriter writer) {
        writeElement(
          "qualifier",
          that,
          writer,
          withNamespace,
          _VisitorWithWriter::writeQualifierAsSequence);
      }

* The families are dual to the readers and named by the same moniker
  grammar: ``write{Cls}AsSequence``, ``writeClass``/``writeUnion``,
  ``writeListOf_{M}``, ``writeTupleOf{N}_{M...}``, ``writeAtV_{M}`` and
  ``writeAtV{i}_{M}``. The list and the tuple writers inline their loop,
  which retires ``serializeItems``, ``asNamedElementSerializer`` and
  ``serializeTuple{N}`` -- Java needs no partial application here, since
  the name to bind is fixed per generated function.

* One property emitter replaces the six. Every property is written by
  ``writeProperty``, or by ``writeOptionalProperty``, which takes the
  ``Optional`` apart once instead of asking it and then unwrapping it --
  the old form called the getter twice, and every call allocates an
  ``Optional``.

* ``_type_moniker``, ``_leaf_moniker``, ``_is_instance_type``,
  ``_is_dispatched`` and ``_collect_needed`` are reused unchanged: the
  reading and the writing are composed out of the same shapes, so one
  gating pass answers for both. Only the two dispatching writers needed
  a walk of their own, because the *position* matters on the write side
  while it does not on the read side: an item of any class is written as
  its own element, whereas a property of a concrete class without
  descendants writes its properties straight into the property element.

Two things are genuinely asymmetric, and both save work:

* There is no ``writeTextAs_{primitive}`` family. Once the value is at
  hand nothing type-specific is left to do with it, so the four
  ``toString`` primitives share ``writeStringifiedContent`` and the byte
  array ``writeByteArrayContent``. An enumeration does have something of
  its own to write and keeps one function each -- named
  ``writeTextAs_{Enum}`` and not ``write{Enum}Content``, so that it stays
  in the moniker-keyed name space: were it keyed on the symbol, an
  enumeration which somebody names ``Stringified`` or ``ByteArray`` would
  silently take a shared writer's name.

* There is no element writer per class. Which element an instance writes
  follows from its run-time type, and one virtual call answers that for
  every class at once. The reading needs a dispatcher per interface
  because it has to decide what to construct before it has read anything.

Two defects of the write side are fixed along the way, both of which the
decomposition makes one-line changes rather than sweeping ones.

``Serialize.to`` never flushed, so the serialization was not finished when
it returned, and a failure of the underlying stream surfaced at the
caller's own flush as a raw ``XMLStreamException`` -- outside the one
method whose ``throws`` clause promises a ``SerializeException``. Probed
with a ``Writer`` which accepts every character and fails on ``flush``,
``to`` used to return normally and now throws. The flush goes in
``Serialize.to`` and nowhere else, as in Go's #689, so
``XMLStreamWriter`` keeps buffering in between. Mind that this does push
the underlying stream: a caller serializing many small documents into one
buffered writer now flushes once per document rather than once per
buffer. That is the right trade, since the alternative leaves the output
silently incomplete whenever the caller forgets.

Every ``SerializeException`` carried an empty path, so every failure
read ``... at: the beginning`` no matter where it happened. The path is
now built as the stack unwinds, each container prepending the one segment
it knows, exactly as on the read side where ``read<Cls>FromSequence``
prepends the property after its ``switch`` and ``readList`` prepends the
index. Probed with an ``XMLStreamWriter`` which refuses to write one
chosen value:

    a property of the root instance   idShort
    nested through two lists          submodels/*[1]/submodelElements/*[1]/value
    through a class-typed property    semanticId/keys/*[0]/value

``SerializeException`` renders its message in its constructor, so its path
has to be complete by the time it is built and it can not be the thing
thrown from inside; a private ``_SerializeFailure`` carries
a ``Reporting.Error`` instead, and ``Serialize.to`` renders that with
``Reporting.generateRelativeXPath`` and converts. The public exception is
untouched. ``writeElement`` contributes no segment and ``writeProperty``
is ``writeElement`` plus the one thing a property knows which nothing
below it does -- its own name -- so there is no ``try`` around any of the
276 property writes. A list and a tuple put their ``try`` outside the
loop and advance the index only once an item has been written, so the
item named is the one which failed and nothing is paid per item.

Two segments are deliberately absent, and the reasoning is written into
the generated Javadoc, since both look like oversights. The outermost
element is not named, because ``Serialize.to`` takes any ``IClass`` and
the name would say nothing the caller does not already know; the read
side does prepend it, having one facade per class. Neither is the
discriminator element of a polymorphic property, which the read side does
prepend -- for a failure in an embedded data specification it reports
``submodel/embeddedDataSpecifications/*[0]/dataSpecificationContent/``
``dataSpecificationIec61360/preferredName/*[0]`` where the write side
reports ``embeddedDataSpecifications/*[0]/dataSpecificationContent/``
``preferredName/*[0]/text``. The reading points into a document it is
parsing, where that element is a real extra level; the writing points
into the instance the caller handed over, where it is not, and
``value/idShort`` there is exactly ``getValue().getIdShort()``. Go's #689
made the same choice when it rendered its serialization paths through
``aasreporting.ToGolangPath`` rather than as XPaths.

For the AAS meta-model the write region of ``Xmlization.java`` goes 3,281
-> 2,873 lines (-12.4%), error paths included. The write region no longer
contains a single lambda.

Unlike the read side, that shows up in the measurement. Serializing an
``Environment`` of 200 submodels (1.14 MB of XML) into a writer which
counts the characters and discards them, 100 runs after 100 warm-up runs,
eight rounds of each variant run alternately: the allocation goes 0.76 ->
0.36 MB per serialization, and the time 7.69 -> 7.49 ms, which is inside
the round-to-round noise. ``XMLStreamWriter`` writes straight through, so
the whole write path allocates ~0.8 MB against the ~49 MB the *reading*
of the same document allocates -- there is no StAX floor here for the
closures to hide behind, which is precisely why #690 could not measure
its own.

The snippet contract for an implementation-specific class changes:
``Xmlization/VisitorWithWriter/visit_{cls}.java`` is no longer read, as
the element around the sequence is now written identically for every
class, and the remaining ``{cls}_to_sequence.java`` has to define
a ``static`` ``write{Cls}AsSequence``.
mristin added a commit that referenced this pull request Sep 15, 2026
Write-side dual of #690, as #687 was of #686 in C# and #689 of #688 in Go.

``visit{Cls}`` was a verbatim copy of ``serializeElement``: 38 methods of
18 lines each spelling out the start tag, the namespace, the content and
the end tag, around a call which already did exactly that. Six
near-identical emitters -- one per property kind -- produced the very same
``serializeElement`` call, differing only in the content serializer. And
everything was an *instance* method, because of one mutable ``topLevel``
flag.

That flag is what cost the most. It made ``this::writeStringifiedContent``,
``this::visit`` and ``(value, w) -> this.xToSequence(value, w)`` capturing,
and ``serializeItems(...)`` and ``asNamedElementSerializer(...)`` returned
capturing lambdas on top of them, so writing one instance allocated
a small pile of closures. It is replaced by a ``final boolean
withNamespace`` pinned in the constructor of two immutable visitors,
``ROOT`` and ``NESTED``: only the outermost element declares the
namespace, and an element written from ``writeClass`` is nested by
construction. Everything else is then ``static``, and a method reference
to a static method captures nothing.

The design is otherwise the read side's, mirrored:

* One shape, not two. ``ContentWriter<T>`` writes a value where the
  writer already is; ``writeElement`` frames it in a start and an end tag.
  There is deliberately no second delegate for a whole element, as there
  is on the reading side -- an element differs from a content only in what
  it writes, never in its shape. The reading needs the distinction because
  a content reader has to be told whether its element was self-closing,
  and a writer has nothing to be told. Use sites take
  ``ContentWriter<? super T>``, which is how Java spells the
  contravariance C# writes as ``in T``, so the single writer of an
  ``IClass`` serves wherever the writer of a more specific interface is
  expected.

* ``visit{Cls}`` collapses to one call, and the visitor becomes
  a dispatch shim::

      @OverRide
      public void visitQualifier(
        IQualifier that,
        XMLStreamWriter writer) {
        writeElement(
          "qualifier",
          that,
          writer,
          withNamespace,
          _VisitorWithWriter::writeQualifierAsSequence);
      }

* The families are dual to the readers and named by the same moniker
  grammar: ``write{Cls}AsSequence``, ``writeClass``/``writeUnion``,
  ``writeListOf_{M}``, ``writeTupleOf{N}_{M...}``, ``writeAtV_{M}`` and
  ``writeAtV{i}_{M}``. The list and the tuple writers inline their loop,
  which retires ``serializeItems``, ``asNamedElementSerializer`` and
  ``serializeTuple{N}`` -- Java needs no partial application here, since
  the name to bind is fixed per generated function.

* One property emitter replaces the six. Every property is written by
  ``writeProperty``, or by ``writeOptionalProperty``, which takes the
  ``Optional`` apart once instead of asking it and then unwrapping it --
  the old form called the getter twice, and every call allocates an
  ``Optional``.

* ``_type_moniker``, ``_leaf_moniker``, ``_is_instance_type``,
  ``_is_dispatched`` and ``_collect_needed`` are reused unchanged: the
  reading and the writing are composed out of the same shapes, so one
  gating pass answers for both. Only the two dispatching writers needed
  a walk of their own, because the *position* matters on the write side
  while it does not on the read side: an item of any class is written as
  its own element, whereas a property of a concrete class without
  descendants writes its properties straight into the property element.

Two things are genuinely asymmetric, and both save work:

* There is no ``writeTextAs_{primitive}`` family. Once the value is at
  hand nothing type-specific is left to do with it, so the four
  ``toString`` primitives share ``writeStringifiedContent`` and the byte
  array ``writeByteArrayContent``. An enumeration does have something of
  its own to write and keeps one function each -- named
  ``writeTextAs_{Enum}`` and not ``write{Enum}Content``, so that it stays
  in the moniker-keyed name space: were it keyed on the symbol, an
  enumeration which somebody names ``Stringified`` or ``ByteArray`` would
  silently take a shared writer's name.

* There is no element writer per class. Which element an instance writes
  follows from its run-time type, and one virtual call answers that for
  every class at once. The reading needs a dispatcher per interface
  because it has to decide what to construct before it has read anything.

Two defects of the write side are fixed along the way, both of which the
decomposition makes one-line changes rather than sweeping ones.

``Serialize.to`` never flushed, so the serialization was not finished when
it returned, and a failure of the underlying stream surfaced at the
caller's own flush as a raw ``XMLStreamException`` -- outside the one
method whose ``throws`` clause promises a ``SerializeException``. Probed
with a ``Writer`` which accepts every character and fails on ``flush``,
``to`` used to return normally and now throws. The flush goes in
``Serialize.to`` and nowhere else, as in Go's #689, so
``XMLStreamWriter`` keeps buffering in between. Mind that this does push
the underlying stream: a caller serializing many small documents into one
buffered writer now flushes once per document rather than once per
buffer. That is the right trade, since the alternative leaves the output
silently incomplete whenever the caller forgets.

Every ``SerializeException`` carried an empty path, so every failure
read ``... at: the beginning`` no matter where it happened. The path is
now built as the stack unwinds, each container prepending the one segment
it knows, exactly as on the read side where ``read<Cls>FromSequence``
prepends the property after its ``switch`` and ``readList`` prepends the
index. Probed with an ``XMLStreamWriter`` which refuses to write one
chosen value:

    a property of the root instance   idShort
    nested through two lists          submodels/*[1]/submodelElements/*[1]/value
    through a class-typed property    semanticId/keys/*[0]/value

``SerializeException`` renders its message in its constructor, so its path
has to be complete by the time it is built and it can not be the thing
thrown from inside; a private ``_SerializeFailure`` carries
a ``Reporting.Error`` instead, and ``Serialize.to`` renders that with
``Reporting.generateRelativeXPath`` and converts. The public exception is
untouched. ``writeElement`` contributes no segment and ``writeProperty``
is ``writeElement`` plus the one thing a property knows which nothing
below it does -- its own name -- so there is no ``try`` around any of the
276 property writes. A list and a tuple put their ``try`` outside the
loop and advance the index only once an item has been written, so the
item named is the one which failed and nothing is paid per item.

Two segments are deliberately absent, and the reasoning is written into
the generated Javadoc, since both look like oversights. The outermost
element is not named, because ``Serialize.to`` takes any ``IClass`` and
the name would say nothing the caller does not already know; the read
side does prepend it, having one facade per class. Neither is the
discriminator element of a polymorphic property, which the read side does
prepend -- for a failure in an embedded data specification it reports
``submodel/embeddedDataSpecifications/*[0]/dataSpecificationContent/``
``dataSpecificationIec61360/preferredName/*[0]`` where the write side
reports ``embeddedDataSpecifications/*[0]/dataSpecificationContent/``
``preferredName/*[0]/text``. The reading points into a document it is
parsing, where that element is a real extra level; the writing points
into the instance the caller handed over, where it is not, and
``value/idShort`` there is exactly ``getValue().getIdShort()``. Go's #689
made the same choice when it rendered its serialization paths through
``aasreporting.ToGolangPath`` rather than as XPaths.

For the AAS meta-model the write region of ``Xmlization.java`` goes 3,281
-> 2,873 lines (-12.4%), error paths included. The write region no longer
contains a single lambda.

Unlike the read side, that shows up in the measurement. Serializing an
``Environment`` of 200 submodels (1.14 MB of XML) into a writer which
counts the characters and discards them, 100 runs after 100 warm-up runs,
eight rounds of each variant run alternately: the allocation goes 0.76 ->
0.36 MB per serialization, and the time 7.69 -> 7.49 ms, which is inside
the round-to-round noise. ``XMLStreamWriter`` writes straight through, so
the whole write path allocates ~0.8 MB against the ~49 MB the *reading*
of the same document allocates -- there is no StAX floor here for the
closures to hide behind, which is precisely why #690 could not measure
its own.

The snippet contract for an implementation-specific class changes:
``Xmlization/VisitorWithWriter/visit_{cls}.java`` is no longer read, as
the element around the sequence is now written identically for every
class, and the remaining ``{cls}_to_sequence.java`` has to define
a ``static`` ``write{Cls}AsSequence``.
mristin added a commit that referenced this pull request Sep 22, 2026
C++ is the last backend to have its XML read side decomposed, after C#
(#686), Go (#688), Java (#690), Python (#694) and TypeScript (#698).

#673 already factored out the framing of an *element*, but every
``<Cls>FromSequence`` still re-emitted the same ~220 lines of property-loop
framing around its ``switch``, so a property cost nineteen lines where it
now costs three:

    case properties::OfExtension::kSemanticId:
      return ReadInto(
        the_semantic_id,
        ReferenceFromSequence<
          types::IReference
        >(reader)
      );

The rationale and the design:

* **The framing does not need ``T``.** Returning the error alone, instead
  of ``NoInstanceAndDeserializationError<std::shared_ptr<T> >`` on every
  failure path, is the whole reason the loop can be lifted out of the class:

      template <
        std::size_t kPropertyCount,
        typename EnumT,
        typename OnPropertyT
      >
      common::optional<DeserializationError> ReadProperties(
        xml_common::ReaderMergingText& reader,
        const std::unordered_map<std::string, EnumT>& map_of_properties,
        const wchar_t* interface_name,
        const OnPropertyT& on_property
      );

* **The ``switch`` stays inline, and in C++ that is free.** Its branches
  assign the locals the constructor is called with -- which is why C#, Go
  and Java had to keep the whole loop inline -- but a lambda passed as
  ``const F&`` to a function template is a stack object which is inlined
  away, so the class can hand the ``switch`` over and nothing is allocated:

      common::optional<DeserializationError> error(
        ReadProperties<
          properties::kPropertyCountOfExtension
        >(
          reader,
          properties::kMapOfExtension,
          L"IExtension",
          [&](
            properties::OfExtension property
          ) -> common::optional<DeserializationError> {
            switch (property) {
              ...
            }
          }
        )
      );

* **The duplicate check belongs to the loop, not to the property.** It is
  now a bit set sized to the class and living on the stack, so a class may
  have arbitrarily many properties and a read allocates nothing for it; the
  check still precedes the read, so a duplicate is refused without its
  content ever being looked at, and a five-line guard leaves all 276 cases:

      std::bitset<kPropertyCount> seen;
      ...
      if (seen[index]) {
        error = DuplicatePropertyError(name);
        PrependElementSegmentToDeserializationError(name, *error);
        return error;
      }

      seen[index] = true;

* **``ReadInto`` takes the results of a read, not the reader.** The item
  readers of a tuple vary both in number and in type, so no signature
  taking the reader could serve every read, while one taking the result
  serves all of them:

      template <typename T>
      common::optional<DeserializationError> ReadInto(
        common::optional<T>& target,
        std::pair<
          common::optional<T>,
          common::optional<DeserializationError>
        >&& read
      );

* **A named union is framed exactly like a class.**
  ``DeserializeUnionFromElement`` and ``DeserializeClassFromElement`` were
  ~160 near-identical lines which differed only in the value type, and
  every error factory they called was already generic in it:

      template <typename ValueT, typename DispatchT>
      std::pair<
        common::optional<ValueT>,
        common::optional<DeserializationError>
      > DeserializeFromElement(
        xml_common::ReaderMergingText& reader,
        const wchar_t* value_name,
        const DispatchT& dispatch
      );

* **An element with a sole model type does not dispatch.** 40 of the 51
  ``<Cls>FromElement`` spelled out a one-case ``switch`` inside a
  fifteen-line lambda signature; the remaining 11 keep theirs, since
  a dense jump table beats a lookup:

      std::pair<
        common::optional<
          std::shared_ptr<types::IExtension>
        >,
        common::optional<DeserializationError>
      > ExtensionFromElement(
        xml_common::ReaderMergingText& reader
      ) {
        return DeserializeSoleFromElement<
          std::shared_ptr<types::IExtension>
        >(
          reader,
          L"IExtension",
          types::ModelType::kExtension,
          ExtensionFromSequence<types::IExtension>
        );
      }

* **The public ``<Cls>From`` are one function.** 51 copies of the same 54
  lines -- open the reader, read the root element, check that nothing but
  whitespace follows it -- became a single call each:

      return DeserializeFrom<
        std::shared_ptr<types::IExtension>
      >(
        is,
        options,
        ExtensionFromElement
      );

* **``interface_name`` is a ``const wchar_t*``.** As a
  ``const std::wstring&`` it made every call site construct a
  ``std::wstring`` on *every* element read, success path included, for
  a message which is almost never built:

      const wchar_t* interface_name,  // was: const std::wstring&

Everything is written for C++11, which the library targets outside of its
tests: no generic lambda, no deduced return type, and every lambda spells
its parameters and its return type out.

For the AAS meta-model, ``xmlization.cpp`` goes from 34,543 to 24,617
lines and its read region from 23,584 to 13,658 (-42%). The
``<Cls>FromSequence`` band goes 13,400 -> 6,076 over 37 classes (-54.7%):
``Key`` 257 -> 95, ``AdministrativeInformation`` 295 -> 112, ``Extension``
331 -> 141. The public ``<Cls>From`` go 2,751 -> 813 (-70%). The
``unions`` fixture goes 7,950 -> 5,795. ``empty_class`` goes 1,622 ->
1,674, since a meta-model with a single property-less class pays for the
shared loop without having anything to save.

Two warts are cleaned up along the way: a class without properties used to
get an enumeration and a map spelled out over a blank line with trailing
whitespace, and the dispatch lambdas now separate their closing angle
brackets the way the rest of the generated code does.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
mristin added a commit that referenced this pull request Sep 22, 2026
C++ is the last backend to have its XML read side decomposed, after C#
(#686), Go (#688), Java (#690), Python (#694) and TypeScript (#698). #673
factored out the framing of an *element*; this factors out the framing of
the property *loop*, which every ``<Cls>FromSequence`` re-emitted in full.

**A property** loses its duplicate check and its ``std::tie``, and names
only itself and what reads it.

    // before
    case properties::OfExtension::kSemanticId: {
      if (the_semantic_id.has_value()) {
        error = DuplicatePropertyError(name);
        break;
      }

      std::tie(
        the_semantic_id,
        error
      ) = ReferenceFromSequence<
        types::IReference
      >(reader);
      break;
    }

    // after
    case properties::OfExtension::kSemanticId:
      return ReadInto(
        the_semantic_id,
        ReferenceFromSequence<
          types::IReference
        >(reader)
      );

**The loop** around it becomes one call. It could be lifted out at all only
because the framing now returns the error alone, instead of
``NoInstanceAndDeserializationError<std::shared_ptr<T> >``, and so no longer
needs ``T``.

    // before, ~220 lines in each of the 38 ``<Cls>FromSequence``
    while (true) {
      error = SkipWhitespace(reader);
      if (error.has_value()) { ...6 lines... }
      if (reader.node().kind() == xml_common::NodeKind::Stop) {
        break;
      } else if (reader.node().kind() != xml_common::NodeKind::Start) {
        ...10 lines naming IExtension...
      }
      const std::string name(...);
      reader.Read();
      auto it = properties::kMapOfExtension.find(name);
      if (it == properties::kMapOfExtension.end()) { ...12 lines... }
      switch (it->second) { ... }
      ...~60 lines of prepending and stop-element handling...
    }

    // after
    common::optional<DeserializationError> error(
      ReadProperties<
        properties::kPropertyCountOfExtension
      >(
        reader,
        properties::kMapOfExtension,
        L"IExtension",
        [&](
          properties::OfExtension property
        ) -> common::optional<DeserializationError> {
          switch (property) { ... }
        }
      )
    );

The ``switch`` stays inline because its branches assign the locals the
constructor is called with -- the very reason C#, Go and Java had to keep
the whole loop inline. The lambda is a template argument taken by
``const&``, so the call is statically bound and nothing is type-erased or
allocated; ``g++ -O2`` keeps ``ReadProperties`` itself out of line, one
instantiation per class, which is as many function bodies as the loops it
replaces.

**The duplicate check** is bookkeeping of the loop, not of the property, so
it moved there, as a bit set sized to the class. It still precedes the read,
so a duplicate is refused without its content ever being looked at.

    std::bitset<kPropertyCount> seen;
    ...
    if (seen[index]) {
      error = DuplicatePropertyError(name);
      PrependElementSegmentToDeserializationError(name, *error);
      return error;
    }

    seen[index] = true;

**``ReadInto`` takes the results of a read, not the reader.** A tuple's item
readers vary in number *and* in type, so no reader-taking signature serves
them all, while one taking the result serves every read.

    template <typename T>
    common::optional<DeserializationError> ReadInto(
      common::optional<T>& target,
      std::pair<
        common::optional<T>,
        common::optional<DeserializationError>
      >&& read
    );

**A named union is framed like a class.** The two framings were ~160
near-identical lines differing only in the value type, which every error
factory they call is already generic in.

    // before
    template <typename T, typename DispatchT>
    ... DeserializeClassFromElement(reader, const std::wstring&, dispatch);
    template <typename VariantT, typename DispatchT>
    ... DeserializeUnionFromElement(reader, const std::wstring&, dispatch);

    // after
    template <typename ValueT, typename DispatchT>
    ... DeserializeFromElement(reader, const wchar_t*, dispatch);

**An element with a sole model type does not dispatch.** 40 of the 51
``<Cls>FromElement`` spelled out a one-case ``switch`` inside a fifteen-line
lambda signature; the other 11 keep theirs.

    return DeserializeSoleFromElement<
      std::shared_ptr<types::IExtension>
    >(
      reader,
      L"IExtension",
      types::ModelType::kExtension,
      ExtensionFromSequence<types::IExtension>
    );

**The public ``<Cls>From``** were 51 copies of the same 54 lines: open the
reader, read the root element, check that nothing but whitespace follows.

    return DeserializeFrom<
      std::shared_ptr<types::IExtension>
    >(
      is,
      options,
      ExtensionFromElement
    );

**``interface_name`` is a ``const wchar_t*``**, not a ``const std::wstring&``:
every call site used to construct a ``std::wstring`` on *every* element read,
success path included, for a message which is almost never built.

Written for C++11, which the library targets outside of its tests: no
generic lambda, no deduced return type, every lambda spells out its
parameters and its return type.

``xmlization.cpp`` for the AAS meta-model goes 34,543 -> 24,634 lines and
its read region 23,584 -> 13,675 (-42%). The ``<Cls>FromSequence`` band goes
13,400 -> 6,076 over 37 classes (-54.7%), with ``Key`` 257 -> 95 and
``Extension`` 331 -> 141, and the public ``<Cls>From`` 2,751 -> 813 (-70%).
``unions`` goes 7,950 -> 5,795; ``empty_class`` goes 1,622 -> 1,674, since
one property-less class pays for the shared loop with nothing to save.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
mristin added a commit that referenced this pull request Sep 22, 2026
C++ is the last backend to have its XML read side decomposed, after C#
(#686), Go (#688), Java (#690), Python (#694) and TypeScript (#698). #673
factored out the framing of an *element*; this factors out the framing of
the property *loop*, which every ``<Cls>FromSequence`` re-emitted in full.

**A property** loses its duplicate check and its ``std::tie``, and names
only itself and what reads it.

    // before
    case properties::OfExtension::kSemanticId: {
      if (the_semantic_id.has_value()) {
        error = DuplicatePropertyError(name);
        break;
      }

      std::tie(
        the_semantic_id,
        error
      ) = ReferenceFromSequence<
        types::IReference
      >(reader);
      break;
    }

    // after
    case properties::OfExtension::kSemanticId:
      return ReadInto(
        the_semantic_id,
        ReferenceFromSequence<
          types::IReference
        >(reader)
      );

**The loop** around it becomes one call. It could be lifted out at all only
because the framing now returns the error alone, instead of
``NoInstanceAndDeserializationError<std::shared_ptr<T> >``, and so no longer
needs ``T``.

    // before, ~220 lines in each of the 38 ``<Cls>FromSequence``
    while (true) {
      error = SkipWhitespace(reader);
      if (error.has_value()) { ...6 lines... }
      if (reader.node().kind() == xml_common::NodeKind::Stop) {
        break;
      } else if (reader.node().kind() != xml_common::NodeKind::Start) {
        ...10 lines naming IExtension...
      }
      const std::string name(...);
      reader.Read();
      auto it = properties::kMapOfExtension.find(name);
      if (it == properties::kMapOfExtension.end()) { ...12 lines... }
      switch (it->second) { ... }
      ...~60 lines of prepending and stop-element handling...
    }

    // after
    common::optional<DeserializationError> error(
      ReadProperties<
        properties::kPropertyCountOfExtension
      >(
        reader,
        properties::kMapOfExtension,
        L"IExtension",
        [&](
          properties::OfExtension property
        ) -> common::optional<DeserializationError> {
          switch (property) { ... }
        }
      )
    );

The ``switch`` stays inline because its branches assign the locals the
constructor is called with. The lambda is a template argument taken by
``const&``, so the call is statically bound and nothing is type-erased or
allocated; ``g++ -O2`` keeps ``ReadProperties`` itself out of line, one
instantiation per class, which is as many function bodies as the loops it
replaces.

**The duplicate property element check** is bookkeeping of the loop,
not of the property, so it moved there, as a bit set sized to
the class. It still precedes the read, so a duplicate is refused without
its content ever being looked at.

    std::bitset<kPropertyCount> seen;
    ...
    if (seen[index]) {
      error = DuplicatePropertyError(name);
      PrependElementSegmentToDeserializationError(name, *error);
      return error;
    }

    seen[index] = true;

**``ReadInto`` takes the results of a read, not the reader.** A tuple's item
readers vary in number *and* in type, so no reader-taking signature serves
them all, while one taking the result serves every read.

    template <typename T>
    common::optional<DeserializationError> ReadInto(
      common::optional<T>& target,
      std::pair<
        common::optional<T>,
        common::optional<DeserializationError>
      >&& read
    );

**A named union is framed like a class.** The two framings were ~160
near-identical lines differing only in the value type, which every error
factory they call is already generic in.

    // before
    template <typename T, typename DispatchT>
    ... DeserializeClassFromElement(reader, const std::wstring&, dispatch);
    template <typename VariantT, typename DispatchT>
    ... DeserializeUnionFromElement(reader, const std::wstring&, dispatch);

    // after
    template <typename ValueT, typename DispatchT>
    ... DeserializeFromElement(reader, const wchar_t*, dispatch);

**An element with a sole model type does not dispatch.** 40 of the 51
``<Cls>FromElement`` spelled out a one-case ``switch`` inside a fifteen-line
lambda signature; the other 11 keep theirs.

    return DeserializeSoleFromElement<
      std::shared_ptr<types::IExtension>
    >(
      reader,
      L"IExtension",
      types::ModelType::kExtension,
      ExtensionFromSequence<types::IExtension>
    );

**The public ``<Cls>From``** were 51 copies of the same 54 lines: open the
reader, read the root element, check that nothing but whitespace follows.

    return DeserializeFrom<
      std::shared_ptr<types::IExtension>
    >(
      is,
      options,
      ExtensionFromElement
    );

**``interface_name`` is a ``const wchar_t*``**, not a ``const std::wstring&``:
every call site used to construct a ``std::wstring`` on *every* element read,
success path included, for a message which is almost never built.

``xmlization.cpp`` for the AAS meta-model goes 34,543 -> 24,634 lines and
its read region 23,584 -> 13,675 (-42%).
mristin added a commit that referenced this pull request Sep 22, 2026
C++ is the last backend to have its XML read side decomposed, after C#
(#686), Go (#688), Java (#690), Python (#694) and TypeScript (#698). #673
factored out the framing of an *element*; this factors out the framing of
the property *loop*, which every ``<Cls>FromSequence`` re-emitted in full.

**A property** loses its duplicate check and its ``std::tie``, and names
only itself and what reads it.

    // before
    case properties::OfExtension::kSemanticId: {
      if (the_semantic_id.has_value()) {
        error = DuplicatePropertyError(name);
        break;
      }

      std::tie(
        the_semantic_id,
        error
      ) = ReferenceFromSequence<
        types::IReference
      >(reader);
      break;
    }

    // after
    case properties::OfExtension::kSemanticId:
      return ReadInto(
        the_semantic_id,
        ReferenceFromSequence<
          types::IReference
        >(reader)
      );

**The loop** around it becomes one call. It could be lifted out at all only
because the framing now returns the error alone, instead of
``NoInstanceAndDeserializationError<std::shared_ptr<T> >``, and so no longer
needs ``T``.

    // before, ~220 lines in each of the 38 ``<Cls>FromSequence``
    while (true) {
      error = SkipWhitespace(reader);
      if (error.has_value()) { ...6 lines... }
      if (reader.node().kind() == xml_common::NodeKind::Stop) {
        break;
      } else if (reader.node().kind() != xml_common::NodeKind::Start) {
        ...10 lines naming IExtension...
      }
      const std::string name(...);
      reader.Read();
      auto it = properties::kMapOfExtension.find(name);
      if (it == properties::kMapOfExtension.end()) { ...12 lines... }
      switch (it->second) { ... }
      ...~60 lines of prepending and stop-element handling...
    }

    // after
    common::optional<DeserializationError> error(
      ReadProperties<
        properties::kPropertyCountOfExtension
      >(
        reader,
        properties::kMapOfExtension,
        L"IExtension",
        [&](
          properties::OfExtension property
        ) -> common::optional<DeserializationError> {
          switch (property) { ... }
        }
      )
    );

The ``switch`` stays inline because its branches assign the locals the
constructor is called with. The lambda is a template argument taken by
``const&``, so the call is statically bound and nothing is type-erased or
allocated; ``g++ -O2`` keeps ``ReadProperties`` itself out of line, one
instantiation per class, which is as many function bodies as the loops it
replaces.

**The duplicate property element check** is bookkeeping of the loop,
not of the property, so it moved there, as a bit set sized to
the class. It still precedes the read, so a duplicate is refused without
its content ever being looked at.

    std::bitset<kPropertyCount> seen;
    ...
    if (seen[index]) {
      error = DuplicatePropertyError(name);
      PrependElementSegmentToDeserializationError(name, *error);
      return error;
    }

    seen[index] = true;

**``ReadInto`` takes the results of a read, not the reader.** A tuple's item
readers vary in number *and* in type, so no reader-taking signature serves
them all, while one taking the result serves every read.

    template <typename T>
    common::optional<DeserializationError> ReadInto(
      common::optional<T>& target,
      std::pair<
        common::optional<T>,
        common::optional<DeserializationError>
      >&& read
    );

**A named union is framed like a class.** The two framings were ~160
near-identical lines differing only in the value type, which every error
factory they call is already generic in.

    // before
    template <typename T, typename DispatchT>
    ... DeserializeClassFromElement(reader, const std::wstring&, dispatch);
    template <typename VariantT, typename DispatchT>
    ... DeserializeUnionFromElement(reader, const std::wstring&, dispatch);

    // after
    template <typename ValueT, typename DispatchT>
    ... DeserializeFromElement(reader, const wchar_t*, dispatch);

**An element with a sole model type does not dispatch.** 40 of the 51
``<Cls>FromElement`` spelled out a one-case ``switch`` inside a fifteen-line
lambda signature; the other 11 keep theirs.

    return DeserializeSoleFromElement<
      std::shared_ptr<types::IExtension>
    >(
      reader,
      L"IExtension",
      types::ModelType::kExtension,
      ExtensionFromSequence<types::IExtension>
    );

**The public ``<Cls>From``** were 51 copies of the same 54 lines: open the
reader, read the root element, check that nothing but whitespace follows.

    return DeserializeFrom<
      std::shared_ptr<types::IExtension>
    >(
      is,
      options,
      ExtensionFromElement
    );

**``interface_name`` is a ``const wchar_t*``**, not a ``const std::wstring&``:
every call site used to construct a ``std::wstring`` on *every* element read,
success path included, for a message which is almost never built.

``xmlization.cpp`` for the AAS meta-model goes 34,543 -> 24,634 lines and
its read region 23,584 -> 13,675 (-42%).
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants