Compose the C# XML de-serialization out of readers - #686
Merged
mristin merged 1 commit intoSep 13, 2026
Merged
Conversation
The generated ``xmlization.cs`` repeated itself at three levels. Every
property of every concrete class inlined its whole reading procedure into
its own ``case`` block -- the self-closing check, the end-of-file check,
the ``try``/``catch`` around the conversion, an error message naming that
property and its class, the ``PrependSegment`` marking the path, and, for
a list or a tuple, the loop or the positional ``<v>`` elements on top.
Every concrete class then got its own ``...FromElement`` function, and all
of those were the same 43 lines with four tokens substituted in. And every
``...FromSequence`` re-emitted the ~70 lines of framing around its property
loop, which said nothing about the class it belonged to.
Everything is now read through one composable shape, and the pieces that
do not vary are generated once:
case "category":
theCategory = ReadString(
reader, isEmptyProperty, out error);
break;
The rationale and the design:
* **One shape for every reader.** Reading a property's content, a list
item's and a tuple item's are the same operation at different nesting
depths, so they are one delegate:
private delegate T ContentReader<T>(
Xml.XmlReader reader, bool isEmpty, out Reporting.Error? error);
Because the shape is uniform a reader can be an argument to another
reader, so a list of tuples -- or anything deeper the meta-model may
grow -- falls out of the existing pieces instead of needing a generated
helper per combination.
* **A plain ``T`` return, not ``T?`` and not ``out T``.** Failure travels
in the error alone and the value is then ``default!``, which the caller
never reads. ``T?`` would have split the machinery in two, since
a nullable return is ``Nullable<T>`` for a value type but a nullable
reference otherwise and no unconstrained parameter covers both. ``out T``
would have compiled, but ``out`` is invariant, so a reader of a concrete
class could no longer serve as the item reader of an interface-typed
list; a return type is covariant in a method-group conversion.
* **Combinators return delegates**, because C# has no partial application:
``AsText``, ``AsEnum``, ``AsList``, ``AsElement`` and ``AsTuple1..N`` each
take what varies and return a ``ContentReader``.
* **``AtElement`` is the hinge, and the element name is data.** It binds
a name to a ``ContentReader`` and yields an ``ElementReader``, which is
what a list and a tuple take for their items -- and also what a class is.
A name cannot be a type argument (``v`` in a list, ``v1``, ``v2``, ... by
position in a tuple, its own XML name for a class), so it is passed. With
the messages phrased in terms of that name, the combinator needs nothing
further to say what it expected, which is what lets one combinator serve
both a class and a ``<v>`` element.
* **A class's element reader is therefore a binding, not a function.**
A class's ``...FromSequence`` already is a ``ContentReader``, so its
element reader is ``AtElement(ExtensionFromSequence, "extension")`` and
the 43-line function per class is gone. Call sites are unchanged: a field
is invoked exactly like the method was.
* **The composed readers live in ``static readonly`` fields.** This is not
cosmetic. Composing at the call site allocates a closure on every read --
measured at 152 bytes against zero for the code being replaced, a real
regression for a de-serialization library. Composing once costs nothing
per read and de-duplicates: 276 properties collapse onto 40 readers.
Two consequences worth knowing. Field initialisers run in declaration
order, so a reader is emitted after everything it composes. And
``ElementReader`` had to become covariant and ``internal``: as a method
group a class's reader widened to its interface for free, but a field is
a value, and a field may not be less accessible than its type.
* **Looking ahead an element's name is one helper**, ``PeekElementName``,
shared by ``AtElement`` and by the thirteen readers that dispatch on
a discriminator element -- they only ever needed the name, and consuming
nothing is what lets the dispatch hand the whole element on.
* **The property loop keeps only its ``switch``.** ``TryNextProperty``
reads the next property's start tag or reports that the sequence ended,
and ``ConsumeEndElement`` -- shared with ``AtElement``, which concludes
identically -- consumes the end tag. Only the ``switch`` stays inline,
because its branches assign the locals the constructor is called with.
* **The property is marked on the error path once per class, after the
``switch``**, rather than in every case: for a matched case the element
name already is the property's XML name.
* **The type in a message comes from ``typeof(T).Name``**, so the shared
skeletons do not grow with the number of types in the meta-model.
* **Only what a model reaches is emitted**, gated along the call graph
rather than by property kind -- a list of enumerations needs the
enumeration combinator even when no property is one. Without that,
a small meta-model would pay for helpers it never calls.
For the AAS meta-model, ``xmlization.cs`` goes from 22,572 to 12,471 lines
(-44.8%), the per-property blocks from 19.7 to 3.7 lines on average, and
the read side loses 38 generated functions along with four helpers.
Coverage Report for CI Build 34729308442Coverage increased (+0.007%) to 85.249%Details
Uncovered Changes
Coverage Regressions2 previously-covered lines in 1 file lost coverage.
Coverage Stats
💛 - Coveralls |
mristin
deleted the
mristin/Decompose-csharp-xmlization-property-reading
branch
September 13, 2026 06:08
mristin
added a commit
that referenced
this pull request
Sep 13, 2026
This is the write-side dual of #686, and finishes the decomposition of ``csharp/lib/_generate_xmlization.py``: every property is now written by one call whose content writer has been composed once into a field, instead of by a lambda built at the call site. Measured on the AAS meta-model fixture, the writer region of ``xmlization.cs`` goes 3,908 -> 2,464 lines (-37%), and the file as a whole 12,471 -> 11,023. Per property 8.1 -> ~5 lines: if (that.SemanticId != null) { WriteElement( "semanticId", that.SemanticId, writer, WriteIReference); } The line count is the lesser win. Every content serializer used to be a capturing lambda composed at the call site -- 212 ``this.`` references in that region -- which allocates a closure on every single write; the read side measured that at 152 bytes/op and rejected it. All 276 call sites now read a ``static readonly`` field instead, so the write path allocates nothing regardless of the consumer's C# version. One delegate, ``ContentWriter<in T>``, replaces ``ElementContentSerializer<T>``. ``in T`` is the dual of the read side's ``out T``: it is what lets the single static ``WriteIClass(Aas.IClass, ...)`` serve as the item writer of a list or a tuple of any interface. Deliberately *one* delegate, not a content/element pair: on the write side both would have identical signatures, since the content/element split is a read-side concern (``isEmpty``). ``WriteElement`` and ``WrapInElement`` are one body with two entry points. The applied form is what the 276 property sites call; ``WrapInElement`` is the partial application C# does not give for free, and is used *only* for the ``<v>`` element of a list item and the positional ``v1``, ``v2``, ... of a tuple item. It carries a doc comment saying so, because its payoff is invisible from its call count: it is what lets ``WriteList`` and ``WriteTupleN`` take a *single* item-writer type even though a class item writes its own run-time-named element while a primitive item needs a fixed positional one. The ``tuples`` fixture's ``Tuple[int, Some_item, Abstract_item, Some_item, Positive_int, Result]`` mixes the two item by item; with two item-writer types it would have needed a combinator per combination. Names are kept distinct from the read side's on purpose: ``WrapInElement`` / ``WriteEnum`` / ``WriteList`` / ``WriteTupleN`` against ``AtElement`` / ``AsEnum`` / ``AsList`` / ``AsTupleN``. There is no dual of ``AsText`` (``WriteValue`` cannot fail) and none of ``AsElement``. The visitor stays, reduced to a dispatch shim: a single ``_instance`` plus static ``WriteIClass``, which one virtual call resolves for every abstract class, every concrete class with descendants and -- through ``WriteIUnion`` -- every named union. That is why the write side needs neither the read side's 13 per-interface dispatchers nor a combinator to invoke one. Along the way, two things shared with the read side: * The gating pass is now recursive and serves both emitters. ``_needed_content_readers`` walked only one level, so a model whose only ``str`` sat inside ``List[List[str]]`` would have emitted ``ReadString = AsText<string>(ReadContentAsString, "")`` while generating neither ``AsText`` nor ``ReadContentAsString`` -- output that does not compile. Latent only because the per-property gate blocks such models today. * ``_content_types_in_initialization_order`` replaces the two copies of the registration recursion, so the readers and the writers are de-duplicated and ordered by one walk. Nesting composes for free in this design -- ``WriteListOfListOfString = WriteList<List<string>>(WrapInElement(WriteListOfString, "v"))`` falls out with no new combinator -- but stays gated, as on the read side, since ``copying``, ``enhancing``, ``jsonization`` and the test generators bail on it in C# and in all five other backends. The write side now reports that as a proper ``Error`` instead of raising ``NotImplementedError``. The snippet contract for an implementation-specific class changes: ``Xmlization/VisitorWithWriter/{cls}_to_sequence.cs`` must now be ``static``, and ``visit_{cls}.cs`` must call it without ``this.`` -- the same change the read side already imposed on ``..._from_sequence.cs``.
mristin
added a commit
that referenced
this pull request
Sep 13, 2026
This is the write-side dual of #686, and finishes the decomposition of ``csharp/lib/_generate_xmlization.py``: every property is now written by one call whose content writer has been composed once into a field, instead of by a lambda built at the call site. Measured on the AAS meta-model fixture, the writer region of ``xmlization.cs`` goes 3,908 -> 2,464 lines (-37%), and the file as a whole 12,471 -> 11,023. Per property 8.1 -> ~5 lines: if (that.SemanticId != null) { WriteElement( "semanticId", that.SemanticId, writer, WriteIReference); } The line count is the lesser win. Every content serializer used to be a capturing lambda composed at the call site -- 212 ``this.`` references in that region -- which allocates a closure on every single write; the read side measured that at 152 bytes/op and rejected it. All 276 call sites now read a ``static readonly`` field instead, so the write path allocates nothing regardless of the consumer's C# version. One delegate, ``ContentWriter<in T>``, replaces ``ElementContentSerializer<T>``. ``in T`` is the dual of the read side's ``out T``: it is what lets the single static ``WriteIClass(Aas.IClass, ...)`` serve as the item writer of a list or a tuple of any interface. Deliberately *one* delegate, not a content/element pair: on the write side both would have identical signatures, since the content/element split is a read-side concern (``isEmpty``). ``WriteElement`` and ``WrapInElement`` are one body with two entry points. The applied form is what the 276 property sites call; ``WrapInElement`` is the partial application C# does not give for free, and is used *only* for the ``<v>`` element of a list item and the positional ``v1``, ``v2``, ... of a tuple item. It carries a doc comment saying so, because its payoff is invisible from its call count: it is what lets ``WriteList`` and ``WriteTupleN`` take a *single* item-writer type even though a class item writes its own run-time-named element while a primitive item needs a fixed positional one. The ``tuples`` fixture's ``Tuple[int, Some_item, Abstract_item, Some_item, Positive_int, Result]`` mixes the two item by item; with two item-writer types it would have needed a combinator per combination. Names are kept distinct from the read side's on purpose: ``WrapInElement`` / ``WriteEnum`` / ``WriteList`` / ``WriteTupleN`` against ``AtElement`` / ``AsEnum`` / ``AsList`` / ``AsTupleN``. There is no dual of ``AsText`` (``WriteValue`` cannot fail) and none of ``AsElement``. The visitor stays, reduced to a dispatch shim: a single ``_instance`` plus static ``WriteIClass``, which one virtual call resolves for every abstract class, every concrete class with descendants and -- through ``WriteIUnion`` -- every named union. That is why the write side needs neither the read side's 13 per-interface dispatchers nor a combinator to invoke one. Along the way, two things shared with the read side: * The gating pass is now recursive and serves both emitters. ``_needed_content_readers`` walked only one level, so a model whose only ``str`` sat inside ``List[List[str]]`` would have emitted ``ReadString = AsText<string>(ReadContentAsString, "")`` while generating neither ``AsText`` nor ``ReadContentAsString`` -- output that does not compile. Latent only because the per-property gate blocks such models today. * ``_content_types_in_initialization_order`` replaces the two copies of the registration recursion, so the readers and the writers are de-duplicated and ordered by one walk. Nesting composes for free in this design -- ``WriteListOfListOfString = WriteList<List<string>>(WrapInElement(WriteListOfString, "v"))`` falls out with no new combinator -- but stays gated, as on the read side, since ``copying``, ``enhancing``, ``jsonization`` and the test generators bail on it in C# and in all five other backends. The write side now reports that as a proper ``Error`` instead of raising ``NotImplementedError``. The snippet contract for an implementation-specific class changes: ``Xmlization/VisitorWithWriter/{cls}_to_sequence.cs`` must now be ``static``, and ``visit_{cls}.cs`` must call it without ``this.`` -- the same change the read side already imposed on ``..._from_sequence.cs``.
mristin
added a commit
that referenced
this pull request
Sep 13, 2026
This is the Go counterpart of #686, but the C# design does not carry over: the shape that works there fails to compile in Go, so the read side is composed out of generated dispatchers instead of composed readers. Measured on the AAS meta-model fixture, the read region of ``xmlization.go`` goes 11,738 -> 5,584 lines (-52%), and the file as a whole 22,846 -> 16,692. The ``read<X>WithLookahead`` band -- 50 functions, 3,057 lines -- goes to **zero**; 20 ``read<X>Dispatched`` functions stand in its place. The whole design turns on one asymmetry in the wire format. An instance element is self-describing: its local name *is* its type. A scalar element is not: its name is its *position*, ``v`` in a list and ``v1``, ``v2``, ... in a tuple. So: every instance is read by dispatch on its own name, every scalar by a name its container supplies, and nothing else in the read path knows the difference. ``readElementDispatched`` is the only place which frames an XML element around a value. ``readListOf`` and every ``readTupleN`` delegate the framing to it, which is what deletes ``readListOfScalars``, ``readScalarWithName``, ``asScalarTupleItemReader`` and ``asInstanceTupleItemReader`` outright: one list helper, one item shape, no adapters. It also returns the token *past* the end element, so the "``read<X>WithLookahead`` stops at the end element, so we look ahead" comment and its hand-written ``readNext`` vanish from every call site: readListOf(decoder, current, readReferenceDispatched) readListOf(decoder, current, readSubmodelElementDispatched) readListOf(decoder, current, readLongAtV) readTuple6( decoder, current, readLongAtV1, readSomeItemDispatched, readAbstractItemDispatched, readSomeItemDispatched, readLongAtV5, readResultAtV6, ) Per class, ``read<X>AsSequence`` has for every property case one assignment of the same ``(value, current, valueErr)`` triple: case "supplementalSemanticIds": theSupplementalSemanticIDs, current, valueErr = readListOf( decoder, current, readReferenceDispatched, ) case "valueType": theValueType, current, valueErr = readOptional( readTextAsDataTypeDefXSD(decoder, current), ) ``readOptional`` absorbs the ``var value T; ...; theX = &value`` dance. It takes the *results* of a read rather than the reader, so Go's multi-value pass-through lets one helper serve any read however many arguments that read takes -- including ``readOptional(readTuple2(...))``, which a reader-taking signature cannot express, since a tuple's item readers vary in number and in type. That also fixes an optional tuple property, which is a pointer (a Go struct is not nilable) and used to be assigned a non-pointer ``readTupleN`` result: output that does not compile.
mristin
added a commit
that referenced
this pull request
Sep 13, 2026
Write-side dual of #688, as #687 was of #686 in C#. Measured on the AAS meta-model fixture, the serialization region of ``xmlization.go`` goes 11,080 -> 4,727 lines (-57%), and the file as a whole 16,692 -> 10,330. ``writeExtensionAsSequence`` and its six properties go 180 -> 69 lines. Everything is written through one shape, ``func(encoder, value T) error``. A stringification, a class's own ``write<X>AsSequence`` and the generated item, list and tuple writers all have it, so ``writeElement`` -- which frames an XML element around a value written that way -- is the only writer a property needs, whatever kind of value it holds. What is left per property is that one call and a decision on presence: err = finishProperty( "SemanticID()", writeOptionalInstance( encoder, "semanticId", that.SemanticID(), writeReferenceAsSequence, ), ) if err != nil { return } ``finishProperty`` takes the *result* of the write rather than the write itself -- the same trick as the read side's ``readOptional``, since Go evaluates the argument first -- so one function concludes every property, prepending the getter to the path of the error. Mind that the getter is spelled out as it is in Go, ``SemanticID()`` and not ``semanticId``, since a serialization error renders its path through ``aasreporting.ToGolangPath``, while the de-serialization prepends the XML name, its errors being XPaths. Go has no ``nameof``. The ``writeOptional*`` family has one member per way Go spells an optional, each named after that spelling, as the check for the presence is the only thing which differs between them: a pointer to the value (``writeOptionalPointer``: a scalar, an enumeration, a tuple), a value which is nil on its own (``writeOptionalInstance``: an instance, or a named union, which is a pointer to a struct) and a slice (``writeOptionalSlice``: a list, the bytes). They can not be collapsed into one, since Go decides nil-ness by the representation: a type parameter can not be compared against ``nil``, ``any(that) == nil`` is false for a nil *pointer* boxed in an ``any`` -- hence the comparison against the zero value in ``writeOptionalInstance``, which in turn panics on a slice, which is not comparable. Each of them writes nothing at all when the value is absent, which is what keeps the absence out of the generated code. No closure is allocated on the write path any more: every writer passed on is a top-level function or an instantiated generic, hence a static funcval, exactly as on the read side. A list and a tuple get one generated content writer as well, per item type resp. per tuple type -- ``writeListOfIReference`` and ``writeTupleOfStringLong`` and their kin -- since ``writeList`` and ``writeTupleN`` take a writer per item, and Go gives no partial application to bind those in without allocating. Finally, the flush. ``write<X>AsSequence`` used to flush the encoder after every property, and ``write<X>`` again at its end element, which defeated the buffering of ``xml.Encoder`` on every write; only the flush after the last token is needed, as ``EncodeToken`` deliberately leaves it to the caller. ``Marshal`` is therefore split into ``writeClass``, which dispatches on the model type, and ``Marshal`` itself, which calls it and flushes exactly once; ``writeInstance`` and ``writeUnion`` go to ``writeClass``, so a nested instance does not flush either. Serializing an instance holding 1,000 instances (42 KB of XML) from the ``list_of_classes`` fixture, the writes which actually reach the underlying writer go 2,003 -> 11, one per 4 KB of buffer instead of one per property and per instance; to an ``os.File`` that is 5.29 -> 0.76 ms per serialization, the median of five.
mristin
added a commit
that referenced
this pull request
Sep 14, 2026
Java is the third backend to have its XML read side decomposed, after C# (#686) and Go (#688), and it carried the repetition #686 described at four levels rather than three. Every property inlined its whole reading procedure into its own ``case`` block. Everything is now read through one composable shape, and the pieces which do not vary are generated once: case "contentType": { final Reporting.Result<String> value = readTextAs_string(reader, isEmptyProperty); if (value.isError()) { valueError = value.getError(); } else { theContentType = value.getResult(); } break; } The rationale and the design: * A method reference to a ``static`` method captures nothing, so its ``invokedynamic`` call site is linked once and hands back the same instance afterwards; passing ``_DeserializeImplementation::readReference`` costs nothing. A capturing lambda is not free, and the old code had one in every ``try<X>FromElement`` and in every dispatcher, both closing over the enclosing ``reader``: an allocation on every single instance read. The read side now contains no lambdas at all -- every reader passed anywhere is a method reference to a static method -- so nothing can accidentally capture. * Two shapes, and two only. ``ContentReader<T>`` reads the content of an element which has already been opened, ``ElementReader<T>`` reads a whole element. They are duals, and the wire format forces both, because it is asymmetric: an instance element is self-describing -- its local name *is* its type -- whereas a scalar element is not, its name being its position (the property's name, ``v`` in a list, ``v1``, ``v2``, ... in a tuple). * One primitive, ``peekElementName``. It consumes nothing, and it is the single answer to "we are at an element, and this is its name". The framer ``readNamedElement`` checks that name against the one its container supplied, the dispatchers switch on it, and every property loop selects the property with it. That is also why no third shape is needed to carry the name into a dispatcher: a dispatcher peeks, switches, and hands the *whole* element on to that type's own ``read<Cls>FromElement``, which re-checks the name and consumes it. * The property loop keeps only its ``switch``, because its branches assign the locals the constructor is called with. ``atEndOfSequence`` is the loop condition, ``consumeEndElement`` -- shared with ``readNamedElement``, which concludes identically -- consumes the end tag, and ``unexpectedProperty`` and ``missingRequiredProperty`` are the two errors which used to be spelled out per class and per property. * The property is marked on the error path once per class, after the ``switch``, rather than in every case: for a matched case the element name already is the property's XML name. That also retires the per-case ``castTo(<Cls>.class)`` re-threading, and fixes an inconsistency along the way -- a primitive and an enumeration used to be marked with the *Java* property name while a list and a class used the XML name. * The names are unique by construction, so no uniqueness check is needed. A reader is keyed either on one symbol or on a type moniker, and the two name spaces are kept apart by a single rule: a moniker-keyed name contains ``_``, a symbol-keyed name never does. That holds because ``naming.capitalized_camel_case``, which every one of our type names goes through, emits letters and digits only, with an upper-case initial. The monikers are then a Polish notation over ``_``-separated tokens -- ``ListOf_string``, ``TupleOf2_string_long``, ``ListOf_TupleOf2_string_long`` -- with the arity spelled out, so the encoding is injective. The primitives are lower-cased for the same reason: an enumeration which somebody names ``String`` can then never collide with the primitive ``str``. * ``readNestedElement`` looks redundant and is not. A polymorphic property and a list item leave the reader at the very same position, just before a self-describing start element, so it is tempting to give ``read<X>FromElement`` an ``isEmpty`` parameter and call it from the property ``case`` directly. The error path is what forbids it. A property wraps its instance in an element of its own, so the discriminator's name has to be prepended and the path reads ``value/property/idShort``; a list item is not wrapped -- the item element *is* the indexed child -- so the same prepend would give ``annotations/*[0]/property/idShort``, which walks one level past the element ``*[0]`` already selects and resolves to nothing. The reasoning is written into the generated doc comment as well, because it is exactly the kind of layer the next reader deletes. * Only what a model reaches is emitted, gated along the call graph rather than by property kind -- a list of enumerations needs the enumeration reader even when no property is one, and an enumeration needs the text path because it is built on it. The pass recurses into the items of a list and of a tuple, without which a model whose only ``str`` sits in a ``List[List[str]]`` would emit a reader built from helpers which were never generated. The snippet contract for an implementation-specific class changes: ``Xmlization/DeserializeImplementation/{cls}_from_element.java`` is no longer read, and the remaining ``{cls}_from_sequence.java`` has to define ``read<Cls>FromSequence`` rather than ``try<Cls>FromSequence``. For the AAS meta-model, ``Xmlization.java`` goes from 16,647 to 11,455 lines (-31.2%) and its read region from 12,187 to 6,995 (-42.6%). Finally, the speed, since the read side is where a de-serialization library is judged: there is no measurable difference, and that is the expected answer. De-serializing a 1.4 MB AAS environment (200 submodels, 131,806 StAX events; 400 runs after 200 warm-up runs, three measurements of each variant, run alternately), the medians are 32.2/33.2/33.2 ms before and 32.1/32.7/33.2 ms after -- the same, within the noise of repeated runs of either one. Draining that very same document through ``XMLEventReader`` without building anything at all already costs 27.4 ms and allocates 41.2 MB, so StAX accounts for ~83% of the time and ~84% of the allocation, and everything the generated code does is the remaining ~6 ms and ~8.2 MB. Against that, the lambdas which are gone are worth ~0.15 MB per run at the median (49.48-49.57 MB before, 49.14-49.44 MB after): consistently in that direction, but under 1% of the total. The decomposition is a maintainability change, not a performance one, and it was worth measuring rather than assuming, since the C# commit's 152 bytes/op is what a reader would expect to carry over.
mristin
added a commit
that referenced
this pull request
Sep 15, 2026
Write-side dual of #690, as #687 was of #686 in C# and #689 of #688 in Go. ``visit{Cls}`` was a verbatim copy of ``serializeElement``: 38 methods of 18 lines each spelling out the start tag, the namespace, the content and the end tag, around a call which already did exactly that. Six near-identical emitters -- one per property kind -- produced the very same ``serializeElement`` call, differing only in the content serializer. And everything was an *instance* method, because of one mutable ``topLevel`` flag. That flag is what cost the most. It made ``this::writeStringifiedContent``, ``this::visit`` and ``(value, w) -> this.xToSequence(value, w)`` capturing, and ``serializeItems(...)`` and ``asNamedElementSerializer(...)`` returned capturing lambdas on top of them, so writing one instance allocated a small pile of closures. It is replaced by a ``final boolean withNamespace`` pinned in the constructor of two immutable visitors, ``ROOT`` and ``NESTED``: only the outermost element declares the namespace, and an element written from ``writeClass`` is nested by construction. Everything else is then ``static``, and a method reference to a static method captures nothing. The design is otherwise the read side's, mirrored: * One shape, not two. ``ContentWriter<T>`` writes a value where the writer already is; ``writeElement`` frames it in a start and an end tag. There is deliberately no second delegate for a whole element, as there is on the reading side -- an element differs from a content only in what it writes, never in its shape. The reading needs the distinction because a content reader has to be told whether its element was self-closing, and a writer has nothing to be told. Use sites take ``ContentWriter<? super T>``, which is how Java spells the contravariance C# writes as ``in T``, so the single writer of an ``IClass`` serves wherever the writer of a more specific interface is expected. * ``visit{Cls}`` collapses to one call, and the visitor becomes a dispatch shim:: @OverRide public void visitQualifier( IQualifier that, XMLStreamWriter writer) { writeElement( "qualifier", that, writer, withNamespace, _VisitorWithWriter::writeQualifierAsSequence); } * The families are dual to the readers and named by the same moniker grammar: ``write{Cls}AsSequence``, ``writeClass``/``writeUnion``, ``writeListOf_{M}``, ``writeTupleOf{N}_{M...}``, ``writeAtV_{M}`` and ``writeAtV{i}_{M}``. The list and the tuple writers inline their loop, which retires ``serializeItems``, ``asNamedElementSerializer`` and ``serializeTuple{N}`` -- Java needs no partial application here, since the name to bind is fixed per generated function. * One property emitter replaces the six. Every property is written by ``writeProperty``, or by ``writeOptionalProperty``, which takes the ``Optional`` apart once instead of asking it and then unwrapping it -- the old form called the getter twice, and every call allocates an ``Optional``. * ``_type_moniker``, ``_leaf_moniker``, ``_is_instance_type``, ``_is_dispatched`` and ``_collect_needed`` are reused unchanged: the reading and the writing are composed out of the same shapes, so one gating pass answers for both. Only the two dispatching writers needed a walk of their own, because the *position* matters on the write side while it does not on the read side: an item of any class is written as its own element, whereas a property of a concrete class without descendants writes its properties straight into the property element. Two things are genuinely asymmetric, and both save work: * There is no ``writeTextAs_{primitive}`` family. Once the value is at hand nothing type-specific is left to do with it, so the four ``toString`` primitives share ``writeStringifiedContent`` and the byte array ``writeByteArrayContent``. An enumeration does have something of its own to write and keeps one function each -- named ``writeTextAs_{Enum}`` and not ``write{Enum}Content``, so that it stays in the moniker-keyed name space: were it keyed on the symbol, an enumeration which somebody names ``Stringified`` or ``ByteArray`` would silently take a shared writer's name. * There is no element writer per class. Which element an instance writes follows from its run-time type, and one virtual call answers that for every class at once. The reading needs a dispatcher per interface because it has to decide what to construct before it has read anything. Two defects of the write side are fixed along the way, both of which the decomposition makes one-line changes rather than sweeping ones. ``Serialize.to`` never flushed, so the serialization was not finished when it returned, and a failure of the underlying stream surfaced at the caller's own flush as a raw ``XMLStreamException`` -- outside the one method whose ``throws`` clause promises a ``SerializeException``. Probed with a ``Writer`` which accepts every character and fails on ``flush``, ``to`` used to return normally and now throws. The flush goes in ``Serialize.to`` and nowhere else, as in Go's #689, so ``XMLStreamWriter`` keeps buffering in between. Mind that this does push the underlying stream: a caller serializing many small documents into one buffered writer now flushes once per document rather than once per buffer. That is the right trade, since the alternative leaves the output silently incomplete whenever the caller forgets. Every ``SerializeException`` carried an empty path, so every failure read ``... at: the beginning`` no matter where it happened. The path is now built as the stack unwinds, each container prepending the one segment it knows, exactly as on the read side where ``read<Cls>FromSequence`` prepends the property after its ``switch`` and ``readList`` prepends the index. Probed with an ``XMLStreamWriter`` which refuses to write one chosen value: a property of the root instance idShort nested through two lists submodels/*[1]/submodelElements/*[1]/value through a class-typed property semanticId/keys/*[0]/value ``SerializeException`` renders its message in its constructor, so its path has to be complete by the time it is built and it can not be the thing thrown from inside; a private ``_SerializeFailure`` carries a ``Reporting.Error`` instead, and ``Serialize.to`` renders that with ``Reporting.generateRelativeXPath`` and converts. The public exception is untouched. ``writeElement`` contributes no segment and ``writeProperty`` is ``writeElement`` plus the one thing a property knows which nothing below it does -- its own name -- so there is no ``try`` around any of the 276 property writes. A list and a tuple put their ``try`` outside the loop and advance the index only once an item has been written, so the item named is the one which failed and nothing is paid per item. Two segments are deliberately absent, and the reasoning is written into the generated Javadoc, since both look like oversights. The outermost element is not named, because ``Serialize.to`` takes any ``IClass`` and the name would say nothing the caller does not already know; the read side does prepend it, having one facade per class. Neither is the discriminator element of a polymorphic property, which the read side does prepend -- for a failure in an embedded data specification it reports ``submodel/embeddedDataSpecifications/*[0]/dataSpecificationContent/`` ``dataSpecificationIec61360/preferredName/*[0]`` where the write side reports ``embeddedDataSpecifications/*[0]/dataSpecificationContent/`` ``preferredName/*[0]/text``. The reading points into a document it is parsing, where that element is a real extra level; the writing points into the instance the caller handed over, where it is not, and ``value/idShort`` there is exactly ``getValue().getIdShort()``. Go's #689 made the same choice when it rendered its serialization paths through ``aasreporting.ToGolangPath`` rather than as XPaths. For the AAS meta-model the write region of ``Xmlization.java`` goes 3,281 -> 2,873 lines (-12.4%), error paths included. The write region no longer contains a single lambda. Unlike the read side, that shows up in the measurement. Serializing an ``Environment`` of 200 submodels (1.14 MB of XML) into a writer which counts the characters and discards them, 100 runs after 100 warm-up runs, eight rounds of each variant run alternately: the allocation goes 0.76 -> 0.36 MB per serialization, and the time 7.69 -> 7.49 ms, which is inside the round-to-round noise. ``XMLStreamWriter`` writes straight through, so the whole write path allocates ~0.8 MB against the ~49 MB the *reading* of the same document allocates -- there is no StAX floor here for the closures to hide behind, which is precisely why #690 could not measure its own. The snippet contract for an implementation-specific class changes: ``Xmlization/VisitorWithWriter/visit_{cls}.java`` is no longer read, as the element around the sequence is now written identically for every class, and the remaining ``{cls}_to_sequence.java`` has to define a ``static`` ``write{Cls}AsSequence``.
mristin
added a commit
that referenced
this pull request
Sep 15, 2026
Write-side dual of #690, as #687 was of #686 in C# and #689 of #688 in Go. ``visit{Cls}`` was a verbatim copy of ``serializeElement``: 38 methods of 18 lines each spelling out the start tag, the namespace, the content and the end tag, around a call which already did exactly that. Six near-identical emitters -- one per property kind -- produced the very same ``serializeElement`` call, differing only in the content serializer. And everything was an *instance* method, because of one mutable ``topLevel`` flag. That flag is what cost the most. It made ``this::writeStringifiedContent``, ``this::visit`` and ``(value, w) -> this.xToSequence(value, w)`` capturing, and ``serializeItems(...)`` and ``asNamedElementSerializer(...)`` returned capturing lambdas on top of them, so writing one instance allocated a small pile of closures. It is replaced by a ``final boolean withNamespace`` pinned in the constructor of two immutable visitors, ``ROOT`` and ``NESTED``: only the outermost element declares the namespace, and an element written from ``writeClass`` is nested by construction. Everything else is then ``static``, and a method reference to a static method captures nothing. The design is otherwise the read side's, mirrored: * One shape, not two. ``ContentWriter<T>`` writes a value where the writer already is; ``writeElement`` frames it in a start and an end tag. There is deliberately no second delegate for a whole element, as there is on the reading side -- an element differs from a content only in what it writes, never in its shape. The reading needs the distinction because a content reader has to be told whether its element was self-closing, and a writer has nothing to be told. Use sites take ``ContentWriter<? super T>``, which is how Java spells the contravariance C# writes as ``in T``, so the single writer of an ``IClass`` serves wherever the writer of a more specific interface is expected. * ``visit{Cls}`` collapses to one call, and the visitor becomes a dispatch shim:: @OverRide public void visitQualifier( IQualifier that, XMLStreamWriter writer) { writeElement( "qualifier", that, writer, withNamespace, _VisitorWithWriter::writeQualifierAsSequence); } * The families are dual to the readers and named by the same moniker grammar: ``write{Cls}AsSequence``, ``writeClass``/``writeUnion``, ``writeListOf_{M}``, ``writeTupleOf{N}_{M...}``, ``writeAtV_{M}`` and ``writeAtV{i}_{M}``. The list and the tuple writers inline their loop, which retires ``serializeItems``, ``asNamedElementSerializer`` and ``serializeTuple{N}`` -- Java needs no partial application here, since the name to bind is fixed per generated function. * One property emitter replaces the six. Every property is written by ``writeProperty``, or by ``writeOptionalProperty``, which takes the ``Optional`` apart once instead of asking it and then unwrapping it -- the old form called the getter twice, and every call allocates an ``Optional``. * ``_type_moniker``, ``_leaf_moniker``, ``_is_instance_type``, ``_is_dispatched`` and ``_collect_needed`` are reused unchanged: the reading and the writing are composed out of the same shapes, so one gating pass answers for both. Only the two dispatching writers needed a walk of their own, because the *position* matters on the write side while it does not on the read side: an item of any class is written as its own element, whereas a property of a concrete class without descendants writes its properties straight into the property element. Two things are genuinely asymmetric, and both save work: * There is no ``writeTextAs_{primitive}`` family. Once the value is at hand nothing type-specific is left to do with it, so the four ``toString`` primitives share ``writeStringifiedContent`` and the byte array ``writeByteArrayContent``. An enumeration does have something of its own to write and keeps one function each -- named ``writeTextAs_{Enum}`` and not ``write{Enum}Content``, so that it stays in the moniker-keyed name space: were it keyed on the symbol, an enumeration which somebody names ``Stringified`` or ``ByteArray`` would silently take a shared writer's name. * There is no element writer per class. Which element an instance writes follows from its run-time type, and one virtual call answers that for every class at once. The reading needs a dispatcher per interface because it has to decide what to construct before it has read anything. Two defects of the write side are fixed along the way, both of which the decomposition makes one-line changes rather than sweeping ones. ``Serialize.to`` never flushed, so the serialization was not finished when it returned, and a failure of the underlying stream surfaced at the caller's own flush as a raw ``XMLStreamException`` -- outside the one method whose ``throws`` clause promises a ``SerializeException``. Probed with a ``Writer`` which accepts every character and fails on ``flush``, ``to`` used to return normally and now throws. The flush goes in ``Serialize.to`` and nowhere else, as in Go's #689, so ``XMLStreamWriter`` keeps buffering in between. Mind that this does push the underlying stream: a caller serializing many small documents into one buffered writer now flushes once per document rather than once per buffer. That is the right trade, since the alternative leaves the output silently incomplete whenever the caller forgets. Every ``SerializeException`` carried an empty path, so every failure read ``... at: the beginning`` no matter where it happened. The path is now built as the stack unwinds, each container prepending the one segment it knows, exactly as on the read side where ``read<Cls>FromSequence`` prepends the property after its ``switch`` and ``readList`` prepends the index. Probed with an ``XMLStreamWriter`` which refuses to write one chosen value: a property of the root instance idShort nested through two lists submodels/*[1]/submodelElements/*[1]/value through a class-typed property semanticId/keys/*[0]/value ``SerializeException`` renders its message in its constructor, so its path has to be complete by the time it is built and it can not be the thing thrown from inside; a private ``_SerializeFailure`` carries a ``Reporting.Error`` instead, and ``Serialize.to`` renders that with ``Reporting.generateRelativeXPath`` and converts. The public exception is untouched. ``writeElement`` contributes no segment and ``writeProperty`` is ``writeElement`` plus the one thing a property knows which nothing below it does -- its own name -- so there is no ``try`` around any of the 276 property writes. A list and a tuple put their ``try`` outside the loop and advance the index only once an item has been written, so the item named is the one which failed and nothing is paid per item. Two segments are deliberately absent, and the reasoning is written into the generated Javadoc, since both look like oversights. The outermost element is not named, because ``Serialize.to`` takes any ``IClass`` and the name would say nothing the caller does not already know; the read side does prepend it, having one facade per class. Neither is the discriminator element of a polymorphic property, which the read side does prepend -- for a failure in an embedded data specification it reports ``submodel/embeddedDataSpecifications/*[0]/dataSpecificationContent/`` ``dataSpecificationIec61360/preferredName/*[0]`` where the write side reports ``embeddedDataSpecifications/*[0]/dataSpecificationContent/`` ``preferredName/*[0]/text``. The reading points into a document it is parsing, where that element is a real extra level; the writing points into the instance the caller handed over, where it is not, and ``value/idShort`` there is exactly ``getValue().getIdShort()``. Go's #689 made the same choice when it rendered its serialization paths through ``aasreporting.ToGolangPath`` rather than as XPaths. For the AAS meta-model the write region of ``Xmlization.java`` goes 3,281 -> 2,873 lines (-12.4%), error paths included. The write region no longer contains a single lambda. Unlike the read side, that shows up in the measurement. Serializing an ``Environment`` of 200 submodels (1.14 MB of XML) into a writer which counts the characters and discards them, 100 runs after 100 warm-up runs, eight rounds of each variant run alternately: the allocation goes 0.76 -> 0.36 MB per serialization, and the time 7.69 -> 7.49 ms, which is inside the round-to-round noise. ``XMLStreamWriter`` writes straight through, so the whole write path allocates ~0.8 MB against the ~49 MB the *reading* of the same document allocates -- there is no StAX floor here for the closures to hide behind, which is precisely why #690 could not measure its own. The snippet contract for an implementation-specific class changes: ``Xmlization/VisitorWithWriter/visit_{cls}.java`` is no longer read, as the element around the sequence is now written identically for every class, and the remaining ``{cls}_to_sequence.java`` has to define a ``static`` ``write{Cls}AsSequence``.
mristin
added a commit
that referenced
this pull request
Sep 22, 2026
C++ is the last backend to have its XML read side decomposed, after C# (#686), Go (#688), Java (#690), Python (#694) and TypeScript (#698). #673 already factored out the framing of an *element*, but every ``<Cls>FromSequence`` still re-emitted the same ~220 lines of property-loop framing around its ``switch``, so a property cost nineteen lines where it now costs three: case properties::OfExtension::kSemanticId: return ReadInto( the_semantic_id, ReferenceFromSequence< types::IReference >(reader) ); The rationale and the design: * **The framing does not need ``T``.** Returning the error alone, instead of ``NoInstanceAndDeserializationError<std::shared_ptr<T> >`` on every failure path, is the whole reason the loop can be lifted out of the class: template < std::size_t kPropertyCount, typename EnumT, typename OnPropertyT > common::optional<DeserializationError> ReadProperties( xml_common::ReaderMergingText& reader, const std::unordered_map<std::string, EnumT>& map_of_properties, const wchar_t* interface_name, const OnPropertyT& on_property ); * **The ``switch`` stays inline, and in C++ that is free.** Its branches assign the locals the constructor is called with -- which is why C#, Go and Java had to keep the whole loop inline -- but a lambda passed as ``const F&`` to a function template is a stack object which is inlined away, so the class can hand the ``switch`` over and nothing is allocated: common::optional<DeserializationError> error( ReadProperties< properties::kPropertyCountOfExtension >( reader, properties::kMapOfExtension, L"IExtension", [&]( properties::OfExtension property ) -> common::optional<DeserializationError> { switch (property) { ... } } ) ); * **The duplicate check belongs to the loop, not to the property.** It is now a bit set sized to the class and living on the stack, so a class may have arbitrarily many properties and a read allocates nothing for it; the check still precedes the read, so a duplicate is refused without its content ever being looked at, and a five-line guard leaves all 276 cases: std::bitset<kPropertyCount> seen; ... if (seen[index]) { error = DuplicatePropertyError(name); PrependElementSegmentToDeserializationError(name, *error); return error; } seen[index] = true; * **``ReadInto`` takes the results of a read, not the reader.** The item readers of a tuple vary both in number and in type, so no signature taking the reader could serve every read, while one taking the result serves all of them: template <typename T> common::optional<DeserializationError> ReadInto( common::optional<T>& target, std::pair< common::optional<T>, common::optional<DeserializationError> >&& read ); * **A named union is framed exactly like a class.** ``DeserializeUnionFromElement`` and ``DeserializeClassFromElement`` were ~160 near-identical lines which differed only in the value type, and every error factory they called was already generic in it: template <typename ValueT, typename DispatchT> std::pair< common::optional<ValueT>, common::optional<DeserializationError> > DeserializeFromElement( xml_common::ReaderMergingText& reader, const wchar_t* value_name, const DispatchT& dispatch ); * **An element with a sole model type does not dispatch.** 40 of the 51 ``<Cls>FromElement`` spelled out a one-case ``switch`` inside a fifteen-line lambda signature; the remaining 11 keep theirs, since a dense jump table beats a lookup: std::pair< common::optional< std::shared_ptr<types::IExtension> >, common::optional<DeserializationError> > ExtensionFromElement( xml_common::ReaderMergingText& reader ) { return DeserializeSoleFromElement< std::shared_ptr<types::IExtension> >( reader, L"IExtension", types::ModelType::kExtension, ExtensionFromSequence<types::IExtension> ); } * **The public ``<Cls>From`` are one function.** 51 copies of the same 54 lines -- open the reader, read the root element, check that nothing but whitespace follows it -- became a single call each: return DeserializeFrom< std::shared_ptr<types::IExtension> >( is, options, ExtensionFromElement ); * **``interface_name`` is a ``const wchar_t*``.** As a ``const std::wstring&`` it made every call site construct a ``std::wstring`` on *every* element read, success path included, for a message which is almost never built: const wchar_t* interface_name, // was: const std::wstring& Everything is written for C++11, which the library targets outside of its tests: no generic lambda, no deduced return type, and every lambda spells its parameters and its return type out. For the AAS meta-model, ``xmlization.cpp`` goes from 34,543 to 24,617 lines and its read region from 23,584 to 13,658 (-42%). The ``<Cls>FromSequence`` band goes 13,400 -> 6,076 over 37 classes (-54.7%): ``Key`` 257 -> 95, ``AdministrativeInformation`` 295 -> 112, ``Extension`` 331 -> 141. The public ``<Cls>From`` go 2,751 -> 813 (-70%). The ``unions`` fixture goes 7,950 -> 5,795. ``empty_class`` goes 1,622 -> 1,674, since a meta-model with a single property-less class pays for the shared loop without having anything to save. Two warts are cleaned up along the way: a class without properties used to get an enumeration and a map spelled out over a blank line with trailing whitespace, and the dispatch lambdas now separate their closing angle brackets the way the rest of the generated code does. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
mristin
added a commit
that referenced
this pull request
Sep 22, 2026
C++ is the last backend to have its XML read side decomposed, after C# (#686), Go (#688), Java (#690), Python (#694) and TypeScript (#698). #673 factored out the framing of an *element*; this factors out the framing of the property *loop*, which every ``<Cls>FromSequence`` re-emitted in full. **A property** loses its duplicate check and its ``std::tie``, and names only itself and what reads it. // before case properties::OfExtension::kSemanticId: { if (the_semantic_id.has_value()) { error = DuplicatePropertyError(name); break; } std::tie( the_semantic_id, error ) = ReferenceFromSequence< types::IReference >(reader); break; } // after case properties::OfExtension::kSemanticId: return ReadInto( the_semantic_id, ReferenceFromSequence< types::IReference >(reader) ); **The loop** around it becomes one call. It could be lifted out at all only because the framing now returns the error alone, instead of ``NoInstanceAndDeserializationError<std::shared_ptr<T> >``, and so no longer needs ``T``. // before, ~220 lines in each of the 38 ``<Cls>FromSequence`` while (true) { error = SkipWhitespace(reader); if (error.has_value()) { ...6 lines... } if (reader.node().kind() == xml_common::NodeKind::Stop) { break; } else if (reader.node().kind() != xml_common::NodeKind::Start) { ...10 lines naming IExtension... } const std::string name(...); reader.Read(); auto it = properties::kMapOfExtension.find(name); if (it == properties::kMapOfExtension.end()) { ...12 lines... } switch (it->second) { ... } ...~60 lines of prepending and stop-element handling... } // after common::optional<DeserializationError> error( ReadProperties< properties::kPropertyCountOfExtension >( reader, properties::kMapOfExtension, L"IExtension", [&]( properties::OfExtension property ) -> common::optional<DeserializationError> { switch (property) { ... } } ) ); The ``switch`` stays inline because its branches assign the locals the constructor is called with -- the very reason C#, Go and Java had to keep the whole loop inline. The lambda is a template argument taken by ``const&``, so the call is statically bound and nothing is type-erased or allocated; ``g++ -O2`` keeps ``ReadProperties`` itself out of line, one instantiation per class, which is as many function bodies as the loops it replaces. **The duplicate check** is bookkeeping of the loop, not of the property, so it moved there, as a bit set sized to the class. It still precedes the read, so a duplicate is refused without its content ever being looked at. std::bitset<kPropertyCount> seen; ... if (seen[index]) { error = DuplicatePropertyError(name); PrependElementSegmentToDeserializationError(name, *error); return error; } seen[index] = true; **``ReadInto`` takes the results of a read, not the reader.** A tuple's item readers vary in number *and* in type, so no reader-taking signature serves them all, while one taking the result serves every read. template <typename T> common::optional<DeserializationError> ReadInto( common::optional<T>& target, std::pair< common::optional<T>, common::optional<DeserializationError> >&& read ); **A named union is framed like a class.** The two framings were ~160 near-identical lines differing only in the value type, which every error factory they call is already generic in. // before template <typename T, typename DispatchT> ... DeserializeClassFromElement(reader, const std::wstring&, dispatch); template <typename VariantT, typename DispatchT> ... DeserializeUnionFromElement(reader, const std::wstring&, dispatch); // after template <typename ValueT, typename DispatchT> ... DeserializeFromElement(reader, const wchar_t*, dispatch); **An element with a sole model type does not dispatch.** 40 of the 51 ``<Cls>FromElement`` spelled out a one-case ``switch`` inside a fifteen-line lambda signature; the other 11 keep theirs. return DeserializeSoleFromElement< std::shared_ptr<types::IExtension> >( reader, L"IExtension", types::ModelType::kExtension, ExtensionFromSequence<types::IExtension> ); **The public ``<Cls>From``** were 51 copies of the same 54 lines: open the reader, read the root element, check that nothing but whitespace follows. return DeserializeFrom< std::shared_ptr<types::IExtension> >( is, options, ExtensionFromElement ); **``interface_name`` is a ``const wchar_t*``**, not a ``const std::wstring&``: every call site used to construct a ``std::wstring`` on *every* element read, success path included, for a message which is almost never built. Written for C++11, which the library targets outside of its tests: no generic lambda, no deduced return type, every lambda spells out its parameters and its return type. ``xmlization.cpp`` for the AAS meta-model goes 34,543 -> 24,634 lines and its read region 23,584 -> 13,675 (-42%). The ``<Cls>FromSequence`` band goes 13,400 -> 6,076 over 37 classes (-54.7%), with ``Key`` 257 -> 95 and ``Extension`` 331 -> 141, and the public ``<Cls>From`` 2,751 -> 813 (-70%). ``unions`` goes 7,950 -> 5,795; ``empty_class`` goes 1,622 -> 1,674, since one property-less class pays for the shared loop with nothing to save. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
mristin
added a commit
that referenced
this pull request
Sep 22, 2026
C++ is the last backend to have its XML read side decomposed, after C# (#686), Go (#688), Java (#690), Python (#694) and TypeScript (#698). #673 factored out the framing of an *element*; this factors out the framing of the property *loop*, which every ``<Cls>FromSequence`` re-emitted in full. **A property** loses its duplicate check and its ``std::tie``, and names only itself and what reads it. // before case properties::OfExtension::kSemanticId: { if (the_semantic_id.has_value()) { error = DuplicatePropertyError(name); break; } std::tie( the_semantic_id, error ) = ReferenceFromSequence< types::IReference >(reader); break; } // after case properties::OfExtension::kSemanticId: return ReadInto( the_semantic_id, ReferenceFromSequence< types::IReference >(reader) ); **The loop** around it becomes one call. It could be lifted out at all only because the framing now returns the error alone, instead of ``NoInstanceAndDeserializationError<std::shared_ptr<T> >``, and so no longer needs ``T``. // before, ~220 lines in each of the 38 ``<Cls>FromSequence`` while (true) { error = SkipWhitespace(reader); if (error.has_value()) { ...6 lines... } if (reader.node().kind() == xml_common::NodeKind::Stop) { break; } else if (reader.node().kind() != xml_common::NodeKind::Start) { ...10 lines naming IExtension... } const std::string name(...); reader.Read(); auto it = properties::kMapOfExtension.find(name); if (it == properties::kMapOfExtension.end()) { ...12 lines... } switch (it->second) { ... } ...~60 lines of prepending and stop-element handling... } // after common::optional<DeserializationError> error( ReadProperties< properties::kPropertyCountOfExtension >( reader, properties::kMapOfExtension, L"IExtension", [&]( properties::OfExtension property ) -> common::optional<DeserializationError> { switch (property) { ... } } ) ); The ``switch`` stays inline because its branches assign the locals the constructor is called with. The lambda is a template argument taken by ``const&``, so the call is statically bound and nothing is type-erased or allocated; ``g++ -O2`` keeps ``ReadProperties`` itself out of line, one instantiation per class, which is as many function bodies as the loops it replaces. **The duplicate property element check** is bookkeeping of the loop, not of the property, so it moved there, as a bit set sized to the class. It still precedes the read, so a duplicate is refused without its content ever being looked at. std::bitset<kPropertyCount> seen; ... if (seen[index]) { error = DuplicatePropertyError(name); PrependElementSegmentToDeserializationError(name, *error); return error; } seen[index] = true; **``ReadInto`` takes the results of a read, not the reader.** A tuple's item readers vary in number *and* in type, so no reader-taking signature serves them all, while one taking the result serves every read. template <typename T> common::optional<DeserializationError> ReadInto( common::optional<T>& target, std::pair< common::optional<T>, common::optional<DeserializationError> >&& read ); **A named union is framed like a class.** The two framings were ~160 near-identical lines differing only in the value type, which every error factory they call is already generic in. // before template <typename T, typename DispatchT> ... DeserializeClassFromElement(reader, const std::wstring&, dispatch); template <typename VariantT, typename DispatchT> ... DeserializeUnionFromElement(reader, const std::wstring&, dispatch); // after template <typename ValueT, typename DispatchT> ... DeserializeFromElement(reader, const wchar_t*, dispatch); **An element with a sole model type does not dispatch.** 40 of the 51 ``<Cls>FromElement`` spelled out a one-case ``switch`` inside a fifteen-line lambda signature; the other 11 keep theirs. return DeserializeSoleFromElement< std::shared_ptr<types::IExtension> >( reader, L"IExtension", types::ModelType::kExtension, ExtensionFromSequence<types::IExtension> ); **The public ``<Cls>From``** were 51 copies of the same 54 lines: open the reader, read the root element, check that nothing but whitespace follows. return DeserializeFrom< std::shared_ptr<types::IExtension> >( is, options, ExtensionFromElement ); **``interface_name`` is a ``const wchar_t*``**, not a ``const std::wstring&``: every call site used to construct a ``std::wstring`` on *every* element read, success path included, for a message which is almost never built. ``xmlization.cpp`` for the AAS meta-model goes 34,543 -> 24,634 lines and its read region 23,584 -> 13,675 (-42%).
mristin
added a commit
that referenced
this pull request
Sep 22, 2026
C++ is the last backend to have its XML read side decomposed, after C# (#686), Go (#688), Java (#690), Python (#694) and TypeScript (#698). #673 factored out the framing of an *element*; this factors out the framing of the property *loop*, which every ``<Cls>FromSequence`` re-emitted in full. **A property** loses its duplicate check and its ``std::tie``, and names only itself and what reads it. // before case properties::OfExtension::kSemanticId: { if (the_semantic_id.has_value()) { error = DuplicatePropertyError(name); break; } std::tie( the_semantic_id, error ) = ReferenceFromSequence< types::IReference >(reader); break; } // after case properties::OfExtension::kSemanticId: return ReadInto( the_semantic_id, ReferenceFromSequence< types::IReference >(reader) ); **The loop** around it becomes one call. It could be lifted out at all only because the framing now returns the error alone, instead of ``NoInstanceAndDeserializationError<std::shared_ptr<T> >``, and so no longer needs ``T``. // before, ~220 lines in each of the 38 ``<Cls>FromSequence`` while (true) { error = SkipWhitespace(reader); if (error.has_value()) { ...6 lines... } if (reader.node().kind() == xml_common::NodeKind::Stop) { break; } else if (reader.node().kind() != xml_common::NodeKind::Start) { ...10 lines naming IExtension... } const std::string name(...); reader.Read(); auto it = properties::kMapOfExtension.find(name); if (it == properties::kMapOfExtension.end()) { ...12 lines... } switch (it->second) { ... } ...~60 lines of prepending and stop-element handling... } // after common::optional<DeserializationError> error( ReadProperties< properties::kPropertyCountOfExtension >( reader, properties::kMapOfExtension, L"IExtension", [&]( properties::OfExtension property ) -> common::optional<DeserializationError> { switch (property) { ... } } ) ); The ``switch`` stays inline because its branches assign the locals the constructor is called with. The lambda is a template argument taken by ``const&``, so the call is statically bound and nothing is type-erased or allocated; ``g++ -O2`` keeps ``ReadProperties`` itself out of line, one instantiation per class, which is as many function bodies as the loops it replaces. **The duplicate property element check** is bookkeeping of the loop, not of the property, so it moved there, as a bit set sized to the class. It still precedes the read, so a duplicate is refused without its content ever being looked at. std::bitset<kPropertyCount> seen; ... if (seen[index]) { error = DuplicatePropertyError(name); PrependElementSegmentToDeserializationError(name, *error); return error; } seen[index] = true; **``ReadInto`` takes the results of a read, not the reader.** A tuple's item readers vary in number *and* in type, so no reader-taking signature serves them all, while one taking the result serves every read. template <typename T> common::optional<DeserializationError> ReadInto( common::optional<T>& target, std::pair< common::optional<T>, common::optional<DeserializationError> >&& read ); **A named union is framed like a class.** The two framings were ~160 near-identical lines differing only in the value type, which every error factory they call is already generic in. // before template <typename T, typename DispatchT> ... DeserializeClassFromElement(reader, const std::wstring&, dispatch); template <typename VariantT, typename DispatchT> ... DeserializeUnionFromElement(reader, const std::wstring&, dispatch); // after template <typename ValueT, typename DispatchT> ... DeserializeFromElement(reader, const wchar_t*, dispatch); **An element with a sole model type does not dispatch.** 40 of the 51 ``<Cls>FromElement`` spelled out a one-case ``switch`` inside a fifteen-line lambda signature; the other 11 keep theirs. return DeserializeSoleFromElement< std::shared_ptr<types::IExtension> >( reader, L"IExtension", types::ModelType::kExtension, ExtensionFromSequence<types::IExtension> ); **The public ``<Cls>From``** were 51 copies of the same 54 lines: open the reader, read the root element, check that nothing but whitespace follows. return DeserializeFrom< std::shared_ptr<types::IExtension> >( is, options, ExtensionFromElement ); **``interface_name`` is a ``const wchar_t*``**, not a ``const std::wstring&``: every call site used to construct a ``std::wstring`` on *every* element read, success path included, for a message which is almost never built. ``xmlization.cpp`` for the AAS meta-model goes 34,543 -> 24,634 lines and its read region 23,584 -> 13,675 (-42%).
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The generated
xmlization.csrepeated itself at three levels. Every property of every concrete class inlined its whole reading procedure into its owncaseblock -- the self-closing check, the end-of-file check, thetry/catcharound the conversion, an error message naming that property and its class, thePrependSegmentmarking the path, and, for a list or a tuple, the loop or the positional<v>elements on top. Every concrete class then got its own...FromElementfunction, and all of those were the same 43 lines with four tokens substituted in. And every...FromSequencere-emitted the ~70 lines of framing around its property loop, which said nothing about the class it belonged to.Everything is now read through one composable shape, and the pieces that do not vary are generated once:
The rationale and the design:
One shape for every reader. Reading a property's content, a list item's and a tuple item's are the same operation at different nesting depths, so they are one delegate:
Because the shape is uniform a reader can be an argument to another reader, so a list of tuples -- or anything deeper the meta-model may grow -- falls out of the existing pieces instead of needing a generated helper per combination.
A plain
Treturn, notT?and notout T. Failure travels in the error alone and the value is thendefault!, which the caller never reads.T?would have split the machinery in two, since a nullable return isNullable<T>for a value type but a nullable reference otherwise and no unconstrained parameter covers both.out Twould have compiled, butoutis invariant, so a reader of a concrete class could no longer serve as the item reader of an interface-typed list; a return type is covariant in a method-group conversion.Combinators return delegates, because C# has no partial application:
AsText,AsEnum,AsList,AsElementandAsTuple1..Neach take what varies and return aContentReader.AtElementis the hinge, and the element name is data. It binds a name to aContentReaderand yields anElementReader, which is what a list and a tuple take for their items -- and also what a class is. A name cannot be a type argument (vin a list,v1,v2, ... by position in a tuple, its own XML name for a class), so it is passed. With the messages phrased in terms of that name, the combinator needs nothing further to say what it expected, which is what lets one combinator serve both a class and a<v>element.A class's element reader is therefore a binding, not a function. A class's
...FromSequencealready is aContentReader, so its element reader isAtElement(ExtensionFromSequence, "extension")and the 43-line function per class is gone. Call sites are unchanged: a field is invoked exactly like the method was.The composed readers live in
static readonlyfields. This is not cosmetic. Composing at the call site allocates a closure on every read -- measured at 152 bytes against zero for the code being replaced, a real regression for a de-serialization library. Composing once costs nothing per read and de-duplicates: 276 properties collapse onto 40 readers.Two consequences worth knowing. Field initialisers run in declaration order, so a reader is emitted after everything it composes. And
ElementReaderhad to become covariant andinternal: as a method group a class's reader widened to its interface for free, but a field is a value, and a field may not be less accessible than its type.Looking ahead an element's name is one helper,
PeekElementName, shared byAtElementand by the thirteen readers that dispatch on a discriminator element -- they only ever needed the name, and consuming nothing is what lets the dispatch hand the whole element on.The property loop keeps only its
switch.TryNextPropertyreads the next property's start tag or reports that the sequence ended, andConsumeEndElement-- shared withAtElement, which concludes identically -- consumes the end tag. Only theswitchstays inline, because its branches assign the locals the constructor is called with.The property is marked on the error path once per class, after the
switch, rather than in every case: for a matched case the element name already is the property's XML name.The type in a message comes from
typeof(T).Name, so the shared skeletons do not grow with the number of types in the meta-model.Only what a model reaches is emitted, gated along the call graph rather than by property kind -- a list of enumerations needs the enumeration combinator even when no property is one. Without that, a small meta-model would pay for helpers it never calls.
For the AAS meta-model,
xmlization.csgoes from 22,572 to 12,471 lines (-44.8%), the per-property blocks from 19.7 to 3.7 lines on average, and the read side loses 38 generated functions along with four helpers.