Compose the Go XML de-serialization - #688
Merged
Merged
Conversation
This is the Go counterpart of #686, but the C# design does not carry over: the shape that works there fails to compile in Go, so the read side is composed out of generated dispatchers instead of composed readers. Measured on the AAS meta-model fixture, the read region of ``xmlization.go`` goes 11,738 -> 5,584 lines (-52%), and the file as a whole 22,846 -> 16,692. The ``read<X>WithLookahead`` band -- 50 functions, 3,057 lines -- goes to **zero**; 20 ``read<X>Dispatched`` functions stand in its place. The whole design turns on one asymmetry in the wire format. An instance element is self-describing: its local name *is* its type. A scalar element is not: its name is its *position*, ``v`` in a list and ``v1``, ``v2``, ... in a tuple. So: every instance is read by dispatch on its own name, every scalar by a name its container supplies, and nothing else in the read path knows the difference. ``readElementDispatched`` is the only place which frames an XML element around a value. ``readListOf`` and every ``readTupleN`` delegate the framing to it, which is what deletes ``readListOfScalars``, ``readScalarWithName``, ``asScalarTupleItemReader`` and ``asInstanceTupleItemReader`` outright: one list helper, one item shape, no adapters. It also returns the token *past* the end element, so the "``read<X>WithLookahead`` stops at the end element, so we look ahead" comment and its hand-written ``readNext`` vanish from every call site: readListOf(decoder, current, readReferenceDispatched) readListOf(decoder, current, readSubmodelElementDispatched) readListOf(decoder, current, readLongAtV) readTuple6( decoder, current, readLongAtV1, readSomeItemDispatched, readAbstractItemDispatched, readSomeItemDispatched, readLongAtV5, readResultAtV6, ) Per class, ``read<X>AsSequence`` has for every property case one assignment of the same ``(value, current, valueErr)`` triple: case "supplementalSemanticIds": theSupplementalSemanticIDs, current, valueErr = readListOf( decoder, current, readReferenceDispatched, ) case "valueType": theValueType, current, valueErr = readOptional( readTextAsDataTypeDefXSD(decoder, current), ) ``readOptional`` absorbs the ``var value T; ...; theX = &value`` dance. It takes the *results* of a read rather than the reader, so Go's multi-value pass-through lets one helper serve any read however many arguments that read takes -- including ``readOptional(readTuple2(...))``, which a reader-taking signature cannot express, since a tuple's item readers vary in number and in type. That also fixes an optional tuple property, which is a pointer (a Go struct is not nilable) and used to be assigned a non-pointer ``readTupleN`` result: output that does not compile.
Coverage Report for CI Build 34775836793Coverage increased (+0.02%) to 85.272%Details
Uncovered Changes
Coverage RegressionsNo coverage regressions found. Coverage Stats
💛 - Coveralls |
mristin
added a commit
that referenced
this pull request
Sep 13, 2026
Write-side dual of #688, as #687 was of #686 in C#. Measured on the AAS meta-model fixture, the serialization region of ``xmlization.go`` goes 11,080 -> 4,727 lines (-57%), and the file as a whole 16,692 -> 10,330. ``writeExtensionAsSequence`` and its six properties go 180 -> 69 lines. Everything is written through one shape, ``func(encoder, value T) error``. A stringification, a class's own ``write<X>AsSequence`` and the generated item, list and tuple writers all have it, so ``writeElement`` -- which frames an XML element around a value written that way -- is the only writer a property needs, whatever kind of value it holds. What is left per property is that one call and a decision on presence: err = finishProperty( "SemanticID()", writeOptionalInstance( encoder, "semanticId", that.SemanticID(), writeReferenceAsSequence, ), ) if err != nil { return } ``finishProperty`` takes the *result* of the write rather than the write itself -- the same trick as the read side's ``readOptional``, since Go evaluates the argument first -- so one function concludes every property, prepending the getter to the path of the error. Mind that the getter is spelled out as it is in Go, ``SemanticID()`` and not ``semanticId``, since a serialization error renders its path through ``aasreporting.ToGolangPath``, while the de-serialization prepends the XML name, its errors being XPaths. Go has no ``nameof``. The ``writeOptional*`` family has one member per way Go spells an optional, each named after that spelling, as the check for the presence is the only thing which differs between them: a pointer to the value (``writeOptionalPointer``: a scalar, an enumeration, a tuple), a value which is nil on its own (``writeOptionalInstance``: an instance, or a named union, which is a pointer to a struct) and a slice (``writeOptionalSlice``: a list, the bytes). They can not be collapsed into one, since Go decides nil-ness by the representation: a type parameter can not be compared against ``nil``, ``any(that) == nil`` is false for a nil *pointer* boxed in an ``any`` -- hence the comparison against the zero value in ``writeOptionalInstance``, which in turn panics on a slice, which is not comparable. Each of them writes nothing at all when the value is absent, which is what keeps the absence out of the generated code. No closure is allocated on the write path any more: every writer passed on is a top-level function or an instantiated generic, hence a static funcval, exactly as on the read side. A list and a tuple get one generated content writer as well, per item type resp. per tuple type -- ``writeListOfIReference`` and ``writeTupleOfStringLong`` and their kin -- since ``writeList`` and ``writeTupleN`` take a writer per item, and Go gives no partial application to bind those in without allocating. Finally, the flush. ``write<X>AsSequence`` used to flush the encoder after every property, and ``write<X>`` again at its end element, which defeated the buffering of ``xml.Encoder`` on every write; only the flush after the last token is needed, as ``EncodeToken`` deliberately leaves it to the caller. ``Marshal`` is therefore split into ``writeClass``, which dispatches on the model type, and ``Marshal`` itself, which calls it and flushes exactly once; ``writeInstance`` and ``writeUnion`` go to ``writeClass``, so a nested instance does not flush either. Serializing an instance holding 1,000 instances (42 KB of XML) from the ``list_of_classes`` fixture, the writes which actually reach the underlying writer go 2,003 -> 11, one per 4 KB of buffer instead of one per property and per instance; to an ``os.File`` that is 5.29 -> 0.76 ms per serialization, the median of five.
mristin
added a commit
that referenced
this pull request
Sep 14, 2026
Java is the third backend to have its XML read side decomposed, after C# (#686) and Go (#688), and it carried the repetition #686 described at four levels rather than three. Every property inlined its whole reading procedure into its own ``case`` block. Everything is now read through one composable shape, and the pieces which do not vary are generated once: case "contentType": { final Reporting.Result<String> value = readTextAs_string(reader, isEmptyProperty); if (value.isError()) { valueError = value.getError(); } else { theContentType = value.getResult(); } break; } The rationale and the design: * A method reference to a ``static`` method captures nothing, so its ``invokedynamic`` call site is linked once and hands back the same instance afterwards; passing ``_DeserializeImplementation::readReference`` costs nothing. A capturing lambda is not free, and the old code had one in every ``try<X>FromElement`` and in every dispatcher, both closing over the enclosing ``reader``: an allocation on every single instance read. The read side now contains no lambdas at all -- every reader passed anywhere is a method reference to a static method -- so nothing can accidentally capture. * Two shapes, and two only. ``ContentReader<T>`` reads the content of an element which has already been opened, ``ElementReader<T>`` reads a whole element. They are duals, and the wire format forces both, because it is asymmetric: an instance element is self-describing -- its local name *is* its type -- whereas a scalar element is not, its name being its position (the property's name, ``v`` in a list, ``v1``, ``v2``, ... in a tuple). * One primitive, ``peekElementName``. It consumes nothing, and it is the single answer to "we are at an element, and this is its name". The framer ``readNamedElement`` checks that name against the one its container supplied, the dispatchers switch on it, and every property loop selects the property with it. That is also why no third shape is needed to carry the name into a dispatcher: a dispatcher peeks, switches, and hands the *whole* element on to that type's own ``read<Cls>FromElement``, which re-checks the name and consumes it. * The property loop keeps only its ``switch``, because its branches assign the locals the constructor is called with. ``atEndOfSequence`` is the loop condition, ``consumeEndElement`` -- shared with ``readNamedElement``, which concludes identically -- consumes the end tag, and ``unexpectedProperty`` and ``missingRequiredProperty`` are the two errors which used to be spelled out per class and per property. * The property is marked on the error path once per class, after the ``switch``, rather than in every case: for a matched case the element name already is the property's XML name. That also retires the per-case ``castTo(<Cls>.class)`` re-threading, and fixes an inconsistency along the way -- a primitive and an enumeration used to be marked with the *Java* property name while a list and a class used the XML name. * The names are unique by construction, so no uniqueness check is needed. A reader is keyed either on one symbol or on a type moniker, and the two name spaces are kept apart by a single rule: a moniker-keyed name contains ``_``, a symbol-keyed name never does. That holds because ``naming.capitalized_camel_case``, which every one of our type names goes through, emits letters and digits only, with an upper-case initial. The monikers are then a Polish notation over ``_``-separated tokens -- ``ListOf_string``, ``TupleOf2_string_long``, ``ListOf_TupleOf2_string_long`` -- with the arity spelled out, so the encoding is injective. The primitives are lower-cased for the same reason: an enumeration which somebody names ``String`` can then never collide with the primitive ``str``. * ``readNestedElement`` looks redundant and is not. A polymorphic property and a list item leave the reader at the very same position, just before a self-describing start element, so it is tempting to give ``read<X>FromElement`` an ``isEmpty`` parameter and call it from the property ``case`` directly. The error path is what forbids it. A property wraps its instance in an element of its own, so the discriminator's name has to be prepended and the path reads ``value/property/idShort``; a list item is not wrapped -- the item element *is* the indexed child -- so the same prepend would give ``annotations/*[0]/property/idShort``, which walks one level past the element ``*[0]`` already selects and resolves to nothing. The reasoning is written into the generated doc comment as well, because it is exactly the kind of layer the next reader deletes. * Only what a model reaches is emitted, gated along the call graph rather than by property kind -- a list of enumerations needs the enumeration reader even when no property is one, and an enumeration needs the text path because it is built on it. The pass recurses into the items of a list and of a tuple, without which a model whose only ``str`` sits in a ``List[List[str]]`` would emit a reader built from helpers which were never generated. The snippet contract for an implementation-specific class changes: ``Xmlization/DeserializeImplementation/{cls}_from_element.java`` is no longer read, and the remaining ``{cls}_from_sequence.java`` has to define ``read<Cls>FromSequence`` rather than ``try<Cls>FromSequence``. For the AAS meta-model, ``Xmlization.java`` goes from 16,647 to 11,455 lines (-31.2%) and its read region from 12,187 to 6,995 (-42.6%). Finally, the speed, since the read side is where a de-serialization library is judged: there is no measurable difference, and that is the expected answer. De-serializing a 1.4 MB AAS environment (200 submodels, 131,806 StAX events; 400 runs after 200 warm-up runs, three measurements of each variant, run alternately), the medians are 32.2/33.2/33.2 ms before and 32.1/32.7/33.2 ms after -- the same, within the noise of repeated runs of either one. Draining that very same document through ``XMLEventReader`` without building anything at all already costs 27.4 ms and allocates 41.2 MB, so StAX accounts for ~83% of the time and ~84% of the allocation, and everything the generated code does is the remaining ~6 ms and ~8.2 MB. Against that, the lambdas which are gone are worth ~0.15 MB per run at the median (49.48-49.57 MB before, 49.14-49.44 MB after): consistently in that direction, but under 1% of the total. The decomposition is a maintainability change, not a performance one, and it was worth measuring rather than assuming, since the C# commit's 152 bytes/op is what a reader would expect to carry over.
mristin
added a commit
that referenced
this pull request
Sep 15, 2026
Write-side dual of #690, as #687 was of #686 in C# and #689 of #688 in Go. ``visit{Cls}`` was a verbatim copy of ``serializeElement``: 38 methods of 18 lines each spelling out the start tag, the namespace, the content and the end tag, around a call which already did exactly that. Six near-identical emitters -- one per property kind -- produced the very same ``serializeElement`` call, differing only in the content serializer. And everything was an *instance* method, because of one mutable ``topLevel`` flag. That flag is what cost the most. It made ``this::writeStringifiedContent``, ``this::visit`` and ``(value, w) -> this.xToSequence(value, w)`` capturing, and ``serializeItems(...)`` and ``asNamedElementSerializer(...)`` returned capturing lambdas on top of them, so writing one instance allocated a small pile of closures. It is replaced by a ``final boolean withNamespace`` pinned in the constructor of two immutable visitors, ``ROOT`` and ``NESTED``: only the outermost element declares the namespace, and an element written from ``writeClass`` is nested by construction. Everything else is then ``static``, and a method reference to a static method captures nothing. The design is otherwise the read side's, mirrored: * One shape, not two. ``ContentWriter<T>`` writes a value where the writer already is; ``writeElement`` frames it in a start and an end tag. There is deliberately no second delegate for a whole element, as there is on the reading side -- an element differs from a content only in what it writes, never in its shape. The reading needs the distinction because a content reader has to be told whether its element was self-closing, and a writer has nothing to be told. Use sites take ``ContentWriter<? super T>``, which is how Java spells the contravariance C# writes as ``in T``, so the single writer of an ``IClass`` serves wherever the writer of a more specific interface is expected. * ``visit{Cls}`` collapses to one call, and the visitor becomes a dispatch shim:: @OverRide public void visitQualifier( IQualifier that, XMLStreamWriter writer) { writeElement( "qualifier", that, writer, withNamespace, _VisitorWithWriter::writeQualifierAsSequence); } * The families are dual to the readers and named by the same moniker grammar: ``write{Cls}AsSequence``, ``writeClass``/``writeUnion``, ``writeListOf_{M}``, ``writeTupleOf{N}_{M...}``, ``writeAtV_{M}`` and ``writeAtV{i}_{M}``. The list and the tuple writers inline their loop, which retires ``serializeItems``, ``asNamedElementSerializer`` and ``serializeTuple{N}`` -- Java needs no partial application here, since the name to bind is fixed per generated function. * One property emitter replaces the six. Every property is written by ``writeProperty``, or by ``writeOptionalProperty``, which takes the ``Optional`` apart once instead of asking it and then unwrapping it -- the old form called the getter twice, and every call allocates an ``Optional``. * ``_type_moniker``, ``_leaf_moniker``, ``_is_instance_type``, ``_is_dispatched`` and ``_collect_needed`` are reused unchanged: the reading and the writing are composed out of the same shapes, so one gating pass answers for both. Only the two dispatching writers needed a walk of their own, because the *position* matters on the write side while it does not on the read side: an item of any class is written as its own element, whereas a property of a concrete class without descendants writes its properties straight into the property element. Two things are genuinely asymmetric, and both save work: * There is no ``writeTextAs_{primitive}`` family. Once the value is at hand nothing type-specific is left to do with it, so the four ``toString`` primitives share ``writeStringifiedContent`` and the byte array ``writeByteArrayContent``. An enumeration does have something of its own to write and keeps one function each -- named ``writeTextAs_{Enum}`` and not ``write{Enum}Content``, so that it stays in the moniker-keyed name space: were it keyed on the symbol, an enumeration which somebody names ``Stringified`` or ``ByteArray`` would silently take a shared writer's name. * There is no element writer per class. Which element an instance writes follows from its run-time type, and one virtual call answers that for every class at once. The reading needs a dispatcher per interface because it has to decide what to construct before it has read anything. Two defects of the write side are fixed along the way, both of which the decomposition makes one-line changes rather than sweeping ones. ``Serialize.to`` never flushed, so the serialization was not finished when it returned, and a failure of the underlying stream surfaced at the caller's own flush as a raw ``XMLStreamException`` -- outside the one method whose ``throws`` clause promises a ``SerializeException``. Probed with a ``Writer`` which accepts every character and fails on ``flush``, ``to`` used to return normally and now throws. The flush goes in ``Serialize.to`` and nowhere else, as in Go's #689, so ``XMLStreamWriter`` keeps buffering in between. Mind that this does push the underlying stream: a caller serializing many small documents into one buffered writer now flushes once per document rather than once per buffer. That is the right trade, since the alternative leaves the output silently incomplete whenever the caller forgets. Every ``SerializeException`` carried an empty path, so every failure read ``... at: the beginning`` no matter where it happened. The path is now built as the stack unwinds, each container prepending the one segment it knows, exactly as on the read side where ``read<Cls>FromSequence`` prepends the property after its ``switch`` and ``readList`` prepends the index. Probed with an ``XMLStreamWriter`` which refuses to write one chosen value: a property of the root instance idShort nested through two lists submodels/*[1]/submodelElements/*[1]/value through a class-typed property semanticId/keys/*[0]/value ``SerializeException`` renders its message in its constructor, so its path has to be complete by the time it is built and it can not be the thing thrown from inside; a private ``_SerializeFailure`` carries a ``Reporting.Error`` instead, and ``Serialize.to`` renders that with ``Reporting.generateRelativeXPath`` and converts. The public exception is untouched. ``writeElement`` contributes no segment and ``writeProperty`` is ``writeElement`` plus the one thing a property knows which nothing below it does -- its own name -- so there is no ``try`` around any of the 276 property writes. A list and a tuple put their ``try`` outside the loop and advance the index only once an item has been written, so the item named is the one which failed and nothing is paid per item. Two segments are deliberately absent, and the reasoning is written into the generated Javadoc, since both look like oversights. The outermost element is not named, because ``Serialize.to`` takes any ``IClass`` and the name would say nothing the caller does not already know; the read side does prepend it, having one facade per class. Neither is the discriminator element of a polymorphic property, which the read side does prepend -- for a failure in an embedded data specification it reports ``submodel/embeddedDataSpecifications/*[0]/dataSpecificationContent/`` ``dataSpecificationIec61360/preferredName/*[0]`` where the write side reports ``embeddedDataSpecifications/*[0]/dataSpecificationContent/`` ``preferredName/*[0]/text``. The reading points into a document it is parsing, where that element is a real extra level; the writing points into the instance the caller handed over, where it is not, and ``value/idShort`` there is exactly ``getValue().getIdShort()``. Go's #689 made the same choice when it rendered its serialization paths through ``aasreporting.ToGolangPath`` rather than as XPaths. For the AAS meta-model the write region of ``Xmlization.java`` goes 3,281 -> 2,873 lines (-12.4%), error paths included. The write region no longer contains a single lambda. Unlike the read side, that shows up in the measurement. Serializing an ``Environment`` of 200 submodels (1.14 MB of XML) into a writer which counts the characters and discards them, 100 runs after 100 warm-up runs, eight rounds of each variant run alternately: the allocation goes 0.76 -> 0.36 MB per serialization, and the time 7.69 -> 7.49 ms, which is inside the round-to-round noise. ``XMLStreamWriter`` writes straight through, so the whole write path allocates ~0.8 MB against the ~49 MB the *reading* of the same document allocates -- there is no StAX floor here for the closures to hide behind, which is precisely why #690 could not measure its own. The snippet contract for an implementation-specific class changes: ``Xmlization/VisitorWithWriter/visit_{cls}.java`` is no longer read, as the element around the sequence is now written identically for every class, and the remaining ``{cls}_to_sequence.java`` has to define a ``static`` ``write{Cls}AsSequence``.
mristin
added a commit
that referenced
this pull request
Sep 15, 2026
Write-side dual of #690, as #687 was of #686 in C# and #689 of #688 in Go. ``visit{Cls}`` was a verbatim copy of ``serializeElement``: 38 methods of 18 lines each spelling out the start tag, the namespace, the content and the end tag, around a call which already did exactly that. Six near-identical emitters -- one per property kind -- produced the very same ``serializeElement`` call, differing only in the content serializer. And everything was an *instance* method, because of one mutable ``topLevel`` flag. That flag is what cost the most. It made ``this::writeStringifiedContent``, ``this::visit`` and ``(value, w) -> this.xToSequence(value, w)`` capturing, and ``serializeItems(...)`` and ``asNamedElementSerializer(...)`` returned capturing lambdas on top of them, so writing one instance allocated a small pile of closures. It is replaced by a ``final boolean withNamespace`` pinned in the constructor of two immutable visitors, ``ROOT`` and ``NESTED``: only the outermost element declares the namespace, and an element written from ``writeClass`` is nested by construction. Everything else is then ``static``, and a method reference to a static method captures nothing. The design is otherwise the read side's, mirrored: * One shape, not two. ``ContentWriter<T>`` writes a value where the writer already is; ``writeElement`` frames it in a start and an end tag. There is deliberately no second delegate for a whole element, as there is on the reading side -- an element differs from a content only in what it writes, never in its shape. The reading needs the distinction because a content reader has to be told whether its element was self-closing, and a writer has nothing to be told. Use sites take ``ContentWriter<? super T>``, which is how Java spells the contravariance C# writes as ``in T``, so the single writer of an ``IClass`` serves wherever the writer of a more specific interface is expected. * ``visit{Cls}`` collapses to one call, and the visitor becomes a dispatch shim:: @OverRide public void visitQualifier( IQualifier that, XMLStreamWriter writer) { writeElement( "qualifier", that, writer, withNamespace, _VisitorWithWriter::writeQualifierAsSequence); } * The families are dual to the readers and named by the same moniker grammar: ``write{Cls}AsSequence``, ``writeClass``/``writeUnion``, ``writeListOf_{M}``, ``writeTupleOf{N}_{M...}``, ``writeAtV_{M}`` and ``writeAtV{i}_{M}``. The list and the tuple writers inline their loop, which retires ``serializeItems``, ``asNamedElementSerializer`` and ``serializeTuple{N}`` -- Java needs no partial application here, since the name to bind is fixed per generated function. * One property emitter replaces the six. Every property is written by ``writeProperty``, or by ``writeOptionalProperty``, which takes the ``Optional`` apart once instead of asking it and then unwrapping it -- the old form called the getter twice, and every call allocates an ``Optional``. * ``_type_moniker``, ``_leaf_moniker``, ``_is_instance_type``, ``_is_dispatched`` and ``_collect_needed`` are reused unchanged: the reading and the writing are composed out of the same shapes, so one gating pass answers for both. Only the two dispatching writers needed a walk of their own, because the *position* matters on the write side while it does not on the read side: an item of any class is written as its own element, whereas a property of a concrete class without descendants writes its properties straight into the property element. Two things are genuinely asymmetric, and both save work: * There is no ``writeTextAs_{primitive}`` family. Once the value is at hand nothing type-specific is left to do with it, so the four ``toString`` primitives share ``writeStringifiedContent`` and the byte array ``writeByteArrayContent``. An enumeration does have something of its own to write and keeps one function each -- named ``writeTextAs_{Enum}`` and not ``write{Enum}Content``, so that it stays in the moniker-keyed name space: were it keyed on the symbol, an enumeration which somebody names ``Stringified`` or ``ByteArray`` would silently take a shared writer's name. * There is no element writer per class. Which element an instance writes follows from its run-time type, and one virtual call answers that for every class at once. The reading needs a dispatcher per interface because it has to decide what to construct before it has read anything. Two defects of the write side are fixed along the way, both of which the decomposition makes one-line changes rather than sweeping ones. ``Serialize.to`` never flushed, so the serialization was not finished when it returned, and a failure of the underlying stream surfaced at the caller's own flush as a raw ``XMLStreamException`` -- outside the one method whose ``throws`` clause promises a ``SerializeException``. Probed with a ``Writer`` which accepts every character and fails on ``flush``, ``to`` used to return normally and now throws. The flush goes in ``Serialize.to`` and nowhere else, as in Go's #689, so ``XMLStreamWriter`` keeps buffering in between. Mind that this does push the underlying stream: a caller serializing many small documents into one buffered writer now flushes once per document rather than once per buffer. That is the right trade, since the alternative leaves the output silently incomplete whenever the caller forgets. Every ``SerializeException`` carried an empty path, so every failure read ``... at: the beginning`` no matter where it happened. The path is now built as the stack unwinds, each container prepending the one segment it knows, exactly as on the read side where ``read<Cls>FromSequence`` prepends the property after its ``switch`` and ``readList`` prepends the index. Probed with an ``XMLStreamWriter`` which refuses to write one chosen value: a property of the root instance idShort nested through two lists submodels/*[1]/submodelElements/*[1]/value through a class-typed property semanticId/keys/*[0]/value ``SerializeException`` renders its message in its constructor, so its path has to be complete by the time it is built and it can not be the thing thrown from inside; a private ``_SerializeFailure`` carries a ``Reporting.Error`` instead, and ``Serialize.to`` renders that with ``Reporting.generateRelativeXPath`` and converts. The public exception is untouched. ``writeElement`` contributes no segment and ``writeProperty`` is ``writeElement`` plus the one thing a property knows which nothing below it does -- its own name -- so there is no ``try`` around any of the 276 property writes. A list and a tuple put their ``try`` outside the loop and advance the index only once an item has been written, so the item named is the one which failed and nothing is paid per item. Two segments are deliberately absent, and the reasoning is written into the generated Javadoc, since both look like oversights. The outermost element is not named, because ``Serialize.to`` takes any ``IClass`` and the name would say nothing the caller does not already know; the read side does prepend it, having one facade per class. Neither is the discriminator element of a polymorphic property, which the read side does prepend -- for a failure in an embedded data specification it reports ``submodel/embeddedDataSpecifications/*[0]/dataSpecificationContent/`` ``dataSpecificationIec61360/preferredName/*[0]`` where the write side reports ``embeddedDataSpecifications/*[0]/dataSpecificationContent/`` ``preferredName/*[0]/text``. The reading points into a document it is parsing, where that element is a real extra level; the writing points into the instance the caller handed over, where it is not, and ``value/idShort`` there is exactly ``getValue().getIdShort()``. Go's #689 made the same choice when it rendered its serialization paths through ``aasreporting.ToGolangPath`` rather than as XPaths. For the AAS meta-model the write region of ``Xmlization.java`` goes 3,281 -> 2,873 lines (-12.4%), error paths included. The write region no longer contains a single lambda. Unlike the read side, that shows up in the measurement. Serializing an ``Environment`` of 200 submodels (1.14 MB of XML) into a writer which counts the characters and discards them, 100 runs after 100 warm-up runs, eight rounds of each variant run alternately: the allocation goes 0.76 -> 0.36 MB per serialization, and the time 7.69 -> 7.49 ms, which is inside the round-to-round noise. ``XMLStreamWriter`` writes straight through, so the whole write path allocates ~0.8 MB against the ~49 MB the *reading* of the same document allocates -- there is no StAX floor here for the closures to hide behind, which is precisely why #690 could not measure its own. The snippet contract for an implementation-specific class changes: ``Xmlization/VisitorWithWriter/visit_{cls}.java`` is no longer read, as the element around the sequence is now written identically for every class, and the remaining ``{cls}_to_sequence.java`` has to define a ``static`` ``write{Cls}AsSequence``.
mristin
added a commit
that referenced
this pull request
Sep 22, 2026
C++ is the last backend to have its XML read side decomposed, after C# (#686), Go (#688), Java (#690), Python (#694) and TypeScript (#698). #673 already factored out the framing of an *element*, but every ``<Cls>FromSequence`` still re-emitted the same ~220 lines of property-loop framing around its ``switch``, so a property cost nineteen lines where it now costs three: case properties::OfExtension::kSemanticId: return ReadInto( the_semantic_id, ReferenceFromSequence< types::IReference >(reader) ); The rationale and the design: * **The framing does not need ``T``.** Returning the error alone, instead of ``NoInstanceAndDeserializationError<std::shared_ptr<T> >`` on every failure path, is the whole reason the loop can be lifted out of the class: template < std::size_t kPropertyCount, typename EnumT, typename OnPropertyT > common::optional<DeserializationError> ReadProperties( xml_common::ReaderMergingText& reader, const std::unordered_map<std::string, EnumT>& map_of_properties, const wchar_t* interface_name, const OnPropertyT& on_property ); * **The ``switch`` stays inline, and in C++ that is free.** Its branches assign the locals the constructor is called with -- which is why C#, Go and Java had to keep the whole loop inline -- but a lambda passed as ``const F&`` to a function template is a stack object which is inlined away, so the class can hand the ``switch`` over and nothing is allocated: common::optional<DeserializationError> error( ReadProperties< properties::kPropertyCountOfExtension >( reader, properties::kMapOfExtension, L"IExtension", [&]( properties::OfExtension property ) -> common::optional<DeserializationError> { switch (property) { ... } } ) ); * **The duplicate check belongs to the loop, not to the property.** It is now a bit set sized to the class and living on the stack, so a class may have arbitrarily many properties and a read allocates nothing for it; the check still precedes the read, so a duplicate is refused without its content ever being looked at, and a five-line guard leaves all 276 cases: std::bitset<kPropertyCount> seen; ... if (seen[index]) { error = DuplicatePropertyError(name); PrependElementSegmentToDeserializationError(name, *error); return error; } seen[index] = true; * **``ReadInto`` takes the results of a read, not the reader.** The item readers of a tuple vary both in number and in type, so no signature taking the reader could serve every read, while one taking the result serves all of them: template <typename T> common::optional<DeserializationError> ReadInto( common::optional<T>& target, std::pair< common::optional<T>, common::optional<DeserializationError> >&& read ); * **A named union is framed exactly like a class.** ``DeserializeUnionFromElement`` and ``DeserializeClassFromElement`` were ~160 near-identical lines which differed only in the value type, and every error factory they called was already generic in it: template <typename ValueT, typename DispatchT> std::pair< common::optional<ValueT>, common::optional<DeserializationError> > DeserializeFromElement( xml_common::ReaderMergingText& reader, const wchar_t* value_name, const DispatchT& dispatch ); * **An element with a sole model type does not dispatch.** 40 of the 51 ``<Cls>FromElement`` spelled out a one-case ``switch`` inside a fifteen-line lambda signature; the remaining 11 keep theirs, since a dense jump table beats a lookup: std::pair< common::optional< std::shared_ptr<types::IExtension> >, common::optional<DeserializationError> > ExtensionFromElement( xml_common::ReaderMergingText& reader ) { return DeserializeSoleFromElement< std::shared_ptr<types::IExtension> >( reader, L"IExtension", types::ModelType::kExtension, ExtensionFromSequence<types::IExtension> ); } * **The public ``<Cls>From`` are one function.** 51 copies of the same 54 lines -- open the reader, read the root element, check that nothing but whitespace follows it -- became a single call each: return DeserializeFrom< std::shared_ptr<types::IExtension> >( is, options, ExtensionFromElement ); * **``interface_name`` is a ``const wchar_t*``.** As a ``const std::wstring&`` it made every call site construct a ``std::wstring`` on *every* element read, success path included, for a message which is almost never built: const wchar_t* interface_name, // was: const std::wstring& Everything is written for C++11, which the library targets outside of its tests: no generic lambda, no deduced return type, and every lambda spells its parameters and its return type out. For the AAS meta-model, ``xmlization.cpp`` goes from 34,543 to 24,617 lines and its read region from 23,584 to 13,658 (-42%). The ``<Cls>FromSequence`` band goes 13,400 -> 6,076 over 37 classes (-54.7%): ``Key`` 257 -> 95, ``AdministrativeInformation`` 295 -> 112, ``Extension`` 331 -> 141. The public ``<Cls>From`` go 2,751 -> 813 (-70%). The ``unions`` fixture goes 7,950 -> 5,795. ``empty_class`` goes 1,622 -> 1,674, since a meta-model with a single property-less class pays for the shared loop without having anything to save. Two warts are cleaned up along the way: a class without properties used to get an enumeration and a map spelled out over a blank line with trailing whitespace, and the dispatch lambdas now separate their closing angle brackets the way the rest of the generated code does. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
mristin
added a commit
that referenced
this pull request
Sep 22, 2026
C++ is the last backend to have its XML read side decomposed, after C# (#686), Go (#688), Java (#690), Python (#694) and TypeScript (#698). #673 factored out the framing of an *element*; this factors out the framing of the property *loop*, which every ``<Cls>FromSequence`` re-emitted in full. **A property** loses its duplicate check and its ``std::tie``, and names only itself and what reads it. // before case properties::OfExtension::kSemanticId: { if (the_semantic_id.has_value()) { error = DuplicatePropertyError(name); break; } std::tie( the_semantic_id, error ) = ReferenceFromSequence< types::IReference >(reader); break; } // after case properties::OfExtension::kSemanticId: return ReadInto( the_semantic_id, ReferenceFromSequence< types::IReference >(reader) ); **The loop** around it becomes one call. It could be lifted out at all only because the framing now returns the error alone, instead of ``NoInstanceAndDeserializationError<std::shared_ptr<T> >``, and so no longer needs ``T``. // before, ~220 lines in each of the 38 ``<Cls>FromSequence`` while (true) { error = SkipWhitespace(reader); if (error.has_value()) { ...6 lines... } if (reader.node().kind() == xml_common::NodeKind::Stop) { break; } else if (reader.node().kind() != xml_common::NodeKind::Start) { ...10 lines naming IExtension... } const std::string name(...); reader.Read(); auto it = properties::kMapOfExtension.find(name); if (it == properties::kMapOfExtension.end()) { ...12 lines... } switch (it->second) { ... } ...~60 lines of prepending and stop-element handling... } // after common::optional<DeserializationError> error( ReadProperties< properties::kPropertyCountOfExtension >( reader, properties::kMapOfExtension, L"IExtension", [&]( properties::OfExtension property ) -> common::optional<DeserializationError> { switch (property) { ... } } ) ); The ``switch`` stays inline because its branches assign the locals the constructor is called with -- the very reason C#, Go and Java had to keep the whole loop inline. The lambda is a template argument taken by ``const&``, so the call is statically bound and nothing is type-erased or allocated; ``g++ -O2`` keeps ``ReadProperties`` itself out of line, one instantiation per class, which is as many function bodies as the loops it replaces. **The duplicate check** is bookkeeping of the loop, not of the property, so it moved there, as a bit set sized to the class. It still precedes the read, so a duplicate is refused without its content ever being looked at. std::bitset<kPropertyCount> seen; ... if (seen[index]) { error = DuplicatePropertyError(name); PrependElementSegmentToDeserializationError(name, *error); return error; } seen[index] = true; **``ReadInto`` takes the results of a read, not the reader.** A tuple's item readers vary in number *and* in type, so no reader-taking signature serves them all, while one taking the result serves every read. template <typename T> common::optional<DeserializationError> ReadInto( common::optional<T>& target, std::pair< common::optional<T>, common::optional<DeserializationError> >&& read ); **A named union is framed like a class.** The two framings were ~160 near-identical lines differing only in the value type, which every error factory they call is already generic in. // before template <typename T, typename DispatchT> ... DeserializeClassFromElement(reader, const std::wstring&, dispatch); template <typename VariantT, typename DispatchT> ... DeserializeUnionFromElement(reader, const std::wstring&, dispatch); // after template <typename ValueT, typename DispatchT> ... DeserializeFromElement(reader, const wchar_t*, dispatch); **An element with a sole model type does not dispatch.** 40 of the 51 ``<Cls>FromElement`` spelled out a one-case ``switch`` inside a fifteen-line lambda signature; the other 11 keep theirs. return DeserializeSoleFromElement< std::shared_ptr<types::IExtension> >( reader, L"IExtension", types::ModelType::kExtension, ExtensionFromSequence<types::IExtension> ); **The public ``<Cls>From``** were 51 copies of the same 54 lines: open the reader, read the root element, check that nothing but whitespace follows. return DeserializeFrom< std::shared_ptr<types::IExtension> >( is, options, ExtensionFromElement ); **``interface_name`` is a ``const wchar_t*``**, not a ``const std::wstring&``: every call site used to construct a ``std::wstring`` on *every* element read, success path included, for a message which is almost never built. Written for C++11, which the library targets outside of its tests: no generic lambda, no deduced return type, every lambda spells out its parameters and its return type. ``xmlization.cpp`` for the AAS meta-model goes 34,543 -> 24,634 lines and its read region 23,584 -> 13,675 (-42%). The ``<Cls>FromSequence`` band goes 13,400 -> 6,076 over 37 classes (-54.7%), with ``Key`` 257 -> 95 and ``Extension`` 331 -> 141, and the public ``<Cls>From`` 2,751 -> 813 (-70%). ``unions`` goes 7,950 -> 5,795; ``empty_class`` goes 1,622 -> 1,674, since one property-less class pays for the shared loop with nothing to save. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
mristin
added a commit
that referenced
this pull request
Sep 22, 2026
C++ is the last backend to have its XML read side decomposed, after C# (#686), Go (#688), Java (#690), Python (#694) and TypeScript (#698). #673 factored out the framing of an *element*; this factors out the framing of the property *loop*, which every ``<Cls>FromSequence`` re-emitted in full. **A property** loses its duplicate check and its ``std::tie``, and names only itself and what reads it. // before case properties::OfExtension::kSemanticId: { if (the_semantic_id.has_value()) { error = DuplicatePropertyError(name); break; } std::tie( the_semantic_id, error ) = ReferenceFromSequence< types::IReference >(reader); break; } // after case properties::OfExtension::kSemanticId: return ReadInto( the_semantic_id, ReferenceFromSequence< types::IReference >(reader) ); **The loop** around it becomes one call. It could be lifted out at all only because the framing now returns the error alone, instead of ``NoInstanceAndDeserializationError<std::shared_ptr<T> >``, and so no longer needs ``T``. // before, ~220 lines in each of the 38 ``<Cls>FromSequence`` while (true) { error = SkipWhitespace(reader); if (error.has_value()) { ...6 lines... } if (reader.node().kind() == xml_common::NodeKind::Stop) { break; } else if (reader.node().kind() != xml_common::NodeKind::Start) { ...10 lines naming IExtension... } const std::string name(...); reader.Read(); auto it = properties::kMapOfExtension.find(name); if (it == properties::kMapOfExtension.end()) { ...12 lines... } switch (it->second) { ... } ...~60 lines of prepending and stop-element handling... } // after common::optional<DeserializationError> error( ReadProperties< properties::kPropertyCountOfExtension >( reader, properties::kMapOfExtension, L"IExtension", [&]( properties::OfExtension property ) -> common::optional<DeserializationError> { switch (property) { ... } } ) ); The ``switch`` stays inline because its branches assign the locals the constructor is called with. The lambda is a template argument taken by ``const&``, so the call is statically bound and nothing is type-erased or allocated; ``g++ -O2`` keeps ``ReadProperties`` itself out of line, one instantiation per class, which is as many function bodies as the loops it replaces. **The duplicate property element check** is bookkeeping of the loop, not of the property, so it moved there, as a bit set sized to the class. It still precedes the read, so a duplicate is refused without its content ever being looked at. std::bitset<kPropertyCount> seen; ... if (seen[index]) { error = DuplicatePropertyError(name); PrependElementSegmentToDeserializationError(name, *error); return error; } seen[index] = true; **``ReadInto`` takes the results of a read, not the reader.** A tuple's item readers vary in number *and* in type, so no reader-taking signature serves them all, while one taking the result serves every read. template <typename T> common::optional<DeserializationError> ReadInto( common::optional<T>& target, std::pair< common::optional<T>, common::optional<DeserializationError> >&& read ); **A named union is framed like a class.** The two framings were ~160 near-identical lines differing only in the value type, which every error factory they call is already generic in. // before template <typename T, typename DispatchT> ... DeserializeClassFromElement(reader, const std::wstring&, dispatch); template <typename VariantT, typename DispatchT> ... DeserializeUnionFromElement(reader, const std::wstring&, dispatch); // after template <typename ValueT, typename DispatchT> ... DeserializeFromElement(reader, const wchar_t*, dispatch); **An element with a sole model type does not dispatch.** 40 of the 51 ``<Cls>FromElement`` spelled out a one-case ``switch`` inside a fifteen-line lambda signature; the other 11 keep theirs. return DeserializeSoleFromElement< std::shared_ptr<types::IExtension> >( reader, L"IExtension", types::ModelType::kExtension, ExtensionFromSequence<types::IExtension> ); **The public ``<Cls>From``** were 51 copies of the same 54 lines: open the reader, read the root element, check that nothing but whitespace follows. return DeserializeFrom< std::shared_ptr<types::IExtension> >( is, options, ExtensionFromElement ); **``interface_name`` is a ``const wchar_t*``**, not a ``const std::wstring&``: every call site used to construct a ``std::wstring`` on *every* element read, success path included, for a message which is almost never built. ``xmlization.cpp`` for the AAS meta-model goes 34,543 -> 24,634 lines and its read region 23,584 -> 13,675 (-42%).
mristin
added a commit
that referenced
this pull request
Sep 22, 2026
C++ is the last backend to have its XML read side decomposed, after C# (#686), Go (#688), Java (#690), Python (#694) and TypeScript (#698). #673 factored out the framing of an *element*; this factors out the framing of the property *loop*, which every ``<Cls>FromSequence`` re-emitted in full. **A property** loses its duplicate check and its ``std::tie``, and names only itself and what reads it. // before case properties::OfExtension::kSemanticId: { if (the_semantic_id.has_value()) { error = DuplicatePropertyError(name); break; } std::tie( the_semantic_id, error ) = ReferenceFromSequence< types::IReference >(reader); break; } // after case properties::OfExtension::kSemanticId: return ReadInto( the_semantic_id, ReferenceFromSequence< types::IReference >(reader) ); **The loop** around it becomes one call. It could be lifted out at all only because the framing now returns the error alone, instead of ``NoInstanceAndDeserializationError<std::shared_ptr<T> >``, and so no longer needs ``T``. // before, ~220 lines in each of the 38 ``<Cls>FromSequence`` while (true) { error = SkipWhitespace(reader); if (error.has_value()) { ...6 lines... } if (reader.node().kind() == xml_common::NodeKind::Stop) { break; } else if (reader.node().kind() != xml_common::NodeKind::Start) { ...10 lines naming IExtension... } const std::string name(...); reader.Read(); auto it = properties::kMapOfExtension.find(name); if (it == properties::kMapOfExtension.end()) { ...12 lines... } switch (it->second) { ... } ...~60 lines of prepending and stop-element handling... } // after common::optional<DeserializationError> error( ReadProperties< properties::kPropertyCountOfExtension >( reader, properties::kMapOfExtension, L"IExtension", [&]( properties::OfExtension property ) -> common::optional<DeserializationError> { switch (property) { ... } } ) ); The ``switch`` stays inline because its branches assign the locals the constructor is called with. The lambda is a template argument taken by ``const&``, so the call is statically bound and nothing is type-erased or allocated; ``g++ -O2`` keeps ``ReadProperties`` itself out of line, one instantiation per class, which is as many function bodies as the loops it replaces. **The duplicate property element check** is bookkeeping of the loop, not of the property, so it moved there, as a bit set sized to the class. It still precedes the read, so a duplicate is refused without its content ever being looked at. std::bitset<kPropertyCount> seen; ... if (seen[index]) { error = DuplicatePropertyError(name); PrependElementSegmentToDeserializationError(name, *error); return error; } seen[index] = true; **``ReadInto`` takes the results of a read, not the reader.** A tuple's item readers vary in number *and* in type, so no reader-taking signature serves them all, while one taking the result serves every read. template <typename T> common::optional<DeserializationError> ReadInto( common::optional<T>& target, std::pair< common::optional<T>, common::optional<DeserializationError> >&& read ); **A named union is framed like a class.** The two framings were ~160 near-identical lines differing only in the value type, which every error factory they call is already generic in. // before template <typename T, typename DispatchT> ... DeserializeClassFromElement(reader, const std::wstring&, dispatch); template <typename VariantT, typename DispatchT> ... DeserializeUnionFromElement(reader, const std::wstring&, dispatch); // after template <typename ValueT, typename DispatchT> ... DeserializeFromElement(reader, const wchar_t*, dispatch); **An element with a sole model type does not dispatch.** 40 of the 51 ``<Cls>FromElement`` spelled out a one-case ``switch`` inside a fifteen-line lambda signature; the other 11 keep theirs. return DeserializeSoleFromElement< std::shared_ptr<types::IExtension> >( reader, L"IExtension", types::ModelType::kExtension, ExtensionFromSequence<types::IExtension> ); **The public ``<Cls>From``** were 51 copies of the same 54 lines: open the reader, read the root element, check that nothing but whitespace follows. return DeserializeFrom< std::shared_ptr<types::IExtension> >( is, options, ExtensionFromElement ); **``interface_name`` is a ``const wchar_t*``**, not a ``const std::wstring&``: every call site used to construct a ``std::wstring`` on *every* element read, success path included, for a message which is almost never built. ``xmlization.cpp`` for the AAS meta-model goes 34,543 -> 24,634 lines and its read region 23,584 -> 13,675 (-42%).
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
This is the Go counterpart of #686, but the C# design does not carry over: the shape that works there fails to compile in Go, so the read side is composed out of generated dispatchers instead of composed readers.
Measured on the AAS meta-model fixture, the read region of
xmlization.gogoes 11,738 -> 5,584 lines (-52%), and the file as a whole 22,846 -> 16,692. Theread<X>WithLookaheadband -- 50 functions, 3,057 lines -- goes to zero; 20read<X>Dispatchedfunctions stand in its place.The whole design turns on one asymmetry in the wire format. An instance element is self-describing: its local name is its type. A scalar element is not: its name is its position,
vin a list andv1,v2, ... in a tuple. So: every instance is read by dispatch on its own name, every scalar by a name its container supplies, and nothing else in the read path knows the difference.readElementDispatchedis the only place which frames an XML element around a value.readListOfand everyreadTupleNdelegate the framing to it, which is what deletesreadListOfScalars,readScalarWithName,asScalarTupleItemReaderandasInstanceTupleItemReaderoutright: one list helper, one item shape, no adapters. It also returns the token past the end element, so the "read<X>WithLookaheadstops at the end element, so we look ahead" comment and its hand-writtenreadNextvanish from every call site:Per class,
read<X>AsSequencehas for every property case one assignment of the same(value, current, valueErr)triple:readOptionalabsorbs thevar value T; ...; theX = &valuedance. It takes the results of a read rather than the reader, so Go's multi-value pass-through lets one helper serve any read however many arguments that read takes -- includingreadOptional(readTuple2(...)), which a reader-taking signature cannot express, since a tuple's item readers vary in number and in type. That also fixes an optional tuple property, which is a pointer (a Go struct is not nilable) and used to be assigned a non-pointerreadTupleNresult: output that does not compile.