Compose the Java XML serialization (#691) - #691
Merged
Merged
Conversation
Write-side dual of #690, as #687 was of #686 in C# and #689 of #688 in Go. ``visit{Cls}`` was a verbatim copy of ``serializeElement``: 38 methods of 18 lines each spelling out the start tag, the namespace, the content and the end tag, around a call which already did exactly that. Six near-identical emitters -- one per property kind -- produced the very same ``serializeElement`` call, differing only in the content serializer. And everything was an *instance* method, because of one mutable ``topLevel`` flag. That flag is what cost the most. It made ``this::writeStringifiedContent``, ``this::visit`` and ``(value, w) -> this.xToSequence(value, w)`` capturing, and ``serializeItems(...)`` and ``asNamedElementSerializer(...)`` returned capturing lambdas on top of them, so writing one instance allocated a small pile of closures. It is replaced by a ``final boolean withNamespace`` pinned in the constructor of two immutable visitors, ``ROOT`` and ``NESTED``: only the outermost element declares the namespace, and an element written from ``writeClass`` is nested by construction. Everything else is then ``static``, and a method reference to a static method captures nothing. The design is otherwise the read side's, mirrored: * One shape, not two. ``ContentWriter<T>`` writes a value where the writer already is; ``writeElement`` frames it in a start and an end tag. There is deliberately no second delegate for a whole element, as there is on the reading side -- an element differs from a content only in what it writes, never in its shape. The reading needs the distinction because a content reader has to be told whether its element was self-closing, and a writer has nothing to be told. Use sites take ``ContentWriter<? super T>``, which is how Java spells the contravariance C# writes as ``in T``, so the single writer of an ``IClass`` serves wherever the writer of a more specific interface is expected. * ``visit{Cls}`` collapses to one call, and the visitor becomes a dispatch shim:: @OverRide public void visitQualifier( IQualifier that, XMLStreamWriter writer) { writeElement( "qualifier", that, writer, withNamespace, _VisitorWithWriter::writeQualifierAsSequence); } * The families are dual to the readers and named by the same moniker grammar: ``write{Cls}AsSequence``, ``writeClass``/``writeUnion``, ``writeListOf_{M}``, ``writeTupleOf{N}_{M...}``, ``writeAtV_{M}`` and ``writeAtV{i}_{M}``. The list and the tuple writers inline their loop, which retires ``serializeItems``, ``asNamedElementSerializer`` and ``serializeTuple{N}`` -- Java needs no partial application here, since the name to bind is fixed per generated function. * One property emitter replaces the six. Every property is written by ``writeProperty``, or by ``writeOptionalProperty``, which takes the ``Optional`` apart once instead of asking it and then unwrapping it -- the old form called the getter twice, and every call allocates an ``Optional``. * ``_type_moniker``, ``_leaf_moniker``, ``_is_instance_type``, ``_is_dispatched`` and ``_collect_needed`` are reused unchanged: the reading and the writing are composed out of the same shapes, so one gating pass answers for both. Only the two dispatching writers needed a walk of their own, because the *position* matters on the write side while it does not on the read side: an item of any class is written as its own element, whereas a property of a concrete class without descendants writes its properties straight into the property element. Two things are genuinely asymmetric, and both save work: * There is no ``writeTextAs_{primitive}`` family. Once the value is at hand nothing type-specific is left to do with it, so the four ``toString`` primitives share ``writeStringifiedContent`` and the byte array ``writeByteArrayContent``. An enumeration does have something of its own to write and keeps one function each -- named ``writeTextAs_{Enum}`` and not ``write{Enum}Content``, so that it stays in the moniker-keyed name space: were it keyed on the symbol, an enumeration which somebody names ``Stringified`` or ``ByteArray`` would silently take a shared writer's name. * There is no element writer per class. Which element an instance writes follows from its run-time type, and one virtual call answers that for every class at once. The reading needs a dispatcher per interface because it has to decide what to construct before it has read anything. Two defects of the write side are fixed along the way, both of which the decomposition makes one-line changes rather than sweeping ones. ``Serialize.to`` never flushed, so the serialization was not finished when it returned, and a failure of the underlying stream surfaced at the caller's own flush as a raw ``XMLStreamException`` -- outside the one method whose ``throws`` clause promises a ``SerializeException``. Probed with a ``Writer`` which accepts every character and fails on ``flush``, ``to`` used to return normally and now throws. The flush goes in ``Serialize.to`` and nowhere else, as in Go's #689, so ``XMLStreamWriter`` keeps buffering in between. Mind that this does push the underlying stream: a caller serializing many small documents into one buffered writer now flushes once per document rather than once per buffer. That is the right trade, since the alternative leaves the output silently incomplete whenever the caller forgets. Every ``SerializeException`` carried an empty path, so every failure read ``... at: the beginning`` no matter where it happened. The path is now built as the stack unwinds, each container prepending the one segment it knows, exactly as on the read side where ``read<Cls>FromSequence`` prepends the property after its ``switch`` and ``readList`` prepends the index. Probed with an ``XMLStreamWriter`` which refuses to write one chosen value: a property of the root instance idShort nested through two lists submodels/*[1]/submodelElements/*[1]/value through a class-typed property semanticId/keys/*[0]/value ``SerializeException`` renders its message in its constructor, so its path has to be complete by the time it is built and it can not be the thing thrown from inside; a private ``_SerializeFailure`` carries a ``Reporting.Error`` instead, and ``Serialize.to`` renders that with ``Reporting.generateRelativeXPath`` and converts. The public exception is untouched. ``writeElement`` contributes no segment and ``writeProperty`` is ``writeElement`` plus the one thing a property knows which nothing below it does -- its own name -- so there is no ``try`` around any of the 276 property writes. A list and a tuple put their ``try`` outside the loop and advance the index only once an item has been written, so the item named is the one which failed and nothing is paid per item. Two segments are deliberately absent, and the reasoning is written into the generated Javadoc, since both look like oversights. The outermost element is not named, because ``Serialize.to`` takes any ``IClass`` and the name would say nothing the caller does not already know; the read side does prepend it, having one facade per class. Neither is the discriminator element of a polymorphic property, which the read side does prepend -- for a failure in an embedded data specification it reports ``submodel/embeddedDataSpecifications/*[0]/dataSpecificationContent/`` ``dataSpecificationIec61360/preferredName/*[0]`` where the write side reports ``embeddedDataSpecifications/*[0]/dataSpecificationContent/`` ``preferredName/*[0]/text``. The reading points into a document it is parsing, where that element is a real extra level; the writing points into the instance the caller handed over, where it is not, and ``value/idShort`` there is exactly ``getValue().getIdShort()``. Go's #689 made the same choice when it rendered its serialization paths through ``aasreporting.ToGolangPath`` rather than as XPaths. For the AAS meta-model the write region of ``Xmlization.java`` goes 3,281 -> 2,873 lines (-12.4%), error paths included. The write region no longer contains a single lambda. Unlike the read side, that shows up in the measurement. Serializing an ``Environment`` of 200 submodels (1.14 MB of XML) into a writer which counts the characters and discards them, 100 runs after 100 warm-up runs, eight rounds of each variant run alternately: the allocation goes 0.76 -> 0.36 MB per serialization, and the time 7.69 -> 7.49 ms, which is inside the round-to-round noise. ``XMLStreamWriter`` writes straight through, so the whole write path allocates ~0.8 MB against the ~49 MB the *reading* of the same document allocates -- there is no StAX floor here for the closures to hide behind, which is precisely why #690 could not measure its own. The snippet contract for an implementation-specific class changes: ``Xmlization/VisitorWithWriter/visit_{cls}.java`` is no longer read, as the element around the sequence is now written identically for every class, and the remaining ``{cls}_to_sequence.java`` has to define a ``static`` ``write{Cls}AsSequence``.
Coverage Report for CI Build 34977000804Coverage increased (+0.02%) to 85.322%Details
Uncovered Changes
Coverage RegressionsNo coverage regressions found. Coverage Stats
💛 - Coveralls |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Write-side dual of #690, as #687 was of #686 in C# and #689 of #688 in Go.
visit{Cls}was a verbatim copy ofserializeElement: 38 methods of 18 lines each spelling out the start tag, the namespace, the content and the end tag, around a call which already did exactly that. Six near-identical emitters -- one per property kind -- produced the very sameserializeElementcall, differing only in the content serializer. And everything was an instance method, because of one mutabletopLevelflag.That flag is what cost the most. It made
this::writeStringifiedContent,this::visitand(value, w) -> this.xToSequence(value, w)capturing, andserializeItems(...)andasNamedElementSerializer(...)returned capturing lambdas on top of them, so writing one instance allocated a small pile of closures. It is replaced by afinal boolean withNamespacepinned in the constructor of two immutable visitors,ROOTandNESTED: only the outermost element declares the namespace, and an element written fromwriteClassis nested by construction. Everything else is thenstatic, and a method reference to a static method captures nothing.The design is otherwise the read side's, mirrored:
One shape, not two.
ContentWriter<T>writes a value where the writer already is;writeElementframes it in a start and an end tag. There is deliberately no second delegate for a whole element, as there is on the reading side -- an element differs from a content only in what it writes, never in its shape. The reading needs the distinction because a content reader has to be told whether its element was self-closing, and a writer has nothing to be told. Use sites takeContentWriter<? super T>, which is how Java spells the contravariance C# writes asin T, so the single writer of anIClassserves wherever the writer of a more specific interface is expected.visit{Cls}collapses to one call, and the visitor becomes a dispatch shim::The families are dual to the readers and named by the same moniker grammar:
write{Cls}AsSequence,writeClass/writeUnion,writeListOf_{M},writeTupleOf{N}_{M...},writeAtV_{M}andwriteAtV{i}_{M}. The list and the tuple writers inline their loop, which retiresserializeItems,asNamedElementSerializerandserializeTuple{N}-- Java needs no partial application here, since the name to bind is fixed per generated function.One property emitter replaces the six. Every property is written by
writeProperty, or bywriteOptionalProperty, which takes theOptionalapart once instead of asking it and then unwrapping it -- the old form called the getter twice, and every call allocates anOptional._type_moniker,_leaf_moniker,_is_instance_type,_is_dispatchedand_collect_neededare reused unchanged: the reading and the writing are composed out of the same shapes, so one gating pass answers for both. Only the two dispatching writers needed a walk of their own, because the position matters on the write side while it does not on the read side: an item of any class is written as its own element, whereas a property of a concrete class without descendants writes its properties straight into the property element.Two things are genuinely asymmetric, and both save work:
There is no
writeTextAs_{primitive}family. Once the value is at hand nothing type-specific is left to do with it, so the fourtoStringprimitives sharewriteStringifiedContentand the byte arraywriteByteArrayContent. An enumeration does have something of its own to write and keeps one function each -- namedwriteTextAs_{Enum}and notwrite{Enum}Content, so that it stays in the moniker-keyed name space: were it keyed on the symbol, an enumeration which somebody namesStringifiedorByteArraywould silently take a shared writer's name.There is no element writer per class. Which element an instance writes follows from its run-time type, and one virtual call answers that for every class at once. The reading needs a dispatcher per interface because it has to decide what to construct before it has read anything.
Two defects of the write side are fixed along the way, both of which the decomposition makes one-line changes rather than sweeping ones.
Serialize.tonever flushed, so the serialization was not finished when it returned, and a failure of the underlying stream surfaced at the caller's own flush as a rawXMLStreamException-- outside the one method whosethrowsclause promises aSerializeException. Probed with aWriterwhich accepts every character and fails onflush,toused to return normally and now throws. The flush goes inSerialize.toand nowhere else, as in Go's #689, soXMLStreamWriterkeeps buffering in between. Mind that this does push the underlying stream: a caller serializing many small documents into one buffered writer now flushes once per document rather than once per buffer. That is the right trade, since the alternative leaves the output silently incomplete whenever the caller forgets.Every
SerializeExceptioncarried an empty path, so every failure read... at: the beginningno matter where it happened. The path is now built as the stack unwinds, each container prepending the one segment it knows, exactly as on the read side whereread<Cls>FromSequenceprepends the property after itsswitchandreadListprepends the index. Probed with anXMLStreamWriterwhich refuses to write one chosen value:SerializeExceptionrenders its message in its constructor, so its path has to be complete by the time it is built and it can not be the thing thrown from inside; a private_SerializeFailurecarries aReporting.Errorinstead, andSerialize.torenders that withReporting.generateRelativeXPathand converts. The public exception is untouched.writeElementcontributes no segment andwritePropertyiswriteElementplus the one thing a property knows which nothing below it does -- its own name -- so there is notryaround any of the 276 property writes. A list and a tuple put theirtryoutside the loop and advance the index only once an item has been written, so the item named is the one which failed and nothing is paid per item.Two segments are deliberately absent, and the reasoning is written into the generated Javadoc, since both look like oversights. The outermost element is not named, because
Serialize.totakes anyIClassand the name would say nothing the caller does not already know; the read side does prepend it, having one facade per class. Neither is the discriminator element of a polymorphic property, which the read side does prepend -- for a failure in an embedded data specification it reportssubmodel/embeddedDataSpecifications/*[0]/dataSpecificationContent/dataSpecificationIec61360/preferredName/*[0]where the write side reportsembeddedDataSpecifications/*[0]/dataSpecificationContent/preferredName/*[0]/text. The reading points into a document it is parsing, where that element is a real extra level; the writing points into the instance the caller handed over, where it is not, andvalue/idShortthere is exactlygetValue().getIdShort(). Go's #689 made the same choice when it rendered its serialization paths throughaasreporting.ToGolangPathrather than as XPaths.For the AAS meta-model the write region of
Xmlization.javagoes 3,281 -> 2,873 lines (-12.4%), error paths included. The write region no longer contains a single lambda.Unlike the read side, that shows up in the measurement. Serializing an
Environmentof 200 submodels (1.14 MB of XML) into a writer which counts the characters and discards them, 100 runs after 100 warm-up runs,eight rounds of each variant run alternately: the allocation goes 0.76 ->
0.36 MB per serialization, and the time 7.69 -> 7.49 ms, which is inside the round-to-round noise.
XMLStreamWriterwrites straight through, so the whole write path allocates ~0.8 MB against the ~49 MB the reading of the same document allocates -- there is no StAX floor here for the closures to hide behind, which is precisely why #690 could not measure its own.The snippet contract for an implementation-specific class changes:
Xmlization/VisitorWithWriter/visit_{cls}.javais no longer read, as the element around the sequence is now written identically for every class, and the remaining{cls}_to_sequence.javahas to define astaticwrite{Cls}AsSequence.