Skip to content

feat: fork the OpenSearch classes this client needs into org.codelibs.fesen.opensearch - #43

Merged
marevol merged 38 commits into
mainfrom
fork/opensearch-import
Sep 13, 2026
Merged

marevol merged 38 commits into
mainfrom
fork/opensearch-import

Conversation

@marevol

@marevol marevol commented Sep 13, 2026 •

Copy link
Copy Markdown
Contributor

Fess ships org.opensearch:opensearch 3.8.0 — 17.09 MiB, 8,578 classes — purely for its request/response DTOs and query builders. It has never run an OpenSearch node; the embedded runner went away in 15.9. This forks the classes it needs into org.codelibs.fesen.opensearch.* and prunes the rest.

The forked jar is 4.69 MiB against the 17.09 MiB it replaces. Across the Fess distribution that is 23.28 MiB off WEB-INF/lib, 229 jars down to 203. Apache Lucene stays as the real artifact. All 101 Http*Action classes are intact — this is still a general-purpose HTTP implementation of the OpenSearch client, not a Fess-private adapter.

before after
classes 8,578 2,491
source files 5,435 1,574
LOC 976,951 275,304
third-party dependencies 31 17

Why not shrink the jar mechanically instead

Both cheaper options were measured first.

Class-level shrinking is useless: the reference graph is one blob, and a single seed class — SearchResponse, QueryBuilders, anything — drags in 95% of the jar. Even emptying every method body leaves 59.6% standing.

Member-level tree shaking with ProGuard does better — 8.45 MiB — but produces something that cannot ship. It silently drops 6 of the 10 META-INF/services providers, including XContentOpenSearchExtension and ServerCompressorProvider, because the ServiceLoader.load calls that need them live in opensearch-core, a library jar it does not trace into; XContentBuilder then fails on every request. It removes enum constants until Setting$Property is no longer an enum and FeatureFlags dies in its static initialiser. And it deletes public DTO methods production code happens not to call: 74 of this repository's own unit tests fail with NoSuchMethodError on things like BulkRequest.pipeline. Fess has a plugin API, and third-party plugins cannot be in anyone's root set. A fork moves those failures to compile time, where they belong.

What was cut

Node-side machinery, in layers, each severed at a load-bearing reference:

  • ClusterSettings — HttpNodesStatsAction instantiated three BUILT_IN_* registries only to build a SearchRequestStats. Those arrays name Setting constants in almost every server package, which is what kept gateway, plugins, telemetry, repositories, node, persistent, identity, discovery and http alive. Behaviour-neutral: the value replaced round-tripped through a constructor that pre-populates every field with zeros.
  • rest/** — 181 files, kept alive by four response DTOs reading one constant, RestSearchAction.TOTAL_HITS_AS_INT_PARAM, inside toXContent.
  • QueryShardContext — AbstractQueryBuilder.doToQuery and 48 overrides. A client serialises query builders to JSON; it never compiles them into Lucene queries.
  • The aggregation value-source path — what actually collects index/mapper/**.
  • META-INF/services — all six files gone. Three deleted outright (ErrorOnUnknown only reworded an error; IsoCalendarDataProvider registered a JVM-wide java.util.Calendar override that has been inert since JDK 9 made CLDR,COMPAT the default — verified byte-identical output across fourteen week-based formats; our PostingsFormat file duplicated lucene-suggest's). The other three became static registration, which keeps the behaviour and removes the failure mode that bit this project twice: a provider silently pruned, build green, registry empty at runtime.
  • The binary TopDocs transport format — Lucene.readTopDocs/writeTopDocs, uncalled, taking CollapseTopFieldDocs and TopDocsAndMaxScore with it. org/apache/lucene is now empty and gone, one less split package against lucene-core.
  • Dozens of smaller cases of the same shape, found by ranking every class by bytes that die with it ÷ members holding it alive: Transports held by one assert; JarHell (14 KB) by one static version-format check; MoreLikeThisQuery + XMoreLikeThis (23 KB) by seven copied constants; AbstractClient (185 KB — OpenSearch's node-side Client) by the twelve-line filterWithHeader, which now wraps HttpAbstractClient instead.

Dependencies removed (4.34 MiB off the distribution): log4j-core, jts-core, spatial4j, RoaringBitmap, HdrHistogram, t-digest, jopt-simple, jzlib, jsr305, and four Lucene modules. zstd-jni also leaves this pom, though it stays in the war via commons-compress. joda-time stays deliberately: replacing Joda.forPattern in Fess's taglib changes date parsing for admin-configured patterns — hh without a zeroes the time, ww is ignored, YYYY-'W'ww-e shifts a day — and silently, so 0.61 MiB is the price of not doing that.

Six methods are stubbed with UnsupportedOperationException, all coordinator-side reduce paths a client never reaches. Nothing else is stubbed.

Licensing

Forked sources keep their Apache-2.0 headers; license-maven-plugin and formatter-maven-plugin are scoped to this project's own sources so they cannot rewrite them. META-INF/NOTICE.txt carries a derivation statement above OpenSearch's own NOTICE; META-INF/LICENSE.txt is verbatim.

Verification

  • 373 tests, 0 failures. No test deleted or weakened.
  • CI, against real servers: OpenSearch 3 91, OpenSearch 2 45, OpenSearch 1 41, Elasticsearch 8 43, Elasticsearch 7 43 — 636 integration tests, 0 failures.
  • A runtime probe of the built jar — the registries, XContentBuilder extension writers, and serialisation of query, aggregation, sort, highlight, suggest and collapse builders. Static reachability alone had deleted the ServiceLoader providers with the build still green; this is what caught it.
  • A behaviour snapshot feeding a _nodes/stats fixture through HttpNodesStatsAction — byte-identical to the pre-fork output through every round, which is what proves the SearchRequestStats removal changed nothing.
  • End to end: a Fess distribution built on this, booted against a real OpenSearch 3.8.0 from a virgin cluster — 40 indices created, a crawl through the crawler child process, search with correct term and filter facet counts, suggest indexing and alias switch, _nodes/stats with all 30 sections. Zero NoClassDefFoundError, ClassNotFoundException, NoSuchMethodError or ServiceConfigurationError across all 13 log files.
  • All eight downstream repositories green: fess 7,469, fess-suggest 689, fess-crawler-opensearch 53, and the five plugins 1,198.

Dead members

A member-level reachability pass with the full consumer root set found 7,436 dead members inside classes that must stay. 4,577 of them are gone, along with 174 StreamInput constructors that had no referrer at all — a DTO rebuilt from the binary transport format is of no use to a client that speaks JSON — and 57 whole files plus 48 nested types that nothing reached.

The 2,859 declined each have a reason, and the four largest are worth recording because they are ways bytecode reachability over-reports against source:

  • 738 compile-time constants. javac inlines the read, so the reference is invisible in the class file while the source still needs the field.
  • 418 @Override plus 81 overridden elsewhere, and 345 implicit constructors with no source to delete.
  • 166 marker interfaces named in an implements clause — CompositeIndicesRequest, RealtimeRequest, WriteResponse. ProGuard prunes an interface list; source cannot without changing what a type implements.
  • 7 annotation types, which are outright false positives: nothing references @PublicApi, but 587 files are annotated with it.

The raw list also wanted 56 Client/IndicesAdminClient default methods — this library's own API — which is why it was treated as a starting point and not an instruction.

What remains and stays: writeTo (679 implementations, held by one virtual call on Writeable.writeTo), the server-side reduce members and the Version-gated overrides. Removing those means changing what a type implements or giving an interface a default that throws, which trades a compile-time guarantee for a runtime failure. That is a design change, not pruning, and it is deliberately not in this PR.

Across the whole of this work, 139 implements/extends lines are touched and every one is a deletion — not a single clause was modified, no default or stub was added, and no signature was relaxed.

Ordering

Depends on #42. Downstream repositories — fess-suggest, fess-crawler, fess and five plugin repositories — have matching branches and must be released together with this.

Shinsuke Sugaya and others added 9 commits September 13, 2026 06:14
Import the org.codelibs.fesen.opensearch.* fork of OpenSearch 3.8.0 into
src/main/java, carry the needed resources (merged META-INF/services,
task-index-mapping.json), migrate the client sources to the forked
package, and delete the node-only trees.
…it compile

Keep only the classes reachable from what fesen-httpclient's own sources use,
drop the node-extension and bootstrap trees (they need protobuf, JNA and the
opensearch-agent-bootstrap module), drop the reactive HTTP streaming path
(it needs reactor-core), and drop the vectorised B-tree rounding searcher
(it needs --add-modules jdk.incubator.vector).
Declare the third-party libraries the forked sources actually compile
against, restore the ServiceLoader provider implementations that static
reachability cannot see, scope the license and formatter plugins to the
client sources only, and ship the OpenSearch NOTICE/LICENSE with a
statement of the CodeLibs modifications.
Node keeps only the settings constants other classes reference; its
construction, lifecycle and dependency-injection wiring are gone.
With that, common/inject (the bundled Guice fork) is removed except the
one-method Provider interface that two aggregation/analysis types use,
and the @Inject annotations and Guice binding methods that referenced it
are removed from the modules and plugin extension points.
Drop lucene-spatial3d (nothing references it) and declare the JUnit 4
assertions the existing tests import directly instead of relying on the
Testcontainers transitive.
Only the Jackson streaming API is used, so exclude the databind
transitive from the three dataformat modules, as OpenSearch does.
fesen-httpclient only ever speaks HTTP to a remote cluster, so everything that
exists to run an OpenSearch node is dead weight. Sever the four bridges that
kept it reachable and drop what falls out:

* ClusterSettings / IndexScopedSettings / FeatureFlagSettings / SettingsModule /
  AbstractScopedSettings. HttpNodesStatsAction no longer builds a node-scoped
  ClusterSettings just to construct a SearchRequestStats whose counters are all
  zero; NodeIndicesStats gains a constructor without it. The rendered
  _nodes/stats output is byte-identical.
* rest/**. The four request parameters SearchHit, SearchHits, Suggest and
  InternalAggregation read are inlined with their exact string values, the
  _shards header renderer moves to
  action.support.broadcast.BroadcastShardsHeader, AliasesNotFoundException
  moves to action.admin.indices.alias (same simple name, so same wire and
  XContent identity), and the error-body decoder moves into HttpAction.
* QueryShardContext: doToQuery/toQuery/toFilter on all 48 query builders, plus
  InnerHitContextBuilder and the shard-side doRewrite branches.
* The aggregation value-source bridge: AggregationBuilder.build, innerBuild,
  ValuesSourceConfig and AggregatorFactories.create*, which is what kept
  index/mapper/** alive.

Classes that survive only for a nested request type (Aggregator, RangeAggregator,
TermsAggregator, FiltersAggregator, AdjacencyMatrixAggregator, HeapUsageTracker,
NativeMemoryUsageTracker, MatchQuery, FunctionScoreQuery, MultiValueMode,
IndexModule, MapperService, DateFieldMapper, SearchContext and friends) keep that
type and lose the node-side body. Every exception in the OpenSearchServerException
wire registry is kept so error deserialisation is unchanged.

The Lucene Codec and PostingsFormat service entries and the classes behind them
go with index/codec/** and index/compositeindex/**: a client never opens a Lucene
index, and removing the provider names along with the classes leaves no dangling
ServiceLoader entry.

ScriptService.compile, ScriptService.isLangSupported and three coordinator-side
aggregation reduce paths that compile scripts now throw UnsupportedOperationException;
a client never reduces aggregations or compiles scripts.

4604 -> 2401 source files, 931k -> 434k lines, jar 17.51 MB -> 7.09 MB.
370 tests still pass and the _nodes/stats fixture round-trip is unchanged.
HttpCreateIndexAction set is_write_index=false on every alias whose
writeIndex() was unset, so an index created with
prepareCreate(name).addAlias(new Alias(a)) came up with a read-only alias.
Indexing through that alias then failed with

    no write index is defined for alias [a]

even though the alias pointed at a single index, which both the transport
client and the REST API treat as writable. Leave the flag unset unless the
caller set it.

Found while moving fess-suggest onto this client: Suggester.createIndexIfNothing()
creates its search and update aliases exactly this way, and every write through
them failed.
Cross-checking the 452 distinct OpenSearch types that Fess, fess-suggest,
fess-crawler-opensearch and the five plugin repositories name against this fork
left three gaps. Two are real and are restored here.

ScoreFunctionBuilders is the factory for the function-score builders. Every
builder it returns was already here; only the factory itself had been pruned,
because nothing inside this library calls it. Fess calls it from KeyMatchHelper,
QueryHelper and SuggestHelper, and fess-suggest from SuggestQueryBuilder and
PopularWordsRequest.

common/joda carries Joda, JodaDateFormatter, JodaDateMathParser and
JodaDeprecationPatterns. Fess uses Joda.forPattern(...).parseMillis(...) in its
taglib to parse a date string, so the parsing semantics are user-visible and
worth keeping exactly as they are rather than substituting another formatter.

The third gap, QueryShardContext, stays gone. It is the bridge that was cut to
drop the mapper and shard trees, and its only consumers are three Fess query
builders whose doToQuery overrides throw UnsupportedOperationException.

370 tests pass. The jar grows by 23 KB.
Alias.toXContent wrote is_write_index unconditionally while every other
optional field in that method is null-guarded. Two failure modes followed:

  - forcing false made every alias created through HttpCreateIndexAction
    read-only, so indexing through an alias failed with "no write index is
    defined for alias [...]" even when the alias pointed at a single index,
    which the transport client accepts as writable;
  - leaving it unset then rendered is_write_index: null, which Elasticsearch 8
    rejects outright with "Unknown token [VALUE_NULL] in alias [...]".

Omitting the field is the only rendering both engines read as "the caller did
not decide". The fork owns Alias, so the guard goes at the source and
HttpCreateIndexAction keeps delegating to it.

373 tests pass.
Client, ClusterAdminClient and IndicesAdminClient advertised 122 operations.
Only 99 of them have an HTTP action registered in HttpClient; the other 23
reached HttpClient.doExecute, found no handler and threw

    UnsupportedOperationException: Action: <name>

so they were never usable over HTTP. Two independent signals agree exactly on
that set: no consumer references any of them, and none has a registered handler.

The used-set was derived mechanically, not by inspection. For each operation the
scan looked for prepareX( call sites, references to the request/response/builder
types that are unique to it (types shared between operations, such as
AcknowledgedResponse, are not evidence), and references to its ActionType --
across this repository's main and test sources, fess, fess-suggest,
fess-crawler-opensearch and the five plugin repositories.

Removed (23 operations, 153 method declarations across the interfaces,
AbstractClient, HttpAbstractClient and HttpIndicesAdminClient):

  dangling indices  deleteDanglingIndex, importDanglingIndex, listDanglingIndices
  decommission      decommission, deleteDecommissionState, getDecommissionState,
                    prepareDeleteDecommissionRequest
  weighted routing  putWeightedRouting, getWeightedRouting, deleteWeightedRouting,
                    weightedRouting
  data streams      createDataStream, deleteDataStream, getDataStreams
  search pipelines  putSearchPipeline, getSearchPipeline, deleteSearchPipeline
  snapshots         cloneSnapshot, cleanupRepository
  remote store      restoreRemoteStore
  other             addBlock, pitSegments, reloadSecureSettings

Kept every operation that has a handler, plus the structural accessors
admin()/cluster()/indices()/settings()/filterWithHeader()/getRemoteClusterClient().
Surviving signatures are untouched, so downstream code needs no change beyond
deleting overrides of methods that no longer exist.

No behaviour change: every removed method threw on call. 373 tests pass.
…umers

Class-level reachability over the built class files, iterated to convergence.
Edges come from each class file's constant pool (CONSTANT_Class entries plus the
type descriptors in Utf8 entries, so Signature attributes and annotations count).

The root set is deliberately wider than this repository, because a prune rooted
only in org/codelibs/fesen/client/** deletes types that nothing here mentions but
downstream code does, leaving this build green and the downstream one broken:

  * org/codelibs/fesen/client/** and every test class
  * every class named by a META-INF/services provider file
  * every org.codelibs.fesen.opensearch type named by fess, fess-suggest,
    fess-crawler-opensearch and the five plugin repositories, main and test
    sources both, plus any fork class whose simple name appears there at all
    (an over-approximation that covers Outer.Inner written after import Outer)
  * public API reachable only through the named-writeable / named-XContent
    registries -- every *AggregationBuilder, *QueryBuilder, *SuggestionBuilder,
    *ScoreFunctionBuilder, *SortBuilder, *SignificanceHeuristic, *SmoothingModel,
    the MovAvg models, and the Internal*/Parsed* aggregation results. Static
    reachability deletes these with the build still green: they are the library's
    published surface, so they are roots rather than an afterthought. That keep
    covers 234 classes, among them RareTermsAggregationBuilder,
    VariableWidthHistogramAggregationBuilder and MovAvgModel.

Inner classes are pinned to their outer class, and a .java file is deleted only
when every class it produces is unreachable.

What went: 122 files of action/admin request/response/builder types orphaned by
the previous commit, 53 node-side transport channel/handler/connection classes,
26 of common/util, and the Lucene package-private wrappers that existed to serve
the already-deleted index/codec and index/compositeindex trees --
Lucene90DocValues{Consumer,Producer}Wrapper, DocValuesProducerUtil,
DocValuesWriterWrapper, MergeIndexWriter, Sorted{Numeric,Set}DocValuesWriterWrapper,
XQueryParser and DocIdStreamHelper. CollapseTopFieldDocs and
StrictISODateTimeFormat stay: both are genuinely reached, the latter from
common/joda/Joda, which Fess's taglib calls.

2408 -> 1839 files, 3751 -> 2992 classes, jar 7,451,679 -> 6,361,718 bytes.
373 tests pass, the 28-check runtime probe passes, the nodes-stats behaviour
snapshot is byte-identical, and every one of the 221 fork types the consumers
name is still present.
  RoaringBitmap  459,295 B  zero references in the fork; a stale declaration
  jzlib           71,976 B  zero references; likewise stale
  jopt-simple     78,146 B  zero references; it was the CLI option parser
  jsr305          19,936 B  only common/Nullable used it
  zstd-jni     6,795,212 B  only ZstdCompressor and its SPI provider used it

common/Nullable was meta-annotated @TypeQualifierNickname and @checkfornull.
Both are static-analysis metadata for tools that are not in this build, nothing
reads them at runtime, and the @nullable uses across the tree are unaffected.

ZstdCompressor and compress/spi/CompressionProvider are gone with the
dependency, and the provider is dropped from the CompressorProvider service
file. This client never compresses or decompresses through OpenSearch's
abstraction -- org/codelibs/fesen/client/** has zero references to anything
matching Compress, and the http.compression setting it does expose is
HTTP-level and handled by curl4j. CompressorRegistry now offers NONE and
DEFLATE, and the runtime probe asserts exactly that rather than dropping the
check.

Honest note on size: zstd-jni still ships in the Fess war, because
commons-compress and httpclient5 both depend on it. Removing it here is a
simplification and drops a spurious declaration; it does not shrink the
distribution. The other four do come off the distribution.

373 tests pass, the runtime probe passes 27/27, the nodes-stats behaviour
snapshot is byte-identical, and all 221 consumer-named types are present.
…pendencies

  HdrHistogram  177,206 B  only the Internal*HDRPercentile* classes used it
  t-digest       81,626 B  only TDigestState and the Internal*TDigest* classes
  log4j-core    ~1.8 MiB   only common/logging/Loggers reached into it

Internal* aggregation classes are the node-side reduce implementations. This
client parses JSON responses, and Parsed* is a separate hierarchy that does not
extend them -- ParsedMax extends ParsedSingleValueNumericMetricsAggregation, not
InternalMax -- so nothing a client does can reach one. The only thing anything
needed from them was a NAME string, in HttpClient.getDefaultNamedXContents() and
in six Parsed* classes. Those constants are inlined ("hdr_percentiles",
"tdigest_percentiles" and so on) and the classes are gone, which is the
governing rule for this prune: when a class you want to delete is referenced
only for a constant, inline the constant and delete the class.

The reachability keep-list is narrowed to match: Parsed* stays a root, Internal*
no longer is. No consumer names any Internal* or Parsed* type directly.

Loggers kept setLevel, addAppender, removeAppender and findAppender, which were
its only reach into log4j-core and which nothing here calls -- a client has no
business reconfiguring the logging implementation. The fork still logs through
the log4j API. GeoWKTParser borrowed Loggers.SPACE, a single-space string, now
inlined.

Kept, with reasons: the jackson cbor/smile/yaml dataformats and snakeyaml-engine,
because these formats are genuinely reachable -- Fess's SearchEngineUtil writes
YAML (SearchEngineUtilTest exercises XContentType.YAML) and this repository's own
HttpActionTest asserts YAML, CBOR and SMILE resolve.

1837 -> 1787 files, 2992 -> 2896 classes, jar 6,361,718 -> 6,089,304 bytes.
373 tests pass, probe 27/27, stats snapshot byte-identical, 221/221 consumer
types present.
…wo libraries

  jts-core   797,736 B
  spatial4j  204,833 B

Both existed only for common/geo/builders/** -- the ShapeBuilder hierarchy,
whose whole purpose is turning a shape into a spatial4j Shape or a JTS Geometry
so a node can compute intersections. A client serialises a geo query; it never
evaluates one. The fork already carries its own jts-free geometry model under
geometry/**, and GeoJson/GeometryParser already parse GeoJSON objects and WKT
strings into it.

Four links kept the old hierarchy alive, each removed by changing the keeper
rather than keeping the class:

  * QueryBuilders had four @deprecated ShapeBuilder overloads of geoShapeQuery,
    geoIntersectionQuery, geoWithinQuery and geoDisjointQuery, each already
    superseded by a Geometry form. Removed.
  * AbstractGeometryQueryBuilder and GeoShapeQueryBuilder each had an
    @deprecated ShapeBuilder constructor, likewise superseded. Removed.
  * ParsedGeometryQueryParams.shape was a ShapeBuilder; it is now a Geometry,
    and GeoShapeQueryBuilder's inline-shape parse goes through
    new GeometryParser(true, true, true) instead of ShapeParser.parse. That
    parser accepts the same GeoJSON-object and WKT-string inputs, and the flags
    are the ones AbstractGeometryQueryBuilder already uses a few lines away when
    it parses an indexed shape fetched from the server.
  * GeoJson referred to ShapeParser.FIELD_COORDINATES sixteen times while
    declaring its own identical FIELD_COORDINATES; the references are now local.

geoDistanceQuery, which is the only geo query Fess actually issues
(entity/GeoInfo), never touched either library and is unchanged.

That left common/geo/builders/**, common/geo/parsers/**, GeoShapeType and
XShapeCollection unreachable -- 17 files.

1787 -> 1770 files, 2896 -> 2866 classes, jar 6,089,304 -> 6,009,739 bytes.
373 tests pass, probe 27/27, stats snapshot byte-identical, 221/221 consumer
types present.
A pipeline aggregation builder describes an aggregation for the request body.
It was also carrying the machinery to instantiate the node-side reducer that
executes it: PipelineAggregationBuilder.create() ->
AbstractPipelineAggregationBuilder.createInternal() -> a concrete
PipelineAggregator, and AggregatorFactories.Builder.buildPipelineTree() ->
AggregationBuilder.buildPipelineTree(), which assembled those reducers into a
tree for the final reduce phase.

None of it can run in this client: it has no shards, no buckets to reduce, and
no aggregator factory path -- that was severed when QueryShardContext went.
Removing create/createInternal from the base classes and their 17 overrides,
and the two buildPipelineTree methods, drops 26 files: the concrete
Avg/Max/Min/Sum/Stats/ExtendedStats/Percentiles bucket reducers, BucketScript,
BucketSelector, BucketSort, CumulativeSum, Derivative, MovAvg, MovFn and
SerialDiff, plus BucketMetricsPipelineAggregator and SiblingPipelineAggregator.

Every *PipelineAggregationBuilder stays -- they are the published API, they
still serialise identically, and they are reachable only through the
named-writeable registry, so they remain explicit roots of the prune.
PipelineAggregator itself stays too: its nested PipelineTree is part of the
InternalAggregation reduce contract that the remaining Internal* result classes
still reference.

1770 -> 1744 files, 2866 -> 2837 classes, jar 6,009,739 -> 5,953,346 bytes.
373 tests pass, probe 27/27, stats snapshot byte-identical, 221/221 consumer
types present.
  lucene-backward-codecs  690,161 B  old on-disk codec formats, SPI-loaded
  lucene-sandbox          232,897 B
  lucene-misc             129,786 B
  lucene-memory            58,051 B  MemoryIndex, for percolation

All four are at zero references in src/main/java, and at zero in fess,
fess-suggest, fess-crawler and the five plugin repositories. They were declared
directly here rather than arriving transitively, so removing the declarations
removes them from the resolved tree -- confirmed with dependency:tree -- and
therefore from the Fess distribution, which gets its Lucene through this
artifact.

backward-codecs deserves a word: Lucene loads it through the codec SPI rather
than by symbol, so a reference count of zero is not by itself proof. It reads
superseded on-disk index formats, and nothing in this stack opens a Lucene index
-- the client speaks HTTP to a remote engine, and Fess does the same through it.

Lucene stays otherwise: core, analysis-common, queries, queryparser, join,
grouping, highlighter, suggest and spatial-extras are all still referenced,
and the downstream repositories compile against analysis, index, queryparser,
search and search.join through this artifact.

One consequence worth recording: fess-suggest inherits a provided-scope
lucene-facet from fess-parent, which used to be satisfied while resolving the
modules removed here, and now has to resolve on its own. It does, and nothing in
fess or fess-suggest references org.apache.lucene.facet, so there is no runtime
exposure -- but a build with a cold local repository now fetches one more
artifact.

373 tests pass, probe 27/27, stats snapshot byte-identical. Downstream re-run
against this jar: fess 7469, fess-suggest 689, fess-crawler-opensearch 53, all
green.
… edges

A library that serialises requests and parses responses has no reason to run
ServiceLoader at class-init. All six META-INF/services files are gone and the
directory with them; three were dead or redundant, three are now registered in
code.

Deleted outright:

  * java.util.spi.CalendarDataProvider / IsoCalendarDataProvider. It installed a
    process-wide java.util.Calendar override forcing MONDAY / minimal-days-4 for
    Locale.ROOT -- a client library embedded in Fess changing Calendar semantics
    for the whole JVM. It was also already inert: since JDK 9 the default
    java.locale.providers is CLDR,COMPAT, SPI is not consulted, and
    WeekFields.of(Locale.ROOT) reports SUNDAY/1 with the provider on the
    classpath. A differential probe over fourteen week-based formats is
    byte-identical before and after, so this is a no-op removal that also drops
    a JVM-wide side effect.
  * ErrorOnUnknown / SuggestingErrorOnUnknown. The interface already carried a
    default implementation; the provider only appended a Levenshtein "did you
    mean" hint. The lookup is now a constant and the provider is gone, taking
    its lucene-suggest LevenshteinDistance use with it.
  * org.apache.lucene.codecs.PostingsFormat. Ours named only
    Completion50PostingsFormat, which lucene-suggest already registers in its own
    service file along with the 84, 90 and 99 variants. Strictly redundant.

Registered statically instead of discovered, same contents:

  * CompressorRegistry -> NoneCompressor + DeflateCompressor
  * MediaTypeRegistry  -> XContentType.values() plus the application/* and
    application/x-ndjson aliases
  * XContentBuilder    -> XContentOpenSearchExtension

Each had exactly one or two providers, all shipping in this same artifact, so
the indirection bought nothing and carried a real failure mode this fork has hit
twice: a provider pruned as unreachable, the build still green, the registry
silently empty at runtime. The contract is now compile-time. The runtime probe
keeps checking it, but now asserts registry contents and the absence of the
service files rather than that discovery happened.

Two node-side edges cut, in the same spirit -- deleting the member rather than
accepting the reference:

  * BaseFuture.blockingAllowed() existed only to feed `assert blockingAllowed()`,
    asserting that a blocking get is not running on a transport thread. This
    client has no transport threads. That one call was the only thing keeping
    transport/Transports alive.
  * The reachability root set no longer over-approximates downstream references
    by bare simple name. Verified that no file downstream, and none under
    client/**, wildcard-imports a fork package, so every downstream reference is
    fully qualified on its import line; inner classes stay pinned to their outer.
    The old approximation rooted any fork class whose simple name appeared
    anywhere downstream -- Point, Table, Index, Task -- keeping 27 files alive
    for no reason, among them task/commons/**, the transport protocol headers,
    the server-side http request/response abstraction, node-side rescoring and
    the _cat Table renderer.

Behaviour change, recorded: "unknown query [x]" and "Unknown aggregation type
[x]" no longer carry a " did you mean [y]?" suffix. No test in this repository
or downstream asserts on it.

1744 -> 1709 files, 2837 -> 2789 classes, jar 5,953,302 -> 5,893,158 bytes.
373 tests pass, probe 21/21, stats snapshot byte-identical, week-format probe
byte-identical, 221/221 consumer types present.
…t types

Found by ranking every reachable class by how few references hold it alive and
how much would die with it -- a class with one incoming edge and a large subtree
is a bridge, and cutting that one member takes the subtree.

  * IndexRequest.extraFieldValues / UpdateRequest.docExtraFieldValues /
    upsertExtraFieldValues. The HTTP client never serialises extra field values:
    they only ever moved over the binary transport streams, so setting them
    through this client was a silent no-op. Nothing downstream or in the tests
    names them. The stream read and write keep their version-gated slots so the
    wire shape is unchanged; the values are now written as absent. 20 classes,
    the whole index/mapper/extrasource packed-array tree.
  * ThreadContext imported http/HttpTransportSettings for exactly two members,
    SETTING_HTTP_MAX_WARNING_HEADER_COUNT and _SIZE. That class is the node's
    HTTP *server* configuration -- bind hosts, port ranges, CORS, pipelining. The
    two settings are now defined verbatim in ThreadContext, same keys, same
    defaults, same validation.
  * IndexRequest and TaskResult used transport/client/Requests for one constant,
    INDEX_CONTENT_TYPE, which is MediaTypeRegistry.JSON. Inlined. Requests is a
    request-factory class nothing downstream calls.
  * SmoothingModel.buildWordScorerFactory() and its three overrides in Laplace,
    StupidBackoff and LinearInterpolation. Same shape as the pipeline
    aggregators: the builder describes the smoothing model for the request body
    and was also carrying the factory that instantiates the node-side scorer.
    That took WordScorer, the Laplace/StupidBackoff/LinearInterpolating scorers,
    CandidateScorer, NoisyChannelSpellChecker, Correction and FreqTermsEnum.
  * CompletionSuggestionBuilder.parseContextBytes(), a package-private helper
    with zero callers that resolved context bytes against a field's
    ContextMappings into node-side InternalQueryContexts. It was the only thing
    holding the completion context-mapping tree.

Checked and deliberately kept: FilterClient, the largest single candidate at
188 KB. filterWithHeader() is real -- SearchEngineClient and FesenClient both
override and delegate it, and FesenClientTest covers it. Also kept every
response type the ranking surfaced (SnapshotStatus, RepositoriesStats,
NodeCacheStats, RepositoryStatsSnapshot, SegmentReplicationPerGroupStats,
RemoteConnectionInfo, IndexingPressurePerShardStats, StatusCounterStats): those
parse real responses, several of them the nodes-stats document the behaviour
snapshot pins.

1709 -> 1671 files, 2789 -> 2738 classes, jar 5,893,158 -> 5,795,335 bytes.
373 tests pass, probe 21/21, stats snapshot byte-identical, 221/221 consumer
types present.
  * PhraseSuggestionBuilder held its tokenLimit default as
    NoisyChannelSpellChecker.DEFAULT_TOKEN_LIMIT. That one constant reference was
    what survived the previous commit's removal of buildWordScorerFactory and
    kept the whole node-side spell-checker alive. Inlined as 10, the value it
    already had, and WordScorer, CandidateScorer, NoisyChannelSpellChecker,
    Correction and FreqTermsEnum go with it.
  * OpenSearchExecutors.newSinglePrioritizing() and newAutoQueueFixed() had no
    callers. They built the cluster-state applier's priority executor and the
    node's auto-sizing search queue. Removing them takes
    PrioritizedOpenSearchThreadPoolExecutor, PrioritizedCallable,
    PrioritizedRunnable and QueueResizingOpenSearchThreadPoolExecutor.

The client's own ThreadPool is untouched: newFixed, newScaling, newResizable,
newDirectExecutorService, the thread factories and all three ExecutorBuilders
stay, because ThreadPool genuinely builds executors with them.

Considered and left: BigArrays' CircuitBreakerService plumbing. BigArrays is only
ever instantiated as NON_RECYCLING_INSTANCE with a null breaker service, so the
breaker path is inert -- but it sits in the array allocation that BytesStreamOutput
really uses, and the whole subtree is 6 classes and 11 KB. Not worth the risk of
disturbing serialisation for that.

1671 -> 1657 files, 2738 -> 2704 classes, jar 5,795,335 -> 5,735,739 bytes.
373 tests pass, probe 21/21, stats snapshot byte-identical, 221/221 consumer
types present.
This library no longer aims to cover the whole OpenSearch API surface. The
keep-rule is now: an operation survives only if one of the eight consumer
repositories uses it. This repository's own integration tests are no longer
sufficient grounds on their own, so several operations that OpenSearch3ClientTest
exercised go, and their tests go with them.

68 of the 99 operations are removed; 31 remain:

  reads     get, multiGet is gone, search, searchView is gone, streamSearch,
            createPit, deletePits, explain/fieldCaps/termVectors are gone
  writes    index, update, delete, bulk
  indices   create, exists, open, close, refresh, flush, getIndex, getMappings,
            putMapping, getSettings, aliases, getAliases, analyze
  cluster   health, nodesStats, nodesHotThreads
  structure admin, cluster, indices, settings, filterWithHeader,
            getRemoteClusterClient

Whole families go: snapshots and repositories, stored scripts, ingest pipelines,
index templates, tasks, cluster state/stats/reroute/settings/allocation-explain,
search shards, views, data streams, ingestion state, remote store, segment
replication, wlm stats, recovery, shard stores, indices stats, force merge,
clear cache, validate query, resize/rollover/scale, upgrade and upgrade status,
resolve index, scroll (searchScroll/clearScroll), multi-search, multi-get,
multi-term-vectors, term vectors, explain, field caps, get field mappings,
reindex, update-by-query, delete-by-query, simulate pipeline, remote info and
the main/info action.

How the used-set was derived, and the trap in it: SearchEngineClient and
FesenClient implement Client and override ~68 and ~64 methods, and an override
is not evidence that Fess calls the operation. Those two files also contain real
logic that does, so they could not simply be excluded. The scan blanks out their
pure delegation overrides -- an @OverRide whose body only forwards to the wrapped
client -- and then drops the imports left behind, which otherwise still match a
type pattern and make every delegated operation look used. Before that second
step the scan reported 43 operations used; after it, 33. searchView and
listViewNames were a third case: their overrides throw
UnsupportedOperationException("Not implemented yet"), which is not usage either.

Spot-checked the operations the admin screens and jobs reach, since those are
easy to miss: refresh (19 call sites), aliases (7), flush, open and close
(AdminMaintenanceAction, and SearchEngineClient.close()), getSettings,
nodesStats (LoadControlMonitorTarget, SystemMonitorTarget), nodesHotThreads
(HotThreadMonitorTarget) -- all kept. forceMerge, reindex, indices stats,
cluster state, update settings, validate query and rollover have zero call sites
and zero request/response type references anywhere in the eight repositories.
Scroll is kept out on the same evidence: everything named "scroll" downstream --
SearchEngineUtil.scroll, EsAbstractBehavior.scrollForDelete,
FesenClient.scrollForDelete -- goes through pitSearch. The Scroll type itself
stays; it is used as a keep-alive holder.

Tests: the 24 Http*ActionTest classes for removed actions are deleted outright.
The five integration suites keep their classes and lose only the test methods
that exercised a removed operation, found by compiling and mapping each error to
its enclosing method rather than by name. PIT coverage survives in
test_pit_search_after_pagination and test_pit_search_after_sort_value_round_trip;
the two PIT tests that went were the ones calling getAllPits.

Also removed Lucene.readTopDocs/writeTopDocs, the binary TopDocs transport
format, which had no callers and was the only thing holding
org/apache/lucene/search/grouping/CollapseTopFieldDocs. The org/apache/lucene
tree is now empty and gone.

1657 -> 1275 files, 2704 -> 2204 classes, jar 5,735,739 -> 4,645,600 bytes.
101 -> 27 Http*Action classes. This repository's own tests are 373 -> 271.
Probe 21/21, stats snapshot byte-identical, 219/219 consumer types present.
Downstream, all green: fess 7469, fess-suggest 689, fess-crawler-opensearch 53,
gsuite 240, mcp 597, multimodal 74, classic-api 80, v1-api 207.
This reverts commit e687609.

The scope was wrong. Removing what Fess does not use was premised on this
library still standing as itself, and the Http*Action classes are what that
means: they are the contract that makes fesen-httpclient an HTTP implementation
of the OpenSearch client rather than a Fess-private adapter. The
"keep only what a consumer uses" rule applies to OpenSearch internals this
library never exercises -- node-side execution, reduce phases, transport
plumbing -- not to the library's own action implementations.

So all 101 Http*Action classes come back, with their ActionTypes, their
request/response/builder types, the client interface methods and the tests: the
24 Http*ActionTest classes and the integration-suite methods. upgrade comes back
too; it was dropped on the belief that nothing implemented it, and
HttpUpgradeAction does.

The analysis behind the reverted commit still stands as a fact about the
consumers -- only 31 of the 99 operations are reachable from fess, fess-suggest,
fess-crawler-opensearch or the five plugin repositories, and the other 68 exist
for callers outside this workspace. That is a statement about who uses the
library, not a licence to shrink its API.

373 tests, probe 21/21, stats snapshot byte-identical, 221/221 consumer types.
Lucene.readTopDocs and Lucene.writeTopDocs serialise a TopDocs plus its max
score to and from a StreamInput/StreamOutput -- the format the node's search
phases use to ship per-shard results between nodes. Nothing calls either of
them: this client parses search responses from JSON. Verified with the full set
of 101 Http*Action classes and every client operation present, so this is not a
consequence of anything the reverted commit removed.

Removing them takes with them:

  org/apache/lucene/search/grouping/CollapseTopFieldDocs   256 lines
  common/lucene/search/TopDocsAndMaxScore                   55 lines
  common/lucene/Lucene                                      89 lines removed

CollapseTopFieldDocs existed only for package-private access to Lucene's
grouping internals and was reached from nothing else, so org/apache/lucene is
now empty and gone -- one less package this artifact splits with lucene-core,
lucene-grouping and lucene-queryparser. TopDocsAndMaxScore had one other
mention, an unused import in NestedQueryBuilder, now dropped.

indices/replication/common/ReplicationLuceneIndex stays. It looked like part of
the same group, but it is reached from SegmentReplicationState, RecoveryState
and ReplicationState -- the response types of the recoveries and
segmentReplicationStats operations, both of which this library implements.

1657 -> 1655 files, 2704 -> 2700 classes, jar 5,735,739 -> 5,726,805 bytes.
373 tests pass, probe 21/21, stats snapshot byte-identical, 221/221 consumer
types present.
Reachability calls a class "reachable" whether a hundred things use it or one
assert mentions it, so it never surfaces the shape that has dominated this prune:
a large class kept alive by one small member. This makes the hunt mechanical.
For every class, rank by the bytes that would die with it divided by the number
of distinct (referrer, member) pairs holding it alive, re-ranking after each cut
because every cut changes the weights below it.

Three shapes the class graph structurally cannot see, all now swept:

RESOURCES -- not in the class graph at all.
  * tasks/task-index-mapping.json, the mapping for a node's .tasks index. Its
    only mention anywhere was a javadoc sentence in Task.java. Nothing loads it.
    src/main/resources is now just the two META-INF licence files, and the jar's
    only other non-class entries are the manifest and the maven descriptor.

A LARGE CLASS HELD BY ONE SMALL MEMBER -- the ratio hunt.
  * JarHell, 14,345 bytes of classpath duplicate/split-package scanner, alive
    because PluginInfo called JarHell.checkVersionFormat -- thirteen lines that
    parse a version string. Moved to PluginInfo; JarHell deleted.
  * InFlightShardSnapshotStates, held only by SnapshotsInProgress's private
    assertConsistentEntries, itself called only from `assert`.
  * SecureSetting, held only by `assert this instanceof SecureSetting || ...`
    in Setting's constructor.
  * RoutingTableIncrementalDiff, held only by RoutingTable.incrementalDiff,
    which nothing calls -- cluster-state diffs are a node-side publication
    concern.
  * BulkShardRequest, held only by BulkRequest's `timeout =
    BulkShardRequest.DEFAULT_TIMEOUT`, inherited from ReplicationRequest.
  * MoreLikeThisQuery and XMoreLikeThis, 23 KB of Lucene more-like-this term
    selection, held only by seven default constants MoreLikeThisQueryBuilder
    copied out of them.
  * TermVectorsWriter, held only by TermVectorsResponse.setFields -- the
    server-side path that *builds* a term-vectors response.
  * InternalExtendedStats and InternalGeoCentroid, node-side reduce
    implementations held only by their nested Fields classes of JSON field
    names, which ParsedExtendedStats and ParsedGeoCentroid import. The names
    moved to the Parsed* classes that read them.
  * TransportRequestOptions, held only by ActionType.transportOptions and one
    override in BulkAction. That method configures the binary transport channel
    -- compression, timeouts, request type. Nothing calls it; an HTTP client has
    no channel to configure.
  * RemoteConnectionStrategy, held only by RemoteConnectionInfo reading its enum
    off a stream and discarding it, and by modeType() in the ModeInfo interface,
    which nothing implements. HttpRemoteInfoAction says as much in a comment: it
    returns an empty connection list because RemoteConnectionInfo cannot be
    constructed externally.

A CLASS HELD BY AN INHERITANCE EDGE -- invisible to both of the above.
  * TransportMessage. Every response DTO inherits from ActionResponse and every
    request from ActionRequest, so the transport base classes were structurally
    reachable and always would be. TransportMessage contributed exactly one
    thing: remoteAddress, a TransportAddress, which is meaningless to an HTTP
    client and read nowhere. TransportResponse and TransportRequest now
    implement Writeable directly and TransportMessage is gone.

    The wider flattening was measured and rejected. Dropping TransportResponse
    and TransportRequest as well would remove 6 classes and 7,321 bytes but
    costs the TaskAwareRequest contract, and that is genuinely used --
    GetTaskRequest, ReplicationRequest and AbstractBulkByScrollRequest all read
    or copy the parent task. Kept.

Also measured and rejected: removing the aggregation reduce() family. 8 classes
and 37,972 bytes for 34 members edited, including a public abstract method on
InternalAggregation -- about 1.1 KB per member, an order of magnitude worse than
anything above.

Kept with reasons: RemoteClusterAware, because an earlier round had already
reduced it to 67 lines holding buildRemoteIndexName and
REMOTE_CLUSTER_INDEX_SEPARATOR -- it *is* the small utility, and the cluster:index
convention is real over HTTP since the server returns such names.
Lucene.readSortField/writeSortField, because SearchHits uses them.
BulkRequestParser, because BulkRequest.add parses an NDJSON bulk body, which is
real client API.

No package-info.java files exist anywhere in the tree, and the remaining asserts
that call into a static helper all name classes kept for other reasons.

1657 -> 1639 files, 2704 -> 2663 classes, jar 5,735,739 -> 5,644,760 bytes.
373 tests pass, probe 21/21, stats snapshot byte-identical, 221/221 consumer
types present. All 101 Http*Action classes untouched.
…stractClient

HttpAbstractClient implements Client directly -- it has never extended OpenSearch's
AbstractClient. The only thing in the whole fork that referenced AbstractClient or
FilterClient was one twelve-line method:

    public Client filterWithHeader(Map<String, String> headers) {
        return new FilterClient(this) { ... };      // FilterClient extends AbstractClient
    }

FilterClient is a 3 KB delegating wrapper built on AbstractClient, which is
OpenSearch's node-side implementation of every Client method in terms of
execute() -- 185,187 bytes of class files, 91,731 of source. This class already
implements all of those methods the same way, so the wrapper is now an anonymous
subclass of HttpAbstractClient and the two classes are gone.

Behaviour is unchanged. doExecute still merges the headers into the thread
context with ThreadContext.stashAndMergeHeaders and runs the original client's
doExecute inside that scope, and close() still closes the wrapped client, which
is what FilterClient.close() did.

Basing the wrapper here means every Client method needs an implementation on this
class, so searchView, listViewNames and prepareStreamSearch move down from
HttpClient. They were plain execute() calls with nothing HttpClient-specific
about them.

Why the previous round kept this: the survivor list said "filterWithHeader is
delegated by SearchEngineClient and FesenClient and covered by FesenClientTest".
That justifies keeping the *operation*, which is part of the Client contract and
stays. It says nothing about the implementation -- and a delegating override is
exactly the false positive this prune has already been caught by once.
SearchEngineClient:3396 and FesenClient:608 both just forward to HttpClient,
which is where the work actually happens.

1639 -> 1634 files, 2663 -> 2652 classes, jar 5,644,760 -> 5,615,308 bytes
(the 185 KB is uncompressed class bytes; the jar stores them deflated).
373 tests pass, probe 21/21, stats snapshot byte-identical, 221/221 consumer
types present. All 101 Http*Action classes untouched.
Suggest kept a hasScoreDocs flag, recomputed in both constructors as

    filter(CompletionSuggestion.class).stream().anyMatch(CompletionSuggestion::hasScoreDocs)

and exposed through hasScoreDocs(). It exists for the node's fetch phase, which
decides whether suggestions carry score docs that still need fetching. Nothing
reads it -- not this repository, not its tests, not any of the eight consumer
repositories.

That one flag was the only thing holding CompletionSuggestion, 32 KB across five
classes. Completion suggestions are not parseable here in any case:
HttpClient.getDefaultNamedXContents() has the suggestion entries commented out,
so Suggestion.fromXContent never produces one.

Kept, with the reason, since it looked like the same shape: TermVectorsFields,
held only by TermVectorsResponse.getFields(). Over HTTP that always returns null
-- HttpTermVectorsAction does not parse term_vectors and nothing calls
setTermVectorsField -- but the binary path is internally consistent: the
StreamInput constructor reads the blob, writeTo writes it, and toXContent renders
it through getFields(). Removing it would mean deleting a rendering branch from a
public response type and foreclosing the HTTP action ever learning to parse term
vectors. That is a gap in the action, not dead weight in the response.

1634 -> 1633 files, 2652 -> 2647 classes, jar 5,615,308 -> 5,602,124 bytes.
373 tests pass, probe 21/21, stats snapshot byte-identical, 221/221 consumer
types present.
… check

Two more of the structural shape, found while measuring whether the empty
ClusterState that HttpClusterStateAction fabricates was holding a large subtree.
It is not (see below), but two of its neighbours were holding one.

  * ResponseCollectorService implemented ClusterStateListener. That is how a node
    maintained the statistics: clusterChanged(ClusterChangedEvent) added and
    removed per-node entries as the cluster state moved. Everything in this tree
    uses only the nested ComputedNodeStats, which AdaptiveSelectionStats carries
    in the _nodes/stats payload -- AdaptiveSelectionStats names
    ResponseCollectorService.ComputedNodeStats six times and the outer class
    never. Dropping the interface and clusterChanged takes ClusterStateListener
    and ClusterChangedEvent with them.
  * ActiveShardCount.enoughShardsActive, in both overloads, one taking a
    ClusterState and one an IndexShardRoutingTable. ActiveShardCount is a real
    request parameter -- wait_for_active_shards, set on a dozen request types and
    read by as many Http*Actions -- but evaluating whether enough shards are
    actually active is the node's job, and nothing here calls it.

1633 -> 1631 files, 2647 -> 2645 classes, jar 5,602,124 -> 5,595,626 bytes.
373 tests pass, probe 21/21, stats snapshot byte-identical, 221/221 consumer
types present.
A member-level reachability run over the packaged jar - rooted at this repo's main
and test sources plus the main and test classes of all eight downstream consumers -
found members of the forked OpenSearch types that nothing reaches. They are residue
of the import, not behaviour: a dead setter means no Http*Action ever read the field,
so the parameter was never sent.

This slice removes the first batch. The deletion set is closed under the bytecode
reference graph, so nothing left behind can still point into it, and it excludes
every member that could change behaviour silently:

- anything overriding or overridden anywhere in the hierarchy
- abstract and interface declarations, including the Client interface family, which
  is this library's own API surface rather than fork residue
- compile-time constants, whose reads javac inlines out of the bytecode
- a type's last constructor, which javac would replace with an implicit no-arg one
- field initialisers, which bytecode attributes to the constructor

No functional change. The 101 Http*Action classes and the HTTP behaviour are untouched.

Verified: 373 tests pass; the 21-check runtime probe passes; the _nodes/stats
rendering snapshot is byte-identical (405 lines, md5 10f57efde06c6745e3c724cdcd31273f);
all 221 fork types named by the eight downstream repos still resolve.
Jar 5,595,626 -> 5,351,504 bytes.
Second batch of the same reachability sweep, under the same guards: the set is closed
under the bytecode reference graph, and overrides, abstract declarations, the Client
interface family, compile-time constants, last constructors and field initialisers are
all held back.

No functional change.

Verified: 373 tests pass; the 21-check runtime probe passes; the _nodes/stats
rendering snapshot is byte-identical (405 lines, md5 10f57efde06c6745e3c724cdcd31273f);
all 221 fork types named by the eight downstream repos still resolve.
Jar 5,351,504 -> 5,224,897 bytes.
Third batch of the same reachability sweep, under the same guards.

No functional change.

Verified: 373 tests pass; the 21-check runtime probe passes; the _nodes/stats
rendering snapshot is byte-identical (405 lines, md5 10f57efde06c6745e3c724cdcd31273f);
all 221 fork types named by the eight downstream repos still resolve.
Jar 5,224,897 -> 5,164,069 bytes.
Final batch of the reachability sweep.

One extra guard earned by this batch: HttpScaleIndexAction reads ScaleIndexRequest
via reflection (getMethod("getIndex") / getMethod("isScaleDown")), which no bytecode
reachability analysis can see. Deleting those two getters compiled cleanly and broke
four tests. Every reflective lookup in the repo was then audited by resolving the
looked-up name against the types named in the same file; these two are the only ones
that reach a forked type, and both are now held back.

No functional change.

Verified: 373 tests pass; the 21-check runtime probe passes; the _nodes/stats
rendering snapshot is byte-identical (405 lines, md5 10f57efde06c6745e3c724cdcd31273f);
all 221 fork types named by the eight downstream repos still resolve.
Jar 5,164,069 -> 5,141,158 bytes.
…ds reflectively

HttpGetFieldMappingsAction does not call this constructor; it looks it up with
getDeclaredConstructor(Map.class) and calls setAccessible(true), because the
constructor is package-private. No reachability analysis can see that, and the
earlier audit of reflective access only matched name literals passed to getMethod
and getDeclaredField, so a constructor selected by a Class<?>[] of parameter types
slipped through.

It is functionality: it is how the action turns a parsed field-mappings response
into its response object. The local suites do not cover it - only the Docker suites,
which run against real servers, caught it, failing test_get_field_mappings and
test_create_knn_field across all five engine versions.

Restored byte-identically. The reflective surface of the library's main sources is
now three sites in total and all are accounted for: getMethod on ScaleIndexRequest
(two getters, already held back), this constructor, and Array.newInstance in
HttpAnalyzeAction, which creates an array rather than looking up a member.

Verified: 373 tests pass; the 21-check runtime probe passes; the _nodes/stats
rendering snapshot is byte-identical (405 lines, md5 10f57efde06c6745e3c724cdcd31273f).
These are not dead by reachability - they are purposeless. They exist to rebuild a DTO
from the binary transport format, and this client speaks JSON over HTTP; it never reads
a node-to-node stream. They are the subset whose deletion needs no other change: nothing
in the jar references them, by call or by method handle, so no Writeable.Reader and no
super(in) chain is left pointing at a missing constructor.

Three things are deliberately excluded, and being referrer-free happens to establish all
three on its own:

- the QueryBuilder family (48 of the 182 referrer-free constructors). AbstractQueryBuilder
  implements NamedWriteable and Fess depends on the contract: HybridQueryBuilder calls
  readNamedWriteableList and writeNamedWriteableList, and KnnQueryBuilder serialises too.
- every DTO the actions that really do round-trip the wire format touch. A constructor
  reachable from one of those actions has a referrer by definition, so none is in here.
- ActionType.getResponseReader and anything its readers point at, for the same reason:
  a stored method handle is a reference and is counted as one.

writeTo is not touched. It cannot be removed the same way: 1,017 types inherit Writeable,
so dropping the method means changing what those types implement, which is interface
surgery rather than member pruning and belongs in its own change.

Verified: 373 tests pass; the 21-check runtime probe passes; the _nodes/stats rendering
snapshot is byte-identical (405 lines, md5 10f57efde06c6745e3c724cdcd31273f) - that
snapshot renders a real response through these DTOs, so it is the direct check that
parsing is unaffected; all 221 fork types named by the eight downstream repos resolve.
Jar 5,141,216 -> 5,115,943 bytes.
… slice

Removing a subclass constructor removes its super(in) call, which is what kept the
parent's StreamInput constructor anchored. Re-running the same referrer analysis over
the rebuilt jar frees another round. Same exclusions as before; the QueryBuilder family
and everything the wire-format actions reach are still held back.

Verified: 373 tests pass; the 21-check runtime probe passes; the _nodes/stats rendering
snapshot is byte-identical (405 lines, md5 10f57efde06c6745e3c724cdcd31273f).
Jar 5,115,943 -> 5,110,902 bytes.
Third and final round of the same referrer analysis; it converges here. 517 StreamInput
constructors remain, all of them still anchored - by a Writeable.Reader method handle, a
super(in) chain, or a live call - and removing those needs the Writeable interface work
rather than member deletion.

Verified: 373 tests pass; the 21-check runtime probe passes; the _nodes/stats rendering
snapshot is byte-identical (405 lines, md5 10f57efde06c6745e3c724cdcd31273f).
Jar 5,110,902 -> 5,109,407 bytes.
…hable

Member-by-member surgery inside a class that is entirely dead is the wrong operation, so
these go as whole files. The compiler decided which ones: all 165 candidates were deleted
at once and any whose type a surviving file still names was restored, repeated to a
fixed point. 116 came back; 49 stayed out.

Three files lost an import line rather than a member: their only use of a deleted type had
already gone with an earlier slice, and an unused import of a package that no longer exists
is an error rather than a warning.

Verified: 373 tests pass; the 21-check runtime probe passes; the _nodes/stats rendering
snapshot is byte-identical (405 lines, md5 10f57efde06c6745e3c724cdcd31273f); all 221 fork
types named by the eight downstream repos still resolve.
Files 1,631 -> 1,582. Jar 5,109,407 -> 4,977,043 bytes.
The rest of the fully-dead classes are nested inside types that are still used, so the file
has to stay and only the nested declaration goes. Same oracle as the whole-file slice:
delete all 297 candidates, drop the ones javac still resolves, repeat to a fixed point.
249 came back; 48 stayed out.

Targets contained inside another target are skipped, so a dead type nested in a dead type
is removed once with its parent rather than twice.

Verified: 373 tests pass; the 21-check runtime probe passes; the _nodes/stats rendering
snapshot is byte-identical (405 lines, md5 10f57efde06c6745e3c724cdcd31273f); all 221 fork
types named by the eight downstream repos still resolve.
Jar 4,977,043 -> 4,920,282 bytes.
Deleting a class removes its references, so the candidate set has to be re-derived rather
than assumed. Running the whole-file and nested-type candidates together through the same
compile oracle frees eight more: the BigArrays element interfaces whose only implementations
had already gone, and three interfaces left without implementors.

Re-running the oracle after this converges - nothing further falls out.

Verified: 373 tests pass; the 21-check runtime probe passes; the _nodes/stats rendering
snapshot is byte-identical (405 lines, md5 10f57efde06c6745e3c724cdcd31273f); all 221 fork
types named by the eight downstream repos still resolve.
Jar 4,920,282 -> 4,915,121 bytes.
@marevol marevol self-assigned this Sep 13, 2026
@marevol
marevol merged commit 5b376ea into main Sep 13, 2026
1 check passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant