feat: fork the OpenSearch classes this client needs into org.codelibs.fesen.opensearch - #43
Merged
Merged
Conversation
Import the org.codelibs.fesen.opensearch.* fork of OpenSearch 3.8.0 into src/main/java, carry the needed resources (merged META-INF/services, task-index-mapping.json), migrate the client sources to the forked package, and delete the node-only trees.
…it compile Keep only the classes reachable from what fesen-httpclient's own sources use, drop the node-extension and bootstrap trees (they need protobuf, JNA and the opensearch-agent-bootstrap module), drop the reactive HTTP streaming path (it needs reactor-core), and drop the vectorised B-tree rounding searcher (it needs --add-modules jdk.incubator.vector).
Declare the third-party libraries the forked sources actually compile against, restore the ServiceLoader provider implementations that static reachability cannot see, scope the license and formatter plugins to the client sources only, and ship the OpenSearch NOTICE/LICENSE with a statement of the CodeLibs modifications.
Node keeps only the settings constants other classes reference; its construction, lifecycle and dependency-injection wiring are gone. With that, common/inject (the bundled Guice fork) is removed except the one-method Provider interface that two aggregation/analysis types use, and the @Inject annotations and Guice binding methods that referenced it are removed from the modules and plugin extension points.
Drop lucene-spatial3d (nothing references it) and declare the JUnit 4 assertions the existing tests import directly instead of relying on the Testcontainers transitive.
Only the Jackson streaming API is used, so exclude the databind transitive from the three dataformat modules, as OpenSearch does.
fesen-httpclient only ever speaks HTTP to a remote cluster, so everything that exists to run an OpenSearch node is dead weight. Sever the four bridges that kept it reachable and drop what falls out: * ClusterSettings / IndexScopedSettings / FeatureFlagSettings / SettingsModule / AbstractScopedSettings. HttpNodesStatsAction no longer builds a node-scoped ClusterSettings just to construct a SearchRequestStats whose counters are all zero; NodeIndicesStats gains a constructor without it. The rendered _nodes/stats output is byte-identical. * rest/**. The four request parameters SearchHit, SearchHits, Suggest and InternalAggregation read are inlined with their exact string values, the _shards header renderer moves to action.support.broadcast.BroadcastShardsHeader, AliasesNotFoundException moves to action.admin.indices.alias (same simple name, so same wire and XContent identity), and the error-body decoder moves into HttpAction. * QueryShardContext: doToQuery/toQuery/toFilter on all 48 query builders, plus InnerHitContextBuilder and the shard-side doRewrite branches. * The aggregation value-source bridge: AggregationBuilder.build, innerBuild, ValuesSourceConfig and AggregatorFactories.create*, which is what kept index/mapper/** alive. Classes that survive only for a nested request type (Aggregator, RangeAggregator, TermsAggregator, FiltersAggregator, AdjacencyMatrixAggregator, HeapUsageTracker, NativeMemoryUsageTracker, MatchQuery, FunctionScoreQuery, MultiValueMode, IndexModule, MapperService, DateFieldMapper, SearchContext and friends) keep that type and lose the node-side body. Every exception in the OpenSearchServerException wire registry is kept so error deserialisation is unchanged. The Lucene Codec and PostingsFormat service entries and the classes behind them go with index/codec/** and index/compositeindex/**: a client never opens a Lucene index, and removing the provider names along with the classes leaves no dangling ServiceLoader entry. ScriptService.compile, ScriptService.isLangSupported and three coordinator-side aggregation reduce paths that compile scripts now throw UnsupportedOperationException; a client never reduces aggregations or compiles scripts. 4604 -> 2401 source files, 931k -> 434k lines, jar 17.51 MB -> 7.09 MB. 370 tests still pass and the _nodes/stats fixture round-trip is unchanged.
HttpCreateIndexAction set is_write_index=false on every alias whose
writeIndex() was unset, so an index created with
prepareCreate(name).addAlias(new Alias(a)) came up with a read-only alias.
Indexing through that alias then failed with
no write index is defined for alias [a]
even though the alias pointed at a single index, which both the transport
client and the REST API treat as writable. Leave the flag unset unless the
caller set it.
Found while moving fess-suggest onto this client: Suggester.createIndexIfNothing()
creates its search and update aliases exactly this way, and every write through
them failed.
Cross-checking the 452 distinct OpenSearch types that Fess, fess-suggest, fess-crawler-opensearch and the five plugin repositories name against this fork left three gaps. Two are real and are restored here. ScoreFunctionBuilders is the factory for the function-score builders. Every builder it returns was already here; only the factory itself had been pruned, because nothing inside this library calls it. Fess calls it from KeyMatchHelper, QueryHelper and SuggestHelper, and fess-suggest from SuggestQueryBuilder and PopularWordsRequest. common/joda carries Joda, JodaDateFormatter, JodaDateMathParser and JodaDeprecationPatterns. Fess uses Joda.forPattern(...).parseMillis(...) in its taglib to parse a date string, so the parsing semantics are user-visible and worth keeping exactly as they are rather than substituting another formatter. The third gap, QueryShardContext, stays gone. It is the bridge that was cut to drop the mapper and shard trees, and its only consumers are three Fess query builders whose doToQuery overrides throw UnsupportedOperationException. 370 tests pass. The jar grows by 23 KB.
This was referenced Sep 13, 2026
refactor: compile against the forked OpenSearch package in fesen-httpclient
codelibs/fess-suggest#98
Merged
Alias.toXContent wrote is_write_index unconditionally while every other
optional field in that method is null-guarded. Two failure modes followed:
- forcing false made every alias created through HttpCreateIndexAction
read-only, so indexing through an alias failed with "no write index is
defined for alias [...]" even when the alias pointed at a single index,
which the transport client accepts as writable;
- leaving it unset then rendered is_write_index: null, which Elasticsearch 8
rejects outright with "Unknown token [VALUE_NULL] in alias [...]".
Omitting the field is the only rendering both engines read as "the caller did
not decide". The fork owns Alias, so the guard goes at the source and
HttpCreateIndexAction keeps delegating to it.
373 tests pass.
Client, ClusterAdminClient and IndicesAdminClient advertised 122 operations.
Only 99 of them have an HTTP action registered in HttpClient; the other 23
reached HttpClient.doExecute, found no handler and threw
UnsupportedOperationException: Action: <name>
so they were never usable over HTTP. Two independent signals agree exactly on
that set: no consumer references any of them, and none has a registered handler.
The used-set was derived mechanically, not by inspection. For each operation the
scan looked for prepareX( call sites, references to the request/response/builder
types that are unique to it (types shared between operations, such as
AcknowledgedResponse, are not evidence), and references to its ActionType --
across this repository's main and test sources, fess, fess-suggest,
fess-crawler-opensearch and the five plugin repositories.
Removed (23 operations, 153 method declarations across the interfaces,
AbstractClient, HttpAbstractClient and HttpIndicesAdminClient):
dangling indices deleteDanglingIndex, importDanglingIndex, listDanglingIndices
decommission decommission, deleteDecommissionState, getDecommissionState,
prepareDeleteDecommissionRequest
weighted routing putWeightedRouting, getWeightedRouting, deleteWeightedRouting,
weightedRouting
data streams createDataStream, deleteDataStream, getDataStreams
search pipelines putSearchPipeline, getSearchPipeline, deleteSearchPipeline
snapshots cloneSnapshot, cleanupRepository
remote store restoreRemoteStore
other addBlock, pitSegments, reloadSecureSettings
Kept every operation that has a handler, plus the structural accessors
admin()/cluster()/indices()/settings()/filterWithHeader()/getRemoteClusterClient().
Surviving signatures are untouched, so downstream code needs no change beyond
deleting overrides of methods that no longer exist.
No behaviour change: every removed method threw on call. 373 tests pass.
…umers
Class-level reachability over the built class files, iterated to convergence.
Edges come from each class file's constant pool (CONSTANT_Class entries plus the
type descriptors in Utf8 entries, so Signature attributes and annotations count).
The root set is deliberately wider than this repository, because a prune rooted
only in org/codelibs/fesen/client/** deletes types that nothing here mentions but
downstream code does, leaving this build green and the downstream one broken:
* org/codelibs/fesen/client/** and every test class
* every class named by a META-INF/services provider file
* every org.codelibs.fesen.opensearch type named by fess, fess-suggest,
fess-crawler-opensearch and the five plugin repositories, main and test
sources both, plus any fork class whose simple name appears there at all
(an over-approximation that covers Outer.Inner written after import Outer)
* public API reachable only through the named-writeable / named-XContent
registries -- every *AggregationBuilder, *QueryBuilder, *SuggestionBuilder,
*ScoreFunctionBuilder, *SortBuilder, *SignificanceHeuristic, *SmoothingModel,
the MovAvg models, and the Internal*/Parsed* aggregation results. Static
reachability deletes these with the build still green: they are the library's
published surface, so they are roots rather than an afterthought. That keep
covers 234 classes, among them RareTermsAggregationBuilder,
VariableWidthHistogramAggregationBuilder and MovAvgModel.
Inner classes are pinned to their outer class, and a .java file is deleted only
when every class it produces is unreachable.
What went: 122 files of action/admin request/response/builder types orphaned by
the previous commit, 53 node-side transport channel/handler/connection classes,
26 of common/util, and the Lucene package-private wrappers that existed to serve
the already-deleted index/codec and index/compositeindex trees --
Lucene90DocValues{Consumer,Producer}Wrapper, DocValuesProducerUtil,
DocValuesWriterWrapper, MergeIndexWriter, Sorted{Numeric,Set}DocValuesWriterWrapper,
XQueryParser and DocIdStreamHelper. CollapseTopFieldDocs and
StrictISODateTimeFormat stay: both are genuinely reached, the latter from
common/joda/Joda, which Fess's taglib calls.
2408 -> 1839 files, 3751 -> 2992 classes, jar 7,451,679 -> 6,361,718 bytes.
373 tests pass, the 28-check runtime probe passes, the nodes-stats behaviour
snapshot is byte-identical, and every one of the 221 fork types the consumers
name is still present.
RoaringBitmap 459,295 B zero references in the fork; a stale declaration jzlib 71,976 B zero references; likewise stale jopt-simple 78,146 B zero references; it was the CLI option parser jsr305 19,936 B only common/Nullable used it zstd-jni 6,795,212 B only ZstdCompressor and its SPI provider used it common/Nullable was meta-annotated @TypeQualifierNickname and @checkfornull. Both are static-analysis metadata for tools that are not in this build, nothing reads them at runtime, and the @nullable uses across the tree are unaffected. ZstdCompressor and compress/spi/CompressionProvider are gone with the dependency, and the provider is dropped from the CompressorProvider service file. This client never compresses or decompresses through OpenSearch's abstraction -- org/codelibs/fesen/client/** has zero references to anything matching Compress, and the http.compression setting it does expose is HTTP-level and handled by curl4j. CompressorRegistry now offers NONE and DEFLATE, and the runtime probe asserts exactly that rather than dropping the check. Honest note on size: zstd-jni still ships in the Fess war, because commons-compress and httpclient5 both depend on it. Removing it here is a simplification and drops a spurious declaration; it does not shrink the distribution. The other four do come off the distribution. 373 tests pass, the runtime probe passes 27/27, the nodes-stats behaviour snapshot is byte-identical, and all 221 consumer-named types are present.
…pendencies
HdrHistogram 177,206 B only the Internal*HDRPercentile* classes used it
t-digest 81,626 B only TDigestState and the Internal*TDigest* classes
log4j-core ~1.8 MiB only common/logging/Loggers reached into it
Internal* aggregation classes are the node-side reduce implementations. This
client parses JSON responses, and Parsed* is a separate hierarchy that does not
extend them -- ParsedMax extends ParsedSingleValueNumericMetricsAggregation, not
InternalMax -- so nothing a client does can reach one. The only thing anything
needed from them was a NAME string, in HttpClient.getDefaultNamedXContents() and
in six Parsed* classes. Those constants are inlined ("hdr_percentiles",
"tdigest_percentiles" and so on) and the classes are gone, which is the
governing rule for this prune: when a class you want to delete is referenced
only for a constant, inline the constant and delete the class.
The reachability keep-list is narrowed to match: Parsed* stays a root, Internal*
no longer is. No consumer names any Internal* or Parsed* type directly.
Loggers kept setLevel, addAppender, removeAppender and findAppender, which were
its only reach into log4j-core and which nothing here calls -- a client has no
business reconfiguring the logging implementation. The fork still logs through
the log4j API. GeoWKTParser borrowed Loggers.SPACE, a single-space string, now
inlined.
Kept, with reasons: the jackson cbor/smile/yaml dataformats and snakeyaml-engine,
because these formats are genuinely reachable -- Fess's SearchEngineUtil writes
YAML (SearchEngineUtilTest exercises XContentType.YAML) and this repository's own
HttpActionTest asserts YAML, CBOR and SMILE resolve.
1837 -> 1787 files, 2992 -> 2896 classes, jar 6,361,718 -> 6,089,304 bytes.
373 tests pass, probe 27/27, stats snapshot byte-identical, 221/221 consumer
types present.
…wo libraries jts-core 797,736 B spatial4j 204,833 B Both existed only for common/geo/builders/** -- the ShapeBuilder hierarchy, whose whole purpose is turning a shape into a spatial4j Shape or a JTS Geometry so a node can compute intersections. A client serialises a geo query; it never evaluates one. The fork already carries its own jts-free geometry model under geometry/**, and GeoJson/GeometryParser already parse GeoJSON objects and WKT strings into it. Four links kept the old hierarchy alive, each removed by changing the keeper rather than keeping the class: * QueryBuilders had four @deprecated ShapeBuilder overloads of geoShapeQuery, geoIntersectionQuery, geoWithinQuery and geoDisjointQuery, each already superseded by a Geometry form. Removed. * AbstractGeometryQueryBuilder and GeoShapeQueryBuilder each had an @deprecated ShapeBuilder constructor, likewise superseded. Removed. * ParsedGeometryQueryParams.shape was a ShapeBuilder; it is now a Geometry, and GeoShapeQueryBuilder's inline-shape parse goes through new GeometryParser(true, true, true) instead of ShapeParser.parse. That parser accepts the same GeoJSON-object and WKT-string inputs, and the flags are the ones AbstractGeometryQueryBuilder already uses a few lines away when it parses an indexed shape fetched from the server. * GeoJson referred to ShapeParser.FIELD_COORDINATES sixteen times while declaring its own identical FIELD_COORDINATES; the references are now local. geoDistanceQuery, which is the only geo query Fess actually issues (entity/GeoInfo), never touched either library and is unchanged. That left common/geo/builders/**, common/geo/parsers/**, GeoShapeType and XShapeCollection unreachable -- 17 files. 1787 -> 1770 files, 2896 -> 2866 classes, jar 6,089,304 -> 6,009,739 bytes. 373 tests pass, probe 27/27, stats snapshot byte-identical, 221/221 consumer types present.
A pipeline aggregation builder describes an aggregation for the request body. It was also carrying the machinery to instantiate the node-side reducer that executes it: PipelineAggregationBuilder.create() -> AbstractPipelineAggregationBuilder.createInternal() -> a concrete PipelineAggregator, and AggregatorFactories.Builder.buildPipelineTree() -> AggregationBuilder.buildPipelineTree(), which assembled those reducers into a tree for the final reduce phase. None of it can run in this client: it has no shards, no buckets to reduce, and no aggregator factory path -- that was severed when QueryShardContext went. Removing create/createInternal from the base classes and their 17 overrides, and the two buildPipelineTree methods, drops 26 files: the concrete Avg/Max/Min/Sum/Stats/ExtendedStats/Percentiles bucket reducers, BucketScript, BucketSelector, BucketSort, CumulativeSum, Derivative, MovAvg, MovFn and SerialDiff, plus BucketMetricsPipelineAggregator and SiblingPipelineAggregator. Every *PipelineAggregationBuilder stays -- they are the published API, they still serialise identically, and they are reachable only through the named-writeable registry, so they remain explicit roots of the prune. PipelineAggregator itself stays too: its nested PipelineTree is part of the InternalAggregation reduce contract that the remaining Internal* result classes still reference. 1770 -> 1744 files, 2866 -> 2837 classes, jar 6,009,739 -> 5,953,346 bytes. 373 tests pass, probe 27/27, stats snapshot byte-identical, 221/221 consumer types present.
lucene-backward-codecs 690,161 B old on-disk codec formats, SPI-loaded lucene-sandbox 232,897 B lucene-misc 129,786 B lucene-memory 58,051 B MemoryIndex, for percolation All four are at zero references in src/main/java, and at zero in fess, fess-suggest, fess-crawler and the five plugin repositories. They were declared directly here rather than arriving transitively, so removing the declarations removes them from the resolved tree -- confirmed with dependency:tree -- and therefore from the Fess distribution, which gets its Lucene through this artifact. backward-codecs deserves a word: Lucene loads it through the codec SPI rather than by symbol, so a reference count of zero is not by itself proof. It reads superseded on-disk index formats, and nothing in this stack opens a Lucene index -- the client speaks HTTP to a remote engine, and Fess does the same through it. Lucene stays otherwise: core, analysis-common, queries, queryparser, join, grouping, highlighter, suggest and spatial-extras are all still referenced, and the downstream repositories compile against analysis, index, queryparser, search and search.join through this artifact. One consequence worth recording: fess-suggest inherits a provided-scope lucene-facet from fess-parent, which used to be satisfied while resolving the modules removed here, and now has to resolve on its own. It does, and nothing in fess or fess-suggest references org.apache.lucene.facet, so there is no runtime exposure -- but a build with a cold local repository now fetches one more artifact. 373 tests pass, probe 27/27, stats snapshot byte-identical. Downstream re-run against this jar: fess 7469, fess-suggest 689, fess-crawler-opensearch 53, all green.
… edges
A library that serialises requests and parses responses has no reason to run
ServiceLoader at class-init. All six META-INF/services files are gone and the
directory with them; three were dead or redundant, three are now registered in
code.
Deleted outright:
* java.util.spi.CalendarDataProvider / IsoCalendarDataProvider. It installed a
process-wide java.util.Calendar override forcing MONDAY / minimal-days-4 for
Locale.ROOT -- a client library embedded in Fess changing Calendar semantics
for the whole JVM. It was also already inert: since JDK 9 the default
java.locale.providers is CLDR,COMPAT, SPI is not consulted, and
WeekFields.of(Locale.ROOT) reports SUNDAY/1 with the provider on the
classpath. A differential probe over fourteen week-based formats is
byte-identical before and after, so this is a no-op removal that also drops
a JVM-wide side effect.
* ErrorOnUnknown / SuggestingErrorOnUnknown. The interface already carried a
default implementation; the provider only appended a Levenshtein "did you
mean" hint. The lookup is now a constant and the provider is gone, taking
its lucene-suggest LevenshteinDistance use with it.
* org.apache.lucene.codecs.PostingsFormat. Ours named only
Completion50PostingsFormat, which lucene-suggest already registers in its own
service file along with the 84, 90 and 99 variants. Strictly redundant.
Registered statically instead of discovered, same contents:
* CompressorRegistry -> NoneCompressor + DeflateCompressor
* MediaTypeRegistry -> XContentType.values() plus the application/* and
application/x-ndjson aliases
* XContentBuilder -> XContentOpenSearchExtension
Each had exactly one or two providers, all shipping in this same artifact, so
the indirection bought nothing and carried a real failure mode this fork has hit
twice: a provider pruned as unreachable, the build still green, the registry
silently empty at runtime. The contract is now compile-time. The runtime probe
keeps checking it, but now asserts registry contents and the absence of the
service files rather than that discovery happened.
Two node-side edges cut, in the same spirit -- deleting the member rather than
accepting the reference:
* BaseFuture.blockingAllowed() existed only to feed `assert blockingAllowed()`,
asserting that a blocking get is not running on a transport thread. This
client has no transport threads. That one call was the only thing keeping
transport/Transports alive.
* The reachability root set no longer over-approximates downstream references
by bare simple name. Verified that no file downstream, and none under
client/**, wildcard-imports a fork package, so every downstream reference is
fully qualified on its import line; inner classes stay pinned to their outer.
The old approximation rooted any fork class whose simple name appeared
anywhere downstream -- Point, Table, Index, Task -- keeping 27 files alive
for no reason, among them task/commons/**, the transport protocol headers,
the server-side http request/response abstraction, node-side rescoring and
the _cat Table renderer.
Behaviour change, recorded: "unknown query [x]" and "Unknown aggregation type
[x]" no longer carry a " did you mean [y]?" suffix. No test in this repository
or downstream asserts on it.
1744 -> 1709 files, 2837 -> 2789 classes, jar 5,953,302 -> 5,893,158 bytes.
373 tests pass, probe 21/21, stats snapshot byte-identical, week-format probe
byte-identical, 221/221 consumer types present.
…t types
Found by ranking every reachable class by how few references hold it alive and
how much would die with it -- a class with one incoming edge and a large subtree
is a bridge, and cutting that one member takes the subtree.
* IndexRequest.extraFieldValues / UpdateRequest.docExtraFieldValues /
upsertExtraFieldValues. The HTTP client never serialises extra field values:
they only ever moved over the binary transport streams, so setting them
through this client was a silent no-op. Nothing downstream or in the tests
names them. The stream read and write keep their version-gated slots so the
wire shape is unchanged; the values are now written as absent. 20 classes,
the whole index/mapper/extrasource packed-array tree.
* ThreadContext imported http/HttpTransportSettings for exactly two members,
SETTING_HTTP_MAX_WARNING_HEADER_COUNT and _SIZE. That class is the node's
HTTP *server* configuration -- bind hosts, port ranges, CORS, pipelining. The
two settings are now defined verbatim in ThreadContext, same keys, same
defaults, same validation.
* IndexRequest and TaskResult used transport/client/Requests for one constant,
INDEX_CONTENT_TYPE, which is MediaTypeRegistry.JSON. Inlined. Requests is a
request-factory class nothing downstream calls.
* SmoothingModel.buildWordScorerFactory() and its three overrides in Laplace,
StupidBackoff and LinearInterpolation. Same shape as the pipeline
aggregators: the builder describes the smoothing model for the request body
and was also carrying the factory that instantiates the node-side scorer.
That took WordScorer, the Laplace/StupidBackoff/LinearInterpolating scorers,
CandidateScorer, NoisyChannelSpellChecker, Correction and FreqTermsEnum.
* CompletionSuggestionBuilder.parseContextBytes(), a package-private helper
with zero callers that resolved context bytes against a field's
ContextMappings into node-side InternalQueryContexts. It was the only thing
holding the completion context-mapping tree.
Checked and deliberately kept: FilterClient, the largest single candidate at
188 KB. filterWithHeader() is real -- SearchEngineClient and FesenClient both
override and delegate it, and FesenClientTest covers it. Also kept every
response type the ranking surfaced (SnapshotStatus, RepositoriesStats,
NodeCacheStats, RepositoryStatsSnapshot, SegmentReplicationPerGroupStats,
RemoteConnectionInfo, IndexingPressurePerShardStats, StatusCounterStats): those
parse real responses, several of them the nodes-stats document the behaviour
snapshot pins.
1709 -> 1671 files, 2789 -> 2738 classes, jar 5,893,158 -> 5,795,335 bytes.
373 tests pass, probe 21/21, stats snapshot byte-identical, 221/221 consumer
types present.
* PhraseSuggestionBuilder held its tokenLimit default as
NoisyChannelSpellChecker.DEFAULT_TOKEN_LIMIT. That one constant reference was
what survived the previous commit's removal of buildWordScorerFactory and
kept the whole node-side spell-checker alive. Inlined as 10, the value it
already had, and WordScorer, CandidateScorer, NoisyChannelSpellChecker,
Correction and FreqTermsEnum go with it.
* OpenSearchExecutors.newSinglePrioritizing() and newAutoQueueFixed() had no
callers. They built the cluster-state applier's priority executor and the
node's auto-sizing search queue. Removing them takes
PrioritizedOpenSearchThreadPoolExecutor, PrioritizedCallable,
PrioritizedRunnable and QueueResizingOpenSearchThreadPoolExecutor.
The client's own ThreadPool is untouched: newFixed, newScaling, newResizable,
newDirectExecutorService, the thread factories and all three ExecutorBuilders
stay, because ThreadPool genuinely builds executors with them.
Considered and left: BigArrays' CircuitBreakerService plumbing. BigArrays is only
ever instantiated as NON_RECYCLING_INSTANCE with a null breaker service, so the
breaker path is inert -- but it sits in the array allocation that BytesStreamOutput
really uses, and the whole subtree is 6 classes and 11 KB. Not worth the risk of
disturbing serialisation for that.
1671 -> 1657 files, 2738 -> 2704 classes, jar 5,795,335 -> 5,735,739 bytes.
373 tests pass, probe 21/21, stats snapshot byte-identical, 221/221 consumer
types present.
This library no longer aims to cover the whole OpenSearch API surface. The
keep-rule is now: an operation survives only if one of the eight consumer
repositories uses it. This repository's own integration tests are no longer
sufficient grounds on their own, so several operations that OpenSearch3ClientTest
exercised go, and their tests go with them.
68 of the 99 operations are removed; 31 remain:
reads get, multiGet is gone, search, searchView is gone, streamSearch,
createPit, deletePits, explain/fieldCaps/termVectors are gone
writes index, update, delete, bulk
indices create, exists, open, close, refresh, flush, getIndex, getMappings,
putMapping, getSettings, aliases, getAliases, analyze
cluster health, nodesStats, nodesHotThreads
structure admin, cluster, indices, settings, filterWithHeader,
getRemoteClusterClient
Whole families go: snapshots and repositories, stored scripts, ingest pipelines,
index templates, tasks, cluster state/stats/reroute/settings/allocation-explain,
search shards, views, data streams, ingestion state, remote store, segment
replication, wlm stats, recovery, shard stores, indices stats, force merge,
clear cache, validate query, resize/rollover/scale, upgrade and upgrade status,
resolve index, scroll (searchScroll/clearScroll), multi-search, multi-get,
multi-term-vectors, term vectors, explain, field caps, get field mappings,
reindex, update-by-query, delete-by-query, simulate pipeline, remote info and
the main/info action.
How the used-set was derived, and the trap in it: SearchEngineClient and
FesenClient implement Client and override ~68 and ~64 methods, and an override
is not evidence that Fess calls the operation. Those two files also contain real
logic that does, so they could not simply be excluded. The scan blanks out their
pure delegation overrides -- an @OverRide whose body only forwards to the wrapped
client -- and then drops the imports left behind, which otherwise still match a
type pattern and make every delegated operation look used. Before that second
step the scan reported 43 operations used; after it, 33. searchView and
listViewNames were a third case: their overrides throw
UnsupportedOperationException("Not implemented yet"), which is not usage either.
Spot-checked the operations the admin screens and jobs reach, since those are
easy to miss: refresh (19 call sites), aliases (7), flush, open and close
(AdminMaintenanceAction, and SearchEngineClient.close()), getSettings,
nodesStats (LoadControlMonitorTarget, SystemMonitorTarget), nodesHotThreads
(HotThreadMonitorTarget) -- all kept. forceMerge, reindex, indices stats,
cluster state, update settings, validate query and rollover have zero call sites
and zero request/response type references anywhere in the eight repositories.
Scroll is kept out on the same evidence: everything named "scroll" downstream --
SearchEngineUtil.scroll, EsAbstractBehavior.scrollForDelete,
FesenClient.scrollForDelete -- goes through pitSearch. The Scroll type itself
stays; it is used as a keep-alive holder.
Tests: the 24 Http*ActionTest classes for removed actions are deleted outright.
The five integration suites keep their classes and lose only the test methods
that exercised a removed operation, found by compiling and mapping each error to
its enclosing method rather than by name. PIT coverage survives in
test_pit_search_after_pagination and test_pit_search_after_sort_value_round_trip;
the two PIT tests that went were the ones calling getAllPits.
Also removed Lucene.readTopDocs/writeTopDocs, the binary TopDocs transport
format, which had no callers and was the only thing holding
org/apache/lucene/search/grouping/CollapseTopFieldDocs. The org/apache/lucene
tree is now empty and gone.
1657 -> 1275 files, 2704 -> 2204 classes, jar 5,735,739 -> 4,645,600 bytes.
101 -> 27 Http*Action classes. This repository's own tests are 373 -> 271.
Probe 21/21, stats snapshot byte-identical, 219/219 consumer types present.
Downstream, all green: fess 7469, fess-suggest 689, fess-crawler-opensearch 53,
gsuite 240, mcp 597, multimodal 74, classic-api 80, v1-api 207.
This reverts commit e687609. The scope was wrong. Removing what Fess does not use was premised on this library still standing as itself, and the Http*Action classes are what that means: they are the contract that makes fesen-httpclient an HTTP implementation of the OpenSearch client rather than a Fess-private adapter. The "keep only what a consumer uses" rule applies to OpenSearch internals this library never exercises -- node-side execution, reduce phases, transport plumbing -- not to the library's own action implementations. So all 101 Http*Action classes come back, with their ActionTypes, their request/response/builder types, the client interface methods and the tests: the 24 Http*ActionTest classes and the integration-suite methods. upgrade comes back too; it was dropped on the belief that nothing implemented it, and HttpUpgradeAction does. The analysis behind the reverted commit still stands as a fact about the consumers -- only 31 of the 99 operations are reachable from fess, fess-suggest, fess-crawler-opensearch or the five plugin repositories, and the other 68 exist for callers outside this workspace. That is a statement about who uses the library, not a licence to shrink its API. 373 tests, probe 21/21, stats snapshot byte-identical, 221/221 consumer types.
Lucene.readTopDocs and Lucene.writeTopDocs serialise a TopDocs plus its max score to and from a StreamInput/StreamOutput -- the format the node's search phases use to ship per-shard results between nodes. Nothing calls either of them: this client parses search responses from JSON. Verified with the full set of 101 Http*Action classes and every client operation present, so this is not a consequence of anything the reverted commit removed. Removing them takes with them: org/apache/lucene/search/grouping/CollapseTopFieldDocs 256 lines common/lucene/search/TopDocsAndMaxScore 55 lines common/lucene/Lucene 89 lines removed CollapseTopFieldDocs existed only for package-private access to Lucene's grouping internals and was reached from nothing else, so org/apache/lucene is now empty and gone -- one less package this artifact splits with lucene-core, lucene-grouping and lucene-queryparser. TopDocsAndMaxScore had one other mention, an unused import in NestedQueryBuilder, now dropped. indices/replication/common/ReplicationLuceneIndex stays. It looked like part of the same group, but it is reached from SegmentReplicationState, RecoveryState and ReplicationState -- the response types of the recoveries and segmentReplicationStats operations, both of which this library implements. 1657 -> 1655 files, 2704 -> 2700 classes, jar 5,735,739 -> 5,726,805 bytes. 373 tests pass, probe 21/21, stats snapshot byte-identical, 221/221 consumer types present.
Reachability calls a class "reachable" whether a hundred things use it or one
assert mentions it, so it never surfaces the shape that has dominated this prune:
a large class kept alive by one small member. This makes the hunt mechanical.
For every class, rank by the bytes that would die with it divided by the number
of distinct (referrer, member) pairs holding it alive, re-ranking after each cut
because every cut changes the weights below it.
Three shapes the class graph structurally cannot see, all now swept:
RESOURCES -- not in the class graph at all.
* tasks/task-index-mapping.json, the mapping for a node's .tasks index. Its
only mention anywhere was a javadoc sentence in Task.java. Nothing loads it.
src/main/resources is now just the two META-INF licence files, and the jar's
only other non-class entries are the manifest and the maven descriptor.
A LARGE CLASS HELD BY ONE SMALL MEMBER -- the ratio hunt.
* JarHell, 14,345 bytes of classpath duplicate/split-package scanner, alive
because PluginInfo called JarHell.checkVersionFormat -- thirteen lines that
parse a version string. Moved to PluginInfo; JarHell deleted.
* InFlightShardSnapshotStates, held only by SnapshotsInProgress's private
assertConsistentEntries, itself called only from `assert`.
* SecureSetting, held only by `assert this instanceof SecureSetting || ...`
in Setting's constructor.
* RoutingTableIncrementalDiff, held only by RoutingTable.incrementalDiff,
which nothing calls -- cluster-state diffs are a node-side publication
concern.
* BulkShardRequest, held only by BulkRequest's `timeout =
BulkShardRequest.DEFAULT_TIMEOUT`, inherited from ReplicationRequest.
* MoreLikeThisQuery and XMoreLikeThis, 23 KB of Lucene more-like-this term
selection, held only by seven default constants MoreLikeThisQueryBuilder
copied out of them.
* TermVectorsWriter, held only by TermVectorsResponse.setFields -- the
server-side path that *builds* a term-vectors response.
* InternalExtendedStats and InternalGeoCentroid, node-side reduce
implementations held only by their nested Fields classes of JSON field
names, which ParsedExtendedStats and ParsedGeoCentroid import. The names
moved to the Parsed* classes that read them.
* TransportRequestOptions, held only by ActionType.transportOptions and one
override in BulkAction. That method configures the binary transport channel
-- compression, timeouts, request type. Nothing calls it; an HTTP client has
no channel to configure.
* RemoteConnectionStrategy, held only by RemoteConnectionInfo reading its enum
off a stream and discarding it, and by modeType() in the ModeInfo interface,
which nothing implements. HttpRemoteInfoAction says as much in a comment: it
returns an empty connection list because RemoteConnectionInfo cannot be
constructed externally.
A CLASS HELD BY AN INHERITANCE EDGE -- invisible to both of the above.
* TransportMessage. Every response DTO inherits from ActionResponse and every
request from ActionRequest, so the transport base classes were structurally
reachable and always would be. TransportMessage contributed exactly one
thing: remoteAddress, a TransportAddress, which is meaningless to an HTTP
client and read nowhere. TransportResponse and TransportRequest now
implement Writeable directly and TransportMessage is gone.
The wider flattening was measured and rejected. Dropping TransportResponse
and TransportRequest as well would remove 6 classes and 7,321 bytes but
costs the TaskAwareRequest contract, and that is genuinely used --
GetTaskRequest, ReplicationRequest and AbstractBulkByScrollRequest all read
or copy the parent task. Kept.
Also measured and rejected: removing the aggregation reduce() family. 8 classes
and 37,972 bytes for 34 members edited, including a public abstract method on
InternalAggregation -- about 1.1 KB per member, an order of magnitude worse than
anything above.
Kept with reasons: RemoteClusterAware, because an earlier round had already
reduced it to 67 lines holding buildRemoteIndexName and
REMOTE_CLUSTER_INDEX_SEPARATOR -- it *is* the small utility, and the cluster:index
convention is real over HTTP since the server returns such names.
Lucene.readSortField/writeSortField, because SearchHits uses them.
BulkRequestParser, because BulkRequest.add parses an NDJSON bulk body, which is
real client API.
No package-info.java files exist anywhere in the tree, and the remaining asserts
that call into a static helper all name classes kept for other reasons.
1657 -> 1639 files, 2704 -> 2663 classes, jar 5,735,739 -> 5,644,760 bytes.
373 tests pass, probe 21/21, stats snapshot byte-identical, 221/221 consumer
types present. All 101 Http*Action classes untouched.
…stractClient
HttpAbstractClient implements Client directly -- it has never extended OpenSearch's
AbstractClient. The only thing in the whole fork that referenced AbstractClient or
FilterClient was one twelve-line method:
public Client filterWithHeader(Map<String, String> headers) {
return new FilterClient(this) { ... }; // FilterClient extends AbstractClient
}
FilterClient is a 3 KB delegating wrapper built on AbstractClient, which is
OpenSearch's node-side implementation of every Client method in terms of
execute() -- 185,187 bytes of class files, 91,731 of source. This class already
implements all of those methods the same way, so the wrapper is now an anonymous
subclass of HttpAbstractClient and the two classes are gone.
Behaviour is unchanged. doExecute still merges the headers into the thread
context with ThreadContext.stashAndMergeHeaders and runs the original client's
doExecute inside that scope, and close() still closes the wrapped client, which
is what FilterClient.close() did.
Basing the wrapper here means every Client method needs an implementation on this
class, so searchView, listViewNames and prepareStreamSearch move down from
HttpClient. They were plain execute() calls with nothing HttpClient-specific
about them.
Why the previous round kept this: the survivor list said "filterWithHeader is
delegated by SearchEngineClient and FesenClient and covered by FesenClientTest".
That justifies keeping the *operation*, which is part of the Client contract and
stays. It says nothing about the implementation -- and a delegating override is
exactly the false positive this prune has already been caught by once.
SearchEngineClient:3396 and FesenClient:608 both just forward to HttpClient,
which is where the work actually happens.
1639 -> 1634 files, 2663 -> 2652 classes, jar 5,644,760 -> 5,615,308 bytes
(the 185 KB is uncompressed class bytes; the jar stores them deflated).
373 tests pass, probe 21/21, stats snapshot byte-identical, 221/221 consumer
types present. All 101 Http*Action classes untouched.
Suggest kept a hasScoreDocs flag, recomputed in both constructors as
filter(CompletionSuggestion.class).stream().anyMatch(CompletionSuggestion::hasScoreDocs)
and exposed through hasScoreDocs(). It exists for the node's fetch phase, which
decides whether suggestions carry score docs that still need fetching. Nothing
reads it -- not this repository, not its tests, not any of the eight consumer
repositories.
That one flag was the only thing holding CompletionSuggestion, 32 KB across five
classes. Completion suggestions are not parseable here in any case:
HttpClient.getDefaultNamedXContents() has the suggestion entries commented out,
so Suggestion.fromXContent never produces one.
Kept, with the reason, since it looked like the same shape: TermVectorsFields,
held only by TermVectorsResponse.getFields(). Over HTTP that always returns null
-- HttpTermVectorsAction does not parse term_vectors and nothing calls
setTermVectorsField -- but the binary path is internally consistent: the
StreamInput constructor reads the blob, writeTo writes it, and toXContent renders
it through getFields(). Removing it would mean deleting a rendering branch from a
public response type and foreclosing the HTTP action ever learning to parse term
vectors. That is a gap in the action, not dead weight in the response.
1634 -> 1633 files, 2652 -> 2647 classes, jar 5,615,308 -> 5,602,124 bytes.
373 tests pass, probe 21/21, stats snapshot byte-identical, 221/221 consumer
types present.
… check
Two more of the structural shape, found while measuring whether the empty
ClusterState that HttpClusterStateAction fabricates was holding a large subtree.
It is not (see below), but two of its neighbours were holding one.
* ResponseCollectorService implemented ClusterStateListener. That is how a node
maintained the statistics: clusterChanged(ClusterChangedEvent) added and
removed per-node entries as the cluster state moved. Everything in this tree
uses only the nested ComputedNodeStats, which AdaptiveSelectionStats carries
in the _nodes/stats payload -- AdaptiveSelectionStats names
ResponseCollectorService.ComputedNodeStats six times and the outer class
never. Dropping the interface and clusterChanged takes ClusterStateListener
and ClusterChangedEvent with them.
* ActiveShardCount.enoughShardsActive, in both overloads, one taking a
ClusterState and one an IndexShardRoutingTable. ActiveShardCount is a real
request parameter -- wait_for_active_shards, set on a dozen request types and
read by as many Http*Actions -- but evaluating whether enough shards are
actually active is the node's job, and nothing here calls it.
1633 -> 1631 files, 2647 -> 2645 classes, jar 5,602,124 -> 5,595,626 bytes.
373 tests pass, probe 21/21, stats snapshot byte-identical, 221/221 consumer
types present.
A member-level reachability run over the packaged jar - rooted at this repo's main and test sources plus the main and test classes of all eight downstream consumers - found members of the forked OpenSearch types that nothing reaches. They are residue of the import, not behaviour: a dead setter means no Http*Action ever read the field, so the parameter was never sent. This slice removes the first batch. The deletion set is closed under the bytecode reference graph, so nothing left behind can still point into it, and it excludes every member that could change behaviour silently: - anything overriding or overridden anywhere in the hierarchy - abstract and interface declarations, including the Client interface family, which is this library's own API surface rather than fork residue - compile-time constants, whose reads javac inlines out of the bytecode - a type's last constructor, which javac would replace with an implicit no-arg one - field initialisers, which bytecode attributes to the constructor No functional change. The 101 Http*Action classes and the HTTP behaviour are untouched. Verified: 373 tests pass; the 21-check runtime probe passes; the _nodes/stats rendering snapshot is byte-identical (405 lines, md5 10f57efde06c6745e3c724cdcd31273f); all 221 fork types named by the eight downstream repos still resolve. Jar 5,595,626 -> 5,351,504 bytes.
Second batch of the same reachability sweep, under the same guards: the set is closed under the bytecode reference graph, and overrides, abstract declarations, the Client interface family, compile-time constants, last constructors and field initialisers are all held back. No functional change. Verified: 373 tests pass; the 21-check runtime probe passes; the _nodes/stats rendering snapshot is byte-identical (405 lines, md5 10f57efde06c6745e3c724cdcd31273f); all 221 fork types named by the eight downstream repos still resolve. Jar 5,351,504 -> 5,224,897 bytes.
Third batch of the same reachability sweep, under the same guards. No functional change. Verified: 373 tests pass; the 21-check runtime probe passes; the _nodes/stats rendering snapshot is byte-identical (405 lines, md5 10f57efde06c6745e3c724cdcd31273f); all 221 fork types named by the eight downstream repos still resolve. Jar 5,224,897 -> 5,164,069 bytes.
Final batch of the reachability sweep.
One extra guard earned by this batch: HttpScaleIndexAction reads ScaleIndexRequest
via reflection (getMethod("getIndex") / getMethod("isScaleDown")), which no bytecode
reachability analysis can see. Deleting those two getters compiled cleanly and broke
four tests. Every reflective lookup in the repo was then audited by resolving the
looked-up name against the types named in the same file; these two are the only ones
that reach a forked type, and both are now held back.
No functional change.
Verified: 373 tests pass; the 21-check runtime probe passes; the _nodes/stats
rendering snapshot is byte-identical (405 lines, md5 10f57efde06c6745e3c724cdcd31273f);
all 221 fork types named by the eight downstream repos still resolve.
Jar 5,164,069 -> 5,141,158 bytes.
…ds reflectively HttpGetFieldMappingsAction does not call this constructor; it looks it up with getDeclaredConstructor(Map.class) and calls setAccessible(true), because the constructor is package-private. No reachability analysis can see that, and the earlier audit of reflective access only matched name literals passed to getMethod and getDeclaredField, so a constructor selected by a Class<?>[] of parameter types slipped through. It is functionality: it is how the action turns a parsed field-mappings response into its response object. The local suites do not cover it - only the Docker suites, which run against real servers, caught it, failing test_get_field_mappings and test_create_knn_field across all five engine versions. Restored byte-identically. The reflective surface of the library's main sources is now three sites in total and all are accounted for: getMethod on ScaleIndexRequest (two getters, already held back), this constructor, and Array.newInstance in HttpAnalyzeAction, which creates an array rather than looking up a member. Verified: 373 tests pass; the 21-check runtime probe passes; the _nodes/stats rendering snapshot is byte-identical (405 lines, md5 10f57efde06c6745e3c724cdcd31273f).
These are not dead by reachability - they are purposeless. They exist to rebuild a DTO from the binary transport format, and this client speaks JSON over HTTP; it never reads a node-to-node stream. They are the subset whose deletion needs no other change: nothing in the jar references them, by call or by method handle, so no Writeable.Reader and no super(in) chain is left pointing at a missing constructor. Three things are deliberately excluded, and being referrer-free happens to establish all three on its own: - the QueryBuilder family (48 of the 182 referrer-free constructors). AbstractQueryBuilder implements NamedWriteable and Fess depends on the contract: HybridQueryBuilder calls readNamedWriteableList and writeNamedWriteableList, and KnnQueryBuilder serialises too. - every DTO the actions that really do round-trip the wire format touch. A constructor reachable from one of those actions has a referrer by definition, so none is in here. - ActionType.getResponseReader and anything its readers point at, for the same reason: a stored method handle is a reference and is counted as one. writeTo is not touched. It cannot be removed the same way: 1,017 types inherit Writeable, so dropping the method means changing what those types implement, which is interface surgery rather than member pruning and belongs in its own change. Verified: 373 tests pass; the 21-check runtime probe passes; the _nodes/stats rendering snapshot is byte-identical (405 lines, md5 10f57efde06c6745e3c724cdcd31273f) - that snapshot renders a real response through these DTOs, so it is the direct check that parsing is unaffected; all 221 fork types named by the eight downstream repos resolve. Jar 5,141,216 -> 5,115,943 bytes.
… slice Removing a subclass constructor removes its super(in) call, which is what kept the parent's StreamInput constructor anchored. Re-running the same referrer analysis over the rebuilt jar frees another round. Same exclusions as before; the QueryBuilder family and everything the wire-format actions reach are still held back. Verified: 373 tests pass; the 21-check runtime probe passes; the _nodes/stats rendering snapshot is byte-identical (405 lines, md5 10f57efde06c6745e3c724cdcd31273f). Jar 5,115,943 -> 5,110,902 bytes.
Third and final round of the same referrer analysis; it converges here. 517 StreamInput constructors remain, all of them still anchored - by a Writeable.Reader method handle, a super(in) chain, or a live call - and removing those needs the Writeable interface work rather than member deletion. Verified: 373 tests pass; the 21-check runtime probe passes; the _nodes/stats rendering snapshot is byte-identical (405 lines, md5 10f57efde06c6745e3c724cdcd31273f). Jar 5,110,902 -> 5,109,407 bytes.
…hable Member-by-member surgery inside a class that is entirely dead is the wrong operation, so these go as whole files. The compiler decided which ones: all 165 candidates were deleted at once and any whose type a surviving file still names was restored, repeated to a fixed point. 116 came back; 49 stayed out. Three files lost an import line rather than a member: their only use of a deleted type had already gone with an earlier slice, and an unused import of a package that no longer exists is an error rather than a warning. Verified: 373 tests pass; the 21-check runtime probe passes; the _nodes/stats rendering snapshot is byte-identical (405 lines, md5 10f57efde06c6745e3c724cdcd31273f); all 221 fork types named by the eight downstream repos still resolve. Files 1,631 -> 1,582. Jar 5,109,407 -> 4,977,043 bytes.
The rest of the fully-dead classes are nested inside types that are still used, so the file has to stay and only the nested declaration goes. Same oracle as the whole-file slice: delete all 297 candidates, drop the ones javac still resolves, repeat to a fixed point. 249 came back; 48 stayed out. Targets contained inside another target are skipped, so a dead type nested in a dead type is removed once with its parent rather than twice. Verified: 373 tests pass; the 21-check runtime probe passes; the _nodes/stats rendering snapshot is byte-identical (405 lines, md5 10f57efde06c6745e3c724cdcd31273f); all 221 fork types named by the eight downstream repos still resolve. Jar 4,977,043 -> 4,920,282 bytes.
Deleting a class removes its references, so the candidate set has to be re-derived rather than assumed. Running the whole-file and nested-type candidates together through the same compile oracle frees eight more: the BigArrays element interfaces whose only implementations had already gone, and three interfaces left without implementors. Re-running the oracle after this converges - nothing further falls out. Verified: 373 tests pass; the 21-check runtime probe passes; the _nodes/stats rendering snapshot is byte-identical (405 lines, md5 10f57efde06c6745e3c724cdcd31273f); all 221 fork types named by the eight downstream repos still resolve. Jar 4,920,282 -> 4,915,121 bytes.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Fess ships
org.opensearch:opensearch3.8.0 — 17.09 MiB, 8,578 classes — purely for its request/response DTOs and query builders. It has never run an OpenSearch node; the embedded runner went away in 15.9. This forks the classes it needs intoorg.codelibs.fesen.opensearch.*and prunes the rest.The forked jar is 4.69 MiB against the 17.09 MiB it replaces. Across the Fess distribution that is 23.28 MiB off
WEB-INF/lib, 229 jars down to 203. Apache Lucene stays as the real artifact. All 101Http*Actionclasses are intact — this is still a general-purpose HTTP implementation of the OpenSearch client, not a Fess-private adapter.Why not shrink the jar mechanically instead
Both cheaper options were measured first.
Class-level shrinking is useless: the reference graph is one blob, and a single seed class —
SearchResponse,QueryBuilders, anything — drags in 95% of the jar. Even emptying every method body leaves 59.6% standing.Member-level tree shaking with ProGuard does better — 8.45 MiB — but produces something that cannot ship. It silently drops 6 of the 10
META-INF/servicesproviders, includingXContentOpenSearchExtensionandServerCompressorProvider, because theServiceLoader.loadcalls that need them live inopensearch-core, a library jar it does not trace into;XContentBuilderthen fails on every request. It removes enum constants untilSetting$Propertyis no longer an enum andFeatureFlagsdies in its static initialiser. And it deletes public DTO methods production code happens not to call: 74 of this repository's own unit tests fail withNoSuchMethodErroron things likeBulkRequest.pipeline. Fess has a plugin API, and third-party plugins cannot be in anyone's root set. A fork moves those failures to compile time, where they belong.What was cut
Node-side machinery, in layers, each severed at a load-bearing reference:
ClusterSettings—HttpNodesStatsActioninstantiated threeBUILT_IN_*registries only to build aSearchRequestStats. Those arrays nameSettingconstants in almost every server package, which is what keptgateway,plugins,telemetry,repositories,node,persistent,identity,discoveryandhttpalive. Behaviour-neutral: the value replaced round-tripped through a constructor that pre-populates every field with zeros.rest/**— 181 files, kept alive by four response DTOs reading one constant,RestSearchAction.TOTAL_HITS_AS_INT_PARAM, insidetoXContent.QueryShardContext—AbstractQueryBuilder.doToQueryand 48 overrides. A client serialises query builders to JSON; it never compiles them into Lucene queries.index/mapper/**.META-INF/services— all six files gone. Three deleted outright (ErrorOnUnknownonly reworded an error;IsoCalendarDataProviderregistered a JVM-widejava.util.Calendaroverride that has been inert since JDK 9 madeCLDR,COMPATthe default — verified byte-identical output across fourteen week-based formats; ourPostingsFormatfile duplicated lucene-suggest's). The other three became static registration, which keeps the behaviour and removes the failure mode that bit this project twice: a provider silently pruned, build green, registry empty at runtime.Lucene.readTopDocs/writeTopDocs, uncalled, takingCollapseTopFieldDocsandTopDocsAndMaxScorewith it.org/apache/luceneis now empty and gone, one less split package against lucene-core.Transportsheld by oneassert;JarHell(14 KB) by one static version-format check;MoreLikeThisQuery+XMoreLikeThis(23 KB) by seven copied constants;AbstractClient(185 KB — OpenSearch's node-sideClient) by the twelve-linefilterWithHeader, which now wrapsHttpAbstractClientinstead.Dependencies removed (4.34 MiB off the distribution):
log4j-core,jts-core,spatial4j,RoaringBitmap,HdrHistogram,t-digest,jopt-simple,jzlib,jsr305, and four Lucene modules.zstd-jnialso leaves this pom, though it stays in the war via commons-compress.joda-timestays deliberately: replacingJoda.forPatternin Fess's taglib changes date parsing for admin-configured patterns —hhwithoutazeroes the time,wwis ignored,YYYY-'W'ww-eshifts a day — and silently, so 0.61 MiB is the price of not doing that.Six methods are stubbed with
UnsupportedOperationException, all coordinator-side reduce paths a client never reaches. Nothing else is stubbed.Licensing
Forked sources keep their Apache-2.0 headers;
license-maven-pluginandformatter-maven-pluginare scoped to this project's own sources so they cannot rewrite them.META-INF/NOTICE.txtcarries a derivation statement above OpenSearch's own NOTICE;META-INF/LICENSE.txtis verbatim.Verification
XContentBuilderextension writers, and serialisation of query, aggregation, sort, highlight, suggest and collapse builders. Static reachability alone had deleted the ServiceLoader providers with the build still green; this is what caught it._nodes/statsfixture throughHttpNodesStatsAction— byte-identical to the pre-fork output through every round, which is what proves theSearchRequestStatsremoval changed nothing._nodes/statswith all 30 sections. ZeroNoClassDefFoundError,ClassNotFoundException,NoSuchMethodErrororServiceConfigurationErroracross all 13 log files.Dead members
A member-level reachability pass with the full consumer root set found 7,436 dead members inside classes that must stay. 4,577 of them are gone, along with 174
StreamInputconstructors that had no referrer at all — a DTO rebuilt from the binary transport format is of no use to a client that speaks JSON — and 57 whole files plus 48 nested types that nothing reached.The 2,859 declined each have a reason, and the four largest are worth recording because they are ways bytecode reachability over-reports against source:
@Overrideplus 81 overridden elsewhere, and 345 implicit constructors with no source to delete.implementsclause —CompositeIndicesRequest,RealtimeRequest,WriteResponse. ProGuard prunes an interface list; source cannot without changing what a type implements.@PublicApi, but 587 files are annotated with it.The raw list also wanted 56
Client/IndicesAdminClientdefault methods — this library's own API — which is why it was treated as a starting point and not an instruction.What remains and stays:
writeTo(679 implementations, held by one virtual call onWriteable.writeTo), the server-sidereducemembers and theVersion-gated overrides. Removing those means changing what a type implements or giving an interface a default that throws, which trades a compile-time guarantee for a runtime failure. That is a design change, not pruning, and it is deliberately not in this PR.Across the whole of this work, 139
implements/extendslines are touched and every one is a deletion — not a single clause was modified, no default or stub was added, and no signature was relaxed.Ordering
Depends on #42. Downstream repositories —
fess-suggest,fess-crawler,fessand five plugin repositories — have matching branches and must be released together with this.