From 8853548f625f20c313c3de9626ef1f845ede15a3 Mon Sep 17 00:00:00 2001 From: Gabriel Cruz Date: Mon, 11 May 2026 15:07:32 -0300 Subject: [PATCH 1/2] feat: provider record spillover --- kad-dht/provider-record-spillover.md | 241 +++++++++++++++++++++++++++ 1 file changed, 241 insertions(+) create mode 100644 kad-dht/provider-record-spillover.md diff --git a/kad-dht/provider-record-spillover.md b/kad-dht/provider-record-spillover.md new file mode 100644 index 00000000..69891ad4 --- /dev/null +++ b/kad-dht/provider-record-spillover.md @@ -0,0 +1,241 @@ +# Provider Record Spillover + +| Lifecycle Stage | Maturity | Status | Latest Revision | +|-----------------|---------------|--------|-----------------| +| 1A | Working Draft | Active | r0, 2026-05-11 | + +Authors: [@gmelodie] + +Interest Group: [@mxinden, @guillaumemichel, @MarcoPolo] + +[@gmelodie]: https://github.com/gmelodie +[@mxinden]: https://github.com/mxinden +[@guillaumemichel]: https://github.com/guillaumemichel +[@MarcoPolo]: https://github.com/MarcoPolo + +See the [lifecycle document][lifecycle-spec] for context about the maturity level +and spec status. + +[lifecycle-spec]: https://github.com/libp2p/specs/blob/master/00-framework-01-spec-lifecycle.md + +--- + +## Overview + +This document specifies an optional extension to the [libp2p Kademlia DHT +specification][kad-spec] that addresses provider record hotspots. When a key is +popular, the `k` closest nodes concentrate all `ADD_PROVIDER` traffic for it +and can become overloaded. This extension lets nodes enforce per-key provider +limits and signal rejection to advertisers, and lets advertisers spill over to +progressively farther peers from their lookup path rather than repeatedly +hammering the same overloaded nodes. + +The extension is opt-in and backward compatible. Nodes that have not enabled it +behave exactly as before. + +## Definitions + +**Provider record capacity**: the maximum number of distinct providers a node +is willing to store for a single key. + +**Rejection**: a node signalling to an advertiser that it will not store the +provider record for the given key. + +**Spillover**: the behavior of an advertiser that, upon failing to reach the +replication target `k` from a batch of peers, continues advertising to +progressively farther peers discovered during the iterative lookup. + +**Spillover round**: one batch of `ADD_PROVIDER` requests sent to the next +group of candidates during a spillover. + +## Motivation + +In the base Kademlia DHT protocol, provider advertisement targets the `k` +closest peers to the content key unconditionally. For popular keys, these peers +accumulate unbounded provider records and act as a permanent bottleneck. The +base spec provides no mechanism for a peer to decline an `ADD_PROVIDER` or for +the advertiser to store the record elsewhere. + +This extension solves both sides of the problem: + +1. **Server-side protection**: nodes can enforce a per-key limit and reject + `ADD_PROVIDER` requests once the limit is reached. +2. **Client-side resilience**: advertisers react to rejections by backtracking + through the lookup path, widening the set of peers that store the record + until the replication target is met. + +## Provider Record Limits + +A node MAY enforce a maximum number of distinct providers per key, +`maxProvidersPerKey`. When set, the node counts the distinct provider peer IDs +already stored for the key. If that count is at or above `maxProvidersPerKey` +and the `ADD_PROVIDER` sender is not already a known provider for that key +(i.e., it is not a re-advertisement), the node MUST reject the request. + +Re-advertisements from a provider that is already stored for the key are always +accepted, regardless of the limit, so that existing providers can refresh their +records. + +Nodes SHOULD also enforce coarser bounds such as total provider records stored +(`providerRecordCapacity`) and total distinct keys for which records are held +(`providedKeyCapacity`), but those are outside the scope of the rejection +signalling defined here. + +Nodes MAY also reject an `ADD_PROVIDER` due to other policies, +not only capacity, which is outside of the scope of the document. + +## ADD_PROVIDER Rejection + +### Response signalling + +Support for this extension is optional. A node that supports it MUST include a +`providerStatus` field (field 11, see [Protobuf](#protobuf)) in its +`ADD_PROVIDER` response: + +- `accepted (0)` — the record was stored. +- `rejected (1)` — the record was not stored. + +An absent `providerStatus` field — whether because the responding node does not +support this extension or due to a timeout — MUST be interpreted as `accepted` +by the advertiser. + +`providerStatus` is a response-only field. Advertisers MUST NOT set it in +`ADD_PROVIDER` requests. Receiving nodes MUST ignore `providerStatus` if it is +present in an incoming request, to avoid interoperability ambiguity. + +## Spillover Algorithm + +### Overview + +When an advertiser sends `ADD_PROVIDER` to a batch of peers and the replication +target `k` has not been reached after that batch, it performs a **spillover +round**: it moves to the next group of peers that are farther from the key, as +discovered during the initial iterative lookup. This continues until either: + +- the replication target `k` has been reached (counting peers that accepted or + did not respond), or +- the advertiser has exhausted all peers discovered during the lookup. + +### Lookup phase + +Before advertising, the advertiser performs a full iterative lookup (using +`FIND_NODE`) for the content key as usual, collecting all peers encountered. +The resulting candidate set is sorted in ascending order of XOR distance to +the key. + +### Advertisement phase + +The advertiser splits the sorted candidate list into chunks of size `α` (the +concurrency parameter). It then iterates over these chunks from closest to +farthest: + +1. Send `ADD_PROVIDER` to all peers in the current chunk concurrently. +2. Collect responses. Count peers that accepted or did not respond (absent + `providerStatus` or `providerStatus = accepted`) towards the replication + target. Peers that explicitly rejected do not count. +3. If the replication target `k` has been reached, stop. +4. If the target has not been reached, continue to the next chunk (spillover + round). +5. If no more chunks remain, stop. + +**Note:** The timeout per peer in a spillover round SHOULD be slightly larger +than the base timeout to account for dial overhead to less-familiar peers. + +### Relationship to base advertisement + +This algorithm is a generalisation of the base `ADD_PROVIDER` procedure. When +the closest chunk alone satisfies the replication target, behaviour is identical +to the base spec. Spillover only occurs when `k` has not been reached after a +chunk. + +## Protobuf + +The following changes extend the `Message` type defined in the [kad-dht +spec][kad-spec]: + +```protobuf +syntax = "proto2"; + +message Record { + bytes key = 1; + bytes value = 2; + string timeReceived = 5; +} + +message Message { + enum MessageType { + PUT_VALUE = 0; + GET_VALUE = 1; + ADD_PROVIDER = 2; + GET_PROVIDERS = 3; + FIND_NODE = 4; + PING = 5; + } + + enum ConnectionType { + NOT_CONNECTED = 0; + CONNECTED = 1; + CAN_CONNECT = 2; + CANNOT_CONNECT = 3; + } + + // Added by this extension. + enum AddProviderStatus { + ACCEPTED = 0; + REJECTED = 1; + } + + message Peer { + bytes id = 1; + repeated bytes addrs = 2; + ConnectionType connection = 3; + } + + MessageType type = 1; + bytes key = 2; + Record record = 3; + repeated Peer closerPeers = 8; + repeated Peer providerPeers = 9; + int32 clusterLevelRaw = 10; // NOT USED + + // Added by this extension. Absent field MUST be treated as accepted. + optional AddProviderStatus providerStatus = 11; +} +``` + +Field 11 is optional; older implementations that do not know about it ignore +it, maintaining full wire-level backward compatibility. + +## Backward Compatibility + +- Nodes that do not support this extension never write `providerStatus` and + are unaffected by receiving it. +- Advertisers that do not implement spillover continue to advertise to the `k` + closest peers; absent `providerStatus` is treated as acceptance. +- No change is made to `GET_PROVIDERS` or any lookup message. + +## Security Considerations + +**False rejections**: an adversary controlling the closest peers to a key could +reject all `ADD_PROVIDER` requests to suppress its advertisement. Spillover +mitigates unanimous rejection: any shortfall below `k` causes the advertiser +to route around those peers and store the record on nodes beyond the adversary's +controlled set. An adversary can limit—but not prevent—replication by mixing accepts +and rejects across chunks, since spillover continues as long as the target is unmet. + +**Slot monopolisation without an eviction policy**: because re-advertisements +from already-stored providers are always accepted regardless of the limit, the +first `maxProvidersPerKey` providers to register for a key can hold their slots +indefinitely simply by refreshing their records. Nodes that fill up later deny +new providers entry, so the stored provider set becomes permanently frozen around +whoever arrived first. Implementations MAY therefore pair `maxProvidersPerKey` +with an eviction policy — for example, evicting the record with the oldest +`timeReceived` when the limit is reached and a new (non-incumbent) provider +advertises — to ensure the stored set can rotate over time and is not captured +by early registrants. + +--- + +## References + +[kad-spec]: https://github.com/libp2p/specs/blob/master/kad-dht/README.md From 52ba189ae794a8b7186991bc9a6b4ac0393f7ca4 Mon Sep 17 00:00:00 2001 From: Gabriel Cruz Date: Tue, 1 Sep 2026 09:08:22 -0300 Subject: [PATCH 2/2] fix: pr comments --- kad-dht/provider-record-spillover.md | 241 ------------------------ src/ipips/ipip-0537.md | 180 ++++++++++++++++++ src/routing/kad-dht.md | 265 ++++++++++++++++++++++++++- 3 files changed, 438 insertions(+), 248 deletions(-) delete mode 100644 kad-dht/provider-record-spillover.md create mode 100644 src/ipips/ipip-0537.md diff --git a/kad-dht/provider-record-spillover.md b/kad-dht/provider-record-spillover.md deleted file mode 100644 index 69891ad4..00000000 --- a/kad-dht/provider-record-spillover.md +++ /dev/null @@ -1,241 +0,0 @@ -# Provider Record Spillover - -| Lifecycle Stage | Maturity | Status | Latest Revision | -|-----------------|---------------|--------|-----------------| -| 1A | Working Draft | Active | r0, 2026-05-11 | - -Authors: [@gmelodie] - -Interest Group: [@mxinden, @guillaumemichel, @MarcoPolo] - -[@gmelodie]: https://github.com/gmelodie -[@mxinden]: https://github.com/mxinden -[@guillaumemichel]: https://github.com/guillaumemichel -[@MarcoPolo]: https://github.com/MarcoPolo - -See the [lifecycle document][lifecycle-spec] for context about the maturity level -and spec status. - -[lifecycle-spec]: https://github.com/libp2p/specs/blob/master/00-framework-01-spec-lifecycle.md - ---- - -## Overview - -This document specifies an optional extension to the [libp2p Kademlia DHT -specification][kad-spec] that addresses provider record hotspots. When a key is -popular, the `k` closest nodes concentrate all `ADD_PROVIDER` traffic for it -and can become overloaded. This extension lets nodes enforce per-key provider -limits and signal rejection to advertisers, and lets advertisers spill over to -progressively farther peers from their lookup path rather than repeatedly -hammering the same overloaded nodes. - -The extension is opt-in and backward compatible. Nodes that have not enabled it -behave exactly as before. - -## Definitions - -**Provider record capacity**: the maximum number of distinct providers a node -is willing to store for a single key. - -**Rejection**: a node signalling to an advertiser that it will not store the -provider record for the given key. - -**Spillover**: the behavior of an advertiser that, upon failing to reach the -replication target `k` from a batch of peers, continues advertising to -progressively farther peers discovered during the iterative lookup. - -**Spillover round**: one batch of `ADD_PROVIDER` requests sent to the next -group of candidates during a spillover. - -## Motivation - -In the base Kademlia DHT protocol, provider advertisement targets the `k` -closest peers to the content key unconditionally. For popular keys, these peers -accumulate unbounded provider records and act as a permanent bottleneck. The -base spec provides no mechanism for a peer to decline an `ADD_PROVIDER` or for -the advertiser to store the record elsewhere. - -This extension solves both sides of the problem: - -1. **Server-side protection**: nodes can enforce a per-key limit and reject - `ADD_PROVIDER` requests once the limit is reached. -2. **Client-side resilience**: advertisers react to rejections by backtracking - through the lookup path, widening the set of peers that store the record - until the replication target is met. - -## Provider Record Limits - -A node MAY enforce a maximum number of distinct providers per key, -`maxProvidersPerKey`. When set, the node counts the distinct provider peer IDs -already stored for the key. If that count is at or above `maxProvidersPerKey` -and the `ADD_PROVIDER` sender is not already a known provider for that key -(i.e., it is not a re-advertisement), the node MUST reject the request. - -Re-advertisements from a provider that is already stored for the key are always -accepted, regardless of the limit, so that existing providers can refresh their -records. - -Nodes SHOULD also enforce coarser bounds such as total provider records stored -(`providerRecordCapacity`) and total distinct keys for which records are held -(`providedKeyCapacity`), but those are outside the scope of the rejection -signalling defined here. - -Nodes MAY also reject an `ADD_PROVIDER` due to other policies, -not only capacity, which is outside of the scope of the document. - -## ADD_PROVIDER Rejection - -### Response signalling - -Support for this extension is optional. A node that supports it MUST include a -`providerStatus` field (field 11, see [Protobuf](#protobuf)) in its -`ADD_PROVIDER` response: - -- `accepted (0)` — the record was stored. -- `rejected (1)` — the record was not stored. - -An absent `providerStatus` field — whether because the responding node does not -support this extension or due to a timeout — MUST be interpreted as `accepted` -by the advertiser. - -`providerStatus` is a response-only field. Advertisers MUST NOT set it in -`ADD_PROVIDER` requests. Receiving nodes MUST ignore `providerStatus` if it is -present in an incoming request, to avoid interoperability ambiguity. - -## Spillover Algorithm - -### Overview - -When an advertiser sends `ADD_PROVIDER` to a batch of peers and the replication -target `k` has not been reached after that batch, it performs a **spillover -round**: it moves to the next group of peers that are farther from the key, as -discovered during the initial iterative lookup. This continues until either: - -- the replication target `k` has been reached (counting peers that accepted or - did not respond), or -- the advertiser has exhausted all peers discovered during the lookup. - -### Lookup phase - -Before advertising, the advertiser performs a full iterative lookup (using -`FIND_NODE`) for the content key as usual, collecting all peers encountered. -The resulting candidate set is sorted in ascending order of XOR distance to -the key. - -### Advertisement phase - -The advertiser splits the sorted candidate list into chunks of size `α` (the -concurrency parameter). It then iterates over these chunks from closest to -farthest: - -1. Send `ADD_PROVIDER` to all peers in the current chunk concurrently. -2. Collect responses. Count peers that accepted or did not respond (absent - `providerStatus` or `providerStatus = accepted`) towards the replication - target. Peers that explicitly rejected do not count. -3. If the replication target `k` has been reached, stop. -4. If the target has not been reached, continue to the next chunk (spillover - round). -5. If no more chunks remain, stop. - -**Note:** The timeout per peer in a spillover round SHOULD be slightly larger -than the base timeout to account for dial overhead to less-familiar peers. - -### Relationship to base advertisement - -This algorithm is a generalisation of the base `ADD_PROVIDER` procedure. When -the closest chunk alone satisfies the replication target, behaviour is identical -to the base spec. Spillover only occurs when `k` has not been reached after a -chunk. - -## Protobuf - -The following changes extend the `Message` type defined in the [kad-dht -spec][kad-spec]: - -```protobuf -syntax = "proto2"; - -message Record { - bytes key = 1; - bytes value = 2; - string timeReceived = 5; -} - -message Message { - enum MessageType { - PUT_VALUE = 0; - GET_VALUE = 1; - ADD_PROVIDER = 2; - GET_PROVIDERS = 3; - FIND_NODE = 4; - PING = 5; - } - - enum ConnectionType { - NOT_CONNECTED = 0; - CONNECTED = 1; - CAN_CONNECT = 2; - CANNOT_CONNECT = 3; - } - - // Added by this extension. - enum AddProviderStatus { - ACCEPTED = 0; - REJECTED = 1; - } - - message Peer { - bytes id = 1; - repeated bytes addrs = 2; - ConnectionType connection = 3; - } - - MessageType type = 1; - bytes key = 2; - Record record = 3; - repeated Peer closerPeers = 8; - repeated Peer providerPeers = 9; - int32 clusterLevelRaw = 10; // NOT USED - - // Added by this extension. Absent field MUST be treated as accepted. - optional AddProviderStatus providerStatus = 11; -} -``` - -Field 11 is optional; older implementations that do not know about it ignore -it, maintaining full wire-level backward compatibility. - -## Backward Compatibility - -- Nodes that do not support this extension never write `providerStatus` and - are unaffected by receiving it. -- Advertisers that do not implement spillover continue to advertise to the `k` - closest peers; absent `providerStatus` is treated as acceptance. -- No change is made to `GET_PROVIDERS` or any lookup message. - -## Security Considerations - -**False rejections**: an adversary controlling the closest peers to a key could -reject all `ADD_PROVIDER` requests to suppress its advertisement. Spillover -mitigates unanimous rejection: any shortfall below `k` causes the advertiser -to route around those peers and store the record on nodes beyond the adversary's -controlled set. An adversary can limit—but not prevent—replication by mixing accepts -and rejects across chunks, since spillover continues as long as the target is unmet. - -**Slot monopolisation without an eviction policy**: because re-advertisements -from already-stored providers are always accepted regardless of the limit, the -first `maxProvidersPerKey` providers to register for a key can hold their slots -indefinitely simply by refreshing their records. Nodes that fill up later deny -new providers entry, so the stored provider set becomes permanently frozen around -whoever arrived first. Implementations MAY therefore pair `maxProvidersPerKey` -with an eviction policy — for example, evicting the record with the oldest -`timeReceived` when the limit is reached and a new (non-incumbent) provider -advertises — to ensure the stored set can rotate over time and is not captured -by early registrants. - ---- - -## References - -[kad-spec]: https://github.com/libp2p/specs/blob/master/kad-dht/README.md diff --git a/src/ipips/ipip-0537.md b/src/ipips/ipip-0537.md new file mode 100644 index 00000000..3d6ed5d0 --- /dev/null +++ b/src/ipips/ipip-0537.md @@ -0,0 +1,180 @@ +--- +title: "IPIP-0537: Provider Record Spillover" +date: 2026-09-01 +ipip: proposal +editors: + - name: Gabriel Cruz + github: gmelodie +relatedIssues: + - https://github.com/libp2p/go-libp2p-kad-dht/issues/316 +order: 537 +tags: ['ipips'] +--- + +## Summary + +Let a DHT Server limit the number of providers that it stores for a single key, +and let it tell an advertising node that it rejected a Provider Record. The +advertising node then stores the record on peers that are farther along its +lookup path, so that a popular CID stays resolvable. + +## Motivation + +A node that advertises content sends `ADD_PROVIDER` to the `k` closest DHT +Servers to the Kademlia Identifier of the CID, and to no other peer. For a +popular CID, those `k` servers receive every `ADD_PROVIDER` for that CID, and +store one Provider Record per provider, with no limit. They become a permanent hotspot, +and they carry the storage cost, the CPU cost and the bandwidth cost of that CID +for the whole network. + +:cite[kad-dht] gives a server no way out of this. It has no way to decline an +`ADD_PROVIDER`, because `ADD_PROVIDER` is fire and forget: the server echoes the +request on success, and current implementations write no response at all. A +server under load can only drop the record silently. The advertising node +learns nothing, and it has nowhere else to put the record, because the base +advertisement stops at the `k` closest servers. + +## Detailed design + +This IPIP modifies :cite[kad-dht]. It adds these sections: + +* **Protocol Versions**: version `2.0.0` of the Kademlia protocol identifier, + for example `/ipfs/kad/2.0.0`. It is identical to version `1.0.0`, except for + the `ADD_PROVIDER` response. A DHT Server that implements it advertises both + versions, and libp2p protocol negotiation tells a sender which version a peer + supports. +* **Provider Record Limits**: the optional `maxProvidersPerKey` limit, its + RECOMMENDED value of `1000`, and the rule that a server always accepts a + re-advertisement from a provider that it already stores. +* **Eviction**: the optional eviction policy, the precedence between eviction + and rejection, and the restriction of the eviction candidates to the records + that are older than the republish interval. +* **`ADD_PROVIDER` Response**: the exact response message on version `2.0.0`, + its fields, and the `ACCEPTED`, `REJECTED` and `INVALID` status values. +* **Outcome Classification**: which outcome of an `ADD_PROVIDER` attempt counts + towards the replication factor `k`. A transport failure, a timeout and a + missing status count as a failure. +* **Spillover**: the advertisement walks the sorted lookup results in chunks of + `α` peers, from the closest to the farthest, until the number of stored + records reaches `k`. +* **Deployment**: the order in which implementations roll the extension out. + +It also modifies these sections: + +* **Amino DHT**: a server that implements version `2.0.0` mounts the swarm under + `/ipfs/kad/2.0.0` as well. +* **Content Provider Advertisement**: the advertising node keeps every peer that + the lookup discovered, and not only the `k` closest ones. The response on a + version `2.0.0` stream is the `ADD_PROVIDER` response. +* **Content Provider Lookup**: a client that holds fewer providers than it wants + continues past the `k` closest servers, so that it finds a record that spilled + over. +* **RPC Messages**: the `AddProviderStatus` enum, the optional + `providerStatus` field 11 of `Message`, and the `ADD_PROVIDER` rules. + +## Design rationale + +The two halves of the design match the two halves of the problem. The limit and +the rejection protect a server against a hotspot. The spillover keeps the CID +resolvable once a server rejects a record, and it needs the rejection signal to +know when to continue. + +The spillover reuses the peers that the lookup already discovered, so it costs +no extra `FIND_NODE` round. It also degrades into the base behavior: the first +`⌈k/α⌉` chunks are the `k` closest peers, so a node that gets `k` records from +them stops exactly where the base specification stops. + +### User benefit + +A popular CID stays resolvable, and its providers stop concentrating on the same +`k` servers. Nodes that host popular content still get a full set of records, +spread over more servers. An operator of a DHT Server can cap the cost of a +hotspot without dropping records silently, which makes a server on modest +hardware viable. + +### Compatibility + +The wire format does not change. Field 11 is optional and new, so a node that +does not know it ignores it. + +The behavior of a legacy peer is the problem, not the wire format. Today +`ADD_PROVIDER` is fire and forget in the deployed network: [go-libp2p-kad-dht +writes no response](https://github.com/libp2p/go-libp2p-kad-dht/blob/10e0adf9859ef86ba08d8493d8313869a2e83d8a/handlers.go#L278), +and clients read none. A client that waits for a status from such a server waits +for its full timeout, on every request. + +Protocol version `2.0.0` removes that cost. A server that implements the +extension advertises both versions, so negotiation tells the sender which +exchange to use before it writes the request. These are the combinations: + +| Sender | DHT Server | Behavior | +|--------|------------|----------| +| legacy | legacy | Fire and forget on version `1.0.0`. The server stores the record. | +| legacy | new | The sender opens version `1.0.0`, the only version that it knows. The server accepts the record while the limits stay unset. Once an operator sets `maxProvidersPerKey`, the server can drop the record, and the sender cannot learn this. | +| new | legacy | The server advertises version `1.0.0` only, so the sender opens that version and waits for no response. It spends no timeout. | +| new | new | The peers negotiate version `2.0.0`. The rejection and the spillover work as specified. | + +One combination loses records: a legacy sender against a server that enforces a +limit. The Deployment section keeps that combination safe. Nodes that advertise +and nodes that look up ship first, and servers set `maxProvidersPerKey` only +once most of the `ADD_PROVIDER` requests that they see arrive on version +`2.0.0`. Adoption at that scale takes months. While it runs, a server that +enforces a limit applies it to version `2.0.0` requests only. + +### Security + +**An acknowledgment is not proof of storage.** With `ACCEPTED`, a server claims +that it stored the record. The protocol cannot prove the storage of any record, +and it cannot detect a server that answers `ACCEPTED` and stores nothing. Such a +server absorbs records and stops the spillover, exactly as a server that drops +records under load does. If the `k` closest servers collude, they suppress an +advertisement completely. The base protocol has the same weakness, because the +same servers can discard a record. A node that needs a stronger assurance checks +the result out of band, for example with a `GET_PROVIDERS` to each server that +it advertised to. + +**False rejections.** An adversary that controls the closest servers to a key +can reject every `ADD_PROVIDER` and suppress an advertisement. The spillover +limits a unanimous rejection: every record short of `k` moves the node outwards, +to servers outside the set that the adversary controls. An adversary that mixes +acceptances and rejections across the chunks reduces the replication. Rejections +alone cannot stop the storage of the record, because the spillover continues +while the count stays below `k`. + +**Eviction and Peer ID rotation.** A policy that always drops the oldest record +lets an attacker flush the honest providers out of the `k` closest servers, for +one request per eviction. The attacker only has to rotate its Peer ID, and the +base protocol has no such censorship vector. The Eviction section removes the +gain: the candidates are the records that are older than the republish interval, +which their providers had to refresh already. + +**Slot monopolisation.** With no eviction policy, the first +`maxProvidersPerKey` providers of a key hold their slots while they republish. +The spillover moves every later provider outwards. This is the safe default. A +later provider loses proximity to the key, and keeps reachability, because the +lookup finds a record that spilled over. + +### Alternatives + +**A limit with no signal.** A server can cap the records that it stores today, +and drop the rest silently. This is what an overloaded server does. The +advertising node keeps its count of `k`, believes that the record is stored, and +the CID becomes harder to resolve with no way to detect it. + +**A retry against the same servers.** A node that treats a rejection as a +transient error and retries reaches the same overloaded servers, and adds load +to the hotspot that the limit protects. + +**A dedicated error message type instead of a status field.** A new message type +carries the same information, and every implementation has to route it. The +optional field 11 on the existing response keeps the exchange to one +request and one response. + +**A new field with no protocol version bump.** A sender then cannot tell a +server that implements the extension from one that does not, so it waits for a +full timeout on every legacy server. Version `2.0.0` moves that detection into +libp2p protocol negotiation, which happens before the request. + +### Copyright + +Copyright and related rights waived via [CC0](https://creativecommons.org/publicdomain/zero/1.0/). diff --git a/src/routing/kad-dht.md b/src/routing/kad-dht.md index 498f104d..f8530309 100644 --- a/src/routing/kad-dht.md +++ b/src/routing/kad-dht.md @@ -5,7 +5,7 @@ description: > overlay network used for peer and content routing in the InterPlanetary File System (IPFS). It extends the libp2p Kademlia DHT specification, adapting and adding features to support IPFS-specific requirements. -date: 2025-11-20 +date: 2026-09-01 maturity: reliable editors: - name: Guillaume Michel @@ -107,6 +107,9 @@ The Amino DHT is utilized by multiple IPFS implementations, including and can be joined by using the [public good Amino DHT Bootstrappers](https://docs.ipfs.tech/concepts/public-utilities/#amino-dht-bootstrappers). ::: +Amino DHT Servers that implement [protocol version +`2.0.0`](#protocol-versions) mount the swarm under `/ipfs/kad/2.0.0` as well. + #### IPFS LAN DHTs _IPFS LAN DHTs_ are DHT swarms operating exclusively within a local network. @@ -140,6 +143,33 @@ Dedicated bootstrapper nodes MAY be used to facilitate this process. They SHOULD be publicly reachable, maintain high availability and possess sufficient resources to support the network. +### Protocol Versions + +The protocol identifier of a swarm ends with a version, for example +`/ipfs/kad/1.0.0`. This document defines two versions. + +Version `1.0.0` is the base protocol. Version `2.0.0` adds the [`ADD_PROVIDER` +response](#add_provider-response) that [Provider Record +Limits](#provider-record-limits) and [Spillover](#spillover) need. The two +versions are identical in every other respect. `PUT_VALUE`, `GET_VALUE`, +`GET_PROVIDERS`, `FIND_NODE` and `PING` keep the same message formats and the +same requirements on both versions. + +A swarm that runs both versions follows these rules: +* A DHT Server that implements version `2.0.0` MUST advertise both versions +through the libp2p identify protocol, and MUST accept an incoming stream on +both versions. +* A node that sends an `ADD_PROVIDER` MUST open the stream on version `2.0.0` +when the remote peer advertises that version, and MUST open it on version +`1.0.0` otherwise. +* On a version `1.0.0` stream, the sender MUST NOT wait for an `ADD_PROVIDER` +response, and the DHT Server MUST NOT set `providerStatus`. + +Protocol negotiation thus tells a sender which version a peer supports before +the sender writes the request, so the sender spends no timeout on a peer that +implements version `1.0.0` only. A swarm that drops version `1.0.0` at a later +date only changes the list of advertised identifiers, and needs no flag day. + ### Client and Server Mode A node operating in Server Mode (or DHT Server) is responsible for responding @@ -481,7 +511,9 @@ When a node wants to indicate that it provides the content associated with a given CID, it first finds the `k` closest DHT Servers to the Kademlia Identifier associated with the CID using [`GetClosestPeers`](#getclosestpeers). The `key` in the `FIND_NODE` payload is set to the multihash contained in the -CID. +CID. The node keeps every peer that the lookup discovered, sorted by ascending +XOR distance to the Kademlia Identifier, and not only the `k` closest ones. +[Spillover](#spillover) uses the rest of that list. Once the `k` closest DHT Servers are found, the node sends each of them an `ADD_PROVIDER` RPC, using the same `key` and setting its own Peer ID as @@ -494,9 +526,14 @@ datastore: 2. Discard `providerPeers` whose Peer ID is not matching the sender's Peer ID Upon successful verification, the DHT Server stores the Provider Record in its -datastore, and caches the provided public multiaddresses. It responds by -echoing the request to confirm success. If verification fails, the server MUST -close the stream without sending a response. +datastore, and caches the provided public multiaddresses, unless [Provider +Record Limits](#provider-record-limits) make it reject the record. + +On a version `1.0.0` stream, the DHT Server responds by echoing the request to +confirm success. If verification fails, the server MUST close the stream without +sending a response. On a version `2.0.0` stream, the server answers with the +[`ADD_PROVIDER` response](#add_provider-response) instead, both for a success +and for a failure. #### Provide Validity @@ -519,6 +556,183 @@ content provider alongside the provide record, avoiding an additional DHT walk for the Client ([rationale](https://github.com/probe-lab/network-measurements/blob/master/results/rfm17.1-sharing-prs-with-multiaddresses.md)). +#### Provider Record Limits + +For a popular CID, the `k` closest DHT Servers to its Kademlia Identifier +receive every `ADD_PROVIDER` for that CID, and they store one Provider Record +per provider, with no limit. A DHT Server MAY set a maximum number of distinct providers +per key, `maxProvidersPerKey`, to limit that load. + +The server counts the distinct provider Peer IDs that it stores for the key. If +that count reaches `maxProvidersPerKey`, and the server holds no record for the +key from the sender of the `ADD_PROVIDER`, the server MUST reject the request or +evict a stored record. [Eviction](#eviction) gives the rules. + +A server MUST always accept a re-advertisement from a provider that it already +stores for the key, whatever the limit is, so that an existing provider can +refresh its record. + +`maxProvidersPerKey` has no default value. A server that sets it MUST keep the +value at the replication factor `k` or above. With a value below `k`, the `k` +closest servers alone cannot resolve even an unpopular key. The RECOMMENDED +value is `1000`. A client stops after a few dozen usable providers, so a limit +three orders of magnitude above `k` caps the storage cost and the CPU cost of a +hotspot, and keeps a popular key fully resolvable. This value is provisional. +Implementations SHOULD measure the live network and correct it, as they do for +the [republish interval](#provider-record-republish-interval) and the [provide +validity](#provide-validity). + +DHT Servers SHOULD also enforce coarser limits, such as the total number of +Provider Records stored, and the total number of keys that they hold records +for. + +A DHT Server MAY reject an `ADD_PROVIDER` for another reason than a limit, for +example a local policy. The rules below apply to every rejection. + +#### Eviction + +A DHT Server always accepts a re-advertisement, so the first +`maxProvidersPerKey` providers of a key can hold their slots forever, and they +only have to refresh their records. The stored set then freezes around the +providers that arrived first. A DHT Server MAY add an eviction policy to +`maxProvidersPerKey` to let the set rotate. + +Eviction and rejection are exclusive, and eviction takes precedence: +* If the server evicts a record, it MUST store the new record, and it MUST +answer `ACCEPTED`. +* If the server evicts no record, it MUST keep every stored record, and it MUST +answer `REJECTED`. + +A server MUST NOT evict a record and answer `REJECTED`. That combination drops a +provider and stores no replacement. + +A policy that evicts the record with the oldest `timeReceived` is unsafe on its +own. An attacker rotates its Peer ID, looks like a new provider on every +request, and flushes the honest providers out of the `k` closest servers at a +cost of one request per eviction. The base protocol has no such censorship +vector, because it removes no record before its expiration. + +A DHT Server that evicts MUST select the eviction candidates only among the +records whose `timeReceived` is older than the [republish +interval](#provider-record-republish-interval). Such a record passed the moment +at which its provider had to refresh it. A provider that republishes on schedule +thus keeps its slot, and an attacker that rotates its Peer ID gains nothing over +an attacker that waits. + +#### `ADD_PROVIDER` Response + +On a version `2.0.0` stream, a DHT Server that receives an `ADD_PROVIDER` MUST +write one response message, and MUST then close its side of the stream. The +response is a new message, and not a copy of the request. It carries these +fields: + +| Field | Presence | Value | +|-------|----------|-------| +| `type` | MUST | `ADD_PROVIDER` | +| `key` | MUST | the `key` of the request | +| `providerStatus` | MUST | see below | + +The server MUST leave every other field empty, and returns no `providerPeers` +and no `closerPeers`. If another field is present, the sender MUST ignore it. + +`providerStatus` is a response-only field. A sender MUST NOT set it in an +`ADD_PROVIDER` request, and a DHT Server MUST ignore it when an incoming request +carries it. + +The status values are: +* `ACCEPTED`: the server stored the Provider Record. +* `REJECTED`: the server did not store the Provider Record, because of a limit +or of a local policy. The request is well formed, so the same request MAY +succeed at another server. +* `INVALID`: the request is malformed, and no other server accepts it either. A +server MUST answer `INVALID` when a check of [Content Provider +Advertisement](#content-provider-advertisement) fails, which means that `key` is +absent, that `key` exceeds `80` bytes, or that no entry of `providerPeers` +matches the Peer ID of the sender. + +#### Outcome Classification + +A node that advertises classifies each `ADD_PROVIDER` attempt as exactly one +outcome. Only the first two outcomes count towards the replication factor `k`. + +| Outcome | Counts towards `k` | +|---------|--------------------| +| version `2.0.0`, response with `providerStatus = ACCEPTED` | yes | +| version `1.0.0`, request written successfully | yes | +| version `2.0.0`, response with `providerStatus = REJECTED` | no | +| version `2.0.0`, response with `providerStatus = INVALID` | no, see below | +| version `2.0.0`, response without `providerStatus` | no | +| dial failure, stream reset, or write failure | no | +| stream closed before a response arrived | no | +| response timeout | no | + +A version `1.0.0` request counts as soon as the write succeeds, because the base +protocol gives the sender no other information. It is the only outcome that +counts without a response. + +A response without `providerStatus` on a version `2.0.0` stream violates this +specification. The sender MUST count that attempt as a failure, and MUST NOT +assume that the server stored the record. + +A transport failure MUST NOT count towards `k`. Silence tells the sender nothing +about storage. If silence counted, a peer that drops streams would absorb +placements and hold no record. + +If a server answers `INVALID`, the sender SHOULD stop the advertisement for that +key, and SHOULD report the error to the caller. Every other server applies the +same checks and answers `INVALID` too, so another attempt gains nothing. + +#### Spillover + +When the `k` closest DHT Servers do not all store the Provider Record, the +advertising node continues with the peers that its lookup found farther from the +Kademlia Identifier. Each extra batch of `ADD_PROVIDER` requests is a spillover +round. + +The node splits the sorted candidate list of the [Content Provider +Advertisement](#content-provider-advertisement) into chunks of `α` peers, and +walks the chunks from the closest to the farthest: +1. Send `ADD_PROVIDER` to every peer of the current chunk at the same time, on +the version that each peer advertises. +2. Classify each attempt with [Outcome +Classification](#outcome-classification), and add the successful ones to the +count. +3. Stop when the count reaches `k`. +4. Continue with the next chunk while the count stays below `k`. Each of these +chunks is a spillover round. +5. Stop when no chunk remains. The node stored fewer than `k` records, and it +SHOULD report how many it stored. + +The first `⌈k/α⌉` chunks hold the `k` closest peers, which is the set that a node +advertises to without this extension. With `k` = 20 and `α` = 10, these are the +first two chunks. If those peers store every record, the node stops there, and +no spillover round happens. + +A node SHOULD use a larger request timeout in a spillover round than in the +first chunks, because it is less likely to already hold a connection to a peer +that is farther from the key. + +#### Deployment + +If DHT Servers enforce `maxProvidersPerKey` before the advertising nodes can +read a rejection, those nodes lose Provider Records with no signal and no +fallback. Implementations SHOULD deploy this extension in this order. +1. **Advertising nodes first.** Add version `2.0.0`: the response reader, the +outcome classification and the spillover. Leave `maxProvidersPerKey` unset. Only +the negotiated version changes on the wire, and every server still stores every +record. +2. **Lookups next.** Add the [Content Provider Lookup](#content-provider-lookup) +change, so that a client finds a record that spilled over before the first +record spills over. +3. **Servers last.** Set `maxProvidersPerKey` only when most of the incoming +`ADD_PROVIDER` requests that a server sees arrive on version `2.0.0`. Adoption +on this scale takes months. An implementation SHOULD measure that share, and +decide from it. +4. **Version `1.0.0` requests.** While many nodes still advertise on version +`1.0.0`, a server that enforces `maxProvidersPerKey` SHOULD apply the limit to +version `2.0.0` requests only, because it cannot tell a version `1.0.0` sender +that it dropped the record. + ### Content Provider Lookup To find providers for a given CID, a node initiates a lookup using the @@ -531,6 +745,21 @@ providers. If a node does not find any provider records and is unable to discover closer DHT servers after querying the `β` closest reachable servers, the request is considered a failure. +A Provider Record that [spilled over](#spillover) sits on a DHT Server outside +the `k` closest servers to the Kademlia Identifier, so a client that queries +only the `k` closest servers never finds it. `GET_PROVIDERS` is unchanged, and +only the point at which a client stops changes. A client that runs an iterative +lookup already moves outwards from the closest servers. While it holds fewer +providers than it wants, it SHOULD continue past the `k` closest servers, and +query the next chunk of `α` candidates in ascending distance order, as +[Spillover](#spillover) does. It stops when it holds enough providers, or when +no candidate remains. + +Some clients skip the iterative lookup, and take the `k` closest servers +directly from a full routing table. Such a client SHOULD extend its query set in +the same way, because the `k` closest servers return the providers that arrived +first, and hide every provider that spilled over. + ## Value Storage and Retrieval The IPFS Kademlia DHT allows users to store and retrieve records directly @@ -690,6 +919,18 @@ message Message { CANNOT_CONNECT = 3; } + enum AddProviderStatus { + // the DHT Server stored the provider record + ACCEPTED = 0; + + // the DHT Server did not store the provider record, because of a limit + // or of a local policy + REJECTED = 1; + + // the request is malformed, and every other DHT Server rejects it too + INVALID = 2; + } + message Peer { // ID of a given peer. bytes id = 1; @@ -723,6 +964,12 @@ message Message { // Used to return Providers // GET_VALUE, ADD_PROVIDER, GET_PROVIDERS repeated Peer providerPeers = 9; + + // Used to report whether the provider record was stored. + // ADD_PROVIDER responses on protocol version 2.0.0 only. + // The field is optional because a sender distinguishes an absent status + // from ACCEPTED. + optional AddProviderStatus providerStatus = 11; } ``` @@ -747,13 +994,17 @@ the `k` closest known `closerPeers`. * `ADD_PROVIDER`: In the request `key` is set to the multihash contained in the target CID. The target node verifies `key` is a valid multihash, all -`providerPeers` matching the RPC sender's PeerID are recorded as providers. +`providerPeers` matching the RPC sender's PeerID are recorded as providers. On +protocol version `2.0.0`, the target node reports the outcome in +`providerStatus`, see [`ADD_PROVIDER` response](#add_provider-response). * `PING`: Deprecated message type replaced by the dedicated [ping protocol](https://github.com/libp2p/specs/blob/master/ping/ping.md). If a DHT server receives an invalid request, it simply closes the libp2p stream -without responding. +without responding. An `ADD_PROVIDER` on protocol version `2.0.0` is the +exception: the server answers `INVALID` instead of closing the stream, see +[`ADD_PROVIDER` response](#add_provider-response). # Appendix: Notes for Implementers