diff --git a/src/ipips/ipip-0537.md b/src/ipips/ipip-0537.md new file mode 100644 index 00000000..3d6ed5d0 --- /dev/null +++ b/src/ipips/ipip-0537.md @@ -0,0 +1,180 @@ +--- +title: "IPIP-0537: Provider Record Spillover" +date: 2026-09-01 +ipip: proposal +editors: + - name: Gabriel Cruz + github: gmelodie +relatedIssues: + - https://github.com/libp2p/go-libp2p-kad-dht/issues/316 +order: 537 +tags: ['ipips'] +--- + +## Summary + +Let a DHT Server limit the number of providers that it stores for a single key, +and let it tell an advertising node that it rejected a Provider Record. The +advertising node then stores the record on peers that are farther along its +lookup path, so that a popular CID stays resolvable. + +## Motivation + +A node that advertises content sends `ADD_PROVIDER` to the `k` closest DHT +Servers to the Kademlia Identifier of the CID, and to no other peer. For a +popular CID, those `k` servers receive every `ADD_PROVIDER` for that CID, and +store one Provider Record per provider, with no limit. They become a permanent hotspot, +and they carry the storage cost, the CPU cost and the bandwidth cost of that CID +for the whole network. + +:cite[kad-dht] gives a server no way out of this. It has no way to decline an +`ADD_PROVIDER`, because `ADD_PROVIDER` is fire and forget: the server echoes the +request on success, and current implementations write no response at all. A +server under load can only drop the record silently. The advertising node +learns nothing, and it has nowhere else to put the record, because the base +advertisement stops at the `k` closest servers. + +## Detailed design + +This IPIP modifies :cite[kad-dht]. It adds these sections: + +* **Protocol Versions**: version `2.0.0` of the Kademlia protocol identifier, + for example `/ipfs/kad/2.0.0`. It is identical to version `1.0.0`, except for + the `ADD_PROVIDER` response. A DHT Server that implements it advertises both + versions, and libp2p protocol negotiation tells a sender which version a peer + supports. +* **Provider Record Limits**: the optional `maxProvidersPerKey` limit, its + RECOMMENDED value of `1000`, and the rule that a server always accepts a + re-advertisement from a provider that it already stores. +* **Eviction**: the optional eviction policy, the precedence between eviction + and rejection, and the restriction of the eviction candidates to the records + that are older than the republish interval. +* **`ADD_PROVIDER` Response**: the exact response message on version `2.0.0`, + its fields, and the `ACCEPTED`, `REJECTED` and `INVALID` status values. +* **Outcome Classification**: which outcome of an `ADD_PROVIDER` attempt counts + towards the replication factor `k`. A transport failure, a timeout and a + missing status count as a failure. +* **Spillover**: the advertisement walks the sorted lookup results in chunks of + `α` peers, from the closest to the farthest, until the number of stored + records reaches `k`. +* **Deployment**: the order in which implementations roll the extension out. + +It also modifies these sections: + +* **Amino DHT**: a server that implements version `2.0.0` mounts the swarm under + `/ipfs/kad/2.0.0` as well. +* **Content Provider Advertisement**: the advertising node keeps every peer that + the lookup discovered, and not only the `k` closest ones. The response on a + version `2.0.0` stream is the `ADD_PROVIDER` response. +* **Content Provider Lookup**: a client that holds fewer providers than it wants + continues past the `k` closest servers, so that it finds a record that spilled + over. +* **RPC Messages**: the `AddProviderStatus` enum, the optional + `providerStatus` field 11 of `Message`, and the `ADD_PROVIDER` rules. + +## Design rationale + +The two halves of the design match the two halves of the problem. The limit and +the rejection protect a server against a hotspot. The spillover keeps the CID +resolvable once a server rejects a record, and it needs the rejection signal to +know when to continue. + +The spillover reuses the peers that the lookup already discovered, so it costs +no extra `FIND_NODE` round. It also degrades into the base behavior: the first +`⌈k/α⌉` chunks are the `k` closest peers, so a node that gets `k` records from +them stops exactly where the base specification stops. + +### User benefit + +A popular CID stays resolvable, and its providers stop concentrating on the same +`k` servers. Nodes that host popular content still get a full set of records, +spread over more servers. An operator of a DHT Server can cap the cost of a +hotspot without dropping records silently, which makes a server on modest +hardware viable. + +### Compatibility + +The wire format does not change. Field 11 is optional and new, so a node that +does not know it ignores it. + +The behavior of a legacy peer is the problem, not the wire format. Today +`ADD_PROVIDER` is fire and forget in the deployed network: [go-libp2p-kad-dht +writes no response](https://github.com/libp2p/go-libp2p-kad-dht/blob/10e0adf9859ef86ba08d8493d8313869a2e83d8a/handlers.go#L278), +and clients read none. A client that waits for a status from such a server waits +for its full timeout, on every request. + +Protocol version `2.0.0` removes that cost. A server that implements the +extension advertises both versions, so negotiation tells the sender which +exchange to use before it writes the request. These are the combinations: + +| Sender | DHT Server | Behavior | +|--------|------------|----------| +| legacy | legacy | Fire and forget on version `1.0.0`. The server stores the record. | +| legacy | new | The sender opens version `1.0.0`, the only version that it knows. The server accepts the record while the limits stay unset. Once an operator sets `maxProvidersPerKey`, the server can drop the record, and the sender cannot learn this. | +| new | legacy | The server advertises version `1.0.0` only, so the sender opens that version and waits for no response. It spends no timeout. | +| new | new | The peers negotiate version `2.0.0`. The rejection and the spillover work as specified. | + +One combination loses records: a legacy sender against a server that enforces a +limit. The Deployment section keeps that combination safe. Nodes that advertise +and nodes that look up ship first, and servers set `maxProvidersPerKey` only +once most of the `ADD_PROVIDER` requests that they see arrive on version +`2.0.0`. Adoption at that scale takes months. While it runs, a server that +enforces a limit applies it to version `2.0.0` requests only. + +### Security + +**An acknowledgment is not proof of storage.** With `ACCEPTED`, a server claims +that it stored the record. The protocol cannot prove the storage of any record, +and it cannot detect a server that answers `ACCEPTED` and stores nothing. Such a +server absorbs records and stops the spillover, exactly as a server that drops +records under load does. If the `k` closest servers collude, they suppress an +advertisement completely. The base protocol has the same weakness, because the +same servers can discard a record. A node that needs a stronger assurance checks +the result out of band, for example with a `GET_PROVIDERS` to each server that +it advertised to. + +**False rejections.** An adversary that controls the closest servers to a key +can reject every `ADD_PROVIDER` and suppress an advertisement. The spillover +limits a unanimous rejection: every record short of `k` moves the node outwards, +to servers outside the set that the adversary controls. An adversary that mixes +acceptances and rejections across the chunks reduces the replication. Rejections +alone cannot stop the storage of the record, because the spillover continues +while the count stays below `k`. + +**Eviction and Peer ID rotation.** A policy that always drops the oldest record +lets an attacker flush the honest providers out of the `k` closest servers, for +one request per eviction. The attacker only has to rotate its Peer ID, and the +base protocol has no such censorship vector. The Eviction section removes the +gain: the candidates are the records that are older than the republish interval, +which their providers had to refresh already. + +**Slot monopolisation.** With no eviction policy, the first +`maxProvidersPerKey` providers of a key hold their slots while they republish. +The spillover moves every later provider outwards. This is the safe default. A +later provider loses proximity to the key, and keeps reachability, because the +lookup finds a record that spilled over. + +### Alternatives + +**A limit with no signal.** A server can cap the records that it stores today, +and drop the rest silently. This is what an overloaded server does. The +advertising node keeps its count of `k`, believes that the record is stored, and +the CID becomes harder to resolve with no way to detect it. + +**A retry against the same servers.** A node that treats a rejection as a +transient error and retries reaches the same overloaded servers, and adds load +to the hotspot that the limit protects. + +**A dedicated error message type instead of a status field.** A new message type +carries the same information, and every implementation has to route it. The +optional field 11 on the existing response keeps the exchange to one +request and one response. + +**A new field with no protocol version bump.** A sender then cannot tell a +server that implements the extension from one that does not, so it waits for a +full timeout on every legacy server. Version `2.0.0` moves that detection into +libp2p protocol negotiation, which happens before the request. + +### Copyright + +Copyright and related rights waived via [CC0](https://creativecommons.org/publicdomain/zero/1.0/). diff --git a/src/routing/kad-dht.md b/src/routing/kad-dht.md index 498f104d..f8530309 100644 --- a/src/routing/kad-dht.md +++ b/src/routing/kad-dht.md @@ -5,7 +5,7 @@ description: > overlay network used for peer and content routing in the InterPlanetary File System (IPFS). It extends the libp2p Kademlia DHT specification, adapting and adding features to support IPFS-specific requirements. -date: 2025-11-20 +date: 2026-09-01 maturity: reliable editors: - name: Guillaume Michel @@ -107,6 +107,9 @@ The Amino DHT is utilized by multiple IPFS implementations, including and can be joined by using the [public good Amino DHT Bootstrappers](https://docs.ipfs.tech/concepts/public-utilities/#amino-dht-bootstrappers). ::: +Amino DHT Servers that implement [protocol version +`2.0.0`](#protocol-versions) mount the swarm under `/ipfs/kad/2.0.0` as well. + #### IPFS LAN DHTs _IPFS LAN DHTs_ are DHT swarms operating exclusively within a local network. @@ -140,6 +143,33 @@ Dedicated bootstrapper nodes MAY be used to facilitate this process. They SHOULD be publicly reachable, maintain high availability and possess sufficient resources to support the network. +### Protocol Versions + +The protocol identifier of a swarm ends with a version, for example +`/ipfs/kad/1.0.0`. This document defines two versions. + +Version `1.0.0` is the base protocol. Version `2.0.0` adds the [`ADD_PROVIDER` +response](#add_provider-response) that [Provider Record +Limits](#provider-record-limits) and [Spillover](#spillover) need. The two +versions are identical in every other respect. `PUT_VALUE`, `GET_VALUE`, +`GET_PROVIDERS`, `FIND_NODE` and `PING` keep the same message formats and the +same requirements on both versions. + +A swarm that runs both versions follows these rules: +* A DHT Server that implements version `2.0.0` MUST advertise both versions +through the libp2p identify protocol, and MUST accept an incoming stream on +both versions. +* A node that sends an `ADD_PROVIDER` MUST open the stream on version `2.0.0` +when the remote peer advertises that version, and MUST open it on version +`1.0.0` otherwise. +* On a version `1.0.0` stream, the sender MUST NOT wait for an `ADD_PROVIDER` +response, and the DHT Server MUST NOT set `providerStatus`. + +Protocol negotiation thus tells a sender which version a peer supports before +the sender writes the request, so the sender spends no timeout on a peer that +implements version `1.0.0` only. A swarm that drops version `1.0.0` at a later +date only changes the list of advertised identifiers, and needs no flag day. + ### Client and Server Mode A node operating in Server Mode (or DHT Server) is responsible for responding @@ -481,7 +511,9 @@ When a node wants to indicate that it provides the content associated with a given CID, it first finds the `k` closest DHT Servers to the Kademlia Identifier associated with the CID using [`GetClosestPeers`](#getclosestpeers). The `key` in the `FIND_NODE` payload is set to the multihash contained in the -CID. +CID. The node keeps every peer that the lookup discovered, sorted by ascending +XOR distance to the Kademlia Identifier, and not only the `k` closest ones. +[Spillover](#spillover) uses the rest of that list. Once the `k` closest DHT Servers are found, the node sends each of them an `ADD_PROVIDER` RPC, using the same `key` and setting its own Peer ID as @@ -494,9 +526,14 @@ datastore: 2. Discard `providerPeers` whose Peer ID is not matching the sender's Peer ID Upon successful verification, the DHT Server stores the Provider Record in its -datastore, and caches the provided public multiaddresses. It responds by -echoing the request to confirm success. If verification fails, the server MUST -close the stream without sending a response. +datastore, and caches the provided public multiaddresses, unless [Provider +Record Limits](#provider-record-limits) make it reject the record. + +On a version `1.0.0` stream, the DHT Server responds by echoing the request to +confirm success. If verification fails, the server MUST close the stream without +sending a response. On a version `2.0.0` stream, the server answers with the +[`ADD_PROVIDER` response](#add_provider-response) instead, both for a success +and for a failure. #### Provide Validity @@ -519,6 +556,183 @@ content provider alongside the provide record, avoiding an additional DHT walk for the Client ([rationale](https://github.com/probe-lab/network-measurements/blob/master/results/rfm17.1-sharing-prs-with-multiaddresses.md)). +#### Provider Record Limits + +For a popular CID, the `k` closest DHT Servers to its Kademlia Identifier +receive every `ADD_PROVIDER` for that CID, and they store one Provider Record +per provider, with no limit. A DHT Server MAY set a maximum number of distinct providers +per key, `maxProvidersPerKey`, to limit that load. + +The server counts the distinct provider Peer IDs that it stores for the key. If +that count reaches `maxProvidersPerKey`, and the server holds no record for the +key from the sender of the `ADD_PROVIDER`, the server MUST reject the request or +evict a stored record. [Eviction](#eviction) gives the rules. + +A server MUST always accept a re-advertisement from a provider that it already +stores for the key, whatever the limit is, so that an existing provider can +refresh its record. + +`maxProvidersPerKey` has no default value. A server that sets it MUST keep the +value at the replication factor `k` or above. With a value below `k`, the `k` +closest servers alone cannot resolve even an unpopular key. The RECOMMENDED +value is `1000`. A client stops after a few dozen usable providers, so a limit +three orders of magnitude above `k` caps the storage cost and the CPU cost of a +hotspot, and keeps a popular key fully resolvable. This value is provisional. +Implementations SHOULD measure the live network and correct it, as they do for +the [republish interval](#provider-record-republish-interval) and the [provide +validity](#provide-validity). + +DHT Servers SHOULD also enforce coarser limits, such as the total number of +Provider Records stored, and the total number of keys that they hold records +for. + +A DHT Server MAY reject an `ADD_PROVIDER` for another reason than a limit, for +example a local policy. The rules below apply to every rejection. + +#### Eviction + +A DHT Server always accepts a re-advertisement, so the first +`maxProvidersPerKey` providers of a key can hold their slots forever, and they +only have to refresh their records. The stored set then freezes around the +providers that arrived first. A DHT Server MAY add an eviction policy to +`maxProvidersPerKey` to let the set rotate. + +Eviction and rejection are exclusive, and eviction takes precedence: +* If the server evicts a record, it MUST store the new record, and it MUST +answer `ACCEPTED`. +* If the server evicts no record, it MUST keep every stored record, and it MUST +answer `REJECTED`. + +A server MUST NOT evict a record and answer `REJECTED`. That combination drops a +provider and stores no replacement. + +A policy that evicts the record with the oldest `timeReceived` is unsafe on its +own. An attacker rotates its Peer ID, looks like a new provider on every +request, and flushes the honest providers out of the `k` closest servers at a +cost of one request per eviction. The base protocol has no such censorship +vector, because it removes no record before its expiration. + +A DHT Server that evicts MUST select the eviction candidates only among the +records whose `timeReceived` is older than the [republish +interval](#provider-record-republish-interval). Such a record passed the moment +at which its provider had to refresh it. A provider that republishes on schedule +thus keeps its slot, and an attacker that rotates its Peer ID gains nothing over +an attacker that waits. + +#### `ADD_PROVIDER` Response + +On a version `2.0.0` stream, a DHT Server that receives an `ADD_PROVIDER` MUST +write one response message, and MUST then close its side of the stream. The +response is a new message, and not a copy of the request. It carries these +fields: + +| Field | Presence | Value | +|-------|----------|-------| +| `type` | MUST | `ADD_PROVIDER` | +| `key` | MUST | the `key` of the request | +| `providerStatus` | MUST | see below | + +The server MUST leave every other field empty, and returns no `providerPeers` +and no `closerPeers`. If another field is present, the sender MUST ignore it. + +`providerStatus` is a response-only field. A sender MUST NOT set it in an +`ADD_PROVIDER` request, and a DHT Server MUST ignore it when an incoming request +carries it. + +The status values are: +* `ACCEPTED`: the server stored the Provider Record. +* `REJECTED`: the server did not store the Provider Record, because of a limit +or of a local policy. The request is well formed, so the same request MAY +succeed at another server. +* `INVALID`: the request is malformed, and no other server accepts it either. A +server MUST answer `INVALID` when a check of [Content Provider +Advertisement](#content-provider-advertisement) fails, which means that `key` is +absent, that `key` exceeds `80` bytes, or that no entry of `providerPeers` +matches the Peer ID of the sender. + +#### Outcome Classification + +A node that advertises classifies each `ADD_PROVIDER` attempt as exactly one +outcome. Only the first two outcomes count towards the replication factor `k`. + +| Outcome | Counts towards `k` | +|---------|--------------------| +| version `2.0.0`, response with `providerStatus = ACCEPTED` | yes | +| version `1.0.0`, request written successfully | yes | +| version `2.0.0`, response with `providerStatus = REJECTED` | no | +| version `2.0.0`, response with `providerStatus = INVALID` | no, see below | +| version `2.0.0`, response without `providerStatus` | no | +| dial failure, stream reset, or write failure | no | +| stream closed before a response arrived | no | +| response timeout | no | + +A version `1.0.0` request counts as soon as the write succeeds, because the base +protocol gives the sender no other information. It is the only outcome that +counts without a response. + +A response without `providerStatus` on a version `2.0.0` stream violates this +specification. The sender MUST count that attempt as a failure, and MUST NOT +assume that the server stored the record. + +A transport failure MUST NOT count towards `k`. Silence tells the sender nothing +about storage. If silence counted, a peer that drops streams would absorb +placements and hold no record. + +If a server answers `INVALID`, the sender SHOULD stop the advertisement for that +key, and SHOULD report the error to the caller. Every other server applies the +same checks and answers `INVALID` too, so another attempt gains nothing. + +#### Spillover + +When the `k` closest DHT Servers do not all store the Provider Record, the +advertising node continues with the peers that its lookup found farther from the +Kademlia Identifier. Each extra batch of `ADD_PROVIDER` requests is a spillover +round. + +The node splits the sorted candidate list of the [Content Provider +Advertisement](#content-provider-advertisement) into chunks of `α` peers, and +walks the chunks from the closest to the farthest: +1. Send `ADD_PROVIDER` to every peer of the current chunk at the same time, on +the version that each peer advertises. +2. Classify each attempt with [Outcome +Classification](#outcome-classification), and add the successful ones to the +count. +3. Stop when the count reaches `k`. +4. Continue with the next chunk while the count stays below `k`. Each of these +chunks is a spillover round. +5. Stop when no chunk remains. The node stored fewer than `k` records, and it +SHOULD report how many it stored. + +The first `⌈k/α⌉` chunks hold the `k` closest peers, which is the set that a node +advertises to without this extension. With `k` = 20 and `α` = 10, these are the +first two chunks. If those peers store every record, the node stops there, and +no spillover round happens. + +A node SHOULD use a larger request timeout in a spillover round than in the +first chunks, because it is less likely to already hold a connection to a peer +that is farther from the key. + +#### Deployment + +If DHT Servers enforce `maxProvidersPerKey` before the advertising nodes can +read a rejection, those nodes lose Provider Records with no signal and no +fallback. Implementations SHOULD deploy this extension in this order. +1. **Advertising nodes first.** Add version `2.0.0`: the response reader, the +outcome classification and the spillover. Leave `maxProvidersPerKey` unset. Only +the negotiated version changes on the wire, and every server still stores every +record. +2. **Lookups next.** Add the [Content Provider Lookup](#content-provider-lookup) +change, so that a client finds a record that spilled over before the first +record spills over. +3. **Servers last.** Set `maxProvidersPerKey` only when most of the incoming +`ADD_PROVIDER` requests that a server sees arrive on version `2.0.0`. Adoption +on this scale takes months. An implementation SHOULD measure that share, and +decide from it. +4. **Version `1.0.0` requests.** While many nodes still advertise on version +`1.0.0`, a server that enforces `maxProvidersPerKey` SHOULD apply the limit to +version `2.0.0` requests only, because it cannot tell a version `1.0.0` sender +that it dropped the record. + ### Content Provider Lookup To find providers for a given CID, a node initiates a lookup using the @@ -531,6 +745,21 @@ providers. If a node does not find any provider records and is unable to discover closer DHT servers after querying the `β` closest reachable servers, the request is considered a failure. +A Provider Record that [spilled over](#spillover) sits on a DHT Server outside +the `k` closest servers to the Kademlia Identifier, so a client that queries +only the `k` closest servers never finds it. `GET_PROVIDERS` is unchanged, and +only the point at which a client stops changes. A client that runs an iterative +lookup already moves outwards from the closest servers. While it holds fewer +providers than it wants, it SHOULD continue past the `k` closest servers, and +query the next chunk of `α` candidates in ascending distance order, as +[Spillover](#spillover) does. It stops when it holds enough providers, or when +no candidate remains. + +Some clients skip the iterative lookup, and take the `k` closest servers +directly from a full routing table. Such a client SHOULD extend its query set in +the same way, because the `k` closest servers return the providers that arrived +first, and hide every provider that spilled over. + ## Value Storage and Retrieval The IPFS Kademlia DHT allows users to store and retrieve records directly @@ -690,6 +919,18 @@ message Message { CANNOT_CONNECT = 3; } + enum AddProviderStatus { + // the DHT Server stored the provider record + ACCEPTED = 0; + + // the DHT Server did not store the provider record, because of a limit + // or of a local policy + REJECTED = 1; + + // the request is malformed, and every other DHT Server rejects it too + INVALID = 2; + } + message Peer { // ID of a given peer. bytes id = 1; @@ -723,6 +964,12 @@ message Message { // Used to return Providers // GET_VALUE, ADD_PROVIDER, GET_PROVIDERS repeated Peer providerPeers = 9; + + // Used to report whether the provider record was stored. + // ADD_PROVIDER responses on protocol version 2.0.0 only. + // The field is optional because a sender distinguishes an absent status + // from ACCEPTED. + optional AddProviderStatus providerStatus = 11; } ``` @@ -747,13 +994,17 @@ the `k` closest known `closerPeers`. * `ADD_PROVIDER`: In the request `key` is set to the multihash contained in the target CID. The target node verifies `key` is a valid multihash, all -`providerPeers` matching the RPC sender's PeerID are recorded as providers. +`providerPeers` matching the RPC sender's PeerID are recorded as providers. On +protocol version `2.0.0`, the target node reports the outcome in +`providerStatus`, see [`ADD_PROVIDER` response](#add_provider-response). * `PING`: Deprecated message type replaced by the dedicated [ping protocol](https://github.com/libp2p/specs/blob/master/ping/ping.md). If a DHT server receives an invalid request, it simply closes the libp2p stream -without responding. +without responding. An `ADD_PROVIDER` on protocol version `2.0.0` is the +exception: the server answers `INVALID` instead of closing the stream, see +[`ADD_PROVIDER` response](#add_provider-response). # Appendix: Notes for Implementers