Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
180 changes: 180 additions & 0 deletions src/ipips/ipip-0537.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,180 @@
---
title: "IPIP-0537: Provider Record Spillover"
date: 2026-09-01
ipip: proposal
editors:
- name: Gabriel Cruz
github: gmelodie
relatedIssues:
- https://github.com/libp2p/go-libp2p-kad-dht/issues/316
order: 537
tags: ['ipips']
---

## Summary

Let a DHT Server limit the number of providers that it stores for a single key,
and let it tell an advertising node that it rejected a Provider Record. The
advertising node then stores the record on peers that are farther along its
lookup path, so that a popular CID stays resolvable.

## Motivation

A node that advertises content sends `ADD_PROVIDER` to the `k` closest DHT
Servers to the Kademlia Identifier of the CID, and to no other peer. For a
popular CID, those `k` servers receive every `ADD_PROVIDER` for that CID, and
store one Provider Record per provider, with no limit. They become a permanent hotspot,
and they carry the storage cost, the CPU cost and the bandwidth cost of that CID
for the whole network.

:cite[kad-dht] gives a server no way out of this. It has no way to decline an
`ADD_PROVIDER`, because `ADD_PROVIDER` is fire and forget: the server echoes the
request on success, and current implementations write no response at all. A
server under load can only drop the record silently. The advertising node
learns nothing, and it has nowhere else to put the record, because the base
advertisement stops at the `k` closest servers.

## Detailed design

This IPIP modifies :cite[kad-dht]. It adds these sections:

* **Protocol Versions**: version `2.0.0` of the Kademlia protocol identifier,
for example `/ipfs/kad/2.0.0`. It is identical to version `1.0.0`, except for
the `ADD_PROVIDER` response. A DHT Server that implements it advertises both
versions, and libp2p protocol negotiation tells a sender which version a peer
supports.
* **Provider Record Limits**: the optional `maxProvidersPerKey` limit, its
RECOMMENDED value of `1000`, and the rule that a server always accepts a
re-advertisement from a provider that it already stores.
* **Eviction**: the optional eviction policy, the precedence between eviction
and rejection, and the restriction of the eviction candidates to the records
that are older than the republish interval.
* **`ADD_PROVIDER` Response**: the exact response message on version `2.0.0`,
its fields, and the `ACCEPTED`, `REJECTED` and `INVALID` status values.
* **Outcome Classification**: which outcome of an `ADD_PROVIDER` attempt counts
towards the replication factor `k`. A transport failure, a timeout and a
missing status count as a failure.
* **Spillover**: the advertisement walks the sorted lookup results in chunks of
`α` peers, from the closest to the farthest, until the number of stored
records reaches `k`.
* **Deployment**: the order in which implementations roll the extension out.

It also modifies these sections:

* **Amino DHT**: a server that implements version `2.0.0` mounts the swarm under
`/ipfs/kad/2.0.0` as well.
* **Content Provider Advertisement**: the advertising node keeps every peer that
the lookup discovered, and not only the `k` closest ones. The response on a
version `2.0.0` stream is the `ADD_PROVIDER` response.
* **Content Provider Lookup**: a client that holds fewer providers than it wants
continues past the `k` closest servers, so that it finds a record that spilled
over.
* **RPC Messages**: the `AddProviderStatus` enum, the optional
`providerStatus` field 11 of `Message`, and the `ADD_PROVIDER` rules.

## Design rationale

The two halves of the design match the two halves of the problem. The limit and
the rejection protect a server against a hotspot. The spillover keeps the CID
resolvable once a server rejects a record, and it needs the rejection signal to
know when to continue.

The spillover reuses the peers that the lookup already discovered, so it costs
no extra `FIND_NODE` round. It also degrades into the base behavior: the first
`⌈k/α⌉` chunks are the `k` closest peers, so a node that gets `k` records from
them stops exactly where the base specification stops.

### User benefit

A popular CID stays resolvable, and its providers stop concentrating on the same
`k` servers. Nodes that host popular content still get a full set of records,
spread over more servers. An operator of a DHT Server can cap the cost of a
hotspot without dropping records silently, which makes a server on modest
hardware viable.

### Compatibility

The wire format does not change. Field 11 is optional and new, so a node that
does not know it ignores it.

The behavior of a legacy peer is the problem, not the wire format. Today
`ADD_PROVIDER` is fire and forget in the deployed network: [go-libp2p-kad-dht
writes no response](https://github.com/libp2p/go-libp2p-kad-dht/blob/10e0adf9859ef86ba08d8493d8313869a2e83d8a/handlers.go#L278),
and clients read none. A client that waits for a status from such a server waits
for its full timeout, on every request.

Protocol version `2.0.0` removes that cost. A server that implements the
extension advertises both versions, so negotiation tells the sender which
exchange to use before it writes the request. These are the combinations:

| Sender | DHT Server | Behavior |
|--------|------------|----------|
| legacy | legacy | Fire and forget on version `1.0.0`. The server stores the record. |
| legacy | new | The sender opens version `1.0.0`, the only version that it knows. The server accepts the record while the limits stay unset. Once an operator sets `maxProvidersPerKey`, the server can drop the record, and the sender cannot learn this. |
| new | legacy | The server advertises version `1.0.0` only, so the sender opens that version and waits for no response. It spends no timeout. |
| new | new | The peers negotiate version `2.0.0`. The rejection and the spillover work as specified. |

One combination loses records: a legacy sender against a server that enforces a
limit. The Deployment section keeps that combination safe. Nodes that advertise
and nodes that look up ship first, and servers set `maxProvidersPerKey` only
once most of the `ADD_PROVIDER` requests that they see arrive on version
`2.0.0`. Adoption at that scale takes months. While it runs, a server that
enforces a limit applies it to version `2.0.0` requests only.

### Security

**An acknowledgment is not proof of storage.** With `ACCEPTED`, a server claims
that it stored the record. The protocol cannot prove the storage of any record,
and it cannot detect a server that answers `ACCEPTED` and stores nothing. Such a
server absorbs records and stops the spillover, exactly as a server that drops
records under load does. If the `k` closest servers collude, they suppress an
advertisement completely. The base protocol has the same weakness, because the
same servers can discard a record. A node that needs a stronger assurance checks
the result out of band, for example with a `GET_PROVIDERS` to each server that
it advertised to.

**False rejections.** An adversary that controls the closest servers to a key
can reject every `ADD_PROVIDER` and suppress an advertisement. The spillover
limits a unanimous rejection: every record short of `k` moves the node outwards,
to servers outside the set that the adversary controls. An adversary that mixes
acceptances and rejections across the chunks reduces the replication. Rejections
alone cannot stop the storage of the record, because the spillover continues
while the count stays below `k`.

**Eviction and Peer ID rotation.** A policy that always drops the oldest record
lets an attacker flush the honest providers out of the `k` closest servers, for
one request per eviction. The attacker only has to rotate its Peer ID, and the
base protocol has no such censorship vector. The Eviction section removes the
gain: the candidates are the records that are older than the republish interval,
which their providers had to refresh already.

**Slot monopolisation.** With no eviction policy, the first
`maxProvidersPerKey` providers of a key hold their slots while they republish.
The spillover moves every later provider outwards. This is the safe default. A
later provider loses proximity to the key, and keeps reachability, because the
lookup finds a record that spilled over.

### Alternatives

**A limit with no signal.** A server can cap the records that it stores today,
and drop the rest silently. This is what an overloaded server does. The
advertising node keeps its count of `k`, believes that the record is stored, and
the CID becomes harder to resolve with no way to detect it.

**A retry against the same servers.** A node that treats a rejection as a
transient error and retries reaches the same overloaded servers, and adds load
to the hotspot that the limit protects.

**A dedicated error message type instead of a status field.** A new message type
carries the same information, and every implementation has to route it. The
optional field 11 on the existing response keeps the exchange to one
request and one response.

**A new field with no protocol version bump.** A sender then cannot tell a
server that implements the extension from one that does not, so it waits for a
full timeout on every legacy server. Version `2.0.0` moves that detection into
libp2p protocol negotiation, which happens before the request.

### Copyright

Copyright and related rights waived via [CC0](https://creativecommons.org/publicdomain/zero/1.0/).
Loading