Repository navigation
ToyOS's TCP offers a scaled window that grows by the receiver's own round trip, learned from the peer's own segment size, and RFC 8985's loss probe keeps a lost window update from stalling a sender - #820
Conversation
…so ToyOS's TCP offers a scaled window toyos-net-tcp already negotiated RFC 7323 window scaling and timestamps with PAWS; it took the smallest shift that fits the receive buffer, and netstack's 65,535-byte buffer made that shift 0, so a sender never had more than 64 KiB in flight. The T14's download read 35.1 Mb/s, 65,535 bytes per 14.9 ms. - netstack's TCP_BUFFER is 4 MiB each way: gigabit to a 33 ms round trip, at window scale 7. PLACE_BYTES is now a stream's (two pipes and two 4 MiB buffers, 12 MiB) rather than a listener's. - The receive buffer starts at 65,535 (limits::RECEIVE_BUFFER_INITIAL) and, once a round trip has passed, grows to twice what the user read per round trip, up to the configured buffer. Only reads grow it: a peer sending into a connection nobody reads, a listener's unaccepted children included, finds 65,535 bytes of room and no more. A window field at the negotiated shift bounds it, so an unscaled connection never grows past 65,535. - A ring's storage grows, by doubling, with the furthest offset written, so a 4 MiB send buffer costs what it holds. - After a read, a window update leaves when the window could at least double. netstack, like the test network, transmits a batch's ACKs before its reader drains the batch, so the ACKs offer the window less that batch; without the update the sender learned of the room only with the next data's ACK, one window per two round trips, and the growth rule, reading half a window per round trip, never grew the buffer past about 133 KB. Linux's tcp_cleanup_rbuf sends the same update. - A scaled window that would round to a zero field is offered as one unit where the buffer has room, as Linux's __tcp_select_window does. With a hole open the edge holds, and at shift 7 a window under 128 bytes read as zero: the sender persisted and resent one byte of the hole per backed-off probe. A property run at the new buffer found it. s_net_005 runs at netstack's buffer: at 65,535 a window-limited sender's window reopens each round on one update ACK, which its every-second-ACK loss takes. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017cSFvbD35xJ2kGANVdm23C
|
Mutation patches at m1-no-scale-offereddiff --git a/toyos-net-shard/tcp/src/open.rs b/toyos-net-shard/tcp/src/open.rs
index 7322f8503..af20ede15 100644
--- a/toyos-net-shard/tcp/src/open.rs
+++ b/toyos-net-shard/tcp/src/open.rs
@@ -160,7 +160,7 @@ fn syn_options(local: &Local, peer: Option<&Negotiated>, now: Instant) -> SynOpt
mss: Some(local.mss),
sack_permitted: peer.is_none_or(|n| n.sack),
timestamps,
- window_scale: if peer.is_none_or(|n| n.scaled) { WindowShift::new(local.shift).ok() } else { None },
+ window_scale: None,
}
}
m2-absent-option-ignoreddiff --git a/toyos-net-shard/tcp/src/open.rs b/toyos-net-shard/tcp/src/open.rs
index 7322f8503..278ba060c 100644
--- a/toyos-net-shard/tcp/src/open.rs
+++ b/toyos-net-shard/tcp/src/open.rs
@@ -38,7 +38,7 @@ pub fn negotiate(seg: &In<'_>, local: &Local, first_tsval: u32, ctx: &mut Ctx<'_
}
Some(mss) => mss,
};
- let scaled = options.window_scale().is_some();
+ let scaled = true;
let (snd_shift, rcv_shift) = match options.window_scale() {
Some(scale) => {
if scale.raw() > scale.effective() {
@@ -46,7 +46,7 @@ pub fn negotiate(seg: &In<'_>, local: &Local, first_tsval: u32, ctx: &mut Ctx<'_
}
(scale.effective(), local.shift)
}
- None => (0, 0),
+ None => (0, local.shift),
};
let ts = options.timestamps().map(|t| Ts { recent: t.value, recent_at: ctx.now, offset: local.ts_offset, first: first_tsval });
Negotiated { peer_mss, snd_shift, rcv_shift, scaled, sack: options.sack_permitted(), ts }m3-no-growthdiff --git a/toyos-net-shard/tcp/src/rx.rs b/toyos-net-shard/tcp/src/rx.rs
index a23e09cc2..a12195bc7 100644
--- a/toyos-net-shard/tcp/src/rx.rs
+++ b/toyos-net-shard/tcp/src/rx.rs
@@ -392,7 +392,7 @@ impl Rx {
let read = u128::try_from(read).unwrap_or(u128::MAX);
let per_rtt = read.saturating_mul(rtt.as_nanos()).checked_div(now.since(began).as_nanos()).unwrap_or(0);
let want = usize::try_from(per_rtt.saturating_mul(2)).unwrap_or(usize::MAX);
- self.buf.grow(want.min(self.max));
+ let _ = want;
self.round = (now, 0);
}
if let Some(candidate) = self.candidate(mss) {m4-fixed-full-bufferdiff --git a/toyos-net-shard/tcp/src/rx.rs b/toyos-net-shard/tcp/src/rx.rs
index a23e09cc2..54e3b37b2 100644
--- a/toyos-net-shard/tcp/src/rx.rs
+++ b/toyos-net-shard/tcp/src/rx.rs
@@ -90,7 +90,7 @@ impl Rx {
next,
edge: next.add(window),
shift,
- buf: Ring::new(max.min(initial)),
+ buf: Ring::new(max),
ranges: Vec::new(),
stamp: 0,
fin: None,m5-no-window-updatediff --git a/toyos-net-shard/tcp/src/rx.rs b/toyos-net-shard/tcp/src/rx.rs
index a23e09cc2..1dd7f64cc 100644
--- a/toyos-net-shard/tcp/src/rx.rs
+++ b/toyos-net-shard/tcp/src/rx.rs
@@ -396,7 +396,7 @@ impl Rx {
self.round = (now, 0);
}
if let Some(candidate) = self.candidate(mss) {
- if self.last_window < self.threshold(mss) || candidate.since(self.next) >= self.window().saturating_mul(2) {
+ if self.last_window < self.threshold(mss) {
self.ack_now = true;
}
}m6-no-round-updiff --git a/toyos-net-shard/tcp/src/rx.rs b/toyos-net-shard/tcp/src/rx.rs
index a23e09cc2..a60e8a25b 100644
--- a/toyos-net-shard/tcp/src/rx.rs
+++ b/toyos-net-shard/tcp/src/rx.rs
@@ -274,7 +274,7 @@ impl Rx {
// room, as Linux does: else a sender owing the text of a hole waits on a window not shut.
let unit = 1u32.checked_shl(u32::from(self.shift)).unwrap_or(u32::MAX);
let window = edge.since(self.next);
- let edge = if window > 0 && window < unit && unit <= self.free() { self.next.add(unit) } else { edge };
+ let _ = (window, unit);
let field = (edge.since(self.next) >> self.shift).min(u32::from(u16::MAX));
(edge, u16::try_from(field).unwrap_or(u16::MAX))
}m7-no-relayoutdiff --git a/toyos-net-shard/tcp/src/ring.rs b/toyos-net-shard/tcp/src/ring.rs
index aaf45daab..e47345fce 100644
--- a/toyos-net-shard/tcp/src/ring.rs
+++ b/toyos-net-shard/tcp/src/ring.rs
@@ -53,7 +53,7 @@ impl Ring {
/// offset, laid out again from the head.
fn reserve(&mut self, end: usize) {
if end > self.bytes.len() {
- self.bytes.rotate_left(self.head);
+ let _ = 0;
self.head = 0;
let len = end.max(self.bytes.len().saturating_mul(2)).min(self.capacity);
self.bytes.resize(len, 0);m8-no-offerable-capdiff --git a/toyos-net-shard/tcp/src/rx.rs b/toyos-net-shard/tcp/src/rx.rs
index a23e09cc2..65693a591 100644
--- a/toyos-net-shard/tcp/src/rx.rs
+++ b/toyos-net-shard/tcp/src/rx.rs
@@ -85,7 +85,7 @@ impl Rx {
pub fn new(next: Seq, max: usize, shift: u8, window: u32, now: Instant) -> Self {
let initial = usize::try_from(crate::limits::RECEIVE_BUFFER_INITIAL).unwrap_or(usize::MAX);
let offerable = usize::from(u16::MAX).checked_shl(u32::from(shift)).unwrap_or(usize::MAX);
- let max = max.min(offerable);
+ let _ = offerable;
Self {
next,
edge: next.add(window), |
|
T14 at |
|
Review of #820 at Net lines, own change: production +102/−37 (net +65: BLOCKER
NOTE
SEND BACK |
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017cSFvbD35xJ2kGANVdm23C
…loss probe, and the receive buffer grows by the receiver's own round trip Review of #820 at d286ece sent four things back; each is answered here. s_net_005_lost_acks runs at 65,535 again, beside netstack's 4 MiB, and at both drop parities. At 65,535 with the read-triggered window update, the update that reopens a window-limited round is at times the tail's only ACK; with every second ACK lost, the sender then sat on a sub-segment window with one segment out until the RTO, which cut cwnd to one segment, and the same loss pattern took each later round's single ACK: 1 MiB had not arrived after 582 s. No receiver rule fixes that, because the lost ACK is the last one the peer sends; a sender timer must. So TCP now sends RFC 8985 §7's tail loss probe: with SACK, outside recovery and with nothing SACKed, two SRTTs after the last send or advancing ACK (plus WCDelAckT with one segment out, plus Linux's 2 ms otherwise, and never after the RTO, which it then stands in for once), it sends a segment of new data where the peer's window takes a whole one, ignoring cwnd, and otherwise the last segment again. Per §7.4 a resent probe acknowledged without a D-SACK repaired a real loss and reduces cwnd once. The arm now asserts no RTO, no probe-inferred loss, and that every byte sent twice was a probe's and came back as a D-SACK. The receive buffer grows by the receiver's own round-trip estimate, not the sender's SRTT, which a downloader samples only from its own data and so keeps at the handshake's value: with timestamps, the age of each new TSecr on a full-sized in-order segment, averaged with gain 1/8 (Linux's tcp_rcv_rtt_measure_ts); without, the least time the peer took to fill the window offered (tcp_rcv_rtt_measure). the_receive_window_grows_by_what_is_read_per_round_trip pins the growth rate both ways: the path's RTT quadruples after the handshake, the reader takes two 40 KB bursts per 20 ms, and the window stays within twice the read rate over the estimated round trip plus a burst and a segment, while the reader is never kept waiting once it has grown. Both of the review's mutations (x64, and growth on every read) red it. a_sub_unit_window_rounds_up_only_into_free_room pins the round-up's `unit <= free` guard at shift 7 with a hole open, and props.rs's check now asserts rx's invariant unread + (edge - next) <= capacity after every event. Info carries the receive capacity and the receiver's estimate. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017cSFvbD35xJ2kGANVdm23C
|
Mutation patches of round 2, at g1-growth-x64diff --git a/toyos-net-shard/tcp/src/rx.rs b/toyos-net-shard/tcp/src/rx.rs
index 76e6b3979..dd076a25b 100644
--- a/toyos-net-shard/tcp/src/rx.rs
+++ b/toyos-net-shard/tcp/src/rx.rs
@@ -442,7 +442,7 @@ impl Rx {
if let Some(rtt) = self.rtt.filter(|&rtt| now.since(began) >= rtt) {
let read = u128::try_from(read).unwrap_or(u128::MAX);
let per_rtt = read.saturating_mul(rtt.as_nanos()).checked_div(now.since(began).as_nanos()).unwrap_or(0);
- let want = usize::try_from(per_rtt.saturating_mul(2)).unwrap_or(usize::MAX);
+ let want = usize::try_from(per_rtt.saturating_mul(64)).unwrap_or(usize::MAX);
self.buf.grow(want.min(self.max));
self.round = (now, 0);
}g2-grown-on-every-readdiff --git a/toyos-net-shard/tcp/src/rx.rs b/toyos-net-shard/tcp/src/rx.rs
index 76e6b3979..9c3d39c92 100644
--- a/toyos-net-shard/tcp/src/rx.rs
+++ b/toyos-net-shard/tcp/src/rx.rs
@@ -439,7 +439,7 @@ impl Rx {
let (began, read) = self.round;
let read = read.saturating_add(n);
self.round = (began, read);
- if let Some(rtt) = self.rtt.filter(|&rtt| now.since(began) >= rtt) {
+ if let Some(rtt) = self.rtt {
let read = u128::try_from(read).unwrap_or(u128::MAX);
let per_rtt = read.saturating_mul(rtt.as_nanos()).checked_div(now.since(began).as_nanos()).unwrap_or(0);
let want = usize::try_from(per_rtt.saturating_mul(2)).unwrap_or(usize::MAX);g3-no-echo-samplediff --git a/toyos-net-shard/tcp/src/rx.rs b/toyos-net-shard/tcp/src/rx.rs
index 9bbe2dec2..122194be1 100644
--- a/toyos-net-shard/tcp/src/rx.rs
+++ b/toyos-net-shard/tcp/src/rx.rs
@@ -145,7 +145,7 @@ impl Rx {
Sampler::Echo(last) => {
let Some((echo, Some(age))) = echo.filter(|&(echo, _)| len >= mss && *last != Some(echo)) else { return };
*last = Some(echo);
- self.rtt = Some(self.rtt.map_or(age, |rtt| rtt.saturating_mul(7).saturating_add(age).checked_div(8).unwrap_or(age)));
+ let _ = age;
}
Sampler::Fill(mark) => {
if let Some((edge, since)) = *mark {g4-fill-keeps-the-largestdiff --git a/toyos-net-shard/tcp/src/rx.rs b/toyos-net-shard/tcp/src/rx.rs
index 76e6b3979..9df0dfbf5 100644
--- a/toyos-net-shard/tcp/src/rx.rs
+++ b/toyos-net-shard/tcp/src/rx.rs
@@ -153,7 +153,7 @@ impl Rx {
return;
}
let took = now.since(since).max(Duration::from_micros(1));
- self.rtt = Some(self.rtt.map_or(took, |rtt| rtt.min(took)));
+ self.rtt = Some(self.rtt.map_or(took, |rtt| rtt.max(took)));
}
*mark = Some((self.edge, now));
}p1-no-loss-probediff --git a/toyos-net-shard/tcp/src/conn.rs b/toyos-net-shard/tcp/src/conn.rs
index 7de89bf2e..9d12ce79d 100644
--- a/toyos-net-shard/tcp/src/conn.rs
+++ b/toyos-net-shard/tcp/src/conn.rs
@@ -739,6 +739,9 @@ impl Sync {
/// RTO, which it then stands in for.
fn schedule_probe(&mut self, now: Instant) {
self.probe_at = None;
+ if self.probe_at.is_none() {
+ return;
+ }
let Some(srtt) = self.rtt.srtt() else { return };
let Some(rto) = self.rtx_timer else { return };
if !self.sack_ok || self.recovery != Recovery::None || !self.tx.sacked().is_empty() || self.probe.is_some() || self.persist.is_some() {p2-no-loss-responsediff --git a/toyos-net-shard/tcp/src/conn.rs b/toyos-net-shard/tcp/src/conn.rs
index 7de89bf2e..69fb6ef70 100644
--- a/toyos-net-shard/tcp/src/conn.rs
+++ b/toyos-net-shard/tcp/src/conn.rs
@@ -643,7 +643,7 @@ impl Sync {
// RFC 8985 §7.4: a retransmitted probe acknowledged without a D-SACK repaired a loss.
if let Some((_, resent)) = self.probe.filter(|&(end, _)| ack.at_or_after(end) && acked <= flight) {
self.probe = None;
- if resent && !dsack {
+ if resent && !dsack && false {
self.cc.on_loss(flight);
self.cc.cwnd = self.cc.cwnd.min(self.cc.ssthresh);
self.cc.end_recovery();p3-dsack-ignoreddiff --git a/toyos-net-shard/tcp/src/conn.rs b/toyos-net-shard/tcp/src/conn.rs
index 7de89bf2e..b8cb1df52 100644
--- a/toyos-net-shard/tcp/src/conn.rs
+++ b/toyos-net-shard/tcp/src/conn.rs
@@ -643,7 +643,7 @@ impl Sync {
// RFC 8985 §7.4: a retransmitted probe acknowledged without a D-SACK repaired a loss.
if let Some((_, resent)) = self.probe.filter(|&(end, _)| ack.at_or_after(end) && acked <= flight) {
self.probe = None;
- if resent && !dsack {
+ if resent {
self.cc.on_loss(flight);
self.cc.cwnd = self.cc.cwnd.min(self.cc.ssthresh);
self.cc.end_recovery();p4-no-delayed-ack-allowancediff --git a/toyos-net-shard/tcp/src/conn.rs b/toyos-net-shard/tcp/src/conn.rs
index 7de89bf2e..c529e08d0 100644
--- a/toyos-net-shard/tcp/src/conn.rs
+++ b/toyos-net-shard/tcp/src/conn.rs
@@ -744,7 +744,7 @@ impl Sync {
if !self.sack_ok || self.recovery != Recovery::None || !self.tx.sacked().is_empty() || self.probe.is_some() || self.persist.is_some() {
return;
}
- let delayed = if self.tx.flight() <= self.smss() { WORST_DELAYED_ACK } else { PROBE_SLACK };
+ let delayed = if self.tx.flight() <= self.smss() { PROBE_SLACK } else { PROBE_SLACK };
self.probe_at = Some(now.after(srtt.saturating_mul(2).saturating_add(delayed)).min(rto));
}
r1-round-up-unguardeddiff --git a/toyos-net-shard/tcp/src/rx.rs b/toyos-net-shard/tcp/src/rx.rs
index 76e6b3979..316c6ef97 100644
--- a/toyos-net-shard/tcp/src/rx.rs
+++ b/toyos-net-shard/tcp/src/rx.rs
@@ -325,7 +325,7 @@ impl Rx {
// room, as Linux does: else a sender owing the text of a hole waits on a window not shut.
let unit = 1u32.checked_shl(u32::from(self.shift)).unwrap_or(u32::MAX);
let window = edge.since(self.next);
- let edge = if window > 0 && window < unit && unit <= self.free() { self.next.add(unit) } else { edge };
+ let edge = if window > 0 && window < unit { self.next.add(unit) } else { edge };
let field = (edge.since(self.next) >> self.shift).min(u32::from(u16::MAX));
(edge, u16::try_from(field).unwrap_or(u16::MAX))
} |
|
Review of #820, round 2, at Net lines. Whole branch against Round 1's BLOCKERs
BLOCKER
NOTE
SEND BACK |
|
T14 at |
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017cSFvbD35xJ2kGANVdm23C
…ur send MSS, and the loss probe follows RFC 8985 §7.1 to §7.4 The T14 read 24.1 Mb/s at round 2's head, against 328.2 at round 1's on the same machine and URL. The cause, reproduced in the crate's host network at the T14's shape (gigabit, 15 ms round trip, timestamps and SACK): the receiver sampled its round trip only from segments at least this end's *send* MSS long. A peer whose segments are shorter, here one sending 1,380 or 1,428 bytes to a downloader whose send MSS is 1,448, was never sampled, so the buffer never grew past 65,535: 34.3 Mb/s for 128 MiB at round 2's crate, 1202.4 at round 1's and at this one. The loss probe is a sender's mechanism and a downloader sends only its request; it is not the cause. - Full-sized is learned from what arrives, as Linux's tcp_measure_rcv_mss: a segment at least as long as the largest so far raises it, up to what our SYN offered less the timestamp option; two in a row of one shorter length, not under the floor MSS, lower it. - An echo's age is at least one tick (Linux's tcp_rtt_tsopt_us) and at most twice the estimate, so the echo of an ACK sent before the peer fell silent moves the estimate by an eighth, not by the silence. - RFC 8985 §7.4.2: only an ACK past the probe's end, without a D-SACK matching that end or a duplicate ACK without SACK first, infers that a resent probe repaired a loss; an ACK at the end leaves the episode open. - §7.3: no probe is scheduled without an RTT sample since the last probe. - §7.2: none in RTO recovery, until SND.UNA passes the point the RTO set. - §7.1: entering fast recovery ends the probe's episode. - A probe that came due but had not left gives way when the window shuts. - Info's rcv_capacity and rcv_rtt and Rx::rtt go: nothing shipping read them. Rx::capacity is the property checker's alone, cfg(test). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017cSFvbD35xJ2kGANVdm23C
|
Mutation patches of round 3, at a1-probe-loss-at-its-enddiff --git a/toyos-net-shard/tcp/src/conn.rs b/toyos-net-shard/tcp/src/conn.rs
index b2cadac..1508dc4 100644
--- a/toyos-net-shard/tcp/src/conn.rs
+++ b/toyos-net-shard/tcp/src/conn.rs
@@ -649,7 +649,7 @@ impl Sync {
if let Some((end, resent)) = self.probe.filter(|&(end, _)| ack.at_or_after(end) && acked <= flight) {
if !resent || dsack == Some(end) || (same && seg.options.sack_blocks().len() == 0) {
self.probe = None;
- } else if ack.after(end) {
+ } else if ack.at_or_after(end) {
self.probe = None;
self.cc.on_loss(flight);
self.cc.cwnd = self.cc.cwnd.min(self.cc.ssthresh);a2-any-dsack-ends-the-episodediff --git a/toyos-net-shard/tcp/src/conn.rs b/toyos-net-shard/tcp/src/conn.rs
index b2cadac..ecfd397 100644
--- a/toyos-net-shard/tcp/src/conn.rs
+++ b/toyos-net-shard/tcp/src/conn.rs
@@ -647,7 +647,7 @@ impl Sync {
// a duplicate without SACK ends the episode with nothing lost; only an ACK past the end
// without either says a resent probe repaired a loss.
if let Some((end, resent)) = self.probe.filter(|&(end, _)| ack.at_or_after(end) && acked <= flight) {
- if !resent || dsack == Some(end) || (same && seg.options.sack_blocks().len() == 0) {
+ if !resent || dsack.is_some() || (same && seg.options.sack_blocks().len() == 0) {
self.probe = None;
} else if ack.after(end) {
self.probe = None;a3-no-duplicate-casediff --git a/toyos-net-shard/tcp/src/conn.rs b/toyos-net-shard/tcp/src/conn.rs
index b2cadac..9f7e403 100644
--- a/toyos-net-shard/tcp/src/conn.rs
+++ b/toyos-net-shard/tcp/src/conn.rs
@@ -647,7 +647,7 @@ impl Sync {
// a duplicate without SACK ends the episode with nothing lost; only an ACK past the end
// without either says a resent probe repaired a loss.
if let Some((end, resent)) = self.probe.filter(|&(end, _)| ack.at_or_after(end) && acked <= flight) {
- if !resent || dsack == Some(end) || (same && seg.options.sack_blocks().len() == 0) {
+ if !resent || dsack == Some(end) {
self.probe = None;
} else if ack.after(end) {
self.probe = None;e1-echo-unclampeddiff --git a/toyos-net-shard/tcp/src/rx.rs b/toyos-net-shard/tcp/src/rx.rs
index 8e98cde..e1999d9 100644
--- a/toyos-net-shard/tcp/src/rx.rs
+++ b/toyos-net-shard/tcp/src/rx.rs
@@ -160,7 +160,7 @@ impl Rx {
*last = Some(echo);
let Some(age) = age.filter(|_| len >= self.rcv_mss) else { return };
let age = age.max(TICK);
- self.rtt = Some(self.rtt.map_or(age, |rtt| rtt.saturating_mul(7).saturating_add(age.min(rtt.saturating_mul(2))).checked_div(8).unwrap_or(rtt)));
+ self.rtt = Some(self.rtt.map_or(age, |rtt| rtt.saturating_mul(7).saturating_add(age).checked_div(8).unwrap_or(rtt)));
}
Sampler::Fill(mark) => {
if let Some((edge, since)) = *mark {e2-zero-tick-takendiff --git a/toyos-net-shard/tcp/src/rx.rs b/toyos-net-shard/tcp/src/rx.rs
index 8e98cde..ed5c430 100644
--- a/toyos-net-shard/tcp/src/rx.rs
+++ b/toyos-net-shard/tcp/src/rx.rs
@@ -159,7 +159,7 @@ impl Rx {
let Some((echo, age)) = echo.filter(|&(echo, _)| *last != Some(echo)) else { return };
*last = Some(echo);
let Some(age) = age.filter(|_| len >= self.rcv_mss) else { return };
- let age = age.max(TICK);
+ let age = age.max(Duration::ZERO);
self.rtt = Some(self.rtt.map_or(age, |rtt| rtt.saturating_mul(7).saturating_add(age.min(rtt.saturating_mul(2))).checked_div(8).unwrap_or(rtt)));
}
Sampler::Fill(mark) => {g1-growth-x64diff --git a/toyos-net-shard/tcp/src/rx.rs b/toyos-net-shard/tcp/src/rx.rs
index 8e98cde..0511578 100644
--- a/toyos-net-shard/tcp/src/rx.rs
+++ b/toyos-net-shard/tcp/src/rx.rs
@@ -472,7 +472,7 @@ impl Rx {
if let Some(rtt) = self.rtt.filter(|&rtt| now.since(began) >= rtt) {
let read = u128::try_from(read).unwrap_or(u128::MAX);
let per_rtt = read.saturating_mul(rtt.as_nanos()).checked_div(now.since(began).as_nanos()).unwrap_or(0);
- let want = usize::try_from(per_rtt.saturating_mul(2)).unwrap_or(usize::MAX);
+ let want = usize::try_from(per_rtt.saturating_mul(64)).unwrap_or(usize::MAX);
self.buf.grow(want.min(self.max));
self.round = (now, 0);
}g2-grown-on-every-readdiff --git a/toyos-net-shard/tcp/src/rx.rs b/toyos-net-shard/tcp/src/rx.rs
index 8e98cde..2c4bcaa 100644
--- a/toyos-net-shard/tcp/src/rx.rs
+++ b/toyos-net-shard/tcp/src/rx.rs
@@ -469,7 +469,7 @@ impl Rx {
let (began, read) = self.round;
let read = read.saturating_add(n);
self.round = (began, read);
- if let Some(rtt) = self.rtt.filter(|&rtt| now.since(began) >= rtt) {
+ if let Some(rtt) = self.rtt {
let read = u128::try_from(read).unwrap_or(u128::MAX);
let per_rtt = read.saturating_mul(rtt.as_nanos()).checked_div(now.since(began).as_nanos()).unwrap_or(0);
let want = usize::try_from(per_rtt.saturating_mul(2)).unwrap_or(usize::MAX);g3-no-echo-samplediff --git a/toyos-net-shard/tcp/src/rx.rs b/toyos-net-shard/tcp/src/rx.rs
index 8e98cde..3d455be 100644
--- a/toyos-net-shard/tcp/src/rx.rs
+++ b/toyos-net-shard/tcp/src/rx.rs
@@ -160,7 +160,7 @@ impl Rx {
*last = Some(echo);
let Some(age) = age.filter(|_| len >= self.rcv_mss) else { return };
let age = age.max(TICK);
- self.rtt = Some(self.rtt.map_or(age, |rtt| rtt.saturating_mul(7).saturating_add(age.min(rtt.saturating_mul(2))).checked_div(8).unwrap_or(rtt)));
+ let _ = age;
}
Sampler::Fill(mark) => {
if let Some((edge, since)) = *mark {g4-fill-keeps-the-largestdiff --git a/toyos-net-shard/tcp/src/rx.rs b/toyos-net-shard/tcp/src/rx.rs
index 8e98cde..210c521 100644
--- a/toyos-net-shard/tcp/src/rx.rs
+++ b/toyos-net-shard/tcp/src/rx.rs
@@ -168,7 +168,7 @@ impl Rx {
return;
}
let took = now.since(since).max(Duration::from_micros(1));
- self.rtt = Some(self.rtt.map_or(took, |rtt| rtt.min(took)));
+ self.rtt = Some(self.rtt.map_or(took, |rtt| rtt.max(took)));
}
*mark = Some((self.edge, now));
}m1-full-sized-is-our-mssdiff --git a/toyos-net-shard/tcp/src/rx.rs b/toyos-net-shard/tcp/src/rx.rs
index 8e98cde..fac8b7f 100644
--- a/toyos-net-shard/tcp/src/rx.rs
+++ b/toyos-net-shard/tcp/src/rx.rs
@@ -158,7 +158,7 @@ impl Rx {
Sampler::Echo(last) => {
let Some((echo, age)) = echo.filter(|&(echo, _)| *last != Some(echo)) else { return };
*last = Some(echo);
- let Some(age) = age.filter(|_| len >= self.rcv_mss) else { return };
+ let Some(age) = age.filter(|_| len >= self.mss_bounds.1) else { return };
let age = age.max(TICK);
self.rtt = Some(self.rtt.map_or(age, |rtt| rtt.saturating_mul(7).saturating_add(age.min(rtt.saturating_mul(2))).checked_div(8).unwrap_or(rtt)));
}m2-full-sized-never-lowereddiff --git a/toyos-net-shard/tcp/src/rx.rs b/toyos-net-shard/tcp/src/rx.rs
index 8e98cde..c7d56a4 100644
--- a/toyos-net-shard/tcp/src/rx.rs
+++ b/toyos-net-shard/tcp/src/rx.rs
@@ -184,9 +184,7 @@ impl Rx {
self.rcv_mss = len.min(offered);
} else if len >= floor {
self.short = len;
- if len == short {
- self.rcv_mss = len;
- }
+ let _ = short;
}
}
p1-no-loss-probediff --git a/toyos-net-shard/tcp/src/conn.rs b/toyos-net-shard/tcp/src/conn.rs
index b2cadac..765065a 100644
--- a/toyos-net-shard/tcp/src/conn.rs
+++ b/toyos-net-shard/tcp/src/conn.rs
@@ -746,6 +746,9 @@ impl Sync {
/// slack when more are, and never after the RTO, which it then stands in for.
fn schedule_probe(&mut self, now: Instant) {
self.probe_at = None;
+ if self.probe_at.is_none() {
+ return;
+ }
let Some(srtt) = self.rtt.srtt() else { return };
let Some(rto) = self.rtx_timer else { return };
let recovering = self.recovery != Recovery::None || (self.episode && self.tx.una.at_or_before(self.recover));p2-no-loss-responsediff --git a/toyos-net-shard/tcp/src/conn.rs b/toyos-net-shard/tcp/src/conn.rs
index b2cadac..3922098 100644
--- a/toyos-net-shard/tcp/src/conn.rs
+++ b/toyos-net-shard/tcp/src/conn.rs
@@ -649,7 +649,7 @@ impl Sync {
if let Some((end, resent)) = self.probe.filter(|&(end, _)| ack.at_or_after(end) && acked <= flight) {
if !resent || dsack == Some(end) || (same && seg.options.sack_blocks().len() == 0) {
self.probe = None;
- } else if ack.after(end) {
+ } else if ack.after(end) && false {
self.probe = None;
self.cc.on_loss(flight);
self.cc.cwnd = self.cc.cwnd.min(self.cc.ssthresh);p3-dsack-ignoreddiff --git a/toyos-net-shard/tcp/src/conn.rs b/toyos-net-shard/tcp/src/conn.rs
index b2cadac..158a1a8 100644
--- a/toyos-net-shard/tcp/src/conn.rs
+++ b/toyos-net-shard/tcp/src/conn.rs
@@ -647,7 +647,7 @@ impl Sync {
// a duplicate without SACK ends the episode with nothing lost; only an ACK past the end
// without either says a resent probe repaired a loss.
if let Some((end, resent)) = self.probe.filter(|&(end, _)| ack.at_or_after(end) && acked <= flight) {
- if !resent || dsack == Some(end) || (same && seg.options.sack_blocks().len() == 0) {
+ if !resent || false || (same && seg.options.sack_blocks().len() == 0) {
self.probe = None;
} else if ack.after(end) {
self.probe = None;p4-no-delayed-ack-allowancediff --git a/toyos-net-shard/tcp/src/conn.rs b/toyos-net-shard/tcp/src/conn.rs
index b2cadac..43f94da 100644
--- a/toyos-net-shard/tcp/src/conn.rs
+++ b/toyos-net-shard/tcp/src/conn.rs
@@ -752,7 +752,7 @@ impl Sync {
if !self.sack_ok || recovering || !self.tx.sacked().is_empty() || self.probe.is_some() || !self.sampled || self.persist.is_some() {
return;
}
- let delayed = if self.tx.flight() <= self.smss() { WORST_DELAYED_ACK } else { PROBE_SLACK };
+ let delayed = if self.tx.flight() <= self.smss() { PROBE_SLACK } else { PROBE_SLACK };
self.probe_at = Some(now.after(srtt.saturating_mul(2).saturating_add(delayed)).min(rto));
}
r1-round-up-unguardeddiff --git a/toyos-net-shard/tcp/src/rx.rs b/toyos-net-shard/tcp/src/rx.rs
index 8e98cde..21c4577 100644
--- a/toyos-net-shard/tcp/src/rx.rs
+++ b/toyos-net-shard/tcp/src/rx.rs
@@ -355,7 +355,7 @@ impl Rx {
// room, as Linux does: else a sender owing the text of a hole waits on a window not shut.
let unit = 1u32.checked_shl(u32::from(self.shift)).unwrap_or(u32::MAX);
let window = edge.since(self.next);
- let edge = if window > 0 && window < unit && unit <= self.free() { self.next.add(unit) } else { edge };
+ let edge = if window > 0 && window < unit { self.next.add(unit) } else { edge };
let field = (edge.since(self.next) >> self.shift).min(u32::from(u16::MAX));
(edge, u16::try_from(field).unwrap_or(u16::MAX))
}s1-no-sample-conditiondiff --git a/toyos-net-shard/tcp/src/conn.rs b/toyos-net-shard/tcp/src/conn.rs
index b2cadac..b7ee310 100644
--- a/toyos-net-shard/tcp/src/conn.rs
+++ b/toyos-net-shard/tcp/src/conn.rs
@@ -749,7 +749,7 @@ impl Sync {
let Some(srtt) = self.rtt.srtt() else { return };
let Some(rto) = self.rtx_timer else { return };
let recovering = self.recovery != Recovery::None || (self.episode && self.tx.una.at_or_before(self.recover));
- if !self.sack_ok || recovering || !self.tx.sacked().is_empty() || self.probe.is_some() || !self.sampled || self.persist.is_some() {
+ if !self.sack_ok || recovering || !self.tx.sacked().is_empty() || self.probe.is_some() || self.persist.is_some() {
return;
}
let delayed = if self.tx.flight() <= self.smss() { WORST_DELAYED_ACK } else { PROBE_SLACK };s2-probe-in-rto-recoverydiff --git a/toyos-net-shard/tcp/src/conn.rs b/toyos-net-shard/tcp/src/conn.rs
index b2cadac..4a66310 100644
--- a/toyos-net-shard/tcp/src/conn.rs
+++ b/toyos-net-shard/tcp/src/conn.rs
@@ -748,7 +748,7 @@ impl Sync {
self.probe_at = None;
let Some(srtt) = self.rtt.srtt() else { return };
let Some(rto) = self.rtx_timer else { return };
- let recovering = self.recovery != Recovery::None || (self.episode && self.tx.una.at_or_before(self.recover));
+ let recovering = self.recovery != Recovery::None;
if !self.sack_ok || recovering || !self.tx.sacked().is_empty() || self.probe.is_some() || !self.sampled || self.persist.is_some() {
return;
}s3-probe-kept-into-fast-recoverydiff --git a/toyos-net-shard/tcp/src/conn.rs b/toyos-net-shard/tcp/src/conn.rs
index b2cadac..e37d459 100644
--- a/toyos-net-shard/tcp/src/conn.rs
+++ b/toyos-net-shard/tcp/src/conn.rs
@@ -815,8 +815,7 @@ impl Sync {
return;
}
// RFC 8985 §7.1: fast recovery starts the loss probe's state afresh.
- (self.probe_at, self.probe_due, self.probe) = (None, false, None);
- self.arm(ctx.now);
+ self.probe_at = None;
let flight = self.tx.flight();
self.cc.on_loss(flight.saturating_sub(self.lt_bytes));
self.recover = self.tx.nxt.sub(1);s4-due-probe-kept-into-persistdiff --git a/toyos-net-shard/tcp/src/conn.rs b/toyos-net-shard/tcp/src/conn.rs
index b2cadac..f030145 100644
--- a/toyos-net-shard/tcp/src/conn.rs
+++ b/toyos-net-shard/tcp/src/conn.rs
@@ -846,7 +846,7 @@ impl Sync {
if self.persist.is_none() {
let rto = self.rtt.rto();
self.persist = Some(Persist { at: now.after(rto), interval: rto, from: self.tx.nxt, due: false, unanswered: false });
- (self.rtx_timer, self.probe_at, self.probe_due) = (None, None, false);
+ (self.rtx_timer, self.probe_at) = (None, None);
}
} else if let Some(persist) = self.persist.take() {
if self.tx.nxt.after(persist.from) { |
|
Review of #820, round 3, at Net lines. Whole branch against Earlier BLOCKERs
Round 2's NOTEs are answered. The idle-aged echo is fixed (e1 and e2 red Gates at BLOCKER
NOTE
SEND BACK |
|
T14 at
|
…s cwnd once Round 3's review: one ACK past a resent loss probe's end that also SACKs enough to enter recovery ran RFC 8985 §7.4.2's reduction and then recovery's own on_loss, cutting twice for one congestion event (and lowering CUBIC's W_max and cwnd_prior twice). Recovery entry now skips its reduction when the same ACK already took the probe's, as Linux's tcp_enter_recovery skips tcp_init_cwnd_reduction in CWR; the probe's cut stands, from the flight before the ACK. rfc_8985_7_4_an_ack_that_infers_the_probes_loss_and_enters_recovery_cuts_once reds before (ssthresh 4054, the second cut from the flight after the ACK) and greens after (5068). Also restores recovery.rs's first doc line, broken by a stray paste. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017cSFvbD35xJ2kGANVdm23C
|
Mutation patch of round 4, at c1-recovery-cuts-againdiff --git a/toyos-net-shard/tcp/src/conn.rs b/toyos-net-shard/tcp/src/conn.rs
--- a/toyos-net-shard/tcp/src/conn.rs
+++ b/toyos-net-shard/tcp/src/conn.rs
@@ -822,3 +822,3 @@
// One ACK is one congestion event, cut once, as Linux's `tcp_enter_recovery` in CWR.
- if !cut {
+ if !cut || true {
self.cc.on_loss(flight.saturating_sub(self.lt_bytes)); |
|
Review of #820, round 4, at Net lines. Whole branch against Earlier BLOCKERs
Round 3's NOTEs:
BLOCKER
NOTE
SEND BACK |
…s cwnd twice, as Linux does Round 4's review withdrew round 3's premise. In Linux v6.12, tcp_ack runs tcp_process_tlp_ack (tcp_input.c:4038) before tcp_fastretrans_alert (:4048). Its loss branch (:3846-3849) runs tcp_init_cwnd_reduction, enters CWR, ends the reduction and calls tcp_try_keep_open (:2760), which leaves CWR for Open or Disorder. tcp_enter_recovery (:2897) then finds tcp_in_cwnd_reduction (tcp.h:1333, CWR or Recovery) false and runs tcp_init_cwnd_reduction again (:2918). Linux cuts twice on that ACK. The `cut` flag guarded one ACK, not one congestion event: the probe's reduction sets no recovery point, so an ACK later in the same window that entered recovery cut again regardless. It is deleted and duplicate(ack, ctx) restored. The test stays and now asserts Linux's sequence: ssthresh and cwnd 4054, the second cut taken from the flight after the ACK (4 * 1448 * 7 / 10). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017cSFvbD35xJ2kGANVdm23C
|
Round 5 negative control, diff --git a/toyos-net-shard/tcp/src/conn.rs b/toyos-net-shard/tcp/src/conn.rs
index b2cadac0b..83e025b7e 100644
--- a/toyos-net-shard/tcp/src/conn.rs
+++ b/toyos-net-shard/tcp/src/conn.rs
@@ -646,6 +646,7 @@ impl Sync {
// RFC 8985 §7.4.2: at or past the probe's end, a probe of new data, a D-SACK of the probe or
// a duplicate without SACK ends the episode with nothing lost; only an ACK past the end
// without either says a resent probe repaired a loss.
+ let mut cut = false;
if let Some((end, resent)) = self.probe.filter(|&(end, _)| ack.at_or_after(end) && acked <= flight) {
if !resent || dsack == Some(end) || (same && seg.options.sack_blocks().len() == 0) {
self.probe = None;
@@ -655,19 +656,20 @@ impl Sync {
self.cc.cwnd = self.cc.cwnd.min(self.cc.ssthresh);
self.cc.end_recovery();
ctx.log.count(Counter::LossProbeRecovery);
+ cut = true;
}
}
if acked > 0 && acked <= flight {
self.new_ack(seg, ack, acked, flight, ctx);
if newly {
- self.duplicate(ack, ctx);
+ self.duplicate(ack, cut, ctx);
}
} else if self.sack_ok {
if newly {
- self.duplicate(ack, ctx);
+ self.duplicate(ack, cut, ctx);
}
} else if flight > 0 && seg.payload.is_empty() && !seg.syn() && !seg.fin() && ack == self.tx.una && window == self.tx.wnd {
- self.duplicate(ack, ctx);
+ self.duplicate(ack, cut, ctx);
}
if ack.at_or_after(self.tx.una) && (self.tx.wl1.before(seg.seq) || (self.tx.wl1 == seg.seq && self.tx.wl2.at_or_before(ack))) {
self.tx.wnd = window;
@@ -788,8 +790,9 @@ impl Sync {
}
}
- /// A duplicate acknowledgment: RFC 5681 §2's without SACK, RFC 6675 §2's with it.
- fn duplicate(&mut self, ack: Seq, ctx: &mut Ctx<'_>) {
+ /// A duplicate acknowledgment: RFC 5681 §2's without SACK, RFC 6675 §2's with it. `cut` says
+ /// this ACK already reduced cwnd for the loss a resent probe repaired.
+ fn duplicate(&mut self, ack: Seq, cut: bool, ctx: &mut Ctx<'_>) {
let smss = self.smss();
match &mut self.recovery {
Recovery::Fast { inflations } => {
@@ -818,7 +821,10 @@ impl Sync {
(self.probe_at, self.probe_due, self.probe) = (None, false, None);
self.arm(ctx.now);
let flight = self.tx.flight();
- self.cc.on_loss(flight.saturating_sub(self.lt_bytes));
+ // One ACK is one congestion event, cut once, as Linux's `tcp_enter_recovery` in CWR.
+ if !cut {
+ self.cc.on_loss(flight.saturating_sub(self.lt_bytes));
+ }
self.recover = self.tx.nxt.sub(1);
self.episode = true;
self.urgent = Some(self.tx.una); |
|
Review of #820, round 5, at Net lines. The whole branch against Earlier BLOCKERs
BLOCKERNone. NOTENone. LAND |
…ed by one pass over the streams Node::receive took one frame and ended in a pass over every stream, so a bulk download cost netstack one write of the client's pipe and one empty read of its send pipe per frame, and the client woke once per write. A QEMU e1000e profile of a 64 MiB download from the host at 6df8222 counted, per 64 MiB: 46,609 frames in 1,081 passes (43 a pass), 46,604 pipe writes, 46,669 empty send-pipe reads and 1,077 ACKs; the reader made 21,748 calls of a 16,389-byte buffer (3,086 bytes each) and 19,956 of a 64 KiB one (3,363 each), so its read size was set by netstack's writes and not by its buffer. Node::receive now takes the pass's frames through a pull closure, each frame handed to the stack and settled as before, and one pass over the streams follows the last. netstack passes Card::rx to it. The pass still precedes the transmit opportunity, so the ACK carries the window the drain opened, as before. What it buys is set by how many frames a pass holds. Under QEMU's TCG, 43 a pass, the same download made 2,103 pipe writes over 1,089 passes. On the T14's I219, which is programmed with no interrupt moderation, almost every pass held 0 to 2 frames: one warm download counted 71,384 pipe writes over 101,403 passes and 119,890 frames, about 15% fewer pipe calls than one a frame. No CPU or bandwidth change is measured: one boot, CPU per MB 13.15 ms warm and 13.49 cold, against 13.5 to 15.4 across #820's. The batches grow only once the card holds its interrupt for more frames, which is a change of its own. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017cSFvbD35xJ2kGANVdm23C
…812), TCP window scaling (#820) and the stop's hold of the console wire (#805), into virtio-sound and the shared PCI claim Both conflicts are two additions at one site, and both sides are kept whole: - tests/common/qemu.rs: this branch's Profile::HeadlessVirtioGpu and main's Profile::HeadlessUsbSpare, each in the enum, the x86-64 arm of arch() and shape(). - tests/toyos.rs: RUST_SKIP takes both virtio_sound_counts and usbd_spare; run_machine_test takes this branch's virtio_sound_counts arm beside main's machine_shutdown_wire_* , usbd_drives_the_spare and usb_keyboard_rollover. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017cSFvbD35xJ2kGANVdm23C
…h toyos-sha2 - userland/update/src/main.rs: the branch's stream_root over the block service's Disk (STREAM_BLOCKS of BLOCK bytes), hashed with main's toyos_sha2::Sha256 in place of sha2. - tests/toyos.rs: both sides' RUST_SKIP entries, MACHINE_TESTS entries, dispatch arms and functions kept, the branch's first; the branch's update_writes_the_idle_slot_through_the_block_service keeps its own close. - Cargo.lock: main's, regenerated by cargo against the merged manifests. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017cSFvbD35xJ2kGANVdm23C
… batch: the icons' and wallpaper's digests hash with toyos-sha2 Cargo.toml keeps both sides' entries: main's toyos-sha2 and usbd members and toyos-sha2 dependency, and the batch's removal of `image`. `sha2` was replaced by toyos-sha2 on main and left in place on the batch, whose #813 and #814 added two tests hashing with it; as main meant every SHA-256 the build takes to be toyos-sha2's, those two tests now hash with toyos-sha2 and the root manifest drops `sha2`. Cargo.lock is the batch's, re-resolved by `cargo metadata --offline`; its delta from the batch's head is exactly main's delta from the merge base. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017cSFvbD35xJ2kGANVdm23C
, #826, #835, #833, #831 and #822, into wt/toyos-netperf No hunk conflicted. userland/netstack/src/main.rs took both sides: main's removal of `mod device` and the branch's batched `node.receive` and its module-doc line. The TCP window-scaling and loss-probe commits main carries were already in the branch from #820, so their files merged to main's text plus the branch's own delta. Both lockfiles are main's and pass `cargo metadata --locked`. The branch's new issue still cites `VirtioNet::poll_rx` and `toyos_i219::RX_BUDGET` as they stand on main. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017cSFvbD35xJ2kGANVdm23C
This branch merged
origin/mainlast at49dca4e02(which carries #810, #816 and #823). Its own change isd286ece79(round 1),cd34232c9(round 2),6df8222c1(round 3),a700899ee(round 4) ande7efa0a87(round 5, which deletes round 4's production change:conn.rsis again6df8222c1's, and against round 3 the branch adds only one test,recovery.rs+20/−1):git diff origin/main...49dca4e02, 13 files, +824/−67. The figures that follow are round 3's. Production (toyos-net-shard/tcp/srcwithout thecfg(test)props.rsandrx.rs's test module, plus netstack): +294/−44. Tests: +512/−24.What changed and why
toyos-net-tcpalready negotiated RFC 7323 window scaling (both ways, never unless both SYNs offer it, the shift clamped at 14 and countedWscaleClamped) and timestamps with PAWS. It took the smallest shift that fits the receive buffer, and netstack'sTCP_BUFFERof 65,535 made that shift 0. So the work is the buffer, what its growth needs, and what a sender needs once windows reopen by update.TCP_BUFFERis 4 MiB each way, window scale 7. At 1 Gb/s that covers a 33.5 ms round trip; Ubuntu on the T14 autotunes to 6 MB (tcp_rmemmax).PLACE_BYTESbecomes a stream's: two 2 MiB pipes and two 4 MiB buffers, 12 MiB, up from the listener's 10 MiB: about 170 places on 16 GiB instead of about 200, under the same one-eighth share.rx's module doc). It starts atlimits::RECEIVE_BUFFER_INITIAL= 65,535, the most a SYN offers. Once a round trip has passed it grows to twice what the user read per round trip, up to the configured buffer, and never shrinks (RFC 7323 §2.4). Text thetoyos-net-tcpuser does not read grows nothing. At netstack this does not bound a client that stops reading:sockets.bridgedrains a connection into its client's 2 MiB pipe whether or not the client reads, so such a connection's buffer can still grow toward 4 MiB; it stays insidePLACE_BYTES. A window field at the negotiated shift bounds the growth, so an unscaled connection stays at 65,535.tcp_rcv_rtt_measure_ts); without them, the least time the peer took to fill the window offered (Linux'stcp_rcv_rtt_measure).tcp_measure_rcv_mss): a segment at least as long as the largest so far raises it, up to what our SYN offered less the timestamp option; two in a row of one shorter length, not under the floor MSS less the option, lower it. Round 2 measured against this end's send MSS, which is the T14 regression below.tcp_rtt_tsopt_us). The ceiling is ours: the echo of the last ACK before the peer fell silent ages by the silence, and now moves the estimate by an eighth of itself, not by an eighth of the silence; a real rise is followed at ×1.125 per new echo, about one a millisecond in a transfer.tcp_cleanup_rbufrule). netstack transmits a received batch's ACKs beforesockets.bridgedrains the batch; without the update the sender learned of the room one round trip late and the growth rule settled at about 133 KB.s_net_005) the sender sat until the RTO, cwnd fell to one segment, and 1 MiB had not arrived at t = 582,893 ms. Now, with SACK, two SRTTs after the last new data or advancing ACK (plus WCDelAckT, 200 ms, with one segment out; plus Linux'sTCP_TIMEOUT_MIN, 2 ms, otherwise; never later than the RTO, which it then stands in for once), the sender sends a segment of new data where the peer's window takes a whole one, ignoring cwnd, and otherwise the last segment again (§7.3). A duplicate ACK or new SACK information cancels it. Counterstcp.loss-probeandtcp.loss-probe-recovery. This is RFC 8985's TLP without RACK: loss below the tail is still found by RFC 6675. Round 3 brings it to the RFC's text:tcp_process_tlp_ack). An ACK at the end leaves the episode open, since the original's ACK reads the same.tx::read_sacknow returns the D-SACK's right edge.recoverpoint the expiry set.tcp_ackrunstcp_process_tlp_ack(tcp_input.c:4038) beforetcp_fastretrans_alert(:4048); its loss branch (:3846-3849) runstcp_init_cwnd_reduction, enters CWR, runstcp_end_cwnd_reductionand callstcp_try_keep_open(:2760-2772), which leaves CWR for Open or Disorder; sotcp_enter_recovery(:2897) findstcp_in_cwnd_reduction(tcp.h:1333, CWR or Recovery only) false and runstcp_init_cwnd_reductionagain (:2915-2919). Read fromtcp_input.cat tag v6.12, sha2568007c66e…5d859749d. The flag also guarded one ACK, not one congestion event: the probe's reduction sets norecover, so an ACK later in the same window that entered recovery cut again anyway.__tcp_select_window). With a hole open the edge holds; at shift 7 a 93-byte window read as zero and a property run stalled 600 s.Info.rcv_capacity,Info.rcv_rttandRx::rtt()are gone (round 3): nothing shipping read them.Rx::capacity()iscfg(test), read byprops.rs's capacity invariant alone. The tests observe the window throughrcv_edge − rcv_nxt, and the estimator throughrx's own unit tests.The T14 regression at round 2, and its cause
The orchestrator's T14 reading at
0350817cd(round 2 plus the measurement commit): 24.1 Mb/s, the download cut at the window after 120 MB, where round 1's head did 328.2 Mb/s on the same machine and URL. 24.1 Mb/s is under the 65,535 ceiling at itsrtt_ms(65,535 · 8 / 15.0 ms = 35.0 Mb/s), and every CPU sat under 5% busy: a window that never grew, not a CPU bound. (Its 31.9 ms of CPU per MB is idle overhead spread over a slow transfer.)Reproduced in the crate's host network at the T14's shape: node 1 sends 128 MiB to node 0 at 125,000 bytes per ms of transmit credit, a 15 ms round trip, timestamps and SACK both ways, 4 MiB buffers; node 0's SYN rewritten to offer an MSS of 1,460, 1,440 or 1,392, so node 1's segments are 1,448, 1,428 or 1,380 bytes against node 0's own send MSS of 1,448. The test file is
repro-t14.rsbeside the logs; same file, each crate copied whole at its head,cargo test --release --test t14, each EXIT=0:d286ece79cd34232c96df8222c1(The harness's credit is per millisecond with no queue, so its rate reads above the gigabit; the rows compare crates, not links.) The cause is round 2's echo sampler: it took a sample only from a segment at least
smss()long, this end's send MSS. A timestamped peer whose segments are shorter, by an MSS it clamps lower than ours or a path it sizes for, was never sampled,rcv_rttstayedNone, and the buffer stayed at 65,535 for the connection's life. The peer of the T14's URL is CloudFront; probed from this development machine on the same day, it negotiates timestamps, SACK and window scale and offers an MSS of 1,440 (macOS'sTCP_CONNECTION_INFO:maxseg1,428). Its segments' size on the T14's path is not read here; the measurement commit now prints the receiver's learnedrcv_mssto read it. The other two candidates are ruled out: the loss probe is a sender's mechanism and the downloader sends only its request (serverrto0 and no retransmitted byte in every row); the measurement commit changes nothing in the stack and the regression reproduces without it.Tests and their controls
Host tests, ours against ours (the crate's consistency control):
a_peer_sending_segments_shorter_than_our_send_mss_still_grows_the_window(new,net.rs): the T14 shape above with 1,380-byte segments, 32 MiB; asserts the peer's longest segment is 1,380, both shifts 7, and the receiver's window reaches the whole 4 MiB.rfc_8985_7_4_a_resent_probe_acknowledged_past_its_end_without_a_dsack_repaired_a_loss(replaces round 2's…acknowledged_without_a_dsack…, which ACKed exactly the end): the ACK at the end leaves the episode open and reduces nothing; the ACK past it reduces cwnd once.rfc_8985_7_4_a_dsack_after_the_ack_at_the_probes_end_infers_no_loss(the review's): ACK at the end, then the D-SACK, then an ACK past the end:tcp.loss-probe-recovery0, cwnd stands.rfc_8985_7_4_a_dsack_of_another_segment_still_infers_the_lossandrfc_8985_7_4_a_duplicate_without_sack_at_the_probes_end_infers_no_loss(Case 2).rfc_8985_7_3_no_probe_without_an_rtt_sample_since_the_last(the review's,net.rs): no timestamps, SACK, a 20 ms handshake and then a 100 ms path, ten whole segments a round, 150 ms between rounds so the last probe's D-SACK has ended its episode. Probes leave in rounds 1, 3, 5 and 7, every other round, then none; SRTT reaches 99.3 ms. Asserted: no two rounds in a row probe, none after round 20, SRTT in [90, 110] ms, no RTO, no inferred loss. The review asked for at most one probe: under RFC 6298 one sample a round moves SRTT an eighth of the way, so SRTT needs four samples to pass RTT/2, and §7.3 permits a probe in each round that follows one. Without the condition (mutation s1) a probe leaves every round and SRTT stays at 20 ms.rfc_8985_7_4_an_ack_that_infers_the_probes_loss_and_enters_recovery_cuts_twice(round 3's NOTE, its assertion corrected in round 5): ten out, the probe resends 14033, five more go out on the ACK at its end, then one ACK of 16929 SACKing 18377–22721: oneloss-probe-recovery, onesack-recovery, ssthresh and cwnd 4,054, the second cut's, from the 4 × 1,448 in flight after the ACK. Its oracle is the Linux sequence above: two reductions on the one ACK, the probe's then recovery's.rfc_8985_7_2_no_probe_in_rto_recovery(the review's):s_lr_019's timeline, then a plainack(2449)at 232 ms and time to past 2·SRTT + 2 ms and to just before the RTO: one probe, one RTO.rfc_8985_7_1_fast_recovery_ends_the_probes_episode: the resent probe outstanding as SACK recovery begins; recovery's own reduction is the only one.a_probe_due_when_the_window_shuts_gives_way_to_persist(the NOTE): the PTO fires with the next hop pending, an ACK shuts the window, the hop wakes: nothing leaves, no probe counted.rx's unit tests (new):full_sized_is_learned_from_what_arrives(1,380 sampled; one 1,380 after a 1,448 not, the second in a row lowers full-sized and is),an_echo_aged_by_the_peers_silence_moves_the_estimate_an_eighth(16 ms, then a 10 s echo moves it to 18 ms, then a 0 ms echo counts as 1 ms),without_timestamps_the_least_fill_time_is_kept.the_receive_window_grows_by_what_is_read_per_round_trip(round 2's, now withoutrcv_rtt): handshake at 20 ms, then an 80 ms path; the downloader's SRTT asserted to stay the handshake's; the reader takes 4 KB/ms in two bursts every 20 ms. Asserts, with timestamps and without, that the window never exceeds2·(R·120 ms + burst) + one MSS, an estimate at most half again the path's, and that over the last 2 s the reader got exactly its rate, which an estimate under the path's would not give.s_net_005_lost_acksat 65,535 and 4 MiB, each drop parity;a_sub_unit_window_rounds_up_only_into_free_roomandprops.rs's capacity invariant;rfc_8985_7_3_the_probe_is_new_data…,rfc_8985_7_4_…reported_as_a_duplicate…,rfc_8985_7_2_with_one_segment_out…;s_lr_018/s_lr_019meeting the probe at 22 ms before their RTO at 222 ms;rfc_7323_2_…in_flight_each_way,rfc_7323_2_2_no_scaling_unless_both_syns_offer_it,a_receive_buffer_nobody_reads_never_grows.Negative control. Production
src/at round 2's headcd34232c9under this round'stests/: EXIT=101, nine reds, every new or changed scripted and network test above (control-round2-src.log).Mutations, round 3, each a checked patch on a copy of the crate at
6df8222c1,cargo test --offline --no-fail-fast,git apply -R; every one EXIT=101, RESTORED=0, the copy clean; the unmutated copy EXIT=0, 405 passed. Patches in the comment of this round. a2, a3 and g4 stayed green in a first run, and s1 and s4 did against the first form of their tests; the tests above were added or changed for them and the whole set run again.rfc_8985_7_4_…above,rfc_8985_7_3_no_probe…rfc_8985_7_4_a_dsack_of_another_segment…rfc_8985_7_4_a_duplicate_without_sack…rfc_8985_7_3_no_probe_without_an_rtt_sample…rfc_8985_7_2_no_probe_in_rto_recoveryrfc_8985_7_1_fast_recovery_ends…a_probe_due_when_the_window_shuts…full_sized_is_learned…,a_peer_sending_segments_shorter…full_sized_is_learned…an_echo_aged_by_the_peers_silence…the_receive_window_grows…the_receive_window_grows…,a_peer_sending_segments_shorter…,rfc_7323_2_…in_flight_each_wayrxtests and four network tests on growthwithout_timestamps_the_least_fill_time_is_keptunit <= freedeleteda_sub_unit_window…, 12 property tests through the capacity invariants_net_005_lost_acks, everyrfc_8985_*but7_3_no_probe…,s_lr_018,s_lr_019rfc_8985_7_4_…past_its_end…,rfc_8985_7_4_a_dsack_of_another…s_net_005_lost_acks,rfc_8985_7_4_a_dsack_after…,rfc_8985_7_3_no_probe…rfc_8985_7_2_with_one_segment_out…,s_rt_010_a_retransmission_echoedMutation, round 5,
cut1: round 4'scutflag re-applied at49dca4e02as a checked patch (git diff 6df8222c1 a700899ee -- toyos-net-shard/tcp/src/conn.rs, posted in this round's comment).git apply --check, applied,cargo test --offline --no-fail-fastin the crate: EXIT=101, only…enters_recovery_cuts_twicered,left: (5068, 5068),right: (4054, 4054);git apply -RRESTORED=0,git status --porcelainempty (r5/logs/mut-cut-once.log).Round 1's m1–m8 and its whole-change revert are in the first mutation comment and stand for
d286ece79's part.Independent oracle. RFC 8985 §7's text (§7.1–§7.4.2's pseudocode) and RFC 7323; Linux v6.12's
tcp_input.c(tcp_process_tlp_ack,tcp_try_keep_open,tcp_enter_recovery,tcp_measure_rcv_mss,tcp_rcv_rtt_measure_ts,tcp_rtt_tsopt_us) andtcp_output.c(tcp_schedule_loss_probe,tcp_send_loss_probe) as read; and the T14 against a CDN for the scaled path. The host's TCP cannot be the peer without a TAP device and privileges, and QEMU's slirp never offers Window Scale (tcp_dooptionsparses MSS alone). A smoltcp peer, which round 1's review proposed, is declined: the owner ruled "no more smoltcp. fast track make it gone", and #801 removed it.No new guest test. A QEMU guest's peer is slirp, which never scales, never loses an ACK and never sends short segments on its own; the probe, the growth and the segment-size learning are reached by the host network's impairment, rewriting and clock, which a guest cannot drive. The whole guest suite checks the no-offer path against a third-party TCP.
Gates at
49dca4e02(round 5)Logs under the orchestrator's
wscale/r5/logs/, each from one script run on the committed head: its first linehead 49dca4e02…, its second the command, its last line the command's ownEXIT=. The worktree'sgit status --porcelain --ignore-submodules=nonewas empty after the last (gates.done).cargo test --offline --no-fail-fastintoyos-net-shard/tcp: EXIT=0, 406 passed (crate-tests.log).cargo run -- --ci host: EXIT=0,[ci] Host: 78 step(s), all green(ci-host.log).cargo run -- --build-only: EXIT=0 (build-only.log).cargo test: EXIT=0,44 passed, 44 total(and the harness's own 348 passed, 15 ignored), withnetstack_streams,netstack_streams_e1000e,netstack_socket_churnandlibc_socketsall PASS; host load averages 66.92 52.41 52.41 before, 52.17 53.40 52.93 after (guest.log).49dca4e02.git diff 6df8222c1 49dca4e02 -- toyos-net-shard userland/netstackisrecovery.rsalone, so the production code the reading ran is byte-identical. Round 4'sa700899eeand its gates are superseded: its production change is deleted.red-before.logandcrate-tests.logare no longer cited. The first had no EXIT line, and the second had neither an EXIT line nor the head.Gates at
6df8222c1(round 3)cargo test --offline --no-fail-fastintoyos-net-shard/tcp: EXIT=0, 405 passed (tcp-tests.log).cargo run -- --ci host: EXIT=0,[ci] Host: 78 step(s), all green(ci-host.log).cargo run -- --build-only: EXIT=0 (build-only.log).cargo test: EXIT=0,43 passed, 43 total,netstack_streams,netstack_streams_e1000e,netstack_socket_churn,libc_socketsamong them; host load averages 66.37 before, 64.66 after (guest.log).wt/toyos-wscale-metalat30c6678ec(this head merged ate1b14a5e4onto round 2's measurement branch, plus one measurement-only commit that never lands: the download twice back to back withcpu_sandcpu_ms_per_mbper run, and netstack saying each connection'srcv_shift,rcv_capacity,rcv_wnd,rcv_rtt,rcv_mss,srtt,cwndand the stack's recovery counts, on the client's close and now also when the client goes without one, which is how round 2's cut run said nothing).cargo test --test toyos-build -- --metal --metal-readback <dir> internet_download: EXIT=2, staged, the machine not touched; image sha256e2c09588…36b8a6fb. The orchestrator's reading (PR comment): judge EXIT=0,PASS internet_download; first download 107.1 Mb/s, second 662.4 Mb/s, bothrcv_shift=7 rcv_capacity=4194304 rcv_mss=1424,loss_probe=0 retransmit_bytes=0;rtt_ms15.0–17.8. Round 2's 24.1 Mb/s regression is gone.What I am unsure of
rcv_mss=1424), shorter than our send MSS less the option, the shape the host reproduction showed round 2 never sampled.on_losstakes ssthresh from FlightSize (RFC 5681 §3.2, RFC 6675), so the second cut is 0.7 × the 5,792 in flight after the ACK = 4,054, where Linux's CUBIC (tcp_cubic.c:341-356) takes it from cwnd in segments, which the first cut has already lowered. The test asserts the count and order of the cuts against Linux; its value is our FlightSize rule's, not compared with Linux's.RX_RINGof 256 × 2 KiB descriptors is about 3.1 ms of gigabit frames beforemissed; netstack takesRX_BUDGET64 frames a pass;TX_RINGis 16. Round 1 used about 13.7 ms of CPU per MB where Ubuntu uses about 2.🤖 Generated with Claude Code
https://claude.ai/code/session_017cSFvbD35xJ2kGANVdm23C