Skip to content

fix(server): route a session's requests to the worker that holds it - #3905

Draft
ananthsub wants to merge 2 commits into
mainfrom
ananthsub/fix-session-worker-routing
Draft

ananthsub wants to merge 2 commits into
mainfrom
ananthsub/fix-session-worker-routing

Conversation

@ananthsub

Copy link
Copy Markdown
Contributor

What changed and why

uvicorn's workers share one listening socket, and each request goes to whichever worker accepts the connection. Session state, keyed by the signed session cookie, lives in the memory of the worker that created the session. Nothing keeps a session's later requests on that worker.

Measured with the stateful counter resources server (example_session_state_mgmt) and 200 concurrent episodes through /run:

Workers Reward 1.0 Reward 0.0
1 200 0
4, before this change 7 to 27 173 to 193
4, after this change 200 0

Servers on main with this shape:

  • swe_rebench runs 8 workers. It stores each session's sandbox in an in-memory dict at /seed_session and pops it at /verify. A /verify on another worker finds nothing, the patch is recorded as empty, and the reward is 0.
  • Agents with 4 workers are affected wherever they keep per-session state across requests.
  • CLI harnesses in sandboxes reach a resources server's tools over /mcp. They send the signed MCP session token minted at /seed_session, and no Gym session cookie. With 4 workers and 200 concurrent counter sequences over MCP, 65 of 200 ended with the right count.

This change, for resources servers and agent servers running with num_workers > 1:

  • Owner in the cookie and the MCP token. Each new session is stamped with the ID of the worker that created it. The ID is stored inside the signed session cookie, next to the session ID. For a server that exposes its tools over MCP, the session token minted at /seed_session carries the same ID.
  • Private sockets. Each worker also serves the same app on a private Unix socket. The sockets live in a directory the main process creates under /tmp, which keeps paths short.
  • Routing.
    • A request without a session cookie or MCP token is handled locally, and its new session belongs to this worker.
    • On the MCP path, and on any request without a session cookie, the owner comes from the MCP token header. The token is verified with the same serializer and salt as the MCP endpoint uses. A token without an owner, or with a bad signature, is handled locally, where the MCP endpoint's own check applies as before.
    • A request for a session another worker owns is forwarded to that worker's private socket unchanged (method, path, query, headers including cookies, and body). The reply is streamed back unchanged.
    • A request that arrives over a private socket is always handled locally, so nothing is forwarded twice.
  • Lost sessions are explicit. If the owner's socket is gone, the request gets a 410 with a clear message. This happens when the worker exited and uvicorn replaced it, or when the server restarted. Before this change, the request ran silently against empty state.

example_session_state_mgmt also gains the module-level app that multi-worker uvicorn imports. Without it, the server could not run with more than one worker.

How it works

The router sits outside Starlette's SessionMiddleware. It reads the owner by decoding the signed cookie with the same secret. A forwarded request never reaches the receiving worker's SessionMiddleware, so the reply carries only the owner's Set-Cookie, built from the session as the owner left it. If the router sat inside SessionMiddleware, the receiving worker would add its own stale copy of the session as a second Set-Cookie.

The owner travels in the cookie or token, so there is no central session table. Creating or ending a session needs no extra round trip, and nothing grows with the number of sessions a server has served.

sequenceDiagram
    participant C as Caller
    participant B as Worker B
    participant A as "Worker A (owner)"
    C->>B: POST /verify with session cookie
    B->>B: decode cookie and read owner A
    B->>A: same request over A's private socket
    A->>A: SessionMiddleware and handler with A's state
    A-->>B: reply with A's Set-Cookie
    B-->>C: reply streamed back unchanged
Loading

An MCP tool call takes the same path, with the token header instead of the cookie:

sequenceDiagram
    participant H as "CLI harness (no cookies)"
    participant B as Worker B
    participant A as "Worker A (owner)"
    H->>A: POST /seed_session
    A-->>H: MCP token signed with session ID and owner A
    H->>B: POST /mcp tools/call with token header
    B->>B: verify token and read owner A
    B->>A: same request over A's private socket
    A->>A: MCP endpoint runs the tool on A's state
    A-->>B: tool result
    B-->>H: tool result streamed back unchanged
Loading
flowchart LR
    P["Main process (serves no requests)"] -->|creates| D["Socket directory under /tmp"]
    P -->|spawns| W1[Worker 1]
    P -->|spawns| W2[Worker 2]
    L[Shared listening port] --> W1
    L --> W2
    W1 -->|serves| S1["Private socket of worker 1"]
    W2 -->|serves| S2["Private socket of worker 2"]
    S1 --- D
    S2 --- D
    W1 -->|forwards sessions owned by worker 2| S2
    W2 -->|forwards sessions owned by worker 1| S1
    T["Session cookie or MCP token, naming the owner"] -.-> W1
    T -.-> W2
Loading

Details:

  • The routing is on by default for SimpleResourcesServer and SimpleResponsesAPIAgent, through a class attribute. Model servers keep no session state, so they opt out. Routing them would only concentrate load.
  • The forwarding client is aiohttp with one UnixConnector per owner socket. Its keep-alive is shorter than the private server's, so the forwarding side always closes idle connections first.
  • The private server runs on the worker's own event loop. It does not run the app's lifespan again, and it does not add its own date or server headers. It starts after the app's startup finishes and stops before the app's shutdown.
  • A client that disconnects cancels the forward, and the owner's own cancellation middleware then cancels its handler.

Related work

This fix stands on its own: it targets main and does not depend on the partial-rollout checkpointing stack (#3882 onward). That stack will build on it to support environment, agent, and resources servers with several workers; restoring a checkpoint will then also re-assign each restored session to a worker.

Issue

None known. The bug was found while measuring multi-worker behavior for partial-rollout checkpointing. This is the fix for main, separate from that work.

Validation

  • pytest tests/unit_tests/test_session_routing.py: 13 passed. Two in-process workers share a real socket directory. The tests cover:
    • a new session is handled locally and owned by the worker that handled it;
    • a session owned by another worker is forwarded there, with a single Set-Cookie;
    • method, path, query, and headers arrive unchanged;
    • a request over a private socket is handled locally;
    • a session whose worker shut down, or crashed and left its socket file, gets a 410;
    • a badly signed cookie is handled locally;
    • a cookie from a single-worker server is adopted;
    • a single-worker server has no routing;
    • an /mcp tool call with a token minted on worker A, received by worker B with no cookies, is forwarded to A, and A's tool sees the session state;
    • an MCP token without an owner is handled locally;
    • a badly signed MCP token is handled locally and refused by the MCP endpoint as before;
    • a single-worker server mints tokens without an owner.
  • pytest tests/unit_tests -n 8: 6233 passed, 160 skipped, 13 failed. The same 13 fail on main without this change: 11 in test_opensandbox_compose.py, 1 in test_sandbox_stop_retry.py, and 1 in test_slurm_script.py.
  • pytest for the example_session_state_mgmt, simple_agent, and legacy_agent server tests: 1, 20, and 5 passed. gym env test --resources-server example_session_state_mgmt: 1 passed.
  • NEMO_GYM_MULTIWORKER_E2E=1 pytest tests/e2e/session_routing: 4 passed, in about 22 seconds. Before this change, the same test gave 7 of 200 for the /run episodes, 127 of 200 for seed then verify, and 65 of 200 for the MCP sequences.
  • pre-commit run --files <changed files>: passed.

Rollout evidence

The opt-in e2e test runs real server processes against a fake OpenAI-compatible backend:

  • The legacy /run relay, Simple Agent, and the vLLM model server each run with 1 worker. The counter resources server runs with 4 workers.

  • 200 concurrent /run episodes, each scripted to increment the counter by 1 and then by 2: 200 of 200 score 1.0. Before this change, 7 of 200 did.

  • The swe_rebench shape: 200 concurrent sessions, where state is stored at /seed_session and checked at /verify. All 200 score 1.0, and the sessions are spread over several workers.

  • MCP without cookies, on the same server with expose_tools_over_mcp. 200 concurrent clients each:

    • seed a session over HTTP and keep only the MCP token;
    • call increment_counter by 1 and by 2 over /mcp;
    • read the count with get_counter_value.

    All 200 counts are right. On main, 65 of 200 were. With cookie routing alone, 72 of 200 were.

  • One worker is killed with SIGKILL after seeding. Every session of that worker gets a 410, and every other session still scores 1.0.

Pending:

  • a run with a real model;
  • a swe_rebench run at 8 workers with real sandboxes;
  • a multi-worker agent that keeps per-session state across requests.

Compatibility

  • Servers with one worker, and model servers, are unchanged. They get no middleware, no sockets, and no extra session key.
  • With several workers, resources and agent session cookies carry one more key, nemo_gym_worker, and so do MCP session tokens. Handlers that read request.session[SESSION_ID_KEY], and the MCP endpoint's token check, are unaffected. Tokens minted before this change still work and are handled by whichever worker receives them.
  • A request for a session whose worker is gone now gets a 410 instead of running against empty state.
  • A cookie minted before a multi-worker server restarted gets a 410, because its worker no longer exists. Before this change, it silently started over with empty state.
  • A request that lands on the wrong worker costs one local Unix-socket hop, which is about (N-1)/N of a session's requests after its first. Forwarded requests no longer see the caller's client address. The Host header and everything else are unchanged.
  • Some servers key state by an ID in the request body (resources_session_id and agent_session_id). They are routed because every Gym caller also sends the session cookie. A caller that sends only the body ID is not routed. The duplicate-seed and already-closed checks on a first, cookieless seed also see only the local worker.

uvicorn's workers share one listening socket, so each request goes to
whichever worker accepts the connection, while session state lives in the
memory of the worker that created the session. With the stateful counter
resources server at 4 workers, 200 concurrent episodes through /run scored
1.0 on only 7; swe_rebench (8 workers) can lose a session's sandbox at
/verify and score 0.

With num_workers > 1, resources and agent servers now:

- stamp each new session with the ID of the worker that created it, inside
  the signed session cookie;
- serve the same app on a private Unix socket per worker, in a directory
  the main process creates under /tmp;
- forward a request for a session another worker owns to that worker's
  private socket, unchanged, and stream the reply back unchanged. A
  forwarded request is always handled locally;
- answer 410 when the owner's socket is gone, instead of running against
  empty state.

The router sits outside SessionMiddleware and decodes the cookie with the
same secret, so a forwarded reply carries only the owner's Set-Cookie.
Single-worker servers and model servers are unchanged.

The counter resources server gains the module-level app that multi-worker
uvicorn imports, and an opt-in e2e test (NEMO_GYM_MULTIWORKER_E2E=1) runs
it with 4 workers behind the legacy /run relay and Simple Agent.

Signed-off-by: Ananth Subramaniam <ansubramania@nvidia.com>
CLI harnesses in sandboxes reach a resources server's tools over /mcp with
the signed session token minted at /seed_session, and send no Gym session
cookie. With several workers, those calls ran on whichever worker accepted
them: 4 workers, 200 concurrent counter sequences over MCP, 65 of 200
correct on main and 72 of 200 with cookie routing alone.

With routing active, the MCP token now also carries the owning worker's ID.
The router reads the owner from the token on the MCP path, and on any
request without a session cookie, verifying it with the same serializer
and salt. A token without an owner (minted by a single-worker server or
before this change) or with a bad signature is handled locally, where the
MCP endpoint's own check applies as before. Single-worker tokens are
unchanged. With this change, 200 of 200 are correct.

Signed-off-by: Ananth Subramaniam <ansubramania@nvidia.com>
@ananthsub ananthsub added bug Something isn't working area:core Shared APIs, servers, telemetry, health, and registries labels Oct 1, 2026
@copy-pr-bot

copy-pr-bot Bot commented Oct 1, 2026

Copy link
Copy Markdown

Auto-sync is disabled for draft pull requests in this repository. Workflows must be run manually.

Contributors can view more details about this message here.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

area:core Shared APIs, servers, telemetry, health, and registries bug Something isn't working

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant