feat(aio): run cluster labeling on the flex service tier - #92166
Merged
Merged
Conversation
The labeling agents are daily batch jobs, so they tolerate flex-tier queueing for tokens billed at half the standard rate. Only gpt-5 family models get the field; gpt-4.1 rejects it. The SDK's existing retries cover flex refusals and gateway timeouts. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
|
😎 Merged successfully - details. |
Contributor
🤖 CI report
|
Each cached workflow holds its full history, payloads included, so queues running short-lived, high-volume workflows need a bound far below the SDK default of 1000. Settable per deploy via the --max-cached-workflows flag or MAX_CACHED_WORKFLOWS. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…r rerun OpenAI decides flex eligibility per model (pro tiers excluded), so the family-prefix check becomes an explicit allowlist. Flex clients get a short per-call timeout with SDK retries off, and a shared runner reruns the agent on the standard tier when flex fails, since the labeling callers otherwise turn a flex outage into silent default labels. Logs the served tier and fallbacks. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Client construction moves outside the callers' catch-all so config errors fail the activity instead of shipping default labels, each attempt gets a fresh state copy, the standard rerun keeps the short per-call timeout, LLMA_LABELING_FLEX_ENABLED gives a no-deploy kill switch, and the worker test no longer binds a real Prometheus socket. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
A run where both tiers fail now returns the last attempt's partial labels instead of discarding them. Standard-tier labeling calls are capped at 240s with one SDK retry so two attempts fit the activity budget, including when the flex kill switch is pulled. Drops the kwargs dict spread and updates both READMEs. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
… labels A failed run raises LabelingAgentError carrying the fullest label set any attempt wrote; the graphs log the real cause and keep those labels before filling defaults, so failures no longer log as completed runs. Also corrects the rerun timeout in both READMEs, adds team_id and status_code to the labeling logs, and annotates make_agent. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…itch Drops the standard-tier rerun machinery: the SDK's own retries recover transient flex refusals, a still-failing run falls to the existing default-label handling and the next day's batch, and the env kill switch covers a sustained outage. Graphs and their tests return to master untouched. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The worker sticky-cache flag moves to a dedicated worker-tuning PR covering the kubernetes and temporal settings together. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…on and labeling posthog/llm/flex.py mirrors OpenAI's flex pricing table; both callers import it instead of keeping product-local copies in sync. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
… kill switch Flex is an OpenAI-specific tier, so the module name says so. Rolling labeling flex back is a one-line revert plus a deploy, matching the summarization rollout's trade-off, so the env lever goes. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
bernatixer
marked this pull request as ready for review
September 3, 2026 12:09
Radu-Raicea
requested changes
Sep 3, 2026
Radu-Raicea
left a comment
Member
There was a problem hiding this comment.
Let's do the retries on the non-flex version of the model, because it might be unavailable more often on flex than on non-flex.
The labeling agent also makes multiple multiple LLM calls, so if they all get significantly slower, I wonder if 10 min is still a good timeout.
… tier A recoverable flex failure (429, 408, 5xx, connection) re-issues only the failing chat completion with service_tier=default, so the agent keeps its completed turns. Flex clients drop SDK retries; the tier switch is the retry. The recoverability predicate moves to posthog/llm/openai_flex.py so both flex callers share it. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The first fallback latches the client to standard, so a flex brownout costs one 120s timeout per agent run instead of one per call. The recoverability predicate regains the SDK's 409 retry, the fallback log drops the response body, and the evaluation child workflow gets the 45min execution timeout its activity budgets already need. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Every request path flows through _get_request_payload, so the tier override lives once there; the fallback timeout field goes away since standard calls fit the flex client's 120s budget. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Radu-Raicea
approved these changes
Sep 4, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Problem
Changes
FlexFirstChatOpenAI(inllm_endpoint.py) overrides_generate/_agenerate: the openai SDK's own retry loop re-sends a byte-identical request, so it can only re-roll the refused tier; the override is the seam where the tier can change per call. The agent keeps its completed turns.flowchart TD A[agent turn N] --> B{{flex call<br/>120s, no SDK retries}} B -->|success| N[turn N+1] B -->|429 / 408 / 409 / 5xx / connection loss| C{{same call, standard tier}} C -->|success| L[latched: remaining turns go straight to standard] --> N C -->|failure| D[exception → existing default-label handling] classDef phBlue fill:#1d4aff,stroke:#1d4aff,color:#fff; classDef phRed fill:#f54e00,stroke:#f54e00,color:#fff; classDef phYellow fill:#f9bd2b,stroke:#f9bd2b,color:#000; class B,C phBlue; class D phRed; class A,N,L phYellow;is_flex_recoverable(429, connection loss and timeouts, 5xx, bare 408/409) moved from the summarization fallback (feat(aio): switch batch trace summarization to gpt-5-nano on flex #91639) toposthog/llm/openai_flex.py, next to the flex model allowlist both callers already share. 409 is new there: turning off the SDK's retry loop had silently dropped its 409 retry for both callers.gpt-5.4-prois excluded whilegpt-5.4is included, per the flex pricing table).max_retries=0(the tier switch is the retry), standard clients 240s × 2. The previous shape (600s × 3 per call) let one parked call outlive the labeling activity's whole 600s budget.LLM_SCHEDULE_TO_CLOSE_TIMEOUTrises 900s → 1260s: a first Temporal attempt that times out at 600s now leaves the retry a full 600s instead of ~295s. The evaluation child workflow gets the 45min execution timeout its activity budgets already needed, replacing a hardcoded 30min that predates this PR.Note
Cost ingestion prices the standard tier, so
$ai_total_cost_usddashboards show no change from this PR (#94200 fixes that at ingestion). The discount is visible only on the provider invoice. Same caveat, mechanism, and rollout check (service_tierpassthrough on the deployed gateway) as #91639.How did you test this code?
service_tier="default"— removing any member of the except tuple fails its row. A 400 propagates with no fallback; a standard-tier client never falls back; the async path falls back too.mypy --cache-fine-grained .is clean.Automatic notifications
Docs update
None: the labeling agent READMEs describe behavior that is unchanged by this PR.
🤖 Agent context
Autonomy: Human-driven (agent-assisted)
--max-cached-workflowsworker flag also lived here for a while and moves to a dedicated worker-tuning PR.