Skip to content

feat(aio): run cluster labeling on the flex service tier - #92166

Merged
trunk-io[bot] merged 16 commits into
masterfrom
feat/aio-clustering-labeling-flex
Sep 7, 2026
Merged

trunk-io[bot] merged 16 commits into
masterfrom
feat/aio-clustering-labeling-flex

Conversation

@bernatixer

@bernatixer bernatixer commented Sep 1, 2026 •

Copy link
Copy Markdown
Contributor

Problem

  • The cluster labeling agents run gpt-5.4 on the standard tier. They are daily batch jobs, so nobody waits on their latency.
  • OpenAI's flex service tier bills the same synchronous calls at half the token rate.
  • Flex is refused more often than standard (capacity), and a labeling run is a ReAct agent making tens of sequential LLM calls: one refused call must not throw away the run or silently degrade labels.

Changes

  • The clustering labeling agents (trace and evaluation clusters) now request the flex tier, halving their token bill. The agents themselves are untouched — only client construction changes.
  • A recoverable flex failure retries only the failing call on the standard tier, and latches the client to standard for the rest of that agent run — a flex brownout costs one 120s timeout per run, not one per call. FlexFirstChatOpenAI (in llm_endpoint.py) overrides _generate/_agenerate: the openai SDK's own retry loop re-sends a byte-identical request, so it can only re-roll the refused tier; the override is the seam where the tier can change per call. The agent keeps its completed turns.
flowchart TD
    A[agent turn N] --> B{{flex call<br/>120s, no SDK retries}}
    B -->|success| N[turn N+1]
    B -->|429 / 408 / 409 / 5xx / connection loss| C{{same call, standard tier}}
    C -->|success| L[latched: remaining turns go straight to standard] --> N
    C -->|failure| D[exception → existing default-label handling]
    classDef phBlue fill:#1d4aff,stroke:#1d4aff,color:#fff;
    classDef phRed fill:#f54e00,stroke:#f54e00,color:#fff;
    classDef phYellow fill:#f9bd2b,stroke:#f9bd2b,color:#000;
    class B,C phBlue;
    class D phRed;
    class A,N,L phYellow;
Loading
  • What counts as recoverable is shared, not re-derived: is_flex_recoverable (429, connection loss and timeouts, 5xx, bare 408/409) moved from the summarization fallback (feat(aio): switch batch trace summarization to gpt-5-nano on flex #91639) to posthog/llm/openai_flex.py, next to the flex model allowlist both callers already share. 409 is new there: turning off the SDK's retry loop had silently dropped its 409 retry for both callers.
  • Flex is gated by that explicit allowlist, not the gpt-5 family prefix: OpenAI decides eligibility per model (gpt-5.4-pro is excluded while gpt-5.4 is included, per the flex pricing table).
  • Per-call budgets close on every path: flex 120s with max_retries=0 (the tier switch is the retry), standard clients 240s × 2. The previous shape (600s × 3 per call) let one parked call outlive the labeling activity's whole 600s budget.
  • LLM_SCHEDULE_TO_CLOSE_TIMEOUT rises 900s → 1260s: a first Temporal attempt that times out at 600s now leaves the retry a full 600s instead of ~295s. The evaluation child workflow gets the 45min execution timeout its activity budgets already needed, replacing a hardcoded 30min that predates this PR.
  • A call that fails on both tiers falls to the existing default-label handling and self-heals with the next day's batch. Rolling flex back stays a one-line revert plus a deploy (fix(aio): remove the summarization model env lever #94013 removed the summarization env lever for the same reason).

Note

Cost ingestion prices the standard tier, so $ai_total_cost_usd dashboards show no change from this PR (#94200 fixes that at ingestion). The discount is visible only on the provider invoice. Same caveat, mechanism, and rollout check (service_tier passthrough on the deployed gateway) as #91639.

How did you test this code?

  • Fallback matrix, one row per recoverable error shape (429, bare-408, bare-409, 5xx, timeout): each asserts the second call carries service_tier="default" — removing any member of the except tuple fails its row. A 400 propagates with no fallback; a standard-tier client never falls back; the async path falls back too.
  • The allowlist rows guard granting flex to a pro model by family prefix; the budget pins guard reverting to 600s × 3 per call, which starves the activity.
  • The labeling agents, their tests, and their READMEs are byte-identical to master; repo-wide mypy --cache-fine-grained . is clean.
  • Timeout sizing checked against production telemetry: the slowest standard-tier labeling call of the past week ran 37s (p99 9.5s), and under 0.005% of three days of internal flex summarization traffic exceeded 120s — those were parked requests that also died at summarization's 180s cap, which is the case the fallback rescues.
  • Not run: a live labeling run through the deployed gateway; a failed flex call falls back to standard, and a misconfigured one fails to default labels as today.

Automatic notifications

  • Publish to changelog?

Docs update

None: the labeling agent READMEs describe behavior that is unchanged by this PR.

🤖 Agent context

Autonomy: Human-driven (agent-assisted)

  • Claude Code session directed by the assignee, carrying the review lessons from feat(aio): switch batch trace summarization to gpt-5-nano on flex #91639: flex allowlist over family prefix, retry × timeout accounting against the enclosing activity budget, one test row per except-tuple member.
  • The per-call standard fallback lands on review feedback (retries belong on the non-flex tier). A whole-agent standard rerun was built first and replaced: with tens of calls per run, a failure at call N discarded and rebilled N−1 successful flex calls.
  • Skills invoked: /writing-pr-descriptions, /writing-tests, /writing-code-comments.
  • History note for reviewers: earlier revisions carried a standard-tier rerun with state isolation and partial-label salvage, removed as over-engineering; the per-call fallback is the narrow replacement. The --max-cached-workflows worker flag also lived here for a while and moves to a dedicated worker-tuning PR.

The labeling agents are daily batch jobs, so they tolerate flex-tier
queueing for tokens billed at half the standard rate. Only gpt-5 family
models get the field; gpt-4.1 rejects it. The SDK's existing retries
cover flex refusals and gateway timeouts.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@bernatixer bernatixer self-assigned this Sep 1, 2026
@trunk-io

trunk-io Bot commented Sep 1, 2026 •

Copy link
Copy Markdown

😎 Merged successfully - details.

@github-actions

github-actions Bot commented Sep 1, 2026 •

Copy link
Copy Markdown
Contributor

🤖 CI report

⚠️ Trunk lane — backend Python lane

This PR is assigned to the backend Python lane. It runs backend Python tests and may merge in parallel with PRs in other lanes.

⚠️ Playwright — 1 flaky

🎭 Playwright report · View test results →

⚠️ 1 flaky test:

  • Duplicating a dashboard preserves text cards, date filter, and variables (chromium)

These issues are not necessarily caused by your changes.
Annoyed by this section? Help fix flakies and failures and it will go green!

✅ ClickHouse migration SQL — none

No ClickHouse migrations in the latest push.

bernatixer and others added 10 commits September 1, 2026 10:58
Each cached workflow holds its full history, payloads included, so
queues running short-lived, high-volume workflows need a bound far
below the SDK default of 1000. Settable per deploy via the
--max-cached-workflows flag or MAX_CACHED_WORKFLOWS.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…r rerun

OpenAI decides flex eligibility per model (pro tiers excluded), so the
family-prefix check becomes an explicit allowlist. Flex clients get a
short per-call timeout with SDK retries off, and a shared runner reruns
the agent on the standard tier when flex fails, since the labeling
callers otherwise turn a flex outage into silent default labels. Logs
the served tier and fallbacks.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Client construction moves outside the callers' catch-all so config
errors fail the activity instead of shipping default labels, each
attempt gets a fresh state copy, the standard rerun keeps the short
per-call timeout, LLMA_LABELING_FLEX_ENABLED gives a no-deploy kill
switch, and the worker test no longer binds a real Prometheus socket.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
A run where both tiers fail now returns the last attempt's partial
labels instead of discarding them. Standard-tier labeling calls are
capped at 240s with one SDK retry so two attempts fit the activity
budget, including when the flex kill switch is pulled. Drops the
kwargs dict spread and updates both READMEs.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
… labels

A failed run raises LabelingAgentError carrying the fullest label set
any attempt wrote; the graphs log the real cause and keep those labels
before filling defaults, so failures no longer log as completed runs.
Also corrects the rerun timeout in both READMEs, adds team_id and
status_code to the labeling logs, and annotates make_agent.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…itch

Drops the standard-tier rerun machinery: the SDK's own retries recover
transient flex refusals, a still-failing run falls to the existing
default-label handling and the next day's batch, and the env kill
switch covers a sustained outage. Graphs and their tests return to
master untouched.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The worker sticky-cache flag moves to a dedicated worker-tuning PR
covering the kubernetes and temporal settings together.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…on and labeling

posthog/llm/flex.py mirrors OpenAI's flex pricing table; both callers
import it instead of keeping product-local copies in sync.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
… kill switch

Flex is an OpenAI-specific tier, so the module name says so. Rolling
labeling flex back is a one-line revert plus a deploy, matching the
summarization rollout's trade-off, so the env lever goes.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@trunk-io

trunk-io Bot commented Sep 3, 2026 •

Copy link
Copy Markdown

Static Badge   Static Badge   Static Badge

View Full Report ↗︎ ⋅ Docs

@bernatixer
bernatixer marked this pull request as ready for review September 3, 2026 12:09
@pr-assigner-resolver-posthog
pr-assigner-resolver-posthog Bot requested a review from a team September 3, 2026 12:10

@Radu-Raicea Radu-Raicea left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Let's do the retries on the non-flex version of the model, because it might be unavailable more often on flex than on non-flex.

The labeling agent also makes multiple multiple LLM calls, so if they all get significantly slower, I wonder if 10 min is still a good timeout.

bernatixer and others added 4 commits September 4, 2026 10:07
… tier

A recoverable flex failure (429, 408, 5xx, connection) re-issues only the
failing chat completion with service_tier=default, so the agent keeps its
completed turns. Flex clients drop SDK retries; the tier switch is the retry.
The recoverability predicate moves to posthog/llm/openai_flex.py so both
flex callers share it.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The first fallback latches the client to standard, so a flex brownout
costs one 120s timeout per agent run instead of one per call. The
recoverability predicate regains the SDK's 409 retry, the fallback log
drops the response body, and the evaluation child workflow gets the
45min execution timeout its activity budgets already need.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Every request path flows through _get_request_payload, so the tier
override lives once there; the fallback timeout field goes away since
standard calls fit the flex client's 120s budget.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@trunk-io
trunk-io Bot merged commit f65a9b6 into master Sep 7, 2026
237 checks passed
@trunk-io
trunk-io Bot deleted the feat/aio-clustering-labeling-flex branch September 7, 2026 09:14
@deployment-status-posthog

deployment-status-posthog Bot commented Sep 7, 2026 •

Copy link
Copy Markdown

Deploy status

Environment Status Deployed At Workflow
dev ✅ Deployed 2026-09-07 09:47 UTC Run
prod-us ✅ Deployed 2026-09-07 10:04 UTC Run
prod-eu ✅ Deployed 2026-09-07 10:07 UTC Run

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants