Skip to content

feat(ai-rate-limiting): expose nested usage fields to cost_expr - #13984

Merged
shreemaan-abhishek merged 4 commits into
apache:masterfrom
shreemaan-abhishek:feat/cost-expr-nested-usage
Sep 28, 2026
Merged

shreemaan-abhishek merged 4 commits into
apache:masterfrom
shreemaan-abhishek:feat/cost-expr-nested-usage

Conversation

@shreemaan-abhishek

@shreemaan-abhishek shreemaan-abhishek commented Sep 23, 2026 •

Copy link
Copy Markdown
Contributor

Description

Some LLM APIs report parts of usage in nested objects. The OpenAI Responses API puts cached_tokens under input_tokens_details and reasoning_tokens under output_tokens_details:

"usage": {
  "input_tokens": 17,
  "input_tokens_details": { "cache_write_tokens": 0, "cached_tokens": 0 },
  "output_tokens": 726,
  "output_tokens_details": { "reasoning_tokens": 39 },
  "total_tokens": 743
}

cost_expr only saw top-level numbers, so these fields read as 0 and could not be priced.

This PR exposes nested numeric usage fields to cost_expr by joining the parent and child keys with __:

(input_tokens * 0.38
 + input_tokens_details__cached_tokens * 0.04
 + input_tokens_details__cache_write_tokens * 0.48
 + output_tokens * 1.71) / 10000 * 100
  • Top-level fields keep their plain names, so existing expressions behave the same.
  • Deeper fields join every level, e.g. prompt_tokens_details__cached_tokens_details__text_tokens.
  • A nested field is only available under its joined name. The same leaf name in sibling objects stays distinct: OpenAI Chat Completions' audio_tokens is prompt_tokens_details__audio_tokens or completion_tokens_details__audio_tokens.
  • __ is used because usage keys already contain single underscores, so the parent/child boundary stays visible.
  • Arrays are skipped. Missing fields still default to 0, and the schema is unchanged.

Tests: TEST 14-28 in t/plugin/ai-rate-limiting-expression.t, with new fixtures openai/responses-usage-clash.json, openai/chat-usage-audio.json and openai/chat-usage-deep.json. They cover joined names for OpenAI Responses (streaming and non-streaming), nested fields not bound by their bare name, a top-level field next to a nested field with the same name, sibling fields with the same leaf name, missing joined names, and two-level nesting with arrays skipped. Docs (en, zh) updated.

Which issue(s) this PR fixes:

N/A

Checklist

  • I have explained the need for this PR and the problem it solves
  • I have explained the changes or the new features added to this PR
  • I have added tests corresponding to this change
  • I have updated the documentation to reflect this change
  • I have verified that this change is backward compatible (If not, please discuss on the APISIX mailing list first)

Comment thread apisix/plugins/ai-rate-limiting.lua Outdated
end


-- expose usage leaves by bare name, breadth first so shallower keys win

@nic-6443 nic-6443 Sep 24, 2026 •

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Explicit usage similar to input_tokens_details.cached_tokens needs to be supported to avoid ambiguity when duplicate field names exist in nested data.

@membphis membphis left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

@shreemaan-abhishek
shreemaan-abhishek merged commit c6b2adc into apache:master Sep 28, 2026
22 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants