Skip to content

✨ feat: Add the Classification Port - #561

Open
danny-avila wants to merge 12 commits into
mainfrom
feat/classification-port
Open

danny-avila wants to merge 12 commits into
mainfrom
feat/classification-port

Conversation

@danny-avila

@danny-avila danny-avila commented Sep 24, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

Keep Classifier as a small SDK contract rather than a chat-model subclass. Jev and self-hosted Laya use the same System One HTTP adapter; an already configured chat model can be injected through a separate strict structured-output adapter. Laya is a data-only preset with an operator-supplied endpoint, optional bearer authentication, and no forced checkpoint.

Semantics and safety

  • Measured System One booleans have probability: number. A chat-only boolean has { decision: boolean, probability: null }; choice distributions and token usage are null when unmeasured. Missing HTTP answers are typed as undefined, while malformed or unexpected answers fail explicitly.
  • Scores remain System One expected rubric values. The chat adapter rejects score questions rather than inventing expected values. Jev and Laya have different confidence definitions, so thresholds cannot be transferred without evaluating the actual checkpoint and use case.
  • Strict JSON-schema or strict tool-calling mode is explicit; unsupported modes fail without prompting for JSON. Compatible questions share a provider-limited invocation. Provider responses are validated locally even when the model's parser accepts them.
  • A single HTTP deadline covers request preparation, credential minting, fetch, bounded response reading, retries, and backoff. Bearer-key redirects fail closed. Non-2xx response bodies and request credentials are not reflected in SDK errors or hooks. Concurrent calls retain independent signals and credential refreshes.

Verification and review

Current pushed head: f38061e8e1de5eab8c51b9d90ab1b647946d8770, incorporating main 64177c706f69d08a38565525437f896e8cfd0b14 and package version 4.0.0.

Verified locally on this head:

Check Result
npx jest src/classification langfuse deterministic-trace-id activity-label-trace-masking --runInBand --silent 291 passed across 14 focused suites
npx tsc --noEmit -p tsconfig.json --pretty false Passed
Zero-warning touched-file ESLint, import order, Prettier, and git diff --check Passed
Package build, circular dependencies, ESM/CJS root classification exports Passed

CI for this exact head passed all 13 validation jobs. Independent review of f38061e8e1de5eab8c51b9d90ab1b647946d8770 completed with no new findings. The reviewer verified 57 native ledger assertions and focused transport invariants; it did not rerun Jest or live provider pipelines in its isolated lane. The parent task's 291 focused tests exercised the real locked LangChain pipelines with controlled provider responses. Earlier CI/reviews are not used as coverage of this head.

Final finding ledger:

Finding Severity Disposition
4128694610: choice contradicts measured probabilities P2 Fixed in 7411317e; final review confirmed
4128694615: malformed boolean criteria reach transport P2 Fixed in 7411317e; final review confirmed
IR-1: provider adapter silently ignores strict mode P2 Fixed in f38061e8; final review confirmed
IR-2: score parser accepts invalid keys without question metadata P2 Fixed in f38061e8; final review confirmed

The earlier eight inline fixes remain present and were independently rechecked. All ten GitHub review threads are resolved. No finding was rejected. This PR is ready to merge; no merge or publication has been performed.

The last regression subset reproduced 11 failures before the current fixes. Independent review found two P2 defects at 7411317e: Bedrock's locked adapter silently ignored strict mode, and exported score parsers accepted nonnumeric keys without expected question metadata. Both are fixed in f38061e8 with regressions. The verified strict adapters are OpenAI (including Azure's inherited implementation) in either mode and Anthropic in strict tool-calling mode; unverified adapters fail locally with unsupported_mode. Score distributions always require canonical consecutive numeric levels starting at zero, independent of an optional expected question.

Self-review follow-up

Four inline review findings from e69549331079ccdd0416bd7060110cc8221ebce4 were reproduced and fixed in d46ae34835e4c73d05aafe85fdc93c14a4088938: presets and their registry are immutable across tenants; per-call bearer minting and the one 401 refresh persist over retries; marked structured-chat prompts keep the original provider request but preserve Langfuse tool-output redaction before trace export; measured choice/score distributions are complete and normalized within a bounded tolerance. Independent review also moved strict-chat question preparation inside the request deadline, so pre-aborted calls never inspect dynamic questions. New regressions cover all five issues.

Further invariant review found two gaps fixed in f88615413542161cd8db44dd7bf406bff3f47b0d. A private tool result copied into free-form classifier state or instructions has no tool identity after stringification, so selective field redaction could export it. The trace processor now drops the whole marked classifier prompt under any active tool-output redaction policy (provider requests are unchanged); a regression first reproduced the leak. A monotonic deadline can also expire before the timer callback runs, returning while fetch stays active. The deadline now aborts its controller on expiry, honors expiry during synchronous preparation errors, and observes late promise rejections. Regressions reproduced both paths before the fix.

The latest Codex review on 0ac0cf9d identified four further correctness gaps. Commits 953216cb and 5fccb734 address them: a measured score must match its rounded distribution; structured-chat failures preserve sanitized HTTP categories and Retry-After; Bedrock cache token counts are included exactly once; and HTTP/chat validation use per-call snapshots of the questions sent. New regressions exercise real OpenAI and Anthropic failure paths, Bedrock usage metadata, and delayed responses during caller mutation.

The two remaining Codex findings on 5fccb734 are addressed in 7411317e: measured choices must select a maximum-probability option (ties are valid), and malformed boolean criteria fail locally before credential minting or HTTP/chat provider invocation. Regressions cover both wire dialects, missing choices, ties, unmeasured choices, rejected criteria, and supported string/one-sided/structured-text criteria. Choice consistency is checked during the existing probability validation pass.

Rollout

Merge and publish this shared SDK separately. Migrate the duplicate LibreChat port in #16180, then adjust the probability consumers and fallbacks in #16181 separately. Live Jev/Laya comparisons, calibration, latency/cost evaluation, live Langfuse verification, and LibreChat consumer tests have not been run in this agents PR. No merge or package publication has been performed.

danny-avila and others added 3 commits September 24, 2026 09:46
A typed question in, a calibrated answer out: `src/classification/` carries the port that
LibreChat PR #16180 introduced under `packages/api` and that codegraph mirrors in ESM, so the
product, the graph and any other consumer share one implementation of the contract a System One
host (TypeSafe's Jev, directly or through a gateway) answers.

- types: boolean / choice / score questions, answers with a probability or a calibrated
  confidence and distribution, `Classifier`, `ClassificationError` with typed failures,
  `ClassificationDialect`, `ClassificationProviderSettings`
- dialect: boolean ↔ `noul`; a string yes-criterion becomes the `{true}` pair a System One host wants
- transport: one deadline for the whole call, bounded retries on 429/5xx/network honouring
  retry-after, an `onAnswered` hook instead of a logger dependency
- http: the host over HTTP, with request/response wrapping for hosts that nest the envelope
- presets: typesafe, openrouter, cloudflare, http; `createClassifier(settings, apiKey)`
- questions: `booleanQuestion`, `choiceQuestion`, `scoreQuestion`
- seven jest tests with a fake fetch; `tsc --noEmit` clean

No LibreChat type is imported: the SDK holds the port, consumers hold their configuration.
…assifier

A host whose bearer expires (the ClickHouse inference gateway mints an hourly Okta token) can
be given a function instead of a key. The transport calls it before each request and once more
with refresh: true after a 401, then retries that request; a 403 is a scope refusal and is not
retried. The `clickhouse` preset points at the gateway's System One route.
@danny-avila

Copy link
Copy Markdown
Collaborator Author

@codex review

@chatgpt-codex-connector

chatgpt-codex-connector Bot commented Sep 27, 2026 •

Copy link
Copy Markdown

Codex Review Summary

This comment shows the latest Codex review activity on this pull request.

Review Status Commit Review trigger
📝 Code Review ✅ Completed 2026-09-29T01:28:16.356760Z 5fccb73 Manual request
ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review" or "@codex security review".

Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: e695493310

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread src/classification/presets.ts
Comment thread src/classification/transport.ts Outdated
Comment thread src/classification/structuredChat.ts Outdated
Comment thread src/classification/dialect.ts
@lia-by-librechat

Copy link
Copy Markdown
Contributor

Self-review handoff for PR #561 at exact pushed head d46ae34835e4c73d05aafe85fdc93c14a4088938.

This head resolves all four inline findings on the earlier head: immutable presets across tenants, cached and refreshed per-call credentials across retries, nested classifier-prompt tool-output redaction before Langfuse export, and complete normalized measured distributions. It also moves structured-chat validation inside the abortable deadline. Tests cover each previously failing case. Local checks passed: 27 classification tests, 204 tracing tests, full workspace TypeScript typecheck, touched-file lint/import order/formatting, circular-dependency check, and package build.

CI for this exact head failed before creating jobs because the current main reusable workflow has duplicate YAML keys. No CI checks ran on this head. A maintainer can trigger a new Codex review for this SHA if desired; a Lia GitHub App comment cannot trigger one.

@lia-by-librechat

Copy link
Copy Markdown
Contributor

Review handoff for draft agents PR #561 at exact remote head f88615413542161cd8db44dd7bf406bff3f47b0d.

Further invariant review found and fixed two issues missed at the preceding head: private tool results copied into free-form classifier state or question text escaped selective Langfuse redaction, and a monotonic timeout could settle before its timer callback fired without aborting the fetch signal. Marked classifier prompts now fail closed in traces under any active tool-output redaction policy. Deadline checks now abort the in-flight signal and observe late rejected tasks, including when synchronous preparation crosses the deadline. No request content or provider behavior is changed by trace redaction.

Local verification on this head: 234 passed tests across 13 focused classification and Langfuse suites, workspace TypeScript typecheck, touched-file lint/import order/formatting, circular-dependency check, ESM/CJS exports, package build, and diff checks. CI for this head failed before creating any jobs; the current main workflow has duplicate YAML keys. A maintainer can trigger a new Codex review for this exact SHA. A Lia GitHub App comment does not initiate that review.

@lia-by-librechat
lia-by-librechat Bot marked this pull request as ready for review September 28, 2026 19:42
@danny-avila

Copy link
Copy Markdown
Collaborator Author

@codex review the latest head

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 0ac0cf9d52

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread src/classification/dialect.ts
Comment thread src/classification/structuredChat.ts Outdated
Comment thread src/classification/structuredChat.ts Outdated
Comment thread src/classification/http.ts
@danny-avila

Copy link
Copy Markdown
Collaborator Author

@codex review the latest head

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 5fccb734ed

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread src/classification/dialect.ts
Comment thread src/classification/questions.ts
@lia-by-librechat

lia-by-librechat Bot commented Sep 30, 2026 •

Copy link
Copy Markdown
Contributor

Review handoff for exact remote head 7411317e3d800de8a11506463c5eb096746e7d06.

This head incorporates main 64177c706f69d08a38565525437f896e8cfd0b14 and addresses the two latest inline findings: measured choices must select a probability maximum, allowing ties, and boolean criteria are validated locally before credential minting or provider invocation. New regressions reproduced both findings before the fixes. All ten existing review threads have recorded fix dispositions and are resolved.

Verified local checks:

  • Classification and tracing: 276 tests passed across 14 focused suites.
  • Workspace npx tsc --noEmit, zero-warning touched-file ESLint, import-order and formatting checks, and git diff --check: passed.
  • Package build, circular dependencies, and ESM/CJS root classification exports: passed.

CI for this exact head passed all 13 validation jobs, including the Anthropic summarization lane. Independent review of this exact head remains in progress. Earlier reviews do not cover this head.

Not run: live Jev/Laya comparisons, calibration, latency/cost evaluation, live Langfuse verification, and LibreChat consumer tests. Nothing merged or published.

@lia-by-librechat

lia-by-librechat Bot commented Sep 30, 2026 •

Copy link
Copy Markdown
Contributor

Final verification for exact remote and independently reviewed head f38061e8e1de5eab8c51b9d90ab1b647946d8770, based on main 64177c706f69d08a38565525437f896e8cfd0b14.

Ready to merge. CI for this head passed all 13 jobs. Fresh independent review completed with no new findings and confirmed all ledger fixes.

Finding Severity Disposition
4128694610: contradictory measured choice P2 Fixed in 7411317e, verified at final head
4128694615: malformed boolean criteria P2 Fixed in 7411317e, verified at final head
IR-1: silent strict-mode downgrade P2 Fixed in f38061e8, verified at final head
IR-2: invalid score keys without expected-question metadata P2 Fixed in f38061e8, verified at final head

The prior eight inline fixes were rechecked and retained. All ten GitHub threads are resolved. No findings were rejected.

Local checks on this exact head passed: 291 tests across 14 focused classification/tracing suites; workspace npx tsc --noEmit; zero-warning touched-file ESLint, import order, formatting, and diff checks; package build; circular dependencies; ESM/CJS root exports. Independent review separately passed 57 native ledger assertions plus focused transport checks. Its isolated lane did not rerun Jest or live provider pipelines.

Strict chat support is explicitly bounded to verified OpenAI modes (including Azure's inherited implementation) and Anthropic strict tool calling. Bedrock and other unverified adapters fail locally instead of silently downgrading. The HTTP Jev/Laya adapter is unchanged.

Not run: live Jev/Laya comparisons, calibration, latency/cost evaluation, live Langfuse verification, and LibreChat consumer tests. Nothing merged or published. Downstream port migration and consumers remain separate PRs.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants