Skip to content

docs(codex): propose native remote-list policy with executable probes - #6157

Draft
luvs01 wants to merge 2 commits into
lidge-jun:devfrom
luvs01:rfc/5848-native-remote-list-policy-20260928
Draft

luvs01 wants to merge 2 commits into
lidge-jun:devfrom
luvs01:rfc/5848-native-remote-list-policy-20260928

Conversation

@luvs01

@luvs01 luvs01 commented Sep 28, 2026

Copy link
Copy Markdown
Collaborator

Summary

Draft RFC and executable specifications only — this is not a production fix, does not change OpenCodex runtime behavior, and does not close #5848.

Related: #5848, duplicate #5906, openai/codex#48358, and the existing #6007/#6070 mitigation.

This PR publishes the investigated alternatives in one isolated research unit, devlog/_plan/260928_remote_thread_provider_policy/, so the next implementation can be reviewed against a concrete contract rather than adding an unverified backend relay to the runtime.

Preferred implementation proposal

When host-side control is needed without changing the mobile app, prefer an operator-opt-in, native Codex app-server policy for remote thread/list requests. The actual Rust configuration, schema, and request-handler plumbing remain to be implemented upstream; the Python code here is an executable specification, not that implementation.

  • Keep the existing default-provider behavior when no operator policy is configured and for every non-remote connection.
  • Use the native server's trusted ConnectionOrigin::RemoteControl, never a client name or caller-supplied field.
  • Preserve every explicit client provider array, including []; preserve the existing parent/ancestor-query exception.
  • A configured nonempty list selects exactly those provider ids; an explicitly empty policy selects all providers. Omission and JSON null follow native typed Option<Vec<String>> semantics.
  • Leave authentication, managed remote-control restrictions, routing, compaction, resume behavior, and native history unchanged. Listing a conversation is not evidence that cross-provider resume is safe.

Alternatives included

  1. Mobile client sends an explicit provider array: still the simplest upstream client correction.
  2. Native remote-only policy: preferred host-side proposal because it adds no extra credential or connection-handling service.
  3. Local backend relay via chatgpt_base_url: a concrete research fallback, with the existing loopback-only mock probe retained. The shared base URL, enrollment identity, token forwarding, multi-segment messages, reconnects, and non-remote backend consumers make it inappropriate to ship as a small default-on workaround. Live ChatGPT/mobile operation is unverified.

The original provider-isolation rationale in openai/codex#5658 is retained. Global omission-to-all changes and history retagging are not proposed. The illustrative configuration name in the design is not an existing supported setting. ADR-5848 and the current warning remain unchanged pending an accepted and released implementation.

Verification

Base: dev at eb7f0f0970c2298f8b2d66d170c4d4be869f301b. Authored head: 1e3a1c94ea0c74f6ca448d1460ba67d187c823bb.

Executed on Linux with Python 3.13.5 and aiohttp 3.13.3:

cd devlog/_plan/260928_remote_thread_provider_policy/probes
python -m unittest -v test_native_policy test_probe test_loopback_bridge

60 tests passed, 0 failures: 22 proposed native-policy contract tests, 29 raw-frame transformation tests, and 9 localhost HTTP/WebSocket mock-relay tests. The synthetic database contains 5,200 openai rows, one opencodex row, and one unrelated-provider row. Fixture filtering/pagination produces 1 / 5,201 / 5,202 results as appropriate, without mutating its dump. This is not a native Codex database or native cursor test.

Also checked:

  • git diff --cached --check: passed for the authored files.
  • Published Git blob hashes match all eight locally tested/authored files.
  • GitHub comparison against the base contains exactly eight added research files in one commit, with no unrelated changes.
  • No GUI, runtime, workflow, dependency manifest, installed configuration, user history, or real credentials are changed. The mock relay accepts only literal-loopback HTTP upstreams and synthetic credentials.

Not run / not established: native Rust implementation/build/tests, actual mobile pairing/list/resume, production multi-chunk/reconnect handling, OpenCodex Bun typecheck/full tests/structure/privacy gates, and independent security review. Bun and a full checkout were unavailable in the execution environment; the checkout attempt failed at DNS resolution. These Python tests do not substitute for those gates, so this PR remains draft. Detailed commands and limitations are in 020_verification.md.

Checklist

  • Scope stays focused and avoids unrelated cleanup.
  • Docs or release notes were updated when needed (research documentation; no release claim).
  • Security-sensitive changes were reviewed for secrets, auth, and unsafe defaults (independent review remains outstanding; no production auth change).

Review readiness

  • Required local validation passed with its scope documented; repository/native gates remain outstanding.
  • Branch starts from the latest dev commit observed at publication.
  • All correct Codex and CodeRabbit findings have been addressed after review.
  • Ready-for-review confirmation for this exact head.

… probes

Publish the lidge-jun#5848 implementation alternatives and executable specifications.
Prefer an opt-in native remote-only list policy over a shared-backend relay.
Keep the existing runtime, authentication, routing and history untouched.

Validation: 60 isolated Python tests passed; native/Bun/mobile validation remains outstanding.
This is an RFC and research unit, not a production fix for lidge-jun#5848.
@coderabbitai

coderabbitai Bot commented Sep 28, 2026

Copy link
Copy Markdown
Contributor

Important

Draft PR not reviewed

Draft PRs are not automatically reviewed by default.

  • Trigger a manual review

To automatically review draft PRs, update your CodeRabbit configuration:

reviews:
  auto_review:
    drafts: true

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@github-actions

Copy link
Copy Markdown
Contributor

✅ Deterministic PR hygiene checks passed.

@github-actions github-actions Bot added the documentation Improvements or additions to documentation label Sep 28, 2026
@lidge-jun

Copy link
Copy Markdown
Owner

리뷰 · 우선순위 36 / 80

이 PR은 연구 문서와 파이썬 시험만 추가합니다. 휴대폰 목록에서 openai 대화가 빠지는 #5848은 그대로입니다. 런타임, 설정, 인증, 대화 기록은 이 커밋에서 바뀌지 않습니다. 파일은 devlog/_plan/260928_remote_thread_provider_policy/ 여덟 개입니다. OpenCodex 실행 경로는 이 폴더를 부르지 않습니다.

다음에 만들 방법으로 적힌 안은 이렇습니다. 휴대폰 앱을 그대로 두고, 서버가 원격 접속이라고 확인한 thread/list에만 운영자가 고른 제공자 목록을 씁니다. 클라이언트가 배열을 보내면 그 배열이 이깁니다. 빈 배열 []은 제공자를 전부 보여 줍니다. 설정을 빼 두면 지금처럼 기본 제공자만 나옵니다. 부모 대화나 조상 대화를 묻는 요청은 지금 서버처럼 기본 필터 없이 갑니다. 글에 나온 thread_list_model_providers는 아직 없는 설정 이름입니다.

옆에 남겨 둔 다른 안은 로컬 중계입니다. 앱 서버로 가기 전에, 빠뜨린 modelProviders만 채웁니다. 시험은 127.0.0.1 목업만 받습니다. 설계는 이 중계를 기본 해법으로 두지 않습니다. chatgpt_base_url이 로그인에 쓰는 주소와 같기 때문입니다. 작성자는 자기 환경에서 파이썬 시험 60개가 통과했다고 적습니다. 이 PR의 CI는 그 시험을 실행하지 않았고, draft라 저장소 게이트 대부분이 건너뛰어졌습니다. 베이스는 dev입니다. 이 설계를 올린 다른 열린 PR은 없습니다. #6007의 경고는 이미 들어가 있고, 이 글은 그 경고를 끄지 않습니다.

라인 - devlog/_plan/260928_remote_thread_provider_policy/probes/native_policy.py resolve_provider_filter 66행. parent_thread_id나 ancestor_thread_id가 None이 아니면 제공자 필터를 없앱니다. 목록은 제공자를 전부 보여 줍니다. 빈 문자열 ""도 그렇게 됩니다. 부모가 있고 조상도 있으면 역시 필터가 사라집니다. 업스트림 thread_list_response_inner(코덱스 1cc7e236, 2574행 근처)는 아이디가 잘못되면 요청 오류를 내고, 부모와 조상을 같이 주면 오류를 냅니다. 이 함수에는 그 검사가 없습니다. 시험은 "fixture-parent"처럼 올바른 문자열만 봅니다. 이 함수를 서버에 그대로 옮기면 잘못된 아이디가 필터를 풉니다.

라인 - devlog/_plan/260928_remote_thread_provider_policy/probes/remote_list_probe.py _patch_message 78행, Policy.providers 22행. 키가 있으면 null이어도 프레임을 그대로 둡니다. 네이티브 명세는 생략과 null을 같게 보고, 원격 접속에 운영자 목록이 있으면 그 목록을 씁니다. 휴대폰이 null을 보내면 이 중계는 openai와 opencodex를 넣지 않습니다. 기본값은 ("openai", "opencodex")입니다. 설계 문서는 이름만으로 같은 제공자라고 단정하지 말라고 적습니다. 두 파일이 고정한 규칙이 서로 다릅니다.

메인테이너의 판단이 필요한 지점

이 PR로 #5848을 닫을지. 닫으면 목록이 고쳐진 것으로 남습니다. 작성자는 고침이 아니라고 적었고 PR은 draft입니다.

다음 구현을 네이티브 서버의 원격 전용 설정으로 갈지, chatgpt_base_url 중계로 갈지. 중계 시험은 목업 루프백만 막습니다.

예시 TOML 키를 지금 설정에 넣을지. 업스트림 앱 서버에는 그 키가 없습니다.

너의 추천

draft로 유지하세요. #5848과 중복 #5906은 열어 두세요. 베이스는 dev입니다. 닫을 중복 PR은 없습니다. resolve_provider_filter에는 업스트림과 같이, 잘못된 스레드 아이디를 거절하고 부모와 조상이 같이 오면 거절하게 넣으세요. 릴레이 프로브는 연구 폴더에만 두세요. thread_list_model_providers는 업스트림이 그 설정을 내보낸 뒤에만 설정 파일에 넣으세요.

이 댓글은 grok-bot이 작성했습니다

@luvs01

luvs01 commented Sep 28, 2026

Copy link
Copy Markdown
Collaborator Author

Applied at a0f933c: resolve_provider_filter now rejects empty/blank related-thread ids and rejects parent_thread_id + ancestor_thread_id together, matching upstream thread_list_response_inner validation instead of letting a malformed id widen the list. Covered by a new spec test (23/23 pass locally; loopback-bridge tests need aiohttp, unavailable here — unchanged code path).

On the relay's null-vs-omission difference: remote_list_probe._patch_message preserving a present JSON null is the documented, deliberate distinction between the raw-frame relay and the typed native proposal (010_design.md), kept as a research contrast rather than a defect — the native policy itself treats null like omission.

Keep-as-draft, keep #5848/#5906 open: agreed, unchanged.

@Ingwannu Ingwannu left a comment

Copy link
Copy Markdown
Owner

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed exact head a0f933c. The probe does not yet match the upstream ThreadId contract claimed by this PR. Upstream parses ThreadId as a UUID; the Python probe rejects only blank strings and its own fixtures accept non-UUID values such as fixture-parent, p, and a. This can certify behavior the upstream implementation would reject.

Please validate UUID syntax in the probe and change the fixtures/negative cases accordingly. Also update the verification counts: the current static inventory is 23 native + 29 relay + 9 loopback = 61 cases, while the documentation/PR text still says 22/60. Hosted CI did not execute these Python probes on this head; most relevant jobs were skipped, so the corrected probes need an actually executed CI path before approval.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

documentation Improvements or additions to documentation

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants