Skip to content

refactor(hermes): remove standalone sandboxed agent - #3966

Open
yaoyu-33 wants to merge 3 commits into
NVIDIA-NeMo:mainfrom
yaoyu-33:codex/remove-hermes-sandboxed-agent
Open

yaoyu-33 wants to merge 3 commits into
NVIDIA-NeMo:mainfrom
yaoyu-33:codex/remove-hermes-sandboxed-agent

Conversation

@yaoyu-33

@yaoyu-33 yaoyu-33 commented Oct 3, 2026 •

Copy link
Copy Markdown
Contributor

Hermes now runs in task sandboxes through EnvironmentServer sessions (#3720, hardened in #3554). Remove the separate hermes_sandboxed_agent implementation introduced by #3542, including its runner, runtime preparation scripts, configs, and adapter-specific tests. SWE-Pro advertises the retained hermes_agent, and its launch documentation points to the existing native recipe.

This is a focused follow-up to those merged PRs; a separate issue is unnecessary.

Migration and compatibility

  • This removes the old component/config paths; it is not an alias or a transparent /run replacement. Use benchmarks/swebench/pro/hermes.yaml and route collection through its EnvironmentServer. Existing prepared rows remain usable through single_agent_turn_legacy.
  • The agent-integration docs provide a shared migration workflow for sandboxed harnesses: replacement capabilities, EnvironmentServer routing, sandbox ownership, configuration, and rollout validation. Hermes-specific obsolete keys, budgets, reporting fields, Apptainer settings, and runtime pins are documented in the Hermes README. Only the Hermes adapter is removed by this PR. The standalone /opt/hermes portable glibc/musl packaging is removed; existing offline and mixed-libc deployments need runtime preparation and validation before switching.
  • Shared sandbox providers, SWE-Pro verification, and the unified Hermes implementation are unchanged. Scores and failure coverage should be re-baselined after switching adapters because their runtime/settings/outcome contracts differ.

Validation

Python 3.13.14 on Linux, using existing component dependencies with package namespaces explicitly bound to the isolated source checkout: 447 passed.

python -m pytest responses_api_agents/hermes_agent/tests \
  resources_servers/swebench_pro/tests \
  environment_servers/single_agent_turn/tests \
  environment_servers/single_agent_turn_legacy/tests \
  tests/unit_tests/test_agent_sessions.py \
  tests/unit_tests/test_agent_session_retries.py \
  tests/unit_tests/test_apptainer_provider.py \
  tests/unit_tests/test_hermes_reasoning_replay.py \
  --import-mode=importlib -q -o addopts= -o log_cli=false --tb=short -rs
  • python -m pytest tests/unit_tests/test_fern_docs_links.py -q -o addopts='' -o log_cli=false: 18 passed, 17 subtests passed on macOS.
  • python -m pre_commit run --all-files and git diff --check: passed. Commit is DCO signed off.
  • cd fern && npm run check: 0 errors, 1 warning. The warning is that the remote missing-redirect check is skipped without Fern authentication.
  • Real-model HTTP smoke: EnvironmentServer → Hermes → OpenSandbox → SWE-Pro, using nvidia/nemotron-3-super-120b-a12b. Three model calls; inspected terminal actions locating and reading the target source. The 15-second runner limit yielded incomplete / stop_reason=wall_time, no patch, reward 0, mask_sample=false, and evaluation_completed=true, with 18 verifier outcomes (15 passed, 3 failed). Agent close preceded verification. Both sessions closed; no owned containers remained, and the temporary network was removed. This is lifecycle evidence, not an accuracy baseline. Detailed artifacts are retained privately.

Kept as a draft for migration review. No fresh Apptainer, Alpine/musl, offline-runtime, full-benchmark, full-repository, or per-server clean-venv validation is claimed; external Slurm deployment configs are not changed here.

Signed-off-by: yaoyu-33 <yaoyu.094@gmail.com>
@yaoyu-33 yaoyu-33 added the area:agent Agent harnesses and Responses API agent behavior label Oct 3, 2026
@copy-pr-bot

copy-pr-bot Bot commented Oct 3, 2026

Copy link
Copy Markdown

This pull request requires additional validation before any workflows can run on NVIDIA's runners.

Pull request vetters can view their responsibilities here.

Contributors can view more details about this message here.

Signed-off-by: yaoyu-33 <yaoyu.094@gmail.com>
Signed-off-by: yaoyu-33 <yaoyu.094@gmail.com>
@yaoyu-33
yaoyu-33 marked this pull request as ready for review October 3, 2026 20:40
@yaoyu-33 yaoyu-33 added complexity:medium Single-domain change with interacting parts or a moderate review surface feature New capabilities, enhancements, or enablement work needs-review PR is ready for code review and waiting on a reviewer labels Oct 3, 2026

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

area:agent Agent harnesses and Responses API agent behavior complexity:medium Single-domain change with interacting parts or a moderate review surface feature New capabilities, enhancements, or enablement work needs-review PR is ready for code review and waiting on a reviewer

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant