Conversation
Signed-off-by: Alex Steiner <asteiner@nvidia.com>
Signed-off-by: Alex Steiner <asteiner@nvidia.com>
979c498 to
c6f2dfc
Compare
|
I reviewed the latest rebased head. The fresh-evidence and same-category confirmation design makes sense, and the transcript-only Terminus deduplication is a good fit for the observed false positive. A few comments before merge:
Overall, the #637 direction is sound. The two build errors and logging level are the actionable items I found. |
Summary
Benchmark evidence
On 20 RF and 20 TW tasks, the patched router scored 15/40 versus 10/40 for the original router and reduced switches from 9 to 5. Under the documented cached-token-aware synthetic rates, estimated cost fell from $1.171/task to $0.889/task. Direct GLM also scored 15/40 at $0.501/task, so this PR claims a false-positive fix and improvement over the original escalation policy, not superiority over the best fixed model.
Full methodology, task-level observations, formulas, limitations, and validation evidence are in benchmark/SWE_ATLAS_ESCALATION_REPORT.md.
Validation
All passed on Rust 1.96.1. Pytest reported 115 passed, 2 deselected, and 2 subtests passed.