Bounded local-first experiment for the AMD + lablab.ai AI Academy Challenge.
This repository exercises the resident Radeon AI PRO R9700 (gfx1201) + ROCm 10 + vLLM + Qwen3-Coder stack without restarting the serving process and without cloud spend.
The loop measures a baseline, asks Qwen to select one allowed concurrency candidate through a strict JSON action, benchmarks the candidate against the same live vLLM endpoint, and applies an objective KEEP/REJECT gate. Rejected candidates roll back to the baseline configuration. At most two candidate values are tested.
This is experimental evidence for the InnerChispa R9700/Hyperloom research path. It does not claim that upstream Hyperloom officially supports the R9700.
Default bounded workload:
- baseline concurrency: 1
- candidate concurrency: 2 or 4, selected by Qwen
- 6 inference requests per measurement
- 48 max output tokens per request
- KEEP requires at least 3% output-token throughput gain
- candidate p95 latency must remain within 2.5x baseline
- no server lifecycle mutation
- no model download
- no package installation
- no cloud
Artifacts are written to evidence/latest.json and a timestamped evidence file. A completed experiment exits successfully even when all candidates are objectively rejected, because REJECT is a valid optimization result rather than a runtime failure.