Skip to content

mlpstorage submit: check, package, record, upload instructions; status Next -> submit --dry-run (status-and-submit PR 3) - #881

Merged
FileSystemGuy merged 2 commits into
mainfrom
status-submit-pr3-submit-verb
Sep 23, 2026
Merged

FileSystemGuy merged 2 commits into
mainfrom
status-submit-pr3-submit-verb

Conversation

@FileSystemGuy

Copy link
Copy Markdown
Contributor

Third PR of the status-and-submit series (design agreed 2026-09-23; #879 = vocabulary + readiness evaluator, #880 = mlpstorage status). Adds the submitter's verb: mlpstorage submit [--dry-run] [--out PATH].

What it does

submit runs over the submitter's own initialized results-dir and does four things in order:

  1. Regenerates the rollups — reports reportgen in-process (per-model and per-org results.{csv,json}, each org's submission.yaml). Happens in --dry-run too, because 2.1.16 / 2.1.22 / RPT-01 / PROV-02 can only be judged on fresh files. reportgen's own stdout report is captured to debug so the status table is the first thing printed; its warnings still reach stderr.
  2. Runs the checker once — the same checker behind mlpstorage validate (readiness.collect_findings), scored per result by readiness.evaluate(findings=...). Never a second implementation.
  3. Refuses with exit 1 while any error remains: the status table (rows, paperwork footer, tree problems), plus a "Rollup tables (regenerated, still failing)" list for rule errors the table has no row for, then Not submittable: 1 result short, 1 needs paperwork; 9 checker errors in all. and Next: fix the above, then mlpstorage submit --dry-run. Warnings never block. whatif results are never packaged and nothing the checker says about them counts. No --force.
  4. On a pass, writes the package, a .sha256 (sha256sum format) and a .manifest.json beside it, appends one packaged line to <results-dir>/.mlps/submissions.jsonl, and prints the manual upload instructions (submit.UPLOAD_INSTRUCTIONS, the plug-in point for a future uploader; wording follows Submission_guidelines §11: MLCommons submission UI, each upload replaces the last, upload every day or two).

--dry-run stops after step 3 and prints what it would write:

closed/Acme   rules edition 3.0   1 result, 6 runs

SYSTEM  BENCHMARK  MODEL   ACCEL  RUNS  SUBMIT  NOTE
sys-1   training   unet3d  b200   6/6   ready

1 of 1 result ready.
Would write /home/me/results/.mlps/packages/Acme-3.0-20260923_161504.tar.gz
  closed/Acme: 1 system, 1 result, 6 runs; code-images: 1 image; 40 files, 91.2 KiB
Dry run: no package written. `mlpstorage submit` builds it.

The package

  • Root directory is <org>/ (Rules.md 2.1.1); inside it closed/<org>/** and open/<org>/** exactly as they stand, and code-images/ holding the pool marker (copied from the tree's pool, synthesized if none) plus only the images the packaged leaves point at — every .mlps-code-image under the packaged org dirs, including datagen/datasize leaves. Pointers resolve against the tree-wide pool first and the v3.0 per-org pool second (2.1.6); an image from the per-org pool is placed under code-images/ in the package.
  • Excluded: whatif/, .mlps/, mlperf-results.yaml, and pool images nothing packaged points at (a whatif-only image or an orphan). The package is what 2.1.2 names and nothing else, so validate on the extracted tarball sees no orphans and no unexpected top-level entries.
  • Default path <results-dir>/.mlps/packages/<org>-<edition>-<YYYYMMDD_HHMMSS>.tar.gz; --out PATH is the tarball when it ends in .tar.gz/.tgz, a directory otherwise. Written as .part and renamed into place.
  • Manifest schema mlps-submission-package/1: org, edition, created_at/by, source results-dir, package name, sha256, size, file count, divisions, systems per division, code images, every result (division/system/benchmark/model/accelerator, run count, ledger run IDs, leaves), and every file with its size and sha256.
  • Ledger line: {"event":"packaged","id":n,"at","package","sha256","size_bytes","file_count","orgname","rules_edition","divisions","results","runs","tool"}; IDs count up per tree like runs.jsonl.

Other changes

  • readiness.SubmissionReadiness.errors — every non-whatif checker error, whether or not the table carries it; submittable now also requires it empty; status --json gains checker_errors. evaluate() accepts pre-collected findings so submit runs the checker once.
  • status "Next:" line when everything is ready is now mlpstorage submit --dry-run (the tree-problems case still points at validate, which lists them in full).
  • Result NOTE collapses identical workload-level findings the checker repeats per run ([3.1.1] dataset parameters not found (x6) instead of six copies).
  • --help_all: SYNOPSIS line, tree entry, SUBMIT block, root context tokens; parity map gains ('submit',). README utility table and CLAUDE.md gain the command.
  • submit is recorded in .mlps/history (it writes to the tree); status stays unrecorded.

Tests

tests/unit/test_submit_command.py — 34 tests, RED first (2eb6c3f): parser and help; refusal on short/paperwork, on a rollup-rule error the table cannot show, on a whatif-only or empty tree, on reportgen failure; warnings and whatif errors never block; reportgen runs before the checker; dry-run writes nothing; tarball layout, exclusions, legacy per-org pool image placement, open division; checksum and manifest contents against the tarball; ledger IDs; --out as directory and as file; output lines and upload text; history recording; evaluate(findings=); errors filtering; NOTE dedupe; status Next line; results-dir gates; and one unstubbed run over the definition-of-done fixture (real reportgen + real checker → refused, submission.yaml written, no package). test_status_command.py / test_help_behavior.py updated for the Next line and the root token list.

All four CI suites green locally. The checker itself is untouched (no validator-diff run).

Calls worth a look

  • Rollups are regenerated on --dry-run too, i.e. a dry run writes results.{csv,json}, submission.yaml and (first time) public_ids.json into the tree. Alternative was a dry run that skips reportgen and therefore cannot judge the rollup rules; I chose the real check.
  • The sentinel is not packaged. mlperf-results.yaml is a results-dir artifact, not a 2.1.2 entry; the manifest carries the org/edition/layout instead.
  • Only referenced images are packaged, rather than the whole pool, so a whatif-only or orphan image cannot make the package fail CHECK-03.
  • No --force, no --json on submit (per the design; status --json is the machine-readable side).
  • PR 4 (ManPage.md) documents all of this; nothing in Rules.md changes here.

…m, ledger), status Next line (RED; status-and-submit PR 3)
…/submissions.jsonl ledger, upload instructions; status Next -> submit --dry-run (status-and-submit PR 3)
@FileSystemGuy
FileSystemGuy requested a review from a team September 23, 2026 21:53
@github-actions

Copy link
Copy Markdown

MLCommons CLA bot All contributors have signed the MLCommons CLA ✍️ ✅

@FileSystemGuy
FileSystemGuy merged commit 31885d0 into main Sep 23, 2026
4 checks passed
@FileSystemGuy
FileSystemGuy deleted the status-submit-pr3-submit-verb branch September 23, 2026 21:57
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant