mlpstorage submit: check, package, record, upload instructions; status Next -> submit --dry-run (status-and-submit PR 3) - #881
Merged
Conversation
…m, ledger), status Next line (RED; status-and-submit PR 3)
…/submissions.jsonl ledger, upload instructions; status Next -> submit --dry-run (status-and-submit PR 3)
|
MLCommons CLA bot All contributors have signed the MLCommons CLA ✍️ ✅ |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Third PR of the status-and-submit series (design agreed 2026-09-23; #879 = vocabulary + readiness evaluator, #880 =
mlpstorage status). Adds the submitter's verb:mlpstorage submit [--dry-run] [--out PATH].What it does
submitruns over the submitter's own initialized results-dir and does four things in order:reports reportgenin-process (per-model and per-orgresults.{csv,json}, each org'ssubmission.yaml). Happens in--dry-runtoo, because 2.1.16 / 2.1.22 / RPT-01 / PROV-02 can only be judged on fresh files. reportgen's own stdout report is captured to debug so thestatustable is the first thing printed; its warnings still reach stderr.mlpstorage validate(readiness.collect_findings), scored per result byreadiness.evaluate(findings=...). Never a second implementation.statustable (rows, paperwork footer, tree problems), plus a "Rollup tables (regenerated, still failing)" list for rule errors the table has no row for, thenNot submittable: 1 result short, 1 needs paperwork; 9 checker errors in all.andNext: fix the above, then mlpstorage submit --dry-run. Warnings never block.whatifresults are never packaged and nothing the checker says about them counts. No--force..sha256(sha256sum format) and a.manifest.jsonbeside it, appends onepackagedline to<results-dir>/.mlps/submissions.jsonl, and prints the manual upload instructions (submit.UPLOAD_INSTRUCTIONS, the plug-in point for a future uploader; wording follows Submission_guidelines §11: MLCommons submission UI, each upload replaces the last, upload every day or two).--dry-runstops after step 3 and prints what it would write:The package
<org>/(Rules.md 2.1.1); inside itclosed/<org>/**andopen/<org>/**exactly as they stand, andcode-images/holding the pool marker (copied from the tree's pool, synthesized if none) plus only the images the packaged leaves point at — every.mlps-code-imageunder the packaged org dirs, including datagen/datasize leaves. Pointers resolve against the tree-wide pool first and the v3.0 per-org pool second (2.1.6); an image from the per-org pool is placed undercode-images/in the package.whatif/,.mlps/,mlperf-results.yaml, and pool images nothing packaged points at (a whatif-only image or an orphan). The package is what 2.1.2 names and nothing else, sovalidateon the extracted tarball sees no orphans and no unexpected top-level entries.<results-dir>/.mlps/packages/<org>-<edition>-<YYYYMMDD_HHMMSS>.tar.gz;--out PATHis the tarball when it ends in.tar.gz/.tgz, a directory otherwise. Written as.partand renamed into place.mlps-submission-package/1: org, edition, created_at/by, source results-dir, package name, sha256, size, file count, divisions, systems per division, code images, every result (division/system/benchmark/model/accelerator, run count, ledger run IDs, leaves), and every file with its size and sha256.{"event":"packaged","id":n,"at","package","sha256","size_bytes","file_count","orgname","rules_edition","divisions","results","runs","tool"}; IDs count up per tree likeruns.jsonl.Other changes
readiness.SubmissionReadiness.errors— every non-whatif checker error, whether or not the table carries it;submittablenow also requires it empty;status --jsongainschecker_errors.evaluate()accepts pre-collected findings sosubmitruns the checker once.status"Next:" line when everything is ready is nowmlpstorage submit --dry-run(the tree-problems case still points atvalidate, which lists them in full).[3.1.1] dataset parameters not found (x6)instead of six copies).--help_all: SYNOPSIS line, tree entry,SUBMITblock, root context tokens; parity map gains('submit',). README utility table and CLAUDE.md gain the command.submitis recorded in.mlps/history(it writes to the tree);statusstays unrecorded.Tests
tests/unit/test_submit_command.py— 34 tests, RED first (2eb6c3f): parser and help; refusal on short/paperwork, on a rollup-rule error the table cannot show, on a whatif-only or empty tree, on reportgen failure; warnings and whatif errors never block; reportgen runs before the checker; dry-run writes nothing; tarball layout, exclusions, legacy per-org pool image placement, open division; checksum and manifest contents against the tarball; ledger IDs;--outas directory and as file; output lines and upload text; history recording;evaluate(findings=);errorsfiltering; NOTE dedupe; status Next line; results-dir gates; and one unstubbed run over the definition-of-done fixture (real reportgen + real checker → refused,submission.yamlwritten, no package).test_status_command.py/test_help_behavior.pyupdated for the Next line and the root token list.All four CI suites green locally. The checker itself is untouched (no validator-diff run).
Calls worth a look
--dry-runtoo, i.e. a dry run writesresults.{csv,json},submission.yamland (first time)public_ids.jsoninto the tree. Alternative was a dry run that skips reportgen and therefore cannot judge the rollup rules; I chose the real check.mlperf-results.yamlis a results-dir artifact, not a 2.1.2 entry; the manifest carries the org/edition/layout instead.--force, no--jsononsubmit(per the design;status --jsonis the machine-readable side).