From b9a95601b32f46e488fb420feaabb6701a419ca4 Mon Sep 17 00:00:00 2001 From: Mike Sulsenti Date: Sun, 4 Oct 2026 10:32:44 -0400 Subject: [PATCH 1/7] fuzz: initialize the current C decoder options The e2e target stopped building after 90cd9a2 added reserved2 and frame_size_limit to WPDDecoderOptions. Initialize the padding field to zero and forward the Rust options limit so the external-buffer path exercises the same configuration as the Rust driver. Validation: all four cargo-fuzz targets build with ASan and debug assertions on nightly-2026-07-10; scripts/stylecheck.sh passes on the current head. --- fuzz/fuzz_targets/e2e.rs | 2 ++ 1 file changed, 2 insertions(+) diff --git a/fuzz/fuzz_targets/e2e.rs b/fuzz/fuzz_targets/e2e.rs index 9832126..baaebf3 100644 --- a/fuzz/fuzz_targets/e2e.rs +++ b/fuzz/fuzz_targets/e2e.rs @@ -123,6 +123,8 @@ fn decode_external(data: &[u8], options: Options) { flip: i32::from(options.flip), reserved: 0, n_threads: options.n_threads, + reserved2: 0, + frame_size_limit: options.frame_size_limit, }; let mut frame = WPDFrame { struct_size: mem::size_of::(), From 443705cca542a103c7c98a3c2bbbe7b9a9415bf7 Mon Sep 17 00:00:00 2001 From: Mike Sulsenti Date: Sun, 4 Oct 2026 10:38:05 -0400 Subject: [PATCH 2/7] build: name the optional Wuffs dependency Meson 1.12.1 drops empty dependency names before looking up a fallback. The Wuffs dependency introduced in 9748ced therefore makes setup crash with an IndexError, including when Wuffs is disabled. Use the conventional wuffs dependency name and keep the existing pinned fallback. This allows the project to configure with current Meson without changing which optional decoder is built. Validation: meson setup build --wipe -Dtrim_dsp=false -Dtestdata_tests=true -Dlibwebp=subproject -Dwuffs=disabled succeeds with Meson 1.12.1; the same command failed before the change. scripts/stylecheck.sh passes. --- meson.build | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/meson.build b/meson.build index f3eb390..66af0e2 100644 --- a/meson.build +++ b/meson.build @@ -517,7 +517,7 @@ custom_target( message('imagewebpdec: enabled (build explicitly with target imagewebpdec)') wuffs_dep = dependency( - '', + 'wuffs', fallback: ['wuffs', 'wuffs_dep'], required: get_option('wuffs'), ) From 24efec1f5db4f8c3b727e41f017dc136104bcb60 Mon Sep 17 00:00:00 2001 From: Mike Sulsenti Date: Sun, 4 Oct 2026 11:15:07 -0400 Subject: [PATCH 3/7] ci: run decoder checks and seeded fuzz smoke on GitHub Actions Add one workflow for pushes, pull requests and manual runs. Test the full pinned corpus with assembly enabled and disabled, and require checkasm in the assembly job. Run the existing format/clippy, CPU-mask, animation, threading, forced-32-bit range-coder, C/Rust sanitizer and container Miri checks. Build every declared cargo-fuzz target with ASan and debug assertions, then run each for 15 seconds with container and raw VP8/VP8L seeds extracted from the corpus. Keep generated seeds separate from developer corpora and upload failure artifacts. Pin the actions, test corpus, Meson, cargo-fuzz and clang-format; the formatter pin avoids changing existing C macro layout on Ubuntu. Compare assembly and fallback CLIs with the existing md5 and CLI scripts. Keep historical comparisons and performance timing manual, and document the bounded Miri and fuzz coverage in README. Validation: actionlint and shellcheck pass. The style and fallback Meson jobs pass under act on Ubuntu 24.04. Native Meson suites pass (235 asm, 234 fallback), as do full CPU-mask, rac32, animation, thread-count, checksum, CLI, C sanitizer and damaged-input scripts. Rust ASan checks 432 decodes per build; TSan passes 215 unit tests and 720 decodes. Miri passes 17 container tests and the C API storage regression. All four fuzz targets build and complete the seeded smoke with no failures. scripts/stylecheck.sh passes. Record build and CI changes in the draft changelog. Compatibility regression seeds are added with the compatibility-mode fuzzing change. Bound raw VP8/VP8L header geometry and e2e frame allocations to 1,048,576 pixels in the fuzz harnesses. A seeded VP8L smoke run exceeded its RSS budget on a 16384x14336 header; keep malformed header coverage while avoiding allocations outside that smoke budget. This does not change decoder limits. --- .github/workflows/ci.yml | 197 +++++++++++++++++++++++++++++++++ CHANGELOG.md | 13 +++ README.md | 18 +++ fuzz/budget.rs | 13 +++ fuzz/fuzz_targets/container.rs | 1 - fuzz/fuzz_targets/e2e.rs | 7 ++ fuzz/fuzz_targets/vp8.rs | 7 +- fuzz/fuzz_targets/vp8l.rs | 7 +- scripts/fuzz-smoke.sh | 65 +++++++++++ 9 files changed, 325 insertions(+), 3 deletions(-) create mode 100644 .github/workflows/ci.yml create mode 100644 CHANGELOG.md create mode 100644 fuzz/budget.rs create mode 100755 scripts/fuzz-smoke.sh diff --git a/.github/workflows/ci.yml b/.github/workflows/ci.yml new file mode 100644 index 0000000..5076716 --- /dev/null +++ b/.github/workflows/ci.yml @@ -0,0 +1,197 @@ +name: CI + +on: + push: + pull_request: + workflow_dispatch: + +permissions: + contents: read + +concurrency: + group: ci-${{ github.workflow }}-${{ github.ref }} + cancel-in-progress: true + +env: + CARGO_BUILD_JOBS: 3 + CARGO_TERM_COLOR: always + +jobs: + style: + runs-on: ubuntu-24.04 + timeout-minutes: 15 + steps: + - uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1 + with: + persist-credentials: false + - uses: dtolnay/rust-toolchain@7e38f4b43b4db5c8dd498af069a4f6196df1d067 + with: + toolchain: stable + components: rustfmt,clippy + - name: Install tools + run: | + sudo apt-get update + sudo apt-get install -y nasm pipx + pipx install clang-format==23.1.1 + echo "$HOME/.local/bin" >> "$GITHUB_PATH" + - name: Format and lint + run: | + ./scripts/stylecheck.sh + git diff --exit-code + + tests: + runs-on: ubuntu-24.04 + timeout-minutes: 25 + strategy: + fail-fast: false + matrix: + asm: ['true', 'false'] + steps: + - uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1 + with: + persist-credentials: false + - uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1 + with: + repository: halidecx/wpd-test-data + ref: f8c31341db3ab4400f048a96e7b3736fed303b34 + path: wpd-test-data + persist-credentials: false + - uses: dtolnay/rust-toolchain@7e38f4b43b4db5c8dd498af069a4f6196df1d067 + with: + toolchain: stable + - name: Install tools + run: | + sudo apt-get update + sudo apt-get install -y pipx ninja-build nasm cmake + pipx install meson==1.12.1 + echo "$HOME/.local/bin" >> "$GITHUB_PATH" + - name: Tests, C API, parity and checkasm + env: + ASM: ${{ matrix.asm }} + run: | + meson setup build -Denable_asm="$ASM" -Dtrim_dsp=false \ + -Dtestdata_tests=true -Dlibwebp=subproject -Dwuffs=disabled + if [ "$ASM" = true ]; then + meson test -C build --list | grep -q checkasm + fi + meson test -C build --print-errorlogs --num-processes 3 + + scripts: + runs-on: ubuntu-24.04 + timeout-minutes: 45 + steps: + - uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1 + with: + persist-credentials: false + - uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1 + with: + repository: halidecx/wpd-test-data + ref: f8c31341db3ab4400f048a96e7b3736fed303b34 + path: wpd-test-data + persist-credentials: false + - uses: dtolnay/rust-toolchain@7e38f4b43b4db5c8dd498af069a4f6196df1d067 + with: + toolchain: stable + - name: Install tools + run: | + sudo apt-get update + sudo apt-get install -y pipx ninja-build nasm cmake webp + pipx install meson==1.12.1 + echo "$HOME/.local/bin" >> "$GITHUB_PATH" + - name: Build assembly and fallback tools + run: | + meson setup build -Dtrim_dsp=false -Dtestdata_tests=true \ + -Dlibwebp=subproject -Dwuffs=disabled + meson compile -C build -j 3 libwebpdec + meson setup build-noasm -Denable_asm=false \ + -Dlibwebp=disabled -Dwuffs=disabled + meson compile -C build-noasm -j 3 + - name: Existing correctness checks + run: | + ./scripts/testdata.sh + ./scripts/animcheck.sh + ./scripts/threadcheck.sh + ./scripts/md5check.sh ./build-noasm/wpd ./build/wpd + ./scripts/clicheck.sh ./build-noasm/wpd ./build/wpd + ./scripts/rac32.sh --print-errorlogs --num-processes 3 + + c-sanitizers: + runs-on: ubuntu-24.04 + timeout-minutes: 30 + steps: + - uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1 + with: + persist-credentials: false + - uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1 + with: + repository: halidecx/wpd-test-data + ref: f8c31341db3ab4400f048a96e7b3736fed303b34 + path: wpd-test-data + persist-credentials: false + - uses: dtolnay/rust-toolchain@7e38f4b43b4db5c8dd498af069a4f6196df1d067 + with: + toolchain: stable + - name: Install tools + run: | + sudo apt-get update + sudo apt-get install -y pipx ninja-build nasm cmake + pipx install meson==1.12.1 + echo "$HOME/.local/bin" >> "$GITHUB_PATH" + - name: C harnesses and damaged-input smoke + run: | + ./scripts/sanitize.sh --print-errorlogs --num-processes 3 + ./scripts/fuzz.sh 8 wpd-test-data/odd_lossy.webp \ + wpd-test-data/odd_a_lossy.webp wpd-test-data/palette_rgb.webp \ + wpd-test-data/mixed_codecs.webp + + nightly: + runs-on: ubuntu-24.04 + timeout-minutes: 45 + strategy: + fail-fast: false + matrix: + check: [rustsan, tsan, miri, fuzz] + steps: + - uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1 + with: + persist-credentials: false + - uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1 + with: + repository: halidecx/wpd-test-data + ref: f8c31341db3ab4400f048a96e7b3736fed303b34 + path: wpd-test-data + persist-credentials: false + - uses: dtolnay/rust-toolchain@7e38f4b43b4db5c8dd498af069a4f6196df1d067 + with: + toolchain: nightly + components: rust-src,miri + - name: Install tools + run: | + sudo apt-get update + sudo apt-get install -y pipx ninja-build nasm cmake clang + pipx install meson==1.12.1 + echo "$HOME/.local/bin" >> "$GITHUB_PATH" + - name: Nightly check + env: + CHECK: ${{ matrix.check }} + run: | + case "$CHECK" in + rustsan) + meson setup build -Dlibwebp=disabled -Dwuffs=disabled + meson compile -C build -j 3 + ./scripts/rustsan.sh + ;; + tsan) ./scripts/tsan.sh ;; + miri) ./scripts/miri.sh --lib container::tests ;; + fuzz) + cargo install cargo-fuzz --version 0.13.2 --locked + ./scripts/fuzz-smoke.sh + ;; + esac + - name: Save fuzz failures + if: failure() && matrix.check == 'fuzz' + uses: actions/upload-artifact@043fb46d1a93c77aae656e7c1c64a875d1fc6a0a # v7.0.1 + with: + name: fuzz-failures + path: fuzz/artifacts + if-no-files-found: ignore diff --git a/CHANGELOG.md b/CHANGELOG.md new file mode 100644 index 0000000..c83f2b6 --- /dev/null +++ b/CHANGELOG.md @@ -0,0 +1,13 @@ +# Changelog + +## Unreleased — 0.2.0 + +### Build and CI + +- Repair the end-to-end fuzz target's decoder options initializer so all four + coverage-guided targets build again. +- Correct the optional Wuffs dependency name in the Meson build. +- Add one GitHub Actions workflow for tests, rustfmt/clippy, fuzz target builds, + seeded smoke runs, and the existing correctness and sanitizer checks. +- Bound fuzz-harness pictures to one megapixel so mutated dimensions fit the + smoke run's memory budget without changing decoder limits. diff --git a/README.md b/README.md index 4da01c4..a789d16 100644 --- a/README.md +++ b/README.md @@ -90,6 +90,24 @@ Test data is maintained at into the `wpd/` root. `./scripts/testdata.sh` runs end-to-end assembly and fallback checks. +GitHub Actions runs assembly and fallback tests, checkasm, format/lint checks, +the correctness scripts, C and Rust sanitizers, and a container Miri smoke +check. The corpus revision is pinned in `.github/workflows/ci.yml`. CI compares +the assembly and fallback tools with `md5check.sh` and `clicheck.sh`; comparing +an older release still needs an explicit baseline binary. Timing scripts +(`bench.sh` and `cmpbench.sh`) remain manual because shared CI runners do not +provide stable performance measurements. + +`./scripts/fuzz-smoke.sh [seconds-per-target] [corpus-directory]` builds every +fuzz target and runs each for 15 seconds by default. It requires nightly Rust, +Python 3 and cargo-fuzz 0.13.2. It derives container and raw VP8/VP8L seeds from +the test corpus without changing its WebP files, and leaves generated seeds and +failure artifacts under `fuzz/`. Longer fuzzing and the full safe-core Miri +suite (`./scripts/miri.sh`) are useful local checks before releases. The +decoding harnesses bound pictures to 1,048,576 pixels so mutated dimensions fit +the smoke run's memory budget; larger pictures remain in ordinary corpus tests. +This is a harness limit, not a decoder limit. + For libwebp parity testing and benchmarking against alternative WebP decoders, build the optional third-party test binaries: diff --git a/fuzz/budget.rs b/fuzz/budget.rs new file mode 100644 index 0000000..fe1e0c6 --- /dev/null +++ b/fuzz/budget.rs @@ -0,0 +1,13 @@ +pub const MAX_PIXELS: u32 = 1 << 20; + +pub fn fits(data: &[u8]) -> bool { + // Keep malformed headers in coverage; only a successfully read size can + // exceed the harness budget. Raw codecs have no decoder options limit. + wpd::api::info(data).map_or(true, |info| { + wpd::api::Options { + frame_size_limit: MAX_PIXELS, + ..Default::default() + } + .fits(info.width, info.height) + }) +} diff --git a/fuzz/fuzz_targets/container.rs b/fuzz/fuzz_targets/container.rs index ff1825b..85373cf 100644 --- a/fuzz/fuzz_targets/container.rs +++ b/fuzz/fuzz_targets/container.rs @@ -1,4 +1,3 @@ - #![no_main] use libfuzzer_sys::fuzz_target; diff --git a/fuzz/fuzz_targets/e2e.rs b/fuzz/fuzz_targets/e2e.rs index baaebf3..70c0a75 100644 --- a/fuzz/fuzz_targets/e2e.rs +++ b/fuzz/fuzz_targets/e2e.rs @@ -11,6 +11,9 @@ use wpd_capi::decoder::{wpd_decode_into, WPDOutputBuffer}; use wpd_capi::frame::{WPDFrame, WPDOutputPlane}; use wpd_capi::options::WPDDecoderOptions; +#[path = "../budget.rs"] +mod budget; + const FORMATS: [Format; 16] = [ Format::Yuv420p, Format::Yuva420p, @@ -40,6 +43,7 @@ fn decode_options(data: &[u8]) -> (Options, bool) { let flags = byte(data, 1); let subframe = flags & 4 != 0; let mut options = Options { + frame_size_limit: budget::MAX_PIXELS, n_threads: [1, 2, 3, 8][usize::from(flags >> 6)], bypass_filtering: flags & 8 != 0, no_fancy_upsampling: flags & 16 != 0, @@ -156,6 +160,9 @@ fn decode_external(data: &[u8], options: Options) { } fuzz_target!(|data: &[u8]| { + if !budget::fits(data) { + return; + } let Some(&first) = data.first() else { return; }; diff --git a/fuzz/fuzz_targets/vp8.rs b/fuzz/fuzz_targets/vp8.rs index dbe3ec7..f2794f6 100644 --- a/fuzz/fuzz_targets/vp8.rs +++ b/fuzz/fuzz_targets/vp8.rs @@ -1,10 +1,15 @@ - #![no_main] use libfuzzer_sys::fuzz_target; use wpd::vp8::Decoder; +#[path = "../budget.rs"] +mod budget; + fuzz_target!(|data: &[u8]| { + if !budget::fits(data) { + return; + } let mut decoder = Decoder::new(); let _ = decoder.decode_frame(data); diff --git a/fuzz/fuzz_targets/vp8l.rs b/fuzz/fuzz_targets/vp8l.rs index 48367bd..fc91499 100644 --- a/fuzz/fuzz_targets/vp8l.rs +++ b/fuzz/fuzz_targets/vp8l.rs @@ -1,14 +1,19 @@ - #![no_main] use libfuzzer_sys::fuzz_target; use wpd::vp8l::{AlphaDst, Decoder, Target}; +#[path = "../budget.rs"] +mod budget; + fuzz_target!(|data: &[u8]| { if data.len() < 2 { return; } let (head, payload) = data.split_at(2); + if !budget::fits(payload) { + return; + } let mut decoder = Decoder::new(); let alpha_chunk = head[0] & 1 != 0; diff --git a/scripts/fuzz-smoke.sh b/scripts/fuzz-smoke.sh new file mode 100755 index 0000000..9ae87e1 --- /dev/null +++ b/scripts/fuzz-smoke.sh @@ -0,0 +1,65 @@ +#!/bin/bash -eu + +SECONDS_PER_TARGET="${1:-15}" +CORPUS="${2:-wpd-test-data}" + +case "$SECONDS_PER_TARGET" in + ''|*[!0-9]*|0*) echo "fuzz-smoke.sh: seconds must be positive" >&2; exit 1 ;; +esac + +# Keep generated seeds separate from saved developer corpora and artifacts. +mkdir -p fuzz/corpus/ci/{container,e2e,vp8,vp8l} fuzz/artifacts +python3 - "$CORPUS" <<'PY' +from pathlib import Path +import sys + +corpus = Path(sys.argv[1]) +out = Path("fuzz/corpus/ci") +files = sorted(corpus.glob("*.webp")) +if not files: + sys.exit(f"no WebP files found in {corpus}") + +def chunks(data): + offset = 0 + while offset + 8 <= len(data): + tag = data[offset:offset + 4] + size = int.from_bytes(data[offset + 4:offset + 8], "little") + end = offset + 8 + size + if end > len(data): + break + payload = data[offset + 8:end] + if tag == b"ANMF": + yield from chunks(payload[16:]) + else: + yield tag, payload + offset = end + (size & 1) + +for path in files: + data = path.read_bytes() + # Large seeds make a smoke run spend its budget replaying images. + if len(data) > 65536: + continue + for target in ("container", "e2e"): + (out / target / path.stem).write_bytes(data) + for index, (tag, payload) in enumerate(chunks(data[12:])): + name = f"{path.stem}-{index}" + if tag == b"VP8 ": + (out / "vp8" / name).write_bytes(payload) + elif tag == b"VP8L": + # The target uses two leading bytes to select ARGB and threading. + for threads in (0, 1): + (out / "vp8l" / f"{name}-{threads}").write_bytes( + bytes((0, threads)) + payload) + +for target in ("container", "e2e", "vp8", "vp8l"): + if not any((out / target).iterdir()): + sys.exit(f"no seeds prepared for {target}") +PY + +# With no target argument cargo-fuzz builds every declared target. +cargo +nightly fuzz build -O --debug-assertions +for target in container vp8 vp8l e2e; do + cargo +nightly fuzz run -O --debug-assertions "$target" \ + "fuzz/corpus/ci/$target" -- -max_total_time="$SECONDS_PER_TARGET" \ + -max_len=65536 -timeout=5 -rss_limit_mb=1024 -seed=1 +done From d9845efe72e325165f86801cc23559758eb1f221 Mon Sep 17 00:00:00 2001 From: Mike Sulsenti Date: Sun, 4 Oct 2026 10:37:45 -0400 Subject: [PATCH 4/7] cli: report JSON info and extract original metadata Add --info=json while preserving the existing --info text. Report canvas dimensions, frame count, raw durations, loops, alpha, ICC presence and a bounded top-level RIFF chunk list only after successful decoding. --icc-out, --exif-out and --xmp-out write unchanged metadata payloads; missing metadata creates an empty file. The options work with stdin, incremental decoding and replay. Reject stdout collisions and test malformed/truncated input, duration extremes, original binary metadata and output errors without adding dependencies. Validation: cargo test -p wpd-tool with all features and without assembly; scripts/stylecheck.sh; deno fmt README.md. --- README.md | 46 ++++++++++ tools/report.rs | 166 ++++++++++++++++++++++++++++++++++ tools/tests/cli.rs | 217 +++++++++++++++++++++++++++++++++++++++++++++ tools/wpd.rs | 127 ++++++++++++++++++++++---- 4 files changed, 540 insertions(+), 16 deletions(-) create mode 100644 tools/report.rs create mode 100644 tools/tests/cli.rs diff --git a/README.md b/README.md index a789d16..eddcd62 100644 --- a/README.md +++ b/README.md @@ -40,6 +40,52 @@ produces `libwpd-sealed.a` alongside `libwpd.a`. The sealed static lib aborts instead of unwinding on an internal panic. The build merges every object into one monolithic object for downstream consumers. +## CLI metadata + +```sh +build/wpd --info=json input.webp +build/wpd --icc-out profile.icc --exif-out exif.bin --xmp-out xmp.bin input.webp +``` + +`--info` retains its text output. `--info=json` writes one JSON object to stdout +after the complete image has decoded successfully: + +```json +{ + "width": 2, + "height": 1, + "frame_count": 1, + "loop_count": 0, + "has_alpha": false, + "has_icc": false, + "durations_ms": [0], + "chunks": [ + { "fourcc": "VP8L", "offset": 12, "size": 17, "complete": true } + ] +} +``` + +Dimensions describe the original canvas, including with `--scale`. Durations are +the original milliseconds, without a minimum playback delay; a still has one +duration of 0. An animation loop count of 0 means infinite repetition. The +ordered chunk list describes top-level RIFF chunks; offsets point to their +FourCC, sizes exclude the header and padding, and `complete` includes padding. +Raw VP8/VP8L input has an empty chunk list. Arbitrary FourCC bytes outside +printable ASCII are escaped as `\u00XX`. + +Metadata outputs contain the original chunk payload, without its header or +padding. Missing metadata produces an empty file. Extraction alone needs no +pixel output. It does not interpret EXIF orientation or apply colour profiles. +These options also work with `--stream`, `--loops`, and `--repeat`; JSON and +metadata are written once. JSON can accompany a pixel output file, but the pixel +output cannot also use stdout. Metadata paths must name files. + +The existing input and frame size limits apply. `--max-output` limits decoded +pixel output; metadata is bounded by `--max-input`. Exit codes are 0 for +success, 1 for input, decoding, limits or output errors, and 2 for invalid +arguments. Failed decodes write no JSON or metadata, and leave existing metadata +output files untouched. + ## Library `meson install -C build` installs the static/shared libraries, `wpd.h`, and diff --git a/tools/report.rs b/tools/report.rs new file mode 100644 index 0000000..63d9824 --- /dev/null +++ b/tools/report.rs @@ -0,0 +1,166 @@ +use std::io::{self, Write}; + +use wpd::api::ImageInfo; + +struct Chunk<'a> { + tag: &'a [u8], + offset: usize, + size: u32, + payload: &'a [u8], + complete: bool, +} + +struct Chunks<'a> { + data: &'a [u8], + offset: usize, + end: usize, +} + +impl<'a> Chunks<'a> { + fn new(data: &'a [u8]) -> Self { + let end = + if data.len() >= 12 && &data[..4] == b"RIFF" && &data[8..12] == b"WEBP" { + (u32::from_le_bytes(data[4..8].try_into().unwrap()) as u64 + 8) + .min(data.len() as u64) as usize + } else { + 0 + }; + + Self { + data, + offset: 12, + end, + } + } +} + +impl<'a> Iterator for Chunks<'a> { + type Item = Chunk<'a>; + + fn next(&mut self) -> Option { + if self.end.saturating_sub(self.offset) < 8 { + return None; + } + let offset = self.offset; + let size = + u32::from_le_bytes(self.data[offset + 4..offset + 8].try_into().unwrap()); + let padded = size as u64 + u64::from(size & 1); + let available = self.end - (offset + 8); + let complete = padded <= available as u64; + let payload = &self.data[offset + 8..offset + 8 + available.min(size as usize)]; + + self.offset = if complete { + offset + 8 + padded as usize + } else { + self.end + }; + Some(Chunk { + tag: &self.data[offset..offset + 4], + offset, + size, + payload, + complete, + }) + } +} + +/* FourCCs are bytes, including on damaged inputs. Escape each byte outside + * printable ASCII rather than replacing it with a Unicode replacement char. */ +fn write_fourcc(w: &mut impl Write, tag: &[u8]) -> io::Result<()> { + w.write_all(b"\"")?; + for &b in tag { + match b { + b'"' | b'\\' => w.write_all(&[b'\\', b])?, + 0x20..=0x7e => w.write_all(&[b])?, + _ => write!(w, "\\u{b:04x}")?, + } + } + w.write_all(b"\"") +} + +pub fn write_json( + mut w: impl Write, + image: &ImageInfo, + data: &[u8], + has_icc: bool, +) -> io::Result<()> { + write!(w, "{{\"width\":{},\"height\":{},\"frame_count\":{},\"loop_count\":{},\"has_alpha\":{},\"has_icc\":{},\"durations_ms\":[", + image.width, image.height, image.frame_count, image.loop_count, image.has_alpha, has_icc)?; + if image.is_animation { + let mut first = true; + + for chunk in Chunks::new(data) { + if chunk.tag != b"ANMF" || chunk.payload.len() < 16 { + continue; + } + if !first { + w.write_all(b",")?; + } + first = false; + let duration = u32::from_le_bytes([ + chunk.payload[12], + chunk.payload[13], + chunk.payload[14], + 0, + ]); + + write!(w, "{duration}")?; + } + } else { + w.write_all(b"0")?; + } + w.write_all(b"],\"chunks\":[")?; + for (i, chunk) in Chunks::new(data).enumerate() { + if i != 0 { + w.write_all(b",")?; + } + w.write_all(b"{\"fourcc\":")?; + write_fourcc(&mut w, chunk.tag)?; + write!( + w, + ",\"offset\":{},\"size\":{},\"complete\":{}}}", + chunk.offset, chunk.size, chunk.complete + )?; + } + w.write_all(b"]}\n")?; + w.flush() +} + +#[cfg(test)] +mod tests { + use super::*; + + #[test] + fn arbitrary_fourcc_bytes_are_valid_json_strings() { + let mut out = Vec::new(); + + write_fourcc(&mut out, &[0, b'"', b'\\', 255]).unwrap(); + assert_eq!(out, b"\"\\u0000\\\"\\\\\\u00ff\""); + } + + #[test] + fn a_chunk_larger_than_the_remaining_input_is_reported_once() { + let mut data = b"RIFF\x14\0\0\0WEBPTEST\xff\xff\xff\xff".to_vec(); + + data.extend_from_slice(b"tail"); + let chunks: Vec<_> = Chunks::new(&data).collect(); + + assert_eq!(chunks.len(), 1); + assert_eq!(chunks[0].size, u32::MAX); + assert_eq!(chunks[0].payload, b"tail"); + assert!(!chunks[0].complete); + for cut in 0..data.len() { + let _ = Chunks::new(&data[..cut]).count(); + } + } + + #[test] + fn padding_and_the_riff_boundary_do_not_become_chunks() { + let data = b"RIFF\x0e\0\0\0WEBPTEST\x01\0\0\0x\0EXTRA123"; + let chunks: Vec<_> = Chunks::new(data).collect(); + + assert_eq!(chunks.len(), 1); + assert_eq!(chunks[0].payload, b"x"); + assert!(chunks[0].complete); + } +} diff --git a/tools/tests/cli.rs b/tools/tests/cli.rs new file mode 100644 index 0000000..10efbb8 --- /dev/null +++ b/tools/tests/cli.rs @@ -0,0 +1,217 @@ +#![forbid(unsafe_code)] + +use std::fs; +use std::io::Write; +use std::path::PathBuf; +use std::process::{Command, Output, Stdio}; +use std::sync::atomic::{AtomicUsize, Ordering}; + +/* A synthetic two-pixel lossless red/green image, encoded with cwebp 1.6.0 + * from P6\n2 1\n255\n followed by ff0000 00ff00. */ +const STILL: &[u8] = &[ + 0x52, 0x49, 0x46, 0x46, 0x1e, 0, 0, 0, 0x57, 0x45, 0x42, 0x50, 0x56, 0x50, 0x38, + 0x4c, 0x11, 0, 0, 0, 0x2f, 1, 0, 0, 0, 0x0f, 0xb0, 0xff, 0xf3, 0x1f, 0xf3, 0x1f, + 0x15, 0x32, 0xa2, 0xff, 1, 0, +]; + +struct Scratch(PathBuf); + +impl Scratch { + fn new() -> Self { + static NEXT: AtomicUsize = AtomicUsize::new(0); + let root = PathBuf::from(env!("CARGO_MANIFEST_DIR")).join("../wpd-test-data"); + + fs::create_dir_all(&root).unwrap(); + let path = root.join(format!( + "cli-test-{}-{}", + std::process::id(), + NEXT.fetch_add(1, Ordering::Relaxed) + )); + + fs::create_dir(&path).unwrap(); + Self(path) + } + + fn path(&self, name: &str) -> String { + self.0.join(name).to_str().unwrap().to_owned() + } +} + +impl Drop for Scratch { + fn drop(&mut self) { + /* This directory contains only the synthetic inputs and outputs of + * this test, never the corpus's reference files. */ + fs::remove_dir_all(&self.0).unwrap(); + } +} + +fn run(args: &[&str], input: &[u8]) -> Output { + let mut child = Command::new(env!("CARGO_BIN_EXE_wpd")) + .args(args) + .stdin(Stdio::piped()) + .stdout(Stdio::piped()) + .stderr(Stdio::piped()) + .spawn() + .unwrap(); + + child.stdin.take().unwrap().write_all(input).unwrap(); + child.wait_with_output().unwrap() +} + +fn chunk(tag: &[u8; 4], payload: &[u8]) -> Vec { + let mut out = tag.to_vec(); + + out.extend_from_slice(&(payload.len() as u32).to_le_bytes()); + out.extend_from_slice(payload); + if payload.len() & 1 != 0 { + out.push(0); + } + out +} + +fn riff(payload: &[u8]) -> Vec { + let mut out = b"RIFF".to_vec(); + + out.extend_from_slice(&(payload.len() as u32 + 4).to_le_bytes()); + out.extend_from_slice(b"WEBP"); + out.extend_from_slice(payload); + out +} + +fn extended() -> Vec { + let mut out = chunk(b"VP8X", &[0x2c, 0, 0, 0, 1, 0, 0, 0, 0, 0]); + + out.extend(chunk(b"ICCP", &[0, 255, b'"', b'\\'])); + out.extend_from_slice(&STILL[12..]); + out.extend(chunk(b"EXIF", b"Exif\0binary")); + out.extend(chunk(b"XMP ", b"\n")); + riff(&out) +} + +fn animation() -> Vec { + let mut out = chunk(b"VP8X", &[2, 0, 0, 0, 1, 0, 0, 0, 0, 0]); + + out.extend(chunk(b"ANIM", &[0, 0, 0, 0, 3, 0])); + for duration in [0u32, 1, 0xff_ffff] { + let mut frame = [0u8; 16].to_vec(); + + frame[6] = 1; + frame[12..15].copy_from_slice(&duration.to_le_bytes()[..3]); + frame[15] = 2; + frame.extend_from_slice(&STILL[12..]); + out.extend(chunk(b"ANMF", &frame)); + } + riff(&out) +} + +fn success(output: &Output) -> &str { + assert!( + output.status.success(), + "{}", + String::from_utf8_lossy(&output.stderr) + ); + std::str::from_utf8(&output.stdout).unwrap() +} + +#[test] +fn still_json_is_one_complete_object_without_text_info() { + let output = run(&["--info=json", "-"], STILL); + + assert_eq!(success(&output), concat!( + "{\"width\":2,\"height\":1,\"frame_count\":1,\"loop_count\":0,", + "\"has_alpha\":false,\"has_icc\":false,\"durations_ms\":[0],", + "\"chunks\":[{\"fourcc\":\"VP8L\",\"offset\":12,\"size\":17,\"complete\":true}]}\n" + )); + assert!(success(&run(&["--info", "-"], STILL)).starts_with("canvas: 2x1\n")); +} + +#[test] +fn streamed_json_keeps_raw_duration_extremes_and_the_loop_count() { + let data = animation(); + let whole = run(&["--info=json", "-"], &data); + let text = success(&whole); + + assert!(text.contains("\"frame_count\":3,\"loop_count\":3")); + assert!(text.contains("\"durations_ms\":[0,1,16777215]")); + assert_eq!(text.matches("\"fourcc\":\"ANMF\"").count(), 3); + for size in ["1", "7", "13", "997"] { + let stream = run( + &["--info=json", "--stream", size, "--loops", "2", "-"], + &data, + ); + + assert_eq!(success(&stream), text); + } +} + +#[test] +fn metadata_is_written_as_original_bytes_even_after_streaming() { + let dir = Scratch::new(); + let icc = dir.path("profile.icc"); + let exif = dir.path("exif.bin"); + let xmp = dir.path("xmp.bin"); + let args = [ + "--info=json", + "--stream", + "1", + "--icc-out", + &icc, + "--exif-out", + &exif, + "--xmp-out", + &xmp, + "-", + ]; + let output = run(&args, &extended()); + + assert!(success(&output).contains("\"has_icc\":true")); + assert_eq!(fs::read(&icc).unwrap(), [0, 255, b'"', b'\\']); + assert_eq!(fs::read(&exif).unwrap(), b"Exif\0binary"); + assert_eq!(fs::read(&xmp).unwrap(), b"\n"); + assert!(run(&["--icc-out", &icc, "-"], STILL).status.success()); + assert!(fs::read(&icc).unwrap().is_empty()); +} + +#[test] +fn failed_decodes_do_not_publish_json_or_metadata() { + let dir = Scratch::new(); + let icc = dir.path("profile.icc"); + let data = extended(); + + for size in ["1", "997"] { + let output = run( + &["--info=json", "--stream", size, "--icc-out", &icc, "-"], + &data[..data.len() - 1], + ); + + assert!(!output.status.success()); + assert!(output.stdout.is_empty()); + assert!(!dir.0.join("profile.icc").exists()); + } + for cut in 0..STILL.len() - 1 { + let output = run(&["--info=json", "-"], &STILL[..cut]); + + assert!(!output.status.success(), "accepted truncation at {cut}"); + assert!(output.stdout.is_empty()); + } + let output = run(&["--info=json", "--max-input", "1", "-"], STILL); + + assert!(!output.status.success()); + assert!(output.stdout.is_empty()); +} + +#[test] +fn metadata_io_errors_and_stdout_collisions_fail_cleanly() { + let dir = Scratch::new(); + let missing = dir.path("missing/profile.icc"); + let output = run(&["--info=json", "--icc-out", &missing, "-"], STILL); + + assert_eq!(output.status.code(), Some(1)); + assert!(output.stdout.is_empty()); + for args in [vec!["--info=json", "-", "-"], vec!["--icc-out", "-", "-"]] { + let output = run(&args, &[]); + + assert_eq!(output.status.code(), Some(2)); + assert!(output.stdout.is_empty()); + } +} diff --git a/tools/wpd.rs b/tools/wpd.rs index d54e2c5..b6303e6 100644 --- a/tools/wpd.rs +++ b/tools/wpd.rs @@ -2,6 +2,7 @@ mod md5; mod output; +mod report; use std::ffi::{OsStr, OsString}; use std::io::{Read, Write}; @@ -99,6 +100,12 @@ const USAGE_TAIL: &str = concat!( " --info\n", " print canvas, animation, the frame table and per-frame\n", " timing to stdout\n", + " --info=json\n", + " write one JSON object to stdout after successful decoding;\n", + " includes dimensions, durations, loops, alpha and chunks\n", + " --icc-out path, --exif-out path, --xmp-out path\n", + " write the original metadata bytes; absent metadata writes\n", + " an empty file. these paths cannot be stdout\n", " --stream u32\n", " decode incrementally, appending this many bytes at a time,\n", " instead of opening the file whole\n", @@ -280,6 +287,8 @@ struct Options { scale: Option<(i32, i32)>, frame_size_limit: u32, info: bool, + info_json: bool, + metadata_out: [Option; 3], subframe: bool, muxer: Option, verify: Option, @@ -303,6 +312,9 @@ const OPTIONS: &[(&str, Option, bool)] = &[ ("muxer", None, true), ("verify", None, true), ("info", None, false), + ("icc-out", None, true), + ("exif-out", None, true), + ("xmp-out", None, true), ("loops", None, true), ("cpumask", None, true), ("subframe", None, false), @@ -372,7 +384,29 @@ fn set(o: &mut Options, name: &str, value: String) -> Result<(), &'static str> { warn_baseline_cpumask(mask); api::set_cpu_flags_mask(mask); } - "info" => o.info = true, + "info" => match value.as_str() { + "" => { + o.info = true; + o.info_json = false; + } + "json" => { + o.info_json = true; + o.info = false; + } + _ => return Err(BAD_INFO), + }, + "icc-out" | "exif-out" | "xmp-out" => { + if value.is_empty() || value == "-" { + return Err(BAD_METADATA_OUT); + } + let index = match name { + "icc-out" => 0, + "exif-out" => 1, + _ => 2, + }; + + o.metadata_out[index] = Some(value); + } "subframe" => o.subframe = true, _ => return Err(MISSING), } @@ -425,7 +459,7 @@ fn parse_args(argv: &[OsString]) -> Parsed { return Parsed::Bad(MISSING); }; - if !takes_value && attached.is_some() { + if !takes_value && attached.is_some() && name != "info" { return Parsed::Bad(MISSING); } if name == "help" { @@ -438,7 +472,7 @@ fn parse_args(argv: &[OsString]) -> Parsed { None => return Parsed::Bad(MISSING), } } else { - String::new() + attached.unwrap_or_default() }; if let Err(e) = set(&mut o, name, value) { @@ -497,6 +531,8 @@ const BAD_PIXELS: &str = "invalid frame size limit; expected a pixel count or Wx const BAD_FORMAT: &str = "invalid output pixel format"; const BAD_MUXER: &str = "invalid output muxer; expected raw, md5, ppm, pam or y4m"; const BAD_SIZE: &str = "invalid byte count; expected digits with an optional K, M or G"; +const BAD_INFO: &str = "invalid info format; expected --info or --info=json"; +const BAD_METADATA_OUT: &str = "metadata output requires a file path, not stdout"; fn errmsg(e: &std::io::Error) -> String { let text = e.to_string(); @@ -786,7 +822,10 @@ fn main() -> ExitCode { let operands = opts.positional.len(); let max = if verifying { 1 } else { 2 }; - if operands < 1 || operands > max || (!verifying && !opts.info && operands != 2) { + let reporting = + opts.info || opts.info_json || opts.metadata_out.iter().any(Option::is_some); + + if operands < 1 || operands > max || (!verifying && !reporting && operands != 2) { let reason = if verifying { if operands < 1 { "input is required" @@ -810,6 +849,14 @@ fn main() -> ExitCode { Some(opts.positional[1].as_os_str()) }; + if opts.info_json && output_name == Some(OsStr::new("-")) { + usage( + &app, + Some("JSON info and decoded output cannot both use stdout"), + ); + return ExitCode::from(2); + } + run(&opts, input_name, output_name, expected_md5) } @@ -869,6 +916,7 @@ fn run( let writes = opened && !output.is_null(); let mut frames = 0; + let mut last_decoder = None; for iter in 0..opts.repeat { let mut info_printed = false; @@ -954,6 +1002,9 @@ fn run( if ret < 0 { return ExitCode::FAILURE; } + if iter + 1 == opts.repeat { + last_decoder = Some(decoder); + } } if frames == 0 { @@ -961,28 +1012,72 @@ fn run( return ExitCode::FAILURE; } if let Some(expected) = expected_md5 { - return if output.verify(&expected) { - ExitCode::SUCCESS - } else { - ExitCode::FAILURE - }; + if !output.verify(&expected) { + return ExitCode::FAILURE; + } + } else if let Err(e) = output.close() { + let _ = writeln!(std::io::stderr(), "write: {}", errmsg(&e)); + return ExitCode::FAILURE; } - if !opened { - return ExitCode::SUCCESS; + let mut decoder = last_decoder.unwrap(); + + for (which, path) in [Metadata::Iccp, Metadata::Exif, Metadata::Xmp] + .into_iter() + .zip(&opts.metadata_out) + { + if let Some(path) = path { + if let Err(e) = + std::fs::write(path, decoder.metadata(which).unwrap_or_default()) + { + eprintln!("{path}: {}", errmsg(&e)); + return ExitCode::FAILURE; + } + } } - match output.close() { - Ok(()) => ExitCode::SUCCESS, - Err(e) => { - let _ = writeln!(std::io::stderr(), "write: {}", errmsg(&e)); - ExitCode::FAILURE + if opts.info_json { + let Ok(image) = decoder.info() else { + eprintln!("{}: {}", input_name.to_string_lossy(), decoder.error()); + return ExitCode::FAILURE; + }; + let has_icc = decoder.metadata(Metadata::Iccp).is_some(); + + if let Err(e) = + report::write_json(std::io::stdout().lock(), &image, &data, has_icc) + { + eprintln!("write: {}", errmsg(&e)); + return ExitCode::FAILURE; } } + ExitCode::SUCCESS } #[cfg(test)] mod tests { use super::*; + #[test] + fn info_keeps_its_flag_form_and_takes_json_only_after_equals() { + for (args, text, json) in [ + (vec!["wpd", "--info"], true, false), + (vec!["wpd", "--info=json"], false, true), + (vec!["wpd", "--info", "--info=json"], false, true), + ] { + let argv: Vec<_> = args.into_iter().map(OsString::from).collect(); + let Parsed::Ok(o) = parse_args(&argv) else { + panic!("info options must parse"); + }; + + assert_eq!(o.info, text); + assert_eq!(o.info_json, json); + } + let argv: Vec<_> = ["wpd", "--info=xml"] + .into_iter() + .map(OsString::from) + .collect(); + + assert!(matches!(parse_args(&argv), Parsed::Bad(BAD_INFO))); + } + #[test] fn a_size_takes_a_binary_suffix_and_rejects_overflow() { assert_eq!(parse_size("0"), Some(0)); From 353924eb986ea026184a09a320d5493afc4ba7e7 Mon Sep 17 00:00:00 2001 From: Mike Sulsenti Date: Sun, 4 Oct 2026 10:45:48 -0400 Subject: [PATCH 5/7] cli: write PAM frame sequences with a timing manifest Add --muxer frames: a new output directory containing numbered straight-RGBA PAM files and a JSON manifest with original durations, timestamps, loop count, output dimensions and canvas offsets. Only the first replay or repeat pass is written. Refuse existing directories and count both pixels and manifest against --max-output. Publish manifest.json only after complete successful decoding and output flushing; failures retain partial files and manifest.json.part. Test byte-at-a-time streaming, duration extremes, replay, pixel parity with concatenated PAM, subframes, scaling, exact output budgets, truncated input and preservation of existing directories. Validation: wpd-tool tests with all features and without assembly; scripts/stylecheck.sh; deno fmt README.md. --- README.md | 29 ++++++++ tools/output.rs | 85 +++++++++++++++++++++++- tools/tests/cli.rs | 160 +++++++++++++++++++++++++++++++++++++++++++++ tools/wpd.rs | 26 +++++++- 4 files changed, 294 insertions(+), 6 deletions(-) diff --git a/README.md b/README.md index eddcd62..7c8a543 100644 --- a/README.md +++ b/README.md @@ -86,6 +86,35 @@ success, 1 for input, decoding, limits or output errors, and 2 for invalid arguments. Failed decodes write no JSON or metadata, and leave existing metadata output files untouched. +## CLI frame sequences + +```sh +build/wpd --muxer frames input.webp output-frames +``` + +The `frames` muxer creates a new directory, writes `frame-000000.pam`, +`frame-000001.pam`, and so on, and publishes `manifest.json` after successful +decoding and output flushing. Every PAM file has straight RGBA pixels and its +own header. An existing output directory is refused. + +The manifest has `canvas_width`, `canvas_height`, `loop_count`, `composited`, +`frame_count`, and an ordered `frames` array. Each frame has `file`, `width`, +`height`, `duration_ms`, `timestamp_ms`, `x`, and `y`. Durations are the +original milliseconds, including 0; timestamps are the start of each frame in +the first pass, starting at 0. A still has one frame with duration and +timestamp 0. + +Frames are composited canvases by default. `--subframe` keeps raw frame sizes +and canvas offsets, and sets `composited` to false. `--scale` changes the frame +dimensions while the manifest's canvas dimensions describe the original file. +`--stream` is supported; only the first pass of `--loops` or `--repeat` is +written. The muxer requires RGBA output and a directory path, and `--max-output` +covers all PAM bytes and the complete manifest together. + +A failed decode or write leaves partial output with `manifest.json.part`, +without a final `manifest.json`. Consumers should require the final manifest +before using a sequence. + ## Library `meson install -C build` installs the static/shared libraries, `wpd.h`, and diff --git a/tools/output.rs b/tools/output.rs index 1b148c3..ba0a611 100644 --- a/tools/output.rs +++ b/tools/output.rs @@ -1,6 +1,7 @@ use std::ffi::OsStr; use std::fs::File; use std::io::{self, BufWriter, Write}; +use std::path::PathBuf; use wpd::api::{Coding, ImageInfo, Picture}; use wpd::dsp::yuv::{extract_alpha, YuvDsp}; @@ -13,6 +14,7 @@ pub enum Muxer { Raw, Ppm, Pam, + Frames, Y4m, } @@ -22,6 +24,7 @@ impl Muxer { Muxer::Raw => "raw", Muxer::Ppm => "ppm", Muxer::Pam => "pam", + Muxer::Frames => "frames", Muxer::Y4m => "y4m", } } @@ -29,7 +32,7 @@ impl Muxer { fn required(self) -> Option<(&'static str, Format)> { match self { Muxer::Ppm => Some(("rgb", Format::Rgb)), - Muxer::Pam => Some(("rgba", Format::Rgba)), + Muxer::Pam | Muxer::Frames => Some(("rgba", Format::Rgba)), Muxer::Raw | Muxer::Y4m => None, } } @@ -46,6 +49,7 @@ pub struct Output { kind: Kind, pub muxer: Muxer, file: Option>, + directory: Option, /* Bytes written to a file or stdout so far, and the most a decode may * write in total; 0 lifts the limit. A hashed or discarded decode costs * nothing downstream, so only Kind::File is budgeted. */ @@ -138,16 +142,36 @@ impl Output { out.muxer = match chosen.as_str() { "ppm" => Muxer::Ppm, "pam" => Muxer::Pam, + "frames" => Muxer::Frames, "y4m" => Muxer::Y4m, _ => Muxer::Raw, }; if out.kind == Kind::Null { + if out.muxer == Muxer::Frames { + return Err(io::Error::other( + "frames requires an output directory", + )); + } return Ok(out); } } let name = filename.unwrap_or(OsStr::new("")); + if out.muxer == Muxer::Frames { + if name.is_empty() || name == OsStr::new("-") { + return Err(io::Error::other("frames requires an output directory")); + } + let directory = PathBuf::from(name); + + std::fs::create_dir(&directory)?; + out.file = Some(Box::new(BufWriter::new(File::create_new( + directory.join("manifest.json.part"), + )?))); + out.directory = Some(directory); + return Ok(out); + } + out.file = Some(if name == OsStr::new("-") { Box::new(io::stdout()) } else { @@ -161,6 +185,7 @@ impl Output { kind: Kind::Null, muxer: Muxer::Raw, file: None, + directory: None, written: 0, limit: 0, y4m_stash: Y4M_STASH, @@ -200,6 +225,9 @@ impl Output { } pub fn close(mut self) -> io::Result<()> { + if self.muxer == Muxer::Frames { + self.write(format!("],\"frame_count\":{}}}\n", self.frames).as_bytes())?; + } if self.kind == Kind::Md5 { let digest = hex(&std::mem::take(&mut self.md5).finish()); @@ -210,6 +238,53 @@ impl Output { if let Some(mut f) = self.file.take() { f.flush()?; } + if let Some(directory) = self.directory.take() { + std::fs::rename( + directory.join("manifest.json.part"), + directory.join("manifest.json"), + )?; + } + Ok(()) + } + + pub fn begin_sequence( + &mut self, + image: &ImageInfo, + subframe: bool, + ) -> io::Result<()> { + let header = format!( + "{{\"canvas_width\":{},\"canvas_height\":{},\"loop_count\":{},\"composited\":{},\"frames\":[", + image.width, image.height, image.loop_count, !subframe + ); + + self.write(header.as_bytes()) + } + + fn write_sequence_frame( + &mut self, + frame: &Picture<'_>, + header: &[u8], + ) -> io::Result<()> { + let name = format!("frame-{:06}.pam", self.frames); + let path = self.directory.as_ref().unwrap().join(&name); + let file = File::create_new(path)?; + let manifest = self.file.take(); + + self.file = Some(Box::new(BufWriter::new(file))); + self.write(header)?; + self.write_plane(frame, 0)?; + self.file.take().unwrap().flush()?; + self.file = manifest; + + let (x, y) = frame.position(); + let entry = format!( + "{}{{\"file\":\"{name}\",\"width\":{},\"height\":{},\"duration_ms\":{},\"timestamp_ms\":{},\"x\":{x},\"y\":{y}}}", + if self.frames == 0 { "" } else { "," }, + frame.width(), frame.height(), frame.duration(), frame.timestamp() + ); + + self.write(entry.as_bytes())?; + self.frames += 1; Ok(()) } @@ -370,8 +445,12 @@ impl Output { ) }; - self.write(header.as_bytes())?; - self.write_plane(frame, 0) + if self.muxer == Muxer::Frames { + self.write_sequence_frame(frame, header.as_bytes()) + } else { + self.write(header.as_bytes())?; + self.write_plane(frame, 0) + } } (None, Muxer::Y4m) => self.write_y4m(frame), (None, _) => self.write_raw(frame, pixel_format), diff --git a/tools/tests/cli.rs b/tools/tests/cli.rs index 10efbb8..8273ba6 100644 --- a/tools/tests/cli.rs +++ b/tools/tests/cli.rs @@ -215,3 +215,163 @@ fn metadata_io_errors_and_stdout_collisions_fail_cleanly() { assert!(output.stdout.is_empty()); } } + +#[test] +fn a_frame_sequence_keeps_durations_and_matches_the_pam_stream() { + let dir = Scratch::new(); + let data = animation(); + let reference = run(&["--muxer", "pam", "-", "-"], &data); + + assert!( + reference.status.success(), + "{}", + String::from_utf8_lossy(&reference.stderr) + ); + let mut expected_manifest = None; + + for (name, stream) in [("whole", "997"), ("stream", "1")] { + let path = dir.path(name); + let output = run( + &[ + "--muxer", "frames", "--stream", stream, "--loops", "2", "--repeat", + "2", "-", &path, + ], + &data, + ); + + assert!(success(&output).is_empty()); + let manifest = + fs::read_to_string(dir.0.join(name).join("manifest.json")).unwrap(); + + assert!(manifest.contains("\"loop_count\":3,\"composited\":true")); + assert!(manifest.contains("\"duration_ms\":0,")); + assert!(manifest.contains("\"duration_ms\":1,")); + assert!(manifest.contains("\"duration_ms\":16777215,")); + assert!(manifest.ends_with("],\"frame_count\":3}\n")); + assert!(!dir.0.join(name).join("manifest.json.part").exists()); + if let Some(expected) = &expected_manifest { + assert_eq!(&manifest, expected); + } else { + expected_manifest = Some(manifest); + } + let mut pixels = Vec::new(); + + for i in 0..3 { + pixels.extend( + fs::read(dir.0.join(name).join(format!("frame-{i:06}.pam"))).unwrap(), + ); + } + assert_eq!(pixels, reference.stdout); + assert_eq!(fs::read_dir(dir.0.join(name)).unwrap().count(), 4); + } +} + +#[test] +fn sequence_failure_has_no_final_manifest_and_never_overwrites_a_directory() { + let dir = Scratch::new(); + let path = dir.path("frames"); + let output = run( + &["--muxer", "frames", "--max-output", "1", "-", &path], + &animation(), + ); + + assert_eq!(output.status.code(), Some(1)); + assert!(!dir.0.join("frames/manifest.json").exists()); + let sentinel = dir.0.join("frames/owner-file"); + + fs::write(&sentinel, b"preserve this").unwrap(); + let output = run(&["--muxer", "frames", "-", &path], STILL); + + assert_eq!(output.status.code(), Some(1)); + assert_eq!(fs::read(sentinel).unwrap(), b"preserve this"); + let data = animation(); + let truncated = dir.path("truncated"); + let output = run( + &["--muxer", "frames", "--stream", "1", "-", &truncated], + &data[..data.len() - 2], + ); + + assert_eq!(output.status.code(), Some(1)); + assert!(!dir.0.join("truncated/manifest.json").exists()); +} + +#[test] +fn sequence_byte_limit_covers_the_manifest_and_every_pam_file() { + let dir = Scratch::new(); + let path = dir.path("reference"); + + success(&run(&["--muxer", "frames", "-", &path], STILL)); + let size: u64 = fs::read_dir(&path) + .unwrap() + .map(|entry| entry.unwrap().metadata().unwrap().len()) + .sum(); + let exact = dir.path("exact"); + let exact_limit = size.to_string(); + + success(&run( + &[ + "--muxer", + "frames", + "--max-output", + &exact_limit, + "-", + &exact, + ], + STILL, + )); + let short = dir.path("short"); + let short_limit = (size - 1).to_string(); + let output = run( + &[ + "--muxer", + "frames", + "--max-output", + &short_limit, + "-", + &short, + ], + STILL, + ); + + assert_eq!(output.status.code(), Some(1)); + assert!(!dir.0.join("short/manifest.json").exists()); + for args in [ + vec!["--muxer", "frames", "-", "-"], + vec!["--muxer", "frames", "--info", "-"], + ] { + assert_eq!(run(&args, &[]).status.code(), Some(2)); + } +} + +#[test] +fn sequence_subframes_and_scaling_describe_the_actual_output_geometry() { + let dir = Scratch::new(); + + for (name, options, geometry, composited) in [ + ( + "scaled", + vec!["--scale", "4x2"], + "\"width\":4,\"height\":2", + true, + ), + ( + "subframes", + vec!["--subframe"], + "\"width\":2,\"height\":1", + false, + ), + ] { + let path = dir.path(name); + let mut args = vec!["--muxer", "frames"]; + + args.extend(options); + args.extend(["-", &path]); + success(&run(&args, &animation())); + let manifest = + fs::read_to_string(dir.0.join(name).join("manifest.json")).unwrap(); + + assert!(manifest.contains("\"canvas_width\":2,\"canvas_height\":1")); + assert!(manifest.contains(&format!("\"composited\":{composited}"))); + assert_eq!(manifest.matches(geometry).count(), 3); + } +} diff --git a/tools/wpd.rs b/tools/wpd.rs index b6303e6..9aa49eb 100644 --- a/tools/wpd.rs +++ b/tools/wpd.rs @@ -87,8 +87,10 @@ const USAGE_HEAD: &str = concat!( " letter marks the channels alpha is multiplied into, and\n", " the bgr 16-bit ones swap the two bytes of every pixel\n", " --muxer str\n", - " output muxer (raw, md5, ppm, pam, y4m); default is selected\n", + " output muxer (raw, md5, ppm, pam, y4m, frames); default is selected\n", " from a .ppm, .pam or .y4m output extension, or raw\n", + " frames creates a new output directory of RGBA PAM files\n", + " and a JSON manifest with original per-frame durations\n", " --verify md5\n", " verify decoded md5; implies --muxer md5 and no output\n", " --cpumask str\n", @@ -372,7 +374,10 @@ fn set(o: &mut Options, name: &str, value: String) -> Result<(), &'static str> { o.out_format = out_format; } "muxer" => { - if !matches!(value.as_str(), "raw" | "md5" | "ppm" | "pam" | "y4m") { + if !matches!( + value.as_str(), + "raw" | "md5" | "ppm" | "pam" | "y4m" | "frames" + ) { return Err(BAD_MUXER); } o.muxer = Some(value); @@ -529,7 +534,8 @@ const BAD_THREADS: &str = "invalid thread count; expected 0..INT_MAX"; const BAD_SCALE: &str = "invalid scale; expected WxH, either 0 to keep the ratio"; const BAD_PIXELS: &str = "invalid frame size limit; expected a pixel count or WxH"; const BAD_FORMAT: &str = "invalid output pixel format"; -const BAD_MUXER: &str = "invalid output muxer; expected raw, md5, ppm, pam or y4m"; +const BAD_MUXER: &str = + "invalid output muxer; expected raw, md5, ppm, pam, y4m or frames"; const BAD_SIZE: &str = "invalid byte count; expected digits with an optional K, M or G"; const BAD_INFO: &str = "invalid info format; expected --info or --info=json"; const BAD_METADATA_OUT: &str = "metadata output requires a file path, not stdout"; @@ -856,6 +862,14 @@ fn main() -> ExitCode { ); return ExitCode::from(2); } + if opts.muxer.as_deref() == Some("frames") + && output_name.is_none_or(|name| { + name == OsStr::new("-") || name == OsStr::new("/dev/null") + }) + { + usage(&app, Some("frames requires an output directory")); + return ExitCode::from(2); + } run(&opts, input_name, output_name, expected_md5) } @@ -912,6 +926,12 @@ fn run( { return ExitCode::FAILURE; } + if output.muxer == Muxer::Frames { + if let Err(e) = output.begin_sequence(&image, opts.subframe) { + eprintln!("write: {}", errmsg(&e)); + return ExitCode::FAILURE; + } + } } let writes = opened && !output.is_null(); From 9ca31db7db0ca3c3a33c2e0e68d13efdb3cb87fd Mon Sep 17 00:00:00 2001 From: Mike Sulsenti Date: Sun, 4 Oct 2026 10:59:17 -0400 Subject: [PATCH 6/7] cli: keep JSON output separate from stdout aliases Reject decoded output sent to /dev/stdout, /dev/fd/1 or /proc/self/fd/1 alongside --info=json, just as for -. Otherwise these ordinary stdout paths successfully emitted PAM bytes followed by JSON. Reject the same aliases for metadata extraction and frame directories. Cover every recognized alias through command-line integration tests and document the path-based guarantee. Validation: CLI integration tests, scripts/stylecheck.sh, deno fmt README.md. --- README.md | 4 +++- tools/tests/cli.rs | 12 ++++++++++++ tools/wpd.rs | 15 ++++++++++----- 3 files changed, 25 insertions(+), 6 deletions(-) diff --git a/README.md b/README.md index 7c8a543..d82180c 100644 --- a/README.md +++ b/README.md @@ -78,7 +78,9 @@ padding. Missing metadata produces an empty file. Extraction alone needs no pixel output. It does not interpret EXIF orientation or apply colour profiles. These options also work with `--stream`, `--loops`, and `--repeat`; JSON and metadata are written once. JSON can accompany a pixel output file, but the pixel -output cannot also use stdout. Metadata paths must name files. +output cannot also use stdout. Metadata paths must name files. `-`, +`/dev/stdout`, `/dev/fd/1`, and `/proc/self/fd/1` are recognized as stdout; +custom links to stdout must not be used as output paths with JSON info. The existing input and frame size limits apply. `--max-output` limits decoded pixel output; metadata is bounded by `--max-input`. Exit codes are 0 for diff --git a/tools/tests/cli.rs b/tools/tests/cli.rs index 8273ba6..a03ce5f 100644 --- a/tools/tests/cli.rs +++ b/tools/tests/cli.rs @@ -341,6 +341,18 @@ fn sequence_byte_limit_covers_the_manifest_and_every_pam_file() { ] { assert_eq!(run(&args, &[]).status.code(), Some(2)); } + for alias in ["-", "/dev/stdout", "/dev/fd/1", "/proc/self/fd/1"] { + for args in [ + vec!["--info=json", "--muxer", "pam", "-", alias], + vec!["--info=json", "--icc-out", alias, "-"], + vec!["--muxer", "frames", "-", alias], + ] { + let output = run(&args, &[]); + + assert_eq!(output.status.code(), Some(2)); + assert!(output.stdout.is_empty()); + } + } } #[test] diff --git a/tools/wpd.rs b/tools/wpd.rs index 9aa49eb..ee238da 100644 --- a/tools/wpd.rs +++ b/tools/wpd.rs @@ -278,6 +278,12 @@ fn parse_size(value: &str) -> Option { digits.parse::().ok()?.checked_mul(unit) } +fn is_stdout(path: &OsStr) -> bool { + ["-", "/dev/stdout", "/dev/fd/1", "/proc/self/fd/1"] + .iter() + .any(|alias| path == OsStr::new(alias)) +} + #[derive(Default)] struct Options { repeat: i32, @@ -401,7 +407,7 @@ fn set(o: &mut Options, name: &str, value: String) -> Result<(), &'static str> { _ => return Err(BAD_INFO), }, "icc-out" | "exif-out" | "xmp-out" => { - if value.is_empty() || value == "-" { + if value.is_empty() || is_stdout(OsStr::new(&value)) { return Err(BAD_METADATA_OUT); } let index = match name { @@ -855,7 +861,7 @@ fn main() -> ExitCode { Some(opts.positional[1].as_os_str()) }; - if opts.info_json && output_name == Some(OsStr::new("-")) { + if opts.info_json && output_name.is_some_and(is_stdout) { usage( &app, Some("JSON info and decoded output cannot both use stdout"), @@ -863,9 +869,8 @@ fn main() -> ExitCode { return ExitCode::from(2); } if opts.muxer.as_deref() == Some("frames") - && output_name.is_none_or(|name| { - name == OsStr::new("-") || name == OsStr::new("/dev/null") - }) + && output_name + .is_none_or(|name| is_stdout(name) || name == OsStr::new("/dev/null")) { usage(&app, Some("frames requires an output directory")); return ExitCode::from(2); From 92ce5fc1d770ff665810ae454e61a905531fd1ec Mon Sep 17 00:00:00 2001 From: Mike Sulsenti Date: Sun, 4 Oct 2026 13:40:32 -0400 Subject: [PATCH 7/7] docs: draft CLI changelog and first release guidance Complete the draft unreleased 0.2.0 changelog with the existing decoder and new JSON, metadata and frame-sequence CLI features. Build/CI and compatibility entries already accompany the PRs that add them. Suggest v0.2.0 as the first maintainer-created tag after merge and checks, without announcing a release date or crates.io publication. State Rust 1.98/1.99 testing without claiming the declared 1.82 MSRV has been established. Validation: deno fmt README.md CHANGELOG.md; scripts/stylecheck.sh. --- CHANGELOG.md | 34 ++++++++++++++++++++++++++++++++++ README.md | 8 ++++++++ 2 files changed, 42 insertions(+) diff --git a/CHANGELOG.md b/CHANGELOG.md index c83f2b6..ec32dd4 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -1,7 +1,25 @@ # Changelog +This is a draft for the first tagged release. There are no historical release +dates to record. The proposed first tag is `v0.2.0`, matching the current Cargo +workspace version, after the maintainer merges the patch series and completes +the release checks. Nothing here announces a tag or crates.io publication. + ## Unreleased — 0.2.0 +### Existing decoder functionality + +- Decode lossy VP8, lossless VP8L, alpha, and animated WebP in Rust, with + optional handwritten assembly and worker threads. +- Provide Rust and C APIs, incremental decoding, composited animation or raw + subframes, frame timing, metadata access, scaling, cropping, and frame size + limits. +- Provide a CLI with packed RGB/RGBA and planar YUV output, raw, PAM, PPM, + YUV4MPEG2 and MD5 muxers, replay, benchmark repeats, and input/output byte + budgets. +- Build static and shared C libraries with Meson, including headers and + pkg-config metadata. The project uses the BSD-2-Clause license. + ### Build and CI - Repair the end-to-end fuzz target's decoder options initializer so all four @@ -11,3 +29,19 @@ seeded smoke runs, and the existing correctness and sanitizer checks. - Bound fuzz-harness pictures to one megapixel so mutated dimensions fit the smoke run's memory budget without changing decoder limits. + +### CLI additions + +- Add `--info=json` with dimensions, frame count, raw frame durations, loop + count, alpha, ICC presence, and an ordered chunk list. Keep text `--info`. +- Add `--icc-out`, `--exif-out`, and `--xmp-out` for original metadata bytes. +- Add `--muxer frames` for numbered RGBA PAM files and a JSON timing manifest. + Publish the final manifest after successful decoding and output flushing, + refuse existing directories, and include the manifest in the output budget. +- Keep JSON separate from the recognized stdout paths and cover streaming, + truncation, binary metadata, replay, scaling, and output errors in CLI tests. + +### Validation scope + +The local patch series has been tested with Rust 1.98 and 1.99. This does not +establish the declared Rust 1.82 minimum or claim crates.io packaging support. diff --git a/README.md b/README.md index d82180c..2fa334c 100644 --- a/README.md +++ b/README.md @@ -40,6 +40,14 @@ produces `libwpd-sealed.a` alongside `libwpd.a`. The sealed static lib aborts instead of unwinding on an internal panic. The build merges every object into one monolithic object for downstream consumers. +## Release status + +[CHANGELOG.md](CHANGELOG.md) is a draft for the first release. The workspace is +already version 0.2.0; `v0.2.0` is the proposed first tag after the maintainer +merges these changes and completes the checks. No release date or crates.io +publication is claimed. This series has been tested with Rust 1.98 and 1.99; +that does not establish compatibility with the declared Rust 1.82 minimum. + ## CLI metadata ```sh