Skip to content

benchmarking: add Hindi Indic ASR benchmark - #2456

Draft
mohammadaaftabv wants to merge 9 commits into
NVIDIA-NeMo:mainfrom
mohammadaaftabv:aaftabv/nmcur-397-indic-asr-benchmark
Draft

mohammadaaftabv wants to merge 9 commits into
NVIDIA-NeMo:mainfrom
mohammadaaftabv:aaftabv/nmcur-397-indic-asr-benchmark

Conversation

@mohammadaaftabv

@mohammadaaftabv mohammadaaftabv commented Sep 29, 2026 •

Copy link
Copy Markdown
Contributor

Description

Adds paired Xenna and Ray Data benchmarks for the complete Hindi Indic ASR
processor graph, backed by the pinned open Hugging Face dataset
ketav/parakeet-hindi-asr.

The processor/configuration contract is ported from
run_asr_indic_slurm.sh
at Granary-v2 commit 3931b631b1a9f23215abdb240b647ef6a3a54b34.
nkoluguri/integration-test
was used as a processor and engine-profile reference.

This PR intentionally contains no fern/ changes, no tests/ changes, and
no new test files
.

What changed

  • Adds audio_indic_asr_xenna and audio_indic_asr_raydata; both use the same
    immutable 216,169-row / 531.7738-hour Hindi cohort and eight primary plus
    eight fallback GPU workers.

  • Applies the 600–900 second wall-clock acceptance gate only to Xenna. Ray Data
    uses the same data and correctness gates without a timing requirement.

  • Executes all eleven stages:

    ManifestReader
      -> PrepareIndicASRInputStage
      -> InferenceIndicCanaryStage
      -> WhisperHallucinationStage
      -> InferenceParakeetStage
      -> WhisperHallucinationStage
      -> SelectBestPredictionStage
      -> RegexSubstitutionStage
      -> AbbreviationConcatStage
      -> GetPairwiseWerStage
      -> ManifestWriterStage
    
  • Keeps Indic Canary's TensorRT encoder and TensorRT-LLM decoder. TensorRT-LLM
    1.2.1 runs in a separately locked CPython 3.12/CUDA 13 child environment;
    Parakeet remains in Curator's parent audio_tensorrt environment and uses
    plain TensorRT.

  • Adds the user-facing audio_canary_trtllm extra, the packaged child lock,
    isolated authenticated IPC worker, environment validation/provisioning, and
    pinned CUDA 13.3 forward-compatibility layer for supported EOS data-center
    GPUs. The root lock intentionally contains no tensorrt-llm package.

  • Extends the generic YAML runner so resource mappings work for regular and
    composite stages.

  • Adds an ALM/Qwen-style YAML-first tutorial. It executes the exact benchmark
    graph through nemo_curator/config/run.py; there is no tutorial main.py.

  • Adds deterministic data staging plus an optional at-least-one-hour manifest
    for the local functional gate without changing the canonical full manifest.

Frozen data contract

  • Hugging Face repo: ketav/parakeet-hindi-asr
  • Revision: 35376a112c4b79318eeaba0c0dd1b6f1a9bf0ea0
  • Split: train
  • License: Apache-2.0
  • Rows: 216,169 unique mono 16 kHz clips
  • Duration: 1,914,385.701 seconds / 531.7738 hours
  • Source manifest SHA-256:
    407b58ccb9c74c75a5129e882b1fd000970e082e109adf95a1889592c66964a4
  • Audio archive SHA-256:
    9f481545c1fe183eeab3a80c1a170215299c333f1cd754f4fab221eebf517c20

Runtime contract

Parent audio runtime:

  • audio_tensorrt = Curator audio_cuda12 plus TensorRT 10.9.0.34
  • Indic Parakeet runs here through TensorRT

Isolated Indic Canary runtime:

  • CPython 3.12
  • TensorRT-LLM 1.2.1
  • TensorRT 10.14.1.48.post1
  • CUDA toolkit 13.3.1
  • Torch 2.9.1+cu128
  • Transformers 4.57.3

The mixed Torch cu128 / CUDA 13 TensorRT-LLM package set matches the
previously successful EOS H100 runtime. This PR's new isolated-process path
still requires current-head GPU execution proof before benchmark submission.

Validation at current head

Exact DCO-signed head: cab24ed4e2c5ed8ef8d9c40309146b5e18a5fab7
(tree ec221c8a9e795ccca510f612e6da692c25307d86).

Completed:

  • ruff check and ruff format --check over every changed Python file
  • root uv lock --check (634 packages)
  • child-runtime uv lock --check (185 packages)
  • Python byte-compilation for every changed Python file
  • real isolated-runtime provisioning and imports of TensorRT-LLM, TensorRT,
    Torch, Transformers, and CUDA toolkit at the pinned versions above
  • wheel build and verification that the worker, installer, runtime helper,
    nested pyproject.toml, and nested uv.lock are packaged
  • BuildKit parse-only validation for both Dockerfiles
  • notebook JSON validation and compilation of every code cell
  • dynamic Hydra construction of the exact eleven-stage tutorial graph,
    including both ordinary and composite resource mappings
  • deterministic one-hour-subset CLI validation, including idempotence
  • git diff --check and secret scanning over all changed files
  • independent audits of the runtime protocol, benchmark/tutorial parity, and
    full PR diff
  • GitHub pre-commit, Ruff, secrets-detector, and DCO checks

Not yet completed:

  • Current-head GPU inference. The available local GPU is a 12 GB RTX 3080 Ti;
    the Canary engine alone reserves roughly 41 GB before KV-cache allocation,
    and this consumer GPU cannot use CUDA forward compatibility. Reducing the
    manifest does not reduce this static engine allocation.
  • Consequently, no EOS benchmark has been submitted for this head. The user
    requested a successful local one-hour GPU run before EOS, so that gate is
    intentionally preserved rather than bypassed.

Local one-hour functional gate

Prepare the pinned full cohort and deterministic subset:

python benchmarking/data_prep/prepare_audio_indic_asr_data.py \
  --output-path "$DATASETS_PATH/audio_indic_asr_parakeet_hindi_531h_35376a11" \
  --cache-dir "$DATASETS_PATH/_hf_cache/audio_indic_asr" \
  --subset-hours 1 \
  --subset-manifest "$DATASETS_PATH/audio_indic_asr_parakeet_hindi_531h_35376a11/manifest-1h.jsonl"

The exact one-GPU tutorial command is documented in
tutorials/audio/indic_asr/README.md; it retains all eleven stages and reduces
only actor concurrency. It requires a compatible high-memory GPU and matching
TensorRT engine bundles.

EOS launch after the local gate

From an authenticated nemo-ci checkout, the intended submission is exactly:

cd nemo-ci
python3 curator/run_benchmarks.py --pr 2456 \
  --entries-exact audio_indic_asr_xenna,audio_indic_asr_raydata \
  --slack-channel-id <your ID>

Checklist

  • I am familiar with the Contributing Guide.
  • New or existing tests cover these changes. (Test-file changes are intentionally outside this PR's scope.)
  • Benchmark setup and tutorial documentation are included.

@copy-pr-bot

copy-pr-bot Bot commented Sep 29, 2026

Copy link
Copy Markdown

This pull request requires additional validation before any workflows can run on NVIDIA's runners.

Pull request vetters can view their responsibilities here.

Contributors can view more details about this message here.

@mohammadaaftabv
mohammadaaftabv force-pushed the aaftabv/nmcur-397-indic-asr-benchmark branch 3 times, most recently from 450c28f to 6f3eb8a Compare September 29, 2026 23:38
Signed-off-by: aaftaabv@gmail.com <aaftaabv@gmail.com>
Signed-off-by: aaftaabv@gmail.com <aaftaabv@gmail.com>
Signed-off-by: aaftaabv@gmail.com <aaftaabv@gmail.com>
Signed-off-by: aaftaabv@gmail.com <aaftaabv@gmail.com>
Signed-off-by: aaftaabv@gmail.com <aaftaabv@gmail.com>
Signed-off-by: aaftaabv@gmail.com <aaftaabv@gmail.com>
Signed-off-by: aaftaabv@gmail.com <aaftaabv@gmail.com>
Signed-off-by: aaftaabv@gmail.com <aaftaabv@gmail.com>
@mohammadaaftabv
mohammadaaftabv force-pushed the aaftabv/nmcur-397-indic-asr-benchmark branch from 7a564ce to 78655a0 Compare September 30, 2026 17:16
Signed-off-by: aaftaabv@gmail.com <aaftaabv@gmail.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant