benchmarking: add Hindi Indic ASR benchmark - #2456
Draft
mohammadaaftabv wants to merge 9 commits into
Draft
mohammadaaftabv wants to merge 9 commits into
mohammadaaftabv wants to merge 9 commits into
Conversation
mohammadaaftabv
force-pushed
the
aaftabv/nmcur-397-indic-asr-benchmark
branch
3 times, most recently
from
September 29, 2026 23:38
450c28f to
6f3eb8a
Compare
Signed-off-by: aaftaabv@gmail.com <aaftaabv@gmail.com>
Signed-off-by: aaftaabv@gmail.com <aaftaabv@gmail.com>
Signed-off-by: aaftaabv@gmail.com <aaftaabv@gmail.com>
Signed-off-by: aaftaabv@gmail.com <aaftaabv@gmail.com>
Signed-off-by: aaftaabv@gmail.com <aaftaabv@gmail.com>
Signed-off-by: aaftaabv@gmail.com <aaftaabv@gmail.com>
Signed-off-by: aaftaabv@gmail.com <aaftaabv@gmail.com>
Signed-off-by: aaftaabv@gmail.com <aaftaabv@gmail.com>
mohammadaaftabv
force-pushed
the
aaftabv/nmcur-397-indic-asr-benchmark
branch
from
September 30, 2026 17:16
7a564ce to
78655a0
Compare
Signed-off-by: aaftaabv@gmail.com <aaftaabv@gmail.com>
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Description
Adds paired Xenna and Ray Data benchmarks for the complete Hindi Indic ASR
processor graph, backed by the pinned open Hugging Face dataset
ketav/parakeet-hindi-asr.The processor/configuration contract is ported from
run_asr_indic_slurm.shat Granary-v2 commit
3931b631b1a9f23215abdb240b647ef6a3a54b34.nkoluguri/integration-testwas used as a processor and engine-profile reference.
This PR intentionally contains no
fern/changes, notests/changes, andno new test files.
What changed
Adds
audio_indic_asr_xennaandaudio_indic_asr_raydata; both use the sameimmutable 216,169-row / 531.7738-hour Hindi cohort and eight primary plus
eight fallback GPU workers.
Applies the 600–900 second wall-clock acceptance gate only to Xenna. Ray Data
uses the same data and correctness gates without a timing requirement.
Executes all eleven stages:
Keeps Indic Canary's TensorRT encoder and TensorRT-LLM decoder. TensorRT-LLM
1.2.1 runs in a separately locked CPython 3.12/CUDA 13 child environment;
Parakeet remains in Curator's parent
audio_tensorrtenvironment and usesplain TensorRT.
Adds the user-facing
audio_canary_trtllmextra, the packaged child lock,isolated authenticated IPC worker, environment validation/provisioning, and
pinned CUDA 13.3 forward-compatibility layer for supported EOS data-center
GPUs. The root lock intentionally contains no
tensorrt-llmpackage.Extends the generic YAML runner so resource mappings work for regular and
composite stages.
Adds an ALM/Qwen-style YAML-first tutorial. It executes the exact benchmark
graph through
nemo_curator/config/run.py; there is no tutorialmain.py.Adds deterministic data staging plus an optional at-least-one-hour manifest
for the local functional gate without changing the canonical full manifest.
Frozen data contract
ketav/parakeet-hindi-asr35376a112c4b79318eeaba0c0dd1b6f1a9bf0ea0train407b58ccb9c74c75a5129e882b1fd000970e082e109adf95a1889592c66964a49f481545c1fe183eeab3a80c1a170215299c333f1cd754f4fab221eebf517c20Runtime contract
Parent audio runtime:
audio_tensorrt= Curatoraudio_cuda12plus TensorRT10.9.0.34Isolated Indic Canary runtime:
The mixed Torch cu128 / CUDA 13 TensorRT-LLM package set matches the
previously successful EOS H100 runtime. This PR's new isolated-process path
still requires current-head GPU execution proof before benchmark submission.
Validation at current head
Exact DCO-signed head:
cab24ed4e2c5ed8ef8d9c40309146b5e18a5fab7(tree
ec221c8a9e795ccca510f612e6da692c25307d86).Completed:
ruff checkandruff format --checkover every changed Python fileuv lock --check(634 packages)uv lock --check(185 packages)Torch, Transformers, and CUDA toolkit at the pinned versions above
nested
pyproject.toml, and nesteduv.lockare packagedincluding both ordinary and composite resource mappings
git diff --checkand secret scanning over all changed filesfull PR diff
Not yet completed:
the Canary engine alone reserves roughly 41 GB before KV-cache allocation,
and this consumer GPU cannot use CUDA forward compatibility. Reducing the
manifest does not reduce this static engine allocation.
requested a successful local one-hour GPU run before EOS, so that gate is
intentionally preserved rather than bypassed.
Local one-hour functional gate
Prepare the pinned full cohort and deterministic subset:
The exact one-GPU tutorial command is documented in
tutorials/audio/indic_asr/README.md; it retains all eleven stages and reducesonly actor concurrency. It requires a compatible high-memory GPU and matching
TensorRT engine bundles.
EOS launch after the local gate
From an authenticated
nemo-cicheckout, the intended submission is exactly:Checklist