Problem Description
This is an issue related to the running the whisper model with strict open AI compatible connection.
When running FastFlowLM with the ASR pipeline enabled (flm serve --asr 1), the /v1/models endpoint only broadcasts text and vision models. It fails to list the active Whisper/ASR model in the payload.
This creates a routing conflict with strict OpenAI-compatible clients (like TypeWhisper) that rely on /v1/models to populate their UI dropdowns. Because the ASR model is missing, users are forced to select another model. FastFlowLM then receives the audio payload, locks the NPU, attempts to process the audio through the text model, and returns an empty transcription string without throwing an error.
Operating System
Windows 11 10.0.26200
CPU
AMD Ryzen AI 9 HX 370 w/ Radeon 890M
GPU
XDNA 2 NPU, 50 TOPS (AMD Ryzen AI 9 HX 370)
ROCm Version
Ryzen AI NPU (Driver Version: 32.0.203.311)
Installation Method
Other (describe in Additional Information)
Installed ROCm Packages / Versions
Installed ROCm packages / versions
N/A - Running on Windows 11 with Ryzen AI NPU.
No traditional ROCm packages are installed.
- FastFlowLM: Installed via the Windows flm-setup.msi
- AMD NPU Driver: [ 32.0.203.311]
ROCm Component
No response
Steps to Reproduce
Launch FastFlowLM on Windows with the ASR pipeline enabled by running flm serve --asr 1 (default port 52625).
Query the /v1/models endpoint. Observe that only text/vision models (e.g., gemma4-it:12b) are returned in the JSON payload, while the active ASR model is missing.
Open an OpenAI-compatible client (e.g., TypeWhisper) and point its API base URL to http://127.0.0.1:52625/v1.
Attempt to configure the transcription model. Because the ASR model is not broadcast, select an available model from the UI dropdown. It does not allow for typing an arbitrary model in or leaving blank as it is strict.
Initiate a voice dictation. The client sends a POST request with audio data to /v1/audio/transcriptions targeting the selected model.
Check the FastFlowLM terminal logs. Observe that the NPU locks, processes the payload, and releases without throwing an error.
Observe the client UI of typewhisper, which receives an empty transcription string (rather than an API error) and outputs "No speech detected."
(Optional for Linux users) Output of rocminfo --support
rocminfo --support output
Additional Information
I used FastflowLM directly so I don't have details on ROCM installation/version.
Problem Description
This is an issue related to the running the whisper model with strict open AI compatible connection.
When running FastFlowLM with the ASR pipeline enabled (flm serve --asr 1), the /v1/models endpoint only broadcasts text and vision models. It fails to list the active Whisper/ASR model in the payload.
This creates a routing conflict with strict OpenAI-compatible clients (like TypeWhisper) that rely on /v1/models to populate their UI dropdowns. Because the ASR model is missing, users are forced to select another model. FastFlowLM then receives the audio payload, locks the NPU, attempts to process the audio through the text model, and returns an empty transcription string without throwing an error.
Operating System
Windows 11 10.0.26200
CPU
AMD Ryzen AI 9 HX 370 w/ Radeon 890M
GPU
XDNA 2 NPU, 50 TOPS (AMD Ryzen AI 9 HX 370)
ROCm Version
Ryzen AI NPU (Driver Version: 32.0.203.311)
Installation Method
Other (describe in Additional Information)
Installed ROCm Packages / Versions
Installed ROCm packages / versions
ROCm Component
No response
Steps to Reproduce
Launch FastFlowLM on Windows with the ASR pipeline enabled by running flm serve --asr 1 (default port 52625).
Query the /v1/models endpoint. Observe that only text/vision models (e.g., gemma4-it:12b) are returned in the JSON payload, while the active ASR model is missing.
Open an OpenAI-compatible client (e.g., TypeWhisper) and point its API base URL to http://127.0.0.1:52625/v1.
Attempt to configure the transcription model. Because the ASR model is not broadcast, select an available model from the UI dropdown. It does not allow for typing an arbitrary model in or leaving blank as it is strict.
Initiate a voice dictation. The client sends a POST request with audio data to /v1/audio/transcriptions targeting the selected model.
Check the FastFlowLM terminal logs. Observe that the NPU locks, processes the payload, and releases without throwing an error.
Observe the client UI of typewhisper, which receives an empty transcription string (rather than an API error) and outputs "No speech detected."
(Optional for Linux users) Output of rocminfo --support
rocminfo --support output
Additional Information
I used FastflowLM directly so I don't have details on ROCM installation/version.