Skip to content

speech: Add speech input - #3333

Draft
madcodelife wants to merge 1 commit into
mainfrom
speech-button
Draft

madcodelife wants to merge 1 commit into
mainfrom
speech-button

Conversation

@madcodelife

@madcodelife madcodelife commented Sep 30, 2026 •

Copy link
Copy Markdown
Member

Summary

Speech input for GPUI Component: capture audio, recognize it and hand the text to the application.

  • SpeechState runs a session (Idle → Connecting → Recording → Stopping) and tracks the transcript. It emits SpeechEvent::Partial / Final with the whole transcript so far, plus Started, Cancelled and Error.
  • SpeechButton starts and stops the session. It shows a microphone at rest and a pressed stop glyph while capturing. While the final result is pending it shows a spinner.
  • SpeechWaveform draws the recent input levels.
  • Two seams:
    • SpeechRecognizer / RecognitionSession / SpeechSink plug in any speech service, for example a cloud service over WebSocket.
    • AudioInput / AudioSink plug in any audio source.
  • Recognizer order: the application's recognizer, else the platform's SystemRecognizer (unless .system_fallback(false)), else none. With none, has_recognizer() is false and the button renders nothing, unless show_when_unsupported(true) is set.
Demo.mov

Platforms (speech feature)

Capture System recognizer
macOS cpal (Core Audio) SFSpeechRecognizer, on-device only. A language without offline support is unavailable. Needs NSMicrophoneUsageDescription and NSSpeechRecognitionUsageDescription; without the latter it reports unavailable instead of letting the system terminate the process.
Windows cpal (WASAPI) Windows.Media.SpeechRecognition continuous dictation. Needs the language's speech pack and "Online speech recognition". The recognizer captures from the default device itself, so pushed PCM only drives the waveform.
Linux cpal (ALSA) None. An application recognizer is required. Building needs libasound2-dev.
wasm — Microphone and system recognizer are compiled out.

SpeechAnalyzer (macOS 26) is Swift-only and not reachable through objc2, so macOS uses SFSpeechRecognizer on every version.

Also in this PR

  • Story: DemoRecognizer implements the trait with no service, and the story also shows the system recognizer.
  • Docs: website/component/speech.md and the zh-CN page.
  • examples/speech: a dictation notepad that doubles as a test bench.
    • The sidebar shows platform, input device, recognizer availability and session state, and a session log records every event.
    • cargo run -p speech -- --check prints the checks without opening a window.
    • build.rs links the usage descriptions into the macOS executable, so cargo run can ask for access without an app bundle.
  • Dependencies: cpal 0.15; optional objc2 Speech / AVFAudio / AVFoundation bindings on macOS; WinRT features of the existing windows dependency. All come in only with speech. cpal 0.15 brings its own older windows crate on Windows.
  • Icons: Mic and Square join the default icon set.
  • Linux setup: libasound2-dev is added to script/install-linux.sh and to the installation docs.

Public API

gpui-component (gpui_component::speech)

  • SpeechState — the session state. EventEmitter<SpeechEvent>.
    • new(cx: &mut Context<Self>) -> Self
    • recognizer(self, recognizer: impl SpeechRecognizer) -> Self — use this recognizer instead of the system's.
    • input(self, input: impl AudioInput) -> Self — capture from this input instead of the microphone.
    • system_fallback(self, system_fallback: bool) -> Self — fall back to SystemRecognizer (default true).
    • stop_timeout(self, timeout: Duration) -> Self — how long stop waits for the final result (default 3 s).
    • status(&self) -> SpeechStatus, has_recognizer(&self) -> bool, is_available(&self, cx: &App) -> bool, transcript(&self) -> SharedString, levels(&self) -> impl ExactSizeIterator<Item = f32>
    • start / stop / cancel / toggle(&mut self, cx: &mut Context<Self>)
  • SpeechStatus — Idle, Connecting, Recording, Stopping; is_active(self) -> bool, is_capturing(self) -> bool.
  • SpeechEvent — Started, Partial(SharedString), Final(SharedString), Cancelled, Error(SpeechError).
  • SpeechButton — new(state: &Entity<SpeechState>) -> Self, show_when_unsupported(self, show: bool) -> Self; Sizable, Disableable.
  • SpeechWaveform — new(state: &Entity<SpeechState>) -> Self, bars(self, bars: usize) -> Self; Sizable.
  • trait SpeechRecognizer: 'static, also implemented for Rc<T>:
    • audio_format(&self) -> AudioFormat (default 16 kHz mono)
    • is_available(&self, cx: &App) -> bool (default true)
    • start(&self, sink: SpeechSink, cx: &mut App) -> Result<Box<dyn RecognitionSession>, SpeechError>
  • trait RecognitionSession: 'static — push_audio(&mut self, samples: &[i16], cx: &mut App), finish(&mut self, cx: &mut App). Dropping it cancels the session.
  • SpeechSink (Clone) — ready, hypothesis(text), phrase(text), finish, error(SpeechError). Calls are deferred and ignored once the session is over.
  • trait AudioInput: 'static, also implemented for Rc<T> — start(&self, format: AudioFormat, sink: AudioSink, cx: &mut App) -> Result<Subscription, SpeechError>. Capture runs until the Subscription is dropped.
  • AudioSink (Clone) — push(samples: Vec<i16>, cx), error(SpeechError, cx).
  • AudioFormat — new(sample_rate: u32, channels: u16) -> Self, sample_rate(), channels(), Default (16 kHz mono).
  • SpeechError — PermissionDenied, NoInputDevice, Unsupported, Input(Arc<anyhow::Error>), Recognizer(Arc<anyhow::Error>); input(error), recognizer(error); Display and Error.
  • Microphone (speech, native) — the default input device through cpal. Default, AudioInput.
  • SystemRecognizer (speech, native) — the platform recognizer. new(), locale(self, locale: impl Into<SharedString>) -> Self, Default, SpeechRecognizer.
  • Feature speech.
  • IconName::Mic, IconName::Square join the default component icon set.
  • Locale keys Speech.Start, Speech.Stop, Speech.Unavailable (en, zh-CN, zh-HK, zh-TW).

gpui-kit

  • Feature speech, forwarding gpui-component/speech.

Test plan

  • cargo test -p gpui-component --features speech --lib speech — 11 tests:
    • the state machine: full session, stop timeout, cancel with late results ignored, input error, failing start, unsupported;
    • the button builder;
    • the resampler across chunk boundaries.
  • cargo test -p gpui-kit-assets --test icons (106 default icons).
  • cargo clippy -p gpui-component -p gpui-component-story -p gpui-kit -p speech --features gpui-kit/speech --all-targets -- -D warnings
  • cargo check -p gpui-component without the feature.
  • cargo check / clippy -D warnings for x86_64-pc-windows-msvc with --features speech, cross-checked from macOS.
  • cargo check -p gpui-component-story-web --target wasm32-unknown-unknown.
  • cargo run -p speech -- --check on macOS 27: en-US and zh-CN are available; en-GB, zh-HK and ja-JP are not.
  • The example's window, idle and recording (demo recognizer with generated audio), was captured and reviewed.

Not yet verified

  • A real recognition session on macOS (permission prompts, partial and final results, the pause handling in SFSpeechRecognizer).
  • Anything at runtime on Windows.
  • A Linux build (CI covers it).

…dictation

Add a `speech` module to gpui-component: a session state machine that feeds
captured audio to a pluggable `SpeechRecognizer`, a `SpeechButton` and a
`SpeechWaveform` to render it, and two seams, `SpeechRecognizer` and
`AudioInput`, so applications can bring any speech service or audio source.

With the new `speech` feature, a state without its own recognizer falls back to
the platform's: `SFSpeechRecognizer` on macOS (on-device only) and
`Windows.Media.SpeechRecognition` on Windows. Linux has no system recognizer and
works with an application recognizer. Audio is captured with cpal.

Also adds the gallery story, English and Chinese docs, and the `speech` example,
a dictation notepad that reports what the machine supports.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@madcodelife madcodelife changed the title speech: Add SpeechState, SpeechButton and system recognizers for dictation speech: Add speech input Sep 30, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant