speech: Add speech input - #3333
Draft
madcodelife wants to merge 1 commit into
Draft
madcodelife wants to merge 1 commit into
madcodelife wants to merge 1 commit into
Conversation
…dictation Add a `speech` module to gpui-component: a session state machine that feeds captured audio to a pluggable `SpeechRecognizer`, a `SpeechButton` and a `SpeechWaveform` to render it, and two seams, `SpeechRecognizer` and `AudioInput`, so applications can bring any speech service or audio source. With the new `speech` feature, a state without its own recognizer falls back to the platform's: `SFSpeechRecognizer` on macOS (on-device only) and `Windows.Media.SpeechRecognition` on Windows. Linux has no system recognizer and works with an application recognizer. Audio is captured with cpal. Also adds the gallery story, English and Chinese docs, and the `speech` example, a dictation notepad that reports what the machine supports. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
SpeechState, SpeechButton and system recognizers for dictation
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Speech input for GPUI Component: capture audio, recognize it and hand the text to the application.
SpeechStateruns a session (Idle → Connecting → Recording → Stopping) and tracks the transcript. It emitsSpeechEvent::Partial/Finalwith the whole transcript so far, plusStarted,CancelledandError.SpeechButtonstarts and stops the session. It shows a microphone at rest and a pressed stop glyph while capturing. While the final result is pending it shows a spinner.SpeechWaveformdraws the recent input levels.SpeechRecognizer/RecognitionSession/SpeechSinkplug in any speech service, for example a cloud service over WebSocket.AudioInput/AudioSinkplug in any audio source.SystemRecognizer(unless.system_fallback(false)), else none. With none,has_recognizer()is false and the button renders nothing, unlessshow_when_unsupported(true)is set.Demo.mov
Platforms (
speechfeature)SFSpeechRecognizer, on-device only. A language without offline support is unavailable. NeedsNSMicrophoneUsageDescriptionandNSSpeechRecognitionUsageDescription; without the latter it reports unavailable instead of letting the system terminate the process.Windows.Media.SpeechRecognitioncontinuous dictation. Needs the language's speech pack and "Online speech recognition". The recognizer captures from the default device itself, so pushed PCM only drives the waveform.libasound2-dev.SpeechAnalyzer(macOS 26) is Swift-only and not reachable through objc2, so macOS usesSFSpeechRecognizeron every version.Also in this PR
DemoRecognizerimplements the trait with no service, and the story also shows the system recognizer.website/component/speech.mdand the zh-CN page.examples/speech: a dictation notepad that doubles as a test bench.cargo run -p speech -- --checkprints the checks without opening a window.build.rslinks the usage descriptions into the macOS executable, socargo runcan ask for access without an app bundle.windowsdependency. All come in only withspeech. cpal 0.15 brings its own olderwindowscrate on Windows.MicandSquarejoin the default icon set.libasound2-devis added toscript/install-linux.shand to the installation docs.Public API
gpui-component (
gpui_component::speech)SpeechState— the session state.EventEmitter<SpeechEvent>.new(cx: &mut Context<Self>) -> Selfrecognizer(self, recognizer: impl SpeechRecognizer) -> Self— use this recognizer instead of the system's.input(self, input: impl AudioInput) -> Self— capture from this input instead of the microphone.system_fallback(self, system_fallback: bool) -> Self— fall back toSystemRecognizer(defaulttrue).stop_timeout(self, timeout: Duration) -> Self— how longstopwaits for the final result (default 3 s).status(&self) -> SpeechStatus,has_recognizer(&self) -> bool,is_available(&self, cx: &App) -> bool,transcript(&self) -> SharedString,levels(&self) -> impl ExactSizeIterator<Item = f32>start/stop/cancel/toggle(&mut self, cx: &mut Context<Self>)SpeechStatus—Idle,Connecting,Recording,Stopping;is_active(self) -> bool,is_capturing(self) -> bool.SpeechEvent—Started,Partial(SharedString),Final(SharedString),Cancelled,Error(SpeechError).SpeechButton—new(state: &Entity<SpeechState>) -> Self,show_when_unsupported(self, show: bool) -> Self;Sizable,Disableable.SpeechWaveform—new(state: &Entity<SpeechState>) -> Self,bars(self, bars: usize) -> Self;Sizable.trait SpeechRecognizer: 'static, also implemented forRc<T>:audio_format(&self) -> AudioFormat(default 16 kHz mono)is_available(&self, cx: &App) -> bool(defaulttrue)start(&self, sink: SpeechSink, cx: &mut App) -> Result<Box<dyn RecognitionSession>, SpeechError>trait RecognitionSession: 'static—push_audio(&mut self, samples: &[i16], cx: &mut App),finish(&mut self, cx: &mut App). Dropping it cancels the session.SpeechSink(Clone) —ready,hypothesis(text),phrase(text),finish,error(SpeechError). Calls are deferred and ignored once the session is over.trait AudioInput: 'static, also implemented forRc<T>—start(&self, format: AudioFormat, sink: AudioSink, cx: &mut App) -> Result<Subscription, SpeechError>. Capture runs until theSubscriptionis dropped.AudioSink(Clone) —push(samples: Vec<i16>, cx),error(SpeechError, cx).AudioFormat—new(sample_rate: u32, channels: u16) -> Self,sample_rate(),channels(),Default(16 kHz mono).SpeechError—PermissionDenied,NoInputDevice,Unsupported,Input(Arc<anyhow::Error>),Recognizer(Arc<anyhow::Error>);input(error),recognizer(error);DisplayandError.Microphone(speech, native) — the default input device through cpal.Default,AudioInput.SystemRecognizer(speech, native) — the platform recognizer.new(),locale(self, locale: impl Into<SharedString>) -> Self,Default,SpeechRecognizer.speech.IconName::Mic,IconName::Squarejoin the default component icon set.Speech.Start,Speech.Stop,Speech.Unavailable(en, zh-CN, zh-HK, zh-TW).gpui-kit
speech, forwardinggpui-component/speech.Test plan
cargo test -p gpui-component --features speech --lib speech— 11 tests:cargo test -p gpui-kit-assets --test icons(106 default icons).cargo clippy -p gpui-component -p gpui-component-story -p gpui-kit -p speech --features gpui-kit/speech --all-targets -- -D warningscargo check -p gpui-componentwithout the feature.cargo check/clippy -D warningsforx86_64-pc-windows-msvcwith--features speech, cross-checked from macOS.cargo check -p gpui-component-story-web --target wasm32-unknown-unknown.cargo run -p speech -- --checkon macOS 27: en-US and zh-CN are available; en-GB, zh-HK and ja-JP are not.Not yet verified
SFSpeechRecognizer).