Shadoword
Read the source
An abstract spectrogram: hundreds of fine vertical frequency bins burning in magenta and cyan on black, densest at the right and dissolving to darkness at the left, over a solid rail of low-frequency light along the bottom. One column at the centre is perfectly black, carrying no signal at all.

Speech to textlocal by defaultconnected by choice

Run Whisper here, connect to your Shadoword API, or opt into OpenRouter. You choose the path.

Linux releases published · source available.

Read the source

Push to talk

F2

Hold to capture

Default shortcut. Rebindable, or switch to toggle.

Runtime modes

Where transcription runs — select a mode to read what it does

Signal path

Nothing istranscribed speculatively.

  1. Capture

    Microphone input is recorded and voice activity is detected, so silence never reaches a model.

  2. Segment

    Audio accumulates into a segment. Duration, queue length, and idle time are all bounded.

  3. Commit

    A commit is what triggers inference. Until one arrives the segment simply sits there.

  4. Deliver

    The transcript arrives as text at your cursor, or as a payload to the process that asked for it.

You choose the model and the silicon

Model catalog

Whisper models the desktop client ships in its catalog, with download sizes
ModelDownload flagSize
Large v3 TurboDefaultRecommended balance of speed and accuracy.turbo1.51 GiB
Large v3Largest and most accurate catalog model.large-v32.88 GiB
MediumHigh-accuracy multilingual model.medium1.43 GiB
SmallBalanced model for lower-memory systems.small465 MiB
BaseFast model with improved accuracy over Tiny.base141 MiB
TinySmallest and fastest catalog model.tiny74 MiB

Accelerator

Build-time accelerator backends and the cargo feature each one needs
BackendCargo feature
CPUdefaultPlain cargo build. No feature flag.
Vulkanwhisper-vulkanOpt-in cargo feature.
CUDAwhisper-cudaOpt-in cargo feature.

The daemon is yours to host

Authenticated, and pointed wherever you point it. Nothing below is illustrative — it is the router and the socket as shipped.

HTTP surface

  • GET/healthPublic.
  • GET/v1/status
  • GET/v1/overview
  • GET/v1/config
  • PUT/v1/configRestricted DTO.
  • POST/v1/transcribe-wavRaw WAV body, 64 MiB cap.
  • GET/v1/streamWebSocket. Opus packets.
  • GET/v1/models
  • POST/v1/models/{id}/select
  • POST/v1/downloadsExplicit catalog jobs only.
  • GET/v1/downloads/{id}

Streaming, in order

  1. send{"type":"Start","sample_rate":48000,"channels":1,"protocol_version":2}
  2. recvStartedflow_id, credit
  3. send<binary>raw Opus packets
  4. sendCommitSegmentsegment_index required
  5. recvAcceptedoutstanding, remaining_credit
  6. recvPartialone per committed segment
  7. sendFinish
  8. recvDoneexactly once

Opus in, credit granted at Started and returned at Accepted, one partial per committed segment, one Done at the end.

Build it and run it

Desktop — build and run

cd crates/shadoword-desktop
bun install
bun run tauri dev

Desktop — with Vulkan

nix develop
bun run tauri dev -- --features whisper-vulkan

Daemon — run locally

cargo run -p shadoword-api
# add --features whisper-cuda inside nix develop .#cuda

Daemon — install a model first

cargo run -p shadoword-api -- \
  --download-model turbo

Not yet measured

There is no published latency figure, because no run of the bench corpus has been recorded. When there is one it will name the hardware it ran on.

Corpus
bench_corpus/clip_{10,15,20,30}s.wav
Model
Large v3 Turbo (turbo)