Open Source Speech-to-Text

Cadence

A cross-platform desktop application that provides simple, privacy-focused speech transcription. Press a shortcut, speak, and have your words appear in any text field. This happens entirely on your own computer without sending any information to the cloud.

Windows x64 (Unsigned Beta) · macOS Build Coming Soon

Application Interface

Designed for clarity & control.

A clean, minimal interface that stays out of your way until you need it. Configure models, shortcuts, sound inputs, and paste rules in seconds.

Cadence Desktop v0.10.0

Transcription Models

Select or download AI models tailored to your language and performance needs. Includes Parakeet V3, Whisper, Nemotron Streaming, and Canary Flash.

Cadence App - Transcription Models Selection
Why Cadence?

Built to fill the gap for truly open speech-to-text.

Cadence was created to give everyone a lightweight, highly accurate, and extensible voice typing tool without proprietary lock-in.

100% Free & Open

Accessibility tooling belongs in everyone's hands, not behind a paywall. Cadence is free software engineered for daily productivity.

Private by Design

Your voice stays strictly on your computer. Get instant transcriptions without sending audio packets over the network or to cloud servers.

Simple & Focused

One tool, one job. Transcribe what you say and place it into whatever text field has focus — IDEs, web browsers, notes, or email.

Workflow

How Cadence Works

Seamlessly integrates into your existing workflow across any application.

01

Press Shortcut

Press a configurable hotkey (e.g. Ctrl + Space) or hold down push-to-talk mode.

02

Speak Your Mind

Speak naturally while the shortcut is active. Audio recording handles low latency capturing seamlessly.

03

Release to Process

Release the key. Cadence filters silence with Silero VAD and transcribes with your chosen local AI model.

04

Direct Text Paste

Transcribed text is pasted immediately into your active window or app with zero manual copying needed.

Local Processing Engine

Powered by state-of-the-art local models.

The entire transcription process happens on your CPU or GPU using industry-standard machine learning primitives.

Voice Activity Detection (VAD)

Silence and background noise are intelligently filtered using Silero VAD before inference begins. This speeds up processing and eliminates unwanted background noise artifacts.

Silero VAD · Zero background noise hallucination · Fast pre-trimming

Flexible Model Choices

Choose the model that fits your hardware:

  • Whisper Models: Small, Medium, Turbo, Large with GPU acceleration.
  • Parakeet V3: CPU-optimized model with exceptional speed and automatic language detection.
Specifications

Technical Overview

Platform Support

Operating SystemsWindows, macOS, Linux
LicenseOpen Source
PriceFree

Models & Acceleration

Whisper ModelsSmall / Medium / Turbo / Large
Parakeet ModelParakeet V3 (CPU Optimized)
GPU SupportCUDA / Metal / Vulkan

Privacy & Audio

Cloud EgressZero
VAD ProcessingSilero VAD
Hotkey TriggersShortcut & Push-to-Talk

Ready to type at the speed of speech?

Download Cadence for Windows, macOS, or Linux. Free, private, and open source speech-to-text.