On-device AI engine

Put AI models on your computer

Transcribe audio and video, generate speech, and create transcripts and subtitles with precise timing on your own computer.

macOS 14.0+ · Apple Silicon — Windows 10/11 x64

Start freeNo credit card needed
local-transcribe.mov
Transcript preview

Audio understanding and generation

Transcription, speech, and precise timing. All on your device.

Accurate transcription

Lattice-2 Flash / Pro supports more than 40 languages.

Speech generation

Local streaming synthesis starts speaking as audio is generated.

Precise timing and subtitles

Word-level timestamps with SRT, VTT, and karaoke subtitle export.

Voice and text chat

Talk to a local model by voice or text, with personas and live interpretation.

Real desktop app

EdgeSpeak runs right on your computer

Audio, video, and transcripts in one workspace
EdgeSpeak desktop Transcribe screen showing the transcription workspace, local model status, transcript content, playback controls, and recent files.
Talk to a local AI model by voice or text

Chat runs on models on your computer. Switch between Voice and Text, pick a persona, and every turn is saved to the record panel.

EdgeSpeak desktop Chat screen in Voice mode with a start orb, a persona selector, a chime toggle, and a record panel that keeps each turn of the conversation.
Ready-made personas, including Buddy

Pick a preset such as Buddy, Voice companion, Language tutor or Roleplay, or write your own prompt. It applies from the next reply.

EdgeSpeak desktop Chat screen with the persona list open, showing presets Voice companion, Emotional support, Buddy, Interpreter, Language tutor and Roleplay, and a New prompt button.
Live interpretation between two languages

Switch the persona to Interpreter, choose the language pair and a style, then speak. Replies stay in the record panel.

EdgeSpeak desktop Chat screen in Interpreter mode, with a language pair selector from Chinese to English and a Natural style option.
Word-by-word karaoke subtitles, burned into video
A video frame with English karaoke subtitles exported by EdgeSpeak, where the words being spoken are highlighted.
Make every word light up with your audio

Choose ASS subtitles and turn on karaoke in the export panel. Pick a highlight preset, adjust fonts, colors and canvas size, then preview the style and export. Word-by-word highlighting requires word timestamps; run alignment first if needed.

Make every word light up with your audio
Choose the right model in EdgeSpeak
EdgeSpeak desktop Models screen listing local Lattice-2 Lite, Flash and Pro transcription models.
Let other tools use EdgeSpeak
EdgeSpeak desktop Gateway screen showing how tools on the same computer can use the local speech engine.

CLI & local gateway

Install EdgeSpeak CLI

The CLI includes its own local runtime. Transcribe, generate text, and synthesize speech without opening the desktop app, or connect scripts, MCP, and agents.

Install from Terminal on macOS or Linux, or from PowerShell on Windows. On Linux the installer picks the CUDA build when it finds an NVIDIA driver or a CUDA environment and otherwise the CPU build; add --device cuda or --device cpu to choose. Windows offers CUDA as an opt-in install.

macOS · Apple SiliconWindows x64 · CPU / CUDALinux x86_64 · CPU / CUDA
Read the CLI docs

macOS / Linux · Terminal

curl -fsSL https://edgespeak.com/install.sh | sh

Windows · PowerShell

irm https://edgespeak.com/install.ps1 | iex
edgespeak-cli login
edgespeak-cli transcribe media.mp4 -o media.srt
edgespeak-cli generate "Summarize this transcript" --model Qwen/Qwen3.5-4B

Local AI gateway

One address for your agent workflows

Connect tools such as OpenClaw, Claude Code, and Cursor through OpenAI-compatible APIs. Transcription, speech synthesis, Chat Completions, and Realtime share one local gateway.

OpenAI APICLI / MCP / AgentEdgeSpeak API

Realtime API: live voice conversations · Preview

Talk in real time with interruption support. Load a local language model before starting.

ws://127.0.0.1:1117/v1/realtime
POST /v1/audio/transcriptionslocalhost:1117
curl http://127.0.0.1:1117/v1/audio/transcriptions \
  -H "Authorization: Bearer sk-edgespeak-..." \
  -F file=@meeting.m4a \
  -F model="lattice-2-flash"

Agent Skills, same engine

EdgeSpeak Skills teach Claude Code, Cursor, and other Skills-capable agents to transcribe on-device through the CLI. Open source on GitHub.

npx skills add lattifai/EdgeSpeak

LOCAL LM + VLM · 0.6B — 27B

LLMs and VLMs run locally, too

Built-in Qwen and Gemma models span 0.6B–27B. Process text and images through Chat Completions / Responses while inputs and generated results stay on your device.

LLM / VLM

Text + images · OpenAI-compatible · local-first

Import GGUF models

Import compatible GGUF models. Vision requires a matching projector file; model choice depends on your memory and hardware.

Generate with a local model
Qwen/Qwen3-0.6B0.6B0.4 GB
Qwen/Qwen3.5-0.8B0.8B0.5 GB
Qwen/Qwen3.5-2B2B1.3 GB
google/gemma-4-E2B-it2B3.1 GB
Qwen/Qwen3.5-4BDefault4B2.7 GB
google/gemma-4-E4B-it4B5.0 GB
Qwen/Qwen3.5-9B9B5.7 GB
google/gemma-4-12B-it12B7.1 GB
Qwen/Qwen3.6-27B27B16.8 GB

Text + images · OpenAI-compatible · local-first

Choose by workflow

Local software, cloud APIs, or a model you maintain

The right option depends on where media can go, how much workflow you want ready-made, and who should maintain the speech stack.

Choose by workflowEdgeSpeakCloud transcriptionSelf-managed local model
Source mediaStays on this deviceUploaded to a remote serviceStays on the machine you configure
Review workflowDesktop playback, transcript, and exportDepends on the providerYou build the review surface
Tools and agentsLocal CLI and OpenAI-compatible gatewayHosted APIYou build and maintain the integration
OperationsInstall the app and local modelsManage an account, API keys, and network accessMaintain the runtime, models, and dependencies
Read the full comparison

Privacy

Local-first AI processing.

Audio, video, images, conversations, and generated transcripts stay on your device unless you export, upload, or share them yourself.

Privacy

Source media and transcripts stay on your device, ready to export or continue through local tools.

Results keep moving

Export text or subtitles, or let another tool continue from the finished transcript.

FAQ

Frequently asked questions

On-device processing, supported platforms, models, and plans — answered before you buy.

See all questions
01Does EdgeSpeak upload source media?

Imported media and generated transcripts stay on your device unless you export, upload, or share them yourself.

02Which desktop platforms are available?

EdgeSpeak is available for Apple Silicon Macs with macOS 14.0 or later. A Windows 10/11 x64 version is also available and supports in-app automatic updates.

03What is the difference between Flash and Pro?

Flash is tuned for faster everyday transcription. Pro raises the accuracy ceiling for harder audio and uses more local resources.

04Can CLI tools and AI agents use EdgeSpeak?

Yes. The bundled CLI and local OpenAI-compatible gateway let trusted tools on the same computer call the local speech engine.

Pricing

Choose a plan that fits how much you use.

Community feedback

Shape EdgeSpeak with real workflows

Share real workflows, agent integration needs, and product ideas on Discord.

  • Real workflows
  • Agent / API
  • Product ideas

Your feedback goes directly into the product.

On-device AI engine

Put AI models on your computer

Transcribe audio and video, generate speech, and create transcripts and subtitles with precise timing on your own computer.