AI product Open source
Lipflow is an open-source, cross-platform silent lip-reading dictation tool that uses a webcam to turn mouthed words into text pasted at the cursor. It runs locally on macOS, Windows, and Linux: face landmarks are tracked live, 96×96 grayscale mouth crops are sampled at 25 fps, and an Auto-AVSR visual-speech-recognition pipeline combines a 3D-convolutional ResNet front end, Conformer encoder, Transformer decoder, CTC scoring, and a subword language model with beam search. Users can train the language model and lip reader on their phrasing and face, import local Wispr Flow dictation history, and optionally apply cleanup through offline rules, an on-device Qwen model, Claude, Codex CLI, or Ollama. Whisper mode combines lip and audio input with an audio-visual model while the hotkey is held. The repository is MIT-licensed, although its LRS3-trained model weights are identified as for non-commercial research use.
1 use taken from transcripts — each links to the moment in the video.
Lets users type by silently mouthing words toward a webcam and pastes the recognized text at the cursor. It runs on macOS, Windows, and Linux, supports personal training, optional text cleanup, and a lip-and-audio whisper mode.
1 in the library.