AI product Open source

jev-voice-browser

jev-voice-browser is a Node application that controls a headed Chromium browser by voice or typed commands. It streams partial speech transcripts from the browser's Web Speech API to a Node server, which uses TypeSafe's Jev model to classify the intent, select a page element or site, extract candidate text and URL spans, determine whether the command is complete or addressed to the browser, and assess whether an action is destructive. Playwright then executes navigation, searches, typing, clicking, scrolling, history, and tab actions.

View repository Mentioned in 1 video ↓

Overview

The application sends the page URL, title, detected site, a compact snapshot of visible elements, and recent actions to Jev for each transcript update. Code—not the model—constructs URLs, copies text verbatim, and applies policy thresholds to decide whether to act, wait, ignore, request confirmation, or show numbered alternatives. Recent action context supports corrections such as reversing an action, rejecting a selected result, or choosing another candidate.

It requires Node.js, Chrome or Edge for microphone input, Playwright's Chromium installation, and a TypeSafe API key. The server can launch a persistent-profile Chromium window or attach to an existing browser through CDP; it also supports headless operation and typed commands when no microphone is available.

What jev-voice-browser is used for

1 use taken from transcripts — each links to the moment in the video.

  • Controls a Chromium window through spoken commands. It uses partial speech transcripts and typed intent questions to select targets and actions, allowing corrections such as choosing a different result.

Videos mentioning jev-voice-browser

1 in the library.