Adds visual character to desktop voice conversations by listening to a supported application's playback process and switching imported VRM models between idle and speaking animations. It supports custom VRMA actions and a local MCP server, while not saving or transmitting audio.