1,038 tools and products — trending open source, and what gets used in AI and other work.
Pipecat is an open-source Python framework maintained by Daily and the community for building real-time voice and multimodal conversational agents. It orchestrates audio and video, transports such as WebSockets and WebRTC, speech recognition and text-to-speech services, speech preprocessing, turn detection, model calls, tool calls, and streamed audio through composable conversation pipelines. Pipelines can operate as single agents or as multi-agent systems whose specialists hand off, fan out in parallel, or coordinate over a shared bus locally or across processes and machines. The project also provides a CLI for scaffolding, monitoring, and deploying agents, along with client SDKs and related tools for structured conversations and pipeline debugging.
Smallest AI is a voice-AI platform offering text-to-speech, speech-to-text, speech-to-speech, and orchestrated voice-agent models for real-time deployments.
Hydra is a speech-to-speech AI model from Smallest AI for asynchronous speech input and output, with support for text and tool calls.
Deepgram is a speech AI platform providing Speech-to-Text, Text-to-Speech, and Voice Agent APIs. Its services support real-time transcription and voice applications, including the speech transcription component in deployment architectures.
Gradient is a framework for training research agents with reinforcement learning. It exposes the agents’ searches, citations, tool calls, and evaluation results.
Ambient Context is a macOS menu bar app that records the text of the focused window for use by an LLM or other agent. Through the macOS accessibility tree, it reads window text every few seconds and appends deduplicated blocks to one plain Markdown file per day, including the document path or URL and an AGENTS.md file describing the format. It does not use screenshots or video and, in the current build, makes no network calls, accounts, servers, telemetry, or bundled model; password fields, password-manager and private-browsing windows are excluded, and credentials, API keys, and card-shaped numbers are scrubbed before writing. The project is early and unsigned, requires macOS 14 or later on Apple Silicon, and currently must be built from source with Node, Rust, and Xcode Command Line Tools.
System Atlas is an agent skill from Inkboard that turns architecture discussions into an explorable isometric atlas. A single data file serves as the source for an interactive map and a generated SYSTEM.md text view, keeping structures, execution flows, design decisions, and open questions synchronized. The map supports hoverable structures, pinned views, drill-down into execution steps, pan and zoom, progressive-disclosure chapters, and inspectable data packets showing routes and representative JSON payloads. The generated SYSTEM.md includes a decisions table with ADR links, structure descriptions and steps, flow tables, and an index of questions tracked by stable IDs and states such as open, resolved, or routed. The skill can be installed with `npx skills add inkboard/system-atlas`. It generates a self-contained HTML map with no build step or runtime dependencies.
agenttrail is a local, open-source observability layer for AI coding agents. It watches plans, tool calls, file changes, and progress from Claude Code, OpenAI Codex, Cursor, or other agents that edit files, then presents them as a live project map with repository components, dependency arrows, progress, and working, blocked, or completed states. It compares declared agent intent with observed filesystem activity, including renewed changes to supposedly finished work, and displays the current run, task list, streaming tool activity, elapsed time, recent calls, session plans, and a live repository tree. The tool runs locally through commands such as `npx agenttrail`, with no account, global installation, or telemetry. `npx agenttrail init` adds the agenttrail convention to CLAUDE.md and AGENTS.md, creates a starter PLAN.md, and installs additive local Claude Code hooks; the resulting plan and file activity are used to generate a map of roughly 5–9 repository components with dependencies and verifiable statuses. It supports one daemon per repository, a shared board tab switcher, restart via `npx agenttrail up`, and login autostart through `npx agenttrail autostart`.
backpass is a local-first command-line tool that analyzes coding-agent session transcripts and proposes evidence-backed edits to an agent memory file such as AGENTS.md or CLAUDE.md, plus project skills. It treats the memory file as weights, agent sessions as forward passes, and their on-disk transcripts as a loss signal, then collects sessions for the current repository, distills failures, aggregates repeated instruction changes, and produces diffs and skill extractions. The tool reads transcript stores from seven agent harnesses directly from disk, requires evidence from at least two independent sessions for a new instruction, and limits each run to at most five edits. Proposed changes include verbatim session quotes; analysis never writes files, while `backpass apply` presents each edit for human acceptance or rejection before applying it. It is distributed through npm or npx, requires Node.js 22.5 or later and `acpx`, uses no API keys of its own, and sends no transcripts to a separate service; obvious secrets are redacted before model calls through an already authenticated harness.
Sentio is an email inbox API and multi-tenant mail server for AI agents, developed by Truespar. It gives each agent a real email address, authenticates and scans inbound SMTP mail, scores and routes it, then delivers the message as a structured webhook; agents send threaded replies through a REST API, which signs outbound mail with DKIM, queues it, and delivers it over SMTP. Tenant isolation covers domains, mailboxes, API keys, rate limits, suppression lists, sending reputation, and spam profiles. The Rust service implements inbound and outbound mail infrastructure including DKIM, SPF, DMARC, ARC, MTA-STS, DANE, and three-tier anti-spam, and includes an MCP server for exposing email as native agent tools. The repository provides Docker deployment with PostgreSQL, Redis, NATS/JetStream, MinIO, ClamAV, and rspamd, plus an API reference and testing UI.
OwnMem is a repository-owned memory system for AI coding agents. It stores project knowledge as reviewable Markdown in a .ownmem/ directory, shared through Git so it can be cloned, reviewed, and rolled back across Claude Code, Codex, Cursor, Gemini CLI, Grok CLI, and other hosts. Its local recall pipeline compiles schema, graph, lifecycle, and evidence checks into an immutable, content-addressed snapshot. Five deterministic candidate lanes—exact matching, BM25F, n-gram, fuzzy, and graph retrieval—are fused locally; embeddings are optional and remain disabled by default until local evaluation supports them. Four delivery gates assess relevance, epistemic validity, task applicability, and action risk, allowing normal delivery, advisory output, quarantine, or abstention. OwnMem separates repository memory from trust receipts and uses quotas, duplicate gates, lifecycle rules, audits, and reversible changes to bound unattended evolution. Its rules allow automation to promote only replay-proven, quota-bounded R0 retrieval metadata, while prose, policy, and higher-risk changes become review material. The repository describes the project as local, deterministic, git-native, evidence-governed, and licensed under Apache-2.0.
MonoCode is a desktop GUI for running coding-agent command-line interfaces in tabbed sessions. Each tab represents a session, and a shared composer provides the prompt input while the selected CLI uses the user's own logged-in subscription; MonoCode does not sell tokens or act as a model provider. It supports Claude Code, Codex, Cursor CLI, OpenCode, Pi, omp, and fx when they are installed and authenticated. The application supports macOS and Linux, with an Apple Silicon macOS disk-image distribution and source builds requiring Node.js 20 or later and a current stable Rust toolchain. It is licensed under the MIT License and is described by its repository as an early project that may contain bugs.
kimodo.cpp is a C++ and GGML implementation of NVIDIA's Kimodo text-to-motion model for local CPU or Vulkan inference. It accepts UTF-8 prompts or precomputed 4096-value F32 LLM2Vec embeddings and generates root translations plus local XYZW rotations for several control skeletons, including SMPL-X, SOMA, and Unitree G1 formats. The project includes GGUF model loading with validation, Safetensors conversion, DDIM sampling, conditioned multi-prompt transitions, C and C++ APIs, CPU/Vulkan parity tests, skeleton-only GLB export, and a local browser demo. Its native API returns the joints predicted by the selected model; features such as general constraint input, SOMA's 77-joint expansion, skinned-mesh GLB export, and quantized models are not implemented.
open-sheet is a spreadsheet framework for coding agents that represents workbook models as React and TypeScript source. Instead of writing fragile A1-style cell addresses, models can use named references such as `ref('pl').column('revenue')`, which open-sheet resolves at compile time while handling cell addressing, formula references, recalculation, and formula-tree previews. It exports live formulas to XLSX, as well as CSV, HTML, and PDF formats, and reports cells it cannot compute as `#NOT_EVALUATED` rather than emitting a plausible number. The project provides an `npx @open-sheet/cli` initializer, is distributed under the MIT license, and is hosted at open-sheet.dev.
Halofy is an open-source governance layer between organizational knowledge and AI agents. It resolves actor identity, roles, namespaces, and source scope on the server; enforces access policies; and records provenance, supersedence, append-only audit history, policy decisions, and signed erasure certificates for reads and writes. The repository exposes governed memory and context operations over MCP and HTTP, including writing, reading, searching, assembling, faulting, statistics, forgetting, policy, sharing, pinning, export, and manifest operations. Postgres with pgvector is the operational authority, while retrieval engines connect through an ACL-scoped read-only driver interface; encrypted, Git-versioned knowledge provides a cold tier, and embedded PGlite is the zero-setup local default. It includes connectors for filesystem, Postgres, Obsidian, and manual CSV import, plus an offline demo and hermetic test lane using deterministic stub components.
spec-ptc is a Python library and demo implementing speculative programmatic tool calling (sPTC) for code-based tool-use harnesses such as Recursive Language Models (RLMs) and CodeAct. While a language model streams code for a REPL call, it identifies closed tool-call statements in the partial code, queues eligible calls as futures, and runs them in a shadow REPL so submodel or tool execution overlaps with continued code generation. When the completed code executes, speculative calls return stored results while non-speculatable or missed calls run normally; tools can be registered with decorators, including controls for speculation and purity. The repository also provides an out-of-process JSON-lines daemon with turn, token-feed, resolution, and cleanup operations, plus wrappers for Claude Code, OpenCode, and Pi-mono.
Kittl is an AI-first online design platform for creating graphics, apparel designs, and other visual assets. It combines templates, AI image generation, editing tools, mockups, and curated design assets, and is used in template-based print-on-demand workflows.
Flax NNX is a neural-network API for JAX developed by the Flax team in collaboration with the JAX team. Released in 2024, it uses Python reference semantics so models can be expressed as regular Python objects with reference sharing and mutability, supporting inspection, debugging, and analysis. Its API includes layers and modules such as Linear, convolution, normalization, attention, recurrent cells, and dropout, along with patterns for replicated training, serialization, and checkpointing. In distributed training, it can manage model state while JAX handles computation across a device mesh.
Orbax is a Google-developed collection of checkpointing and persistence utilities for JAX models and training state. Its checkpointing API saves and restores pytrees such as model weights and optimizer state, with support for asynchronous checkpointing, standard and custom types, and flexible storage formats. It is distributed through domain-specific PyPI packages including orbax-checkpoint and orbax-export, and is used by JAX-based machine-learning frameworks and model implementations.
StableHLO is a backward-compatible operation set for high-level machine-learning operations and a portability layer between ML frameworks and compilers. Frameworks such as TensorFlow, JAX, and PyTorch can produce StableHLO programs, while compilers such as XLA and IREE can consume them; the programs can also be used for on-device execution through Google AI Edge. Based on the MHLO dialect, StableHLO adds serialization and versioning through MLIR bytecode, with backward- and forward-compatibility guarantees. The project includes the StableHLO specification and an MLIR-based C++ and Python implementation for defining programs and integrating them with compatible compilers.
Kubernetes-sigs Agent Sandbox creates isolated, disposable environments for running untrusted or agent-generated code. It is intended to separate such code execution from other workloads.
Gemini Live API is a Google Gemini API for building voice agents that listen and respond through live audio-to-audio interactions. It supports real-time responses, interruption handling, audio streaming, and calls to application tools.
live-dj is an open-source voice-agent demo in which users talk to Mira, a late-night radio DJ, ask her to play music, and interrupt her while she speaks. It uses the Gemini Live API through the raw `google-genai` SDK rather than an agent framework. The browser handles microphone capture, 16 kHz audio input, 24 kHz playback, client-side barge-in, and music ducking. The server maintains one asynchronous Live API session per browser and runs the core loop of opening a session, sending microphone audio, receiving streamed voice responses, and playing them. The full application adds Mira's persona, transcripts, music-tool dispatch, and controls for playing playlists or tracks, skipping, and pausing; the tools return immediately so the voice response does not stall. A minimal backend exposes the 39-line voice-only primitive, while the full server demonstrates the complete DJ application. The repository also documents a per-turn `session.receive()` behavior that requires an outer loop for continuing conversation, and shows that microphone audio must be sent through `send_realtime_input` rather than `send_client_content`. It runs locally with `uv`, Uvicorn, and a Gemini Developer API key. The repository includes a browser client, four dream-pop tracks, persona assets, and examples of the relevant implementation pitfalls.
Free Claude Code is an independent open-source local proxy for routing Claude Code, Codex, Pi, OpenCode, and other coding-agent requests to selected free, paid, subscription, or local model providers while preserving their existing APIs. It provides a searchable model catalog and an Admin UI for configuring providers, supports automatic fallback to another configured model after provider retries are exhausted, and can be launched from a terminal, desktop app, IDE, Discord, Telegram, or phone. Optional RTK filtering reduces common terminal-output tokens, while voice input can use local Whisper or NVIDIA NIM transcription. The project states that it is not affiliated with or endorsed by Anthropic, and that provider free-tier availability and limits may change.
Agent Plugins for AWS is an AWS Labs collection of plugins for AI coding agents, helping them architect, deploy, and operate on AWS. Supported agents include Claude Code, Codex, and Cursor. Each plugin can package agent skills—structured workflows and best-practice playbooks—alongside MCP servers that provide access to live documentation, pricing data, and other APIs; hooks that validate changes or trigger workflows; and references containing documentation and configuration defaults. The repository describes the plugins as reusable, versioned capabilities intended to reduce prompt context and standardize agent behavior. It warns that generative AI can make mistakes and recommends reviewing generated code, costs, security, and credentials. The repository also identifies Agent Toolkit for AWS as the successor for production use, while stating that this project continues to work and accept contributions.
Originality.ai is an AI-content detection service that scans text, including AI-generated writing that has been manually revised, to assess whether it is likely to have been produced by an AI system or written originally by a person. Its website describes it as an AI detector and checker for text generated with systems such as ChatGPT, GPT-5, and Gemini.
GPTZero is an AI-detection service for assessing whether text was written by a person or generated by an AI system. Its website offers an AI score and sentence-by-sentence detection for pasted text, including text associated with ChatGPT, GPT-5, and Gemini.
QuillBot is an AI writing tool that rewrites or humanizes AI-generated text. In the cited video context, it is used to modify text before the text is tested by GPTZero.
Grammarly is an AI writing assistance service that provides personalized guidance and text generation across apps and websites. In the cited video context, it is used to rewrite or humanize text before the text is checked for AI signals.
Clever AI Humanizer is described as a tool that humanizes test text before the text is evaluated by GPTZero. The available evidence does not specify its developer, interface, distribution, or implementation.
Turnitin is an academic integrity service that detects suspected AI-generated writing in student documents. The videos discuss false-positive results and the importance of preserving document drafts when using its detection results.
An AI text humanizer from UndetectedGPT that transforms AI-generated text into writing described as natural and undetectable by AI detectors. The service claims to bypass detectors including GPTZero and Turnitin and offers free humanizing.
ZeroGPT is an AI-content detection service that tests text for AI-generated or humanized content and reports the percentage detected as AI-generated. Its website describes it as an AI detector and ChatGPT detector for content associated with ChatGPT, GPT-5, and Gemini, with a free checking option.
Droids are Factory’s autonomous software-development agents, designed to carry out engineering work for enterprise teams. They are an AI software-development product rather than a general-purpose category.
arena.ai is a model-discovery platform or mechanism used to find and try newly released AI models. The videos identify it as a resource for exploring available models rather than as a particular model itself.
Marin is an open-source research program, software platform, and community for developing foundation models, with a focus on large language models. It covers data curation, transformation and filtering, tokenization, pretraining, posttraining, and evaluation, while recording processes, experiments, decisions, artifacts, and failed experiments from raw data through final model. Marin experiments are defined as dependent steps executed in topological order, similar to a Makefile, and the platform can scale from small language-model runs to large TPU deployments. The project has also been used for audio-text, DNA, and protein models and provides documentation, experiment code, checkpoints, and reusable training pipelines.
OpenHuman is a local-first personal AI system from tinyhumansai that combines persistent memory, agent orchestration, and research tools. It stores a user's data as scored Markdown trees in SQLite on the local machine and mirrors the result to an editable Obsidian vault; its TokenJuice component compresses tool output before it reaches the language model. Its orchestration layer runs checkpointed, durable agent workflows and worker fleets on graphs, with triggers, approvals, steering, halting, and replay support. The system also provides web search, scraping, coding tools, a browser, native voice through in-process Whisper, model routing, messaging integrations, and support for provider keys or fully local Ollama models. The repository describes it as an early beta under active development.
Markdown-based agent skill that rewrites AI-sounding text to read human-written without changing what it says — works with any agent that supports skills. Rewrites against the 35 patterns from Wikipedia's 'Signs of AI writing' (WikiProject AI Cleanup): a first pass free to restructure, then a check of the draft against those patterns and the original claims before rewriting what still sounds artificial. Its rules bar invented facts — names, numbers, dates and quotes must come from the source or the writer — and preserve a writer's personal style (or follow a provided sample). The skill shows its work: the first rewrite and a critique of what still sounds artificial precede the final version.
TURNR is an AI marketing agent for home service businesses. It learns from customer calls, the business website, and existing content, then creates social media posts, video hooks, SEO blog posts, and ad copy, with a weekly email summarizing the generated materials.
4DAnyone is an open-source generative AI pipeline from ant-research that converts a casual monocular video of one person into synchronized videos from multiple virtual camera views for downstream 3D or 4D Gaussian-splatting reconstruction. It supports configurable camera layouts, including full 360-degree or frontal-arc coverage, multiple pitch layers, adjustable yaw ranges, and flexible view counts. The repository provides inference scripts, automatically downloadable models and examples, motion-recovery outputs, camera metadata, skeleton videos, and generated target-view videos. Input footage should be at least 720p, preferably 1080p, use a 9:16 portrait aspect ratio, contain a single person in a full-body or upper-body shot, include at least 121 frames, and have only mild camera motion.
MemoraX Code is a memory plugin and shared memory layer for AI coding agents, developed by MemoraX. It integrates with Codex, Claude Code, CodeBuddy/WorkBuddy, DeepSeek Harness, and OpenCode to retrieve relevant context for new tasks and capture reusable knowledge from completed work. Its memory is divided into Coding Memory for engineering lessons and design decisions, Repo Memory for repository structure and history evidence, Personal Memory for user preferences, and Procedure Memory for reusable steps and validation gates. Background writeback extracts selected knowledge from trusted workspace turns, while the bundled skill and CLI support explicit search and memory operations; repository and personal/procedure content are maintained under the documented local storage boundaries. The package is distributed through npm and requires Node.js 20 or later, with Python 3 required for Repo Memory operations. Cloud-backed search and storage require a MemoraX account, while guest mode is available for a limited period. Local trace capture is enabled by default for supported clients and may retain prompts, responses, recalled memory, reminder text, and local paths; the project documents settings for metadata-only capture or disabling traces. The repository is licensed under the MIT License.
sepia is a portable Agent Skill for Claude Code, Codex, Grok Build, and Antigravity that revises AI-assisted fiction and professional prose at the narrative-architecture and discourse levels, rather than only changing word choice. For fiction, its three-pass protocol addresses narrative architecture, discourse flow, and surface style; its rules cover issues such as overly tidy causality, explained themes, linear time, sparse character networks, uniform emotional rendering, templated paragraph flow, and predictable endings. Professional-writing rules are matched to document venues including release notes, pull-request and issue replies, postmortems, tickets, and technical articles. The package provides write, review, refactor, and recreate operations, plus a general router. Review diagnoses without editing, refactor makes minimal in-place changes, and recreate rewrites from source facts and intent. It includes a 30-feature diagnosis rubric, model-specific fingerprint corrections, shared professional-prose checks, and research references. The repository distributes the skill as a plugin package with native installation paths for the supported tools and an alternative Skills CLI installation. It is licensed under the MIT License.
OpenInstinct is an open-source, self-hostable personal assistant operated through iMessage. It uses a cloud browser to sign into sites and perform tasks such as booking tickets or handling groceries, while keeping passwords, credit cards, and other secrets outside the model's view through encrypted storage. It can connect to Gmail, Google Calendar, and read-only Google Contacts; sending email and creating confirmed calendar events require approval. The application can be deployed to Vercel, where it provisions Kernel for cloud browsers, Neon for PostgreSQL, Vercel Blob for browser images and per-user memory, and the Vercel AI Gateway for inference. Linq provides the optional iMessage and production phone-sign-in connection, and the system can use any model. Local development uses Docker Compose, PostgreSQL, the same vault, browser, and AI Gateway path, and Better Auth. The repository warns that it is not intended for production use.
An MIT-licensed Claude Code skill that applies concrete design rules from Refactoring UI by Adam Wathan and Steve Schoger when building or reviewing interfaces. It selects spacing, typography, color, shadow, and radius values from fixed scales; creates hierarchy through font weight and color rather than relying on larger text; and translates subjective issues such as an interface looking "off" into specific mechanical fixes. The repository includes a SKILL.md procedure, references for palette construction, diagnosis, depth, typefaces, grids, and images, plus a contrast-verified starter tokens.css file. Claude Code can load it as a personal or project skill, and it does not include the book's text or images.
ask is a command-line tool for asking AI questions from a terminal without giving an agent control over the project. It runs an already installed and authenticated Codex, Claude Code, Pi, or OpenCode CLI in read-only mode in the current directory, prints answers to stdout, and sends prompts and errors to stderr. One-shot questions use `ask [QUESTION...]`; interactive sessions can be continued with `ask -c`, with turns, agent choices, model settings, reasoning settings, and the underlying agent session ID saved for later use in the folder. It can also reopen saved sessions, configure defaults, and upgrade itself. The installer supports Intel and ARM Macs and x86_64 and ARM64 Linux, and the project is licensed under MIT.
A style-neutral Codex skill that turns raw photo collections or finished pages into responsive, page-turning 3D photo books using HTML, CSS, and vanilla JavaScript. For raw collections, it creates and reviews an ordered contact sheet, selects compatible visual skills, curates the strongest images, and plans the book's opener, transitions, pauses, peaks, echoes, and ending while varying composition and preserving a coherent visual language. It generates and reviews complete cover and interior spreads, then assembles accepted artwork into a flipbook runtime with responsive sizing, touch and mouse interaction, buttons, keyboard controls, and page-bound spine shadows. It can be installed from GitHub for use in Codex; the repository also includes dependency-free HTML, React Three Fiber, and Three.js reference implementations.
Goldie is an app-store screenshot and preview-video generator for iOS apps, designed for coding agents and human users. It uses Argent flows to replay app interactions in an iOS simulator, captures the results, adds device bezels, backgrounds, headlines and other design elements, joins the clips into preview videos, and checks the output against Apple's upload rules. It is framework agnostic and can drive SwiftUI, UIKit, Flutter, React Native and Kotlin Multiplatform apps through the simulator. The CLI provides commands to check tools, simulators and flows, run the capture-and-render pipeline, and open a browser-based studio for editing backgrounds, templates, bezels, fonts and per-tile copy. Design settings are saved in goldie.design.json, while generated assets are written to an output directory per locale. Goldie requires macOS with iOS simulators, Node.js 20 or newer and ffmpeg; its previews must be 15 to 30 seconds long. The project is sponsored by Software Mansion, the creator of Argent.
Open Pstack is an unofficial, plugin-based workflow kit that brings Lauren Tan's pstack to Claude Code and Codex. It gives coding agents engineering rules, task-specific workflows, focused skills, and small local tools rather than providing a new model or hosted service. Its main poteto-mode workflow reads a task, selects an appropriate workflow, learns how the existing system works, compares designs when needed, favors small changes, and can use multiple models to challenge important decisions. It runs the code and checks behavior as a user would instead of stopping at passing tests, then can continue through review and continuous integration to prepare a pull request. Additional skills cover system explanation, architectural decisions, competing implementations, design interrogation, verification-skill creation and maintenance, pull-request supervision, and reflection. The repository supports installation as a Claude Code or Codex plugin and shares the same skill set between those applications. It is distributed under the MIT license and tracks the upstream pstack project while adapting its skills for Claude Code and Codex.
Project SuperDex is a unified simulation platform for dexterous manipulation research, developed by Facebook Research. It combines SuperDex Physics, a contact-first physics engine designed for tactile manipulation and stable contact; SuperDex Robotics, an SDK for composing robot definitions, controllers, sensors, actuators, and simulation configurations; SuperDex Studio, a desktop GUI for creating and validating robot, mesh, task-prefab, and scene assets; and SuperDex Lab, a simulation harness with a Gymnasium-style API for reinforcement learning, model-predictive control, and system identification. The project provides Python support with pre-built wheels and can also be built from C++ sources. Its first-party source code is licensed under Apache 2.0, while assets and documentation use CC-BY-4.0 except where otherwise noted; third-party components retain their own licenses. VR-based teleoperation is described as a planned module.
Jot is an Apache-2.0-licensed macOS menu-bar dictation app created by Ammaar Reshi. Holding a configurable hotkey records speech, sends it directly from the Mac to Gemini's specialist transcription model, removes filler words and applies corrections such as changes of mind, then inserts the polished text at the cursor. It supports hands-free recording, cancellation, tone matching, a custom dictionary, and a History window. Audio is written to disk from the start of capture, with crash recovery and offline retry queuing. The insertion pipeline tries the macOS Accessibility API first, then a clipboard-restoring paste, and finally offers a manual paste when focus changes; an optional tone pass is guarded by validation and falls back to the raw transcript. Dictations and retry data are stored in a local SQLite-based history store. The app requires macOS 14+, a Gemini API key, microphone permission, and Accessibility permission, and is explicitly described as not an officially supported Google product.
Whip is a Go-based coding-agent harness distributed as a single binary with no runtime dependency. It runs an LLM tool-use loop for bash commands, reading, writing, editing, and subagents, alongside an interactive Bubble Tea terminal interface. The harness supports parallel tool calls, streaming, background subagents, MCP servers, and provider-routable models with live discovery from provider catalogs; any OpenAI-compatible endpoint can be used as a provider. Prebuilt checksum-verified binaries are available for Linux and macOS on x64 and arm64, and it can also be installed from source with Go.
modelprint is a browser-based tool for fingerprinting OpenAI-compatible API endpoints and investigating which model or serving infrastructure may sit behind an alias. It sends controlled probes and compares results side by side, including normalized tokenizer counts for English, CJK, code, and rare Unicode inputs; template offsets; validation messages and error codes; maximum-output refusals; finish-reason vocabulary; router metadata; generation records; response headers; provider path behavior; context-window classes; and logprob geometry. The tool runs without installation, a build step, or a server: API calls go directly from the browser to the provider, with keys retained in the browser tab. It can pin models to provider hosts to reduce routing noise and labels signals reported or validated by a router. Its documentation cautions that matching fingerprints indicate shared infrastructure or a model family rather than proving model identity, and that probes may report missing signals when headers, logprobs, or other telemetry are unavailable. The project supports community-authored probes through a defined JavaScript contract and is released under the MIT license.
Open Steps is an open-source pack of agent skills that translates coding-agent output into plain-language reports, verdicts, questions, and next steps. Its skills include done-or-not reports, step-by-step instructions for nontechnical users, simple explanations of agent questions, premortems for hard-to-reverse decisions, verification of another session's claims, next-task recommendations, and plain-language rewrites. The pack is built and measured primarily for Claude Code, where it installs as a plugin with two shell hooks. The session-start hook injects the routing table and the latest report, while the stop hook checks for completed work and requests a report when appropriate. Reports are stored outside project repositories. Skills and routing instructions can also be installed for Codex, Cursor, and Gemini CLI, although the repository notes that hook support differs across those tools. Open Steps separates measured results from assumptions and marks unchecked information as "not checked." It is free software under the MIT license and is developed by Pavlo Kharmanskyi.
headcount is an agent organization for Claude Code, structured as a company with independently installable departmental plugins containing named skills for engineering, business, and operational work. Projects install only the departments they need, address skills as department:skill to avoid name collisions, and can invoke skills directly or have them load when a request matches their territory. Each department includes an agent charter for delegation as a subagent with an exclusive write surface. The repository organizes agents by ownership boundaries rather than topic, and documents cross-department workflows, decision logs, surface ownership, and an interactive searchable organization chart. Reviewer-class security and legal-risk departments can block work under review. headcount is distributed as a Claude Code plugin marketplace and is licensed under the MIT License. The repository identifies Chris Brock as its builder and includes a validation script and CI checks for skill front matter, manifests, department references, license text, and ownership-surface consistency.
Lemmalog is a Datalog engine for LLM-agent memory, distributed as a Rust crate with an MCP server, REPL, and agent skill. It treats extracted statements as provenance-tracked base facts, then uses runtime-parsed stratified Datalog to derive temporal projections, closures, contradiction candidates, relevance relationships, and aggregates. The engine supports seminaive incremental fixpoint evaluation, stratified negation with negative-cycle rejection, bi-temporal facts, confidence and provenance annotations, proof trees through why(), scoped retraction recomputation, demand-driven ask_deep queries using magic sets, and persistence of episodes, rules, and base facts while rebuilding derived relations on load. Its AgentMemory facade connects an extractor to deterministic ADD, UPDATE, NOOP, and escalation policies, and assembles budgeted contexts from distilled facts and their source episodes. The MCP server exposes the engine to agent harnesses such as Claude Code and Kimi CLI through tools for observing facts, querying, explaining proofs, retrieving context, saving state, installing rules, and running hypotheticals.
Epic Infographics is an open-source skill for AI agents that generates data-driven infographics as HTML/CSS scenes rather than stock dashboard templates. It guides the agent through audience and story-angle selection, visual-metaphor and design-language selection, scene composition, chart construction, rendering, and self-review. Charts are computed from explicit arithmetic, palettes undergo color-blind-safety checks, and a headless preflight script checks text collisions, clipping, canvas boundaries, and readable sizes. The skill can render still PNGs and animate the same HTML through CSS keyframes into MP4 or GIF output using Playwright, headless Chromium, and ffmpeg; its motion workflow treats animation as a layer on an approved still and reviews contact sheets of the result. It includes named design-language specifications, composition and chart references, HTML templates, render, and animation workflows.
Fire Your SEO Agency is a Claude Code skill that audits and optimizes websites across five search-visibility lanes: SEO for Google and Bing crawling, AEO for AI answer boxes such as Google AI Overviews and Bing Copilot, GEO for generative engines such as ChatGPT, Perplexity, and Claude, LLMO for how language models represent a brand, and NEO for Naver Search and AI Briefing. Its operating procedure audits sites without JavaScript, checks rendering, sitemaps, metadata, structured data, server-side rendering, machine-readable content, brand consistency, crawler policies, and Naver requirements, then proposes and implements fixes, creates intent-focused landing pages and citation-ready content structures, and schedules remeasurement. It is distributed as a Claude Code plugin or project/personal skill and is licensed under the MIT License.
ACRYL is an agent-agnostic Agentic Development Environment and continuity layer for software work. It provides a persistent project workspace with a canonical event stream, durable tasks and artifacts, agent identities and sessions, context projections, structured handoffs, workspaces, and checkpoints, allowing replaceable coding agents to work on the same project context. ACRYL is built on the Cordis meta-framework, which supplies lifecycle-managed plugins, named services and replaceable providers, reactive dependency injection, typed events, reversible effects, scoped composition, and configuration-driven application profiles. Agents, models, memory systems, code graphs, tools, workflows, terminals, and user-interface surfaces are treated as composable capabilities. Its Development Canvas can host PTY terminals, coding-agent sessions, files and editors, browser tabs, and other capability-provided views. The project is distributed as separate Desktop GUI, terminal TUI, and local web installations; the web surface runs locally rather than as a hosted cloud service. ACRYL is in active early development, is licensed under the MIT License, and is developed independently of the DeepSeek Harness project while retaining architectural influences from it.
LiveKit Agents is an open-source Python framework for building programmable, real-time multimodal voice agents that run as server-side participants. It combines speech-to-text, large language models, text-to-speech, and realtime APIs with LiveKit's WebRTC clients, telephony stack, RPC and data APIs, and MCP tool support. The framework provides agent sessions, server-side job scheduling and dispatch, semantic turn detection based on a transformer model, multi-agent handoffs, and tools for voice, text-only, transcription, vision, and video-avatar applications. Agents can be tested with native test integration, including event assertions and LLM-based judges, and can run locally in console mode, in development mode with hot reloading, or in production mode. The framework can use LiveKit Cloud or a self-hosted LiveKit server and is distributed as a Python package with plugins for model providers. The Agents framework is licensed under Apache-2.0; LiveKit turn-detection models use the LiveKit Model License.
Paseo is a self-hosted platform for orchestrating multiple coding agents, including Claude Code, Codex, GitHub Copilot, OpenCode, and Pi. Its local daemon manages agent processes on users' machines, while desktop, mobile, web, and CLI clients connect to run agents in parallel, stream output, send follow-up tasks, and work in specified directories or worktrees. Paseo supports voice task dictation and control, agent handoffs, advisor and committee workflows, local tools, configurations, skills, and development environments, and provides an MCP server, WebSocket API, and TypeScript SDK for integrations, dashboards, and orchestration services. Remote connections use an end-to-end encrypted relay, TCP, Tailscale, or another VPN. Paseo can run as an installed application, headlessly, or as a Dockerized daemon with a self-hosted web UI. The repository is licensed under Apache-2.0 and states that Paseo has no telemetry, tracking, or forced log-ins.