704 tools and products — trending open source, and what gets used in AI and other work.
ask is a command-line tool for asking AI questions from a terminal without giving an agent control over the project. It runs an already installed and authenticated Codex, Claude Code, Pi, or OpenCode CLI in read-only mode in the current directory, prints answers to stdout, and sends prompts and errors to stderr. One-shot questions use `ask [QUESTION...]`; interactive sessions can be continued with `ask -c`, with turns, agent choices, model settings, reasoning settings, and the underlying agent session ID saved for later use in the folder. It can also reopen saved sessions, configure defaults, and upgrade itself. The installer supports Intel and ARM Macs and x86_64 and ARM64 Linux, and the project is licensed under MIT.
Goldie is an app-store screenshot and preview-video generator for iOS apps, designed for coding agents and human users. It uses Argent flows to replay app interactions in an iOS simulator, captures the results, adds device bezels, backgrounds, headlines and other design elements, joins the clips into preview videos, and checks the output against Apple's upload rules. It is framework agnostic and can drive SwiftUI, UIKit, Flutter, React Native and Kotlin Multiplatform apps through the simulator. The CLI provides commands to check tools, simulators and flows, run the capture-and-render pipeline, and open a browser-based studio for editing backgrounds, templates, bezels, fonts and per-tile copy. Design settings are saved in goldie.design.json, while generated assets are written to an output directory per locale. Goldie requires macOS with iOS simulators, Node.js 20 or newer and ffmpeg; its previews must be 15 to 30 seconds long. The project is sponsored by Software Mansion, the creator of Argent.
Open Pstack is an unofficial, plugin-based workflow kit that brings Lauren Tan's pstack to Claude Code and Codex. It gives coding agents engineering rules, task-specific workflows, focused skills, and small local tools rather than providing a new model or hosted service. Its main poteto-mode workflow reads a task, selects an appropriate workflow, learns how the existing system works, compares designs when needed, favors small changes, and can use multiple models to challenge important decisions. It runs the code and checks behavior as a user would instead of stopping at passing tests, then can continue through review and continuous integration to prepare a pull request. Additional skills cover system explanation, architectural decisions, competing implementations, design interrogation, verification-skill creation and maintenance, pull-request supervision, and reflection. The repository supports installation as a Claude Code or Codex plugin and shares the same skill set between those applications. It is distributed under the MIT license and tracks the upstream pstack project while adapting its skills for Claude Code and Codex.
Whip is a Go-based coding-agent harness distributed as a single binary with no runtime dependency. It runs an LLM tool-use loop for bash commands, reading, writing, editing, and subagents, alongside an interactive Bubble Tea terminal interface. The harness supports parallel tool calls, streaming, background subagents, MCP servers, and provider-routable models with live discovery from provider catalogs; any OpenAI-compatible endpoint can be used as a provider. Prebuilt checksum-verified binaries are available for Linux and macOS on x64 and arm64, and it can also be installed from source with Go.
Open Steps is an open-source pack of agent skills that translates coding-agent output into plain-language reports, verdicts, questions, and next steps. Its skills include done-or-not reports, step-by-step instructions for nontechnical users, simple explanations of agent questions, premortems for hard-to-reverse decisions, verification of another session's claims, next-task recommendations, and plain-language rewrites. The pack is built and measured primarily for Claude Code, where it installs as a plugin with two shell hooks. The session-start hook injects the routing table and the latest report, while the stop hook checks for completed work and requests a report when appropriate. Reports are stored outside project repositories. Skills and routing instructions can also be installed for Codex, Cursor, and Gemini CLI, although the repository notes that hook support differs across those tools. Open Steps separates measured results from assumptions and marks unchecked information as "not checked." It is free software under the MIT license and is developed by Pavlo Kharmanskyi.
headcount is an agent organization for Claude Code, structured as a company with independently installable departmental plugins containing named skills for engineering, business, and operational work. Projects install only the departments they need, address skills as department:skill to avoid name collisions, and can invoke skills directly or have them load when a request matches their territory. Each department includes an agent charter for delegation as a subagent with an exclusive write surface. The repository organizes agents by ownership boundaries rather than topic, and documents cross-department workflows, decision logs, surface ownership, and an interactive searchable organization chart. Reviewer-class security and legal-risk departments can block work under review. headcount is distributed as a Claude Code plugin marketplace and is licensed under the MIT License. The repository identifies Chris Brock as its builder and includes a validation script and CI checks for skill front matter, manifests, department references, license text, and ownership-surface consistency.
Lemmalog is a Datalog engine for LLM-agent memory, distributed as a Rust crate with an MCP server, REPL, and agent skill. It treats extracted statements as provenance-tracked base facts, then uses runtime-parsed stratified Datalog to derive temporal projections, closures, contradiction candidates, relevance relationships, and aggregates. The engine supports seminaive incremental fixpoint evaluation, stratified negation with negative-cycle rejection, bi-temporal facts, confidence and provenance annotations, proof trees through why(), scoped retraction recomputation, demand-driven ask_deep queries using magic sets, and persistence of episodes, rules, and base facts while rebuilding derived relations on load. Its AgentMemory facade connects an extractor to deterministic ADD, UPDATE, NOOP, and escalation policies, and assembles budgeted contexts from distilled facts and their source episodes. The MCP server exposes the engine to agent harnesses such as Claude Code and Kimi CLI through tools for observing facts, querying, explaining proofs, retrieving context, saving state, installing rules, and running hypotheticals.
Epic Infographics is an open-source skill for AI agents that generates data-driven infographics as HTML/CSS scenes rather than stock dashboard templates. It guides the agent through audience and story-angle selection, visual-metaphor and design-language selection, scene composition, chart construction, rendering, and self-review. Charts are computed from explicit arithmetic, palettes undergo color-blind-safety checks, and a headless preflight script checks text collisions, clipping, canvas boundaries, and readable sizes. The skill can render still PNGs and animate the same HTML through CSS keyframes into MP4 or GIF output using Playwright, headless Chromium, and ffmpeg; its motion workflow treats animation as a layer on an approved still and reviews contact sheets of the result. It includes named design-language specifications, composition and chart references, HTML templates, render, and animation workflows.
Boop is a tiny, self-hosted notification inbox for developers. Applications submit events to a Go server using a project API key; the server redacts configured sensitive keys, stores the events in SQLite, and sends push notifications directly to Apple's APNs. The notification contains the title, body, and event ID, while the iOS app fetches the full event from the server using a device credential. The bundled web UI manages projects, devices, pairing, retention, silences, and event grouping. A one-time pairing token and QR code connect the iOS app, which must be built and signed by the operator. Events support levels, structured error data, up to three URL actions, fingerprints for grouping repeated occurrences, and a Sentry-compatible ingest endpoint. Boop also exposes an optional read-only, Streamable HTTP MCP endpoint for AI agents to query projects and events; it contains no LLM. Distribution consists of a Go binary with an embedded Svelte web UI, a Docker deployment, and the iOS app source. The project uses one SQLite database file, has no hosted relay, account system, or telemetry, and is released under the MIT license.
ACRYL is an agent-agnostic Agentic Development Environment and continuity layer for software work. It provides a persistent project workspace with a canonical event stream, durable tasks and artifacts, agent identities and sessions, context projections, structured handoffs, workspaces, and checkpoints, allowing replaceable coding agents to work on the same project context. ACRYL is built on the Cordis meta-framework, which supplies lifecycle-managed plugins, named services and replaceable providers, reactive dependency injection, typed events, reversible effects, scoped composition, and configuration-driven application profiles. Agents, models, memory systems, code graphs, tools, workflows, terminals, and user-interface surfaces are treated as composable capabilities. Its Development Canvas can host PTY terminals, coding-agent sessions, files and editors, browser tabs, and other capability-provided views. The project is distributed as separate Desktop GUI, terminal TUI, and local web installations; the web surface runs locally rather than as a hosted cloud service. ACRYL is in active early development, is licensed under the MIT License, and is developed independently of the DeepSeek Harness project while retaining architectural influences from it.
LiveKit Agents is an open-source Python framework for building programmable, real-time multimodal voice agents that run as server-side participants. It combines speech-to-text, large language models, text-to-speech, and realtime APIs with LiveKit's WebRTC clients, telephony stack, RPC and data APIs, and MCP tool support. The framework provides agent sessions, server-side job scheduling and dispatch, semantic turn detection based on a transformer model, multi-agent handoffs, and tools for voice, text-only, transcription, vision, and video-avatar applications. Agents can be tested with native test integration, including event assertions and LLM-based judges, and can run locally in console mode, in development mode with hot reloading, or in production mode. The framework can use LiveKit Cloud or a self-hosted LiveKit server and is distributed as a Python package with plugins for model providers. The Agents framework is licensed under Apache-2.0; LiveKit turn-detection models use the LiveKit Model License.
Paseo is a self-hosted platform for orchestrating multiple coding agents, including Claude Code, Codex, GitHub Copilot, OpenCode, and Pi. Its local daemon manages agent processes on users' machines, while desktop, mobile, web, and CLI clients connect to run agents in parallel, stream output, send follow-up tasks, and work in specified directories or worktrees. Paseo supports voice task dictation and control, agent handoffs, advisor and committee workflows, local tools, configurations, skills, and development environments, and provides an MCP server, WebSocket API, and TypeScript SDK for integrations, dashboards, and orchestration services. Remote connections use an end-to-end encrypted relay, TCP, Tailscale, or another VPN. Paseo can run as an installed application, headlessly, or as a Dockerized daemon with a self-hosted web UI. The repository is licensed under Apache-2.0 and states that Paseo has no telemetry, tracking, or forced log-ins.
A personal AI learning system distributed as a pi configuration. It encodes a teaching philosophy and learning process in skills, including teaching and diagram-based visualization, and adds extensions for question popups, graded quizzes, Markdown session logs, and visualization tools. The configuration delegates research and visual creation to researcher, SVG-maker, and Mermaid-maker subagents; it can also run without subagents, with those delegation-based capabilities omitted. It is intended for one learner and is shared as-is under the repository's stated installation instructions.
LiveKit is an open-source framework and developer platform for building, testing, deploying, scaling, and observing real-time voice, video, and physical AI agents. Its agent pipeline streams user speech from an app, browser, or phone call to an agent, which applies custom business logic and returns a response; the platform supports automatic turn detection and interruption handling, speech-to-text, language-model, and text-to-speech providers, web and mobile applications, and telephony through phone numbers and SIP integrations. LiveKit Cloud provides deployment and scaling on LiveKit's real-time infrastructure, alongside an inference gateway and full-stack observability for agent sessions.
A downloadable web-development project bundle from GreatStack for building a full-stack AI website builder with MongoDB, Express.js, React.js, and Node.js. The project includes starter assets and source files for a React website generator that accepts text prompts, builds websites step by step, exposes generation progress, and supports manual source-code editing, follow-up AI prompts, exporting, and publishing. Its tutorial project includes user authentication, REST APIs, an OpenRouter model integration, and an agent chat API for updating generated projects.
Crawl4AI is an open-source Python web crawler and scraper that converts web pages into structured, LLM-ready Markdown for retrieval-augmented generation, agents, and data pipelines. Its asynchronous Playwright-based crawler supports Chromium, Firefox, and WebKit, dynamic JavaScript pages, sessions, persistent browser profiles, cookies, headers, proxies, screenshots, media, iframes, lazy loading, full-page scanning, caching, and deep crawling with BFS, DFS, and best-first strategies. For extraction, it provides heuristic Markdown filtering including BM25-based relevance filtering, CSS- and XPath-based schema extraction, chunking and cosine-similarity strategies, and optional LLM-driven structured JSON extraction. It also includes adaptive crawling, link analysis, URL seeding, virtual-scroll handling, anti-bot and proxy escalation features, and customizable hooks. Crawl4AI can be installed with pip and used through Python or its command-line interface. It is also distributed as a Dockerized FastAPI server with JWT authentication, browser pooling, monitoring dashboards, a playground, and endpoints for crawling, HTML extraction, screenshots, PDF generation, and JavaScript execution. The repository states that it is licensed under Apache License 2.0.
GitNexus, developed by Akon Labs, is a code-intelligence engine that indexes repositories into a knowledge graph for code exploration and AI-agent context. Its indexing pipeline walks the file tree, parses source with Tree-sitter, resolves imports, calls, inheritance, constructor-inferred receiver types and other relationships, groups symbols into functional communities, traces execution processes, and builds BM25-plus-semantic hybrid search indexes backed by LadybugDB. The resulting graph supports MCP tools and CLI commands for process-grouped search, symbol context, call-path tracing, blast-radius and Git-diff impact analysis, structural checks, coordinated renaming, API and route mapping, taint and dependence queries, and Cypher access; repository groups can link contracts and impacts across multiple repositories. The CLI runs locally and can connect editors such as Claude Code, Cursor, Codex and others through MCP, skills and selected hooks. GitNexus also provides a browser-based WebAssembly UI with an interactive graph explorer and AI chat, plus a local HTTP server and Docker deployment mode for accessing indexed repositories through a backend. The web-only mode keeps repository processing in the browser and is constrained by browser memory, while the native CLI stores indexes locally in each repository's .gitnexus directory and uses a global registry for multi-repository access.
FuXi is a self-contained terminal AI coding agent developed by FUXI. Built in Go and distributed as a static binary, it uses a Think → Act → Verify loop to read and edit code, run shell commands, drive tools, connect to MCP servers, and route requests across multiple LLM providers with automatic failover and cost-aware settings. Its built-in capabilities include file operations, shell execution, code search, web fetching, LSP diagnostics, Jupyter, browser use, background tasks, and parallel sub-agents. Shell commands pass an AST-based safety classifier, while permissions and audit logs govern autonomous actions. Sessions persist to disk, with checkpoints for resuming, rolling back, or forking; the TUI also supports memory consolidation and automatic context compaction. FuXi supports provider API keys or FuXi OAuth, configurable OpenAI-compatible endpoints, MCP clients, hooks, skills, plugins, and slash commands. The repository contains documentation, installers, release information, and issue-tracking materials; it states that the product source is proprietary and not published.
MyContext is a local-first desktop app from openTrinity that builds a private personal work-context layer from sources such as instant-messaging conversations, documents, and meeting records. It stores local copies, indexes, source references, and derived context in an on-disk SQLite vault, then organizes them into a personal context graph linking people, projects, topics, events, conversations, and supporting facts. Its search and answer workflow combines local full-text search, semantic retrieval, and graph queries, with agents assembling answers from traceable source material and falling back to ranked local results when the agent runtime is unavailable. A digital-self workflow recalls relationship-specific context and communication history to draft replies, while sending, deletion, and other consequential actions require explicit user confirmation. The repository describes an Electron and React desktop architecture with source connectors, incremental ingestion, context processing, retrieval, knowledge-graph, persona, and isolated agent-runtime layers. MyContext is in developer preview and under active development; its README warns of compatibility-breaking changes and migrations that may require recollection. It is licensed under the Elastic License 2.0, which permits use, modification, and self-hosting but restricts offering it as a hosted or managed service to third parties.
screenpipe is a local AI-agent memory layer that continuously captures computer history on macOS, Windows, and Linux. It records screen frames with OCR, accessibility data, microphone and system audio with transcripts, and application activity, storing the underlying history locally. Agents can search the history through a local REST API, database, or MCP server and use it for tasks such as meeting summaries, follow-ups, and workflow automations. Users can exclude apps, windows, URLs, or time periods and redact sensitive fields on the device; the page describes the project as source-available.
OpenClaude is an open-source, terminal-first coding-agent CLI for cloud and local model providers. It connects to OpenAI-compatible APIs, Gemini, GitHub Models, Codex, Ollama, Atomic Chat, and other supported backends, providing prompts, streaming output, Bash and file tools, grep, glob, agents, tasks, MCP, slash commands, web search and fetch, and image inputs for compatible providers. The CLI supports guided provider setup with saved profiles, conversation continuation and forking, detached local background sessions, model-specific agent routing, repository maps based on PageRank-ranked code structure, and a headless bidirectional-streaming gRPC server for integrations, CI/CD pipelines, and custom interfaces. A bundled VS Code extension provides launch integration, in-editor chat, provider-aware controls, and theme support. OpenClaude runs on Node.js 22 or newer, is distributed through npm and an Arch Linux AUR package, and can use local inference or remote APIs. Its repository is licensed MIT for the project's modifications and states that it is an independent community project derived from and substantially modified from the Claude Code codebase, without Anthropic affiliation.
A workstation for cybersecurity investigation and forensics that hosts autonomous incident-response agent harnesses. The video states that five winning agents from the SANS Institute's Find Evil! hackathon were made available on it.
Skill Cabinet is a local catalog for agent skills installed on a machine. It scans user-level skill locations such as .agents, .claude, .codex, .cursor plugins, Hermes profiles, and other ~/.* /skills folders, then lets users filter skills by drawer, metadata, risk, invocation mode, and status; inspect rendered or source bodies, YAML frontmatter, extra files, symlinks, duplicates, origins, and broken links; and review disk usage. It runs with Node 20 or later through npx skill-cabinet, starting a server bound to 127.0.0.1 and opening the catalog in a browser. Users can quarantine skills to ~/.skill-cabinet/quarantine and restore them, or delete individual skills or groups; deletion removes folders, files, or symlinks from the scanned locations, while a symlink's target is retained. The project is distributed under the MIT license.
SkillRadar is open-source discovery, security, ranking, and routing infrastructure for Agent Skills and Codex. It discovers public SKILL.md files, parses their contents, performs conservative static checks for capabilities such as shell commands, dynamic execution, secret access, networking, package installation, filesystem writes, and deployment tooling, then classifies and ranks candidates in a safety-gated registry. D and Blocked candidates are kept audit-only and excluded from automatic routing. Its Codex plugin provides task-to-Top-3 routing, skill search, provenance and safety inspection, and a read-only Skill Budget Doctor. Routing can use a bundled offline registry and returns relevance, SkillRadar score, security grade, provenance, reasons, and match details without executing candidate repositories or depending on live GitHub discovery. The repository includes radar data, a matching system, a router-quality benchmark, a local registry UI, and daily bot-refreshed generated data; it is released under the MIT license.
Procedura is an open-source agentic 3D-modeling tool from SpatiaOS that converts text prompts into editable procedural assemblies rather than point clouds or triangle meshes. It generates OpenSCAD source with named parts and typed mates, plans and builds parts incrementally, and can use reference images and optional Blender render feedback during refinement. Optional passes assign per-part PBR materials and plan articulation, exporting motion to OpenUSD and URDF with Isaac-based validation. It runs locally as a Bun/TypeScript pipeline using a configurable OpenAI-compatible, Gemini, or local model endpoint; it does not provide hosted inference or API keys. The pipeline uses a Manifold-capable OpenSCAD build to compile the generated programs and Blender for renders, and includes a web Studio for composing runs and inspecting intermediate artifacts. The repository is MIT-licensed.
CDAF is an open sidecar format and toolkit for video that stores a timestamped plain-text description beside the corresponding video file. Its Python library and CLI can generate, parse, validate, read, and report the status of `.cdaf` files, while an agent skill teaches video agents to check for a matching sidecar before processing footage. Each sidecar contains a minimal versioned header with the video filename, SHA-256 hash, byte size, duration, generator, and creation time, followed by sections such as summary, timestamped segments, transcript, on-screen text, and tags. Conforming tools verify the video's freshness and refuse to use a stale sidecar after the video changes; the format is model-agnostic even though the included generator uses the Gemini Files API, with optional local-model support. The repository includes a normative specification, reproducible sidecar-versus-direct-video benchmarks, an agent skill installable with `npx cdaf-skill`, and CLI commands for generation, validation, reading, and status checks. The core validation functions require only the Python standard library; generation requires Python 3.10 or later and a user-supplied Gemini API key. The project is licensed under MIT.
Claude 5.1 is described in the supplied video evidence as an Anthropic AI model for programming, scientific research, and agentic tasks. The video claims that it can work continuously for dozens of hours on codebases and multistep research, with separate lower pricing for cached, ordinary, and complex agent tasks.
DBOS is a database-oriented operating-system project and application environment that stores important system state in a database. Its practical focus includes durable, recoverable workflows, particularly workflows used by agentic AI systems.
Academic Research Skills for Claude Code is an open-source suite of Claude Code skills for academic research and publication workflows, maintained by Cheng-I Wu. It provides separate deep-research, academic-paper, academic-paper-reviewer, and academic-pipeline skills for literature reviews, systematic reviews, guided research, drafting, citation conversion, revision, peer review, rebuttal auditing, methodology review, and re-review. The suite uses staged, human-in-the-loop workflows with multi-agent orchestration, Socratic checkpoints, style calibration, writing-quality checks, citation formatting, and outputs in Markdown, DOCX when Pandoc is available, and LaTeX/PDF through tectonic. Its pipeline orchestrator connects activities through ten stages, user-confirmation checkpoints, Material Passport handoffs, claim and citation verification, integrity gates, and final process summaries. Citation verification can cross-check references against Semantic Scholar, OpenAlex, Crossref, and arXiv when available.
Claude-Mem is an open-source persistent-memory plugin and service for coding agents. It captures agent activity through lifecycle hooks, stores sessions, observations, and summaries in SQLite, and uses hybrid full-text and Chroma vector search to retrieve relevant context across sessions. Its MCP search workflow uses progressive disclosure: the agent first searches a compact index, then reviews a timeline, and finally fetches full observations for selected result IDs. A local worker service, managed by Bun, provides the HTTP API, search endpoints, and web viewer; the project also supports integrations with Claude Code and other listed agent environments, configurable context injection, private-content exclusion tags, and optional cloud synchronization. The repository states that it is distributed under the Apache License 2.0 and requires Node.js 20 or later, with Bun, uv, and SQLite used by the runtime. It can be installed through its npx installer or Claude Code's plugin marketplace.
CodeBurn is a free, open-source, local-first tool by AgentSeal that reads session files written by AI coding tools and reports token usage and estimated cost by provider, model, project, task, and activity. It provides a terminal dashboard and reports, a localhost web dashboard, desktop and tray or menubar views, exports, model comparisons, subscription-plan tracking, and optional cross-device aggregation. Its deterministic analyzers classify work into task categories from tool usage and message keywords, calculate token costs using LiteLLM pricing cached locally, and scan coding-agent sessions for patterns such as repeated file reads, low read-to-edit ratios, uncapped shell output, unused MCP servers, bloated configuration, and retry-heavy work. The optimize command produces estimated savings and fixes; applicable configuration changes are backed up and journaled so they can be undone and later compared with observed usage. The yield command heuristically correlates sessions with Git commits to classify spend as productive, reverted, abandoned, or ambiguous. The guard feature installs opt-in Claude Code hooks for soft and hard session-spending caps, checkpoints, and status-line reporting. CodeBurn also exposes local usage and savings through an MCP server over stdio. The CLI reads data from the local machine without wrappers, proxies, API keys, or uploads; the optional desktop applications can send anonymous bucketed telemetry after consent. It requires Node.js 22.13 or newer, and the repository is licensed under MIT.
here.now is an agent-oriented hosting and publishing service for publishing files and folders—including websites, documents, dashboards, presentations, prototypes, games, and media—to the web and receiving a live URL. Any AI agent that can make HTTP requests can publish to it; no account is required, but unauthenticated sites expire after 24 hours, while registered accounts can keep sites permanently. Sites are public by default with randomly generated URLs, and can be protected with passwords or restricted to invited email addresses or domains. The service also supports custom domains and team workspaces with member-only visibility and workspace subdomains.
Miles is an enterprise-facing reinforcement learning framework for large-scale post-training of large language and vision-language models. It pairs SGLang for high-throughput, agentic rollouts with Megatron-LM for scalable training, and also provides a PyTorch FSDP2 backend for Hugging Face implementations. Its asynchronous architecture decouples rollout and training workers, supports configurable on- and off-policy schedules, and updates rollout engines in-loop through peer-to-peer RDMA weight transfer. The framework includes token-in-token-out data flow, Rollout Routing Replay for replaying mixture-of-experts routing decisions during training, fault-tolerant recovery of failed SGLang engines, low-precision training with MXFP8 and NVFP4 alongside FP8, INT4 QAT, BF16, and FP16, and LoRA or multi-LoRA training. It supports reinforcement-learning recipes including GRPO, GSPO, PPO, and REINFORCE++, as well as supervised fine-tuning, on-policy distillation, agentic environments, and diffusion-model training. Miles was forked from slime and integrates SGLang, Megatron-LM, and torch_memory_saver; the repository is released as an open-source project, though the provided page text does not state its license.
Speechify is an AI voice and text-to-speech platform that reads written material aloud and also provides speech-to-text, conversational, and business-focused AI tools. Its SpeechifyAI service offers text-to-speech and realtime voice-agent APIs through a single interface, with streaming synthesis, zero-shot voice cloning from a consented reference clip, SSML-based emotion control, and multilingual output for English, German, Mexican Spanish, French, Italian, and Brazilian Portuguese. The service includes the Simba model family, which the site describes as streaming-native speech models designed to model voice identity, expression, and language for realtime conversation. Speechify was co-founded by Cliff Weitzman, who initially built it to consume written material through audio.
Chinese Patent Skill is an MIT-licensed open-source agent skill for Chinese patent workflows, covering invention, utility-model, and industrial-design patents. It mines patentable points from project materials such as Markdown, code, DOCX, PPTX, and optionally STEP/CAD files; performs novelty, bibliographic, and prior-art searches with China National Intellectual Property Administration sources preferred; drafts patent disclosure documents; rewrites disclosures into claims, specifications, and abstracts; produces Mermaid diagrams or planned patent figures, black-and-white drawings, and optional editable DOCX exports; explains published patents in plain language; and assists with patent-office examination responses and policy briefs. It can read technical and product drawings, extract outlines and component references, derive multiple views from CAD models, and organize utility-model and design workflows with schemas, figure plans, views, and component numbering. Patent explanations can be stored in Obsidian as linked notes, graphs, and canvases, while disclosure work supports self-checking, correction, iterative updates, timestamped drafts, multiple saved versions, and conversation records.
OpenWhispr is an open-source, cross-platform desktop voice-to-text application for macOS, Windows, and Linux. A global hotkey captures speech and inserts the resulting text at the cursor in another application; dictation can use local Whisper or NVIDIA Parakeet speech-to-text engines, where audio remains on the device, or cloud providers through user-supplied keys. It also supports dictation translation, AI-agent commands, meeting transcription with speaker diarization and voice fingerprinting, audio and video transcription, and searchable notes with semantic search and optional cloud sync. The application is built with Electron, React, TypeScript, SQLite, whisper.cpp, and sherpa-onnx, and provides an API and MCP server for programmatic access to notes and transcriptions. The repository describes it as having no data collection or telemetry and distributes installers for the three supported desktop platforms. It is licensed under the MIT license.
Reverify is an AI-agent verification toolkit for reverse engineering and related code analysis. It places deterministic tools between an agent and its claims: the model proposes hypotheses about a binary or candidate implementation, while parsers, disassemblers, pattern scanners, emulators, equivalence checks, and optional analysis engines compare them with ground truth and return evidence-backed VERIFIED, REFUTED, INCONCLUSIVE, OBSERVED, or INVALIDATED results. Its pure-Python core handles PE, ELF, and Mach-O parsing, x86/x64 and ARM/ARM64 disassembly, byte and pattern inspection, CPU micro-emulation, Protobuf/TLV dissection, and Frida hook generation; optional installations add Capstone, Unicorn, LIEF, Z3, or angr for deeper analysis. The project exposes the verification loop through a command-line interface and an MCP server that agents such as Claude Code and Cursor can call. It also records verified, observed, proved, and refuted results in a content-keyed local ledger, allowing grounded facts and negative findings to survive context resets, compaction, and new sessions. A rollover and orchestration system hands sessions off through files rather than model-written summaries, keeping verified facts separate from unverified notes. Reverify includes claim types for bytes, typed reads, instructions, patterns, emulation, behavioral or formal equivalence, imports, exports, sections, and—when angr is installed—functions, calls, references, and reachability. It is distributed from PyPI and licensed under the MIT License. The repository states that it is intended for authorized reverse engineering, malware analysis, CTF work, interoperability research, and software the user owns or is permitted to analyze.
ffmpeg-skill is an agent skill that gives Claude Code, Cursor, Codex, and other agents a local video- and audio-editing workflow built on FFmpeg and ffprobe. It uses a probe → edit, preferring stream copying where possible → check → verify process, with typed Python scripts rather than shell strings. Its 21 tools cover cutting, joining, silence removal, duration and aspect-ratio fitting, captions, overlays, graphics, HDR-to-SDR conversion, LUTs, audio cleanup, loudness, synchronization with drift correction, multicamera editing, delivery checks, project rendering, and batch processing. Each tool supports structured results and dry runs, and the machine-readable contract generates an MCP interface whose tool definitions are derived from the scripts. The skill runs locally without cloud services, API keys, or Python dependencies beyond the standard library; it requires FFmpeg 5.0 or later. It is a tool for local, agent-controlled workflows that probe, edit, verify, and render video.
Commerce Agents is Anthropic's reference implementation for building shopping agents for customers and merchant agents for back-office staff with Claude. It defines each agent through prompts, skills, tool contracts, grounding, memory, approval gates, and backend interfaces, and provides runnable examples for retail, travel, telecom, and entertainment. The shopping agent searches and compares products, plans purchases, fills carts, answers order and policy questions, and maintains customer memory. The merchant agent analyzes performance, manages listings and inventory alerts, handles pricing and promotions, and drafts campaigns; its writes are staged until a host approves them. The agents can run through the Anthropic Messages API, Claude Agent SDK, or Managed Agents, and the repository provides a Claude Code plugin for scaffolding or reviewing commerce agents.
Reef is open-source continual-learning infrastructure for self-improving AI agents, developed by Human-Agent-Society. It connects agent inference, interaction records, feedback, learning jobs, evaluation, and versioned delivery, supporting updates to model weights as well as agent harnesses such as prompts, rules, and skills. Each learning cycle has four stages: Reef serves requests and records interactions; matches later scores or structured feedback to those records; produces a candidate update from eligible records; and evaluates the candidate against a configured selection policy before publishing it. Rejected candidates leave the previous release serving, while accepted updates are committed and delivered without restarting the service. The project provides an HTTP API with OpenAI- and Anthropic-compatible inference endpoints, feedback reporting, scenario-based releases, and integrations for training or inference components including Slime and SGLang. It can be installed from PyPI as `reef-infra` or from source; artifact and checkpoint management requires Git LFS.
SlopMonster is a Python linter for detecting formulaic or suspicious AI-generated writing in HTML pages, Markdown, plain text, landing pages, READMEs, emails, and scripts. Its scorer checks AI-associated vocabulary, sentence constructions, punctuation cadence, rule-of-three rhythms, and unsupported sales claims, assigns a score out of 5, and can fail a build when copy scores below 5/5. Its workflow scores the draft, performs a three-pass rewrite that removes suspect vocabulary and sentence shapes and replaces them with specific copy, sends the draft to a different model family for cleansing, and scores the result again. The cleansing script can use a Codex or Claude CLI, refuses to route a draft back to its own model family, and prints a prompt when no supported rival CLI is available. A GitHub Actions workflow is included as a reusable build gate; the scorer itself uses only Python's standard library, while the cleansing step requires an external AI CLI or manual prompt handling. The repository also includes regression tests, a catalogue of detection rules, rewrite principles, and worked before-and-after examples. It is distributed under the MIT license and can also be installed as an agent skill for Claude Code or used by other agents through its plain-Markdown instructions.
FrontierHarness Eval is an open evaluation benchmark and repository for measuring how coding-agent harness configurations affect software-engineering task results when the underlying model and runtime are held constant. It publishes benchmark definitions, task instructions and metadata, harness-version metadata, normalized aggregate and task-level results, and an agent-neutral workflow for evaluating additional harnesses. The benchmark runs the same model through multiple harness configurations on 30 tasks, using a frozen golden checkpoint and a fresh restore for each task. Its scripts provision the environment, run trials, record verifier-based pass/fail outcomes, measure cost, cache behavior, and speed, then normalize results and generate charts and reports. The repository also provides a CLI and skill-based workflow for evaluating a third-party harness, with infrastructure failures marked separately from task failures. The project focuses on terminal-based software-engineering tasks and states that its results may not generalize to other knowledge-work domains. The repository is published under the frontier-harness-eval organization, with isolated runtimes and checkpoint restores provided by Runta.
Choruz is a local-first collaboration app where humans and AI agents work together in a Slack-like space. Each agent runs a real CLI in its own workspace, preserving that CLI's models, tools, and session capabilities while allowing work to be handed to people or other agents through direct messages, groups, mentions, threads, tasks, and files. It provides isolated company and agent workspaces, dedicated directories or Git worktrees, an integrated terminal, file browser and editor, SSH runtime hosts, and browser-based remote control. It supports Claude Code, Codex, Pi, Grok, OpenCode, and webhook-driven external agents, with REST APIs, WebSocket synchronization, webhook agents, Slack and Telegram bridges, and optional plugins. Choruz is in pre-release development and requires Rust, Node.js, pnpm, PostgreSQL, and at least one supported agent CLI when run from source. Its source code and software documentation are licensed under the MIT License; visual assets have separate licensing records.
agent-memory is a local-first, agent-agnostic long-term memory runtime for AI agents. It stores memories as plain Markdown files, with a rebuildable SQLite index used as a cache rather than the source of truth. Conversation-boundary hooks trigger writes, while an independent sleep-time Manage pass consolidates, ages, and forgets memories by value; unattended deletion is limited to proposals that require confirmation, and superseded or archived material remains available. The runtime provides three recall paths: deterministic MEMORY.md injection at session start, BM25 search with progressive disclosure, and direct access through the store's directory tree using tools such as ls and grep. Memory records use frontmatter for names, abstracts, status, timestamps, links, weights, and provenance, with free-form Markdown bodies. Its CLI, MCP server, and host-agent hooks share the same validation, hash-diff, and reindexing path, and it can be used by agents that run shell commands, including Claude Code and Codex CLI. The repository documents a Python 3.12-or-later setup using uv, commands for initialization, recording, recall, rebuilding, sleep-time consolidation, and proposal approval, plus host setup for Claude Code and Codex. It does not include an LLM client; reasoning is delegated to the host agent's CLI, so the library requires no separate model keys or billing surface.
OrcaReplay is an open-source replay and debugging tool for AI-agent executions, built by the OrcaRouter.ai team. It records an agent run as a local trace, including model requests and responses, tool calls and results, shell exit codes and timing, filesystem snapshots, MCP traffic, and selected network activity. It runs as a local proxy and capture layer around an unmodified agent. A proxy records model traffic, while PATH shims, JSON-RPC interception, filesystem snapshots, fetch hooks, and optional TLS interception capture effects that model transcripts cannot show. The resulting timeline can be viewed, exported, queried through JSON or MCP, and analyzed with recorded versus inferred causal edges. The same recorded stream supports offline replay with the network blocked, forking from a derived checkpoint onto another model, and comparison of several models with an optional verification command. Replay uses the recorded conversation and workspace state; interactive terminal sessions can be approximate because prompts and interactive-only tools may not exist on the wire. The CLI is installed with npm, requires Node 20 or newer, stores runs under `.orca/runs/`, and the code, CLI, viewer, adapters, and trace format implementation are released under Apache-2.0; the trace specification is CC BY 4.0. Traces are local and redacted on write, but the project describes them as sensitive and treats redaction as best-effort.
Fable51 Worlds is an open-source, AI-assisted world-generation project that turns a text brief, photograph, or video into a walkable browser scene implemented as a pure Three.js application. Claude Fable 5.1 agent swarms research the location, collecting map geometry, elevation, transit, street information, and storefront data; generation scripts produce procedural façades, street furniture, vehicles, vegetation, fixtures, and other assets; and the runtime assembles terrain, streets, buildings, props, crowds, and traffic from JSON specifications without a game engine, proprietary 3D tiles, or downloaded meshes. Verification uses Playwright to drive the application, capture fixed viewpoints, and compare them with photographs from corresponding locations. Independent reviewer agents covering architecture, geography, technical art, and interaction produce reports for subsequent fix cycles, and each world includes a quality-assurance report.
BoardUI is a React design system for agentic interfaces and dashboards, combining AI-product components such as chat, thinking indicators, streaming agent logs, task lists, web-search trails, composers, and an agent runtime with dashboard components such as tables, charts, forms, cards, navigation, authentication, and design tokens. Its CLI copies individual components and their dependencies, or the complete free catalog, into a Next.js project as source files rather than installing a runtime dependency, so developers can modify the code directly. The system uses React Aria Components for accessible controls, Tailwind CSS v4, semantic design tokens, light and dark theme classes without runtime CSS, and Figma-based typography and styling rules. The repository includes a working AI chat application with a streaming chat endpoint and an agent runtime that accepts provider keys for supported services or an OpenAI-compatible server, reads keys server-side, and streams replies to the chat UI. It also provides an MCP server for browsing and installing components, along with agent rules specifying design tokens, type scale, and conventions.
Kitter is a local-first skill manager for agent skills across project folders. It keeps one canonical copy of each skill in a shared library, links skills into individual projects according to their needs, reports which skills are loaded, and estimates their context cost.
Higgsfield for Blender is an AI add-on that brings Higgsfield's generation tools into Blender, generating and importing 3D scenes, meshes, rigs, textures, images, and video. It adds a floating viewport bar with tabs for scene building, 3D models, character animation, images, video, cameras, and assets; results can be inserted into the open scene as editable geometry, planes, materials, rigs, keyframes, or video outputs. Scene Builder creates editable objects, layouts, and lighting from a prompt. The character-animation workflow produces a fitted weighted rig with timeline keyframes. An optional Higgsfield Bridge MCP endpoint at https://bridge.higgsfield.ai/mcp connects an external AI agent such as Claude to the add-on for operations including building blockouts, generating meshes at the 3D cursor, and requesting Seedance renders in the open Blender scene. Generation runs on Higgsfield's servers, requires an internet connection and a signed-in Higgsfield account, uses the account's existing Higgsfield credits, and supports Blender 4.2 through 5.1.
Context Mode is an MCP server and plugin for AI coding agents that reduces context-window usage by routing large tool outputs through sandboxed subprocesses. Its execution tools run code in supported languages, process files, fetch and analyze URLs, and return selected stdout or search results instead of exposing raw logs, snapshots, API responses, or file contents to the conversation. The project reports up to 98% context reduction in its benchmarks. It also provides session continuity across supported agents. Hooks capture tool calls, edits, prompts, decisions, errors, and other session events in SQLite; before compaction, the system builds a prioritized resume snapshot, and after compaction or session resumption it retrieves relevant events through SQLite FTS5 search with BM25 ranking. Content indexing uses heading-aware chunking, stemming, trigram matching, reciprocal-rank fusion, proximity reranking, fuzzy correction, and smart snippets. The MCP interface includes execution, batching, indexing, search, fetching, statistics, diagnostics, upgrade, and purge tools. Context Mode supports multiple coding-agent platforms through MCP servers, native plugins, and platform-specific hooks, with automatic routing enforcement where hooks are available and instruction files for platforms without them. The repository states that processing and SQLite storage are local, with no account or telemetry requirement. It is licensed under the Elastic License 2.0, which permits use, modification, and distribution but restricts offering the software as a hosted or managed service and removing licensing notices.
OpenHands is a self-hosted platform for running autonomous coding agents that work on GitHub issues. It can use either remote model providers or language models installed locally.
OpenAI Plugins is a curated repository of Codex plugin examples, maintained as a directory of independently packaged plugins and marketplace manifests. Each plugin has a required `.codex-plugin/plugin.json` manifest and may include skills, app and MCP configuration, agents, commands, hooks, assets, and other supporting files. The repository's standard and API-key-login marketplaces index the available plugins, including examples for Figma, Notion, iOS and macOS development, web apps, Expo, Netlify, Remotion, and Google Slides.
AIRUNCODE is a local-first agent runtime for running coding agents on a user's own computer. It supports voice-driven code-agent tasks and access to AI providers at their origin pricing, while its site describes execution as local.
Routines by Databox is a scheduling feature for recurring analytics and reporting. Users save an analysis as a reusable Skill, schedule that Skill to run daily, weekly, or monthly, and receive the results by email, Slack, or in Databox. Databox describes Routines as running Skills automatically and delivering the results, alongside AI agents that can delegate broader analysis-to-action workflows under user oversight.
Tadata is an AI employee for Slack that handles research and repetitive sales and go-to-market operations work. It can prepare call briefings, research prospects and companies, draft outreach and follow-up emails, fill CRM fields, monitor relevant LinkedIn activity, and suggest recurring automations. It connects to tools such as HubSpot, Attio, Notion, Linear, GitHub, Google Sheets, Gmail, and Google Calendar, as well as external web sources and MCP or API integrations. Tadata runs pre-built agents or builds an agent from a described process. It observes repeated work, learns team preferences, and requests approval before automating tasks or sending messages; connected-tool access is scoped to the user's authorization. The service is presented as model-agnostic, with portable automations, preferences, exceptions, and agent behaviors that can be exported and versioned.
Speakeasy is an AI control plane for discovering, securing, and governing enterprise AI agents, MCP servers, Skills, and AI applications. It maintains a catalog of approved AI tools, gives each agent an identity through the organization's existing identity provider, and scopes access by team and role. Every prompt, response, and tool call passes through the control plane for real-time policy checks before reaching internal APIs, MCP servers, or SaaS tools. Policies can distinguish read and write access and individual tools; violations are blocked, while allow-or-deny decisions are recorded in an audit trail. The platform is designed to detect and quarantine unapproved MCP connections and to block threats such as prompt injection, PII exposure, and leaked credentials in flight. Speakeasy deploys through an organization's existing MDM, including Jamf and Intune, and integrates with SAML/OIDC identity providers. The page states that it is SOC 2 Type II audited, ISO 27001 certified, GDPR compliant, and HIPAA ready.
dif.sh is a developer tool for storing feature flags, A/B tests, holdouts, and staged rollouts as Markdown files in a Git repository. Each experiment uses frontmatter for its status, audience, variants, metrics, guardrails, and exclusion group, while Git provides versioning and pull requests provide review and audit history. Its CLI initializes the repository structure, drafts experiments from surface-specific learning logs, validates configurations, previews assignments, builds a typed client, and archives concluded tests. The build resolves exclusion groups and fails on conflicting live experiments, generates a client that evaluates variants in the application without a production service call, and refreshes dif/context.json for coding agents. Audience rules declare runtime attributes such as country, plan, or returning-visitor status without committing customer lists. The optional Dif Cloud integration can calculate lift through dif.track(); alternatively, custom event handlers can forward results to other analytics systems. The project provides an npm CLI package and an SDK, and its site describes a no-signup workflow.
Reflexio is a learning platform for AI agents that turns user corrections, failed paths, and successful outcomes into reusable behavioral changes without retraining the underlying model. Its SDK and integration loop publish interaction outcomes, extract actionable feedback, store learned behaviors, and retrieve only relevant learnings during later inference; integrations are available through Python, REST, a CLI, and a portable coding-agent skill. Reflexio evaluates learnings against control responses and user-defined success criteria, tracking whether they solved the user's problem, required correction, or escalated to a human. It makes learnings auditable and revocable: they can be reviewed, rewritten, approved, rejected, or deleted, with rejected learnings removed from retrieval. Background processes consolidate duplicate signals and resolve conflicts or outdated lessons, while evidence from later sessions is used to revise learnings. The service supports managed, bring-your-own-key, customer-owned database, bring-your-own-cloud, and fully self-hosted deployments. The page identifies an official repository at github.com/ReflexioAI/reflexio and states that the same API can be used across deployment modes.
HyperProbe is an AI-native production debugger for investigating live application state without code changes, redeployment, or service restarts. It reads logs and distributed traces, uses a coding agent to locate a suspected file and line, and places a temporary read-only virtual breakpoint there. When the breakpoint fires on live traffic, it captures the variable state asynchronously without pausing requests, then uses the evidence to confirm a root cause. Probes are non-blocking, cannot write memory or execute code, disappear after capture, and are recorded in an immutable audit trail. The service is designed for silent failures, exceptions whose causes occur earlier in the call chain, incorrect behavior without thrown errors, race conditions, third-party contract changes, and business-metric failures. It supports JavaScript, TypeScript, Java, Kotlin, Python, and Ruby, and can run in a managed cloud, self-hosted deployment, or private VPC. The page states that it integrates with PagerDuty, Datadog, Slack, Cursor, Claude Code, Codex, and Opencode, with approval gates and agent-side PII redaction.
Experiential Labs is an open-source AI gateway that provides an OpenAI-compatible endpoint for hosted model providers, customer-owned provider keys, private GPUs, and self-hosted models. It routes requests through a single API and key while handling provider access, model selection, failover, and streaming responses. Its optional intelligence layer monitors traffic to identify model switches, improve cache hit rates, and route requests to customer-owned fine-tuned models. The service also provides organization-wide usage attribution, request logs, spend reporting, access controls, model allowlists, and per-key spending caps scoped by team, person, agent, or tool. It can be used as a hosted gateway or self-hosted, with hosted inference and Pro offered alongside the open-source gateway.