580 tools and products — trending open source, and what gets used in AI and other work.
Miles is an enterprise-facing reinforcement learning framework for large-scale post-training of large language and vision-language models. It pairs SGLang for high-throughput, agentic rollouts with Megatron-LM for scalable training, and also provides a PyTorch FSDP2 backend for Hugging Face implementations. Its asynchronous architecture decouples rollout and training workers, supports configurable on- and off-policy schedules, and updates rollout engines in-loop through peer-to-peer RDMA weight transfer. The framework includes token-in-token-out data flow, Rollout Routing Replay for replaying mixture-of-experts routing decisions during training, fault-tolerant recovery of failed SGLang engines, low-precision training with MXFP8 and NVFP4 alongside FP8, INT4 QAT, BF16, and FP16, and LoRA or multi-LoRA training. It supports reinforcement-learning recipes including GRPO, GSPO, PPO, and REINFORCE++, as well as supervised fine-tuning, on-policy distillation, agentic environments, and diffusion-model training. Miles was forked from slime and integrates SGLang, Megatron-LM, and torch_memory_saver; the repository is released as an open-source project, though the provided page text does not state its license.
Speechify is an AI voice and text-to-speech platform that reads written material aloud and also provides speech-to-text, conversational, and business-focused AI tools. Its SpeechifyAI service offers text-to-speech and realtime voice-agent APIs through a single interface, with streaming synthesis, zero-shot voice cloning from a consented reference clip, SSML-based emotion control, and multilingual output for English, German, Mexican Spanish, French, Italian, and Brazilian Portuguese. The service includes the Simba model family, which the site describes as streaming-native speech models designed to model voice identity, expression, and language for realtime conversation. Speechify was co-founded by Cliff Weitzman, who initially built it to consume written material through audio.
Chinese Patent Skill is an MIT-licensed open-source agent skill for Chinese patent workflows, covering invention, utility-model, and industrial-design patents. It mines patentable points from project materials such as Markdown, code, DOCX, PPTX, and optionally STEP/CAD files; performs novelty, bibliographic, and prior-art searches with China National Intellectual Property Administration sources preferred; drafts patent disclosure documents; rewrites disclosures into claims, specifications, and abstracts; produces Mermaid diagrams or planned patent figures, black-and-white drawings, and optional editable DOCX exports; explains published patents in plain language; and assists with patent-office examination responses and policy briefs. It can read technical and product drawings, extract outlines and component references, derive multiple views from CAD models, and organize utility-model and design workflows with schemas, figure plans, views, and component numbering. Patent explanations can be stored in Obsidian as linked notes, graphs, and canvases, while disclosure work supports self-checking, correction, iterative updates, timestamped drafts, multiple saved versions, and conversation records.
OpenWhispr is an open-source, cross-platform desktop voice-to-text application for macOS, Windows, and Linux. A global hotkey captures speech and inserts the resulting text at the cursor in another application; dictation can use local Whisper or NVIDIA Parakeet speech-to-text engines, where audio remains on the device, or cloud providers through user-supplied keys. It also supports dictation translation, AI-agent commands, meeting transcription with speaker diarization and voice fingerprinting, audio and video transcription, and searchable notes with semantic search and optional cloud sync. The application is built with Electron, React, TypeScript, SQLite, whisper.cpp, and sherpa-onnx, and provides an API and MCP server for programmatic access to notes and transcriptions. The repository describes it as having no data collection or telemetry and distributes installers for the three supported desktop platforms. It is licensed under the MIT license.
Reverify is an AI-agent verification toolkit for reverse engineering and related code analysis. It places deterministic tools between an agent and its claims: the model proposes hypotheses about a binary or candidate implementation, while parsers, disassemblers, pattern scanners, emulators, equivalence checks, and optional analysis engines compare them with ground truth and return evidence-backed VERIFIED, REFUTED, INCONCLUSIVE, OBSERVED, or INVALIDATED results. Its pure-Python core handles PE, ELF, and Mach-O parsing, x86/x64 and ARM/ARM64 disassembly, byte and pattern inspection, CPU micro-emulation, Protobuf/TLV dissection, and Frida hook generation; optional installations add Capstone, Unicorn, LIEF, Z3, or angr for deeper analysis. The project exposes the verification loop through a command-line interface and an MCP server that agents such as Claude Code and Cursor can call. It also records verified, observed, proved, and refuted results in a content-keyed local ledger, allowing grounded facts and negative findings to survive context resets, compaction, and new sessions. A rollover and orchestration system hands sessions off through files rather than model-written summaries, keeping verified facts separate from unverified notes. Reverify includes claim types for bytes, typed reads, instructions, patterns, emulation, behavioral or formal equivalence, imports, exports, sections, and—when angr is installed—functions, calls, references, and reachability. It is distributed from PyPI and licensed under the MIT License. The repository states that it is intended for authorized reverse engineering, malware analysis, CTF work, interoperability research, and software the user owns or is permitted to analyze.
ffmpeg-skill is an agent skill that gives Claude Code, Cursor, Codex, and other agents a local video- and audio-editing workflow built on FFmpeg and ffprobe. It uses a probe → edit, preferring stream copying where possible → check → verify process, with typed Python scripts rather than shell strings. Its 21 tools cover cutting, joining, silence removal, duration and aspect-ratio fitting, captions, overlays, graphics, HDR-to-SDR conversion, LUTs, audio cleanup, loudness, synchronization with drift correction, multicamera editing, delivery checks, project rendering, and batch processing. Each tool supports structured results and dry runs, and the machine-readable contract generates an MCP interface whose tool definitions are derived from the scripts. The skill runs locally without cloud services, API keys, or Python dependencies beyond the standard library; it requires FFmpeg 5.0 or later. It is a tool for local, agent-controlled workflows that probe, edit, verify, and render video.
Commerce Agents is Anthropic's reference implementation for building shopping agents for customers and merchant agents for back-office staff with Claude. It defines each agent through prompts, skills, tool contracts, grounding, memory, approval gates, and backend interfaces, and provides runnable examples for retail, travel, telecom, and entertainment. The shopping agent searches and compares products, plans purchases, fills carts, answers order and policy questions, and maintains customer memory. The merchant agent analyzes performance, manages listings and inventory alerts, handles pricing and promotions, and drafts campaigns; its writes are staged until a host approves them. The agents can run through the Anthropic Messages API, Claude Agent SDK, or Managed Agents, and the repository provides a Claude Code plugin for scaffolding or reviewing commerce agents.
Reef is open-source continual-learning infrastructure for self-improving AI agents, developed by Human-Agent-Society. It connects agent inference, interaction records, feedback, learning jobs, evaluation, and versioned delivery, supporting updates to model weights as well as agent harnesses such as prompts, rules, and skills. Each learning cycle has four stages: Reef serves requests and records interactions; matches later scores or structured feedback to those records; produces a candidate update from eligible records; and evaluates the candidate against a configured selection policy before publishing it. Rejected candidates leave the previous release serving, while accepted updates are committed and delivered without restarting the service. The project provides an HTTP API with OpenAI- and Anthropic-compatible inference endpoints, feedback reporting, scenario-based releases, and integrations for training or inference components including Slime and SGLang. It can be installed from PyPI as `reef-infra` or from source; artifact and checkpoint management requires Git LFS.
SlopMonster is a Python linter for detecting formulaic or suspicious AI-generated writing in HTML pages, Markdown, plain text, landing pages, READMEs, emails, and scripts. Its scorer checks AI-associated vocabulary, sentence constructions, punctuation cadence, rule-of-three rhythms, and unsupported sales claims, assigns a score out of 5, and can fail a build when copy scores below 5/5. Its workflow scores the draft, performs a three-pass rewrite that removes suspect vocabulary and sentence shapes and replaces them with specific copy, sends the draft to a different model family for cleansing, and scores the result again. The cleansing script can use a Codex or Claude CLI, refuses to route a draft back to its own model family, and prints a prompt when no supported rival CLI is available. A GitHub Actions workflow is included as a reusable build gate; the scorer itself uses only Python's standard library, while the cleansing step requires an external AI CLI or manual prompt handling. The repository also includes regression tests, a catalogue of detection rules, rewrite principles, and worked before-and-after examples. It is distributed under the MIT license and can also be installed as an agent skill for Claude Code or used by other agents through its plain-Markdown instructions.
Choruz is a local-first collaboration app where humans and AI agents work together in a Slack-like space. Each agent runs a real CLI in its own workspace, preserving that CLI's models, tools, and session capabilities while allowing work to be handed to people or other agents through direct messages, groups, mentions, threads, tasks, and files. It provides isolated company and agent workspaces, dedicated directories or Git worktrees, an integrated terminal, file browser and editor, SSH runtime hosts, and browser-based remote control. It supports Claude Code, Codex, Pi, Grok, OpenCode, and webhook-driven external agents, with REST APIs, WebSocket synchronization, webhook agents, Slack and Telegram bridges, and optional plugins. Choruz is in pre-release development and requires Rust, Node.js, pnpm, PostgreSQL, and at least one supported agent CLI when run from source. Its source code and software documentation are licensed under the MIT License; visual assets have separate licensing records.
agent-memory is a local-first, agent-agnostic long-term memory runtime for AI agents. It stores memories as plain Markdown files, with a rebuildable SQLite index used as a cache rather than the source of truth. Conversation-boundary hooks trigger writes, while an independent sleep-time Manage pass consolidates, ages, and forgets memories by value; unattended deletion is limited to proposals that require confirmation, and superseded or archived material remains available. The runtime provides three recall paths: deterministic MEMORY.md injection at session start, BM25 search with progressive disclosure, and direct access through the store's directory tree using tools such as ls and grep. Memory records use frontmatter for names, abstracts, status, timestamps, links, weights, and provenance, with free-form Markdown bodies. Its CLI, MCP server, and host-agent hooks share the same validation, hash-diff, and reindexing path, and it can be used by agents that run shell commands, including Claude Code and Codex CLI. The repository documents a Python 3.12-or-later setup using uv, commands for initialization, recording, recall, rebuilding, sleep-time consolidation, and proposal approval, plus host setup for Claude Code and Codex. It does not include an LLM client; reasoning is delegated to the host agent's CLI, so the library requires no separate model keys or billing surface.
OrcaReplay is an open-source replay and debugging tool for AI-agent executions, built by the OrcaRouter.ai team. It records an agent run as a local trace, including model requests and responses, tool calls and results, shell exit codes and timing, filesystem snapshots, MCP traffic, and selected network activity. It runs as a local proxy and capture layer around an unmodified agent. A proxy records model traffic, while PATH shims, JSON-RPC interception, filesystem snapshots, fetch hooks, and optional TLS interception capture effects that model transcripts cannot show. The resulting timeline can be viewed, exported, queried through JSON or MCP, and analyzed with recorded versus inferred causal edges. The same recorded stream supports offline replay with the network blocked, forking from a derived checkpoint onto another model, and comparison of several models with an optional verification command. Replay uses the recorded conversation and workspace state; interactive terminal sessions can be approximate because prompts and interactive-only tools may not exist on the wire. The CLI is installed with npm, requires Node 20 or newer, stores runs under `.orca/runs/`, and the code, CLI, viewer, adapters, and trace format implementation are released under Apache-2.0; the trace specification is CC BY 4.0. Traces are local and redacted on write, but the project describes them as sensitive and treats redaction as best-effort.
Fable51 Worlds is an open-source, AI-assisted world-generation project that turns a text brief, photograph, or video into a walkable browser scene implemented as a pure Three.js application. Claude Fable 5.1 agent swarms research the location, collecting map geometry, elevation, transit, street information, and storefront data; generation scripts produce procedural façades, street furniture, vehicles, vegetation, fixtures, and other assets; and the runtime assembles terrain, streets, buildings, props, crowds, and traffic from JSON specifications without a game engine, proprietary 3D tiles, or downloaded meshes. Verification uses Playwright to drive the application, capture fixed viewpoints, and compare them with photographs from corresponding locations. Independent reviewer agents covering architecture, geography, technical art, and interaction produce reports for subsequent fix cycles, and each world includes a quality-assurance report.
BoardUI is a React design system for agentic interfaces and dashboards, combining AI-product components such as chat, thinking indicators, streaming agent logs, task lists, web-search trails, composers, and an agent runtime with dashboard components such as tables, charts, forms, cards, navigation, authentication, and design tokens. Its CLI copies individual components and their dependencies, or the complete free catalog, into a Next.js project as source files rather than installing a runtime dependency, so developers can modify the code directly. The system uses React Aria Components for accessible controls, Tailwind CSS v4, semantic design tokens, light and dark theme classes without runtime CSS, and Figma-based typography and styling rules. The repository includes a working AI chat application with a streaming chat endpoint and an agent runtime that accepts provider keys for supported services or an OpenAI-compatible server, reads keys server-side, and streams replies to the chat UI. It also provides an MCP server for browsing and installing components, along with agent rules specifying design tokens, type scale, and conventions.
Kitter is a local-first skill manager for agent skills across project folders. It keeps one canonical copy of each skill in a shared library, links skills into individual projects according to their needs, reports which skills are loaded, and estimates their context cost.
Higgsfield for Blender is an AI add-on that brings Higgsfield's generation tools into Blender, generating and importing 3D scenes, meshes, rigs, textures, images, and video. It adds a floating viewport bar with tabs for scene building, 3D models, character animation, images, video, cameras, and assets; results can be inserted into the open scene as editable geometry, planes, materials, rigs, keyframes, or video outputs. Scene Builder creates editable objects, layouts, and lighting from a prompt. The character-animation workflow produces a fitted weighted rig with timeline keyframes. An optional Higgsfield Bridge MCP endpoint at https://bridge.higgsfield.ai/mcp connects an external AI agent such as Claude to the add-on for operations including building blockouts, generating meshes at the 3D cursor, and requesting Seedance renders in the open Blender scene. Generation runs on Higgsfield's servers, requires an internet connection and a signed-in Higgsfield account, uses the account's existing Higgsfield credits, and supports Blender 4.2 through 5.1.
Context Mode is an MCP server and plugin for AI coding agents that reduces context-window usage by routing large tool outputs through sandboxed subprocesses. Its execution tools run code in supported languages, process files, fetch and analyze URLs, and return selected stdout or search results instead of exposing raw logs, snapshots, API responses, or file contents to the conversation. The project reports up to 98% context reduction in its benchmarks. It also provides session continuity across supported agents. Hooks capture tool calls, edits, prompts, decisions, errors, and other session events in SQLite; before compaction, the system builds a prioritized resume snapshot, and after compaction or session resumption it retrieves relevant events through SQLite FTS5 search with BM25 ranking. Content indexing uses heading-aware chunking, stemming, trigram matching, reciprocal-rank fusion, proximity reranking, fuzzy correction, and smart snippets. The MCP interface includes execution, batching, indexing, search, fetching, statistics, diagnostics, upgrade, and purge tools. Context Mode supports multiple coding-agent platforms through MCP servers, native plugins, and platform-specific hooks, with automatic routing enforcement where hooks are available and instruction files for platforms without them. The repository states that processing and SQLite storage are local, with no account or telemetry requirement. It is licensed under the Elastic License 2.0, which permits use, modification, and distribution but restricts offering the software as a hosted or managed service and removing licensing notices.
OpenHands is a self-hosted platform for running autonomous coding agents that work on GitHub issues. It can use either remote model providers or language models installed locally.
AIRUNCODE is a local-first agent runtime for running coding agents on a user's own computer. It supports voice-driven code-agent tasks and access to AI providers at their origin pricing, while its site describes execution as local.
Routines by Databox is a scheduling feature for recurring analytics and reporting. Users save an analysis as a reusable Skill, schedule that Skill to run daily, weekly, or monthly, and receive the results by email, Slack, or in Databox. Databox describes Routines as running Skills automatically and delivering the results, alongside AI agents that can delegate broader analysis-to-action workflows under user oversight.
Tadata is an AI employee for Slack that handles research and repetitive sales and go-to-market operations work. It can prepare call briefings, research prospects and companies, draft outreach and follow-up emails, fill CRM fields, monitor relevant LinkedIn activity, and suggest recurring automations. It connects to tools such as HubSpot, Attio, Notion, Linear, GitHub, Google Sheets, Gmail, and Google Calendar, as well as external web sources and MCP or API integrations. Tadata runs pre-built agents or builds an agent from a described process. It observes repeated work, learns team preferences, and requests approval before automating tasks or sending messages; connected-tool access is scoped to the user's authorization. The service is presented as model-agnostic, with portable automations, preferences, exceptions, and agent behaviors that can be exported and versioned.
Speakeasy is an AI control plane for discovering, securing, and governing enterprise AI agents, MCP servers, Skills, and AI applications. It maintains a catalog of approved AI tools, gives each agent an identity through the organization's existing identity provider, and scopes access by team and role. Every prompt, response, and tool call passes through the control plane for real-time policy checks before reaching internal APIs, MCP servers, or SaaS tools. Policies can distinguish read and write access and individual tools; violations are blocked, while allow-or-deny decisions are recorded in an audit trail. The platform is designed to detect and quarantine unapproved MCP connections and to block threats such as prompt injection, PII exposure, and leaked credentials in flight. Speakeasy deploys through an organization's existing MDM, including Jamf and Intune, and integrates with SAML/OIDC identity providers. The page states that it is SOC 2 Type II audited, ISO 27001 certified, GDPR compliant, and HIPAA ready.
Reflexio is a learning platform for AI agents that turns user corrections, failed paths, and successful outcomes into reusable behavioral changes without retraining the underlying model. Its SDK and integration loop publish interaction outcomes, extract actionable feedback, store learned behaviors, and retrieve only relevant learnings during later inference; integrations are available through Python, REST, a CLI, and a portable coding-agent skill. Reflexio evaluates learnings against control responses and user-defined success criteria, tracking whether they solved the user's problem, required correction, or escalated to a human. It makes learnings auditable and revocable: they can be reviewed, rewritten, approved, rejected, or deleted, with rejected learnings removed from retrieval. Background processes consolidate duplicate signals and resolve conflicts or outdated lessons, while evidence from later sessions is used to revise learnings. The service supports managed, bring-your-own-key, customer-owned database, bring-your-own-cloud, and fully self-hosted deployments. The page identifies an official repository at github.com/ReflexioAI/reflexio and states that the same API can be used across deployment modes.
HyperProbe is an AI-native production debugger for investigating live application state without code changes, redeployment, or service restarts. It reads logs and distributed traces, uses a coding agent to locate a suspected file and line, and places a temporary read-only virtual breakpoint there. When the breakpoint fires on live traffic, it captures the variable state asynchronously without pausing requests, then uses the evidence to confirm a root cause. Probes are non-blocking, cannot write memory or execute code, disappear after capture, and are recorded in an immutable audit trail. The service is designed for silent failures, exceptions whose causes occur earlier in the call chain, incorrect behavior without thrown errors, race conditions, third-party contract changes, and business-metric failures. It supports JavaScript, TypeScript, Java, Kotlin, Python, and Ruby, and can run in a managed cloud, self-hosted deployment, or private VPC. The page states that it integrates with PagerDuty, Datadog, Slack, Cursor, Claude Code, Codex, and Opencode, with approval gates and agent-side PII redaction.
Experiential Labs is an open-source AI gateway that provides an OpenAI-compatible endpoint for hosted model providers, customer-owned provider keys, private GPUs, and self-hosted models. It routes requests through a single API and key while handling provider access, model selection, failover, and streaming responses. Its optional intelligence layer monitors traffic to identify model switches, improve cache hit rates, and route requests to customer-owned fine-tuned models. The service also provides organization-wide usage attribution, request logs, spend reporting, access controls, model allowlists, and per-key spending caps scoped by team, person, agent, or tool. It can be used as a hosted gateway or self-hosted, with hosted inference and Pro offered alongside the open-source gateway.
GitWarren is a desktop application for reviewing local Git changes made by coding agents before they are committed. It reads the worktree directly and combines staged, unstaged, and untracked files into one diff for local review, with comments and threaded discussions that can be edited or resolved. An MCP server over stdio lets Claude Code, Codex, or another MCP client open reviews, inspect discussions, reply in threads, comment on lines, and resolve comments; agent messages are attributed and separate MCP sessions distinguish concurrent agents. Reviews and comments are stored in a local SQLite file, while repository state is read from Git when displayed. The application requires no account, runs on macOS, Windows, and Linux, and is free and open source under GPL-3.0.
Tucky is a native macOS notes app that keeps notes in a thin stripe at the edge of the screen. Notes are stored locally and encrypted on disk, with support for live Markdown, checkbox tasks, global shortcuts, clean exports, and local-only files. Its optional AI agent can search, write, rewrite, and ask questions about notes from the edge or through the global ⌥Space shortcut. The app can connect to Gmail, Google Calendar, Google Docs, Google Sheets, GitHub, and Notion; voice dictation is supported in more than 60 languages. The page says note titles accompany questions by default, while note bodies are read only when allowed. The free plan supports five notes without AI, voice, or connectors. The Plus plan is listed at $4 per month and adds unlimited notes, AI models, voice, and connectors. Tucky runs on macOS 15 or later on Apple silicon or Intel.
TeamAI is an open-source CLI for managing a team's skills, rules, documentation, agents, hooks, MCP configuration, environment settings, and knowledge across AI coding tools such as Claude Code, Codex, Cursor, CodeBuddy, WorkBuddy, and OpenCode. It is installed with npm as `teamai-cli` and uses a shared Git repository as the team's source of truth. The CLI distributes resources through an administrative `push` → review and merge → `pull` workflow. Session-start hooks can run `teamai pull` to synchronize the shared harness into project- or user-scoped local AI-tool directories; roles, tags, and subscribed source repositories control which resources members receive. TeamAI also provides a knowledge layer: `teamai import` and `teamai codebase` build a searchable team knowledge base and codebase graph, while `teamai recall` uses BM25 search with graph-boosted reranking and can deploy a recall subagent into supported AI tools. Code relationships are extracted with a WebAssembly tree-sitter parser for TypeScript/JavaScript, Python, and Go, with regex-based heuristic extraction as a fallback and for other languages. Additional commands support friction-based learning suggestions, privacy-scrubbed session summaries, maintenance of stale or low-confidence knowledge, usage digests, a web dashboard, member and role management, package and plugin installation, diagnostics, and uninstallation. The repository is maintained by Tencent and contributors and is licensed under MIT.
PI-Desktop is a local-first desktop workspace for AI coding agents, built with Electron, a Rust host core, and the pi Agent Harness. It lets users open local projects, connect cloud or local models through providers such as OpenAI-compatible APIs, manage projects and long-running sessions, and review file changes and command output. Its architecture separates the React renderer, Electron desktop orchestration, Rust-managed permissions, filesystem access, SQLite persistence and secrets, and a pi Agent sidecar responsible for the agent loop, model interaction, and streaming. Agent, Plan, and Goal modes provide different approval boundaries; agents can read and edit files, run commands, and delegate exploration, implementation, research, testing, or review to background Subagents. The workspace supports MCP servers, Skills, installable Plugins, model and provider switching, session imports, notifications, context checkpoints, and a plugin marketplace. Conversations are stored locally as JSONL with a SQLite index, settings and logs remain on the machine, credentials use the operating-system keychain, and model requests go directly to the configured provider or endpoint. Plugin processes are permission-gated and isolated from the renderer, but plugins remain user-trusted code rather than a complete operating-system sandbox. It is an early-preview cross-platform application distributed for macOS, Windows, and Linux under the GNU Lesser General Public License v3.0.
Vals is a code-generation evaluation product that uses a company's GitHub codebase to build an internal coding benchmark. It evaluates coding models and agents against private, continuously evolving tasks to identify which systems perform best on the company's work and which offer the strongest return on investment. The platform is also described as supporting pre-release testing and evaluation of agentic coding work that may unfold over hours, days, or weeks, while accounting for factors such as token cost and latency.
Clodds is a self-hosted, open-source AI trading terminal and autonomous trading agent for prediction markets, cryptocurrency spot and perpetual futures, Solana and EVM decentralized exchanges, token launches, and Bittensor mining. It is built around Claude and can be accessed through a local WebChat interface, CLI, MCP server, or 21 messaging channels. Its architecture combines four agents—main, trading, research, and alerts—with trading skills, market-data tools, semantic memory, strategy execution, and a unified risk layer. The documented capabilities include arbitrage detection, whale and copy trading, DCA bots, backtesting, order and portfolio management, circuit breakers, VaR/CVaR, volatility-regime detection, stress testing, Kelly sizing, daily loss limits, and a kill switch. It connects to prediction-market platforms, futures exchanges, Solana protocols, and EVM networks, and also provides agent-oriented forum, marketplace, token-launch, compute, and x402 USDC payment features. The project is distributed as an npm package or from source, stores local data under ~/.clodds/, supports SQLite, LanceDB, and PostgreSQL components, and is licensed under the MIT license. Its documented quick start is `npm install -g clodds` followed by `clodds onboard`; trading integrations require the relevant credentials and wallet configuration.
Hyperresearch is an MIT-licensed Python command-line research harness that extends Claude Code with an agent-driven, persistent research knowledge base. It classifies a query into light, full, or opt-in dissertation work and runs a tier-adaptive pipeline covering query decomposition, source search and fetching, contradiction and locus analysis, depth investigation, evidence digestion, multi-angle drafting, synthesis, adversarial criticism, targeted gap fetching, surgical patching, citation checks, and readability review. The pipeline stores fetched sources as Markdown notes with a rebuildable SQLite index, provenance links, lifecycle statuses, full-text or optional semantic search, source-quality and retraction metadata, and resumable per-run manifests. It can fetch PDFs, seek disclosed open-access replacements for thin or blocked academic sources through Unpaywall and Europe PMC, and apply structural checks for citation bindings, quote integrity, provenance, retractions, and patch-only editing. The vault can also be accessed through an MCP server or a dependency-free local web UI. It is installed with pip for Python 3.11–3.13 and is intended to run with Claude Code; model assignments and research scale can be configured per project.
OpenResearch is a local-first workspace for research agents and autoresearch, available as a desktop application and a macOS/Linux CLI. It turns Claude Code, Codex, or OpenCode into agents that can review literature, develop hypotheses, run experiments, and produce research artifacts. The workspace assigns independent agent sessions and isolated Git worktrees to parallel research directions. It tracks experiment variants in a Git-native experiment tree, archives each run against its recorded commit, and keeps logs, diffs, files, results, and artifacts tied to the work that produced them. Its autoresearch loop can propose an idea, modify code, launch an experiment, inspect evidence, and choose the next direction. The same committed source snapshot can run locally or through SSH, Slurm, Kubernetes, Ray, Hugging Face Jobs, Modal, Tinker, or managed OpenResearch compute. By default, projects and run data remain on the user's machine in a local SQLite store, with a browser dashboard served on localhost. An OpenResearch account is used for service-owned capabilities such as organizations and managed compute. The CLI also installs an OpenResearch skill into supported coding agents and provides commands for projects, runs, logs, experiments, discovery, and paper lookup.
YuE2 is an open music-generation model and Python pipeline that turns lyrics and a style prompt into a symbolic melody-and-chord plan and then renders that plan as a complete stereo song with vocals and accompaniment. It supports zero-shot covers by using a transcribed melody score, and composition editing by allowing the score, lyrics, style, harmony, melody, tempo, or form to be revised before rendering. Its AR–NAR Mixture-of-Transformers backbone predicts symbolic scores and semantic tokens, generates acoustic latents with flow matching, and uses a VAE to decode them into 48 kHz audio. The staged API exposes plan(), generate_semantic(), synthesize(), and decode() operations; the repository also provides an agent skill for song generation, transcription, cover creation, ABC-score editing, and listening comparisons. The repository provides the YuE2-3B model and requires Linux, Python 3.12, and an NVIDIA GPU with BF16 support and 24 GB of VRAM for the documented quick start. First-party code, documentation, and the agent skill use Apache 2.0, while the model weights use CC BY-NC 4.0. The original YuE implementation is preserved on the repository's YuE-v1 branch.
Worktrunk is a command-line interface for managing Git worktrees, designed for running AI coding agents in parallel. It addresses worktrees by branch name and computes their paths from a configurable template; its core commands switch to, list, merge, and remove worktrees, with an option to launch a command such as Claude Code after switching. It also provides hooks for local workflow automation, LLM-generated commit messages, merge and cleanup workflows, an interactive worktree picker with diff and log previews, shared build caches, CI status and AI-generated branch summaries, pull-request checkout, per-worktree development-server ports, and aliases with branch-scoped variables. It can be installed through Homebrew, Cargo, Winget, Arch Linux, Conda, or Pixi, and shell integration enables commands to change directories.
Dream Loop is an AI agent skill for building games, apps, and 3D scenes toward a visual target. It uses image generation to create a target screenshot, builds the scene—such as a Three.js scene in a browser—and has a separate vision-enabled critic compare the live screenshot with the target and provide feedback; the agent then iterates on the build and can optionally generate an improved target from the current state. Blender can be used for custom 3D modeling through Blender MCP or scripting. The skill can be installed with `npx skills add achimala/dream-loop` or cloned into an agent's skills directory, and requires an agent with image-generation access, vision input, and preferably subagents.
Bang Motion is an agent skill for building browser-based motion graphics such as product openers, promos, bumpers, channel intros, kinetic typography, lower thirds, and explainers. It produces a single index.html that plays in a browser and encodes anti-slide rules: scenes change through camera or world movement rather than section fades, subjects or shared worlds persist, motion occurs continuously, and text and numbers remain part of the scene. Its deterministic timeline makes each frame a pure function of time for exact scrubbing and frame-by-frame export, while a shared camera rig supports flowing-world and camera-through-collage modes. The skill includes five explainer starters—cartoon collage, visual journalism, white catalog, vintage sketch, and continuous action—along with configurable style briefs, background motion, surfaces, transitions, highlight shapes, voice-over re-timing, and optional Puppeteer and FFmpeg scripts for verification and MP4 export. It can be installed as a Claude Code plugin or copied into Claude Code, Codex CLI, Gemini CLI, Cursor, or another agent workflow; watching the output requires only a browser and internet connection. The repository is licensed under MIT.
Browser Use Pi is a TypeScript web agent built on Pi Mono. It combines a persistent V8 REPL with raw Chrome DevTools Protocol control: the agent writes JavaScript, uses accessibility-tree data and screenshots, and builds browser helpers as needed. It supports persistent sessions, saved logins, workspaces, streaming, hooks, typed results, follow-up tasks, and cloud browsers or a user's own Chrome; runs can be limited by steps, time, or cost, with partial work retained. Interaction highlights, recordings, and GIF exports can show the agent's work. It is distributed as the @browser_use/pi package for Node 22.19+ or Bun 1.3.14+, and its quickstart uses Browser Use Cloud without requiring a local Chrome installation.
Astra Advisor is a Codex plugin for capability-routed software delivery from DannyMac180. GPT-6 Astra acts as the primary architect and acceptance owner: it plans work, decides whether bounded independent tasks should be delegated, selects supported native subagent models and reasoning effort, and continues parent-session work while delegations run. It uses the exposed collaboration tool to route tasks among Sol, Terra, and Luna subagents, then inspects the complete diff, reruns requested checks, and sends substantial changes to a fresh read-only reviewer; work is accepted only after the reviewer returns "ship". Each delegation reports its selected and runtime-observed settings, and tasks can end with an API-equivalent cost receipt that compares observed token usage with a versioned Astra pricing snapshot. It can be installed as a plugin in a current Codex CLI or ChatGPT desktop app with plugins enabled. The README notes that ChatGPT Work cloud tasks currently cannot provide arbitrary model or effort controls, so Astra does not dispatch model-pinned requests there by default.
SuperAstra is a SNES-themed desktop companion that uses natural-language prompts to investigate and alter a running game through BizHawk. Its agent receives screenshots, cartridge identity data, emulator registers, memory scans, hardware-domain reads, controller probes, and checkpoint results; it can search cartridge data, identify game-memory structures, create routines or guarded patches, test changes, and retain cartridge-specific findings in a local notebook. Each committed mutation can be undone through emulator checkpoints, while named experiment checkpoints support controlled trials that restore the player's prior state. The prototype runs on Windows or Linux with Python and Tkinter, BizHawk using its BSNES SNES core, a ROM, and an OpenAI API key configured for the stated model; it also includes limited local Mario shortcuts that do not require an API key. It does not patch the original ROM file, create exportable ROM patches, expand ROM capacity, or guarantee success on unfamiliar games. The original source is MIT-licensed.
Agent Skills is Google's repository of installable skills for Google products and technologies, including Google Cloud. The skills provide Markdown-based procedures for tasks such as onboarding and authenticating to Google Cloud, deploying and managing AI agents, working with GKE, databases, networking, security, monitoring, analytics, advertising APIs, Firebase, Android, Dart, Flutter, and Google Maps Platform. Skills can be selected and installed with `npx skills add google/skills`; the repository also bundles plugins and MCP servers for agent harnesses including Claude Code, Codex, and Antigravity CLI. The repository accepts bug reports and contributions, and its skills are distributed under the Apache 2.0 license.
Distilly is an AI agent skill-generation tool by titanwings that turns messages, documents, interviews, and other supplied sources into source-grounded Person Profiles. It models observable experience, decision patterns, expression, and ways of working, rather than claiming to clone a person, and supports colleague, relationship, and celebrity profile families. Each generated profile is packaged as an installable Agent Skill for supported hosts. Its workflow includes family-specific source collection and analysis, a six-dimension research pipeline for celebrity profiles, incremental file analysis and merging, conversational corrections written to a correction layer, and automatic version archiving with rollback to earlier versions. Supported source types include Lark, DingTalk, Slack, public X posts, WeChat exports, PDFs, images, email, Markdown, and pasted text, with source-specific setup requirements. The project documents native local Skill discovery for Claude Code, Hermes, OpenClaw, Codex, DeepSeek Harness, Pi, Grok Build, and OpenCode. It is distributed from its GitHub repository under the MIT License and is described as a demo version.
Agent Reach is an open-source command-line capability layer for AI agents. It selects, installs, configures, checks, and routes tools that let command-line agents read and search websites and services including web pages, YouTube, RSS, GitHub, Twitter/X, Reddit, Bilibili, Facebook, Instagram, Xiaohongshu, LinkedIn, V2EX, and Xueqiu. It uses an ordered list of backends for each channel and probes them for actual availability, selecting the first usable option and reporting failures and repair guidance through `agent-reach doctor`. The repository lists integrations such as Jina Reader for web pages, yt-dlp for YouTube transcripts, `gh` for GitHub, feedparser for RSS, Exa through mcporter for web search, and alternative CLI or browser-session-based routes for sites requiring authentication. Agents invoke the upstream tools directly rather than through a data-wrapping layer. The CLI supports environment checks, explicit system installation, dry runs, updates, and uninstalling. Credentials and cookies are documented as being stored locally in `~/.agent-reach/config.yaml` with owner-only permissions; authenticated services may require user-provided cookies or browser sessions. The project is distributed under the MIT license.
oh-my-hermes (OMH) is a plugin and operating layer for Hermes Agent that adds model routing, coding workflows, specialist skills, project memory, and evidence-gated execution without replacing Hermes as the natural-language interface. It scores requests, selects workflows and model-and-effort categories, applies model-family-specific prompting, and can split accepted plans into parallel worktree-based lanes with typed results and verification gates. Its workflows cover interviewing, research, planning, coding delegation, quality assurance, performance work, and iterative plan-build-review loops; the terminal interface exposes phase-based todos, delegated-lane status, cost and token telemetry, and execution states that distinguish preparation, reported completion, and verified results. OMH also provides a file-backed, reviewer-gated memory store with provenance, review dates, conflict handling, and task-scoped recall packs, while leaving Hermes's native memory untouched. It is distributed as a command-line package and plugin with installation paths including a shell installer, Homebrew, Bun, npm, and Hermes skill tap.
Viserys is a self-contained pack of Markdown engineering workflows, reviewer personas, validation scripts, commands, and lifecycle hooks for AI coding agents. Its workflow routes development through DEFINE, PLAN, BUILD, VERIFY, REVIEW, and SHIP stages, with 28 skills covering requirements clarification, specification and task breakdown, implementation, test-driven development, debugging, browser testing, review, security, performance, migrations, CI/CD, documentation, observability, and release preparation. The pack also includes four reviewer personas, shared checklists, evaluation cases and fixtures, and validators for skills, commands, artifact paths, reference links, and versions. The repository includes an OpenCode agent definition that registers the skills and routes requests through the lifecycle workflow. The Markdown skills can also be supplied to skill-aware agents such as Claude Code, Cursor, and other agents, or read directly as standalone process documents.
3dviz-pro-max is a standalone agent skill for turning plain-language ideas into interactive Three.js or Blender scenes. It is designed for Claude Code and Codex, and provides a ten-step workflow covering intent, object reasoning, representation, construction routing, first-view creation, meaningful behavior, capture, inspection, refinement, and reporting. The skill includes searchable recipes and knowledge records, runnable Three.js templates, reusable scene kits, style and camera guidance, and a capture helper that drives a built scene in a browser and writes PNGs and console logs for inspection. Its catalog covers subjects including environments, anatomy, mathematics, physics, architecture, logistics, robotics, and abstract systems. The supplied Python helpers search the catalog and drive an existing scene; they do not themselves render or judge scenes, and screen capture requires an optional Playwright installation. The repository includes 37 runnable studies and a site at 3dviz.dev. Its own code, data, and authored media are released under the MIT license, while bundled third-party assets retain separate terms. Blender 4.2 LTS or newer is recommended but not required; without Blender, the repository states that scripts still run but kit output is limited to T2.
Infinite World is an open-source system for building persistent worlds with multimodal AI. It combines multimodal understanding, world simulation, and interaction in a continuous loop: prompts, images, or other inputs are interpreted into world state, rules, and possible actions; generated scenes respond to that state; and people, agents, and events can update the world. The runtime records scenes, choices, state changes, history, and branches so worlds can preserve context across scenes and be replayed. The project supports text, voice, and image interaction, local previews, saved-branch replay, and cached generated video. Its current local setup uses a Node.js and pnpm workspace, with FFmpeg and whisper.cpp for voice transcription and live output. Interactive live streaming is in development, while additional inputs such as chat, audio, mouse, and keyboard are planned. The repository describes the project as early development, with APIs and stored data subject to change, and distributes it under the Apache License 2.0.
geiger is a read-only command-line scanner from Atomburst that inventories AI agents, harnesses, MCP servers, plugins, skills, hooks, browser extensions, desktop AI applications, and related IDE configurations on a machine. It reads known configuration files and directories, identifies each finding's origin and evidence path, and labels capabilities such as code execution, secret storage, broad filesystem access, broad web access, and network access. Credential-shaped values are reported by key name, file, and secret type without exposing value contents; the scanner performs no telemetry and writes only an explicitly requested report file. It can emit terminal, HTML, or versioned JSON reports, scan specified project or home directories, enforce a strict exit status for findings that can execute code or hold secrets, and compare JSON snapshots to identify newly appeared, removed, or escalated findings. It runs through npx with Node.js 18 or newer, has no runtime dependencies or account requirement, and is licensed under MIT. The project notes that it reads configuration rather than runtime behavior, covers known locations, and does not determine whether a package is malicious.
Agent Launcher is a local desktop workspace for configuring and running existing coding-agent command-line interfaces. It detects installed CLI binaries, links or installs supported agents, applies account or provider profiles, and runs them in an embedded terminal or chat view without automatically reinstalling or updating already installed CLIs. It supports project-aware sessions by launching an agent in a selected project directory, synchronizes provider settings with CLI-native configuration files and environment variables, and performs a minimal model request to check credentials, endpoints, models, network access, and account status. The application reads local conversation histories in JSONL or SQLite formats, can resume or delete local sessions, displays installed MCP servers and Skills, and summarizes locally stored usage data. Agent Launcher distributes macOS, Windows, and Linux builds and is released under the MIT License. It is local-first rather than offline-only: launched agents send requests to the provider or relay selected by the active profile, while API keys are stored as plaintext in local configuration files and masked only in the interface.
LibreChat is an open-source, self-hosted AI chat platform that provides a ChatGPT-inspired interface for connecting to major AI providers and custom OpenAI-compatible endpoints. It supports model switching, multimodal file interactions, web search, speech-to-text and text-to-speech, image generation, conversation search, presets, branching, resumable streams, and a sandboxed Code Interpreter for languages including Python, JavaScript/TypeScript, Go, C/C++, Java, PHP, Rust, and Fortran. Its agent system supports no-code custom assistants, community-shared agents, MCP servers, tools, file search, code execution, reusable SKILL.md instruction bundles, subagents, and an Agent Management API. Agents can optionally use attached workspaces to inspect, search, edit, and run commands in bounded environments. The platform also provides generative UI artifacts, including React, HTML, and Mermaid content, plus custom actions and integrations with providers such as Anthropic, AWS Bedrock, OpenAI, Azure OpenAI, Google, Vertex AI, Ollama, Groq, Mistral, OpenRouter, and others. LibreChat includes multi-user authentication through OAuth2, LDAP, and email login; browser-based administration of users, groups, roles, and configuration; role-based permissions; observability through OpenTelemetry and Langfuse; Redis-backed synchronization for scaled deployments; and Docker Compose deployment options. It is built in public and intended for local or cloud self-hosting.
security-audit is a Cloudflare coding-agent skill that orchestrates multi-phase security audits of codebases. It uses reconnaissance to map architecture, trust boundaries, input surfaces, prior evidence, and deterministic coverage; assigns isolated agents to coverage-led hunting; sends each unique candidate to a fresh verifier; and records confirmed, needs_validation, and rejected results in machine-readable findings.json validated against a report schema. Independent agents verify final source claims, after which the skill derives target-neutral reports and coverage records. It can be installed with the Skills CLI and requires a tool-using coding agent with parallel sub-agent support, Node.js for its validators, and an OS-enforced sandbox for target-controlled execution. The repository is licensed under MIT.
Knowledge Work Plugins is an open-source repository of role-specific plugins for Claude Cowork and Claude Code. Each plugin bundles domain skills, MCP connector configuration, slash commands, and, where applicable, sub-agents for workflows such as productivity, sales, customer support, product management, marketing, legal, finance, data, enterprise search, and biomedical research. Plugins use file-based Markdown and JSON rather than application code, infrastructure, or build steps; users can install them through Cowork or the Claude Code plugin marketplace and customize their connectors, company context, and workflow instructions.
Cline is an open-source AI coding agent developed by Cline Bot Inc. It is available as a CLI assistant, VS Code extension, JetBrains plugin, native macOS and Windows desktop app, and SDK. Cline reads project structure, edits files across a codebase, runs terminal commands, observes build and test output, and can correct issues such as missing imports, type mismatches, syntax errors, and failed tests. Its Plan and Act modes separate codebase exploration and planning from execution; edits and commands can require human approval, with checkpoints for reviewing or reverting changes. The CLI supports interactive and headless operation for scripting and CI/CD, while the SDK exposes the shared agent engine for custom tools, plugins, lifecycle hooks, multi-agent teams, scheduled agents, connectors, and MCP servers. The project supports models from multiple providers and OpenAI-compatible endpoints, including local-model runtimes, and can connect agent sessions to messaging platforms such as Telegram, Slack, Discord, Google Chat, WhatsApp, and Linear. It is distributed under the Apache 2.0 license.
Nimble Web Search Agents are domain-specific web research agents from Nimble that perform self-learning search, crawling, structured extraction, dataset enrichment, and web monitoring. They use auditable Search Plans, a proprietary index, and memory that adapts to a use case after each run to retrieve targeted data from the live web, including JavaScript-rendered pages, without requiring an LLM to parse raw pages. Nimble exposes these capabilities through APIs for search, complex research workflows, dataset building, monitoring, extraction, and crawling.
BrowserSkill is an open-source browser-automation bridge for AI agents, consisting of the `bsk` CLI and daemon plus a Chrome or Microsoft Edge extension. An agent sends browser tasks through the CLI; the local daemon routes them over WebSocket to the extension, which performs actions in a separate visible Agent Window. Agents can use existing browser login state, borrow an already open tab only with explicit approval, and request human intervention for captchas, confirmations, or other user-only steps. It supports shell-capable agent harnesses, macOS, Linux, and Windows, and is distributed under the MIT license.
ToolReplay is a dependency-free Python CLI for auditing recorded AI-agent tool-call transcripts without executing the tools again. It strictly parses JSON Lines records, compares repeated calls and their canonicalized responses to detect non-determinism and redundant calls, checks tool names against a declared permission scope, and produces deterministic, timestamp-free reports suitable for CI checks and Git diffs. It can seal transcripts into a SHA-256 hash chain, verify the chain for edits or reordered records, replay a session against its recorded responses, and report permission overreach. It requires Python 3.11 or newer, has no third-party runtime dependencies or network access, and is distributed under the MIT license. Its scope checks are name-based, sealing is tamper detection rather than a signature, and replay only detects issues represented in the recorded transcript.
Design Studio AI is an open-source, agent-first design workspace for humans and AI agents. It supports web interfaces, slides, reports, wireframes, 3D scenes, and timeline videos; users start from a brief or template, inspect a preview, revise through chat or a manual editor, and work together on the same versioned document. The workspace provides structured flex/grid layouts, reusable design systems, 2D character motion, editable 3D scenes, BYOK text/image/speech/music/video generation, and exports including JSON, HTML, SVG, PNG, PDF, PowerPoint, WebM, GLB/glTF, React prototype ZIPs, and authorized Google Slides. It exposes REST, authenticated MCP, experimental WebMCP, and a CLI for agent access, and can run on Cloudflare or be self-hosted with Docker and persistent SQLite/files. The repository is distributed under the MIT license.
RSIAgent is an open-source, training-free multi-agent framework for recursive self-improvement in unfamiliar digital environments. It uses a Curriculum Agent, Actor Agent, and Verifier Agent while keeping model parameters fixed. Its learning loop has two stages: broad recursive self-exploration, in which agents perform diverse projects and consolidate verified experience, followed by deep recursive self-exploration, in which the Curriculum Agent selects focused practice around gaps and fragile successes. The Actor Agent interacts with software through executable Python or Bash programs and visual observations; the Verifier Agent independently checks task requirements and resulting environments. Procedures, scripts, successful experiences, and failure lessons are stored in persistent memory, which is frozen and reused for downstream task execution. The repository provides runtimes and benchmark integrations for OSWorld-V2 and Agents' Last Exam, along with setup, smoke-test, batch execution, reporting, recovery, and validation utilities. It requires Python 3.12 and a Linux host with Docker and /dev/kvm for the documented benchmark runs, and is licensed under Apache License 2.0.
gap-trap is an installable coding-agent skill that reads a repository and generates repository-specific contracts, checks, and quality gates for AI-written code. Contracts document the approved way to handle areas such as HTTP, logging, settings, and authentication; gates fail commits or CI runs when those rules are violated. Its Proven Red CI job runs each new test against the old code and rejects tests that pass without the change, while ratchets track known problems such as lint backlogs or oversized files and allow their counts to decrease but not grow. It also creates playbooks from project lessons and installs the related slop-mop skill for agent-generated documentation, commit messages, and pull-request bodies. Node and Python repositories receive gates inside their test suites; other listed languages use shell-based gates requiring git, grep, awk, and the project test command. It is installed through the skills CLI or by copying the skill into Claude Code or Codex skill directories, and provides setup and refine commands. The repository is MIT-licensed.
Panel is a research workspace that places an AI-agent chat alongside files, PDFs, and live Jupyter notebooks. The agent can read and write files, run notebooks against a real kernel, create scratch data, and edit the same notebook as the user; configurable panes can also display images, data, code, chats, and custom viewers or apps. It supports workspaces with separate chats and saved layouts, background commands, and literature reviews. The early build currently has full support for Claude Code, with optional OpenAI API support for chat and tools; it is distributed under the MIT license and runs locally with Node, pnpm, uv, and Python.