1,037 tools and products — trending open source, and what gets used in AI and other work.
Google's developer API for building applications with Gemini large language models. It provides access to models like Gemini 2.0 Flash and 2.5 Pro, and also supports Gemma open models. Developers use it for tasks such as custom model testing, retrieval, memory, and tool-calling systems.
An Anthropic feature in Claude that turns AI-generated content into high-fidelity work products that users can view, iteratively refine, and share. Artifacts can include documents, code, visualizations, and interactive interfaces displayed in a dedicated workspace alongside the conversation.
Baseten is an inference platform and company that provides tools to deploy, serve, and scale open-source and custom AI models in production. It offers managed infrastructure and tooling for model hosting, APIs, and inference at scale.
SGLang is an open-source serving framework and inference engine for large language models and multimodal models, developed by the SGLang project. It provides an alternative serving stack for deployed models, with runtime and kernel optimizations for inference workloads. Its documented capabilities include RadixAttention, a zero-overhead batch scheduler, cache-aware load balancing, structured-output decoding, speculative decoding, disaggregated prefill and decode, model parallelism, and support for GPU and TPU backends. The project also publishes integrations and optimizations for current open models and multimodal, image, and video-generation workloads.
Blotato is a social-media API, marketing, and multi-channel publishing platform that provides an MCP server for connecting marketing operating systems and AI agents or LLMs such as ChatGPT and Claude. It supports programmatic automation of social posts, including scheduling from Claude Code, as well as comment and direct-message management, content creation, and analytics.
Clay Studio is a self-hosted, browser-based team workspace for running and sharing local coding-agent sessions across projects. It connects to supported local runtimes such as Claude Code and Codex, allowing people and AI collaborators to share live sessions, hand off work, respond to permission requests, and retain project knowledge between conversations. The Clay Studio daemon runs on the user's machine and serves the workspace to a browser or mobile PWA. It stores sessions and knowledge on disk as portable JSONL and Markdown, and supports repositories, parallel sessions, Git worktrees, background workers, scheduled tasks, persistent AI collaborators called Mates, project instructions, file editing, diffs, terminals, MCP servers, and push notifications. The request path is browser to the Clay Studio daemon to the selected coding-agent runtime and its model provider; project traffic is not relayed through a Clay-hosted cloud.
Apollo, operating as Apollo.io, is a B2B sales prospecting and engagement platform with a contact and company database, lead-generation tools, outbound and inbound automation, outreach sequencing, CRM integrations, and deal-management capabilities. Its AI-assisted features support contact discovery and workflow automation for sales teams.
Circleback is an AI-powered meeting assistant that generates meeting notes, action items, automations, and searchable conversation records. It integrates with Zoom, Google Meet, Microsoft Teams, Slack huddles, and can capture in-person conversations.
IBM Granite is a family of foundation models developed by IBM for enterprise AI applications, including language and code models for text generation, summarization, question answering, language understanding, and software-development tasks. The models can be combined in multi-model systems that route simpler requests to smaller, lower-cost models and more complex requests to larger ones. They are distributed for local and hosted AI applications through model repositories and enterprise platforms, with several releases available under the Apache 2.0 license.
Rerun is an open-source SDK and data layer for recording, querying, visualizing, transforming, and streaming multimodal physical-robotics data. It ingests multi-rate data such as images, point clouds, transforms, time series, joint states, tensors, and video from robot logs, human-data rigs, simulation, and web video, using formats including MCAP, RRD, and LeRobot. Its built-in viewer synchronizes and renders the data in real time for episode scrubbing, sensor comparison, and live computer-vision pipelines; the same data can be queried with dataframes or SQL and streamed directly into training. Rerun is built in Rust on column-chunk storage designed for multi-rate physical data, with Python, Rust, and C++ SDKs and a separate Viewer binary for network streaming and RRD files. The repository describes the project as being in active development, with an evolving API and possible breaking changes.
Segment Anything Model (SAM) is an image-segmentation model from Meta AI Research's FAIR. It produces object masks from prompts such as points or bounding boxes, or can automatically generate masks for all objects in an image. The repository provides Python and command-line inference interfaces, pretrained checkpoints, example notebooks, and optional ONNX export and COCO-format mask processing. In the cited robotics workflow, SAM is used with Foundation Pose to extract object goal poses from a human video demonstration. The model was trained on 11 million images and 1.1 billion masks, according to the repository.
MLflow is an open-source AI engineering platform for agents, large language models, and machine-learning models, developed in the MLflow repository. It provides experiment tracking for models, parameters, metrics, and evaluation results, along with production observability, evaluation, prompt versioning and optimization, and model lifecycle tools. Its agent and LLM observability captures application traces using OpenTelemetry and supports different LLM providers and agent frameworks; its AI Gateway provides an OpenAI-compatible interface for routing requests, managing rate limits and credentials, handling fallbacks, controlling costs, and applying guardrails. MLflow can be run as a server and accessed through Python, TypeScript/JavaScript, Java, and other programming languages, with integrations including OpenTelemetry and MCP.
Cloudflare Computer is an open-source library that provides AI agents with a virtual filesystem backed by a Cloudflare Durable Object. The Durable Object stores the authoritative state in SQLite, while the workspace exposes a single execution interface through `workspace.runtime.exec(source, { backend })`. It supports three execution backends: Container projects the SQLite state into a sandbox container through a FUSE mount and synchronizes changes over capnweb RPC; Isolate shell runs `just-bash` in a Dynamic Worker; and Isolate JavaScript runs an ECMAScript module in a fresh Dynamic Worker with structured inputs and results, durable relative imports, configured libraries, workspace-backed `node:fs/promises`, and trusted `ws:git` and `ws:artifacts` modules. Workspaces can register multiple backends under stable identifiers, connect to them lazily, or be used solely as a filesystem without an execution backend. The repository includes examples for container, Worker shell, Worker JavaScript, egress policies, MCP, agent workspaces, tutorials, artifacts, and assets. The package is marked preview-only: its APIs are unstable and the repository says it is intended for experiments, exploration, and prototypes rather than production use.
AirLLM is an open-source Python library for running large language model inference on GPUs with limited memory. It reduces GPU memory use by loading and processing one model layer at a time instead of holding the entire model in GPU memory; for sparse mixture-of-experts models, it streams individual experts as needed. The project is designed to run large models without requiring quantization, distillation, or pruning, and provides a pip-installable package with support for multiple model families, CPU inference, macOS, and optional model compression and quantization features.
WARP (Weight-Aware Runtime and Paging), formerly WASTE, is an embeddable inference engine written in dependency-free C by SQLiteAI. It runs large mixture-of-experts models beyond available RAM by keeping the shared model trunk in memory, streaming selected experts from NVMe or other fast storage, and using remaining RAM as a bounded expert cache. Its container layout gives each expert an aligned read; reads overlap computation, while a lookahead router prefetches experts for the next layer without changing the decisions made by the actual router. Experts use 3-bit residual vector quantization, while more sensitive shared weights use 4- or 8-bit storage. The project is designed for local inference of frontier-scale models such as Kimi K3 and GLM-5.3-Flash on consumer hardware.
GenOffice is a free, open-source AI office suite from Genspark AI for macOS, Windows, and Linux. It consists of six Electron applications sharing an engine layer: Docs for .docx files, Sheets for .xlsx spreadsheets, Slides for .pptx presentations, PDF for viewing and editing PDFs, Markdown for plain Markdown files, and a shell application that hosts the editors. It opens and saves Microsoft Office formats, and includes light, dark, and system themes. Its AI panel supports document-aware, block-level edits with version snapshots and diffs; in spreadsheets, presentations, and PDFs it provides a tool-calling agent over document state. Built-in tools include web search, image search, image generation, and media analysis. Users can sign in through Genspark or supply their own keys for providers including Claude, OpenAI, Gemini, DeepSeek, Kimi, GLM, Qwen, Doubao, MiniMax, Grok, Mistral, and OpenRouter, as well as custom OpenAI-compatible endpoints and local servers. GenOffice Docs uses byte-preserving .docx editing in which only modified paragraphs are regenerated, with Word-compatible pagination, tracked changes, comments, styles, equations, and ink. GenOffice Sheets uses the open-source Univer core with in-house extensions and a Rust .xlsx import/export sidecar, plus charts, pivot tables, slicers, conditional formatting, and formula tracing. GenOffice Slides has an in-house .pptx parsing, rendering, and editing engine with masters, layouts, charts, cropping, ink, and text shaping. GenOffice PDF supports annotations, forms, outlines, stamps, signatures, page operations, printing, direct text reflow and image editing, and local conversion to editable Word, PowerPoint, or Excel files; scanned PDFs can use system OCR on macOS and Windows. GenOffice Markdown is a Tiptap block editor that saves headings, lists, tables, images, and code blocks back to plain Markdown. The suite is distributed under the Apache-2.0 license with installers for the three supported desktop platforms.
TencentDB Agent Memory is a self-hostable, team-level memory hub for AI agents. It extracts conversations and tasks into reusable Chat Memory and Skills, and converts documents and code into an LLM Wiki and CodeGraph. These assets can be reviewed, versioned, governed, shared, routed, and reused across agents, frameworks, and team members; existing documents, codebases, and agent sessions can also be imported to reduce cold-start work. The system provides a shared memory server and a proxy that preserves the agent protocol, so supported clients can use the same memory by pointing their base URL at the proxy without plugins, hooks, or an MCP server. The repository deploys memory-core, memory-hub, and the proxy together, with a local management panel and configuration for separate memory and proxy LLM parameters. Documented integrations include DeepSeek Harness, Claude Code, Codex, CodeBuddy, WorkBuddy, Hermes, and OpenClaw.
Desktop Commander MCP is an open-source Model Context Protocol server that lets AI assistants search, read, write, move, and edit files; run terminal commands and code; and manage interactive or long-running processes. It is built on the MCP Filesystem Server and adds search-and-replace editing, streamed command output, session management, process controls, configurable timeouts, and background execution. It also supports file previews and Markdown editing, data analysis for CSV, JSON, and Excel files, and native operations on Excel, PDF, and DOCX documents. The repository documents use with Claude Desktop and other MCP clients, as well as remote AI control through ChatGPT, Claude web, and other AI services. Its associated Desktop Commander App adds a graphical interface, live file-change previews, arbitrary model support, and custom MCP configuration.
OfficeCLI is an open-source command-line Office suite developed by iOfficeAI for AI agents to create, read, edit, render, and automate Word (.docx), Excel (.xlsx), and PowerPoint (.pptx) files without a Microsoft Office installation or external dependencies. It is distributed as a single binary and provides commands for creating documents, adding, modifying, moving, copying, and removing elements, reading text and structure as plain text or JSON, and saving changes. Its built-in HTML rendering engine converts DOCX, XLSX, and PPTX files to HTML or PNG for a render-and-review workflow. The `watch` command provides a live browser preview that refreshes after document changes, while the tool can also analyze formatting and structural issues, evaluate Excel formulas, and install an agent skill for supported AI coding agents.
Waku Agent is an open-source, local-first personal assistant and AI-agent harness built by seanchen.io. It runs on a laptop through a terminal, local browser dashboard, voice input, or Telegram gateway, and is organized around four components: a harness for tool execution, an approximately 95-line plain-Python agent loop, memory, and evaluation/LLM operations. Its memory uses semantic, episodic, and procedural stores in a local SQLite database, with a gate that decides whether to retrieve or save memory and a pass that determines what to retain. The project includes deterministic tests alongside LLM-as-judge evaluation with a release gate. The dashboard displays message flow through the harness, including gate decisions, tool calls, loop iterations, memory updates, traces, costs, and latency; it also includes views for workflows, tools and MCP connectors, memory, and the SQLite state database. It can be installed with pip and exposes the `waku` command. The dashboard runs a local web server at `127.0.0.1:7777`, and the Telegram gateway starts when `TELEGRAM_BOT_TOKEN` is configured. Model providers are selected through a provider setting and API key, with support stated for Anthropic, OpenAI, Gemini, DeepSeek, MiniMax, Kimi, GLM, OpenRouter, OpenCode Zen, and OpenCode Go.
Swiftlet is an open-source Swift and Metal runtime for running Qwen3-Next and Qwen3.5/3.6 mixture-of-experts models locally on Apple devices, including iPhones. It keeps the models' small dense core in memory and streams routed expert weights from storage on demand, allowing 35B and 80B models to run with comparatively low RAM. The project is available as a Swift package, command-line interface, OpenAI-compatible server, and iOS app; its repository describes the runtime as working end to end and notes that only about 3B parameters are active per token in the supported models.
MAGI-2 Preview is the inference implementation for SandAI’s unified audio-video generation model. The 114-billion-parameter architecture, built on MagiMoE, activates about 6 billion parameters per token and generates 10-second clips from text prompts (T2V) or a prompt plus a still image (I2V), with sound generated alongside the video and muxed into the output file. Generation runs in two stages: `magi2_preview` denoises the clip at low resolution, after which `magi2_refiner` increases it to 1080p. The current base release is not step-distilled, so denoising steps account for most of the generation time; a distilled release is described as forthcoming. The repository contains inference code, while the model weights are downloaded separately from the `sand-ai/MAGI-2-preview` Hugging Face repository and occupy hundreds of gigabytes in total. An optional prompt-enhancement stage sends the input to an OpenAI-compatible instruction-following LLM, which rewrites it as a structured JSON caption for the 10-second clip, renders that result as Markdown, and passes it to the generation pipeline. It can be disabled to use the raw prompt. The documented runtime requires eight NVIDIA Hopper GPUs, Python 3.12, a recent CUDA toolkit, and FFmpeg for audio-video muxing; Docker images and a source installation are provided.
Fin is an AI customer-agent product from Intercom that automates support, sales, and e-commerce interactions across chat and voice channels. It is described as able to handle complex customer issues, update accounts, and process payments and refunds while managing end-to-end customer journeys inside Intercom's messaging platform. It is offered as part of Intercom's conversational support suite.
Docling is an open-source document-processing toolkit from the Docling project for preparing files for generative-AI applications. It parses formats including PDF, DOCX, PPTX, XLSX, HTML, EPUB, images, plain text, audio, and video, and represents the results in a unified DoclingDocument structure. For PDFs, it analyzes page layout, reading order, headings, tables, code, formulas, images, and scanned content through OCR. Audio and video inputs can be processed with automatic speech-recognition models; video parsing produces an ASR transcript and representative keyframes. Parsed content can be exported as Markdown, HTML, WebVTT, DocLang, DocTags, or lossless JSON. Docling supports local and air-gapped execution, a command-line interface, Python usage, an API server named docling-serve, and an MCP server for connecting agents. The project also provides integrations for LangChain, LlamaIndex, CrewAI, and Haystack. It is installable with pip and runs on macOS, Linux, and Windows on x86_64 and arm64 systems.
Grok is an AI conversational assistant developed by xAI, the company founded by Elon Musk. It is offered as a chat model integrated with the X platform and designed for open-ended dialogue and real-time web-aware responses. The name has been used with variant version labels for model updates.
Open Science is an open-source, local-first, model-agnostic AI research workbench developed by AIPOCH for scientists and researchers. It runs on macOS, Windows, and Linux and supports computational and data-intensive research in fields including machine learning, statistics, life sciences, chemistry, materials science, physics, and environmental science. Users create projects and sessions, describe research goals in plain language, attach source files, select a model and approval mode, and inspect the agent’s tool activity. Scientific AI agents can read files, search the web, execute Python and R code, query scientific data sources, and produce reports, tables, figures, and other research artifacts. The workbench supports literature review, hypothesis development, code execution, data analysis, simulation, visualization, and related reproducible research workflows. Its first-run setup checks the environment, configures model-provider credentials, optionally prepares Python and R runtimes, and prepares an agent runtime such as Claude Code, OpenCode, or Codex; app-managed runtimes can be installed without Node.js, npm, or administrator privileges. Projects keep sessions, uploads, generated files, and preview state together. Conversations record agent answers and the commands, file reads, edits, searches, and connector calls behind them. Generated artifacts are stored as immutable, checksummed versions; their Provenance view can show producer code and execution history, referenced inputs, the observed environment inventory, the producing conversation branch, and version-scoped reviewer findings, marking unavailable evidence as unavailable rather than guessing.
Vibe-Trading is an open-source AI personal trading agent and quantitative-finance research workspace developed by HKUDS. It turns natural-language finance questions into runnable analysis and backtesting workflows by connecting an agent to market-data loaders, portfolio and backtesting engines, report generation, persistent memory, and multi-agent research teams across asset classes. Its quantitative and risk-analysis capabilities include strategy backtesting, Heston stochastic-volatility pricing, hierarchical risk parity, copulas, and market-microstructure estimators. The project provides API and MCP interfaces, broker and market-data connectors, portfolio-management and broker-authorized trading workflows, examples, documentation, a demo, and a Shadow Account for testing. It supports read-only account sources such as official IBKR MCP and opt-in Binance USD-M snapshots, while scheduling actions require explicit confirmation and live-order safeguards fail closed on contradictory or invalid data. Implemented as a Python project, it uses a connector onboarding system that stores connection secrets in the operating-system keyring and is distributed under an open-source license.
An open-source agent skill for redesigning GitHub README homepages around a repository's actual content. It reads the repository first, identifies the clearest value and supporting proof, and then derives a project-specific visual system rather than applying a shared template. In whole-README mode, it works across content, visual-system, and engineering layers: it removes repetition and internal jargon, moves proof toward the opening, derives typography, color, composition, and project-native motifs, and keeps assets GitHub-safe, accessible, searchable, and copyable. The skill separates visual and content layers by using SVG for editable heroes, section transitions, comparisons, diagrams, and identity while retaining maintainable Markdown for searchable body text. It supports hybrid SVG compositions with optional AI-generated elements, local previews, and an approval step before publishing. The repository documents examples used by eight public repositories, including project-specific heroes and real outputs for slide creation, UI reconstruction, icon generation, agent delegation, vehicle telemetry, interactive mapping, and a solo werewolf game.
Audar-ASR-V1 is a family of Arabic-first generative speech-recognition models developed by AudarAI. It treats transcription as audio-conditioned next-token prediction with a language-model decoder, rather than a CTC or transducer objective, using a Whisper-style 128-mel audio encoder and a Qwen3 decoder with a 30-second context. The models cover Modern Standard Arabic, major Arabic dialects, code-switched Arabic-English, English, and 30 languages overall. The Flash tier is intended for real-time, edge, on-device, or offline use, while Turbo targets lower error on difficult dialectal and long-form audio. Both tiers share an architecture and prompt interface and can be run through Transformers, GGUF with llama.cpp, or vLLM; the repository provides model pointers, benchmarks, and inference examples. The published models use the AudarAI Open v1.0 and AudarAI Community v1.0 licenses.
text-to-cad is a library of agent skills for creating, inspecting, sourcing, slicing, previewing, and handing off CAD, CAE, CAM, robotics, and hardware-design artifacts from local project files. Its skills generate and edit CAD models from plain-language or image requests, primarily producing STEP files with optional STL, 3MF, and GLB exports; preview local CAD and robot files in a browser; source off-the-shelf STEP parts; and create DXF drawings from Python sources or CAD geometry. Additional skills write URDF robot structures, SRDF/MoveIt2 planning and collision data, and SDF simulator models and worlds; check DXF and STEP files for SendCutSend; measure mesh printability by wall thickness, overhangs, support volume, and build orientation; slice supported meshes into validated, printer-profiled FDM G-code using slicer command-line tools; and dry-run, upload, and cautiously start local Bambu Lab print jobs. An experimental implicit-CAD skill creates browser-native models with GLSL signed-distance fields and CAD Viewer raymarch rendering. The library is installed with the Skills CLI, including `npx skills add earthtojake/text-to-cad`, or through provider-native plugins for Codex, Claude Code, and Grok Build.
AI Job Search is an open-source, local job-application framework built around Claude Code or compatible agent tools. Users create a career profile, search and deduplicate job postings, evaluate and rank matches against structured fit criteria, and run `/apply` to produce tailored CVs and cover letters in LaTeX. The application pipeline uses a drafter–reviewer process: an agent evaluates the posting and drafts the materials, a second reviewer critiques them, and the workflow revises the output; `/interview` supports interview preparation, and optional salary benchmarking is included. The included portal-search skills target Danish services such as Jobindex, Jobnet, and Akademikernes Jobbank, while the profiling, fit-evaluation, and application workflow is described as language- and country-agnostic and adaptable to other job boards. It requires Python 3.10 or later, Bun for the job-search CLI tools, and a LaTeX distribution with `lualatex` and `xelatex`; an optional `pdftotext` installation supports the ATS parseability check for compiled CVs. The repository states that the project is independent of and unaffiliated with Anthropic or Claude Code.
An open-source AI agent skill for Claude Code, Codex, and similar coding agents that produces cinematic product, marketing, launch, and demo videos with Remotion. It provides shot recipe cards, motion previews, a reusable video template, and workflows for storyboarding, real page capture, animation, 2.5D camera moves, sound design, beat-synced cuts, and visual checks. The repository includes native Remotion components driven by normalized progress values, a gallery for browsing motion previews, and an optional JianYing project export that separates shot plates, captions, sound effects, and background music into editable tracks. It can be installed through the skills CLI, by cloning the repository, or by linking it into a Claude Code or Codex skills directory.
eve is an AI agent framework for building durable, structured, multi-agent projects from declarative configuration covering instructions, skills, tools, channels, schedules, permissions, and evaluations. Its agents can be developed within a single folder and run through chat surfaces such as Slack or the eve development TUI. Vercel uses eve as the foundation for a marketing-team template with one lead and five specialists: product marketing, content marketing, social media, SEO, and email. The lead loads shared brand context and user preferences, writes self-contained briefs, routes work one level deep, and returns deliverables produced in Notion, Typefully, or Resend; specialists research and edit their own work against written rubrics rather than spawning further agents. The template uses MCP connections for Notion, Typefully, and Resend, Vercel Blob for shared state and files, Vercel AI Gateway for model access, and Vercel Sandbox for reference files and shell commands. Irreversible actions—including email sends, deletes, scheduled social publishing, and Notion page moves—pause for human approval, while the available Resend tools are restricted to keep account administration out of reach. The project is implemented in strict ESM TypeScript and includes commands for local development, validation, type checking, discovery diagnostics, and deployment.
Basedash is an AI-native business intelligence SaaS platform developed by Basedash. It connects to customer data sources and translates natural-language questions into charts, dashboards, and reports, while supporting reporting workflows for teams. Its governed semantic layer defines approved metrics and links them to answers, dashboards, and reports, providing a controlled basis for data queries and outputs without extensive BI setup.
Impeccable is an open-source design skill for AI coding agents, developed by pbakaus. It provides a shared design vocabulary with 23 commands, including init, shape, critique, audit, polish, harden, animate, typeset, layout, and live. The init command gathers product and design context into PRODUCT.md and DESIGN.md files, while other commands review, plan, modify, or refine a frontend project. Its CLI and browser extension run 61 deterministic detector rules for common AI-generated frontend design problems without an LLM or API key; it also supports LLM-based critique checks. The live command provides browser-based visual iteration, and the package can be installed from a project root with npx impeccable install.
Ratel is an in-process context-engineering platform for AI agents. It catalogs tools, skills, and persistent facts, then searches the catalog on each turn and progressively discloses only the capabilities relevant to the request instead of placing every tool schema and instruction in the context window. Its retrieval supports BM25, semantic, and hybrid ranking without requiring a vector database. The project includes a Rust core, TypeScript and Python SDKs, an MCP server, and a CLI. Tools can be registered directly or ingested from an MCP server; skills can associate instructions with tools, while registered facts are re-injected when they are no longer fresh in the conversation. The repository describes the system as reducing token usage and mitigating tool overload.
SimpleEnglish is an open-source agent skill that makes large language models write technical documentation in ASD-STE100 Simplified Technical English, a controlled language used in aerospace. Its rules enforce short sentences, active voice, explicit instructions, simple tenses, and consistent terminology when rewriting READMEs, error messages, incident reports, runbooks, and release notes. The skill is distributed as a dependency-free folder for tools that support the Agent Skills standard, including Claude Code, Cursor, VS Code Copilot, OpenAI Codex, Gemini CLI, Goose, and OpenCode. It can also be installed as a Claude Code plugin and output style, or used by adding its prompt to a system prompt, AGENTS.md, or .cursorrules file. The repository is licensed under MIT.
Prime Agent is an open-source command-line coding and research agent for general and long-running tasks, developed by Prime Intellect. It combines recursive language-model workflows through an RLM harness with a persistent Python REPL, treating prompts as variables and invoking tools and recursive subagents programmatically. A continual harness stores supplemental prompts, memories, skill descriptions, and reusable subagent specifications as durable session state that can be refined through small, evidence-backed updates without changing the immutable base system prompt. The agent supports importable Python skills, parallel or background child agents, messaging between running agents, automatic context compaction, persistent goals, heartbeats, schedules, detached sessions, autonomous operation, and retained subagents.
Superpowers is an open-source agentic skills framework and software development methodology for coding agents. It begins by eliciting and presenting a software specification for approval, then produces an implementation plan emphasizing red/green test-driven development, YAGNI, and DRY. After approval, it uses a subagent-driven-development process in which agents implement engineering tasks, inspect and review the results, and continue according to the plan. It is distributed as a plugin for supported coding-agent harnesses, including Claude Code, Codex, Cursor, and others.
Code Review Graph is an open-source, local-first code intelligence tool that provides a command-line interface and MCP server. It builds a persistent structural map of a codebase with Tree-sitter, representing relationships such as functions, classes, imports, calls, inheritance, and test coverage, while tracking changes incrementally. The graph supplies focused repository context to AI coding tools so they can read relevant code instead of repeatedly scanning large portions of a project. It also supports code-review workflows such as blast-radius analysis and risk-scored pull-request reviews. The Python package can be installed with pip or pipx; its install command detects supported AI coding platforms and configures MCP settings, hooks, skills, and graph-aware instructions.
OpenEdit is an open-source, agent-driven video-editing pipeline from VEED. It has no graphical interface or timeline; an agent-agnostic skill and repository guide drive it through coding agents such as Claude Code, Codex, and Gemini. The pipeline can edit, cut, and reframe footage; create burned-in subtitles, motion graphics, charts, and visual elements; assemble slides, websites, stills, or generated media into video; capture web pages; and connect generation services or MCP servers when needed. Source footage is optional, and VEED Fabric can generate footage only with the user's approval. For captioning, the agent transcribes the video, designs the captions, renders the result, opens a video previewer, and accepts follow-up requests such as changing subtitle color, position, or emphasis. Transcription can use VEED, local offline WhisperX, or a user-supplied service; all providers must produce per-word timings, which are stored in a common transcript format. Caption styles are authored in HTML and CSS. The project includes VEED's closed-source but free-to-use HTML renderer, which does not require a headless browser; Chrome can be used as an alternative rendering backend. V1 primarily targets captions, while motion graphics, charts, and brand-book-matched styling are less exercised. OpenEdit supports Apple Silicon Macs and Windows x64 PCs. Intel Macs are unsupported, Linux support is planned, and Windows requires Git, Node, and ffmpeg. It can be installed from the repository or with `npx skills add veedstudio/open-edit`.
Code-Graph-RAG is an open-source retrieval-augmented generation system for querying, understanding, and editing multilingual codebases. Its Tree-sitter-based parser extracts functions, classes, methods, modules, and relationships, then stores them in Memgraph under a unified graph schema. An interactive CLI converts natural-language questions into Cypher queries, retrieves source code by name or intent, and supports AI-assisted editing and optimization. It provides AST-based surgical patches with diff previews, structural search and rewriting through ast-grep, dead-code discovery by traversing call and reference edges, and runtime-call overlays from test traces or production eBPF profiles. The project also supports MCP integration and publishes the `cgr` command-line tool to PyPI.
ComfyUI is an open-source AI content-creation engine and diffusion-model GUI, API, and backend developed by Comfy-Org. Its visual node-graph interface lets users build and reuse no-code workflows for image, video, audio, 3D, and text generation, editing, and processing, including workflows using diffusion models, VAEs, text encoders, LoRAs, ControlNets, adapters, and upscalers. Workflows can be saved and loaded as JSON, exposed through App Mode, or integrated into applications through local API endpoints. The engine supports asynchronous queueing, partial graph re-execution, VRAM and RAM management, model offloading, quantized models, custom nodes, and fully offline core operation unless optional resources are requested. It can run locally on Windows, Linux, and macOS through desktop, portable, or manual installations, with support for NVIDIA, AMD, Intel, Apple Silicon, and Ascend hardware. Comfy-Org also provides an official paid cloud version.
LoopX is an open-source, provider-neutral state kernel and local-first control plane for long-horizon AI agents and peer agent teams. It runs on top of agent harnesses such as Codex, Claude Code, Cursor, dsh, or a custom harness rather than replacing them, keeping objectives, gates, tasks, evidence, quotas, schedules, recovery state, and handoffs durable across bounded turns, sessions, tools, and agents. Its control layer makes semantic decisions about what happens next and supports governance, recovery, human-agent collaboration, and reviewable work. The accompanying personal agent workspace stores goals, attention, conversations, tasks, files, schedules, and recovery state locally across restarts; its supported browser/PWA dashboard can continue work across registered agent sessions and use typed previews, explicit confirmation, and receipts for protected changes.
Pi Web is an open-source local browser interface for the Pi coding agent, developed in the agegr/pi-web repository. It uses Pi’s local configuration and session files, allowing users to browse, resume, rename, export, delete, and branch project-grouped conversations; run agent turns; and inspect running state, context usage, costs, and compaction details. New sessions create independent session files, while “Edit from here” creates a branch within the current session. Its project workspace supports file browsing and uploads, Git diff inspection, automatic previews for source files, Markdown, images, audio, PDFs, and DOCX files, and Git worktree switching. Configuration panels manage provider logins and API keys, models and model tests, plugin packages, and skills. The interface supports English, Simplified Chinese, and Traditional Chinese, follows the browser language initially, and includes a language switcher. Pi Web is distributed through the `@agegr/pi-web` npm package, requires Node.js 22.19.0 or newer, and listens on `127.0.0.1` by default. Remote binding and HTTP Basic Authentication are available, but the documentation warns that Basic Auth does not encrypt credentials in transit and that the service should not be exposed over plain HTTP without a trusted reverse proxy or VPN. It reads Pi agent data from `~/.pi/agent` by default, shares Pi’s model, settings, and credential storage, and limits file browsing to known project or session roots rather than providing general filesystem access.
Kimi Code CLI is an AI coding agent developed by Moonshot AI that runs in a terminal. It reads and edits code, runs shell commands, searches files, fetches web pages, and selects subsequent actions based on the feedback from those operations. It works with Moonshot AI's Kimi models and can be configured with other compatible model providers. The tool is distributed as a single binary and provides an interactive terminal UI. It accepts video input, supports conversational configuration of Model Context Protocol (MCP) servers, and can install skills, MCP servers, and data sources from its marketplace or GitHub repositories. Built-in coder, explore, and plan subagents can work in isolated contexts, while lifecycle hooks can run local commands at selected points in a session. Kimi Code CLI also supports the Agent Client Protocol (ACP), allowing compatible editors and IDEs such as Zed and JetBrains to drive sessions through the `kimi acp` command. The repository documents installation for macOS, Linux, and Windows and provides OAuth or Moonshot AI Open Platform API-key login options.
Persona is an open-source, cross-platform desktop character that gives voice conversations an animated visual identity. It listens locally to a selected application's playback process rather than the microphone, using PipeWire playback-stream capture on Linux, WASAPI process-loopback capture on Windows, and a Core Audio process tap on macOS. Imported VRM character models switch between idle and speaking animations, with support for custom VRMA actions, model libraries, click-through desktop-pet behavior, and a local MCP server. Persona does not save audio, produce speech, transcribe content, or send audio over the network. It is distributed as an AppImage or DEB on Linux, an NSIS installer on Windows, and DMG or ZIP packages on macOS.
numbat is an open-source endpoint security tool for monitoring AI-agent activity across supported desktop, CLI, IDE, and gateway agents. It combines local hooks and plugins, OTLP/HTTP logs, and on-disk session artifacts, normalizes live and stored activity into a common event model, and evaluates it locally with built-in, sequence-based, or custom YAML CEL rules. The tool can produce NDJSON events, findings, enforcement decisions, indicators, and scan summaries for alerting and forensic reconstruction, including read-only artifact scans, per-session timelines, and portable case bundles with SHA-256 manifests. Optional pre-action blocking is disabled by default and is limited to supported synchronous hooks and explicitly enforced rules. It is distributed as a single cgo-free binary for macOS, Linux, and Windows, with read-only inventory and scanning commands.
kimi-k3-mlx is an MLX port for Apple Silicon of Moonshot’s Kimi K3 native-multimodal mixture-of-experts model. It provides a stock `mlx-lm`-compatible text model definition, a 3D video-capable MoonViT vision tower with a patch-merger projector, and an `mlx-vlm`-shaped wrapper that joins the language and vision components. The model has a 1-million-token context and uses routed and shared experts, Kimi Delta Attention, gated multi-head latent attention, Attention Residuals, latent-dimension experts, and the SiTU-GLU activation; routed experts are distributed in MXFP4 while other weights use bf16. The multimodal processor expands each image placeholder into a feature block rather than performing a same-length one-token-per-image scatter. The repository includes a streaming converter because standard `mlx-lm` conversion would materialize the multi-terabyte bf16 model, along with REAP expert-pruning scripts and expert-overlap analysis using Chinese, English, code, or mixed calibration data. It also contains vision-parity tests against Moonshot’s reference implementation. The repository states that the full model cannot run on a single Mac: even its smallest deployment tier requires hundreds of gigabytes of memory and exceeds the capacity of the largest Apple Silicon machine.
QM is a multiplayer agent harness for startups that operates through Slack and a web interface. It gives each employee and room an isolated scope with its own memory, files, keychain view, permissions, scheduled jobs, web apps, and durable sandbox, while supporting collaboration in channels, group messages, and projects. Shared skills can be granted by scope, promoted by administrators, or imported from Git repositories; background work can run through crons, watches, and inbound webhooks, and internal apps can be published to selected users. A headless TypeScript core runs directly on Node with Fastify and handles identity, policy, scheduling, persistence, and the agent loop. Deployments can select harnesses and models including Pi, OpenCode, Codex, and Claude Code. PostgreSQL stores sessions, memory, queues, and other durable state, while a fixed tool surface includes an execute tool that runs commands in each scope's isolated, persistent sandbox. Slack is an optional in-process Bolt plugin; the web UI, admin panel, and public portal are optional HTTP API plugins built with Vite and Lit. Each deployment keeps organization-specific configuration, tools, skills, sandbox images, and infrastructure in a deployment directory validated and deployed by the qm CLI. Administrators can set organization-wide configuration, available harnesses and models, and a security posture. Strict mode pauses harness tool calls for human approval, Auto mode screens provenance-labelled external data and tool results with a classifier, and Dangerous mode disables content screening and pauses; a predeclared command policy with approvals and hard denials for operations such as recursive deletion or destructive SQL applies in every posture. Deployments run in the operator's own cloud account and are initialized from the @yc-software/qm package without requiring a source checkout; the project documents its threat model, operator assumptions, and known limitations in SECURITY.md.
Headroom is a local-first macOS menu-bar app for monitoring AI coding quotas and software delivery status, with optional iPhone, iPad, Apple Watch, and ESP32 desk-display clients. A Python host running on the Mac reads the authentication data and command-line tools already configured for services such as Claude, Codex, Cursor, GitHub, and Vercel, then serves a single local JSON feed; credentials remain on the machine and no Headroom cloud account is required. Its interfaces show provider quota meters, usage rings, burn rate and spend, failed Actions, deployments, monitors, and other attention items, and the iPhone client can approve, deny, or reply to Claude and Codex requests without starting work. The Mac host uses port 8737, requires macOS 14+ and Python 3.9+, and the optional desk display uses a Waveshare ESP32-S3-Touch-AMOLED-1.8 board.
Bindwidth is a browser-based calculator for sizing on-premises LLM inference infrastructure and estimating total cost of ownership. It evaluates three constraints for a resident text model—KV-cache memory, combined prefill and decode serving capacity, and the runtime session ceiling—and identifies which constraint binds for a workload. The tool accounts for interactive users and autonomous agents as different load types, then compares owned hardware, eligible sovereign rentals, and enterprise subscriptions with per-token APIs for agents. It distinguishes measured from estimated hardware and model data through confidence levels and profile-specific overrides, flags domain and model-capability limits, and exports a decision record as Markdown or the full scenario as JSON. The repository describes it as a directional sensitivity estimator rather than a procurement quote. It is a frontend-only application with no build step, server, database, account, or telemetry; calculations run in the browser and entered data is not transmitted.
Octop is an open-source, self-hosted AI assistant platform for households and small teams. It runs as a single process, storing conversations, workspaces, credentials, and shared state in a SQLite database under ~/.octop/, while serving a web dashboard, CLI, chat integrations, and cron automation. Users can switch among specialized personal agents through its multi-agent architecture. Built on the Harness stack, Octop combines an agent runtime for model routing, tools, skills, and conversation checkpoints with a gateway that normalizes messages from supported chat platforms into one processing pipeline. It also provides a built-in expert library, portable workspace memory, OAuth and MCP connectors, ACP integration for IDE and terminal workflows, browser automation, terminal-assisted command execution, and remote desktop access. The project uses FastAPI and uvicorn, React with TypeScript and Vite, SQLite through aiosqlite, and APScheduler. Its security features include JWT-based multi-user isolation, tool approval, shell-command guardrails, and PII redaction; supported storage backends include local disk, Docker containers, PostgreSQL, and COS/S3.
Phone Harness is an open-source harness that lets AI agents such as Claude Code, Codex, or other LLMs control a real iPhone or Android phone without a jailbreak, injected app, Xcode, WebDriverAgent, or on-device installation. It exposes helpers for opening apps, reading screen content, tapping, scrolling, typing, and verifying actions. On iPhone, it uses the macOS iPhone Mirroring window as the transport: screenshots are captured from the window, Apple's Vision framework performs OCR and returns tap-ready coordinates, and HID-level CGEvents provide taps, long presses, drags, flicks, scrolling, Unicode typing, and shortcuts. On Android, it uses ADB over USB or Wi-Fi, combining screenshots with the phone's accessibility tree for text and element bounding boxes. Actions can be verified by capturing the screen again, with the screenshot serving as the ground truth rather than relying on a DOM. Each invocation is self-contained rather than relying on a daemon.
Ante is a self-contained terminal coding agent and agent harness developed by Antigma Labs. It is distributed as a single Rust executable with no runtime dependencies and is designed to work with provider APIs, subscriptions, or local GGUF models. Its core embeds Grep and git operations in one process and uses a managed version of llama.cpp for local inference. In offline mode, it runs the coding-agent loop entirely on the user's machine without an API key, account, or internet connection; settings profiles can define the agent and its system prompt. The project also supports multiple providers, multi-agent orchestration, MCP skills, and memory. Ante is in beta preview, with breaking changes and incomplete functionality expected. The repository states that macOS and Linux are supported and recommends WSL for Windows.
Airship is a command-line visual editor for existing web applications, developed by Airship Labs. It runs as a reverse proxy in front of a development server and presents the application on an infinite canvas or in an inline editor, with live frames for different device sizes. Developers can select an element, describe a change, review the resulting diff, and have Claude Code, Codex, or OpenCode update the source file that rendered it; changes can also be undone. It works with applications serving HTML over HTTP, including Vite, Next, Remix, and Rails, without plugins, configuration changes, or additions to the project's dependencies or bundle. The CLI can connect to an existing server or start it with a command, and the README states that operation stays on the local machine without an account, telemetry, or separate hosted service.
An open-source collection of Markdown output styles for Claude Code, Codex, and other coding agents. It changes how an agent communicates rather than how it codes: responses are answer-first, written in plain English, and formatted for skimming. The collection includes Attention-kind, which spaces points apart, uses arrows and bold emphasis, and expands only when useful; Spartan, a terse style with no warmth; and Rundown, a brief-briefing style. Each style is a single Markdown file that can be added and switched on independently.
Mkdirs is an open-source Next.js template for building AI-powered directory websites. It provides listings with categories, collections, tags, and search; user authentication and dashboards; AI-assisted submission and content generation; and free, paid, or sponsored submissions through Stripe. Sanity supplies the content management system and blog, while Resend handles transactional email and newsletters. The template also includes SEO, analytics, themes, responsive layouts, and deployment options for Vercel and Docker. It is developed by OpenFox and licensed under Apache License 2.0, with the Mkdirs name, logo, and trademarks excluded from permission to identify derived products.
OpenMausBot is an open-source, local-first chat application for managing a team of AI bots as separate contacts. Each bot can have its own personality, model, conversation thread, computer, and connected applications, using agents launched through locally installed Claude, Codex, or Grok command-line tools and existing user logins. A local harness server on 127.0.0.1 manages agent processes, transcripts, keys, and events in the user's local OpenMausBot data directory. Bots can work through a cloud Linux desktop, an isolated local virtual machine, or—where supported—the host computer. Shell commands, file edits, and questions are routed through a permission broker and shown as approval or response cards. The application includes model selection, bot and conversation management, browser-accessible desktop takeover, and connected applications through Composio, including services such as Gmail, Slack, GitHub, Notion, and Linear. The repository provides desktop distributions for macOS, Windows, and Ubuntu, with Ubuntu Wayland host control disabled according to the README while a stated issue is resolved.
Axolotl is a free, open-source framework for post-training and fine-tuning large language models. It uses YAML configuration to define workflows for preprocessing, training, evaluation, quantization, inference, preference tuning, reinforcement learning, and reward modeling. The project is hosted by axolotl-ai-cloud and documents support for distributed training, mixture-of-experts models, multimodal models, and multiple fine-tuning methods.