703 tools and products — trending open source, and what gets used in AI and other work.
HQBase is an open-source shared email workspace for teams that runs in the customer's own Cloudflare account, keeping the application, mail, and Cloudflare credentials in infrastructure the customer controls. It provides shared mailboxes, team access controls, multi-domain configuration, collaborative drafts, and audit history, along with installation, update, backup, and recovery operations. HQBase also includes an OAuth-protected remote MCP server for controlled access by compatible agents. The project is distributed under the GNU Affero General Public License v3.0 only.
Cumora is a cross-platform team-chat platform in which AI agents participate alongside human teammates. It provides shared rosters, direct messages, group conversations, a Kanban board, and a calendar; agents can maintain personas and memory, claim work, coordinate, and send and receive email. Agents can run in Cumora's cloud in per-agent Kubernetes pods, using a multi-hop tool-calling loop on the OpenAI Responses API, or through its BYOA mode, where a local daemon connects the platform to agent CLIs such as Claude Code, Codex, Grok Build, Cursor Agent, OpenCode, or pi. The frontend uses React, Vite, TypeScript, and Tailwind across web, desktop, and mobile shells; the backend is a stateless Node service using Express and WebSockets, with PostgreSQL as the source of truth and Redis for pub/sub and presence. Agent coordination uses a freshness gate that holds stale replies for reconsideration, atomic claims on work units, and a triage gate intended to reduce unnecessary large-model calls.
herdr is a terminal multiplexer and persistent runtime for coding agents, developed as a Rust binary. It runs a background server that keeps agent terminals and sessions available across terminal disconnections, SSH reconnections, network loss, lid closure, and machine restarts; users can reattach from another terminal and use tmux-style keyboard controls or mouse interactions to split, move, and manage panes. Each pane is marked working, blocked, or idle, and agents can control herdr through its CLI and socket API to spawn panes, prompt other agents, and wait for an agent that is blocked. It hosts existing tools such as Claude Code, Codex, Cursor, OpenCode, and Grok without wrapping or replacing them, and supports plugins for extending panes and workflows. The project provides installation scripts and package-manager installation, documents remote use and session state, and is licensed under the Apache License 2.0.
OpenViking is an open-source context database for AI agents developed by Volcengine. It unifies agent memories, resources for knowledge retrieval, and skills in a virtual filesystem accessed through the viking:// protocol, allowing agents to browse context with filesystem-style operations such as ls, tree, and find instead of querying an opaque vector store. Content is processed into three loading tiers: L0 abstracts for relevance checks, L1 overviews for planning, and L2 full details loaded on demand. Recursive retrieval first locates a relevant directory through vector search and then drills down through its hierarchy, preserving the surrounding context and recording the browsing trajectory for inspection and debugging. After a session is committed, OpenViking asynchronously extracts user preferences and agent experience into long-term memory. The project also provides an OpenViking Studio browser playground, documentation, and a live demo.
Pi is an open-source AI agent harness and toolkit from earendil-works for building and running coding agents. Its packages provide a unified multi-provider LLM API, an agent runtime with tool calling and state management, an interactive coding-agent CLI, a terminal UI library with differential rendering, and vendor-neutral telemetry contracts and adapters. Slack and chat automation are provided through a separate package project. Pi does not provide built-in restrictions for filesystem, process, network, or credential access; it runs with the permissions of its launching user and process. The project documents containerization and sandboxing approaches, including a local Linux micro-VM, Docker, and policy-controlled sandboxing, for stronger isolation.
fx is an open-source coding-agent harness and command-line interface written in Zig by Vercel Labs. It is designed as a compact Unix-like alternative to a terminal IDE, with interactive and one-shot requests for inspecting and modifying repository code, running shell commands, and saving or resuming sessions. The agent supports skills, MCP tools, plugins, subagents, permission rules, and headless requests, and can be embedded natively or through WebAssembly. It is model-agnostic, supports local and cloud inference, is distributed under the Apache-2.0 license, and is marked experimental by its repository.
Apache Maka (Incubating) is a local-first agent workspace developed under the Apache Software Foundation. It inspects projects and runs tools through a shared Runtime Host within a sandbox boundary; tools that leave the sandbox require approval. Model messages, tool calls, tool results, and how a turn ended are stored as recoverable execution facts on the user's machine, while older tool output can be omitted from later prompts without deleting the saved history. Model connections can use cloud APIs, local models, or compatible gateways, with sessions, settings, and run records kept local by default. Maka provides an Electron and React desktop application with streaming sessions, tool timelines, branching, search, recovery, artifacts, and model and sandbox settings. Its TUI/CLI supports work in a project directory and non-interactive turns, while its evaluation surface runs declarative multi-arm experiments across Maka and external subjects. Built-in tools include Read, Write, Edit, Bash, Glob, and Grep; Computer Use and catalog skills are optional. The project is under active development and Apache incubation; the README describes the macOS Apple Silicon desktop build as an early public release and notes that data formats, CLI commands, and experimental capabilities may change.
OpenBot is an open-source AI agent platform from CopilotKit that gives each agent its own computer with a browser, logins, files, tools, and a dedicated interaction channel. It runs in the user's infrastructure and accepts agents that speak the AG-UI protocol, including agents built with LangGraph, Mastra, CrewAI, Pydantic AI, Google ADK, or custom code. A single gateway mediates actions involving the computer, files, MCP servers, and UI components: it decides whether an action is permitted before execution and records it afterward. Agents can operate their own browser screen, hand control to a human for restricted actions, and return component-based responses. Docker Compose runs the application components and PostgreSQL, while the administrator supplies the model credentials; the repository describes the project as alpha software under active development.
GitHub CLI is GitHub's official command-line tool, developed for GitHub.com, GitHub Enterprise Cloud, and GitHub Enterprise Server. The `gh` command brings pull requests, issues, and other GitHub concepts into the terminal alongside Git and source code, including workflows for authenticating an account, creating repositories and branches, pushing changes, and opening draft pull requests. It supports macOS, Windows, and Linux, and an agent skill is available for driving `gh` from coding agents.
Agent Deck is an open-source terminal dashboard for running and monitoring multiple AI coding agents in parallel. It displays each session’s status, active tool, working directory, and last prompt in real time, with Vim-style keyboard navigation for creating, focusing, closing, and renaming panes. It supports Claude Code and OpenCode through automatically installed hooks and runs as the single-binary `dot-agent-deck`, with native embedded terminal panes and no external terminal multiplexer required. The dashboard runs inside terminals including Ghostty, iTerm2, Alacritty, Kitty, and WezTerm, while preserving the agent clients’ existing shortcuts, skills, and configurations. Per-project TOML configuration can pair an agent with side panes for test runs, log tails, or `kubectl` watches. macOS and Linux installation is available through Homebrew, Nix, prebuilt binaries, or source builds; Windows use is currently through WSL. The project is distributed under the MIT license.
Macroscope is an AI-powered code review tool that analyzes pull-request changes and provides feedback on them, including changes generated by coding agents.
Grok Bot is an AI agent application from xAI that acts as an AI teammate on a persistent cloud computer. Its agents can perform research, monitoring, file-based workflows, automations, administrative tasks, and coding fixes through a computer interface, with documented system behavior and bot security boundaries. Subscription plans, usage limits, and billing are managed through Cursor.
Hyperagent is an AI agent platform for assigning work to a fleet of agents that research, build, deliver and maintain outputs using an organization's data, tools and working style. Its agents operate through real browsers and shells, can search, code, make decisions and generate artifacts such as live websites, videos, presentations, documents and dashboards. Agents run in individual computing environments, expose their execution as it happens, and can learn new memories and skills over time. Work is delivered through Slack, email, Telegram, webhooks or scheduled runs, and agents are deployed and managed from a central command center.
Praxist is an autonomous research system by Sapient Intelligence for computer-executable projects with measurable objectives. It turns an existing runnable project into a persistent research run: parallel research agents develop competing implementations or hypotheses, a task-defined evaluator converts their results into structured evidence, and a planning panel uses that evidence to set the agenda for later generations. Successful candidates and supporting evidence can move through incubator, frontier, and Gems retention lanes, while multi-metric evaluation, optional Quality-Diversity and Deep Innovation Gate allocation, resource scheduling, replay, and monitoring support longer runs. The task project remains responsible for its code, evaluator, metrics, baselines, prompts, roles, data, and domain constraints. Praxist can be installed as a Python package and operated through Codex, Claude Code, or its direct CLI; it requires CPython 3.11+ and a runnable project with measurable evaluation. The source is publicly available under the Fair Source License Agreement 1.0, with commercial terms described in the repository's license summary.
Munder Difflin is a free, open-source desktop multi-agent harness that runs multiple terminal-based AI coding agents as coordinated local agents. It wraps agents including Claude Code, OpenAI Codex, Gemini CLI, Qwen, OpenCode, GitHub Copilot CLI, and custom commands, working with users' existing subscriptions and hourly usage limits. Each agent runs as a real node-pty process and is rendered through xterm.js, while an Electron and React interface displays sessions as avatars on a Pixi.js office floor. A GOD orchestrator called Michael assigns and routes work, adjudicates agent messages, and escalates spending, destructive operations, scope changes, and other configured approvals. Agents coordinate through a local Git-backed hive of plain-file memory, atomic mailboxes, a shared blackboard, and an append-only event log; a markdown-first semantic memory layer supports recall across sessions. Optional Git worktrees isolate parallel agents.
OpenMAIC (Open Multi-Agent Interactive Classroom) is an open-source AI learning platform from THU-MAIC that turns topics or uploaded documents into interactive classrooms. It generates slides, quizzes, HTML simulations, and project-based learning activities, with AI teachers and classmates that can speak, draw on a whiteboard, conduct discussions, and respond to learners in real time. Its classic generation pipeline has two stages: an AI-generated lesson outline followed by scene generation for each outline item. The platform also provides a database-backed agent workbench that plans, builds, and revises courses through validated tools, with resumable sessions, follow-up steering, uploaded or web-retrieved materials, reusable skills, and support for importing PPTX files. Multi-agent orchestration uses a LangGraph director graph, while the playback and action engines handle classroom state and actions such as speech, whiteboard drawing, spotlights, and laser effects. OpenMAIC supports browser-only storage by default and can use PostgreSQL or S3-backed storage through its swappable storage packages. It accepts document, image, audio, and video materials through configured extraction providers, supports multiple LLM, media, speech, search, and local-provider configurations, and exports editable PPTX slides, interactive HTML, or classroom ZIP files. The repository is licensed under the MIT License, with separate terms for bundled components including an LGPL-licensed MathML-to-Office-Math package.
ECC is an MIT-licensed open-source agent-harness performance system maintained by affaan-m. It packages reusable skills, specialized agents, project rules, commands, hooks, memory, and security tooling for Claude Code as its primary target, with supported or limited adapters for Codex, Cursor, OpenCode, Gemini, Zed, GitHub Copilot, and other coding harnesses. Its core workflow turns plan, test, implement, review, verify, remember, and improve into reusable agent workflows. Skills are loaded for tasks such as test-driven development, research, security review, end-to-end testing, documentation, and refactoring; agents isolate planning, implementation, and review; rules provide always-loaded project or language standards; and hooks run event-triggered checks and session automation outside the model context. The optional Memory Vault stores inspectable Markdown handoffs and session context, while AgentShield scans agent files, hooks, MCP configurations, permissions, prompts, and secrets for security risks. ECC can be installed from its repository or through the ecc@ecc Claude Code plugin and ecc-universal package. The repository also provides selective installers, native or project-local integrations for several harnesses, a desktop dashboard, and the optional ecc-agentshield security-auditing package. Feature parity varies by harness: GitHub Copilot receives instructions and reusable prompts but not ECC hooks or agent delegation, while Codex has a native marketplace plugin with a narrower hook model.
Comp AI CRM is an open-source, self-hostable customer relationship management system for AI agents, developed by Comp AI and hosted by Try Comp AI. Its agent runs as an independent durable deployment on its own schedule and database-backed work queue rather than waiting for browser requests: it selects records to investigate, researches contacts and companies, spends a research budget, schedules follow-ups and rechecks, records observed evidence, and sends weak or ambiguous matches to a human instead of writing them directly to records. It supports durable sessions, contact and company records, an Agent tab for viewing work and answering questions, authored tools for reading CRM history, searching records, identifying contacts, researching people, enriching companies, recording facts and scheduling rechecks, and versioned Markdown skills. The agent uses file-based tools and Markdown skills on Vercel's eve durable-agent framework, while its queue leases due tasks with PostgreSQL row locking so concurrent workers claim disjoint work and expired leases release tasks from failed runs. A restricted shell sandbox provides bash, grep, glob and a workspace without network egress, database credentials or direct database access. Optional sources include mailbox history, company brand data, LinkedIn and Perplexity web research.
IP as Logo Skill is a compact Agent Skill in the open Agent Skills format that guides compatible AI agents in turning a product brief into highly simplified, rounded mascot-character concepts. It gathers product context, proposes three design directions, and after approval generates six independent full-resolution square candidates using one dominant silhouette of roughly four to seven large shapes, two mascot or IP colors, a named solid background color, thick rounded forms, and lower-left or lower-right corner emergence. Its default batch uses two variants per direction with a three-left, three-right composition split rather than a contact sheet. Familiar animals are the default subjects, while objects, machines, fantasy artifacts, and other unusual subjects require a clear product-related reason. The skill provides instructions and prompts rather than an image generator, can be installed with the Agent Skills CLI, and is intended for agents such as Codex, Coze, Doubao, YouMind, Manus, Gemini Apps, and Replit Agent when paired with a supported image model.
awesome-gpt-image-2 is a curated collection of structured GPT Image 2 prompts, reverse-engineered image-generation cases, reusable templates, and agent tooling. It organizes examples into categories such as interfaces, infographics, posters, products, branding, photography, illustration, characters, storytelling, and documents, and uses an atomic schema that separates subjects, lighting, materials, layout, and visual details into composable prompt components. The repository also includes the GPT-Image-2 Style Library agent skill, whose generated reference data is shared with the visual gallery. The skill helps agents select styles, templates, categories, and scene tags, and can be installed for tools including Claude Code, Codex, and Cursor through the skills CLI, npm, GitHub Packages, or the Claude Code plugin marketplace. A companion website provides searchable gallery browsing, prompt copying, filtering, and signed-in image generation. The repository is released under the MIT License and states that its third-party prompt cases and images are organized for learning, research, and automated testing, with ownership and commercial-use rights remaining subject to their original sources.
Omarchy is an opinionated, preconfigured, terminal-oriented and agentic Arch Linux-based desktop distribution and ready-to-use Linux setup by DHH, developed by Omacom. It provides a tiling desktop workflow with navigation, hotkeys, themes, clipboard history, reminders, text extraction and dictation, screenshots and recording, an Omarchy command-line interface, terminal and Neovim tooling, AI and development tools, shell utilities, browsers, web apps, gaming support, networking, dotfiles, updates, and system snapshots. It includes applications such as Chromium, Obsidian, LibreOffice, Kdenlive, and OBS Studio. Its extensive manual covers configuration, hardware, troubleshooting, and installation, including dual-boot and unattended installation options, and is maintained in the repository's `manual/` directory and mirrored to learn.omacom.io. Omarchy is released under the MIT License.
Composio is a developer platform for orchestrating just-in-time tool calls, secure delegated authentication, sandboxed execution environments, and parallel execution across integrations with 1,000+ applications. It is designed to let applications and AI agents invoke external tools safely and at scale.
LangFlow is an open-source, low-code visual builder for creating and running language-model pipelines, agent workflows, and retrieval-augmented generation (RAG) applications. The project is developed and maintained by the open-source community led by the GitHub user logaretm and integrates with language-model providers and tools such as LangChain.
Blotato is a social-media API, marketing, and multi-channel publishing platform that provides an MCP server for connecting marketing operating systems and AI agents or LLMs such as ChatGPT and Claude. It supports programmatic automation of social posts, including scheduling from Claude Code, as well as comment and direct-message management, content creation, and analytics.
Clay Studio is a self-hosted, browser-based team workspace for running and sharing local coding-agent sessions across projects. It connects to supported local runtimes such as Claude Code and Codex, allowing people and AI collaborators to share live sessions, hand off work, respond to permission requests, and retain project knowledge between conversations. The Clay Studio daemon runs on the user's machine and serves the workspace to a browser or mobile PWA. It stores sessions and knowledge on disk as portable JSONL and Markdown, and supports repositories, parallel sessions, Git worktrees, background workers, scheduled tasks, persistent AI collaborators called Mates, project instructions, file editing, diffs, terminals, MCP servers, and push notifications. The request path is browser to the Clay Studio daemon to the selected coding-agent runtime and its model provider; project traffic is not relayed through a Clay-hosted cloud.
LoopNet is an online marketplace for commercial real estate listings, allowing brokers, agents and property owners to list and search for properties for sale or lease. The service is owned by CoStar Group and focuses on U.S. and international commercial property markets.
Salesforce is an American cloud-based software company that provides customer relationship management (CRM) services and a suite of enterprise applications for sales, service, marketing, commerce, analytics and integration. Its products and brands include Sales Cloud, Service Cloud, Marketing Cloud, Tableau and MuleSoft, and it offers AI capabilities such as Einstein and agentic AI features. The company operates a multi-tenant cloud platform and is headquartered in San Francisco.
MLflow is an open-source AI engineering platform for agents, large language models, and machine-learning models, developed in the MLflow repository. It provides experiment tracking for models, parameters, metrics, and evaluation results, along with production observability, evaluation, prompt versioning and optimization, and model lifecycle tools. Its agent and LLM observability captures application traces using OpenTelemetry and supports different LLM providers and agent frameworks; its AI Gateway provides an OpenAI-compatible interface for routing requests, managing rate limits and credentials, handling fallbacks, controlling costs, and applying guardrails. MLflow can be run as a server and accessed through Python, TypeScript/JavaScript, Java, and other programming languages, with integrations including OpenTelemetry and MCP.
Cloudflare Computer is an open-source library that provides AI agents with a virtual filesystem backed by a Cloudflare Durable Object. The Durable Object stores the authoritative state in SQLite, while the workspace exposes a single execution interface through `workspace.runtime.exec(source, { backend })`. It supports three execution backends: Container projects the SQLite state into a sandbox container through a FUSE mount and synchronizes changes over capnweb RPC; Isolate shell runs `just-bash` in a Dynamic Worker; and Isolate JavaScript runs an ECMAScript module in a fresh Dynamic Worker with structured inputs and results, durable relative imports, configured libraries, workspace-backed `node:fs/promises`, and trusted `ws:git` and `ws:artifacts` modules. Workspaces can register multiple backends under stable identifiers, connect to them lazily, or be used solely as a filesystem without an execution backend. The repository includes examples for container, Worker shell, Worker JavaScript, egress policies, MCP, agent workspaces, tutorials, artifacts, and assets. The package is marked preview-only: its APIs are unstable and the repository says it is intended for experiments, exploration, and prototypes rather than production use.
GenOffice is a free, open-source AI office suite from Genspark AI for macOS, Windows, and Linux. It consists of six Electron applications sharing an engine layer: Docs for .docx files, Sheets for .xlsx spreadsheets, Slides for .pptx presentations, PDF for viewing and editing PDFs, Markdown for plain Markdown files, and a shell application that hosts the editors. It opens and saves Microsoft Office formats, and includes light, dark, and system themes. Its AI panel supports document-aware, block-level edits with version snapshots and diffs; in spreadsheets, presentations, and PDFs it provides a tool-calling agent over document state. Built-in tools include web search, image search, image generation, and media analysis. Users can sign in through Genspark or supply their own keys for providers including Claude, OpenAI, Gemini, DeepSeek, Kimi, GLM, Qwen, Doubao, MiniMax, Grok, Mistral, and OpenRouter, as well as custom OpenAI-compatible endpoints and local servers. GenOffice Docs uses byte-preserving .docx editing in which only modified paragraphs are regenerated, with Word-compatible pagination, tracked changes, comments, styles, equations, and ink. GenOffice Sheets uses the open-source Univer core with in-house extensions and a Rust .xlsx import/export sidecar, plus charts, pivot tables, slicers, conditional formatting, and formula tracing. GenOffice Slides has an in-house .pptx parsing, rendering, and editing engine with masters, layouts, charts, cropping, ink, and text shaping. GenOffice PDF supports annotations, forms, outlines, stamps, signatures, page operations, printing, direct text reflow and image editing, and local conversion to editable Word, PowerPoint, or Excel files; scanned PDFs can use system OCR on macOS and Windows. GenOffice Markdown is a Tiptap block editor that saves headings, lists, tables, images, and code blocks back to plain Markdown. The suite is distributed under the Apache-2.0 license with installers for the three supported desktop platforms.
TencentDB Agent Memory is a self-hostable, team-level memory hub for AI agents. It extracts conversations and tasks into reusable Chat Memory and Skills, and converts documents and code into an LLM Wiki and CodeGraph. These assets can be reviewed, versioned, governed, shared, routed, and reused across agents, frameworks, and team members; existing documents, codebases, and agent sessions can also be imported to reduce cold-start work. The system provides a shared memory server and a proxy that preserves the agent protocol, so supported clients can use the same memory by pointing their base URL at the proxy without plugins, hooks, or an MCP server. The repository deploys memory-core, memory-hub, and the proxy together, with a local management panel and configuration for separate memory and proxy LLM parameters. Documented integrations include DeepSeek Harness, Claude Code, Codex, CodeBuddy, WorkBuddy, Hermes, and OpenClaw.
OfficeCLI is an open-source command-line Office suite developed by iOfficeAI for AI agents to create, read, edit, render, and automate Word (.docx), Excel (.xlsx), and PowerPoint (.pptx) files without a Microsoft Office installation or external dependencies. It is distributed as a single binary and provides commands for creating documents, adding, modifying, moving, copying, and removing elements, reading text and structure as plain text or JSON, and saving changes. Its built-in HTML rendering engine converts DOCX, XLSX, and PPTX files to HTML or PNG for a render-and-review workflow. The `watch` command provides a live browser preview that refreshes after document changes, while the tool can also analyze formatting and structural issues, evaluate Excel formulas, and install an agent skill for supported AI coding agents.
Waku Agent is an open-source, local-first personal assistant and AI-agent harness built by seanchen.io. It runs on a laptop through a terminal, local browser dashboard, voice input, or Telegram gateway, and is organized around four components: a harness for tool execution, an approximately 95-line plain-Python agent loop, memory, and evaluation/LLM operations. Its memory uses semantic, episodic, and procedural stores in a local SQLite database, with a gate that decides whether to retrieve or save memory and a pass that determines what to retain. The project includes deterministic tests alongside LLM-as-judge evaluation with a release gate. The dashboard displays message flow through the harness, including gate decisions, tool calls, loop iterations, memory updates, traces, costs, and latency; it also includes views for workflows, tools and MCP connectors, memory, and the SQLite state database. It can be installed with pip and exposes the `waku` command. The dashboard runs a local web server at `127.0.0.1:7777`, and the Telegram gateway starts when `TELEGRAM_BOT_TOKEN` is configured. Model providers are selected through a provider setting and API key, with support stated for Anthropic, OpenAI, Gemini, DeepSeek, MiniMax, Kimi, GLM, OpenRouter, OpenCode Zen, and OpenCode Go.
Can I Vibecode It? is an open-source catalog that evaluates whether AI coding agents such as Claude Code, Codex, or Cursor can build a usable personal replacement for a paid SaaS application. Each app entry provides a YES, KINDA, or NOT REALLY verdict, the exact prompt for attempting the replacement, and the tradeoffs of leaving the original service, including lost network effects, data, or infrastructure. App entries are contributed as one JSON file per app, with the schema and verdict criteria documented in the repository. The site uses Astro server output with a Node adapter to render fully server-rendered HTML, SQLite via better-sqlite3 for vote counters and waitlist data, and vanilla JavaScript and CSS for interactions. Satori and resvg generate Open Graph images at build time. It runs locally with npm install, npm run dev, npm run build, and npm start, requires no environment variables for local development, and supports optional analytics, sign-in, payments, email, and media-storage configuration. It can be deployed on a VPS behind a reverse proxy, with DATA_DIR used to keep user data outside the repository. The project is released under the MIT License, and its prompts are intended to remain free.
Fin is an AI customer-agent product from Intercom that automates support, sales, and e-commerce interactions across chat and voice channels. It is described as able to handle complex customer issues, update accounts, and process payments and refunds while managing end-to-end customer journeys inside Intercom's messaging platform. It is offered as part of Intercom's conversational support suite.
Docling is an open-source document-processing toolkit from the Docling project for preparing files for generative-AI applications. It parses formats including PDF, DOCX, PPTX, XLSX, HTML, EPUB, images, plain text, audio, and video, and represents the results in a unified DoclingDocument structure. For PDFs, it analyzes page layout, reading order, headings, tables, code, formulas, images, and scanned content through OCR. Audio and video inputs can be processed with automatic speech-recognition models; video parsing produces an ASR transcript and representative keyframes. Parsed content can be exported as Markdown, HTML, WebVTT, DocLang, DocTags, or lossless JSON. Docling supports local and air-gapped execution, a command-line interface, Python usage, an API server named docling-serve, and an MCP server for connecting agents. The project also provides integrations for LangChain, LlamaIndex, CrewAI, and Haystack. It is installable with pip and runs on macOS, Linux, and Windows on x86_64 and arm64 systems.
Hark is a webhook-to-iPhone notification platform for CI jobs, agents, scripts, monitoring tools, and other systems that can send HTTP requests. Each service receives a name, avatar, destination URL, and secret webhook endpoint; JSON requests can create rich notifications with sender, body, images, tap destinations, project grouping, summaries, and optional device routing. The platform also supports approval requests and text replies for agent workflows, tracks delivery attempts and registered devices in a web dashboard, and provides stateful Live Activities for task progress on the iOS Lock Screen and Dynamic Island. Hark provides an iPhone app, a Node.js 22-or-newer CLI package named harkctl, and an agent skill; the repository documents a Pro tier for multiple devices and targeted delivery.
Open Science is an open-source, local-first, model-agnostic AI research workbench developed by AIPOCH for scientists and researchers. It runs on macOS, Windows, and Linux and supports computational and data-intensive research in fields including machine learning, statistics, life sciences, chemistry, materials science, physics, and environmental science. Users create projects and sessions, describe research goals in plain language, attach source files, select a model and approval mode, and inspect the agent’s tool activity. Scientific AI agents can read files, search the web, execute Python and R code, query scientific data sources, and produce reports, tables, figures, and other research artifacts. The workbench supports literature review, hypothesis development, code execution, data analysis, simulation, visualization, and related reproducible research workflows. Its first-run setup checks the environment, configures model-provider credentials, optionally prepares Python and R runtimes, and prepares an agent runtime such as Claude Code, OpenCode, or Codex; app-managed runtimes can be installed without Node.js, npm, or administrator privileges. Projects keep sessions, uploads, generated files, and preview state together. Conversations record agent answers and the commands, file reads, edits, searches, and connector calls behind them. Generated artifacts are stored as immutable, checksummed versions; their Provenance view can show producer code and execution history, referenced inputs, the observed environment inventory, the producing conversation branch, and version-scoped reviewer findings, marking unavailable evidence as unavailable rather than guessing.
Vibe-Trading is an open-source AI personal trading agent and quantitative-finance research workspace developed by HKUDS. It turns natural-language finance questions into runnable analysis and backtesting workflows by connecting an agent to market-data loaders, portfolio and backtesting engines, report generation, persistent memory, and multi-agent research teams across asset classes. Its quantitative and risk-analysis capabilities include strategy backtesting, Heston stochastic-volatility pricing, hierarchical risk parity, copulas, and market-microstructure estimators. The project provides API and MCP interfaces, broker and market-data connectors, portfolio-management and broker-authorized trading workflows, examples, documentation, a demo, and a Shadow Account for testing. It supports read-only account sources such as official IBKR MCP and opt-in Binance USD-M snapshots, while scheduling actions require explicit confirmation and live-order safeguards fail closed on contradictory or invalid data. Implemented as a Python project, it uses a connector onboarding system that stores connection secrets in the operating-system keyring and is distributed under an open-source license.
An open-source agent skill for redesigning GitHub README homepages around a repository's actual content. It reads the repository first, identifies the clearest value and supporting proof, and then derives a project-specific visual system rather than applying a shared template. In whole-README mode, it works across content, visual-system, and engineering layers: it removes repetition and internal jargon, moves proof toward the opening, derives typography, color, composition, and project-native motifs, and keeps assets GitHub-safe, accessible, searchable, and copyable. The skill separates visual and content layers by using SVG for editable heroes, section transitions, comparisons, diagrams, and identity while retaining maintainable Markdown for searchable body text. It supports hybrid SVG compositions with optional AI-generated elements, local previews, and an approval step before publishing. The repository documents examples used by eight public repositories, including project-specific heroes and real outputs for slide creation, UI reconstruction, icon generation, agent delegation, vehicle telemetry, interactive mapping, and a solo werewolf game.
text-to-cad is a library of agent skills for creating, inspecting, sourcing, slicing, previewing, and handing off CAD, CAE, CAM, robotics, and hardware-design artifacts from local project files. Its skills generate and edit CAD models from plain-language or image requests, primarily producing STEP files with optional STL, 3MF, and GLB exports; preview local CAD and robot files in a browser; source off-the-shelf STEP parts; and create DXF drawings from Python sources or CAD geometry. Additional skills write URDF robot structures, SRDF/MoveIt2 planning and collision data, and SDF simulator models and worlds; check DXF and STEP files for SendCutSend; measure mesh printability by wall thickness, overhangs, support volume, and build orientation; slice supported meshes into validated, printer-profiled FDM G-code using slicer command-line tools; and dry-run, upload, and cautiously start local Bambu Lab print jobs. An experimental implicit-CAD skill creates browser-native models with GLSL signed-distance fields and CAD Viewer raymarch rendering. The library is installed with the Skills CLI, including `npx skills add earthtojake/text-to-cad`, or through provider-native plugins for Codex, Claude Code, and Grok Build.
AI Job Search is an open-source, local job-application framework built around Claude Code or compatible agent tools. Users create a career profile, search and deduplicate job postings, evaluate and rank matches against structured fit criteria, and run `/apply` to produce tailored CVs and cover letters in LaTeX. The application pipeline uses a drafter–reviewer process: an agent evaluates the posting and drafts the materials, a second reviewer critiques them, and the workflow revises the output; `/interview` supports interview preparation, and optional salary benchmarking is included. The included portal-search skills target Danish services such as Jobindex, Jobnet, and Akademikernes Jobbank, while the profiling, fit-evaluation, and application workflow is described as language- and country-agnostic and adaptable to other job boards. It requires Python 3.10 or later, Bun for the job-search CLI tools, and a LaTeX distribution with `lualatex` and `xelatex`; an optional `pdftotext` installation supports the ATS parseability check for compiled CVs. The repository states that the project is independent of and unaffiliated with Anthropic or Claude Code.
An open-source AI agent skill for Claude Code, Codex, and similar coding agents that produces cinematic product, marketing, launch, and demo videos with Remotion. It provides shot recipe cards, motion previews, a reusable video template, and workflows for storyboarding, real page capture, animation, 2.5D camera moves, sound design, beat-synced cuts, and visual checks. The repository includes native Remotion components driven by normalized progress values, a gallery for browsing motion previews, and an optional JianYing project export that separates shot plates, captions, sound effects, and background music into editable tracks. It can be installed through the skills CLI, by cloning the repository, or by linking it into a Claude Code or Codex skills directory.
eve is an AI agent framework for building durable, structured, multi-agent projects from declarative configuration covering instructions, skills, tools, channels, schedules, permissions, and evaluations. Its agents can be developed within a single folder and run through chat surfaces such as Slack or the eve development TUI. Vercel uses eve as the foundation for a marketing-team template with one lead and five specialists: product marketing, content marketing, social media, SEO, and email. The lead loads shared brand context and user preferences, writes self-contained briefs, routes work one level deep, and returns deliverables produced in Notion, Typefully, or Resend; specialists research and edit their own work against written rubrics rather than spawning further agents. The template uses MCP connections for Notion, Typefully, and Resend, Vercel Blob for shared state and files, Vercel AI Gateway for model access, and Vercel Sandbox for reference files and shell commands. Irreversible actions—including email sends, deletes, scheduled social publishing, and Notion page moves—pause for human approval, while the available Resend tools are restricted to keep account administration out of reach. The project is implemented in strict ESM TypeScript and includes commands for local development, validation, type checking, discovery diagnostics, and deployment.
Impeccable is an open-source design skill for AI coding agents, developed by pbakaus. It provides a shared design vocabulary with 23 commands, including init, shape, critique, audit, polish, harden, animate, typeset, layout, and live. The init command gathers product and design context into PRODUCT.md and DESIGN.md files, while other commands review, plan, modify, or refine a frontend project. Its CLI and browser extension run 61 deterministic detector rules for common AI-generated frontend design problems without an LLM or API key; it also supports LLM-based critique checks. The live command provides browser-based visual iteration, and the package can be installed from a project root with npx impeccable install.
Ratel is an in-process context-engineering platform for AI agents. It catalogs tools, skills, and persistent facts, then searches the catalog on each turn and progressively discloses only the capabilities relevant to the request instead of placing every tool schema and instruction in the context window. Its retrieval supports BM25, semantic, and hybrid ranking without requiring a vector database. The project includes a Rust core, TypeScript and Python SDKs, an MCP server, and a CLI. Tools can be registered directly or ingested from an MCP server; skills can associate instructions with tools, while registered facts are re-injected when they are no longer fresh in the conversation. The repository describes the system as reducing token usage and mitigating tool overload.
SimpleEnglish is an open-source agent skill that makes large language models write technical documentation in ASD-STE100 Simplified Technical English, a controlled language used in aerospace. Its rules enforce short sentences, active voice, explicit instructions, simple tenses, and consistent terminology when rewriting READMEs, error messages, incident reports, runbooks, and release notes. The skill is distributed as a dependency-free folder for tools that support the Agent Skills standard, including Claude Code, Cursor, VS Code Copilot, OpenAI Codex, Gemini CLI, Goose, and OpenCode. It can also be installed as a Claude Code plugin and output style, or used by adding its prompt to a system prompt, AGENTS.md, or .cursorrules file. The repository is licensed under MIT.
Prime Agent is an open-source command-line coding and research agent for general and long-running tasks, developed by Prime Intellect. It combines recursive language-model workflows through an RLM harness with a persistent Python REPL, treating prompts as variables and invoking tools and recursive subagents programmatically. A continual harness stores supplemental prompts, memories, skill descriptions, and reusable subagent specifications as durable session state that can be refined through small, evidence-backed updates without changing the immutable base system prompt. The agent supports importable Python skills, parallel or background child agents, messaging between running agents, automatic context compaction, persistent goals, heartbeats, schedules, detached sessions, autonomous operation, and retained subagents.
Superpowers is an open-source agentic skills framework and software development methodology for coding agents. It begins by eliciting and presenting a software specification for approval, then produces an implementation plan emphasizing red/green test-driven development, YAGNI, and DRY. After approval, it uses a subagent-driven-development process in which agents implement engineering tasks, inspect and review the results, and continue according to the plan. It is distributed as a plugin for supported coding-agent harnesses, including Claude Code, Codex, Cursor, and others.
OpenEdit is an open-source, agent-driven video-editing pipeline from VEED. It has no graphical interface or timeline; an agent-agnostic skill and repository guide drive it through coding agents such as Claude Code, Codex, and Gemini. The pipeline can edit, cut, and reframe footage; create burned-in subtitles, motion graphics, charts, and visual elements; assemble slides, websites, stills, or generated media into video; capture web pages; and connect generation services or MCP servers when needed. Source footage is optional, and VEED Fabric can generate footage only with the user's approval. For captioning, the agent transcribes the video, designs the captions, renders the result, opens a video previewer, and accepts follow-up requests such as changing subtitle color, position, or emphasis. Transcription can use VEED, local offline WhisperX, or a user-supplied service; all providers must produce per-word timings, which are stored in a common transcript format. Caption styles are authored in HTML and CSS. The project includes VEED's closed-source but free-to-use HTML renderer, which does not require a headless browser; Chrome can be used as an alternative rendering backend. V1 primarily targets captions, while motion graphics, charts, and brand-book-matched styling are less exercised. OpenEdit supports Apple Silicon Macs and Windows x64 PCs. Intel Macs are unsupported, Linux support is planned, and Windows requires Git, Node, and ffmpeg. It can be installed from the repository or with `npx skills add veedstudio/open-edit`.
LoopX is an open-source, provider-neutral state kernel and local-first control plane for long-horizon AI agents and peer agent teams. It runs on top of agent harnesses such as Codex, Claude Code, Cursor, dsh, or a custom harness rather than replacing them, keeping objectives, gates, tasks, evidence, quotas, schedules, recovery state, and handoffs durable across bounded turns, sessions, tools, and agents. Its control layer makes semantic decisions about what happens next and supports governance, recovery, human-agent collaboration, and reviewable work. The accompanying personal agent workspace stores goals, attention, conversations, tasks, files, schedules, and recovery state locally across restarts; its supported browser/PWA dashboard can continue work across registered agent sessions and use typed previews, explicit confirmation, and receipts for protected changes.
Pi Web is an open-source local browser interface for the Pi coding agent, developed in the agegr/pi-web repository. It uses Pi’s local configuration and session files, allowing users to browse, resume, rename, export, delete, and branch project-grouped conversations; run agent turns; and inspect running state, context usage, costs, and compaction details. New sessions create independent session files, while “Edit from here” creates a branch within the current session. Its project workspace supports file browsing and uploads, Git diff inspection, automatic previews for source files, Markdown, images, audio, PDFs, and DOCX files, and Git worktree switching. Configuration panels manage provider logins and API keys, models and model tests, plugin packages, and skills. The interface supports English, Simplified Chinese, and Traditional Chinese, follows the browser language initially, and includes a language switcher. Pi Web is distributed through the `@agegr/pi-web` npm package, requires Node.js 22.19.0 or newer, and listens on `127.0.0.1` by default. Remote binding and HTTP Basic Authentication are available, but the documentation warns that Basic Auth does not encrypt credentials in transit and that the service should not be exposed over plain HTTP without a trusted reverse proxy or VPN. It reads Pi agent data from `~/.pi/agent` by default, shares Pi’s model, settings, and credential storage, and limits file browsing to known project or session roots rather than providing general filesystem access.
Kimi Code CLI is an AI coding agent developed by Moonshot AI that runs in a terminal. It reads and edits code, runs shell commands, searches files, fetches web pages, and selects subsequent actions based on the feedback from those operations. It works with Moonshot AI's Kimi models and can be configured with other compatible model providers. The tool is distributed as a single binary and provides an interactive terminal UI. It accepts video input, supports conversational configuration of Model Context Protocol (MCP) servers, and can install skills, MCP servers, and data sources from its marketplace or GitHub repositories. Built-in coder, explore, and plan subagents can work in isolated contexts, while lifecycle hooks can run local commands at selected points in a session. Kimi Code CLI also supports the Agent Client Protocol (ACP), allowing compatible editors and IDEs such as Zed and JetBrains to drive sessions through the `kimi acp` command. The repository documents installation for macOS, Linux, and Windows and provides OAuth or Moonshot AI Open Platform API-key login options.
numbat is an open-source endpoint security tool for monitoring AI-agent activity across supported desktop, CLI, IDE, and gateway agents. It combines local hooks and plugins, OTLP/HTTP logs, and on-disk session artifacts, normalizes live and stored activity into a common event model, and evaluates it locally with built-in, sequence-based, or custom YAML CEL rules. The tool can produce NDJSON events, findings, enforcement decisions, indicators, and scan summaries for alerting and forensic reconstruction, including read-only artifact scans, per-session timelines, and portable case bundles with SHA-256 manifests. Optional pre-action blocking is disabled by default and is limited to supported synchronous hooks and explicitly enforced rules. It is distributed as a single cgo-free binary for macOS, Linux, and Windows, with read-only inventory and scanning commands.
QM is a multiplayer agent harness for startups that operates through Slack and a web interface. It gives each employee and room an isolated scope with its own memory, files, keychain view, permissions, scheduled jobs, web apps, and durable sandbox, while supporting collaboration in channels, group messages, and projects. Shared skills can be granted by scope, promoted by administrators, or imported from Git repositories; background work can run through crons, watches, and inbound webhooks, and internal apps can be published to selected users. A headless TypeScript core runs directly on Node with Fastify and handles identity, policy, scheduling, persistence, and the agent loop. Deployments can select harnesses and models including Pi, OpenCode, Codex, and Claude Code. PostgreSQL stores sessions, memory, queues, and other durable state, while a fixed tool surface includes an execute tool that runs commands in each scope's isolated, persistent sandbox. Slack is an optional in-process Bolt plugin; the web UI, admin panel, and public portal are optional HTTP API plugins built with Vite and Lit. Each deployment keeps organization-specific configuration, tools, skills, sandbox images, and infrastructure in a deployment directory validated and deployed by the qm CLI. Administrators can set organization-wide configuration, available harnesses and models, and a security posture. Strict mode pauses harness tool calls for human approval, Auto mode screens provenance-labelled external data and tool results with a classifier, and Dangerous mode disables content screening and pauses; a predeclared command policy with approvals and hard denials for operations such as recursive deletion or destructive SQL applies in every posture. Deployments run in the operator's own cloud account and are initialized from the @yc-software/qm package without requiring a source checkout; the project documents its threat model, operator assumptions, and known limitations in SECURITY.md.
Bindwidth is a browser-based calculator for sizing on-premises LLM inference infrastructure and estimating total cost of ownership. It evaluates three constraints for a resident text model—KV-cache memory, combined prefill and decode serving capacity, and the runtime session ceiling—and identifies which constraint binds for a workload. The tool accounts for interactive users and autonomous agents as different load types, then compares owned hardware, eligible sovereign rentals, and enterprise subscriptions with per-token APIs for agents. It distinguishes measured from estimated hardware and model data through confidence levels and profile-specific overrides, flags domain and model-capability limits, and exports a decision record as Markdown or the full scenario as JSON. The repository describes it as a directional sensitivity estimator rather than a procurement quote. It is a frontend-only application with no build step, server, database, account, or telemetry; calculations run in the browser and entered data is not transmitted.
Octop is an open-source, self-hosted AI assistant platform for households and small teams. It runs as a single process, storing conversations, workspaces, credentials, and shared state in a SQLite database under ~/.octop/, while serving a web dashboard, CLI, chat integrations, and cron automation. Users can switch among specialized personal agents through its multi-agent architecture. Built on the Harness stack, Octop combines an agent runtime for model routing, tools, skills, and conversation checkpoints with a gateway that normalizes messages from supported chat platforms into one processing pipeline. It also provides a built-in expert library, portable workspace memory, OAuth and MCP connectors, ACP integration for IDE and terminal workflows, browser automation, terminal-assisted command execution, and remote desktop access. The project uses FastAPI and uvicorn, React with TypeScript and Vite, SQLite through aiosqlite, and APScheduler. Its security features include JWT-based multi-user isolation, tool approval, shell-command guardrails, and PII redaction; supported storage backends include local disk, Docker containers, PostgreSQL, and COS/S3.
Phone Harness is an open-source harness that lets AI agents such as Claude Code, Codex, or other LLMs control a real iPhone or Android phone without a jailbreak, injected app, Xcode, WebDriverAgent, or on-device installation. It exposes helpers for opening apps, reading screen content, tapping, scrolling, typing, and verifying actions. On iPhone, it uses the macOS iPhone Mirroring window as the transport: screenshots are captured from the window, Apple's Vision framework performs OCR and returns tap-ready coordinates, and HID-level CGEvents provide taps, long presses, drags, flicks, scrolling, Unicode typing, and shortcuts. On Android, it uses ADB over USB or Wi-Fi, combining screenshots with the phone's accessibility tree for text and element bounding boxes. Actions can be verified by capturing the screen again, with the screenshot serving as the ground truth rather than relying on a DOM. Each invocation is self-contained rather than relying on a daemon.
Ante is a self-contained terminal coding agent and agent harness developed by Antigma Labs. It is distributed as a single Rust executable with no runtime dependencies and is designed to work with provider APIs, subscriptions, or local GGUF models. Its core embeds Grep and git operations in one process and uses a managed version of llama.cpp for local inference. In offline mode, it runs the coding-agent loop entirely on the user's machine without an API key, account, or internet connection; settings profiles can define the agent and its system prompt. The project also supports multiple providers, multi-agent orchestration, MCP skills, and memory. Ante is in beta preview, with breaking changes and incomplete functionality expected. The repository states that macOS and Linux are supported and recommends WSL for Windows.
PartMode is an open-source, local-first browser-based mechanical CAD application built on OpenCascade WebAssembly, replicad, and three.js. It provides constrained sketches, editable feature history, exact B-rep solid evaluation, multiple bodies, assemblies with mates and exploded views, drawings, inspection tools, and import or export for formats including STEP, STL, AMF, 3MF, SVG, DXF, and PDF. People work through the CAD interface, while permissioned typed agents operate against the same canonical document schema and exact geometry kernel. Browser-approved agents can edit the open project; a separate headless workflow uses an account-owned document. The application runs without a desktop CAD installation, supports anonymous local browser modeling with optional hosted agent connectivity, and is distributed under the GNU AGPL v3 license.