704 tools and products — trending open source, and what gets used in AI and other work.
BrowserOS is an open-source web browser that integrates agentic AI capabilities and runs AI agents locally on users' computers. It is presented as a privacy-first alternative to Chrome.
An agentic retrieval-augmented generation platform for ingesting timestamped Markdown, filtering a corpus, and searching it conversationally.
Kilo Code is an open-source AI coding agent and agentic engineering platform from Kilo, available as extensions for VS Code and JetBrains, a command-line interface, and cloud services. It can generate and edit code across multiple files, provide inline autocomplete, run terminal commands and control a browser, and act as an external coding harness for cross-harness model testing. The platform includes Code, Plan, Ask, Debug, and Review agents, supports custom agents, allows model switching during a task, and uses self-checking to review and correct its work. It also provides MCP server discovery, autonomous CI/CD execution, cloud agents, and automated pull-request code reviews.
LobeHub is a self-hostable platform for organizing teams of AI agents, presented as a chief agent operator. It treats agents as units of work and provides tools to create and configure agents, schedule and report on their operations, coordinate them in agent groups, and monitor shared workspaces and memory. The platform provides access to multiple models and modalities, an agent builder, a library of skills and MCP-compatible plugins, and an IM gateway for interacting with agents through existing chat services. It can be deployed using Docker or listed cloud deployment options, and is described as being under active development.
Cangjie Skill is an open-source pipeline that distills methodologies from books, long-form videos, podcasts, interviews, and courses into callable AI Agent Skills rather than summaries. The resulting skills package evidence, examples, executable steps, and usage boundaries so coding agents can invoke the extracted methodology. Its capability-bundle workflow extracts stable capability cards and metadata before compiling installable output. It supports a single router-style Skill or a compact pack containing a router and promoted standalone Skills. The unified local toolchain includes diagnostics, compilation, output replanning, incremental updates, repair, rollback, evaluation, and benchmarking, while content-addressed preprocessing, source diffs, impact analysis, transactional patches, snapshots, and edit detection support controlled evolution. The repository also provides a standalone installation package for DeepSeek Harness.
Copilot Workshops is an open-source collection of guided, hands-on lessons for exploring GitHub Copilot's agentic capabilities across the software development lifecycle. Its content covers the Copilot CLI, VS Code agent mode, the GitHub Copilot app, and the Copilot cloud agent, with Markdown lessons, images, and an Astro + Starlight site that publishes the material; the workshop's Tailspin Toys demo application is maintained in a separate repository. The repository is MIT-licensed.
An extension for the Pi coding agent that delegates work to focused child-agent sessions. Pi acts as the parent session, assigning tasks to subagents and bringing their results back; foreground runs stream in the conversation, while background runs continue asynchronously and can be checked later. Built-in roles include scout for codebase reconnaissance, researcher for web and documentation research, worker for implementation, reviewer for code review and fixes, oracle for challenging assumptions, and delegate for general delegation. It supports parallel and chained workflows, saved workflows, council-style model-based reviews, artifacts, truncation, session sharing, and optional isolated worktrees. The extension is installed with `pi install npm:pi-subagents` and does not start automatic reviews unless requested in a prompt or project instruction.
clawk is an open-source command-line tool that runs AI coding agents in disposable Linux virtual machines instead of directly on the host. It mounts a repository from the host into the guest, keeps host files and keychains outside the VM, forwards the host's SSH agent, and restricts outbound connections with a per-sandbox network allowlist. It supports agents and shells including Claude Code, Codex, pi, and a standard shell; sandboxes can be destroyed and recreated while host-side code and agent conversations remain available, and sessions can be resumed. The repository describes the project as pre-1.0 and subject to breaking changes; allowed destinations and forwarded credentials remain accessible to the agent and can be used to publish data.
waggle is an open-source reference layer for handing off artifacts between AI agents. It replaces pasted files with compact, approximately 30-byte tokens that resolve into an agent-specific view while preserving attribution and recording which parts were read. The token travels without automatically expanding the underlying artifact. Resolve, read, and search operations return only the requested projection or slice within byte budgets, and corrections can propagate to token holders. waggle is MCP-native and runs as an MCP server for agent tools.
mindwalk is a local visualization tool that replays coding-agent sessions on a 3D map of a codebase. Its Go binary reads Claude Code, Codex, and pi session logs, showing searched, read, and edited files as changing glow and touch states on radial-tree or treemap views; files no longer present remain as wireframe ghosts. A playback deck supports timeline scrubbing, event-bucketed observation and mutation phases, speed control, and client-side WebM export, while timeline marks, subagent lenses, and a file inspector connect session events to individual file histories. It runs locally and sends nothing elsewhere during viewing; optional session evaluation sends a session summary to the model used by the user's local Claude or Codex CLI when explicitly invoked.
CLODEx is an open-source, local-first, zero-trust agentic IDE for long-running software engineering work. It maintains durable tasks with searchable history, workspace-aware context, restart recovery, and continued work across sessions, combining code editing, pending edits, line-level diffs, Git and worktrees, persistent terminal sessions, embedded browsing, screenshots, models, and MCP tools in one desktop workspace. Sensitive actions can require explicit approval and remain reviewable; its governing principle is that model output is input, not authority. CLODEx supports account-backed models, bring-your-own provider keys, compatible endpoints, and local Ollama models, and is available for macOS, Windows, and Linux.
Flawless, also named CISRE (Cloud Infrastructure Site Reliability Engine), is an AI SRE and AgenticOps control plane for Kubernetes and cloud infrastructure. It connects risk discovery, evidence collection, diagnosis, Skill-based remediation, human approval, controlled changes, recovery verification, and auditable records into a remediation loop. Its architecture separates model planning and explanation from domain Skills, composable plugins, a Harness for state, permissions, and orchestration, controlled executors for real changes, and Verifiers that determine whether the target has actually recovered. The documented flow is discovery → evidence collection → diagnosis → Skill routing → change preview → human approval → execution → same-target readback → stability verification → records; if recovery fails, the system retains evidence and can continue with another strategy. Kubernetes is documented as having a complete loop, while database, VM/host, storage, middleware, cloud-resource, and network integrations are described as contract-ready extension areas. The platform includes SRE Run, scheduled or manual AI inspections, topology impact analysis, a Skill library, a plugin center, operational dashboards, and Agent Trace. Real changes must pass through a typed action, policy and blast-radius checks, human approval, an executor, same-target readback, a recovery verifier, and a record. Its plugin-first architecture supports declared service dependencies, event-driven orchestration, reversible loading and hot reload, event-sourced audit records, replay, fork, resume, and resource-domain agents.
Albato is an embedded integration platform (iPaaS) that enables SaaS products to add and white‑label native integrations and automated workflows. The service advertises support for 1,000+ integrations and features targeted support for integrations used by AI agents. It is operated from albato.com.
Nitrosend is an AI-native email automation platform and email layer for agencies, SaaS teams, developers, and autonomous agents. It provides agent onboarding, agent-owned inboxes, marketing and transactional email, customer replies, full-stack workflows, and multi-brand campaigns, with human approval gates and a free-to-start model. Agents can be integrated by pointing them to nitrosend.com/SKILL.md.
In Parallel is a platform that turns the conversations, documents and decisions around an initiative into persistent business context that teams and their AI tools and agents can use. It is positioned to provide living context for collaboration and for AI-driven workflows.
Graft AI is an enterprise AI infrastructure product from Axcelner that enables AI agents to interact with legacy, internal, desktop, and browser-only software workflows. It provides integration layers and connectors to expose those systems to agent-based automation while emphasising secure and stable operation.
Verse is an AI platform that lets teams create autonomous AI “employees” from a single prompt without writing code. The agents can access connected customer tools, retain context across interactions, and run automated tasks continuously.
Kit for AI is a service that provides a persistent memory and grounded-knowledge layer for AI agents. It exposes APIs and native memory tools that let agents call stored knowledge directly, supports ingesting files and URLs, and is intended to avoid building a separate retrieval-augmented generation stack. The site states the product works with any model and offers a free-to-start tier.
Agently is an AI Work OS that integrates with a company's software stack to build a centralized "company brain" and deploy autonomous agents to perform tasks and workflows. The platform describes an orchestration component named Jarvis that coordinates those agents across the stack.
YAGNI is a coding-agent platform for engineering teams, offered as a CLI and desktop client. Its agents are grounded in a company’s data and use a shared decision ledger, while per-developer budget controls govern usage. The platform routes requests across vetted, US-hosted open-weight models and supports developer skills, plugins, and custom components; it advertises lower per-call inference costs than frontier APIs.
Higgsfield AI is an AI-native creative suite from Higgsfield that generates images, videos, and synthetic voice from text prompts or reference media, and provides editing and upscaling tools. It includes an AI agent to automate creative workflows and is accessible via web and mobile; it accepts text and media inputs and produces image, video and audio outputs.
Open Interpreter is a command-line coding agent optimized for open and low-cost models. It is a fork of OpenAI's Codex focused on emulating provider-specific agent harnesses; its Rust-native harnesses can be switched with `/harness`, while providers and models can be changed from its terminal interface. It supports native sandboxing on macOS, Linux, and Windows, MCP, skills, hooks, permissions, AGENTS.md instructions, and local configuration and session state under `~/.openinterpreter`. The agent can run as an Agent Client Protocol agent for compatible editors, speak the Codex exec protocol, and provide a one-line binary override for applications using the Codex SDK. Its built-in QA skill can test web applications through a real browser and native applications through trycua. The project is distributed for macOS, Linux, and Windows through shell or PowerShell installation commands and uses shared agent directories and protocols where available.
UI Skills is a command-line tool and collection of skills for design engineers working with AI coding agents. Its registry can be browsed from a terminal with commands such as `npx ui-skills start`, `categories`, and `list --category motion`, while individual skills such as `baseline-ui` can be fetched with `get`. It also exposes the registry to agents over the Model Context Protocol through `list_skills` and `get_skill`, and provides a playbook of distilled design-engineering lessons.
Graphify is a command-line tool and agent skill that turns a codebase and related project materials—including documentation, SQL schemas, configuration files, PDFs, images, and videos—into a locally queryable knowledge graph. It uses deterministic, local tree-sitter AST parsing for code, while non-code materials can receive a semantic pass from the configured AI assistant or API key; it does not use embeddings or a vector store. Graphify labels graph edges as EXTRACTED when they are explicit in the source or INFERRED when resolved by the tool, and can generate a clickable graph.html visualization, a GRAPH_REPORT.md summary, and a graph.json representation for later queries. The CLI is installed as the graphifyy package, registered with `graphify install`, and invoked through `/graphify` in supported AI coding assistants such as Claude Code, Cursor, Codex, Gemini CLI, and GitHub Copilot. The project is developed by Graphify Labs and provides free, fully local code mapping; its homepage describes a separate Graphify platform available through early access.
OpenMontage is an open-source, agent-driven video production system that turns plain-language instructions into video projects through an AI coding assistant. Its workflow covers research, scripting, production planning, asset generation, editing, and final composition, with production pipelines, approval gates, cost estimates, and post-render checks. The system can produce image-based videos as well as videos assembled from real motion footage: its agent can build a corpus from free stock footage and open archives, retrieve clips, edit them into a timeline, and render the result. It runs locally with FFmpeg and Remotion and includes 12 production pipelines, more than 100 tools, and hundreds of agent-skill and production-knowledge files.
codebase-memory-mcp is an MCP server from DeusData that indexes a codebase into a persistent local knowledge graph for AI coding agents. It parses source with tree-sitter AST grammars and uses Hybrid LSP semantic type resolution for selected languages to model functions, classes, call chains, HTTP routes, and cross-service links. Its MCP tools support code search, structural and Cypher-style queries, call tracing, architecture overviews, impact analysis, dead-code detection, and graph visualization. The project is distributed as native executables for macOS, Linux, and Windows, with no language runtime, hosted service, or API key required. Processing is local; the repository describes a RAM-first pipeline using compressed data, in-memory SQLite, and fused Aho-Corasick pattern matching, and reports sub-millisecond structural queries. A local visualization interface is available at localhost:9749.
Page Agent is a JavaScript in-page GUI agent that lets users control web interfaces with natural-language instructions. It operates through text-based DOM manipulation inside the webpage, without requiring a browser extension, Python, screenshots, multimodal models, or a headless browser. It can be integrated with a script tag or installed from NPM, and supports bring-your-own LLM configurations, including locally deployed models. An optional Chrome extension supports multi-page tasks, while a beta MCP server allows external agent clients to control the browser. The project is distributed under the MIT License and is designed for client-side web enhancement rather than server-side automation.
Cognee is an open-source AI memory platform for agents. It ingests data in varied formats and builds a self-hosted knowledge graph that combines vector embeddings, graph relationships, and ontology generation, allowing agents to search by meaning and connect related information across sessions. Its API provides remember, recall, forget, and improve operations, with session memory synchronized to the graph. The project provides Python, Rust, and TypeScript clients, a CLI and UI, MCP support, and integrations for Claude Code and OpenClaw.
omg.dev is an open-source, self-hosted parallel coding-agent harness and control plane developed by BennyKok. It runs coding-agent sessions on a local computer or hosted Computer, keeps them running when the web interface disconnects, and provides a single web UI for parallel sessions, transcript reading, follow-up instructions, chats, bots, schedules, and notifications. The project supports agents including Claude Code, Codex, Grok, Cursor, OpenCode, Jcode, GitHub Copilot, and Pi, using the user's existing agent subscriptions or API keys. Local installations run through a Bun CLI and expose the interface on localhost; remote phone access can be provided through Tailscale. The repository is licensed under the MIT License.
loop.js is a TypeScript framework and CLI for loop engineering: it runs an agent against a declared goal and verification criterion until a separate, skeptical Verify agent judges the work complete. Each round starts with fresh context and reads state and handoff notes from disk; the worker agent cannot approve its own work, and a failed verification supplies a reason for the next round. The framework provides round, spending, and per-round timeout guards; typed exits; crash recovery; idempotent resumption from a stored cursor; and a compare-and-set lock with heartbeat to prevent overlapping live runs and take over dead ones. It can run from a terminal, be embedded in a product, or be scheduled through crontab, launchd, Task Scheduler, or Modal's cloud.
self-learning-skills is a meta-skill for AI coding agents, including Claude Code, Cursor, Codex, and agents that read AGENTS.md or similar standing-instruction files. It recognizes hard-won, reusable procedures such as non-obvious commands and recurring operational workflows, then captures them as skills, rules, or project instructions for later sessions. The capture includes known dead ends and excludes one-off details and secret values. The project stores learned procedures as new skills/<name>/SKILL.md files for Claude Code, Codex, and Agent Skills clients; as .cursor/rules/learned/<name>.mdc files for Cursor; or in AGENTS.md or project notes for other agents. It can be installed with the community skills CLI, as a Claude Code plugin, or manually by copying the relevant files.
Three.js Game Skills is a set of self-contained Codex and Claude Code agent skills for building playable Three.js browser games. Its threejs-game-director skill routes work across gameplay, graphics, UI, asset generation, audio, debugging, and release verification, so users can request an outcome without manually selecting specialist skills. The package includes SKILL.md files, references, checklists, prompt templates, helper scripts, and a Vite, TypeScript, and Three.js scaffold. Generated games include deterministic test hooks and seeded randomness, with Playwright templates for smoke tests, visual-regression baselines, and bot playtests; the documented workflow also checks browser behavior, mobile viewports, visual output, UI, performance, and release readiness. It is created by Majid Manzarpour and can be installed for Codex or Claude Code through the skills CLI or the repository's installer.
Palmier Pro is a Swift-native macOS video editor built for AI, with generative video and image models integrated into the timeline. It can expose a local MCP server at http://127.0.0.1:19789/mcp for connections from Claude, Codex, and Cursor, and includes an in-app agent that can create and edit within the same project. It requires macOS 26 (Tahoe) or later on Apple Silicon. Releases through v0.7.6 and the corresponding published source are available under GPLv3; later binary releases are proprietary.
deepsec is an open-source, agent-powered vulnerability scanner and security harness from Vercel Labs for on-demand review of large codebases within the user's own infrastructure. It identifies candidate vulnerabilities, uses AI models to investigate them, supports triage and optional revalidation to reduce false positives, and can run work across parallel worker machines. The CLI initializes a repository with `npx deepsec init`, stores state and findings in a `.deepsec/` directory, and resumes interrupted runs while skipping files already analyzed. Subsequent commands separate fast pattern scanning from AI processing and revalidation, and findings can be exported as Markdown directories or JSON. Model access can use Vercel AI Gateway, direct OpenAI or Anthropic credentials, or a custom HTTPS provider; scan cost and duration can be capped with command-line options.
LLM Space v4 is a local-first desktop workbench for prototyping and developing AI agents. It lets users version prompts, system messages, tools, and model settings; trace each model call and tool run in an agent loop; replay historical runs to debug failures; and evaluate agent performance across runs. It can generate prompts and tools with AI and turn a thread into a runnable LangGraph agent. Threads, project files, and API keys remain on the user's computer, while the project describes the app as cloud-ready for managed agents. The desktop application uses Electrobun with a React, Tailwind CSS, and shadcn/ui interface, and is built as a Bun monorepo around Pi Agent Core.
OpenAI Evals is an open-source framework for evaluating large language models and systems built with them, including tool-using agents and prompt chains. It includes a registry of benchmark evals, supports custom model-graded and private evals based on a user's data, and uses a completion-function protocol for advanced workflows. The package can be installed with pip and run locally with an OpenAI API key; the repository also documents optional result logging to Snowflake and configuration through the OpenAI Dashboard.
Claude Opus 4.5 is an AI foundation model developed by Anthropic for complex reasoning, software engineering, and agentic workflows. It is available through Anthropic's API and Claude applications.
Vercel Connect is a Vercel service for connecting autonomous agents and applications on the Vercel platform to external systems, APIs and event sources while mediating and restricting data and event access. Developed by Vercel as part of its agentic infrastructure, it provides connectors and permission controls that determine what data and events an agent can read, transmit, or act on. The service is exposed via Vercel's agent framework and platform APIs and runs on Vercel's hosting infrastructure.
Fluree is an enterprise AI data platform that turns raw, structured, and unstructured data into trusted, queryable knowledge graphs. Its platform uses governed data models, taxonomies, golden records, entity resolution, semantic layers, and GraphRAG-powered retrieval to prepare enterprise data for conversational analytics, enterprise search, decision intelligence, and AI agents. Users can ask questions, create dashboards, build applications, and deploy governed agents over the resulting knowledge graph.
Freesolo is a full-stack platform for post-training of small models. The company offers a managed post-training service (mentioned in videos as 'Flash') that uses an agent-driven workflow combining supervised fine-tuning and reinforcement learning to produce specialized, downloadable small-model weights and to deploy models via an OpenAI-compatible API.
AskCodi is an AI coding assistant and developer platform that provides multi-model AI chat, code generation, refactoring, and custom agent capabilities. It offers an OpenAI-compatible API, IDE integrations, and developer-focused workflows.
Moxie Docs is a hosted service that automatically generates and maintains documentation from GitHub repositories. It re-checks docs on every merge, opens cleanup pull requests when documentation drifts, and provides context for AI agents plus hosted public help centers.
Notte is a browser framework and platform for building and running AI agents, providing tools to simplify web-based LLM tasks, handle CAPTCHAs, and offer developer tooling for agent development and deployment.
AgentLoop is a developer tool that runs unattended coding loops on a user's machine, executing plans authored in ChatGPT by launching Codex-based workers and independent critic processes. It automates iterative code generation and evaluation until configured quality standards are met.
Blaxel is an infrastructure platform for autonomous AI agents that provides isolated microVMs (which boot in milliseconds and resume in ~25ms), persistent shared memory, and programmable networking to run, connect, and operate agents at scale. It offers a distributed shared-storage feature (presented in videos as "Agent Drive") that can be mounted as a POSIX-style drive so agents in separate sandboxes can share files, cache dependencies, exchange context/memory, and coordinate parallel work in real time.
Last 30 Days is an open-source AI-agent research skill developed by mvanhorn. It searches configured sources in parallel—including Reddit, X, YouTube transcripts, TikTok, Hacker News, Polymarket, GitHub, arXiv, Techmeme, and other web and social platforms—then ranks material using signals such as upvotes, likes, views, comments, engagement, and real-money prediction-market odds. An AI-agent judge synthesizes the results into a cited brief focused on recent information, with source-specific outputs such as quotable transcript passages, Reddit comment scores, GitHub activity, and Polymarket probabilities. Reddit, Hacker News, Polymarket, and GitHub work without additional configuration. A setup wizard can enable further sources using user-provided API keys, browser sessions, or locally available source connectors. The skill can be installed through the Claude Code marketplace or with the Skills CLI for Claude Code, Codex, Cursor, Copilot, Gemini CLI, and other agent-skill hosts.
Awesome LLM Apps is an open-source GitHub collection of Python templates for AI agents, agent skills, and retrieval-augmented generation (RAG) applications. Its examples include starter and advanced agents, multi-agent teams, voice and multimodal agents, MCP-based agents, memory-enabled agents, data analysis, web scraping, and other application patterns. The repository provides runnable source projects and step-by-step tutorials. Individual agents can be cloned and run with an API key, while agent skills can be installed with a command such as `npx skills add` and used with coding agents including Claude Code, Codex, and Cursor. The projects support models from providers and ecosystems including Claude, Gemini, GPT, DeepSeek, Llama, and Qwen. The repository describes the collection as free and open source under the Apache-2.0 license.
Hubble.md is a free, open-source Markdown and HTML note-taking app for people and AI agents. It provides a Notion- or Apple Notes-style editor with Markdown shortcuts, slash commands, and frontmatter properties, while allowing an agent to work directly with the notes folder; edits made by an agent are live-reloaded in Hubble. Beyond Markdown notes, Hubble can display HTML-based apps built from a notes folder, including views such as tables, bookshelves, and maps. It is distributed as an Electron desktop app for macOS, Windows, and Linux, and the repository also contains a web app, Markdown editor packages, an HTML-app runtime, a filesystem sync engine, and a `hubble` command-line interface.
Native SDK is an open-source toolkit for building native desktop applications. It uses declarative markup in `.native` files for views, plain TypeScript compiled to native code at build time or Zig for application logic, and an engine that renders directly into real operating-system windows without a browser, WebView, or JavaScript runtime in the application binary. The CLI can scaffold, run, check, and build applications, with live view updates during development and an optimized release-binary build. Its component catalog includes controls such as buttons, tabs, text fields, dialogs, charts, and virtual lists, while styling is organized around replaceable design tokens. The toolkit also includes an embedded automation server for agents.
DwarfStar is a self-contained, model-specific native inference engine for running DeepSeek V4 Flash, GLM 5.2, and, on high-memory systems, DeepSeek V4 PRO locally. It combines model loading, prompt rendering, tool calls, persistent KV state, an HTTP server, and a coding agent rather than serving as a general GGUF runner. The engine supports Metal on Macs, CUDA including multi-GPU systems, and ROCm on Strix Halo hardware; it also provides SSD streaming for machines with insufficient RAM, distributed tensor and pipeline parallelism, and server-side micro-batching for multi-user inference. The repository includes tooling and data for GGUF, imatrix, quality, and speed testing, and acknowledges code and design contributions from llama.cpp and GGML.
Nimbus is an open-source documentation-site scaffolder by Cloudflare that builds Astro sites directly in a repository. It writes layouts, components, styles, routes, content collections, and theme tokens as editable source files, while its plumbing is provided through npm packages. The generated sites include agent-readable Markdown and MDX page twins, llms.txt and llms-full.txt files, JSON-LD, sitemaps, robots.txt, per-page Open Graph images, full-text search, theming, accessible navigation, and optional versioned documentation. Nimbus also provides a registry for adding editable components, utilities, and agent-handoff recipes to an existing project. Sites produce static output for deployment anywhere, with Cloudflare as a first-class target, and include prose, structure, MDX, and configuration validation. The project is built on Astro, Sätteri, Tailwind, and optionally React; its README describes it as work in progress and pre-1.0.
world-model-optimizer converts OpenTelemetry traces from agent workflows into a routing policy. It scores models on held-out tasks, selects a model for each request, and provides world models for testing changes to prompts, tools, and runtime code.
CodeJury is a terminal-first, knowledge-grounded multi-agent software delivery pipeline for scoping requirements, implementing changes, running tests, and gating pull requests. It divides delivery into six terminal-driven stages and uses isolated development branches, human approval gates, deterministic QA, and ensemble code review. The system indexes a repository into a persistent code graph and uses semantic search to retrieve relevant code during later work. Its review stage sends the same change to independent models from different providers; both must approve it for the change to pass, while a rejection returns the reason to the developer. A configurable panel can use up to six specialist reviewers and a foreperson to reconcile their findings. The project is distributed as a Python package installable with pip and documents support for macOS, Linux, and Windows. Its repository describes a one-command run and operation without an API key.
TrueDeck is a terminal-first multi-agent coding workbench for running several coding agents on the same repository in separate live terminal panes. It supports agents including Grok, Codex, Cursor, Claude, and Gemini while retaining their native command-line interfaces rather than presenting them in a chat webview. It provides split panes, tabs, focus control, and workspace restoration, with repository-specific and global memory plus MCP and project-context synchronization handled in the background. Installers and portable builds are distributed through the project's GitHub Releases, and the repository is licensed under the MIT license.
Cynative is an open-source framework and CLI for building security agents with live, read-only access to code, cloud, and runtime infrastructure. It reasons across GitHub, GitLab, AWS, GCP, Azure, and Kubernetes as one system, using frontier models to investigate questions such as exposed resources, privilege escalation paths, infrastructure drift, and leaked credentials. For each investigation, Cynative generates and runs code in an ephemeral sandbox, queries APIs in parallel, cross-checks findings against live evidence, and traces findings to their origins. Its action gate resolves each call to the required IAM actions and applies a read-only policy before attaching credentials; the video also describes an auditable JSONL log of tool calls. Agents are defined as Markdown files containing a description and prompt, and can run interactively, non-interactively, or with piped findings. The project is distributed as a single open-source binary and supports configuring the model provider and model through environment variables.
optim-plans is a human-in-the-loop planning plugin for Claude and Codex. It turns repository-change requests into traceable Markdown plans, asks planning questions one at a time, records decisions, applies reviewer or criticizer refinement passes, and requires explicit scope confirmation before writing versioned plan artifacts. Its five public skills cover small, broad, and high-risk planning, diagnosis before planning, and reference analysis before planning. After approval, it hands the accepted plan back to the current agent session for normal implementation. Version 0.3.0 keeps machine state in the Git common directory and public artifacts in repository documentation, but does not include a separate execution engine, manifest, delegated implementation role, verification role, retry loop, or worktree execution gate.
hwatu is a Linux visual-verification browser and harness for AI coding agents, implemented as a WebKit daemon. It performs one-call page checks, returns pixel-difference scores and heat maps, exposes animation timing numerically, and runs headless windows that can be handed off to a live human-controlled window. It connects to coding agents through MCP, short CLI commands, or newline-delimited JSON over a Unix socket, and is designed for tiling window managers including Hyprland, sway, niri, and i3. The project distributes a static binary that uses the system's WebKitGTK 6.0 and can also be built from source with Cargo.
Skill Recorder is a Microsoft desktop application that records a work session—including screen activity, application and window switches, browser pages, clipboard previews, and optional spoken narration—and uses the GitHub Copilot CLI to reconstruct the session as an overall intent and ordered steps. Users can review and edit the analysis, then generate either a reusable SKILL.md procedure for an AI agent or an Automation that runs on a schedule or trigger. Generated procedures prefer an agent’s native tools, such as the gh CLI or web_fetch, over replaying interface clicks and can generalize from the recorded example. Recording, frame extraction, local storage, and optional narration transcription occur on the computer; choosing Analyze sends the event timeline, extracted screen images, and narration text to GitHub’s cloud for Copilot processing. Narration is transcribed on-device using a Whisper model. The application is distributed as a source release that downloads a pinned Node.js runtime and builds the release locally without a global installation. macOS is the primary target, with Windows 11 support for x64 and ARM64 and Ubuntu installation instructions. A GitHub account with Copilot access is required.
Ratchet is a code-auditing tool for examining coding-agent edits. It checks changes for new dependencies, duplicate helper code, thin wrappers, and reimplemented standard-library features, while tracking file, dependency, and line-count budgets in an auditable ledger.
Cata-centavo is a self-hosted MCP server that gives compatible AI agents read-only access to a user's Brazilian Open Finance bank, credit-card, and investment accounts through Pluggy. It runs locally over stdio, reads existing Pluggy connections identified by environment variables, and supports queries about transactions, spending categories, account balances, current card statements, unfamiliar charges, and investments. It has no hosted or multi-user service. The package runs with Node.js 22.13 or newer and can be launched with npx. Its main commands include the MCP server itself, `init` for checking credentials and configured connections, and `doctor` for diagnosing connection status, consent, local cache, and learned categorization data. Provider transaction categories are copied into a persistent local `data.db`; the server also learns merchant-category associations from the user's transactions, and manually corrected or newly assigned categories are retained locally and applied retroactively.