579 tools and products — trending open source, and what gets used in AI and other work.
world-model-optimizer converts OpenTelemetry traces from agent workflows into a routing policy. It scores models on held-out tasks, selects a model for each request, and provides world models for testing changes to prompts, tools, and runtime code.
CodeJury is a terminal-first, knowledge-grounded multi-agent software delivery pipeline for scoping requirements, implementing changes, running tests, and gating pull requests. It divides delivery into six terminal-driven stages and uses isolated development branches, human approval gates, deterministic QA, and ensemble code review. The system indexes a repository into a persistent code graph and uses semantic search to retrieve relevant code during later work. Its review stage sends the same change to independent models from different providers; both must approve it for the change to pass, while a rejection returns the reason to the developer. A configurable panel can use up to six specialist reviewers and a foreperson to reconcile their findings. The project is distributed as a Python package installable with pip and documents support for macOS, Linux, and Windows. Its repository describes a one-command run and operation without an API key.
TrueDeck is a terminal-first multi-agent coding workbench for running several coding agents on the same repository in separate live terminal panes. It supports agents including Grok, Codex, Cursor, Claude, and Gemini while retaining their native command-line interfaces rather than presenting them in a chat webview. It provides split panes, tabs, focus control, and workspace restoration, with repository-specific and global memory plus MCP and project-context synchronization handled in the background. Installers and portable builds are distributed through the project's GitHub Releases, and the repository is licensed under the MIT license.
Cynative is an open-source framework and CLI for building security agents with live, read-only access to code, cloud, and runtime infrastructure. It reasons across GitHub, GitLab, AWS, GCP, Azure, and Kubernetes as one system, using frontier models to investigate questions such as exposed resources, privilege escalation paths, infrastructure drift, and leaked credentials. For each investigation, Cynative generates and runs code in an ephemeral sandbox, queries APIs in parallel, cross-checks findings against live evidence, and traces findings to their origins. Its action gate resolves each call to the required IAM actions and applies a read-only policy before attaching credentials; the video also describes an auditable JSONL log of tool calls. Agents are defined as Markdown files containing a description and prompt, and can run interactively, non-interactively, or with piped findings. The project is distributed as a single open-source binary and supports configuring the model provider and model through environment variables.
optim-plans is a human-in-the-loop planning plugin for Claude and Codex. It turns repository-change requests into traceable Markdown plans, asks planning questions one at a time, records decisions, applies reviewer or criticizer refinement passes, and requires explicit scope confirmation before writing versioned plan artifacts. Its five public skills cover small, broad, and high-risk planning, diagnosis before planning, and reference analysis before planning. After approval, it hands the accepted plan back to the current agent session for normal implementation. Version 0.3.0 keeps machine state in the Git common directory and public artifacts in repository documentation, but does not include a separate execution engine, manifest, delegated implementation role, verification role, retry loop, or worktree execution gate.
hwatu is a Linux visual-verification browser and harness for AI coding agents, implemented as a WebKit daemon. It performs one-call page checks, returns pixel-difference scores and heat maps, exposes animation timing numerically, and runs headless windows that can be handed off to a live human-controlled window. It connects to coding agents through MCP, short CLI commands, or newline-delimited JSON over a Unix socket, and is designed for tiling window managers including Hyprland, sway, niri, and i3. The project distributes a static binary that uses the system's WebKitGTK 6.0 and can also be built from source with Cargo.
Skill Recorder is a Microsoft desktop application that records a work session—including screen activity, application and window switches, browser pages, clipboard previews, and optional spoken narration—and uses the GitHub Copilot CLI to reconstruct the session as an overall intent and ordered steps. Users can review and edit the analysis, then generate either a reusable SKILL.md procedure for an AI agent or an Automation that runs on a schedule or trigger. Generated procedures prefer an agent’s native tools, such as the gh CLI or web_fetch, over replaying interface clicks and can generalize from the recorded example. Recording, frame extraction, local storage, and optional narration transcription occur on the computer; choosing Analyze sends the event timeline, extracted screen images, and narration text to GitHub’s cloud for Copilot processing. Narration is transcribed on-device using a Whisper model. The application is distributed as a source release that downloads a pinned Node.js runtime and builds the release locally without a global installation. macOS is the primary target, with Windows 11 support for x64 and ARM64 and Ubuntu installation instructions. A GitHub account with Copilot access is required.
Ratchet is a code-auditing tool for examining coding-agent edits. It checks changes for new dependencies, duplicate helper code, thin wrappers, and reimplemented standard-library features, while tracking file, dependency, and line-count budgets in an auditable ledger.
Cata-centavo is a self-hosted MCP server that gives compatible AI agents read-only access to a user's Brazilian Open Finance bank, credit-card, and investment accounts through Pluggy. It runs locally over stdio, reads existing Pluggy connections identified by environment variables, and supports queries about transactions, spending categories, account balances, current card statements, unfamiliar charges, and investments. It has no hosted or multi-user service. The package runs with Node.js 22.13 or newer and can be launched with npx. Its main commands include the MCP server itself, `init` for checking credentials and configured connections, and `doctor` for diagnosing connection status, consent, local cache, and learned categorization data. Provider transaction categories are copied into a persistent local `data.db`; the server also learns merchant-category associations from the user's transactions, and manually corrected or newly assigned categories are retained locally and applied retroactively.
Scopey is a lightweight Rust command-line tool for keeping Claude Code, Codex, Grok, Pi, and OpenCode coding-agent sessions aligned with the user's current intent. It turns prompts into a current scope, evaluates agent tool activity against that scope, and surfaces sessions that need attention when work drifts. It can inject a correction into an active session and send notifications, while ignoring subagent sessions by default. Scopey is pre-1.0 software and is distributed as a Homebrew package or Rust source installation.
Paritok-4B-v1 is an open-source 4B model for compressing coding-agent context, including tool outputs, file reads, and older conversation history. Used through a local proxy, it replaces compressed blocks with reference IDs so an agent can request the exact original content when needed. The repository says the model was trained on 45K teacher-distilled agent trajectories using a Qwen3-4B backbone.
Kiro Crew is an open-source persistent workspace for development work, developed by Kiro. It runs locally or remotely on the user's hardware and preserves ongoing work beyond a single session. Users can access it through a desktop app, web dashboard, CLI, Slack, or Discord; multi-step tasks can run unattended, recurring jobs follow schedules, and heartbeats monitor systems until attention is needed. Kiro Crew Apps combine purpose-built interfaces with agents, skills, schedules, integrations, and backend services. The project can run through a desktop application, a local or remote installation, a Docker image, or a source build, with kiro-cli underneath.
auteur is an agent skill for directing landing-page development like a film. It requires an art-direction commit sheet before markup, then runs a pipeline of asset generation, site construction, and quality gates. The skill includes reference recipes and runnable scripts for generating assets with local CLIs, building WebGL or scroll-driven scenes, and checking styling, accessibility, motion, browser errors, and anti-slop rules with an executable linter and real-motion check. It is distributed as a dependency-free SKILL.md package with installation commands for Claude Code and other agent environments.
Cambium is a governance standard and reference toolset for knowledge repositories maintained by LLM agents. It governs repository work rather than supplying a knowledge corpus, performing retrieval-augmented generation, scheduling agents, or deciding domain meaning. Its governance model combines a kernel of cross-domain rules and invariants, one adopter-selected profile, and adopter-owned runtime state containing task state, queues, plans, deltas, receipts, and recovery evidence. The toolset uses deterministic checks, controlled writers, schemas, and generated projections to manage task and batch transitions, amendments, standards adoption, interruption recovery, and work closure. It provides a profile template and scaffolder, a machine-readable adoption interview, an onboarding status view, persistent coverage and progress state, append-only receipts, terminal-proof bindings, and validation for global maps, capability matrices, and gap registers.
Codex iOS Assistant is an open-source system for controlling and inspecting an iPhone from Codex running on a Mac. Its `iphone` CLI sends private commands through iMessage to an iOS Shortcut; the Shortcut performs native iOS actions and returns device context through an authenticated HTTPS response routed via Cloudflare Tunnel to a receiver on the Mac. A per-user macOS LaunchAgent runs the Messages sender outside the Codex sandbox over a Unix socket. The system can read visible screen text, save screenshots, access or replace the clipboard, manage alarms, open specified apps, control timers, the flashlight, and Low Power Mode, place calls, and read Mac Contacts and the Mac Messages database. Message composition creates a draft for review; the CLI does not send ordinary messages, make purchases, install apps, order food, or request rides. It requires macOS 14 or newer, Python 3.11 or newer, Messages, Xcode Command Line Tools, an iPhone reachable at a configured iMessage address, iCloud sync for Shortcuts, and a Cloudflare-managed domain. Installation uses the repository scripts and creates an iOS Shortcut containing 95 native Shortcuts actions, with a manually configured iPhone message automation.
KADATH (Kernel for Agentic Darwinian Adaptation, Tooling, and Heredity) is an evolutionary multi-agent runtime developed by i3t4an. It evolves populations of complete agents across reproducible epochs: an Architect turns a goal into a measurable benchmark, organisms attempt the goal, an independent Grader scores their evidence, and selection, mutation, reproduction, and replacement produce later generations. Each evolvable genome includes an agent's system prompt, Python implementation, tools, dependency declarations, and supporting files, rather than only a prompt or response. The kernel keeps the objective, benchmark, scoring, scheduler, lineage, database, deadlines, containers, and isolation rules outside the genome; agents can change only during the post-grade mutation phase. The interactive runtime asks for an OpenAI API key and model ID, then collects the goal, epoch duration, population size, and epoch count. Its local runtime prepares Docker-based PostgreSQL, MinIO, LiteLLM, and SearXNG services.
soundshuman is an open-source writing-analysis and repository-auditing toolkit for identifying patterns associated with AI-generated prose. It combines a humanizer skill, a machine-readable rule pack, and a zero-dependency CLI: the skill uses a 41-pattern catalog, statistical signals, voice calibration from writing samples, and a draft–audit–final workflow; the rule pack stores vocabulary tiers, phrase fixes, and weighted regular-expression detectors in JSON; and the CLI scores text from 0 to 100, explains findings with line numbers and triggered rules, applies selected mechanical fixes, scans repositories, and can gate CI. The project states that its goal is inspectable editing rather than defeating AI detectors, with rewrites delivered as reviewable Git diffs and a no-invented-facts rule. The CLI runs on Node 18 or later without dependencies and supports pasted text, files, repository audits, and directory scans. The skill is distributed as plain Markdown for agent harnesses and as a Claude Code plugin; the repository also includes reference materials, a pre-commit hook, CI integration, and installation instructions. It is MIT licensed, with upstream copyrights preserved in the license.
Vibe Watch is a wearable M5Stack StopWatch controller for monitoring and operating multiple AI coding-agent sessions. Its ESP32-S3 drives a round touchscreen that shows six agent states, while physical controls provide selection, approve/reject actions, FAST actions, Plan mode, assistant access, and hold-to-talk voice input. The device communicates over Bluetooth Low Energy HID and combines visual state indicators with configurable sound and haptic feedback.
DeepSeek and Destroy (DSD) is a long-horizon coding-agent skill for executing large implementation plans through persistent worker loops. A premium parent agent handles architecture, decomposition, decisions, and phase gates, while fresh lower-cost specialist workers implement, review, repair, and recheck repository-scale changes. The default worker backend is external OpenCode using opencode-go/deepseek-v4-flash, and the skill is designed for Codex, Claude Code, OpenCode, Kilo Code, and comparable coding harnesses. Each task is rendered into an immutable JSON contract and managed through attempt directories containing prompts, reports, logs, state, and scope baselines. Its normal loop uses an Implementer, a fresh independent Reviewer, and—when needed—a fresh Fixer followed by another Reviewer. Python-based gates check objective integrity such as hashes, lifecycle, and exact source movement; an optional read-only Evidence Clerk interprets already-produced evidence for the parent but cannot invent proof, rerun verification, repair code, waive integrity failures, approve work, or recurse into another Clerk. Documentation is split between always-relevant parent rules and cold-loaded workspace, harness, transport, compaction, prompt, and specialist-role materials.
wallfacer is a terminal session manager for AI coding agents, developed by pradipta. It indexes sessions from Claude Code, Cursor CLI, Kiro CLI, and Codex without modifying their transcript files, storing titles, tags, and projects separately in a local SQLite database. A full-screen terminal user interface and command-line subcommands provide session browsing, renaming, tagging, project grouping, searching across titles, prompts, directories, projects, and tags, launching or resuming sessions, and safe deletion with a trash step before permanent purge. It is distributed as a Go binary, with installation through Homebrew, Go, or pre-built releases, and supports JSON output for scripts.
NeuroArxiv is an agent skill that makes Claude Code or Codex CLI check arXiv prior art before proposing non-trivial architectures, algorithms, or systems techniques. It maps a problem to relevant arXiv categories, fetches papers over HTTP, reads sources independently to avoid cross-source anchoring, and converges them into a cited recommendation with a first implementation step and documented limitations or prior failures. The repository describes evaluations across physics, applied mathematics, quantitative biology, machine learning, and statistics, including checks for weaknesses in cited papers. It is distributed as a bundled skill that can be installed with npx without cloning or adding a build step.
Waku is a native desktop app for working with local coding agents. Built in Rust with GPUI, it keeps projects, sessions, transcripts, and application state on the user's machine, while detecting supported agent CLIs and using each provider's native structured protocol and session continuity. Its interface supports independent project sessions, model and reasoning settings, queued or directed follow-up messages, and Git-backed task rewind with conversation-aware checkpoints. The desktop app communicates with a standalone waku-daemon over an authenticated, versioned WebSocket protocol. The daemon manages task data in SQLite, attachments, provider-native session forks, workspace files, and Git operations; a browser client uses the same generated protocol. Waku is distributed for macOS, Linux, and Windows, with no Waku account or remote service required.
loomfeed is a self-hosted, open-source Reddit alternative for AI agents and humans, developed as a community platform with communities, posts, threaded comments, voting, and agent debates. AI agents have identities, API keys, trust scores, and reputations, and can publish, reply, and vote through REST, MCP, or the A2A protocol. The platform records provenance for agent-generated content, including sources, confidence scores, model information, and generation methods. A typed citation graph relates claims as supporting, contradicting, extending, or quoting one another. Epistemic status labels such as Hypothesis, Supported, Contested, Refuted, and Consensus, along with source checking and research-depth scoring, are used to assess content; only human participants can grant an agent post the Human Seal of Approval. It also provides structured head-to-head agent debates and falsifiable predictions with resolve-by times and Brier-scored accuracy records. The repository is MIT-licensed and supports self-hosting with Docker Compose. Its documented stack includes a Go backend, Next.js frontend, PostgreSQL with pgvector, Redis, an API, an MCP gateway, and a web frontend; optional LLM providers, OAuth, analytics, and email services remain disabled until configured.
pi-peer is a pi extension for discovering and messaging coding-agent sessions. It provides `list_peers` to show sessions, working directories, and statuses, and `message_peer` to send short plain-text messages to a named session. Messages carry no conversation history or files and arrive marked as peer text with no authority: they cannot approve actions, change configuration, or execute slash commands. Local messages use disk-backed mailboxes that survive restarts and queue mail for inactive sessions; sessions can also communicate across team machines through a shared Redis server. The extension supports inbound-message policies and stores local peer data with restricted filesystem permissions.
Hiring Agent is an open-source Python resume-to-score pipeline. It extracts structured data from resume PDFs, enriches candidate information with GitHub signals, and evaluates projects, work experience, and technical skills with supporting evidence to produce an explainable evaluation. The repository describes it as a ranking aid for prioritizing resume review, not an applicant-tracking system or a product for HackerRank customers. It includes local-model configuration, with Gemma 4 listed as the default demonstration model, and provides command-line setup and usage instructions.
oh-my-pi is an open-source terminal coding agent and fork of Pi by Mario Zechner, with an IDE-oriented interface wired into the agent. It supports context-aware code edits, code search, language-server operations such as renames, debugging through debugger operations, persistent Python and Bun code-execution workers, web browsing, sub-agents, and MCP; the project describes support for dozens of model providers and includes built-in tools. It runs on macOS, Linux, and Windows, with installation options including a shell installer, Homebrew, Bun, Nix, and PowerShell.
Klaat Code is an open-source, terminal-native AI coding agent from KlaatAI. It reads and edits project files, runs shell commands, performs type checks after edits, and asks for permission before risky operations. It supports planning, sub-agents, MCP, hooks, and skills, and indexes projects into a code knowledge graph containing symbols, callers, callees, and semantic-search data for targeted exploration. Its hosted Klaatu-o1 router classifies each request and dispatches it across six model-cost tiers—nano, fast, code, reason, heavy, and titan—rather than using one model for every task. The CLI is a thin client to this service; the repository states that routing decisions, model health tracking, pricing, and the code-graph index run server-side. It supports providers and models including Claude, GPT, Gemini, and DeepSeek, and can be installed with npm.
Frontman is an open-source, browser-based AI coding agent for visual frontend editing. It lets users select an element in a running application, describe a change in plain language, and edit the underlying source files with hot reload. The agent uses the live DOM, component tree, computed CSS, routes, source maps, and server logs to identify the relevant code, then presents the resulting diff for review. It supports frontend projects using frameworks and tools including Next.js, Astro, Vite, React, Svelte, and Vue, and is distributed under Apache 2.0 / AGPL-3.0 as listed in its repository.
Exo is an AI agent harness for recursive self-improvement, providing tools, tasks, integrations, and a layered architecture that separates executive policy, protected state, and sandbox execution. It can inspect its own code and runtime logs, incrementally modify prompts, memory, tools, and harness policy, clone itself, and manage a lineage of clones. An event log remains outside these modifications as a canonical record intended to prevent recursive loops.
Human Review is a local visual review tool and agent skill for editing HTML and Markdown files, reviewing localhost pages, and sending feedback to an AI coding agent. It opens a file or URL in a browser, where users can edit text and basic formatting, rearrange or remove page elements, resize or paste images, and attach comments to exact phrases, images, charts, or sections. Direct HTML edits and image changes save automatically; Markdown and localhost reviews are sent to the agent, which applies the changes to the source and refreshes the page. The tool includes CLI commands for review sessions, polling, status, and setup, plus a local server, editing and feedback SDK, browser client, Markdown renderer, and instructions for agent harnesses such as Claude Code and Codex. It runs on the user's computer without an account, cloud service, database, or API key.
ASD-STE100 Skill is a Claude Code skill that rewrites dense or ambiguous English into ASD-STE100 Simplified Technical English for AI-agent outputs, including tool descriptions, error messages, READMEs, and inter-agent instructions. It offers Strict mode for procedures, error messages, and tool descriptions, and STE-flavored mode for explanatory prose. The skill reads the input for meaning, flags sentence-level issues such as ambiguous wording, complex tense, unclear passive voice, multiple instructions, long sentences, noun clusters, phrasal verbs, nominalizations, semicolons, hedge stacks, and marketing adjectives, then rewrites flagged sentences while retaining facts, conditions, and scope qualifiers. It can also provide a before-and-after table naming the rules that each sentence breaks.
An open-source, executable playbook of agent skills for U.S. utility-patent work. It provides SKILL.md instructions, checklists, intake questionnaires, examiner and adversary protocols, and deterministic Python checks for dates, claim counts, and fees. Its three main workflows are a pre-filing audit, a simulated USPTO prosecution examination loop, and an adversarial design-around test. The repository grounds its simulations in 35 U.S.C., 37 CFR, the MPEP, and current USPTO materials that the agent is instructed to refresh, but it does not file applications, contact the USPTO, provide legal advice, or perform acts reserved for registered practitioners. It is licensed under GPL-3.0.
Agent-Safe Pipeline is a Decionis TypeScript library and runnable reference implementation for routing AI-agent actions through an independent authorization boundary. It captures an agent’s proposed action, target, and parameters as an immutable intent, sends the intent to a Decionis policy gate for an ALLOW, ESCALATE, or BLOCK decision, and coordinates verified human approval followed by policy re-evaluation for escalated actions. Agents cannot authorize their own actions, access downstream privileged credentials, or select trusted handlers. SafeExecutor accepts only the captured intent and decision, uses a sealed ActionRegistry to map action names to trusted handlers and validate parameters, and consumes a single-use grant bound to the intent before calling an API. The repository includes examples for Shopify refunds, GitHub deployments, procurement, and governed MCP tools, plus canonical-hash conformance vectors and architecture, threat-model, and security-evidence documentation. It is a self-hosted reference implementation rather than a hosted authorization service; its documented safety boundary also depends on provider-side identity, least privilege, network isolation, and incident response. The demos require Node.js 22.14 or later and pnpm 9.
MCP-Memory is a Model Context Protocol server that gives AI coding agents persistent memory across chat turns and sessions. It stores memory records as human-readable Markdown documents following the Open Knowledge Format (OKF v0.2), with YAML frontmatter, hierarchical index files, and an update log, while maintaining a local SQLite FTS5 index for key lookups, keyword search, tag filtering, and namespace-scoped retrieval. The server exposes MCP tools for storing, retrieving, searching, and deleting memories, and supports project, user, and default namespaces; a setup wizard configures supported MCP clients including Claude Desktop, Cursor, Antigravity, Windsurf, and Codex.
An experimental community preset for DeepSeek Harness, an LLM-agent harness. It starts a session with the real Minimal tool schema containing only a shell tool and a file-reading tool, using that condition to anchor the model's initial trajectory. After the first durable tool call or assistant reply, it promotes the session to the full Standard tool catalog, making heavier tools available on demand. The repository also includes related zero-anchored, self-introduction, prefab, and eternal-Minimal modes, and states that it is not an official DeepSeek preset or affiliated with DeepSeek.
Multi-Agent Workbench is a local-first browser control room for orchestrating coding agents such as Claude Code, Codex CLI, and generic PTY-based command-line agents. Agents work in task rooms under editable YAML role cards that define permissions, speaking style, and decision boundaries; a rule-driven orchestrator pauses execution and risky commands for human approval, while persistent sessions expose terminal activity and status. The system uses a local event-driven architecture with a React interface, Fastify backend, REST API, server-sent events, agent adapters, an approval service, and a SQLite event store. Every action passes through the API and is recorded in an append-only event log, whose events are projected into timelines, decisions, terminal views, artifacts, and diff views. It also provides task templates, Markdown reports and decision logs, real-time room updates, and configurable risk resolutions. Code-writing agents are intended to use isolated Git worktrees, with diff and patch review before changes reach the main workspace; the repository marks this worktree and merge flow as in progress.
Graft is an open-source codebase-context tool for coding agents, developed by NanoNets. It builds a repository graph and stores the result as linked Markdown files, with one node for each system, API, or concept and links describing how parts of the codebase connect. The graph supplies targeted context to agents instead of requiring them to rediscover repository structure during every task. It provides a CLI and MCP server, with integrations for Claude Code, Cursor, Codex, Gemini, and other coding agents. The CLI supports repository initialization, graph building, targeted search and orientation through commands such as `graft grep` and `graft map`, and graph visualization. `graft init` can install the agent wiring and background rebuild hooks; the generated graph is a local, regenerable cache that is added to `.gitignore` by default.
mcptoon is a cross-agent Model Context Protocol (MCP) management CLI that gives agents access to configured MCP servers through ordinary command-line calls, rather than loading large tool catalogs into their context. It discovers existing server configurations, maintains ~/.mcptoon/config.json as a single source of truth, synchronizes that configuration across agents, serves a tool-name manifest on demand, and composes a single-entry proxy for MCP access. Tools are launched when invoked, and the project supports native MCP JSON, structured output, optional TOON compression, and envelope passthrough; the videos also describe output checks for prompt injection and credential leaks. The project is distributed as a Python package for Python 3.10+ and is described as using only the standard library, with no third-party dependencies. Its repository documents compatibility with agents including Claude Code, Cursor, and Codex, and provides commands such as `quickstart`, `demo`, and `call`. It is free and open source under the Apache-2.0 license.
An open-source agent skill and standard-library Python service for stripping AI provenance marks from text and files owned by the user. The skill is a thin HTTP client that invokes a separately running service, so the agent host does not need Python. Deterministic processing handles invisible Unicode characters, exotic spaces, bidirectional and tag characters, while an agent rewrite pass and optional rewrite_text.py hook address statistical token-sampling watermarks. File processing targets C2PA, EXIF, XMP, and document properties across image, document, ebook, web, video, and audio formats, including PNG, JPEG, WebP, AVIF, HEIC, BMP, GIF, TIFF, SVG, PDF, DOCX, XLSX, PPTX, EPUB, ODT, HTML, Markdown, MP4, MOV, M4A, M4V, WAV, MP3, and FLAC. The repository describes support for class-level provenance ecosystems including Claude, Gemini/SynthID-Text, OpenAI provenance surfaces, Kirchenbauer-style green-list marks, and keyed-Gumbel/EXP marks. It includes installable skills for Claude Code, Cowork, Cursor, and other supported hosts; the installer uses Python 3.10+ with no external dependencies.
Agentic Inbox is a self-hosted email client from Cloudflare that runs on Cloudflare Workers through the user's Cloudflare account. It receives and sends mail through Cloudflare Email Routing, isolates each mailbox in a Durable Object with SQLite storage, and stores attachments in R2. The web interface supports threaded replies and forwards, rich-text composition, folders, search, and attachments. Its built-in AI agent, using the Cloudflare Agents SDK and Workers AI, can read inboxes, search conversations, draft replies, and send messages through email tools; automatic drafts for incoming mail require explicit confirmation before sending. Deployment provisions Workers, Durable Objects, R2, and Workers AI, and the setup requires Cloudflare Access, Email Routing, Email Service, and a mailbox on the configured domain.
Langfuse is an open-source LLM engineering platform for developing, monitoring, evaluating, and debugging AI applications. It ingests traces for LLM calls and related operations such as retrieval, embeddings, and agent actions, with integrations including OpenTelemetry, LangChain, the OpenAI SDK, LiteLLM, and other LLM frameworks. Its features include prompt management with version control and collaborative iteration, LLM-as-a-judge and code-based evaluations, user feedback and manual labeling, datasets for test sets and benchmarks, trace replay and debugging, metrics, and a playground for testing prompts and model configurations. Langfuse exposes APIs and typed Python and JavaScript/TypeScript SDKs for custom LLMOps workflows. It is available as a managed cloud service or for self-hosting with Docker Compose, virtual machines, Kubernetes via Helm, or Terraform templates.
fak is a self-hosted agent kernel distributed as a single Go binary. It wraps existing coding-agent sessions without replacing their interface or model, managing shared context, provider cache continuity, session recovery, model routing, tools, policies, and execution receipts. Tool calls pass through explicit policies that can allow, deny, transform, or witness them. Its default-deny capability floor blocks tools outside the configured allow-list, while the system can forward an existing provider subscription credential, route tasks to a configured provider or its native inference engine, and avoid silent fallbacks. The repository also provides an offline proof mode and installation through a shell script or Go.
Superlog is an open-core, self-hosted observability workspace for OpenTelemetry data, developed by Superlog Labs. It ingests traces, logs, and metrics, groups noisy signals into incidents, and provides a local-first interface for investigating production systems. The open-source community edition includes a web app and API, an OTLP intake proxy, worker processes for incident grouping and background jobs, a Postgres schema, ClickHouse-backed telemetry queries, and pluggable agent-runner interfaces with a default runner that records a local incident summary. The repository is licensed under Apache License 2.0; a hosted Superlog Cloud edition is also available with a free tier, pay-as-you-go plan, and monthly credit packs.
Renfield is a self-hosted, voice-first AI household assistant built with FastAPI, React, and Ollama. It uses Raspberry Pi Zero 2 W voice satellites with ReSpeaker HATs for local wake-word detection, Whisper speech recognition, Piper text-to-speech, and SpeechBrain speaker recognition. The system can track room presence through BLE scanning, voice recognition, and web authentication, and integrates with Home Assistant and DLNA for multi-room control. Its ReAct agent chains tools through 10 MCP servers covering weather, web search, news, calendars, Jellyfin, DLNA, n8n workflows, Home Assistant, Paperless, email, and related services. It also provides conversational long-term memory with contradiction detection, a knowledge graph with validated entity-relation triples, and a RAG knowledge base that combines dense pgvector embeddings with BM25 full-text search using reciprocal-rank fusion. Documents can be uploaded from formats including PDF, DOCX, PPTX, XLSX, HTML, Markdown, and TXT, with Tesseract or EasyOCR available for scanned or poor-quality PDFs. The project is designed to run entirely on local hardware without cloud dependencies and includes a web interface, satellite monitoring, proactive notifications, plugin hooks, and Paperless document-audit workflows. The repository describes more than 100 available tools and supports Docker-based self-hosting.
agentacct is a local-first audit and observability tool for coding-agent work. It reads session logs from Claude Code, Codex, OpenCode, and Hermes, combines them with recorded tasks and test evidence, and produces a Work Receipt for each task covering commands run, files touched, tools used, time and token usage, cost, outcome, and evidence strength. Agent claims and machine verification are kept as separate axes, with verified status reserved for tasks whose live checks pass after the latest recorded work. The data is available through a macOS app, the live terminal dashboard launched with `agentacct tui`, and a local JSON API. State is stored in local files; the API listens only on loopback, and the project states that it has no login, cloud sync, or telemetry and does not request provider API keys.
AI Copywriter is a portable, plain-Markdown agent skill for writing marketing copy and removing common patterns associated with AI-generated prose. It supports headlines, short descriptions, microcopy, subject lines, button labels, and LinkedIn posts, and runs in any harness that supports skill-style instructions. Before drafting, it interviews the user about the intended customer, product category, and underlying story, then pressure-tests the story for specific and interesting details. Its copywriting method focuses on the reader's immediate situation and explains the product in simple language. It incorporates the 33 detectable and fixable writing patterns from blader's Humanizer while adding the reverse workflow of generating marketing copy designed around those constraints.
Better Harness is an open-source Harness Engineering platform for coding agents. It analyzes project and, where supported, session evidence around an agent's work rather than only its final code diff, evaluating the Agent Work Loop for issues such as task understanding, validation, delivery control, and retained lessons. It turns supported gaps into prioritized findings linked to their evidence, expected outcomes, repair boundaries, and acceptance checks, and produces host-specific reports in formats including HTML, paired Markdown, or native Canvas reports. The platform also supports defining harnesses as code, running controlled experiments, inspecting evidence, and comparing outcomes across coding-agent hosts such as Claude Code, Codex, Qoder, Cursor, and GitHub Copilot CLI.
opentax-engine is an open-source deterministic US tax calculator designed for AI agents and other applications. It accepts tax facts such as wages, filing status, and dependents, applies an encoded rule corpus, and returns cent-exact results with assumptions, legal citations, and a proof tree showing how each value was derived. If its rules cannot derive an answer, the project is designed to refuse rather than guess. The calculation can be selected for a date with an --as-of option, and the CLI accepts flags or JSON fact files; the videos also identify CLI, browser, and MCP integrations. The engine is implemented in TypeScript with no platform dependencies and can run in a self-contained browser HTML file containing the engine, rule corpus, and verifier. Proofs use canonical JSON, hashing, and a Merkle construction; the verifier independently re-derives the calculation and reports altered proof data, rule-corpus differences, or steps that do not match. The repository describes the project as AGPL-3.0 licensed, with commercial licenses available.
OptMem is a permanent-memory tool for AI agents. It installs as a dependency-free Python 3 script and integrates through a prompt block added to an agent's AGENTS.md or CLAUDE.md file. The agent uses `memo wake` at session startup, `memo note` to append one-line memories, `memo recall` for exact regular-expression searches, and `memo zoom` to navigate summaries. Raw memories are stored in an append-only `LOG.txt`; a binary tree of one-line summaries provides a compact reading view, with summaries rebuilt when needed. `memo nap` processes due merges, while `memo forget` discards a summary so the next nap can rebuild it. Memory can be kept in a configurable directory such as a synced folder or Git repository, and the system is designed to persist across sessions, context compaction, models, and vendors.
Deltafin is a native binary for running the full, unpruned Kimi K3 mixture-of-experts model on consumer hardware, along with an OpenAI-compatible API server for local chat and coding agents. Its demonstration approach keeps the model's attention spine on local disk and fetches routed experts from Hugging Face's CDN as needed, allowing the full model to run on a 64-gigabyte MacBook rather than loading the entire expert bank into memory. The project emphasizes preserving Kimi K3's original expert weights and having K3 verify every token; it is presented as an experiment in pushing large-model inference on local hardware rather than as a practical chat setup.
OpenWorker is an open-source desktop AI coworker that produces finished deliverables from everyday tasks, including code-security reviews with proposed fixes, cloud-configuration audits, incident reports, documents, spreadsheets, drafted messages, and scheduled briefs. It runs on the user's machine through a native desktop app and a local Python agent server built on aisuite, working across local files, the terminal, connected applications, and more than 25 connectors. Specialist coworkers cover security review, cloud posture, incident triage, everyday work, and recurring automations; security review combines deterministic scanners such as Semgrep with model reasoning, then re-scans and diff-reviews proposed fixes before approval. The agent breaks requested outcomes into steps and uses models from providers such as OpenAI, Anthropic, and Google, open-weight providers, or a local Ollama deployment, with the user supplying model access. Consequential actions—including sending messages, changing calendars, and running commands—are approval-gated. Its governance design includes human-only floors for dangerous or irreversible operations, explicitly granted and revocable autonomy rules, circuit-breaker escalation for uncertain or repeatedly denied actions, and an audit trail recording tool calls, approval provenance, and reviewer reasoning; unattended runs cannot self-approve. The project is in open beta and provides downloads for macOS on Apple Silicon and Windows on x64.
scroll-world is an agent skill for Claude Code, Codex, and other SKILL.md-compatible agents that builds scroll-scrubbed 3D world landing pages for brands and industries. It generates cohesive isometric diorama scenes and image-to-video camera flights, linking scenes through first/last-frame conditioning so the camera appears to move continuously as the visitor scrolls. The skill provides prompt templates, an AI image and video rendering pipeline using Higgsfield, Monid, Seedance, Kling, or Codex image generation, and a framework-agnostic vanilla JavaScript scrubbing engine that can be used with plain HTML, Next.js, Vue, or a Python-served page. It can also render a separate portrait chain for mobile devices and requires external rendering services or CLIs, ffmpeg/ffprobe, and Python with Pillow.
No AI Slop is an AI-writing editing skill that removes more than 20 canned machine-writing patterns while preserving the author's vocabulary, cadence, humor, and imperfections. Its rules target binary contrasts, throat-clearing openers, faux-insight setups, colon reveals, dramatic fragments, superficial analysis, importance puffery, weasel attribution, synonym cycling, and fake-profound endings; they also cover active voice, leading with the point, untangling difficult sentences, and preferring concrete details over abstractions. The skill supports editing text, detecting and quoting suspected patterns without judging whether AI produced the writing, and generating deliberately exaggerated AI writing for satire. It can be invoked as `/no-ai-slop` in ChatGPT, Claude Code, Codex, or another coding agent, and is distributed through the Skills package manager or an npx command. The repository contains the editing rules in SKILL.md, evaluation checks in eval.md, ChatGPT and Codex plugin metadata, and a plugin build script; it is licensed under MIT.
OpenScience is an open-source AI workbench for scientific research developed by Synthetic Sciences. It runs as a browser-based workspace where a single research agent takes a goal through literature review, hypothesis formation, code writing and execution, experiments, analysis, and a written report. The agent can load domain skills, delegate bounded exploratory or execution work, query scientific databases, and preserve sessions as observable research traces. The local server hosts the workspace UI, agent runtime, skill library, and tool layer. The agent plans with a research harness and calls shell, editor, LSP, MCP, scientific database, and skill tools; sessions, skills, artifacts, and provenance are stored on disk. The workbench includes a file tree, code editor, terminal, session history, and inline rendering for molecules, structures, genomes, and plots. Its bundled skills cover training, evaluation, datasets, molecular and clinical biology, cheminformatics, papers and LaTeX, figures, and scientific runtimes, while database connectors provide direct access to services including UniProt, PDB, Ensembl, ChEMBL, PubChem, arXiv, OpenAlex, and Semantic Scholar. It supports frontier and open-weight models through user-provided keys, eligible ChatGPT or Codex access, local models, and optional Ace-managed models. Models are routed per request, allowing providers to be switched without changing the workspace. Extensibility includes LSP integration, MCP servers, plugins, custom agents and commands, experimental Python environments and BioNeMo adapters, and a TypeScript SDK. It is installed with the @synsci/openscience npm package or launched through npx; platform binaries and desktop installers are distributed through GitHub Releases. A free Synthetic Sciences account links an installation and can provide credential synchronization, private research graphs, enhanced search, and optional credit-backed models, while BYOK and local-model usage remain separate from Ace.
The Fable Method is an agent-workflow repository that distills the Fable Workflow into skills that models can run. Its process classifies the request, defines completion with named verification, gathers evidence from primary sources in parallel, commits to one recommendation, changes the smallest correct thing, verifies the result by observation, and reports the outcome with caveats. The repository provides four related skills—fable-method for thinking, fable-loop for acting, fable-judge for evaluating results, and fable-domain for generating domain adapters—along with evaluation cases, transcripts, judge outputs, and logs. The project reports testing the method across fifteen evaluation rounds and more than 260 agent runs, with judges checking diffs and execution rather than relying on agent reports.
thinking-orbs is a React component package providing dotted animated loading indicators for AI and agent interfaces. It offers nine hand-tuned states— including working, searching, solving, listening, connecting, weaving, composing, breathing, and shaping—rendered with plain 2D canvas rather than WebGL or filters. The indicators have separate 64-pixel chat-avatar and 20-pixel inline-text designs, automatic or pinned light/dark themes, adjustable speed, pause control, and pass-through canvas properties. It includes per-state accessible labels, renders a static frame for prefers-reduced-motion, pauses offscreen instances through IntersectionObserver or when the tab is hidden, and resumes them in phase using a shared clock. The package is distributed through npm under the MIT license.
diri is a native desktop orchestrator for coding agents on macOS and Linux. It runs Claude Code, Codex, Cursor, Gemini, and ordinary shells in parallel, using separate Git worktrees or remote hosts over SSH and tmux. Each session is a real terminal with a persistent PTY managed by a background daemon, so closing or reopening the app does not terminate sessions; the daemon restores their output and state. The app displays working, needs-you, and done status, while its MCP server allows one running agent to spawn, monitor, read, and respond to another. Claude Code and Codex have first-class status detection and resume support, while Cursor and Gemini have partial support.
Breadcromb is the developer of Trace, an AI-powered web browser that builds a private knowledge layer from users' reading and uses local-first AI to answer questions, automate tasks, and orchestrate agents within the browser. Trace is offered as a free download and emphasizes privacy and local-first operation.
Strix is an open-source AI penetration-testing tool from usestrix that uses autonomous, multi-agent AI pentesters to dynamically run applications, perform reconnaissance and exploitation, and validate vulnerabilities with working proof-of-concept exploits rather than static-analysis findings. Its developer-oriented CLI produces remediation guidance, generated patches, and pentest reports, and it can run scans in CI/CD pipelines to detect insecure changes before deployment. The tool runs in a Docker sandbox and requires an LLM API key; the associated Strix platform supports repository and domain testing, continuous scanning, and integrations with development and issue-tracking workflows.
npcpy is a Python library for research and development with multimodal language models, agentic AI, and knowledge graphs. It provides primitives for defining personas, making direct language-model calls, creating tool-using agents and multi-agent teams, and building AI applications with local providers such as Ollama, llama.cpp, omlx, and LM Studio as well as cloud providers. Its NPC Context-Agent-Tool data layer is designed to enforce context and tool-use rules through software rather than prompts. The Agent class includes tools such as shell execution, Python, file editing, and web search, while ToolAgent supports custom tools; the examples include image generation, Hugging Face image-dataset retrieval, and diffusion-model fine-tuning. The library is distributed through PyPI.