704 tools and products — trending open source, and what gets used in AI and other work.
CommerceAgentBench is a benchmark developed and maintained by the Accio team at Alibaba International for evaluating whether AI agents can complete long-horizon business workflows in high-fidelity, stateful, reproducible replicas of online services. Its tasks cover browser operations, native-style command-line tools, file and spreadsheet production, API and MCP workflows, public-web research, supplier analysis, product publishing, logistics, storefront configuration, and other commerce operations. Each task runs in a fresh container and is graded by a deterministic or LLM-assisted verifier. The benchmark uses local mock services representing commerce, SaaS, messaging, document, and operational systems; runs preserve the resolved configuration, agent trajectory, verifier result, artifacts, logs, and container metadata for auditing and reproducibility. The repository describes 107 tasks spanning CLI, browser, file, and API/MCP interfaces, with text-only, browser-text-capable, and vision-required capability slices. The project was previously called RealReplicaBench.
diri is a native desktop orchestrator for coding agents on macOS and Linux. It runs Claude Code, Codex, Cursor, Gemini, and ordinary shells in parallel, using separate Git worktrees or remote hosts over SSH and tmux. Each session is a real terminal with a persistent PTY managed by a background daemon, so closing or reopening the app does not terminate sessions; the daemon restores their output and state. The app displays working, needs-you, and done status, while its MCP server allows one running agent to spawn, monitor, read, and respond to another. Claude Code and Codex have first-class status detection and resume support, while Cursor and Gemini have partial support.
Breadcromb is the developer of Trace, an AI-powered web browser that builds a private knowledge layer from users' reading and uses local-first AI to answer questions, automate tasks, and orchestrate agents within the browser. Trace is offered as a free download and emphasizes privacy and local-first operation.
Strix is an open-source AI penetration-testing tool from usestrix that uses autonomous, multi-agent AI pentesters to dynamically run applications, perform reconnaissance and exploitation, and validate vulnerabilities with working proof-of-concept exploits rather than static-analysis findings. Its developer-oriented CLI produces remediation guidance, generated patches, and pentest reports, and it can run scans in CI/CD pipelines to detect insecure changes before deployment. The tool runs in a Docker sandbox and requires an LLM API key; the associated Strix platform supports repository and domain testing, continuous scanning, and integrations with development and issue-tracking workflows.
npcpy is a Python library for research and development with multimodal language models, agentic AI, and knowledge graphs. It provides primitives for defining personas, making direct language-model calls, creating tool-using agents and multi-agent teams, and building AI applications with local providers such as Ollama, llama.cpp, omlx, and LM Studio as well as cloud providers. Its NPC Context-Agent-Tool data layer is designed to enforce context and tool-use rules through software rather than prompts. The Agent class includes tools such as shell execution, Python, file editing, and web search, while ToolAgent supports custom tools; the examples include image generation, Hugging Face image-dataset retrieval, and diffusion-model fine-tuning. The library is distributed through PyPI.
USBridge Remote is a unified client for managing remote machines, combining software-based remote desktop control with integration for USBridge KVM hardware. Its client runs on Windows, macOS, Linux, Android, iOS, and the web, while an agent on the target machine handles screen capture, input injection, and Tailscale networking. The system provides live remote desktop, virtual device passthrough, snapshot management, shared clipboard transfers, encrypted peer-to-peer connections, and Moonlight-based low-latency streaming. It is distributed as beta software, is free without session or connection limits, and does not require an account on the target machine; the web client has feature and performance limitations compared with native applications.
PortalJS is an open-source, AI-native framework maintained by Datopian and its community for building data portals. Its agentic skills advise on storage, compute, catalog, access, hosting, and metadata, then scaffold the result as plain, editable Next.js code. Claude Code skills can create a portal, add CSV or JSON datasets, connect to backends such as CKAN, and deploy it; generated portals include a home page, catalog, and dataset showcase, with support for configurable data providers, charts, maps, and schemas. A bare template is also available without the AI skills, and the project presents Git, object storage, Parquet, and DuckDB as an optional modern architecture while retaining traditional datastores as supported choices.
Kandev is an open-source, self-hostable AI Kanban and development environment for orchestrating coding agents and reviewing their work. It organizes tasks in Kanban and pipeline views, supports parallel execution and multi-step agentic workflows, and isolates concurrent work with Git worktrees. Its integrated workspace combines a file editor, file tree, terminal, browser preview, chat, and Git changes for reviewing and iterating on agent output. Agents can run through local, Docker, SSH, or cloud runtimes, with support for multiple agent providers; the project also provides configurable workflows, agent profiles, prompts, review gates, scheduled or webhook-triggered automations, and no telemetry.
deja-vu is a local memory and search layer for AI coding agents. Its Go binary indexes session histories already stored on a machine by Claude Code, Codex, Cursor, and other supported agents, including sessions recorded before installation, without using an LLM or embeddings. It provides natural-language and direct search, session inspection, context retrieval, JSON output, redaction, usage statistics, and MCP-based cross-agent recall; the index strips keys and tokens as it is built and can synchronize append-only memory between machines over SSH. Recall can be injected automatically at session start and before prompts, file edits, or commands, and a post-command hook can retrieve what followed a matching failure when the agent supports those hooks. It also indexes files opened during each turn, executed commands and exit statuses, and exact spans replaced by edits, rather than only conversation text. Decisions can be marked as accepted or rejected with notes, and results can report when indexed files have changed since a session. The repository reports sub-millisecond lookups over 5 GB of history, with benchmark harnesses for LongMemEval-S and LoCoMo. The project is distributed as a local binary with installation options including a shell script, Homebrew, Go, npm, Scoop, release archives, and agent-specific plugin or MCP bundles. Its optional installation wiring configures supported agents for MCP and session-start recall.
The GitHub Copilot SDK is a multi-platform SDK from GitHub for embedding the agent runtime behind Copilot CLI in applications and services. It provides SDKs for Node.js/TypeScript, Python, Go, .NET, Java, and Rust, allowing applications to invoke Copilot agent workflows while the runtime handles planning, tool invocation, file edits, and related orchestration. Applications can also define custom agents, skills, and tools. Each SDK communicates with Copilot CLI running as a server over JSON-RPC. The SDK manages the CLI process lifecycle automatically and can connect to an externally run CLI server. The repository provides package-specific installation and API instructions; the CLI is bundled automatically for Node.js, Python, and .NET, while Go, Java, and Rust generally require a separately available CLI, with application-level bundling features available for Go and Rust. Standard use requires a GitHub Copilot subscription, and usage follows the Copilot CLI billing model, with prompts counted toward the applicable allowance. A BYOK mode supports API keys from supported providers without GitHub authentication.
Chatwoot is an open-source, self-hosted customer support platform and omnichannel support desk. It centralizes conversations from website live chat, email, Facebook, Instagram, Twitter, WhatsApp, Telegram, Line, SMS, and other channels in a shared inbox, with tools for assignment, automation, labels, canned responses, teams, internal notes, contact management, segmentation, and reporting. It includes a Help Center Portal for publishing help articles, FAQs, and guides, and Captain, its AI support agent, which automates responses and handles common queries. The platform also provides integrations including Slack, Dialogflow, Shopify, Google Translate, and Linear, and is distributed as an open-source, self-hostable alternative to Intercom, Zendesk, and Salesforce Service Cloud.
Speech To Speech is an open-source, modular voice-agent pipeline from Hugging Face. It chains voice activity detection, speech-to-text, a language model, and text-to-speech as separate stages connected by queues; each stage has interchangeable backends. The language-model stage supports OpenAI-compatible protocols and can use hosted providers, Hugging Face Inference Providers, vLLM, or llama.cpp, while the system streams transcripts, generated text, tool calls, and synthesized audio. The project exposes the core OpenAI Realtime event set through WebSocket and WebRTC and includes commands for running a realtime server, a microphone-and-speaker client, or a local setup. It is distributed as a Python package and requires Python 3.10 or newer. The repository says the pipeline is used as the conversation backend for Reachy Mini robots.
Bolcho AI is an Indian company offering a voice and chat AI platform that enables organizations to build, launch, and scale conversational agents for phone calls and websites in Indian languages. The service markets itself as operating at near provider cost and targets Bharat-language use cases.
An open-source AI software-development factory provided as a Vercel template and implemented by Foreman. It takes tasks from GitHub issues, GitHub mentions, Linear Agent Sessions, or a local developer TUI and moves them through four isolated agent stations: a Classifier triages task type, priority, complexity, and actionability; an Analyst creates a plan and acceptance criteria from a live repository checkout; an Implementer executes the plan in a sandbox, runs the repository's checks, and pushes a branch; and a Reviewer independently evaluates the resulting diff with evidence. The pipeline produces a reviewed draft pull request for human approval, after which people mark it ready and merge it. Foreman maintains shared repository notes between runs, can diagnose failing CI on its own factory pull requests, and posts progress to the originating GitHub or Linear task. The deployment flow configures GitHub and Linear connectors, Vercel Blob storage, a target repository, and an issue label.
Embabel is an open-source framework for authoring agentic flows on the JVM. Written in Kotlin with a natural Java usage model and built on Spring, it combines LLM-prompted interactions with manually written code and strongly typed domain models. Flows are expressed through actions, goals, conditions, and domain objects; they can be authored with Spring-style annotations such as @Agent, @Goal, @Condition, and @Action, or with an idiomatic Kotlin DSL. The framework dynamically formulates plans toward goals using a pluggable, non-LLM planning step. Its default approach is Goal Oriented Action Planning (GOAP), while Utility AI is also supported for selecting actions by utility scores rather than strict preconditions and postconditions. Conditions are reassessed after each action, and the system replans based on new information and observed effects, allowing known actions to be combined in novel orders and enabling runtime decisions such as parallelization. The planning model is independent of the actions, goals, and plans themselves, and can extend application capabilities by adding domain objects, actions, goals, and conditions without changing finite-state-machine definitions. Execution is provided through an AgentPlatform with focused, closed, and open modes. Focused mode runs a requested agent from application code; closed mode classifies user intent or an incoming event to select an agent; and open mode searches across available goals and actions to build a custom agent, while allowing developers to restrict goal selection. Open mode can find paths not explicitly envisioned by developers and combine functionality from multiple providers, but individual steps must still be defined by the application. Embabel supports mixing language models, including local models for selected tasks, separates the programming model from platform internals for local or production execution, and integrates with Spring dependency injection, AOP, persistence, transactions, and testing. The framework is designed for unit and end-to-end agent testing.
Decagon is an enterprise AI customer-support platform with autonomous agents for handling customer inquiries and operational workflows across support channels. Its agents resolve issues using connected business systems, escalate cases to human support staff when needed, and provide business-process logic, integrations, testing, monitoring, compliance support, and deployment capabilities.
A feature of Anthropic's Claude platform that packages reusable instructions, scripts, and reference files into skills that Claude can apply to specific tasks. Skills can be used in Claude applications and through Anthropic's agent-development tools.
camelAI is an AI coding assistant platform built on Cloudflare Workers and Durable Objects. Each chat thread runs a coding agent in its own persistent workspace, retaining chat state and project files while providing workspace-aware tools. The platform connects to APIs, databases, email, Slack, Discord, and other services; supports application builds, notebook analysis, and SQL execution in isolated short-lived sandboxes; and provides previews and publishing through Workers for Platforms. Its agent harness is built on pi's lower-level agent-loop and state-management libraries. The agent uses native file tools and writes JavaScript rather than shell commands; Code Mode executes that JavaScript in fresh V8 isolates with explicit platform and connection methods. Project files are stored in WorkspaceFilesystem Durable Objects, using Durable Object SQLite for small files and R2 for larger ones, while Cloudflare Artifacts provides Git history. Linux sandbox containers are reserved for jobs requiring them, including builds, notebook analysis, and database queries, and credentials remain outside the execution sandbox. The platform can use Anthropic, OpenAI, OpenRouter, Bedrock, and custom model endpoints. Its repository is released under the MIT License and includes local development setup using Node.js, Bun, and a Cloudflare account.
LingBot-World 2.0, also called LingBot-World-Infinity, is an AI world-modeling system from the Robbyant Team for generating interactive video environments. Its causal pretraining paradigm is designed for an unbounded interaction horizon, while a distilled real-time variant is claimed to drive 720p video streams at 60 frames per second. The system supports diverse actions, including attacking, archery, spell-casting, and shooting, along with text-driven events. Its agentic harness uses a pilot agent to plan and execute character behavior and a director agent to synthesize new environmental elements as a scene progresses. The causal inference implementation processes video frames chunk by chunk with KV caching rather than processing the entire sequence at once. The release provides inference code and model weights, including a 14B causal-fast model, with a codebase built on Wan2.2; models are distributed through Hugging Face and ModelScope, and the real-time system can be tried through Reactor on the web or LingGuang on mobile. The project is licensed under CC BY-NC-SA 4.0 for non-commercial use. It states that deployment code will not be released and points to SGLang or flashdreams for self-deployment references.
DrDroid is an AI SRE agent for incident response and root-cause analysis. It connects to cloud platforms, source code and telemetry to scan stacks, build a knowledge graph, and assist with incident response, root-cause analysis, and automated remediation.
Pushary is a service for approving or authorizing AI agents from a mobile device when they pause to request permission. It supports agent frameworks and models including Claude Code, Codex, Cursor, and other MCP-compatible agents, allowing users to provide one-tap approvals so the agents can continue operating.
AgentOps is an open-source observability and developer-tools platform for building, evaluating, and monitoring AI agents. Its Python SDK records agent sessions and LLM calls for dashboard-based session replays, step-by-step execution graphs, debugging, analytics, cost tracking, and benchmarking. Developers instrument code with decorators for sessions, agents, operations, tasks, and workflows; these support nested span hierarchies, input and output recording, exception handling, asynchronous functions, generators, and custom attributes. The SDK provides integrations with agent frameworks and LLM systems including CrewAI, AG2/AutoGen, Agno, LangGraph, and the OpenAI Agents SDK, and the platform can be self-hosted with its dashboard and API backend. The app is released under the MIT license.
TinyFish is a company and platform that provides enterprise web infrastructure for AI agents. It offers a fully serverless architecture to run hundreds of websites concurrently for tasks such as navigation, authentication, data extraction, and automating web workflows.
HarnessRouter Community Edition is the open-source edition of HarnessRouter, an agent-harness infrastructure that provides a unified Agent API to integrate language models and agent runtimes. It is maintained by HarnessRouter and is intended for deployment on users' own infrastructure to connect models such as Codex, Claude Code, and Hermes.
Chert is a company offering AI video agents that can programmatically answer and place FaceTime calls; the product is presented as "Vapi for FaceTime." It provides developer-facing tools and APIs to embed video-call agent behavior in applications.
Zetik is a web service that describes itself as a "personal intelligence agent" which continuously monitors internet sources—news, podcasts, papers, and code—and delivers briefings via feed, push, newsletter, or RSS. It is offered under the Zetik brand and positions itself as a tool for keeping users informed about developments it deems relevant.
Codeman is a self-hosted web dashboard for running and monitoring AI coding agents. It launches Claude Code, OpenCode, Codex, Antigravity, Gemini, Pi, or Grok in persistent tmux sessions, with an optional plain shell, and streams their terminals to a browser. Sessions persist across restarts and network drops; idle detection can re-prompt agents, resume them after usage limits reset, and run scheduled jobs, while live floating windows show background agents and subagents. The dashboard is designed for touch devices and includes local echo, QR login, swipe navigation, and push notifications. It can run locally, in Docker, or over SSH, is MIT licensed, runs on the user's machine by default, and states that it has no telemetry.
OpenGeni is an open-source, self-hostable runtime for long-running AI agents rather than an agent itself. It exposes a session-based API for creating, steering, observing, interrupting, and replaying agent runs, storing every event in a PostgreSQL log and streaming live history over SSE so sessions can be resumed or audited. Sessions can run in managed sandboxes or on enrolled Connected Machines such as laptops, build servers, and GPU systems; machine connections are outbound-only and can be revoked. The runtime includes human approval gates for tool use and proposed curated memories, per-session credential brokering, durable pauses for structured responses, and an audit trail. Its control plane, sessions API, web app, and deployment artifacts are distributed under the Apache-2.0 license.
Magic Cloud is an open-source, self-hosted full-stack application builder and AI-agent platform. Its API Wizard connects to databases such as SQLite, PostgreSQL, MySQL, MariaDB, and SQL Server and generates secured REST APIs with authentication, role-based access control, CRUD operations, and business logic; endpoints can then be extended from plain-English descriptions without a separate build or deployment step. It can also generate database-backed applications, wrap APIs described by OpenAPI or Swagger specifications, run scheduled Hyperlambda jobs, and expose endpoints as MCP tools for clients such as Claude, Cursor, and Codex. The runtime uses Hyperlambda, a compiled .NET execution environment. The project says its generator produces an abstract syntax tree rather than unrestricted function-call text, allowing generated invocations to be checked against available functions, while sandboxing and RBAC restrict execution. Magic Cloud is self-hostable, distributed under the MIT license, and is developed in the polterguy/magic repository.
LibreDesk is an open-source, self-hosted omnichannel customer support desk distributed as a single binary. It combines live-chat and email conversations in one inbox, provides an embeddable website chat widget, and supports searchable multilingual help-center content. Its AI assistant answers live-chat questions using the knowledge base and hands conversations to a human when it cannot help; an agent copilot drafts replies, summarizes conversations, and retrieves knowledge-base answers. The system also provides event-based automations, capacity- or criteria-based assignment, SLA tracking, custom roles and permissions, CSAT surveys, analytics, macros, custom attributes, SSO through Google, Microsoft, or OIDC, HTTP/JSON APIs, webhooks, activity logs, and conversation search. It can be deployed through Railway or Docker, with PostgreSQL and Redis provisioned or configured for the deployment.
NutriTrace is a self-hosted personal nutrition tracker developed by TraceApps. It runs as a single Docker container and provides a browser Progressive Web App and a native Android app for food diaries, meal and recipe logging, calorie and nutrition goals, adaptive TDEE tracking, intermittent fasting, manual activities, and wellness data. It can use Open Food Facts and USDA data, import recipes from Mealie, and integrate with Fitbit via Google Health, Withings, Garmin, and Health Connect. Its Trace AI features can read diary and wellness data, log food or activities conversationally, and propose food entries from photos for user confirmation. It also exposes selected diary, goal, and food data to external AI agents through an opt-in, read-only Model Context Protocol endpoint. The project supports multi-user access and OIDC single sign-on, provides migration tools for several nutrition trackers, and distributes its server as a Docker image. The source is licensed under AGPL-3.0; the project states that it has no telemetry or cloud synchronization unless a user opts into a third-party integration.
AI Design Skills is an MIT-licensed collection of Markdown design skills for AI coding tools such as Claude Code, Cursor, Codex, Windsurf, and Cline. Its skills provide coding agents with guidance for landing-page strategy, page structure, conversion copy, typography, spacing, visual systems, border radius, and motion. Each skill is stored as a self-contained folder with a SKILL.md file that can be copied into an agent's project-rules directory, added to a system prompt, or uploaded to a Claude.ai project. The repository currently documents a landing-page-design skill that asks intake questions, builds the page structure and copy, and defines the visual system; its design values are grouped in an editable section so users can change fonts, colors, spacing, and motion. The collection can be forked and used, modified, distributed, or used in commercial work under the MIT license.
arc-code is a minimal harness for testing whether a general coding agent can learn unfamiliar ARC-AGI-3 games during play. It runs a headless coding agent with a shell, filesystem, and action command inside an internet-isolated sandbox; the harness provides no solver, planner, world model, grid tooling, fine-tuning, MCP servers, subagents, plugins, or hooks. The agent interacts with each game only through `act.py`, which executes actions and appends resulting boards and actions to a log. It uses the shell and approved file tools to parse that log, record findings, and build task-specific machinery such as parsers, rule models, simulators, and search procedures, which are discarded when the game ends. `run.py` launches and records one session per game, while the `rig/` directory supplies sandboxing, brokering, auditing, scoring, and export infrastructure. The repository includes the harness and six complete session records.
VibePulse is an always-on ESP32-S3 AMOLED panel and local-first host service for monitoring Claude Code and Codex usage, live agent activity, and prompts that need user input. The panel displays quota information and agent status, raises a full-screen NEEDS YOU alert when an agent is waiting, and can send supported decisions back through a tap on the display. A pure-standard-library Python service runs on a macOS or Windows computer and communicates with the panel over the local network. The project supports optional, default-off relays: a numbers-only relay for quota data, an encrypted interaction relay for supported NEEDS YOU decisions, and an independent status relay for Claude/Codex activity. The repository describes the core panel loop, Windows host lifecycle, and physical prompt-answer flow as verified, with cloud features disabled by default.
Vibe ASO is a Claude Code skill that takes an iOS app from ready-to-submit status through App Store optimization and localization. It performs popularity- and difficulty-driven, intent-matched keyword research; generates App Store metadata in up to 50 locales; renders localized screenshots with Playwright, per-script Noto fonts, right-to-left handling, and automatic text fitting; applies territory-specific pricing through the App Store Connect API with read-back verification; and localizes in-app strings. Its checks detect translation problems such as format-specifier drift, wrong-sense domain terms, script contamination, and register drift. It sets and verifies App Store Connect fields where the API permits, producing an explicit manual-steps list for the remaining submission tasks. The skill has no separate SaaS signup: it uses the user's App Store Connect API key, Claude subagents by default, or a DeepSeek/OpenAI-compatible translation API, and can use ASO keyword data from an external tool such as Astro's MCP.
md2hd is a local CLI that turns a Markdown file or folder of notes into an interactive relationship map in a browser. YAML front matter becomes nodes, while wikilinks and relationship entries become directed edges; an optional map block defines the map's title, types, and relationship inverses. The map supports searching, type-based views, node-focused navigation, degree-based relationship expansion, dragging and saving positions, and editing the source Markdown through a live code pane. It runs a local server bound to 127.0.0.1, reads changes from disk on refresh, and is described as requiring no account or telemetry. The package also includes a writing-md2hd-maps skill for coding agents that author or convert Markdown maps.
RAGFlow is an open-source retrieval-augmented generation (RAG) engine from infiniflow that provides a context layer between documents and large language models, combining document understanding with agent capabilities. It ingests heterogeneous sources including Word, Slides, Excel, text files, images, scanned documents, structured data, and web pages; extracts knowledge from complex formats; and uses configurable template-based chunking, multiple retrieval methods, and fused re-ranking to assemble context for language-model applications. Its workflows provide traceable citations and visualized text chunks for human review, and it supports configurable language and embedding models, agent templates, orchestration, APIs, and self-hosting. A hosted cloud service is also available.
Unsloth is an open-source local desktop and web tool for running, training, fine-tuning, and deploying language, diffusion, embedding, audio, and vision models. It supports Windows, macOS, Linux, WSL, multi-GPU setups, NVIDIA, AMD, and Intel GPUs, CPUs, and the Vulkan backend. The project is distributed as Unsloth Desktop, Unsloth Studio, and Unsloth Core. It can run local models, provide an OpenAI-compatible API, connect models to agents and MCP tools, perform search and RAG, and offer remote or LAN access. Its training functions cover reinforcement learning, LoRA, QLoRA, full fine-tuning, pretraining, GRPO, DPO, and FP8, with export formats including GGUF, NVFP4, and FP8. The repository claims that its fine-tuning workflows use two times less training time and 70% less VRAM.
Paritok is an open-source compression gateway for AI coding agents. It runs as a drop-in proxy between an agent and an upstream LLM API, stripping irrelevant tool schemas, compressing tool results and file reads, and summarizing stale conversation history before forwarding requests. Compressed content remains recoverable on demand through the gateway, while responses pass back to the agent unchanged. It supports agents such as Claude Code, Cursor, Codex, and OpenHands, as well as agents that support BASE_URL configuration, and is powered by an open-source code-focused 4B compression model.
Freebuff is an ad-supported AI coding-agent suite from CodebuffAI. Its CLI runs in a project from the terminal, where specialized agents find relevant files, plan and implement changes, run project checks, and review the results. The suite also includes desktop, web, cloud, and chat products for running parallel local agents, building full-stack applications, operating on GitHub repositories, and conducting research. Freebuff provides a catalog of included AI models without a subscription, credits, or an API key; text ads support access to those models. Desktop can run concurrent agents in separate workspaces, while the web and cloud products provide hosted sandboxes, previews, terminals, browser-based research, and deployment workflows.
Harbor is a framework from the creators of Terminal-Bench for evaluating and optimizing AI agents and language models in container environments. It can run arbitrary agents, including Claude Code, OpenHands, and Codex CLI, against shared or custom benchmarks and environments such as Terminal-Bench, SWE-Bench, and Aider Polyglot. Harbor launches benchmark runs locally with Docker or in parallel through cloud and sandbox providers, and can generate rollouts for reinforcement-learning optimization. It is distributed as a Python package installable with uv or pip and provides command-line tools for running datasets, selecting agents and models, and listing supported benchmarks.
An AI tool that orchestrates agents to synthesize realistic reinforcement-learning environments and tasks from domain descriptions, real usage data, scenarios, and personas.
Firecracker is an open-source virtualization technology developed at Amazon Web Services for running secure, multi-tenant container and function workloads in lightweight virtual machines called microVMs. Its virtual machine monitor uses Linux KVM and a minimalist device set to combine hardware-virtualization isolation with container-like startup speed and resource efficiency. Firecracker exposes a host API for configuring and managing microVMs and is designed for serverless workloads and isolated environments, including agent environments. It is distributed under the Apache License 2.0 and can be built from source or obtained as release binaries.
Moonshine Voice is an open-source AI toolkit for developers building real-time voice agents and applications. It provides on-device speech-to-text, intent recognition, and text-to-speech, including streaming transcription that processes audio while the user is still speaking. The project offers speech-to-text models ranging from higher-accuracy models to approximately 1 MB models and provides one library for Python, JavaScript/WASM, iOS, Android, macOS, Linux, Windows, and Raspberry Pi. It is used by Gemma Translator for speech recognition and speech output, and can run without an account or API keys. The project is distributed under the MIT License, including its models by default, with legacy non-streaming models for non-English languages covered by a non-commercial Moonshine Community License.
M5Stack is a modular hardware and open-source software platform for IoT prototyping and deployment. Its standard 5 × 5 cm system uses stackable functional modules built around an ESP32 microcontroller, with features including a microSD card slot, USB-C port, and expansion connectors. M5Stack also provides multifunctional base modules and ODM services for projects ranging from prototypes to mass production. In the cited video, M5Stack hardware is used by VibeWatch as a smartwatch interface for controlling multiple AI coding agents.
Fractera is an open-source, self-hosted agent-engineering infrastructure platform from Fractera. Given an Ubuntu 24.04 VPS, its installer configures the operating system environment, Nginx routing, HTTPS certificates, role-based authentication, a database, local object storage, browser terminals, and a starter application or other public Git repository. The resulting deployment keeps code, data, and agent execution on the user’s server rather than relying on a managed cloud platform. Its deployment architecture is described as deterministic and MCP-first. The server includes a browser-accessible multi-agent development environment coordinated by the Hermes orchestrator, with five specialized code-generation engines and shared local RAG memory; the platform is also described as supporting agents for coding, marketing, and sales. Fractera can build the infrastructure around multiple application frameworks and Git repositories, rather than requiring a single frontend stack. The current Next.js starter is an approximately 50,000-line enterprise boilerplate with multilingual routing, static SEO trees, a SQLite write-ahead-logging database, NextAuth v5 session state, Next.js parallel routing, and on-demand incremental static regeneration. Its largely static base architecture is intended to limit agents’ repeated scanning and reconstruction of workspace context; the project claims its optimizer can reduce token costs by up to 90% by limiting context-window inflation.
An open-source marketing-agent team from Vercel Labs built on the eve framework. Users submit launch planning, copywriting, or SEO requests through Slack or the eve terminal TUI; a lead agent loads the shared brand-context document, writes a self-contained brief, routes it to one of five specialists—product marketing, content marketing, social media, SEO, or email—and returns the resulting deliverable. The lead does not write deliverables itself, and specialists work independently without shared conversation history; the product marketer alone maintains the brand-context document, which the other specialists read at the start of a task. Newsletter work passes from the content marketer to the email specialist for inbox adaptation and sending. Outputs are delivered through Notion, Typefully, Resend, or stored audit artifacts. The system pauses irreversible operations—including Resend sends and deletes, Typefully deletes and scheduled publishing, and Notion page moves—for approval in Slack or the terminal; its email specialist is restricted to an allowlist of Resend tools. Deployment provisions Notion, Resend, and Slack connectors and a Vercel Blob store, and requests a Typefully API key. The implementation uses TypeScript, Vercel Connect, MCP connections for Notion, Typefully, and Resend, Vercel AI Gateway for model access, and Vercel Sandbox for reference files and shell commands. Email campaigns require a verified Resend sending domain and at least one segment configured in Resend.
NVIDIA-labs Object Oriented Agents (NOOA) is a model-agnostic Python framework for building AI agents as ordinary Python objects. A single agent class combines typed state, capabilities, prompts, and interfaces: fields represent state, methods represent capabilities, docstrings provide prompts, and type annotations define contracts. Methods with an ellipsis body become LLM-driven agentic loops, while methods with ordinary bodies remain deterministic Python. The runtime lets a model act by writing Python in a Jupyter-style REPL with access to the agent object, imports, and helper functions. It supports typed inputs and outputs with automatic retries, live-object arguments passed by reference, and model-callable context and event APIs, and can use hosted or local models through LiteLLM-supported providers. The core framework is distributed as the `nooa` Python package through PyPI or the project repository. Separate packages or extras provide a CLI and trace viewer, Agent Client Protocol support, long-term memory, benchmarking, and evaluation tooling. NOOA is research software whose agents may execute LLM-generated code; its AST checks and module deny-lists are defense-in-depth guardrails rather than a containment boundary, so the documentation recommends running such agents in an OS-level sandbox such as a container, virtual machine, or NVIDIA OpenShell.
Context Ontology Accelerator is an open-source semantic context layer for AWS and AI agents. It combines knowledge graphs, formal ontologies, rule-based systems, and modern AI in a Scan → Model → Serve workflow: it connects data sources, discovers schemas, enriches metadata, and ingests documents; induces and manages ontologies, defines metrics, and builds a unified semantic graph; then serves context through SPARQL federation using a Virtual Knowledge Graph, knowledge-graph traversal, queries, and MCP tools. The repository includes source-ingestion services for databases and documents, ontology induction and reasoning with HermiT and ELK, an Ontop-based Virtual Knowledge Graph, a context manager for query orchestration, and a React frontend. Access is controlled through namespace isolation and role-based authorization, with namespace-scoped roles and platform-level roles. The implementation uses Python and TypeScript, AWS CDK, Smithy-generated API contracts, and an Nx-managed monorepo; the project is licensed under Apache License 2.0.
Prisma ORM is an object-relational mapper and query builder for Node.js and TypeScript applications, developed by Prisma. It supports PostgreSQL, MySQL, MariaDB, SQL Server, SQLite, MongoDB, and CockroachDB. The repository's Prisma Next development line is a TypeScript rewrite with a composable core, database support built through a public extension interface, and agent skills that can be registered in a project so AI coding agents can work from a Prisma contract and workflow-specific references. Prisma Next is described as early access and not recommended for production workloads yet.
MiniMax is an artificial-intelligence company and model provider offering large language models through hosted APIs, including models designed for coding, reasoning, and tool-using agent workflows. It is used as a model provider by agent frameworks such as Hermes Agent.
Overlap is a web-based AI platform for automated video clipping and short-form content production, marketed to creators, podcasts, and media organizations. It ingests uploaded long-form video or podcast audio and uses agentic AI workflows to find high-interest moments, produce clips and shorts, generate captions and on-screen graphics (or B-roll), and export or publish social posts for distribution.
An open-source research simulation for computational agents that model believable human behavior in an interactive game environment. The repository includes a Django-based environment server and a separate agent simulation server; the two run concurrently, with the simulation using an OpenAI API key and the browser-based environment providing the interactive view and replayable demo animations. The project was tested with Python 3.9.12 and includes setup instructions, required packages, and predefined simulation scenarios.
MoneyPrinterTurbo is an AI short-video generation tool that turns a topic or keyword into a high-definition video. Its automated workflow generates or accepts a script, extracts search keywords, matches footage from local materials or supported stock sources, synthesizes narration, creates configurable subtitles, adds background music, and renders the result. It provides WebUI, API, CLI, and AI-agent interfaces; supports batch generation, multiple languages, portrait and landscape formats, adjustable clip durations, and multiple text-to-speech and large-language-model providers. It can also generate visual material through supported text-to-video models and automatically publish completed videos to TikTok, Instagram, and YouTube Shorts.
Career Ops is an open-source AI job-search system that runs locally through compatible AI coding CLIs. It scans Greenhouse, Ashby, Lever, and company career pages; compares listings with a CV and evaluates them in an A–H report with a global 1–5 score. The system uses Playwright to navigate career pages, can process multiple offers with sub-agents, generates ATS-oriented CV PDFs tailored to individual job descriptions, records applications in a tracker, and researches companies and potential contacts. Its separate Block G assesses posting legitimacy, including scam or ghost-job risk, while Block H drafts additional material only for highly scored roles. It is intended to filter opportunities rather than submit applications, which users review and submit themselves.
Microsandbox is a local-first microVM runtime and library for running untrusted workloads, including AI-agent tasks, user code, plugins, CI jobs, development environments, scrapers, and automation. It uses hardware-level isolation and runs standard OCI container images from registries such as Docker Hub and GHCR, with Docker-like image, command, shell, and volume workflows. The project provides SDKs for TypeScript, Rust, Python, Go, and Ruby, along with a CLI for booting and controlling sandboxes. Applications can create microVMs as child processes without a setup server or long-running daemon; sandboxes can also run detached for long-lived sessions. Agent Skills and an MCP server support agent-created sandboxes, and the project documents secret handling intended to keep keys out of the VM. It runs on Linux, macOS, and Windows, with platform-specific virtualization requirements, and is described as beta software.
Remocn is a copy-paste component library for building videos with Remotion. It provides animations, transitions, backgrounds, scenes, fades, wipes, and kinetic titles as Remotion code that developers add to their projects through the shadcn component registry and can modify directly. Its components use Remotion's `useCurrentFrame()`, `interpolate()`, and `spring()` APIs. Component pages provide live previews using `@remotion/player`, allowing frame-by-frame scrubbing. Remotion is required as a prerequisite, and the project also provides an optional agent skill for setting up and working on Remotion video projects.
E2B is a cloud sandbox service for securely running AI-generated code, tools, and agents. It provides infrastructure for executing agent workloads in isolated environments.
Mu is a self-hosted personal-agent platform that provides one home for agents, tools, services, and collected data. Its server can be accessed through a web app, email, MCP, HTTP, and a command-line interface, and supports bringing your own model. Mu exposes services for mail, chat, news, video, markets, search, files, calendars, and small web applications. These services are available through APIs, the CLI, the web, or as MCP tools. Data collected by the services is archived locally so it can be searched across sources. Users can define agents by name, prompt, and permitted tools, then chat with them through the web, mail, or XMPP. The platform also provides a unified inbox for email and chat threads, including access through IMAP.
Scandinavian Design is a Cursor agent skill for applying a restrained Scandinavian visual system to an existing codebase. It can redesign or review an interface, prototype visual directions, and inspect a complete flow while preserving useful brand and accessibility decisions. Its stated design rules use a black-and-white foundation, sans-serif typography, generous spacing, and purposeful product imagery. The repository includes ten demo sites restyled through a CSS-override fallback, with desktop and mobile before-and-after captures, and can be installed with npx skills add ericzakariasson/scandinavian-design.