1,038 tools and products — trending open source, and what gets used in AI and other work.
Harbor is a framework from the creators of Terminal-Bench for evaluating and optimizing AI agents and language models in container environments. It can run arbitrary agents, including Claude Code, OpenHands, and Codex CLI, against shared or custom benchmarks and environments such as Terminal-Bench, SWE-Bench, and Aider Polyglot. Harbor launches benchmark runs locally with Docker or in parallel through cloud and sandbox providers, and can generate rollouts for reinforcement-learning optimization. It is distributed as a Python package installable with uv or pip and provides command-line tools for running datasets, selecting agents and models, and listing supported benchmarks.
An AI tool that orchestrates agents to synthesize realistic reinforcement-learning environments and tasks from domain descriptions, real usage data, scenarios, and personas.
An AI API service from MiniMax that provides access to generation capabilities and context preprocessing associated with the MiniMax-H3 language model. The API includes a 2K generation stage and a context preprocessor that are not included in MiniMax-H3's open-weight release.
Moonshine Voice is an open-source AI toolkit for developers building real-time voice agents and applications. It provides on-device speech-to-text, intent recognition, and text-to-speech, including streaming transcription that processes audio while the user is still speaking. The project offers speech-to-text models ranging from higher-accuracy models to approximately 1 MB models and provides one library for Python, JavaScript/WASM, iOS, Android, macOS, Linux, Windows, and Raspberry Pi. It is used by Gemma Translator for speech recognition and speech output, and can run without an account or API keys. The project is distributed under the MIT License, including its models by default, with legacy non-streaming models for non-English languages covered by a non-commercial Moonshine Community License.
Krea is a web-based AI creative suite for generating, editing, and enhancing images, video, and 3D assets. Its image tools accept text and image prompts and include prompt management, galleries, and team or shared workspaces for organizing and collaborating on generative assets. It is offered as a SaaS platform at krea.ai with free and paid tiers.
Fathom is a commercial AI meeting assistant and call recorder that captures video calls, including Zoom meetings, and generates searchable transcripts, highlights, shareable clips, summaries, and notes. It is provided as a web app with Zoom and other conferencing-platform integrations, and supports exporting meeting artifacts for sharing and archival.
Steno is an open-source, privacy-first desktop AI notepad and meeting notetaker for Windows and macOS, intended for confidential conversations. It records microphone and system audio, provides live or post-meeting transcription with speaker labels, generates summaries, and supports natural-language queries across individual notes or the meeting library. Its local mode uses on-device AI models; on Apple Silicon, live transcription uses Parakeet TDT v3 through MLX, while Whisper supports post-stop transcription in 99 languages. On macOS 14.4 and later, native Core Audio Tap captures both sides of virtual meetings without a virtual audio cable and with selectable microphone input. Steno can detect meeting start and end events, offer to start or summarize recordings, append recordings to existing notes, provide global recording shortcuts, and include user notes in generated summaries. Notes, transcripts, and summaries are stored as editable Markdown, with one-way synchronization to an Obsidian vault and report templates for structured outputs. The project states that local recordings, transcripts, and summaries remain on the device and have no cloud usage limits; users can alternatively configure hosted models through OpenAI, Anthropic, AWS Bedrock, or a custom API endpoint.
Ailin is an open-source, self-hosted collective-intelligence engine and OpenAI-compatible interface in which AI models collaborate inside a collective model rather than being routed to a single model. It indexes models from frontier APIs, open-weight systems, and its own model family, coordinating them through dozens of strategies such as debate, critique, and synthesis. The platform records coordination decisions and decision provenance for each request, can return reasoning and cost information with an audit trail, and includes model discovery, an audit layer, and a closed-loop training pipeline; some components are still maturing.
Fractera is an open-source, self-hosted agent-engineering infrastructure platform from Fractera. Given an Ubuntu 24.04 VPS, its installer configures the operating system environment, Nginx routing, HTTPS certificates, role-based authentication, a database, local object storage, browser terminals, and a starter application or other public Git repository. The resulting deployment keeps code, data, and agent execution on the user’s server rather than relying on a managed cloud platform. Its deployment architecture is described as deterministic and MCP-first. The server includes a browser-accessible multi-agent development environment coordinated by the Hermes orchestrator, with five specialized code-generation engines and shared local RAG memory; the platform is also described as supporting agents for coding, marketing, and sales. Fractera can build the infrastructure around multiple application frameworks and Git repositories, rather than requiring a single frontend stack. The current Next.js starter is an approximately 50,000-line enterprise boilerplate with multilingual routing, static SEO trees, a SQLite write-ahead-logging database, NextAuth v5 session state, Next.js parallel routing, and on-demand incremental static regeneration. Its largely static base architecture is intended to limit agents’ repeated scanning and reconstruction of workspace context; the project claims its optimizer can reduce token costs by up to 90% by limiting context-window inflation.
An open-source marketing-agent team from Vercel Labs built on the eve framework. Users submit launch planning, copywriting, or SEO requests through Slack or the eve terminal TUI; a lead agent loads the shared brand-context document, writes a self-contained brief, routes it to one of five specialists—product marketing, content marketing, social media, SEO, or email—and returns the resulting deliverable. The lead does not write deliverables itself, and specialists work independently without shared conversation history; the product marketer alone maintains the brand-context document, which the other specialists read at the start of a task. Newsletter work passes from the content marketer to the email specialist for inbox adaptation and sending. Outputs are delivered through Notion, Typefully, Resend, or stored audit artifacts. The system pauses irreversible operations—including Resend sends and deletes, Typefully deletes and scheduled publishing, and Notion page moves—for approval in Slack or the terminal; its email specialist is restricted to an allowlist of Resend tools. Deployment provisions Notion, Resend, and Slack connectors and a Vercel Blob store, and requests a Typefully API key. The implementation uses TypeScript, Vercel Connect, MCP connections for Notion, Typefully, and Resend, Vercel AI Gateway for model access, and Vercel Sandbox for reference files and shell commands. Email campaigns require a verified Resend sending domain and at least one segment configured in Resend.
NVIDIA-labs Object Oriented Agents (NOOA) is a model-agnostic Python framework for building AI agents as ordinary Python objects. A single agent class combines typed state, capabilities, prompts, and interfaces: fields represent state, methods represent capabilities, docstrings provide prompts, and type annotations define contracts. Methods with an ellipsis body become LLM-driven agentic loops, while methods with ordinary bodies remain deterministic Python. The runtime lets a model act by writing Python in a Jupyter-style REPL with access to the agent object, imports, and helper functions. It supports typed inputs and outputs with automatic retries, live-object arguments passed by reference, and model-callable context and event APIs, and can use hosted or local models through LiteLLM-supported providers. The core framework is distributed as the `nooa` Python package through PyPI or the project repository. Separate packages or extras provide a CLI and trace viewer, Agent Client Protocol support, long-term memory, benchmarking, and evaluation tooling. NOOA is research software whose agents may execute LLM-generated code; its AST checks and module deny-lists are defense-in-depth guardrails rather than a containment boundary, so the documentation recommends running such agents in an OS-level sandbox such as a container, virtual machine, or NVIDIA OpenShell.
U-Pool is a desktop provider and account switcher for Claude Code, Claude Desktop, Codex, Hermes, OpenCode, and Cursor. It keeps endpoint URLs, API keys, models, and provider accounts in per-tool lists, then makes a selected provider or Cursor account active without manual edits to settings.json, config.toml, other tool configuration files, or Windows user environment variables. Its reachability probe sends a plain GET request to a provider’s models endpoint to measure latency without running a completion or consuming a token. The app uses a Python backend, a statically exported Next.js interface, and a native OS webview rather than Electron or a Node runtime. Its source of truth is ~/.u-pool/config.json; on each switch, an adapter changes only the keys it owns in each target file and preserves other settings such as plugins, hooks, marketplaces, themes, MCP servers, and project trust. Writes use a temporary file followed by os.replace, keep an adjacent backup, and stop rather than overwrite a configuration file that cannot be parsed. Existing provider settings can be imported on first run. The app includes provider presets, an official-vendor mode that clears U-Pool-managed settings, per-provider permission controls, cached installed Claude Code and Codex versions, and controls that erase the transcript and prompt-history files for those CLIs. The Cursor pool stores multiple signed-in Cursor accounts, accepts cookies in common formats, displays account names, email addresses, plans, and usage, and switches accounts by changing the specified Cursor authentication database rows. Claude Desktop is currently preview-only: providers can be saved and tested but are not written to its configuration. Hermes and OpenCode have dedicated configuration writers, while Codex writes its provider, model, approval, sandbox, web-search, and authentication settings. The core configuration operations are cross-platform. Windows-only integrations write owned variables to HKCU\\Environment and broadcast changes, launch the app at sign-in, and provide an in-app updater; Hermes, OpenCode, and Cursor read credentials from their own files and do not use those environment-variable writes.
Context Ontology Accelerator is an open-source semantic context layer for AWS and AI agents. It combines knowledge graphs, formal ontologies, rule-based systems, and modern AI in a Scan → Model → Serve workflow: it connects data sources, discovers schemas, enriches metadata, and ingests documents; induces and manages ontologies, defines metrics, and builds a unified semantic graph; then serves context through SPARQL federation using a Virtual Knowledge Graph, knowledge-graph traversal, queries, and MCP tools. The repository includes source-ingestion services for databases and documents, ontology induction and reasoning with HermiT and ELK, an Ontop-based Virtual Knowledge Graph, a context manager for query orchestration, and a React frontend. Access is controlled through namespace isolation and role-based authorization, with namespace-scoped roles and platform-level roles. The implementation uses Python and TypeScript, AWS CDK, Smithy-generated API contracts, and an Nx-managed monorepo; the project is licensed under Apache License 2.0.
LingBot-Map is a feed-forward 3D foundation model from the Robbyant Team for streaming 3D reconstruction from image sequences and video. Its Geometric Context Transformer unifies coordinate grounding, dense geometric cues, and long-range drift correction through anchor context, a pose-reference window, and trajectory memory. The system uses paged KV-cache attention for streaming inference and supports interactive browser-based reconstruction with point-cloud visualization, keyframe intervals, windowed processing for long sequences, and sky masking. It also provides an offline rendering pipeline for image folders or video, along with evaluation pipelines for datasets including KITTI and Oxford Spires. The repository reports approximately 20 FPS inference at 518×378 resolution on sequences exceeding 10,000 frames, and documents PyTorch installation with FlashInfer as the recommended attention backend and native SDPA as a fallback.
Loki is an open-source log aggregation system developed by Grafana Labs that indexes and queries logs using Prometheus-style labels instead of full-text indexing. It ingests log streams via shippers such as Promtail, Fluentd, or other collectors, stores data in local or object-storage backends, and is queried with the LogQL language; it integrates with Grafana for visualization and is available for self-hosting or via Grafana Cloud.
A serverless AI inference platform from Cloudflare that runs machine-learning models on Cloudflare's network and exposes them through the Workers platform. It supports tasks such as text generation, image generation, speech recognition, and text embeddings, making it usable for applications including semantic search.
MiniMax is an artificial-intelligence company and model provider offering large language models through hosted APIs, including models designed for coding, reasoning, and tool-using agent workflows. It is used as a model provider by agent frameworks such as Hermes Agent.
An MCP server from Zapier that connects authenticated applications so compatible AI tools can access and use them through a single integration layer. It supports clients including Cursor, Claude Code, and Codex.
Blitzy is an enterprise AI software platform that ingests large codebases and generates coordinated code changes across them, with a focus on refactoring and modernizing legacy applications.
Claude is a family of large language models developed and hosted by Anthropic. Available through Anthropic's web chat products and a commercial API, Claude accepts text prompts (and structured inputs or file attachments via the API) and returns natural-language and code outputs for tasks such as coding, summarization, document Q&A, and general business and productivity workflows; it is used inside Anthropic products (for example Claude Code) and by third-party services via Anthropic's API and enterprise agreements.
An AI governance and compliance platform from IBM's watsonx suite that provides audit trails, bias detection, prompt monitoring, policy enforcement, and reporting for enterprise AI deployments. It ingests models, prompts and usage telemetry and produces audit logs, explainability and bias reports, and configurable controls to support regulatory and internal compliance processes.
Overlap is a web-based AI platform for automated video clipping and short-form content production, marketed to creators, podcasts, and media organizations. It ingests uploaded long-form video or podcast audio and uses agentic AI workflows to find high-interest moments, produce clips and shorts, generate captions and on-screen graphics (or B-roll), and export or publish social posts for distribution.
An open-source research simulation for computational agents that model believable human behavior in an interactive game environment. The repository includes a Django-based environment server and a separate agent simulation server; the two run concurrently, with the simulation using an OpenAI API key and the browser-based environment providing the interactive view and replayable demo animations. The project was tested with Python 3.9.12 and includes setup instructions, required packages, and predefined simulation scenarios.
MoneyPrinterTurbo is an AI short-video generation tool that turns a topic or keyword into a high-definition video. Its automated workflow generates or accepts a script, extracts search keywords, matches footage from local materials or supported stock sources, synthesizes narration, creates configurable subtitles, adds background music, and renders the result. It provides WebUI, API, CLI, and AI-agent interfaces; supports batch generation, multiple languages, portrait and landscape formats, adjustable clip durations, and multiple text-to-speech and large-language-model providers. It can also generate visual material through supported text-to-video models and automatically publish completed videos to TikTok, Instagram, and YouTube Shorts.
Career Ops is an open-source AI job-search system that runs locally through compatible AI coding CLIs. It scans Greenhouse, Ashby, Lever, and company career pages; compares listings with a CV and evaluates them in an A–H report with a global 1–5 score. The system uses Playwright to navigate career pages, can process multiple offers with sub-agents, generates ATS-oriented CV PDFs tailored to individual job descriptions, records applications in a tracker, and researches companies and potential contacts. Its separate Block G assesses posting legitimacy, including scam or ghost-job risk, while Block H drafts additional material only for highly scored roles. It is intended to filter opportunities rather than submit applications, which users review and submit themselves.
E2B is a cloud sandbox service for securely running AI-generated code, tools, and agents. It provides infrastructure for executing agent workloads in isolated environments.
Venice AI is a private AI service for generating text, images, characters, and video, with access to AI inference without the intermediary markup associated with services such as OpenRouter.
Mu is a self-hosted personal-agent platform that provides one home for agents, tools, services, and collected data. Its server can be accessed through a web app, email, MCP, HTTP, and a command-line interface, and supports bringing your own model. Mu exposes services for mail, chat, news, video, markets, search, files, calendars, and small web applications. These services are available through APIs, the CLI, the web, or as MCP tools. Data collected by the services is archived locally so it can be searched across sources. Users can define agents by name, prompt, and permitted tools, then chat with them through the web, mail, or XMPP. The platform also provides a unified inbox for email and chat threads, including access through IMAP.
Scandinavian Design is a Cursor agent skill for applying a restrained Scandinavian visual system to an existing codebase. It can redesign or review an interface, prototype visual directions, and inspect a complete flow while preserving useful brand and accessibility decisions. Its stated design rules use a black-and-white foundation, sans-serif typography, generous spacing, and purposeful product imagery. The repository includes ten demo sites restyled through a CSS-override fallback, with desktop and mobile before-and-after captures, and can be installed with npx skills add ericzakariasson/scandinavian-design.
Marketing OS is a Markdown-based agent skill that turns Claude, Codex, or Cursor into a structured marketing workspace. Its 14 modules cover website and funnel audits, AI-search citability, copy, hooks, paid ads, email, social posts, launches, positioning, pricing, competitor research, app-store optimization, analytics, and prose-quality checks. The skill routes each task to the relevant module and can fan out multidimensional work across subagents when the host supports it. It produces scored reports, prioritized fixes, diagnostic briefs, marketing copy and sequences, launch plans, and other finished artifacts rather than advice alone; its reports are designed to identify what could not be determined and label scores as heuristics. It is distributed as pure Markdown under the MIT license and can be installed through Claude Code, uploaded to Claude, or copied into the skills directory of compatible agents.
Zoetrope is a terminal and browser visualization tool for Claude Code sessions. It reads Claude Code's local JSONL transcripts and renders the main agent, spawned subagents, workflow groups, and their tool calls as a live flow graph, updating as a session runs or replaying a completed session from its recorded timestamps. Its event-indexed timeline supports scrubbing, pausing, stepping between prompt eras, following the live edge, and seeking backward to view the graph at an earlier state; agent cards show status, current tools, tool counts, and output tokens, while tool calls resolve to success or failure. The browser version uses the same engine compiled to WebAssembly, and the README states that sessions remain local and read-only. It is distributed through Homebrew, Cargo, prebuilt binaries, source builds, and the browser interface; its command-line executable is `zoe`.
Autoprompt is a coding-agent skill and CLI that turns a development goal into a managed loop of scoping, implementation, testing, review, repair, and verification. It coordinates coding agents with configurable concurrency and model routing, and supports installation into several coding-agent environments, including Claude Code, Codex, OpenCode, Kilo Code, and VS Code. The repository reports an OpenCode comparison in which Autoprompt reduced failures from 29 of 89 tasks to 16, or 45% fewer failures, with an expected trade-off of roughly three times the execution time and twice the token use. It requires Node.js 20+, Python 3.11+ with PyYAML, and Bash 4.3+ on macOS or Linux.
claude-db is a persistent-memory and code-graph system for Claude Code. It captures project context through hooks and injects recalled context into prompts, while its MCP server provides text search, symbol-usage, code-explanation, and dependency-path modes. Memory uses SQLite by default and can be configured to use shared MongoDB or PostgreSQL storage. The `scan` command builds the code graph and hashes files so later scans re-parse only changed files; TypeScript, TSX, JavaScript, Python, Go, Rust, and Ruby receive structured parsing, while other listed languages are analyzed by pattern matching with inferred references. It is distributed as an npm CLI and can install its Claude Code hooks and MCP configuration per project.
Jixu is a TypeScript harness for running a single AI agent in durable, recoverable Threads. An immutable Agent definition is given Tools and Skills, while each Thread records execution, context decisions, and external work as ordered Events; state is deterministically derived from those Events, with external work recorded before dispatch. This allows a Thread to recover after interruption, replay completed work without repeating live side effects, fork, or continue later without a second workflow engine. The public API centers on `createHarness`, `createThread`, and `thread.send`; the project also provides a native terminal UI and SQLite-backed storage packages. Jixu is pre-1.0, and its public API may change before 1.0.
nopus is a deterministic prose checker for coding-agent responses. It evaluates completed answers for uncommon or highly uncommon wording, abstract vocabulary and sentences, noun and modifier stacks, dense phrase load, and formulaic filler, using packaged language-frequency data, human-rated word-concreteness data, a computing glossary, and a small allowlist of literal style cues. When a response exceeds the configured complexity threshold, it requests one clearer rewrite; rejected responses can be hidden from the terminal transcript while remaining in the session and model history. It is distributed as an extension or plugin for Pi, Claude Code, and Codex, with controls for checking, sensitivity, activation, and hiding the original response.
Benjamin-Plus is an instruction-set skill for coding agents that aims to reduce token and tool-call usage without changing the agent's implementation task. It teaches five operating habits: batch repository reconnaissance into one pass, inspect narrow file windows unless full data is needed, probe dependencies in one command, treat the task's specified verification command as the definition of done, and poll unfinished builds at longer intervals rather than repeatedly checking them. The skill is intended to be injected into an agent's instructions rather than installed as a discoverable skill. Its repository provides an injected instruction file for Claude Code hooks, Codex CLI's AGENTS.md, project CLAUDE.md files, or other agents' system prompts. The project reports measured cost reductions ranging from about 10% to 18% depending on baseline behavior, with unchanged quality in its reported evaluation; it also documents a cross-platform Java SWE-bench measurement with reduced cost and tool calls.
Fullstack-agent is a Claude Code installer wizard that assembles a local AI-agent setup from several optional components. It can install ai-memory-vault for persistent memory stored in plain-text files, backtalk for push-to-talk voice input and spoken replies, ai-visualizer for full-screen visual states synchronized with the conversation, and barehands for optional webcam hand tracking that moves notes and images on screen. The wizard presents the components in a guided conversation and installs only the pieces selected by the user. It runs on Claude Code and requires a Claude subscription; macOS and Linux also use git, while the Windows setup downloads the repository as a ZIP and configures git during installation.
Procoder is a Go binary that provides commit gates and quality controls for AI coding agents. It checks formatting, tests, linting, secrets, Git and repository hygiene, CI and infrastructure hygiene, and documentation health, while treating unavailable checks as failures. Its tools can be run by the agent, and lifecycle hooks can invoke checks at fixed points; the binary computes and reports findings while the agent reviews and applies changes rather than modifying repository files itself. A lessons loop records escaped bug classes for future checks, and adapters support Claude Code and other agents that use integrations such as AGENTS.md. It has no runtime dependencies and is distributed as a single binary.
Vomit is a local, offline tool that converts Claude's token output into shorter English text by piping it through a separate local LLM. It can install as a Go binary, configure a connection to a local LLM through `vomit init`, and use Claude hooks to replace the displayed output with the processed version via `vomit scrub -claude`. Vomit also provides a non-invasive tail mode that runs alongside Claude, listing session identifiers and translating the latest or a specified session with `vomit tail`. The repository states that it has no external dependencies or telemetry; the local model sees only what Claude communicates, not its actions or files. It warns that the model may hallucinate, processing can be slow, and the project has been tested only on Mac. It is licensed under GNU GPLv3.
Hermes3D is an open-source, self-hosted visualization and interaction layer for AI agents. It presents connected agents as workers in a live retro 3D office, with surfaces for standups, pull-request reviews, task execution, monitoring, and agent chat; it also provides a 2D pixel-office view for lower-power machines and an office builder for editing layouts. The frontend connects to agent runtimes through a bundled Hermes WebSocket gateway adapter, a direct HTTP custom runtime provider, or a built-in demo gateway. Runtime state remains in the connected backend, while Studio stores local interface preferences. Hermes3D does not build or run the upstream agent runtimes; it supplies the office UI, Studio, and adapter/proxy layer using the Hermes3D gateway protocol. The repository describes it as an independent community project maintained by LukeTheDev and unaffiliated with the backend teams it connects to.
Alvarmethod is a portable collection of agent skills that implements Eero Alvar’s AI learning loop for Codex, Claude Code, Grok, Pi, OpenCode, Cursor, and other skills-CLI agents. Its full `teach` skill probes the learner with graded questions, builds a Mermaid dependency DAG, teaches one reasoning step at a time, and uses a native quiz tool to lock in each node; a failed quiz inserts a prerequisite before continuing rather than merely repeating the same material. The pack also includes separate skills for probing, maintaining a learner profile, generating a visual, and fact-checking claims before they are taught. It stores learner, map, session, and visual files in a `.alvar/` directory and can be installed globally or per project through the `skills` CLI or its installer. The repository states that the skills, installer, and documentation are released under the MIT license.
Habit Hooks is a developer tool that steers AI coding agents toward code-quality and refactoring practices. It runs project linters and replaces each raw rule violation with a short coaching guide that explains an actionable fix, using the linter finding as a cue and the guide as the response rather than presenting the agent with a bare metric. The project is designed to reduce metric gaming, technical debt, and the context needed for later tasks; its repository reports an independent study in which coaching produced genuine fixes more often than bare linter targets. It is installed as a Python package with `uv`, pip, pipx, or Homebrew, then initialized in a project with `habit-hooks init`. Initialization detects the project language and writes `.habit-hooks/config.toml`; separate plugins are provided for Python, TypeScript, PHP, and Java, with language-specific detectors such as Ruff, ESLint, and jscpd enabled separately.
Google Cloud Vertex AI is a managed platform for building, deploying, and using generative AI and machine learning models. It provides access to Google and third-party models, including Anthropic models used with tools such as Claude Code, along with APIs and development services for AI applications.
Vertex AI Model Garden is Google Cloud's catalog for discovering, evaluating, and deploying generative AI and foundation models from Google and third-party providers. It can provide access to models such as Anthropic Claude for use through Vertex AI services and integrations.
Google AI Studio is a web platform for generating web applications from natural-language prompts using Gemini models and moving them toward production deployment. It supports code generation and integrations including Firestore, Cloud SQL, authentication, and database schema creation.
Vertex AI is Google Cloud’s managed platform for developing and deploying machine-learning and generative-AI applications. It provides hosted models and APIs that applications, including agents, can call at runtime, along with tools for model training, evaluation, and serving.
NVIDIA Blackwell is a family of GPU architectures and data-center GPU systems for generative AI and high-performance computing, including the B300 GPU platform. Blackwell systems can be deployed with Google Distributed Cloud to run Gemini and other AI workloads on premises.
Gemini Enterprise Agent Platform, formerly Vertex AI, is a Google Cloud platform for developers to build, run, scale, govern, and optimize autonomous AI agents and agent workflows. It provides session and memory management, persistence, parallel execution, model swapping, and built-in support capabilities.
Google Cloud Agent Observability provides visibility into AI agents’ execution paths, actions, and reasoning, with evaluation capabilities for measures such as accuracy and hallucination detection. It is documented as part of Google Cloud Observability.
Kiro is a proprietary agentic integrated development environment and command-line interface from Amazon Web Services. It supports specification-driven development, steering files, hooks, and cloud deployment, and uses Anthropic models through Amazon Bedrock.
itai is Viktor's AI toolkit for Kubernetes remediation. A controller watches and classifies Kubernetes events, sends repeated warnings to an agent for analysis, and can notify an operator, create a pull request, or apply a fix.
AlexNet is a convolutional neural network architecture developed by Alex Krizhevsky, Ilya Sutskever, and Geoffrey Hinton for large-scale image classification. Its 2012 ImageNet result helped establish deep learning's use of deep neural networks, ReLU activations, dropout, and GPU-based training for computer vision.
DevAssure is an AI-powered software testing platform whose O2 agent reads code changes, maps potentially affected application flows, and validates them automatically, including through browser interactions based on plain-English expected behavior. It generates scriptless, self-healing acceptance tests and reports results for pull requests.
Arcade is a runtime for the Model Context Protocol (MCP) that helps AI agents take actions in real systems. It provides authentication and permissions management, reliable tool calls, and audit trails.
A remote Model Context Protocol (MCP) server from Atlassian that lets MCP-compatible AI clients retrieve and work with live Jira and Confluence content. It applies the user's existing Atlassian permissions, allowing permission-aware access without maintaining a separately embedded retrieval-augmented generation data pipeline.
Muse Code is an AI-powered terminal coding agent for inspecting repositories, investigating code issues, generating HTML reports, modifying code, spawning parallel sub-agents, and auditing pull requests.
An evaluation toolkit for Strands AI agents that measures off-script behavior, violations of behavioral guardrails, and other incorrect behavior against test and live data. It is associated with Strands Agents, an open-source AI agent SDK for Python and TypeScript.
Vois is a local AI voice-generation and audio-processing studio for writing, casting, voice selection and cloning, multi-speaker scripts, editing, and mastering. It provides local text-to-speech, more than 100 natural voices, audio tools for podcasts, audiobooks, and video, and integrations with AI agents; its Pro plan adds Omni, more than 600 languages, and Voice Design.
NVIDIA DGX Spark is a compact AI workstation from NVIDIA for local development, fine-tuning, and inference of generative-AI models. It is built around the NVIDIA GB10 Grace Blackwell Superchip, with 128 GB of unified memory and up to 1 petaflop of FP4 AI performance, and runs NVIDIA DGX OS.
Vapi is a developer platform for building, testing, and deploying conversational voice AI agents. It provides a cascaded voice-agent architecture and supports customer-supplied text-to-speech servers for specialized applications.