1,037 tools and products — trending open source, and what gets used in AI and other work.
Codewhale is an open-source coding agent for the terminal, built in Rust and independently maintained. It reads repositories, edits files, runs commands, inspects results, and works toward user-defined goals. It connects to hosted providers or local models through Ollama, vLLM, or SGLang, supports switching providers and models during a session, and can run interactively in a terminal UI or non-interactively with the `exec` command. Its control mechanisms include read-only planning, Ask, Auto-Review, and Full Access approval modes; `/undo` reverts the last turn and `/restore` returns the workspace to an earlier snapshot. It also supports saved sessions, durable goals, reviewable workflows, agent coordination, MCP servers, skills, hooks, and configurable agent roles. The project runs on the user's machine with the access granted to it, with optional operating-system sandboxing where supported, and is distributed under the MIT license.
TokensBurned is a privacy-first tool that publishes AI coding activity on a GitHub profile through a live SVG card. It collects token counts and model metadata from AI coding harnesses, reduces raw sessions locally, aggregates usage into 15-minute buckets, and serves totals, activity heatmaps, harness/provider/model comparisons, and an optional site-wide rank. It supports native or official integrations for Claude Code, Codex, Gemini CLI, Cline CLI or SDK, and GitHub Copilot CLI, with OTLP or standalone CLI fallbacks for other harnesses. The tool uploads aggregate token counts, harness/provider/model metadata, hashed session identifiers, time buckets, and request counts, while its stated privacy boundary excludes prompts, responses, source code, tool payloads, repository names and paths, transcript files and paths, and API keys. Public display is a separate explicit opt-in tied to a verified GitHub account; users can build cards with selectable layouts, themes, heatmaps, comparisons, and rankings. It is distributed through harness plugins and an npm CLI, installs no cron job, daemon, proxy, or Git synchronization task, and is licensed under the MIT License.
LightNav-0 is an open-source generalist embodied-navigation model from the Light Origins Team. It uses a pretrained Qwen3-VL backbone with a single egocentric RGB stream and natural-language instructions to control humanoid, quadruped, wheeled, and aerial robots, transferring across tasks, embodiments, and scenes without per-task or per-benchmark fine-tuning. The model extends the vocabulary with dual-channel pointing tokens and residual vector-quantized action tokens rather than adding navigation-specific prediction heads. At each step, it emits an affordance point representing a feasible local direction or waypoint and an object point representing the goal, followed by three action tokens that decode to ten future SE(2) waypoints for an embodiment-specific low-level controller. Its temporally aware history compressor samples older frames less frequently, pools them more coarsely, and preserves ordering with timestamp tokens while bounding the visual context. The repository includes a released checkpoint and action decoder, command-line prediction, a vLLM-based server with WebSocket streaming, a MuJoCo simulation demo, evaluation harnesses, and a ROS 2 deployment stack with adapters for the Unitree Go2 and LimX TRON 1. The project is released under the Apache License 2.0; the included EVT-Bench is separately licensed CC BY-NC-SA 4.0.
TrustMeBro is a Go command-interception tool for controlled red-team testing of coding agents such as Codex, Claude Code, and pi. It uses PATH shims to intercept configured command-line tools without requiring a plugin, hook, or MCP integration; rules can spoof generated or fixed output, rewrite stdout from a real command while preserving stderr and its exit status, pass through to the real binary, or reject the call. Each decision is recorded in a timestamped JSONL audit log. Rules match command names, domains, DNS record types, argument globs, and regular expressions. The bundled DNS generators support dig, nslookup, and host, while custom shims can use fixed output, exit codes, and standard streams. On Linux, lab mode uses Bubblewrap to shadow PATH lookups and discovered absolute paths so an agent cannot bypass interception merely by invoking a resolved system binary. TrustMeBro is distributed as a single MIT-licensed Go binary with installation, configuration validation, status, rule-listing, and uninstall commands. Lab mode requires Linux and Bubblewrap and is explicitly an interception namespace rather than a security sandbox: it reuses the host filesystem, workspace, network, environment, and agent credentials. Outside lab mode, absolute paths, changed PATH environments, and in-process DNS clients can bypass command shims.
HexStellar is an agent-first computational platform whose Cortex service executes structured optimization, decision, scientific-computing, and verification requests through a Python CLI and API. An agent submits a JSON formulation for problems such as QUBO/Ising optimization, maximum cut, traveling-salesperson routing, facility assignment, mixed-integer optimization, selection, ranking, scheduling, coloring, and business-rule feasibility; the managed service returns a structured result with execution metadata, a receipt, and an assurance label distinguishing certified optima, heuristics, operations, and abstentions. The CLI also provides free validation, estimation, service re-checks, local witness recomputation for supported families, reproducible seeds and model versions, batch and compressed-binary transport, and an MCP server over standard input/output, with read-only mode for free analysis and verification tools. The public package is a zero-dependency Python thin client: the proprietary solver runs on HexStellar-managed infrastructure rather than inside the package. A separate enterprise runtime is licensed for compatible customer-controlled compute paths under NDA; it has no public runtime download in version 1.0.
Editable Visual Design is an open-source, coding-agent-driven toolkit for creating editable visual artifacts from prompts. A persistent coding agent interprets a brief, plans a composition, generates assets, implements the design in semantic HTML, renders and observes the result, and applies repairs. An image model supplies visual direction for composition, hierarchy, color, and spatial relationships, while the delivered artifact rebuilds typography and layout in HTML rather than shipping reference pixels. The resulting artifact contains real text, independent assets, selectable and movable layers, a visual editor, rendered PNG output, an animated layer breakdown, and a replayable creation process. Deterministic checks cover the canvas, fonts, layer contracts, rendering, and editor round trips. The repository distributes the workflow as two Codex skills, including editable-design for fixed-canvas designs and a separate html-to-pptx skill for converting compatible HTML designs into editable presentations.
Magnitude is an open-source local inference server and CLI for running language models with AI agent harnesses. It profiles a machine's chip, memory, and bandwidth, recommends compatible models, downloads the selected models, tunes inference settings such as speculative decoding and concurrency, and runs models on demand. Models are loaded when needed and unloaded when idle or when memory is constrained; prompts, files, and models remain on the local machine, allowing offline operation after setup. It integrates with Pi, OpenCode, Hermes, OpenClaw, Codex, Claude Code, Oh My Pi, and Cline, and also provides a built-in harness. Magnitude supports macOS and Linux, with Windows support through WSL, and is licensed under Apache 2.0.
Chrome DevTools for agents (chrome-devtools-mcp) is an MCP server from the Chrome DevTools project that lets coding agents control and inspect a live Google Chrome or Chrome for Testing browser. It uses Puppeteer for browser automation, including navigation and actions that wait for results, and exposes DevTools capabilities for recording and inspecting performance traces and insights, analyzing network requests, taking screenshots, and inspecting browser console messages with source-mapped stack traces. A CLI is also provided for use without MCP, with configuration options including headless, isolated, slim, concurrent-session, persistent-profile, WebSocket, and Android-debugging modes. It requires Node.js LTS, a current stable Chrome version or newer, and npm. Browser contents are exposed to MCP clients; performance tools may query the Google CrUX API, and usage statistics and update checks are enabled by default with an opt-out flag.
Monocolor Editorial Print is an agent skill for generating one-ink or controlled two-ink editorial images for posters, zines, portraits, packaging, and visual field notes. It accepts a theme, phrase, object, article idea, or supplied photograph and produces a raster image when image-generation tools are available, the exact production prompt, and a short recipe describing the print mode, palette, layout, typography, process, and originality changes. The skill identifies the subject and intent, selects a layout family, assigns one or two ink plates, composes the page with visible paper and an asymmetric grid, and generates and inspects the result. Its visual system uses adaptive white, gray, or beige substrates; halftone, risograph grain, cyanotype exposure, or photocopy breakup; responsive typography; and controlled negative space. Without image generation, it returns the prompt and states the limitation.
Sierra is an AI agent platform for businesses. Its agents use tool calling to handle customer support, sales, and other pre-sale and post-sale workflows, with an outcome-based pricing model.
HyperFrames is an open-source framework for turning HTML, CSS, media, and seekable animations into deterministic MP4 videos. It can be used locally through a CLI, by AI coding agents through skills, or as a rendering core for hosted authoring workflows. Compositions are defined as HTML with data attributes for timing, tracks, media, and sub-compositions. The renderer seeks each frame in headless Chrome, uses animation adapters such as GSAP, CSS, Lottie, Three.js, Anime.js, or WAAPI, and encodes the result with FFmpeg; its toolchain includes preview, linting, inspection, local rendering, browser-based editing, reusable catalog components, and AWS Lambda rendering. The project requires Node.js 22 or newer and FFmpeg, does not require React or a build step, and is licensed under Apache 2.0. Its agent skills support workflows including product videos, explainers, pull-request videos, captions, motion graphics, slideshows, music-driven videos, and general compositions.
Ponytail is a rule set and plugin for AI coding agents that directs them to implement the smallest working solution. Its decision ladder skips speculative requirements, reuses code already in the repository, prefers the standard library and native platform features, avoids adding dependencies, and then chooses the shortest implementation that works. The project says that validation, error handling, security, and accessibility are not simplified away. It provides chat commands for changing enforcement intensity, reviewing the current diff for over-engineering, auditing an entire repository for bloat, recording deferred shortcuts in a debt ledger, and viewing benchmark results. It can be installed across multiple coding-agent environments, including Claude Code, Codex, Copilot CLI, Gemini CLI, and Pi; the page also lists support for other agents. The project reports lower code volume, token use, cost, and execution time in benchmarked Claude Code sessions editing a FastAPI and React repository, while noting that results vary by task and model.
DeskcommCRM is an open-source, self-hosted CRM and AI sales operating system for businesses that sell through WhatsApp and other chat channels. It combines a multi-tenant CRM and sales pipeline with inbox, Kanban pipelines, contacts, team governance, audit logs, LGPD features, webhooks, automations, scheduling, and human takeover capabilities, and supports WhatsApp through WAHA QR connections or the official Meta Cloud API. Its AI agents use tenant-specific RAG, organizational memory, executable skills, intent routing, sentiment analysis, follow-ups, and audited handoff to human attendants. Agents can operate CRM records such as leads and funnel stages, while the CRM exposes an internal MCP interface for agent operations. Tenants can receive leads through public webhook endpoints and process event-driven QUANDO/SE/ENTÃO rules through an event-log queue drained by scheduled workers. The application is built with Next.js, TypeScript, and Supabase/Postgres.
PR Lens is a pull-request visualization tool by Coldtea that generates animated architecture and data-flow diagrams and posts them inside GitHub pull requests. It maps a change's components and call paths, uses green for new elements, amber for changed elements, and red for removed elements, and provides step-by-step walkthroughs, nested diagrams, and interactive canvases with pan, zoom, and light or dark themes. It is available as a GitHub App, GitHub Action, CLI, or coding-agent skill. The CLI analyzes a diff against a merge base, produces a graph document, renders theme-paired SVGs, validates documents against a schema, and can compose the pull-request comment. The GitHub Action can use Gemini, OpenAI, or an OpenAI-compatible endpoint; the agent skill lets a coding agent generate and attach the diagrams. Repository configuration in .github/pr-lens.yml supports renames, exclusions, lane pins, and groupings. The repository documents Node 20.11+ and pnpm 10 requirements and is licensed under the MIT License.
BankMCP is a self-hosted, read-only Model Context Protocol server that lets AI assistants query a user's bank accounts through Enable Banking's PSD2 open-banking API. The user authenticates at the bank through a consent flow; BankMCP stores consent and account metadata but does not store balances or transactions, make payments, or send telemetry. It exposes MCP tools for managing bank connections and accounts, retrieving balances and paginated transactions, and creating rules that poll accounts for conditions such as low balances or matching payments and send webhook notifications. It can run locally or on a server, is distributed as the npm package `bankmcp`, supports standard MCP clients including Claude and Ollama, and is licensed under the MIT License.
NiubiGEO is an open-source, self-hosted tool for tracking how AI models describe products and brands, which competitors they name, and which sources they cite. It accepts a domain, queries selected OpenRouter models with optional web search, and stores each model's original answer, descriptions, competitor associations, keywords, citations, ordinary URLs, errors, and execution conditions. Users can confirm model-derived competitors and keywords for follow-up tests that omit the target brand's name, repeat measurements, and schedule monitoring runs. Results are organized into separate projects for different domains and retain historical records so findings can be traced back to the underlying answers and evidence. The project states that it does not track traditional search-engine rankings and that offline and web-enabled tests should be interpreted separately. The Community Edition is free, open source, and licensed under Apache-2.0. It runs locally with Node.js 22 or later, requires an OpenRouter API key for testing, and can also be deployed with Docker; users pay the model, search-service, and hosting costs they incur. The project also offers paid AI testing by real people and GEO optimization services.
OKF Agent Memory is a Git-native persistent memory layer for AI coding agents, developed as a pure Go CLI and library. It stores project knowledge as plain Markdown files with YAML front matter in a repository, following the vendor-neutral Open Knowledge Format (OKF) v0.2 rather than relying on an external vector database. The tooling parses and validates OKF bundles, builds an in-memory BM25 index for local lexical search, and provides progressive disclosure through hierarchical index files and link graphs. Its search-before-write workflow queries existing concepts before creating new ones; concepts can include provenance, trust tiers, lifecycle metadata, and code references for linking knowledge to source files. The standalone executable also supports creating and updating concepts, bootstrapping the memory structure into a project, and running an embedded Model Context Protocol (MCP) server over stdio for agent platforms. The repository reports sub-300-microsecond concept search and approximately 4-millisecond graph validation for its benchmark cases, with no external databases or API costs for retrieval. It is distributed under the MIT License.
Bot Crossing is a local web application that visualizes coding-agent threads as an astronaut colony. It reads session records from installed agent harnesses on the user's machine, maps repositories to persistent hex zones, and represents sessions as astronauts whose behavior reflects states such as running, waiting, errored, merged, archived, or inactive. Selecting an astronaut or repository opens its associated thread through the owning harness, while the application can also start sessions, reveal folders, mark threads viewed, and archive them. The application uses per-harness adapters to scan local session files and merge them into a common thread model. It currently documents support for Claude Code, Codex, and Cursor, with adapters for other harnesses able to be added under `server/harnesses/`. Astronaut navigation uses a rasterized navigation grid, A* routing, string-pulling, collision handling, and crowd separation. The browser renders the colony with Three.js techniques including instanced GPU-skinned astronauts, shader-based construction progress, procedural terrain and surfaces, and configurable planets, lighting, and quality settings. Bot Crossing runs locally through a Vite development server or a built Node server, binds to loopback by default, and stores the colony layout in `data/colony.json`. It does not upload session data or use an account; it reads harness records and writes only the colony state plus the documented archive field. The repository requires Node 22.13 or newer, supports macOS, Linux, and Windows, and is licensed under MIT. Bundled Kay Lousberg assets are separately covered by CC0, while bundled Material Design Icons use Apache-2.0.
anything2explainer is a Claude Code and Codex skill for producing narrated explainer videos from a topic or document. It outputs a 1280×720 H.264 MP4 in Chinese or English, with synchronized TTS voiceover, word-aligned subtitles, chapter cards, a top HUD, and a chapter progress bar. Every frame is drawn as code with Remotion, React, and TypeScript rather than generated video or stock footage. Its pipeline researches the subject and records sources, writes the narration, generates voiceover and a frame-accurate timeline, storyboards each shot, dispatches parallel agents to build Remotion components, renders the film, and runs quantitative frame metrics plus chapter-level quality-control and fix passes. The repository includes a compilable Remotion template, visual primitives, style and motion specifications, scripts for voiceover, storyboarding, rendering and QC, and a complete reference film with its research and production records. It is not a standalone CLI; it is a skill and production method intended for AI coding agents. The toolkit is licensed under PolyForm Noncommercial 1.0.0: noncommercial use is free, while commercial use requires prior authorization. The videos produced with it are described as belonging to their creators. It renders through headless Chromium on the CPU, supports landscape output rather than 9:16 video, and provides two backdrops within its specified visual style.
DeepSelect is a high-performance CUDA and PyTorch implementation of TopK kernels for DeepSeek Sparse Attention and sampling workloads. It replaces vanilla `torch.topk` for supported inputs, covering bfloat16 lightning-indexer workloads and float32 sampling workloads, with configurable sorting, index type, value output, variable-length rows, and NaN handling. The repository reports 2–20× speedups against `torch.topk` on its benchmarks. It is distributed as an installable Python package from the `deepseek-ai` GitHub repository. Supported TopK values are limited to 4096 or less, and input tensors must meet specified stride and contiguity requirements; the optimized sampling case targets vocabularies of approximately 128K.
It's Giving is an open-source Python webcam application that recognizes facial expressions, hand gestures, and body positions, then overlays matching meme images or animated GIFs on the user's head. It uses MediaPipe face landmarks and blendshapes, hand points, body landmarks, face-relative measurements, tongue color, and hand speed; its calibrated version records seven seconds of the user's neutral expression and evaluates later expressions as deviations from that baseline. An ordered pose-matching pass requires gestures to persist for a configured number of frames before triggering, and the overlay follows the user's face with alpha compositing and animation support. The application publishes the result through a virtual camera for Zoom, Meet, Teams, Discord, OBS, and other webcam-compatible software. It includes fourteen predefined reactions, preview-only and virtual-camera modes, keyboard testing and recalibration controls, and an assets directory where users can replace or add reaction images. It runs with Python 3.11 or 3.12 and requires a virtual-camera backend such as OBS Studio or v4l2loopback, depending on the operating system.
WeKnora is an open-source, LLM-powered knowledge platform from Tencent for turning documents into queryable knowledge bases, retrieval-augmented Q&A, autonomous reasoning workflows, and self-maintaining Markdown wikis. It supports RAG-based quick Q&A and a ReAct agent that orchestrates retrieval, MCP tools, sandbox skills, and web search for multi-step tasks. Its Wiki Mode distills source documents into structured, interlinked Markdown pages with an interactive knowledge graph, manual editing, revision history, line-level diffs, and rollback. The platform ingests formats including PDF, Word, Markdown, HTML, images, spreadsheets, presentations, and XMind, and can synchronize sources such as Feishu, GitLab, Tencent IMA, Notion, Yuque, and RSS. Its modular pipeline supports interchangeable parsers, LLMs, embedding providers, vector databases, and storage backends, with dense, sparse, hybrid, parent-child, and graph-based retrieval strategies. WeKnora provides a web interface, REST API, command-line client, MCP server, website embed widget, and integrations with messaging channels. It can be deployed locally, with Docker, or on Kubernetes, including private and offline deployments. The repository states that it is licensed under the MIT License and supports workspace RBAC, scoped API keys, audit logs, and Langfuse-based observability.
ARTEMIS is an open-source Google tool for natural-language Android automation and testing. It drives real phones or emulators through end-to-end workflows, generates test plans, captures screenshots and logs, profiles performance, and produces structured reports. Its element-locating system combines accessibility hierarchies and OCR with visual-model and coordinate fallbacks for custom interfaces. ARTEMIS provides a CLI, web visual console, Python SDK, and native Model Context Protocol (MCP) server for connecting Android devices to AI coding assistants such as Antigravity, Claude Code, Codex, and Windsurf. Its Flash profile uses a reactive observe-and-act loop with compressed history, while the Pro profile uses planning, pre-execution checks, action bursts, incident recovery, checkpoint verification, and optional final reports. The project also includes an accessibility helper for reading screen layouts and can fall back to UIAutomator2. It is distributed under the Apache License 2.0; the repository reports 99%+ completion on Google Research's AndroidWorld benchmark.
SoL-Pi is an open-source standalone extension for the Pi coding agent, developed and maintained by NVIDIA. It packages four opt-in efficiency mechanisms: Action Fusion runs a follow-up validation command in the same edit or write call; ObservationPack replaces repeated large tool results with stable handles that support exact paged recall; the Evidence-Preserving Reducer condenses diagnostic logs into receipts only when retained quotations match archived source material; and Online Context Compact marks completed plan steps for Pi's native compaction when economic and context-window checks allow it. The extension operates through Pi's public APIs without patching or vendoring Pi, leaves the mechanisms disabled by default, and preserves original observations locally. It requires Node.js 22.19 or newer and is released under the MIT License.
An open-source Codex configuration that uses Astra as the root orchestrator and independent reviewer, with Luna-based subagents for exploration, implementation, testing, and research. It provides separate Pro and Plus profiles, TOML role configurations, an Astra orchestrator skill, project-level AGENTS.md instructions, and shell and PowerShell installers that copy the selected configuration into a target repository. The profiles set model, reasoning-effort, approval, sandbox, and concurrent-thread settings, while named role files can override the defaults. The repository also includes guides for orchestration patterns, plan-specific setup, iteration, and token-usage reporting from Codex session logs. It is licensed under Apache License 2.0.
Zetta is an open-source closed-loop embodied harness for self-evolving physical intelligence. It keeps a base policy frozen while evolving code-based runtime critics and recovery skills when rollouts fail. Its evolution protocol clusters failed trajectories, uses a diagnostic stage to identify an observable causal failure mechanism, generates Critic-Recovery candidates under fixed tool schemas, and checks them through shadow replay, paired same-seed evaluation, and held-out seeds. Candidates are rejected or promoted based on these gates; the critic may propose recovery actions, a higher-level role approves them, and a bounded recovery actor executes only approved programs. The repository provides campaign management, rollout workers and environment backends for systems including LIBERO, RoboCasa, RoboTwin, and Genie Sim, along with manifests, tool catalogs, audit records, deployment helpers, and contract tests. It supports Python 3.10–3.12, while simulator assets, model weights, credentials, and host runtime files remain outside the repository.
Birdview is an open-source developer tool that makes AI coding agents map a project's architecture before editing. Its Birdview workflow defines modules, responsibilities, file ownership, evidence, relationships, and layout in an architecture JSON file; an agent then declares planned task scope, targets, files, lifecycle phases, and verification records against that map in an activity JSONL stream. Validators check the schemas and cross-record rules before a renderer produces a standalone interactive HTML view with architecture, changes, comparison, module evidence, and activity-history views. Birdview supports automatic or on-demand activation through project-level agent guidance and includes Node.js scripts, JSON Schemas, example maps, and tests. The generated output has no server or network dependency, but the v0.1 implementation is file-based: activity is agent-declared rather than automatically observed, and updates require regenerating the HTML. It is MIT-licensed and is not published to npm.
Coder is a self-hosted platform for cloud development environments and AI coding agents. It defines workspaces with Terraform and can provision them on EC2 virtual machines, Kubernetes pods, Docker containers, and other infrastructure, connecting users through a secure WireGuard tunnel and automatically shutting down idle resources. Coder Agents runs a native AI coding-agent loop in the control plane on the operator's infrastructure. It supports models from Anthropic, OpenAI, Google, Amazon Bedrock, and self-hosted providers without placing LLM credentials in workspaces, while providing centralized model governance, cost tracking, identity on actions, and audit logging. Workspaces can be accessed through existing IDEs including VS Code and JetBrains products.
Jevlike is an open-source research starter for training small models that choose among a variable list of text options. Given a context and an option list, it represents each option as a query, attends to the context tokens to produce an option-specific context vector, applies a shared dot product, and runs a softmax across the options to return probabilities in one pass rather than generating an answer token by token. It provides a trainable byte-level encoder and an optional frozen Hugging Face encoder, with command-line tools for generating synthetic data, training, evaluation, and prediction from JSONL records. The repository also includes a visual option scorer used by Doom and chess controller examples. Training supports CPU, Apple MPS, and CUDA; the default encoder truncates contexts to 192 bytes and options to 32 bytes, and the project requires the complete option list at prediction time. Jevlike is an independent implementation rather than a copy of TypeSafe's commercial Jev model, and its code is released under the MIT License.
@shadcn/lint is an agent-first linter for Tailwind v4 design systems. It lets teams define component and theme rules that coding agents can verify through ESLint or Oxlint, reporting violations with suggested fixes drawn from existing components, variants, and theme values. Its rules include preventing component restyling, raw colors, arbitrary values, inline styles, unknown Tailwind classes, and dynamic classes the linter cannot read; custom contracts, messages, placeholders, component-import settings, and shared configuration are supported. It can be used with custom Tailwind components and does not require shadcn/ui. The package requires Node.js 20.19 or later, with ESLint 9.30 or later or Oxlint 1.80 or later, and is licensed under MIT.
Pizza Bot is a local-first inbox for long-running AI-agent work, developed at Amazon. It lets users start or schedule tasks, switch conversations, and return to completed work in an Unread queue or approval requests in an Action queue. Its stateful DeepAgents/LangGraph runtime checkpoints runs so they can continue after a client disconnects; cron and webhook triggers can also start work. Electron, browser, and terminal clients communicate with an api-server over HTTP/SSE, while skills provide tool-scoped subagents and MCP servers extend the system. The project supports model providers including Amazon Bedrock, Anthropic, Google Gemini, OpenAI, OpenRouter, and Ollama. It stores application state locally, requires explicit folder permissions for local file access, and is distributed under the Apache 2.0 license with desktop installers for macOS, Windows, and Linux.
AgentVerse OS is a personal cloud operating system for a developer and their AI agents, installed on an Ubuntu server and accessed through a browser from a laptop, tablet, or phone. It provides a windowed desktop, isolated project workspaces, browser-based VS Code and a terminal, and supports Claude Code and Codex agents whose work remains on the server. Each project runs in an isolated Incus container with Docker inside, its own network and access gate. A Rust core coordinates projects, workspaces, capabilities, applications, backups, and updates; Coder controls the workspaces, Komodo deploys application stacks, Caddy provides the entry points, and Tailscale supplies private access and certificates. Projects request capabilities such as S3 storage, an LLM gateway, or notifications, which the core connects to provider networks and exposes through environment variables. The system includes a store combining catalogs from Runtipi, Coolify, and Umbrel, with guided installation, logs, snapshots, rollback, and application updates. It also supports ZFS snapshots, restic backups, an embedded Svelte desktop PWA, and one-command installation from a release package or source. The project is in alpha, is designed for one person, has no user accounts or permissions yet, and requires a clean Ubuntu installation and a Tailscale account. The repository is licensed under Apache-2.0; store applications retain their own licenses.
AuK is an open-source 1.5-billion-parameter foundation model for speech generation and editing, developed by Tencent Hunyuan. It exposes zero-shot and instruction-based text-to-speech, content and lyric rewriting, pitch, speed, volume, emotion, timbre, accent, whisper and nonverbal-sound editing, speech enhancement, speaker separation, music separation, and target-speaker extraction through a unified natural-language instruction interface. Inputs can include text instructions and optional source or reference audio; outputs are generated or edited audio. The repository provides an AuK base model for configurable high-quality generation and AuK-Flash, a distilled variant using four-step inference, along with command-line, Python, Gradio, ComfyUI, and fine-tuning interfaces. The code is released under the MIT License, and model weights are distributed through Hugging Face and ModelScope.
Video Gen Client is an asynchronous Python client for video-generation APIs from Seedance, Kling, MiniMax/Hailuo, and Wan. It provides a shared interface for text-to-video and image-to-video generation, with provider and model discovery, task-status polling, and MP4 output. The project includes a command-line interface for listing providers and models, submitting generation jobs, checking task status, and serving the application. Its bundled FastAPI web UI supports provider and model selection, prompts, negative prompts, duration, aspect ratio, resolution, live task polling, video preview, MP4 download, and in-memory task history; it is served from the Python package without an npm, Node.js, or frontend build step. Provider adapters depend on external APIs and model IDs that may change independently of the client and therefore may require updates. The project is licensed under the MIT License.
Fugleramme is a self-hosted bird-display application for Raspberry Pi and web kiosks. It uses BirdNET-Go to detect birds from microphone audio, polls its API for detected species, matches each species to a curated public-domain illustration, and arranges the cut-outs on a textured page with larger birds toward the centre and sizes based on body mass. The display redraws when detections change and can run on a Pimoroni Inky Impression e-ink panel or any network-connected screen; an admin page configures what is shown. The project includes more than 800 illustrations covering over 400 species, with strongest coverage in Scandinavia, the British Isles, and Germany. It can use an existing BirdNET-Go installation, run in a container, or be installed as a systemd service on a Raspberry Pi. Fugleramme's code is MIT-licensed; BirdNET-Go, detection data, artwork, fonts, and taxonomy assets have separate licenses and attribution requirements.
H3 Max is fal’s post-trained video-generation model, based on MiniMax’s open-weight video model. It combines model post-training with systems and hardware optimization to reduce generation time while maintaining video quality, supporting fast generation, interactive streams, and professional video workflows. Its continuous-video experiments can retain previous scene context and respond to new directions while a stream is running, with development focused on controllable camera movement, lighting, characters, motion, and lip sync.
Databricks is a unified platform for data, analytics, and AI. It supports ETL, data warehousing, governance, and AI workflows through a data-centric platform. The videos describe its use for security detection, organizational ontologies, agents, and internal enterprise AI workflows.
SemIf is an independent open-source tool, formerly called OpenJev, for making runtime-defined semantic decisions with open language models. It accepts unstructured state, runtime-defined questions or criteria, and typed options, then reads the model’s native option logits to return conditional probabilities rather than generating an answer sentence, JSON, or running a decoding or parsing loop. It supports direct, serial, and shared-state execution, prefilling common state for evaluation of multiple criteria; results include option scores, timing, the exact model revision, and a prompt hash. It supports local GPU and CPU workflows, Apple Silicon MPS and MLX, and llama.cpp backends, and provides a browser-based WebGPU demo, reproducible fixtures and benchmark runners, calibration methods, and quality evaluations. The project is an independent reproduction of the interface pattern of TypeSafe’s closed Jev service, not its model or training, and is not affiliated with Jev or TypeSafe. The code is released under the MIT License.
typesafe-computer-use is an open-source macOS computer-use tool that drives a Mac toward a plain-English goal. It captures the frontmost window, reads text with OCR, collects labelled controls from the macOS accessibility tree, parses dates, and sends a structured list of possible actions to the TypeSafe decision model instead of sending screenshots to a frontier model for every step. The loop chooses actions such as clicking, scrolling, browser navigation, waiting, or pressing keys; a writing model is called only for free-text entry, proposed URLs, and the final answer, while passwords are never typed. It supports dry runs, confidence and step limits, annotated captures, detailed action and probability logs, and offline replay of saved screenshots. The repository targets macOS 14 or newer with Python 3.12 or newer, requires a TypeSafe API key, and is licensed under the MIT License.
Splash is a local inference engine for Apple silicon developed by Inco AI. It serves selected language models to coding agents and clients compatible with the OpenAI Chat Completions, OpenAI Responses, and Anthropic Messages APIs, with streaming, tool calls, JSON Schema output, image input, and inline PDF support. Its runtime uses model-specific fused Metal kernels, weights packed for those kernels, a dedicated DFlash 2 draft model for speculative decoding, and a startup-computed memory plan for context, KV cache, and batching. The served model packages include the target model and its draft model; plain MLX or Transformers checkpoints are not supported. Splash runs locally on macOS, binds to a local HTTP server by default, and can be installed through Homebrew. The project is licensed under Apache-2.0, while model weights retain their own licenses.
Hindsight is an agent memory system developed by Vectorize.io for storing, retrieving, and synthesizing long-term memories for AI agents. It organizes memories into world facts, experiences, observations, and mental models within isolated memory banks, with support for multilingual data and optional per-bank scanning for secrets and personally identifiable information. Its retain operation uses an LLM to extract facts, entities, relationships, and temporal data, then normalizes them into searchable representations. Recall combines semantic vector search, BM25 keyword matching, entity/temporal/causal graph links, and time-range filtering, merging results with reciprocal-rank fusion and cross-encoder reranking. Reflect performs deeper analysis of stored memories to form connections and answer questions requiring more than retrieval. Mental models and knowledge pages provide continuously updated answers or documents derived from a bank's memories. Hindsight can run as a Docker service, Python package, embedded database, Kubernetes deployment, or hosted Hindsight Cloud service. It provides Python, Node.js, Go, CLI, and REST clients, an OpenAI-compatible LLM wrapper, integrations for agent frameworks and coding agents, and a built-in Model Context Protocol server. Self-hosted storage uses PostgreSQL with pgvector or Oracle AI Database 23ai; the project is MIT-licensed.
mini-AGI is an experimental continual-learning, byte-level language model that trains from scratch on a CUDA-capable GPU with at least 8 GB of VRAM. It reads 256 byte values directly rather than using a conventional tokenizer, processes data one chunk at a time, and uses the same forward path for training and generation. Its architecture combines dense prelude blocks with a recurrent block applied up to 24 times, adaptive PonderNet-style halting, and per-application top-8 routing through a dynamically growing and pruning expert pool. Expert weights and Adam moments are stored as files on disk; a working set is paged through RAM and VRAM as needed, allowing the pool size to exceed available VRAM. The project includes Python commands for building corpora, reading files, streaming training, serving a local web interface, and inspecting or replicating continual-learning experiments. The repository describes the current model as toy-level rather than frontier-capable, and its weights are generated in the local weights directory rather than distributed with the source.
OpenMuse is a personal-agent application built by CopilotKit for iOS, Android, and web. It combines a persistent Chromium browser with an optional isolated Linux container, file and PDF handling, visible task plans, action reviews, and browser or terminal takeover so a person can inspect or continue the agent's work. It also provides Gmail and Google Calendar adapters with review required for sends and calendar changes. The application runs an API, durable task worker, and browser worker. Tasks use stored plans, progress, approvals, pauses, retries, and SQL leases; the browser worker maintains persistent Chromium profiles, while the optional non-root Linux container provides a retained workspace volume with bounded commands and no host-directory mounts or credentials. The interface is built with CopilotKit headless chat and AG-UI events, with inline email, browser, PDF, plan, and finance results. The repository describes OpenMuse as an alpha for self-hosting and building on, with model and Google-account configuration required for live agent and mail/calendar workflows. It is MIT licensed; CopilotKit Intelligence is a separate required service for rich conversation persistence and is not included under the repository's license.
Unreal Agent is an async-first agent harness from Unreal Labs. Its library and command-line executables coordinate persisted agent sessions, LLM turns, tool calls, and asynchronous operations. Inputs carry caller-supplied globally unique IDs for deduplication; sessions store append-only history that can be recovered or forked; and tool translators validate model-produced calls and convert them into serializable operations whose execution state is tracked separately. The harness includes a session inbox, coordinator, session store, context builder, LLM adapter, tool registry, tool translators, and an operation manager. Session and operation data are versioned and serializable, allowing operations to be dispatched to alternative or remote runtimes; unsupported session versions produce an explicit resume error. The repository contains the harness library, executables, and benchmark runners.
Z.ai's coding agent harness and AI programming workspace, offered through a desktop application, browser interface, and terminal agent. Its repository contains the client, backend services, shared React UI, Agent CLI, and runtime; the desktop client uses Electron, while the web and server components communicate through HTTP and WebSocket services. It supports remote project connections over SSH and WSL. The standalone command-line distribution combines a terminal UI, web server, and agent, and can run locally without Electron; the same command starts either the terminal interface or a browser-accessible web interface.
Company Brain is an open-source Slack teammate developed by supermemory. It remembers decisions, projects, owners, and other context from the Slack channels and connected tools it can access, answers questions from that information, and can act through tools such as GitHub, Linear, Notion, and Google Workspace via MCP. It can open issues, read pull requests, run repository scripts in a sandbox, return charts, CSVs, and PDFs in Slack, speak up when relevant, and send scheduled digests or perform research. Its memory is organized as a permissions graph: public-channel queries use the shared organizational memory, private-channel queries add that channel's memory, and direct messages can use the asker's personal memory and private channels. Tool writes run through the user's own connection, while access requiring another teammate's connection requires that teammate's approval. The project runs on a user's Cloudflare account using Workers, Durable Objects, D1, KV, and supermemory for memory; it can be deployed through a setup flow or run locally. The repository states that the formerly paid product was discontinued and released under the Apache 2.0 license.
CLM is an AI decision-making model and Python service that selects among candidate actions by embedding a state and each action separately, then scoring their alignment instead of generating a textual response. It is trained with bidirectional InfoNCE: matched state-action pairs are pulled together while incorrect or hard-negative actions are pushed apart. At deployment, cached state and action embeddings are scored with a projection head, enabling typed decisions such as binary judgments, choices, and ordered scores, as well as ranking candidate answers, tools, trajectories, or next moves. The repository provides CLM-8B through a TypeSafe-compatible HTTP API and a local Python engine. It runs a Qwen3-8B pooling encoder through vLLM, exposes endpoints including `/v1/systemone` and `/v1/rank`, and includes a browser playground, embedding and action caches, checkpoint hot-reloading, and scripts for fine-tuning projection heads on typed decisions or agent trajectories. The reference head can run on CPU or GPU, while the encoder is served separately; states longer than 2,048 tokens are truncated by default unless the limits are raised together. The project is released under the Apache 2.0 License, and the repository's CLM-8B weights are also released under Apache 2.0 on Hugging Face.
Ollaya is an open-source system for running decision models locally. It accepts a message, email, ticket, or other JSON state together with typed questions and returns calibrated probabilities for choices, scores, or binary decisions rather than generated text. The Ollaya daemon pulls models by name and serves them through a CLI and HTTP APIs, including TypeSafe-compatible `/v1/systemone`, `/v1/decisions`, and `/v1/models` endpoints. Its native API adds routing information and timings, while model definitions can embed reusable question sets. It supports ONNX Runtime on CPU and NVIDIA GPUs, and llama.cpp for GGUF models on CPU, CUDA, and Apple silicon. Ollaya also provides an MCP server for clients such as Claude Code, Claude Desktop, and Cursor. It publishes small ONNX graphs while obtaining the original model weights from their authors' Hugging Face repositories, pinned and verified by commit and SHA-256; it does not re-host those weights. The project is distributed under Apache-2.0, while individual models retain their own licenses, and it is available as a CLI, daemon, desktop application, and Docker image.
Directify is an AI-powered directory website builder that generates a complete directory site (listings, design, and basic SEO) from a user description in about 60 seconds. It includes monetization features and positions itself as allowing creators to retain 100% of revenue. The service is offered via its website.
An AI-powered scam-intelligence system that responds to scam emails as though it were a victim, eliciting information about scammers’ tactics and infrastructure for security teams, threat-intelligence feeds, and law enforcement.
QMD is an on-device command-line search engine for Markdown notes, meeting transcripts, documentation, knowledge bases, and other organized information. It combines BM25 full-text search, vector semantic search, and LLM reranking locally through node-llama-cpp and GGUF models. Its hybrid query pipeline expands the original query into lexical, vector, and HyDE variants, routes them to keyword or vector search, combines the results with Reciprocal Rank Fusion, and applies an LLM reranker to produce final ranked results. Collections can be indexed and enriched with hierarchical context; documents can be retrieved individually or in batches, and JSON and file-oriented output formats support agent workflows. QMD also exposes query, retrieval, batch retrieval, and index-status operations through an MCP server.
Duet is an AI agent-development tool that generates agent operating procedures, integrations, tools, tests, and simulations from customer transcripts and documentation.
An AI agent for reviewing large volumes of production conversations, identifying trends and weak topics, creating model variants, and drafting improvements for a primary conversational agent. It is associated with Decagon's conversational AI platform.
An AI capability in IBM Db2 that enables semantic queries over data stored in relational databases through SQL, using large language models to derive insights from database content.
Google Lens is an image recognition and visual search tool developed by Google that uses machine learning to identify objects, landmarks, text, barcodes and other elements in images. It can translate and copy text from photos, perform product and image searches, and provide contextual information about recognized items; it is integrated into the Google Camera app, Google Photos, Google Assistant, and the Google app on iOS.
PredictLeads provides company intelligence data designed for use in AI applications, offering structured company datasets and signals for prospecting and modeling.
LlamaIndex is an open-source data framework and developer platform (originally GPT-Index) that provides connectors, indexing structures, and retrieval utilities to connect documents and external data to large language models. It is maintained by the LlamaIndex team and is used to build LLM-powered pipelines, agents, and document workflows.
AutoGen is an open-source framework from Microsoft for building and orchestrating agent-based workflows that use large language models. It provides components for defining agents, managing conversations and interactions, and integrating models and external tools.
BabyAGI is an experimental framework for building a self-building autonomous agent. Its core, the functionz framework, stores, manages, and executes functions from a database using a graph structure that tracks imports, dependent functions, and authentication secrets; it also supports automatic function loading and logging. A dashboard provides function management, updates, execution, and log viewing. The repository can be installed as a Python package and is intended for experimentation rather than production use. The original repository was archived in September 2024 and moved to a separate archive repository.
CrewAI is an open-source Python framework for building multi-agent workflows. Its Crews use autonomous, role-based agents that collaborate on tasks, while its Flows provide event-driven workflow control combining precise orchestration, individual LLM calls, and Crews. The project also offers CrewAI AMP Suite, a commercial control plane for managed deployment, observability, governance, security, enterprise support, and cloud or on-premise operation. The repository provides high-level abstractions and lower-level APIs, along with documentation, examples, installation guidance, and integrations for connecting agents to language models.