1,038 tools and products — trending open source, and what gets used in AI and other work.
Camera to Blender is an open-source workflow that turns a phone or webcam photo of a real object into a 3D model and imports it into Blender. Its pipeline sends the image through optional Gemini background removal, submits it to the Tripo3D API for model generation, and relays the result to Blender through a WebSocket add-on. The project includes a phone-oriented web camera app, a Python relay server, and a Blender add-on. The server runs on the computer with Blender; phone access uses an HTTPS ngrok URL, while laptop webcam use can run locally. It requires Python 3.10+, Blender 3.0+, and a Tripo3D API key; a Gemini API key is optional. The project is distributed under the MIT license.
agent-memory is a local-first, agent-agnostic long-term memory runtime for AI agents. It stores memories as plain Markdown files, with a rebuildable SQLite index used as a cache rather than the source of truth. Conversation-boundary hooks trigger writes, while an independent sleep-time Manage pass consolidates, ages, and forgets memories by value; unattended deletion is limited to proposals that require confirmation, and superseded or archived material remains available. The runtime provides three recall paths: deterministic MEMORY.md injection at session start, BM25 search with progressive disclosure, and direct access through the store's directory tree using tools such as ls and grep. Memory records use frontmatter for names, abstracts, status, timestamps, links, weights, and provenance, with free-form Markdown bodies. Its CLI, MCP server, and host-agent hooks share the same validation, hash-diff, and reindexing path, and it can be used by agents that run shell commands, including Claude Code and Codex CLI. The repository documents a Python 3.12-or-later setup using uv, commands for initialization, recording, recall, rebuilding, sleep-time consolidation, and proposal approval, plus host setup for Claude Code and Codex. It does not include an LLM client; reasoning is delegated to the host agent's CLI, so the library requires no separate model keys or billing surface.
OrcaReplay is an open-source replay and debugging tool for AI-agent executions, built by the OrcaRouter.ai team. It records an agent run as a local trace, including model requests and responses, tool calls and results, shell exit codes and timing, filesystem snapshots, MCP traffic, and selected network activity. It runs as a local proxy and capture layer around an unmodified agent. A proxy records model traffic, while PATH shims, JSON-RPC interception, filesystem snapshots, fetch hooks, and optional TLS interception capture effects that model transcripts cannot show. The resulting timeline can be viewed, exported, queried through JSON or MCP, and analyzed with recorded versus inferred causal edges. The same recorded stream supports offline replay with the network blocked, forking from a derived checkpoint onto another model, and comparison of several models with an optional verification command. Replay uses the recorded conversation and workspace state; interactive terminal sessions can be approximate because prompts and interactive-only tools may not exist on the wire. The CLI is installed with npm, requires Node 20 or newer, stores runs under `.orca/runs/`, and the code, CLI, viewer, adapters, and trace format implementation are released under Apache-2.0; the trace specification is CC BY 4.0. Traces are local and redacted on write, but the project describes them as sensitive and treats redaction as best-effort.
Fable51 Worlds is an open-source, AI-assisted world-generation project that turns a text brief, photograph, or video into a walkable browser scene implemented as a pure Three.js application. Claude Fable 5.1 agent swarms research the location, collecting map geometry, elevation, transit, street information, and storefront data; generation scripts produce procedural façades, street furniture, vehicles, vegetation, fixtures, and other assets; and the runtime assembles terrain, streets, buildings, props, crowds, and traffic from JSON specifications without a game engine, proprietary 3D tiles, or downloaded meshes. Verification uses Playwright to drive the application, capture fixed viewpoints, and compare them with photographs from corresponding locations. Independent reviewer agents covering architecture, geography, technical art, and interaction produce reports for subsequent fix cycles, and each world includes a quality-assurance report.
Tripo AI is an online AI 3D-model generator that converts text prompts, photographs, images, and sketches into 3D models. The video presents it in a camera-to-Blender workflow for generating a model from a photograph. Its generation modes include detailed geometry and production-oriented quad meshes. The service also provides automatic part segmentation for 3D printing, exports STL, 3MF, and OBJ files, automatic character and creature rigging, and PBR texture generation.
BoardUI is a React design system for agentic interfaces and dashboards, combining AI-product components such as chat, thinking indicators, streaming agent logs, task lists, web-search trails, composers, and an agent runtime with dashboard components such as tables, charts, forms, cards, navigation, authentication, and design tokens. Its CLI copies individual components and their dependencies, or the complete free catalog, into a Next.js project as source files rather than installing a runtime dependency, so developers can modify the code directly. The system uses React Aria Components for accessible controls, Tailwind CSS v4, semantic design tokens, light and dark theme classes without runtime CSS, and Figma-based typography and styling rules. The repository includes a working AI chat application with a streaming chat endpoint and an agent runtime that accepts provider keys for supported services or an OpenAI-compatible server, reads keys server-side, and streams replies to the chat UI. It also provides an MCP server for browsing and installing components, along with agent rules specifying design tokens, type scale, and conventions.
Kitter is a local-first skill manager for agent skills across project folders. It keeps one canonical copy of each skill in a shared library, links skills into individual projects according to their needs, reports which skills are loaded, and estimates their context cost.
Higgsfield for Blender is an AI add-on that brings Higgsfield's generation tools into Blender, generating and importing 3D scenes, meshes, rigs, textures, images, and video. It adds a floating viewport bar with tabs for scene building, 3D models, character animation, images, video, cameras, and assets; results can be inserted into the open scene as editable geometry, planes, materials, rigs, keyframes, or video outputs. Scene Builder creates editable objects, layouts, and lighting from a prompt. The character-animation workflow produces a fitted weighted rig with timeline keyframes. An optional Higgsfield Bridge MCP endpoint at https://bridge.higgsfield.ai/mcp connects an external AI agent such as Claude to the add-on for operations including building blockouts, generating meshes at the 3D cursor, and requesting Seedance renders in the open Blender scene. Generation runs on Higgsfield's servers, requires an internet connection and a signed-in Higgsfield account, uses the account's existing Higgsfield credits, and supports Blender 4.2 through 5.1.
MarkItDown is a Microsoft Python utility and command-line tool that converts files and office documents into Markdown for use with large language models and text-analysis pipelines. It aims to preserve document structure such as headings, lists, tables, links, and other content rather than produce high-fidelity human-facing document conversions. Its built-in converters handle formats including PDF, PowerPoint, Word, Excel, images, audio, HTML, CSV, JSON, XML, ZIP files, YouTube URLs, and EPUBs. It can be used through a Python API, command line, Docker, or third-party plugins; optional integrations support OCR with an LLM vision client, Azure Document Intelligence, and Azure Content Understanding, including multimodal conversion and YAML front-matter field extraction. The repository notes that conversion performs I/O with the current process's privileges, so untrusted inputs must be sanitized and the narrowest conversion method should be used.
Context Mode is an MCP server and plugin for AI coding agents that reduces context-window usage by routing large tool outputs through sandboxed subprocesses. Its execution tools run code in supported languages, process files, fetch and analyze URLs, and return selected stdout or search results instead of exposing raw logs, snapshots, API responses, or file contents to the conversation. The project reports up to 98% context reduction in its benchmarks. It also provides session continuity across supported agents. Hooks capture tool calls, edits, prompts, decisions, errors, and other session events in SQLite; before compaction, the system builds a prioritized resume snapshot, and after compaction or session resumption it retrieves relevant events through SQLite FTS5 search with BM25 ranking. Content indexing uses heading-aware chunking, stemming, trigram matching, reciprocal-rank fusion, proximity reranking, fuzzy correction, and smart snippets. The MCP interface includes execution, batching, indexing, search, fetching, statistics, diagnostics, upgrade, and purge tools. Context Mode supports multiple coding-agent platforms through MCP servers, native plugins, and platform-specific hooks, with automatic routing enforcement where hooks are available and instruction files for platforms without them. The repository states that processing and SQLite storage are local, with no account or telemetry requirement. It is licensed under the Elastic License 2.0, which permits use, modification, and distribution but restricts offering the software as a hosted or managed service and removing licensing notices.
Town is a personal AI assistant that manages inbox tasks, schedules meetings, and handles other work in the background. Its assistants, called Townies, learn from a person's existing work and maintain a wiki describing how they work, including their clients, deadlines, and writing style. Townies can act on to-do lists, draft emails and proposals, prepare briefings, and return updates or request input when a task requires a decision.
9router is a local gateway between AI applications and model providers. It exposes one OpenAI-compatible endpoint, supports provider fallback tiers, tracks usage, and compresses tool outputs.
Diffy is an AI workflow tool with a visual canvas for building workflows that retrieve compatible database records, generate explanations for the matches, and expose the resulting workflow through an API.
OpenHands is a self-hosted platform for running autonomous coding agents that work on GitHub issues. It can use either remote model providers or language models installed locally.
Clipnote is a notepad for saving and organizing content generated in AI conversations. It stores plain text, Markdown, and HTML clips, preserving HTML layout and interactivity, and lets users search, pin, archive, version, restore, and delete them. Clips can be private or shared through public URLs, and collections bundle multiple clips into one shareable page. Through MCP connections with Claude and ChatGPT, an assistant can save a new clip, retrieve a clip's content to inform a response, or update an existing clip. Updates automatically preserve the previous version. Clips can also be created or updated manually by pasting text or uploading a file.
AIRUNCODE is a local-first agent runtime for running coding agents on a user's own computer. It supports voice-driven code-agent tasks and access to AI providers at their origin pricing, while its site describes execution as local.
Routines by Databox is a scheduling feature for recurring analytics and reporting. Users save an analysis as a reusable Skill, schedule that Skill to run daily, weekly, or monthly, and receive the results by email, Slack, or in Databox. Databox describes Routines as running Skills automatically and delivering the results, alongside AI agents that can delegate broader analysis-to-action workflows under user oversight.
AI Toolbox is a Chrome extension, formerly called ChatGPT Toolbox, that adds folders, subfolders, full-text search, prompt management, conversation export, and other organizational features to ChatGPT, Gemini, Claude, and Grok. Its cross-platform search indexes conversations locally, groups results by AI service, and links directly to the matching chat. It also supports saved prompts, a curated prompt library, sequential prompt chains of up to 10 steps, bulk conversation actions, smart tags, bookmarks, audio downloads, and exports in text, Markdown, JSON, PDF, or ZIP formats. The extension runs on Chrome, Edge, Brave, and other Chromium browsers; the site describes conversation content as remaining on the user's device, while folders, prompts, and tags can sync in encrypted form. A free plan is available, with paid plans and team features.
Tadata is an AI employee for Slack that handles research and repetitive sales and go-to-market operations work. It can prepare call briefings, research prospects and companies, draft outreach and follow-up emails, fill CRM fields, monitor relevant LinkedIn activity, and suggest recurring automations. It connects to tools such as HubSpot, Attio, Notion, Linear, GitHub, Google Sheets, Gmail, and Google Calendar, as well as external web sources and MCP or API integrations. Tadata runs pre-built agents or builds an agent from a described process. It observes repeated work, learns team preferences, and requests approval before automating tasks or sending messages; connected-tool access is scoped to the user's authorization. The service is presented as model-agnostic, with portable automations, preferences, exceptions, and agent behaviors that can be exported and versioned.
Speakeasy is an AI control plane for discovering, securing, and governing enterprise AI agents, MCP servers, Skills, and AI applications. It maintains a catalog of approved AI tools, gives each agent an identity through the organization's existing identity provider, and scopes access by team and role. Every prompt, response, and tool call passes through the control plane for real-time policy checks before reaching internal APIs, MCP servers, or SaaS tools. Policies can distinguish read and write access and individual tools; violations are blocked, while allow-or-deny decisions are recorded in an audit trail. The platform is designed to detect and quarantine unapproved MCP connections and to block threats such as prompt injection, PII exposure, and leaked credentials in flight. Speakeasy deploys through an organization's existing MDM, including Jamf and Intune, and integrates with SAML/OIDC identity providers. The page states that it is SOC 2 Type II audited, ISO 27001 certified, GDPR compliant, and HIPAA ready.
Reflexio is a learning platform for AI agents that turns user corrections, failed paths, and successful outcomes into reusable behavioral changes without retraining the underlying model. Its SDK and integration loop publish interaction outcomes, extract actionable feedback, store learned behaviors, and retrieve only relevant learnings during later inference; integrations are available through Python, REST, a CLI, and a portable coding-agent skill. Reflexio evaluates learnings against control responses and user-defined success criteria, tracking whether they solved the user's problem, required correction, or escalated to a human. It makes learnings auditable and revocable: they can be reviewed, rewritten, approved, rejected, or deleted, with rejected learnings removed from retrieval. Background processes consolidate duplicate signals and resolve conflicts or outdated lessons, while evidence from later sessions is used to revise learnings. The service supports managed, bring-your-own-key, customer-owned database, bring-your-own-cloud, and fully self-hosted deployments. The page identifies an official repository at github.com/ReflexioAI/reflexio and states that the same API can be used across deployment modes.
HyperProbe is an AI-native production debugger for investigating live application state without code changes, redeployment, or service restarts. It reads logs and distributed traces, uses a coding agent to locate a suspected file and line, and places a temporary read-only virtual breakpoint there. When the breakpoint fires on live traffic, it captures the variable state asynchronously without pausing requests, then uses the evidence to confirm a root cause. Probes are non-blocking, cannot write memory or execute code, disappear after capture, and are recorded in an immutable audit trail. The service is designed for silent failures, exceptions whose causes occur earlier in the call chain, incorrect behavior without thrown errors, race conditions, third-party contract changes, and business-metric failures. It supports JavaScript, TypeScript, Java, Kotlin, Python, and Ruby, and can run in a managed cloud, self-hosted deployment, or private VPC. The page states that it integrates with PagerDuty, Datadog, Slack, Cursor, Claude Code, Codex, and Opencode, with approval gates and agent-side PII redaction.
Experiential Labs is an open-source AI gateway that provides an OpenAI-compatible endpoint for hosted model providers, customer-owned provider keys, private GPUs, and self-hosted models. It routes requests through a single API and key while handling provider access, model selection, failover, and streaming responses. Its optional intelligence layer monitors traffic to identify model switches, improve cache hit rates, and route requests to customer-owned fine-tuned models. The service also provides organization-wide usage attribution, request logs, spend reporting, access controls, model allowlists, and per-key spending caps scoped by team, person, agent, or tool. It can be used as a hosted gateway or self-hosted, with hosted inference and Pro offered alongside the open-source gateway.
GitWarren is a desktop application for reviewing local Git changes made by coding agents before they are committed. It reads the worktree directly and combines staged, unstaged, and untracked files into one diff for local review, with comments and threaded discussions that can be edited or resolved. An MCP server over stdio lets Claude Code, Codex, or another MCP client open reviews, inspect discussions, reply in threads, comment on lines, and resolve comments; agent messages are attributed and separate MCP sessions distinguish concurrent agents. Reviews and comments are stored in a local SQLite file, while repository state is read from Git when displayed. The application requires no account, runs on macOS, Windows, and Linux, and is free and open source under GPL-3.0.
Compliance by TwelveLabs is an AI video-compliance application from TwelveLabs that applies multimodal retrieval and contextual reasoning to video libraries against configurable regional or custom compliance rules that users can author, version, and control. It identifies and explains potentially problematic moments, provides timestamped evidence and review workflows, and exports signed reports through a managed service and REST API.
Tucky is a native macOS notes app that keeps notes in a thin stripe at the edge of the screen. Notes are stored locally and encrypted on disk, with support for live Markdown, checkbox tasks, global shortcuts, clean exports, and local-only files. Its optional AI agent can search, write, rewrite, and ask questions about notes from the edge or through the global ⌥Space shortcut. The app can connect to Gmail, Google Calendar, Google Docs, Google Sheets, GitHub, and Notion; voice dictation is supported in more than 60 languages. The page says note titles accompany questions by default, while note bodies are read only when allowed. The free plan supports five notes without AI, voice, or connectors. The Plus plan is listed at $4 per month and adds unlimited notes, AI models, voice, and connectors. Tucky runs on macOS 15 or later on Apple silicon or Intel.
BrickForgerAI is an AI web service that generates custom, buildable brick models from text prompts using a library of 55 real, purchasable brick parts, including slope angles and curved pieces. It previews each model in 3D with its part count and dominant colors, then provides a downloadable LDraw .ldr file, full parts list, and step-by-step PDF building instructions for purchased generations. Small, medium, and large build sizes are offered at approximately 15, 22, or 30 studs. The service currently performs best with organic shapes such as animals, plants, and sculptures and is adding techniques including sideways building (SNOT). Its site says generated models commonly achieve full connectivity but may receive stability warnings in BrickLink Studio. Part geometry comes from the LDraw parts library under CCAL 2.0. The listed plan costs £1.50 per month for three generations with the .ldr file and instructions included. BrickForgerAI is not affiliated with LEGO, BrickLink, or Studio.
TeamAI is an open-source CLI for managing a team's skills, rules, documentation, agents, hooks, MCP configuration, environment settings, and knowledge across AI coding tools such as Claude Code, Codex, Cursor, CodeBuddy, WorkBuddy, and OpenCode. It is installed with npm as `teamai-cli` and uses a shared Git repository as the team's source of truth. The CLI distributes resources through an administrative `push` → review and merge → `pull` workflow. Session-start hooks can run `teamai pull` to synchronize the shared harness into project- or user-scoped local AI-tool directories; roles, tags, and subscribed source repositories control which resources members receive. TeamAI also provides a knowledge layer: `teamai import` and `teamai codebase` build a searchable team knowledge base and codebase graph, while `teamai recall` uses BM25 search with graph-boosted reranking and can deploy a recall subagent into supported AI tools. Code relationships are extracted with a WebAssembly tree-sitter parser for TypeScript/JavaScript, Python, and Go, with regex-based heuristic extraction as a fallback and for other languages. Additional commands support friction-based learning suggestions, privacy-scrubbed session summaries, maintenance of stale or low-confidence knowledge, usage digests, a web dashboard, member and role management, package and plugin installation, diagnostics, and uninstallation. The repository is maintained by Tencent and contributors and is licensed under MIT.
PI-Desktop is a local-first desktop workspace for AI coding agents, built with Electron, a Rust host core, and the pi Agent Harness. It lets users open local projects, connect cloud or local models through providers such as OpenAI-compatible APIs, manage projects and long-running sessions, and review file changes and command output. Its architecture separates the React renderer, Electron desktop orchestration, Rust-managed permissions, filesystem access, SQLite persistence and secrets, and a pi Agent sidecar responsible for the agent loop, model interaction, and streaming. Agent, Plan, and Goal modes provide different approval boundaries; agents can read and edit files, run commands, and delegate exploration, implementation, research, testing, or review to background Subagents. The workspace supports MCP servers, Skills, installable Plugins, model and provider switching, session imports, notifications, context checkpoints, and a plugin marketplace. Conversations are stored locally as JSONL with a SQLite index, settings and logs remain on the machine, credentials use the operating-system keychain, and model requests go directly to the configured provider or endpoint. Plugin processes are permission-gated and isolated from the renderer, but plugins remain user-trusted code rather than a complete operating-system sandbox. It is an early-preview cross-platform application distributed for macOS, Windows, and Linux under the GNU Lesser General Public License v3.0.
Vals is a code-generation evaluation product that uses a company's GitHub codebase to build an internal coding benchmark. It evaluates coding models and agents against private, continuously evolving tasks to identify which systems perform best on the company's work and which offer the strongest return on investment. The platform is also described as supporting pre-release testing and evaluation of agentic coding work that may unfold over hours, days, or weeks, while accounting for factors such as token cost and latency.
Clodds is a self-hosted, open-source AI trading terminal and autonomous trading agent for prediction markets, cryptocurrency spot and perpetual futures, Solana and EVM decentralized exchanges, token launches, and Bittensor mining. It is built around Claude and can be accessed through a local WebChat interface, CLI, MCP server, or 21 messaging channels. Its architecture combines four agents—main, trading, research, and alerts—with trading skills, market-data tools, semantic memory, strategy execution, and a unified risk layer. The documented capabilities include arbitrage detection, whale and copy trading, DCA bots, backtesting, order and portfolio management, circuit breakers, VaR/CVaR, volatility-regime detection, stress testing, Kelly sizing, daily loss limits, and a kill switch. It connects to prediction-market platforms, futures exchanges, Solana protocols, and EVM networks, and also provides agent-oriented forum, marketplace, token-launch, compute, and x402 USDC payment features. The project is distributed as an npm package or from source, stores local data under ~/.clodds/, supports SQLite, LanceDB, and PostgreSQL components, and is licensed under the MIT license. Its documented quick start is `npm install -g clodds` followed by `clodds onboard`; trading integrations require the relevant credentials and wallet configuration.
Hyperresearch is an MIT-licensed Python command-line research harness that extends Claude Code with an agent-driven, persistent research knowledge base. It classifies a query into light, full, or opt-in dissertation work and runs a tier-adaptive pipeline covering query decomposition, source search and fetching, contradiction and locus analysis, depth investigation, evidence digestion, multi-angle drafting, synthesis, adversarial criticism, targeted gap fetching, surgical patching, citation checks, and readability review. The pipeline stores fetched sources as Markdown notes with a rebuildable SQLite index, provenance links, lifecycle statuses, full-text or optional semantic search, source-quality and retraction metadata, and resumable per-run manifests. It can fetch PDFs, seek disclosed open-access replacements for thin or blocked academic sources through Unpaywall and Europe PMC, and apply structural checks for citation bindings, quote integrity, provenance, retractions, and patch-only editing. The vault can also be accessed through an MCP server or a dependency-free local web UI. It is installed with pip for Python 3.11–3.13 and is intended to run with Claude Code; model assignments and research scale can be configured per project.
OpenResearch is a local-first workspace for research agents and autoresearch, available as a desktop application and a macOS/Linux CLI. It turns Claude Code, Codex, or OpenCode into agents that can review literature, develop hypotheses, run experiments, and produce research artifacts. The workspace assigns independent agent sessions and isolated Git worktrees to parallel research directions. It tracks experiment variants in a Git-native experiment tree, archives each run against its recorded commit, and keeps logs, diffs, files, results, and artifacts tied to the work that produced them. Its autoresearch loop can propose an idea, modify code, launch an experiment, inspect evidence, and choose the next direction. The same committed source snapshot can run locally or through SSH, Slurm, Kubernetes, Ray, Hugging Face Jobs, Modal, Tinker, or managed OpenResearch compute. By default, projects and run data remain on the user's machine in a local SQLite store, with a browser dashboard served on localhost. An OpenResearch account is used for service-owned capabilities such as organizations and managed compute. The CLI also installs an OpenResearch skill into supported coding agents and provides commands for projects, runs, logs, experiments, discovery, and paper lookup.
YuE2 is an open music-generation model and Python pipeline that turns lyrics and a style prompt into a symbolic melody-and-chord plan and then renders that plan as a complete stereo song with vocals and accompaniment. It supports zero-shot covers by using a transcribed melody score, and composition editing by allowing the score, lyrics, style, harmony, melody, tempo, or form to be revised before rendering. Its AR–NAR Mixture-of-Transformers backbone predicts symbolic scores and semantic tokens, generates acoustic latents with flow matching, and uses a VAE to decode them into 48 kHz audio. The staged API exposes plan(), generate_semantic(), synthesize(), and decode() operations; the repository also provides an agent skill for song generation, transcription, cover creation, ABC-score editing, and listening comparisons. The repository provides the YuE2-3B model and requires Linux, Python 3.12, and an NVIDIA GPU with BF16 support and 24 GB of VRAM for the documented quick start. First-party code, documentation, and the agent skill use Apache 2.0, while the model weights use CC BY-NC 4.0. The original YuE implementation is preserved on the repository's YuE-v1 branch.
Worktrunk is a command-line interface for managing Git worktrees, designed for running AI coding agents in parallel. It addresses worktrees by branch name and computes their paths from a configurable template; its core commands switch to, list, merge, and remove worktrees, with an option to launch a command such as Claude Code after switching. It also provides hooks for local workflow automation, LLM-generated commit messages, merge and cleanup workflows, an interactive worktree picker with diff and log previews, shared build caches, CI status and AI-generated branch summaries, pull-request checkout, per-worktree development-server ports, and aliases with branch-scoped variables. It can be installed through Homebrew, Cargo, Winget, Arch Linux, Conda, or Pixi, and shell integration enables commands to change directories.
Dream Loop is an AI agent skill for building games, apps, and 3D scenes toward a visual target. It uses image generation to create a target screenshot, builds the scene—such as a Three.js scene in a browser—and has a separate vision-enabled critic compare the live screenshot with the target and provide feedback; the agent then iterates on the build and can optionally generate an improved target from the current state. Blender can be used for custom 3D modeling through Blender MCP or scripting. The skill can be installed with `npx skills add achimala/dream-loop` or cloned into an agent's skills directory, and requires an agent with image-generation access, vision input, and preferably subagents.
Bang Motion is an agent skill for building browser-based motion graphics such as product openers, promos, bumpers, channel intros, kinetic typography, lower thirds, and explainers. It produces a single index.html that plays in a browser and encodes anti-slide rules: scenes change through camera or world movement rather than section fades, subjects or shared worlds persist, motion occurs continuously, and text and numbers remain part of the scene. Its deterministic timeline makes each frame a pure function of time for exact scrubbing and frame-by-frame export, while a shared camera rig supports flowing-world and camera-through-collage modes. The skill includes five explainer starters—cartoon collage, visual journalism, white catalog, vintage sketch, and continuous action—along with configurable style briefs, background motion, surfaces, transitions, highlight shapes, voice-over re-timing, and optional Puppeteer and FFmpeg scripts for verification and MP4 export. It can be installed as a Claude Code plugin or copied into Claude Code, Codex CLI, Gemini CLI, Cursor, or another agent workflow; watching the output requires only a browser and internet connection. The repository is licensed under MIT.
Browser Use Pi is a TypeScript web agent built on Pi Mono. It combines a persistent V8 REPL with raw Chrome DevTools Protocol control: the agent writes JavaScript, uses accessibility-tree data and screenshots, and builds browser helpers as needed. It supports persistent sessions, saved logins, workspaces, streaming, hooks, typed results, follow-up tasks, and cloud browsers or a user's own Chrome; runs can be limited by steps, time, or cost, with partial work retained. Interaction highlights, recordings, and GIF exports can show the agent's work. It is distributed as the @browser_use/pi package for Node 22.19+ or Bun 1.3.14+, and its quickstart uses Browser Use Cloud without requiring a local Chrome installation.
Astra Advisor is a Codex plugin for capability-routed software delivery from DannyMac180. GPT-6 Astra acts as the primary architect and acceptance owner: it plans work, decides whether bounded independent tasks should be delegated, selects supported native subagent models and reasoning effort, and continues parent-session work while delegations run. It uses the exposed collaboration tool to route tasks among Sol, Terra, and Luna subagents, then inspects the complete diff, reruns requested checks, and sends substantial changes to a fresh read-only reviewer; work is accepted only after the reviewer returns "ship". Each delegation reports its selected and runtime-observed settings, and tasks can end with an API-equivalent cost receipt that compares observed token usage with a versioned Astra pricing snapshot. It can be installed as a plugin in a current Codex CLI or ChatGPT desktop app with plugins enabled. The README notes that ChatGPT Work cloud tasks currently cannot provide arbitrary model or effort controls, so Astra does not dispatch model-pinned requests there by default.
SuperAstra is a SNES-themed desktop companion that uses natural-language prompts to investigate and alter a running game through BizHawk. Its agent receives screenshots, cartridge identity data, emulator registers, memory scans, hardware-domain reads, controller probes, and checkpoint results; it can search cartridge data, identify game-memory structures, create routines or guarded patches, test changes, and retain cartridge-specific findings in a local notebook. Each committed mutation can be undone through emulator checkpoints, while named experiment checkpoints support controlled trials that restore the player's prior state. The prototype runs on Windows or Linux with Python and Tkinter, BizHawk using its BSNES SNES core, a ROM, and an OpenAI API key configured for the stated model; it also includes limited local Mario shortcuts that do not require an API key. It does not patch the original ROM file, create exportable ROM patches, expand ROM capacity, or guarantee success on unfamiliar games. The original source is MIT-licensed.
Agent Skills is Google's repository of installable skills for Google products and technologies, including Google Cloud. The skills provide Markdown-based procedures for tasks such as onboarding and authenticating to Google Cloud, deploying and managing AI agents, working with GKE, databases, networking, security, monitoring, analytics, advertising APIs, Firebase, Android, Dart, Flutter, and Google Maps Platform. Skills can be selected and installed with `npx skills add google/skills`; the repository also bundles plugins and MCP servers for agent harnesses including Claude Code, Codex, and Antigravity CLI. The repository accepts bug reports and contributions, and its skills are distributed under the Apache 2.0 license.
WebLLM is an in-browser large language model inference engine developed by the MLC community. It runs model inference directly in web browsers using WebGPU hardware acceleration, without server-side processing, and can be used as an npm, Yarn, or CDN package for building web applications. Its MLCEngine exposes an OpenAI-compatible chat-completions interface with streaming, JSON-mode structured generation, seeding, and preliminary function-calling support. It supports model families including Llama, Phi, Gemma, Mistral, and Qwen, and can load custom MLC-format models consisting of model artifacts and WebAssembly model libraries. WebLLM provides Web Worker and Service Worker integrations, Chrome-extension examples, browser caching through Cache API, IndexedDB, or OPFS, and optional Subresource Integrity verification for downloaded model artifacts.
Distilly is an AI agent skill-generation tool by titanwings that turns messages, documents, interviews, and other supplied sources into source-grounded Person Profiles. It models observable experience, decision patterns, expression, and ways of working, rather than claiming to clone a person, and supports colleague, relationship, and celebrity profile families. Each generated profile is packaged as an installable Agent Skill for supported hosts. Its workflow includes family-specific source collection and analysis, a six-dimension research pipeline for celebrity profiles, incremental file analysis and merging, conversational corrections written to a correction layer, and automatic version archiving with rollback to earlier versions. Supported source types include Lark, DingTalk, Slack, public X posts, WeChat exports, PDFs, images, email, Markdown, and pasted text, with source-specific setup requirements. The project documents native local Skill discovery for Claude Code, Hermes, OpenClaw, Codex, DeepSeek Harness, Pi, Grok Build, and OpenCode. It is distributed from its GitHub repository under the MIT License and is described as a demo version.
OpenAI GPT Image 2 is identified in the video as an AI image-generation offering that creates images from prompt templates in a GPT Image 2 library.
Switchyard is an NVIDIA Rust library and proxy for routing LLM requests among models and providers while preserving native OpenAI and Anthropic API compatibility. Its routing algorithms choose among efficient and capable models using mechanisms such as LLM judging, stage-based response evaluation, escalation, advisor approval, random selection, and custom target-selection policies. It can run as a standalone proxy, integrate with NeMo Relay or LiteLLM, or be embedded in a custom gateway or harness; the proxy exposes OpenAI Chat Completions, Anthropic Messages, and OpenAI Responses endpoints, along with statistics and Prometheus metrics for requests, errors, latency, tokens, and routing overhead. The repository provides Rust components for routing, provider-neutral protocols, HTTP model calls, and request, response, and stream translation. It is pre-1.0 software under the Apache 2.0 license, with the standalone server documented as suitable for demonstrations and evaluation rather than production.
Agent Reach is an open-source command-line capability layer for AI agents. It selects, installs, configures, checks, and routes tools that let command-line agents read and search websites and services including web pages, YouTube, RSS, GitHub, Twitter/X, Reddit, Bilibili, Facebook, Instagram, Xiaohongshu, LinkedIn, V2EX, and Xueqiu. It uses an ordered list of backends for each channel and probes them for actual availability, selecting the first usable option and reporting failures and repair guidance through `agent-reach doctor`. The repository lists integrations such as Jina Reader for web pages, yt-dlp for YouTube transcripts, `gh` for GitHub, feedparser for RSS, Exa through mcporter for web search, and alternative CLI or browser-session-based routes for sites requiring authentication. Agents invoke the upstream tools directly rather than through a data-wrapping layer. The CLI supports environment checks, explicit system installation, dry runs, updates, and uninstalling. Credentials and cookies are documented as being stored locally in `~/.agent-reach/config.yaml` with owner-only permissions; authenticated services may require user-provided cookies or browser sessions. The project is distributed under the MIT license.
oh-my-hermes (OMH) is a plugin and operating layer for Hermes Agent that adds model routing, coding workflows, specialist skills, project memory, and evidence-gated execution without replacing Hermes as the natural-language interface. It scores requests, selects workflows and model-and-effort categories, applies model-family-specific prompting, and can split accepted plans into parallel worktree-based lanes with typed results and verification gates. Its workflows cover interviewing, research, planning, coding delegation, quality assurance, performance work, and iterative plan-build-review loops; the terminal interface exposes phase-based todos, delegated-lane status, cost and token telemetry, and execution states that distinguish preparation, reported completion, and verified results. OMH also provides a file-backed, reviewer-gated memory store with provenance, review dates, conflict handling, and task-scoped recall packs, while leaving Hermes's native memory untouched. It is distributed as a command-line package and plugin with installation paths including a shell installer, Homebrew, Bun, npm, and Hermes skill tap.
Viserys is a self-contained pack of Markdown engineering workflows, reviewer personas, validation scripts, commands, and lifecycle hooks for AI coding agents. Its workflow routes development through DEFINE, PLAN, BUILD, VERIFY, REVIEW, and SHIP stages, with 28 skills covering requirements clarification, specification and task breakdown, implementation, test-driven development, debugging, browser testing, review, security, performance, migrations, CI/CD, documentation, observability, and release preparation. The pack also includes four reviewer personas, shared checklists, evaluation cases and fixtures, and validators for skills, commands, artifact paths, reference links, and versions. The repository includes an OpenCode agent definition that registers the skills and routes requests through the lifecycle workflow. The Markdown skills can also be supplied to skill-aware agents such as Claude Code, Cursor, and other agents, or read directly as standalone process documents.
edge0 is an open-source streaming Mixture-of-Experts inference framework for running sparse MoE models on Apple Silicon Macs. Its MLX backend memory-maps expert weights from SSD storage and loads them on demand, so peak memory depends on the active experts rather than the model's total parameter count. A trained prerouter predicts the next expert routes so storage reads can overlap the forward pass, while Recover-LoRA adapters distill from the full-precision model to compensate for quantization loss; the base weights remain read-only and the adapters stay separate. The framework ships two end-to-end model tiers with bundled 4-bit checkpoints, LoRA adapters, and prerouter heads. It provides a command-line interface and Python API, supports demo, chat, and server modes, and exposes an OpenAI-compatible `/v1/chat/completions` endpoint. The current backend supports macOS on Apple Silicon M1 through M4; CUDA support is identified as planned. The repository is licensed under Apache-2.0.
3dviz-pro-max is a standalone agent skill for turning plain-language ideas into interactive Three.js or Blender scenes. It is designed for Claude Code and Codex, and provides a ten-step workflow covering intent, object reasoning, representation, construction routing, first-view creation, meaningful behavior, capture, inspection, refinement, and reporting. The skill includes searchable recipes and knowledge records, runnable Three.js templates, reusable scene kits, style and camera guidance, and a capture helper that drives a built scene in a browser and writes PNGs and console logs for inspection. Its catalog covers subjects including environments, anatomy, mathematics, physics, architecture, logistics, robotics, and abstract systems. The supplied Python helpers search the catalog and drive an existing scene; they do not themselves render or judge scenes, and screen capture requires an optional Playwright installation. The repository includes 37 runnable studies and a site at 3dviz.dev. Its own code, data, and authored media are released under the MIT license, while bundled third-party assets retain separate terms. Blender 4.2 LTS or newer is recommended but not required; without Blender, the repository states that scripts still run but kit output is limited to T2.
Life Recorder is an open-source native iPhone recorder with a private Mac receiver. The iPhone captures roughly one-minute AAC chunks while the screen is locked or other apps are in use, then uploads them over an authenticated HTTPS connection. The Mac receiver runs whisper.cpp locally, removes common stage-direction markers and repetitive hallucinated noise, and appends the results to one continuous Markdown transcript with hourly capture markers; audio is deleted after durable receipt and successful transcription. It requires an iPhone development build installed with Xcode, a macOS system running Python, ffmpeg, whisper-cli, and a downloaded GGML Whisper model. Setup creates a bearer token and self-signed TLS certificate, uses certificate pinning, and stores the transcript in a private runtime directory. The default connection is over the local network; cellular use requires a private VPN. The project is licensed under MIT.
deepseek-recipe is a collection of Rust libraries and Python bindings for converting API requests in Messages, Chat Completions, and Responses formats into a shared conversation representation, encoding conversations into prompts or token IDs for DeepSeek V4 and V4.1 models, and converting model output back into protocol-specific responses. It supports streaming, text and image inputs, thinking and reasoning settings, function tools, tool calls, JSON object output, and stop sequences; its image component provides DeepSeek V4.1 preprocessing with OpenCV. The package parses model output and handles prompt encoding, while inference, tool execution, and HTTP transport are supplied externally. It includes separate components for core conversation types, prompt encoding, image processing, Python bindings, and example Axum and FastAPI servers with mock inference. Project code and public documentation are licensed under the MIT License.
Infinite World is an open-source system for building persistent worlds with multimodal AI. It combines multimodal understanding, world simulation, and interaction in a continuous loop: prompts, images, or other inputs are interpreted into world state, rules, and possible actions; generated scenes respond to that state; and people, agents, and events can update the world. The runtime records scenes, choices, state changes, history, and branches so worlds can preserve context across scenes and be replayed. The project supports text, voice, and image interaction, local previews, saved-branch replay, and cached generated video. Its current local setup uses a Node.js and pnpm workspace, with FFmpeg and whisper.cpp for voice transcription and live output. Interactive live streaming is in development, while additional inputs such as chat, audio, mouse, and keyboard are planned. The repository describes the project as early development, with APIs and stored data subject to change, and distributes it under the Apache License 2.0.
Mural is an open-source native iPhone and Android language-learning app built around live conversations with an AI voice companion. It lets learners choose a learning language and subtitle language, practise through voice or typed replies, receive corrections and optional meaning subtitles, look up words, and review vocabulary from later conversations. Challenge levels and provisional ability observations are adjusted from evidence in the learner’s replies, while each language has separate progress records and conversation themes. The iPhone client uses SwiftUI, SwiftData, Keychain storage and WebRTC; the Android client uses Jetpack Compose. Conversations, vocabulary and preferences remain on the device, and learning data can be exported or imported as JSON. The app connects directly to OpenAI using the learner’s own API key, requires an internet connection, and sends practice audio, selected text, learning context and requested searches to OpenAI. Hosted free conversations, minute purchases, and public App Store or TestFlight distribution are not active according to the repository. Mural supports Norwegian Bokmål, Spanish from Spain, international English, French from France, German, Italian, Brazilian Portuguese and Standard Mandarin with Simplified Chinese. It is released under the MIT License; third-party components retain their own licenses.
geiger is a read-only command-line scanner from Atomburst that inventories AI agents, harnesses, MCP servers, plugins, skills, hooks, browser extensions, desktop AI applications, and related IDE configurations on a machine. It reads known configuration files and directories, identifies each finding's origin and evidence path, and labels capabilities such as code execution, secret storage, broad filesystem access, broad web access, and network access. Credential-shaped values are reported by key name, file, and secret type without exposing value contents; the scanner performs no telemetry and writes only an explicitly requested report file. It can emit terminal, HTML, or versioned JSON reports, scan specified project or home directories, enforce a strict exit status for findings that can execute code or hold secrets, and compare JSON snapshots to identify newly appeared, removed, or escalated findings. It runs through npx with Node.js 18 or newer, has no runtime dependencies or account requirement, and is licensed under MIT. The project notes that it reads configuration rather than runtime behavior, covers known locations, and does not determine whether a package is malicious.
Agent Launcher is a local desktop workspace for configuring and running existing coding-agent command-line interfaces. It detects installed CLI binaries, links or installs supported agents, applies account or provider profiles, and runs them in an embedded terminal or chat view without automatically reinstalling or updating already installed CLIs. It supports project-aware sessions by launching an agent in a selected project directory, synchronizes provider settings with CLI-native configuration files and environment variables, and performs a minimal model request to check credentials, endpoints, models, network access, and account status. The application reads local conversation histories in JSONL or SQLite formats, can resume or delete local sessions, displays installed MCP servers and Skills, and summarizes locally stored usage data. Agent Launcher distributes macOS, Windows, and Linux builds and is released under the MIT License. It is local-first rather than offline-only: launched agents send requests to the provider or relay selected by the active profile, while API keys are stored as plaintext in local configuration files and masked only in the interface.
Eliza OS is a software compatibility target for converted robot demonstration records.
DeepJIT is a lightweight, header-only C++20 library for just-in-time compilation, caching, loading, and launching of device kernels on NVIDIA CUDA GPUs and Huawei Ascend NPUs. It provides a shared runtime interface for both backends while keeping kernel source and compiler options backend-specific. The runtime hashes kernel source, tracked includes, compiler information, effective options, and application-provided dependency signatures to reuse compiled artifacts through in-memory and on-disk caches. Cache directories can be shared across users, processes, and nodes on local or distributed filesystems, with completed entries published through atomic directory renames. Lazy initialization defers device and compiler discovery until first use. DeepJIT integrates with PyTorch CUDA and torch_npu streams and can expose a runtime through pybind11. The CUDA backend uses NVCC and the CUDA Driver API to produce, load, and launch CUBIN files; the Ascend backend uses Bisheng, ld.lld, and ACL to compile, link, load, and launch kernels. It supports per-kernel compiler and launch options, compilation diagnostics, and optional PTX, SASS, or Ascend assembly dumps. The host environment requires Linux, a C++20 compiler with std::format support, Python, pybind11, and the dependencies for the selected backend; it is intended to be embedded in another C++ or Python extension as a header-only dependency.
gpu-time is an experimental local neural parser for English date and time expressions. It converts text into ISO dates and time ranges, RFC 5545 recurrence rules, diagnostics, and source spans. The parser extracts token features without a word list, classifies tokens into time-related roles, detects expression boundaries, and uses TypeScript to build schedules from the predictions. It runs on CPU or WebGPU, with WebGPU processing parallel blocks and overlapping windows for long inputs; a calendar resolver applies the supplied reference instant, time zone, daylight-saving rules, and occurrence limit. The package provides parse, parseMany, and defineParser APIs, and does not send input to a server.
Phyzical is a browser-based robot teleoperation data platform for embodied-AI training. Users control simulated robot arms through a web interface; inverse kinematics converts end-effector movements into joint actions, while sessions record demonstration trajectories containing joint positions, end-effector poses, object poses, and actions at 30–60 Hz. Episodes use a documented JSON Schema format and can be converted with the included Python tool into elizaOS-compatible SQLite trajectory databases for imitation learning and related robot-training pipelines. The project also documents quality checks and on-chain episode provenance using a data ID, content hash, and contributor, and includes human, scripted, and Fly-controller action sources. The repository is MIT-licensed and provides the converter and example episodes; it requires Python 3.10 or later and has no external dependencies for the quickstart converter.