2,417 tools and products — trending open source, and what gets used in AI and other work.
LatticeDB is an embedded, single-file property-graph database written in Zig for local applications. It combines relationship traversal, HNSW vector similarity search, and BM25 full-text search in one query language and engine, allowing graph, semantic, and textual queries over the same dataset. It also provides durable named event streams and a graph changefeed through the same transaction and write-ahead-log path as graph writes. The database operates without a server or configuration and is designed for one owning process on one machine with WAL-backed durability. It is distributed through a CLI and bindings or packages for Python, TypeScript/Node.js, Java, and Go. Graph RAG, agent memory, and local knowledge tools are documented as example workloads rather than the engine's definition.
CarWatch is an open-source, offline in-car AI agent from ThinkOffApp that runs on a Raspberry Pi 5 with a locally hosted language model. It joins chat rooms as a vehicle agent, answers questions from the car's owner's manual using lexical retrieval-augmented generation and page citations, and reads live vehicle and system data while stating when information cannot be sensed. A continuous energy-based voice listener and whisper.cpp provide local speech input without a wake word, with responses played through the car's speakers. The system can monitor OBD and engine data, send departure and arrival messages, trip summaries, and dashcam clips, and expose a phone dashboard for approvals, replies, model selection, and maintenance. The repository describes systemd-managed services for the model server, room agent, voice listener, dashboard, and engine watcher, with Python standard-library code and no cloud subscription required.
SELF is an experimental executable format and Linux/NixOS runtime in which a program is stored as a SQLite database rather than an ELF file. Its rows represent the executable format, and a binfmt_misc interpreter can rebuild an ELF in memory and execute it with execveat, map segments and hand off to the dynamic linker, or act as a dynamic linker that binds through SQL; the project also provides ELF/SELF conversion tools, a selfifying Nix hook, a NixOS module, and an LD_AUDIT library for resolving .self shared libraries through SQL. The repository includes self-httpd, a web server whose program, routes, pages, and visitor log are stored in one database file. It opens its own executable path as a database and serves routes from its routes table, so editing the live site is performed with an SQL UPDATE. SELF databases can also be queried for executable information, such as shared-library names, and modified transactionally, including deleting sections and vacuuming the database.
Factory is a reference software factory for Claude Code and Codex that installs a repeatable, version-controlled software-delivery workflow into an existing GitHub project. GitHub Issues serve as the work queue, while committed policies, skills, labels, handoff comments, pull requests, and run records preserve state between fresh agent sessions. Scheduled agents triage issues, route them to implementation, specification, questions, or blockers, claim bounded work, implement it on a branch, run configured type, lint, test, build, audit, and architecture checks, obtain independent verification, and open draft pull requests. An independent verifier reads the diff and checks that the new test fails without the implementation. Humans retain responsibility for ambiguous requirements, system design, significant changes, and merging pull requests. The repository has no custom orchestrator or queue service. Claude Code routines provide the default scheduling and compute, with a thin Codex adapter using the same policies, gates, and evidence files; GitHub events or an optional API-triggered Action can also start runs. It is configured through files such as a human-owned charter and gates script, and includes a /factory control room backed by live issues, pull requests, and run records.
FrontierAgent is an open-source agent runtime, terminal product, and evaluation suite for long-horizon research and file-based work. Its native terminal TUI provides a ReAct workflow in which one stateful agent researches, reads files, writes deliverables, runs commands, and iterates in a task-scoped sandbox, and an Agent Team workflow in which a coordinator maintains a task board, delegates bounded independent assignments to parallel sub-agents, collects structured reports, and synthesizes the result. Shell and file tools use a shared sandbox with read-only /inputs, working-state /workspace, and persistent /outputs directories. The TUI supports queued instructions during execution, approval-required diffs for mutating operations, local action traces, checkpointed sessions, resumption, and reversion. The same workflow engine powers a subprocess benchmark runner with deterministic artifact collection, concurrency, progress inspection, and reruns of individual failures. The repository also documents connection to the OpenAI-compatible Apodex-1.1 endpoint through the Apodex API.
hayamimi is a local, real-time multilingual speech-to-text tool that produces live subtitles, a browser dashboard, speaker labels, and optional translation without a GPU or cloud API. It detects the language of each utterance and routes it to a dedicated specialist model for Japanese, Chinese, Korean, Cantonese, English, and European languages, with a broad Omnilingual ASR fallback for other languages. The models run as quantized INT8 ONNX models through sherpa-onnx, without PyTorch or CUDA. The pipeline updates partial subtitles while speech is in progress, emits finalized lines shortly after speech stops, and can perform a second batch-decoding pass after silence to produce a refined transcript. Speaker labeling uses CAM++ embeddings and is intended for turn-taking rather than full diarization. The `--serve` option starts a local HTTP server with a live dashboard and an OBS browser-source overlay. Audio can also be supplied over WebSocket, and model eviction keeps resident models below a configurable memory cap, with the default under 2 GB of RAM. The project reports 5.8% character error rate on its Japanese broadcast-audio scorecard and 10–50× real-time processing on a six-core desktop CPU.
XLA (Accelerated Linear Algebra) is an open-source machine-learning compiler for GPUs, CPUs, and other ML accelerators. It takes models from frameworks such as JAX, PyTorch, and TensorFlow, traces and compiles them into optimized executable code for the target hardware. The project is developed under OpenXLA and is primarily used through the documentation and integration of the corresponding ML framework; its repository is intended mainly for compiler contributors and frontend or hardware-backend integrators.
FreeLLMAPI is an open-source proxy that aggregates free tiers from multiple LLM providers behind a single OpenAI-compatible `/v1` endpoint. It can also connect custom OpenAI-compatible chat, embedding, image, and audio endpoints, including local servers and remote gateways. Its router selects an available model, fails over when a provider is rate-limited, stores provider keys encrypted, and tracks per-key usage against free-tier limits. The router refreshes its model catalog from a signed feed; free installations receive periodic snapshots, while the hosted premium offering provides the live catalog. The project is intended for personal experimentation and is available as a self-hosted install with a hosted service at freellmapi.co.
ai-memory is an open-source local server for long-term memory and session handoff among AI coding agents. It captures prompts, tool calls, and agent context through MCP integrations and lifecycle hooks, storing the material in a searchable, Git-backed Markdown wiki. When a session ends or is explicitly finalized, it generates handoff context for a subsequent session, including the project architecture, failed approaches, and open questions, so work can continue across agent vendors without restating the context; clients without a true session-end hook use explicit finalization commands. The project documents integrations for Claude Code, Codex, Command Code, Devin CLI, OpenCode, Cursor, Gemini CLI, Oh My Pi, Pi, and Crush, with agent-specific configuration, generated plugins or extensions, and capture exclusions. It runs on Linux, macOS, and Windows through WSL2, with experimental native Windows support, and is distributed through Docker images and native release binaries.
AI Engineering from Scratch is a free, open-source curriculum developed by rohitg00 for learning to build and ship AI systems end to end. Its 20-phase sequence covers development tooling, mathematical and machine-learning foundations, deep learning, NLP, computer vision, generative AI, LLM applications, agent engineering, Model Context Protocol (MCP), agent skills, reinforcement learning, and related topics, with implementations in Python, TypeScript, Rust, and Julia. The repository describes 511 lessons and approximately 329 hours of material; each lesson produces a reusable artifact such as a prompt, skill, agent, or MCP server. Learners are instructed to read the lesson documentation, type and run the code from the repository root, record command evidence and outputs, and make a small change before continuing. It provides routes for complete foundations, mathematics and machine learning, production LLM applications, agent engineering, MCP, agent skills, and Claude certification preparation. The repository is distributed under the MIT license and is accompanied by a website containing the same lesson content.
Mercor is a data and talent service discussed as providing training information for frontier AI models. The video emphasizes the company’s rapid growth and high valuation.
Bizee is a business-formation and compliance service for creating LLCs and other business entities. Its services include registered-agent management, filing-deadline support, EIN filings, operating agreements, and related business documents.
BetterVoice is an experimental, open-source macOS menu-bar app for local voice dictation with screen context. It transcribes speech locally and captures the full display whenever the user circles an area with the pointer, marking each capture with a blue pointer trail and pulse while preserving multiple references in order. The app inserts the transcript into the selected text field when macOS permits it and pastes captured images when both speech and screen context are available. It supports quick notes and long explanations, configurable microphones, dictation languages, circle-detection thresholds, and keyboard shortcuts. English uses a local Parakeet model by default, while other languages use a separately downloaded multilingual model; an optional English-only grammar-cleanup feature runs locally through quantized ONNX weights.
BookOrbit is a self-hosted library and reading platform for ebooks, PDFs, audiobooks, and comics. It provides web readers and synchronizes reading progress, highlights, annotations, and reading status between BookOrbit, Kobo devices, and KOReader. The platform includes library management with metadata enrichment, multiple libraries, collections and rule-based saved filters, multi-user accounts with OIDC/SSO, OPDS access, Send-to-Kindle delivery, and browser uploads. It also provides reading statistics, goals, achievements, searchable annotation exports, and synchronization with services including Hardcover, Readwise, and StoryGraph. The project is designed to run on infrastructure controlled by the operator.
Workout Guide is an open exercise illustration library and framework-neutral npm package by Bryl Lim. It provides structured metadata for 302 exercises, three consistent frames per exercise, transparent normalized assets, and typed APIs such as `getExercise`, `searchExercises`, and `getAssetUrl`; its searchable static gallery includes exercise detail pages and integration guidance. The repository is an npm-workspace monorepo containing the package API, canonical manifest, asset files, gallery site, documentation, and deterministic catalog import and validation scripts. Code and documentation are licensed under MIT, while the visual assets are licensed under CC BY-SA 4.0, including artwork derived from Everkinetic.
Microduck is the software brain for Pollen Robotics' approximately 25 cm biped duck robot. It runs reinforcement-learning policies on a Rockchip RK3566, with a 50 Hz control loop for fifteen servos, radios, the camera, and behaviors such as walking, rolling, grasping, kicking, quacking, and self-righting; it also accepts gamepad control. The Rust implementation is organized as daemons communicating through a shared JSON-RPC contract over Unix sockets. `robotd` owns the control loop and motor bus, while other services handle updates, configuration and Wi-Fi, Bluetooth, gamepad input, camera streaming over WebRTC, and the depth sensor. Signed updates are health-gated, reversible, and roll-backable. The policies are trained in the companion `microduck_rl` project with MuJoCo and PPO, then exported to ONNX for loading by this repository.
Experiential is an open-source gateway and router for agent workflows from Experiential Labs. Its `exp` CLI provides hosted, bring-your-own-key, local, and custom models through OpenAI-compatible and Anthropic Messages APIs, with model aliases, access controls, use-case restrictions, and spending limits for users and agents. It can persist provider connections and configuration locally, expose API routes on loopback, and load a fitted project router as an OpenAI client through its Python interface. It collects OpenTelemetry traces from agent traffic to build simulations and optimize routing for quality, speed, and cost, and supports harness optimization, endpoint serving, model distillation, and fine-tuning an owned open-source model through Tinker. The hosted gateway is available at `api.experientiallabs.ai`; anonymous aggregate PostHog telemetry is enabled by default locally and can be disabled.
ESLint is a pluggable and configurable linter for statically analyzing JavaScript code. It identifies and reports code patterns, can automatically fix many problems with syntax-aware fixes, and supports custom parsers, preprocessing, and user-defined rules alongside its built-in rules. It can run in text editors and continuous integration pipelines, and the video identifies it as providing linting support for tsrx.
Drizzle ORM is a lightweight TypeScript ORM for defining database schemas and querying relational databases. Its migration tooling can generate, apply, push, pull, export, and check schema migrations; the documentation covers PostgreSQL, MySQL, SQLite, SingleStore, MSSQL, CockroachDB, and related database services. In the cited context, Drizzle owns the PostgreSQL schema and migration workflow for effect-mq.
Doop is an open-source multiplayer design canvas where people and AI agents create and edit designs together. Each canvas contains frames that render real HTML in sandboxed iframes; humans edit in the browser, while agents connect through the built-in Model Context Protocol (MCP) server and can create frames, stream HTML in chunks, inspect screenshots, and revise designs. The app synchronizes cursors, presence, frame edits, agent status, comments, tasks, and activity over WebSocket rooms, and includes a server-side Doop Agent that can process queued cards, mentions, and feedback through specialist roles. Canvases are private by default, with email invitations or optional link sharing, and agents inherit the access of the user who authenticates them. Doop also provides design-memory features for exemplar frames, decisions, and proposed style rules, and can be self-hosted with Docker Compose or Bun using embedded Postgres through PGlite.
OCR It is a Chrome and Firefox browser extension for extracting text from paginated document viewers, including scanned books, slide decks, PDFs, and embedded readers where text selection is unavailable. Users select a screen region, capture it with a hotkey, and append the locally generated OCR result to an ordered transcript; an automatic mode captures each page, advances the viewer, and stops when text repeats, paging fails, OCR fails, or a page limit is reached. OCR runs offline through a bundled Tesseract build without outbound requests, API calls, or image uploads. The extension crops, optionally enlarges, and converts the visible region to grayscale, queues OCR jobs serially, stores page text and verification thumbnails, flags consecutive duplicate pages, and exports the transcript as copied text or a TXT file with page separators. Page turning can use a picked screen point or keyboard event, including controls inside supported frames and shadow roots. It uses narrowly scoped browser permissions, relying on activeTab for single captures and requesting an additional site grant for automatic runs that persist across page loads or turn pages inside cross-origin iframes.
ego lite is a macOS browser from Citro Labs designed for AI-agent browser automation alongside a user's normal browsing. It gives each agent or task an isolated Space within the same browser, allowing multiple tasks to run in parallel without taking over the user's tabs; users can observe, take over, or stop an agent's Space. Through the ego-browser skill, agents such as Claude Code, Codex, Cursor, or custom agents can call in-page JavaScript tools including snapshot, fill, click, wait, navigate, and capture; an agent can compose several operations into one JavaScript execution rather than repeatedly issuing CLI commands. During setup, optional Chrome-data migration can transfer existing logins, cookies, extensions, and bookmarks, while browsing data remains on the user's device. The browser is distributed as a separate free download, the repository is MIT-licensed, and Windows and Linux support are listed as planned.
SwarmForge is a local, tmux-based orchestration platform for coordinating multiple AI agents across Git worktrees. It reads a project-local configuration that assigns roles, agent backends, worktrees, task or batch handling, and handoff propagation; launches each role in an isolated tmux session; and provides a browser-based pack cockpit for projects, tasks, approvals, clarifications, live status, agent-pane inspection, and teardown. Agents communicate through a daemon-delivered handoff protocol and helper scripts rather than direct tmux messages. The repository supplies two-pack, four-pack, and six-pack workflow templates with role prompts and constitution articles, while projects can define their own swarm topology. It runs locally with zsh, Git, tmux, Babashka, and at least one supported agent backend such as Claude, Codex, Copilot, or Grok; swarm state is kept in the working director
Rome is an agentic operating environment from Rome OS for collaboration between humans and AI agents. It provides manifests, typed actions, agents, skills, hooks, tools, workflows, memory, interfaces, policies, and persistent database-backed state in a guardrailed environment where agents can build applications, define standard operating procedures, and orchestrate workflows under human guidance. Rome supports one-off tasks, scheduled work, and long-running follow-through. Rome Apps combine a purpose-built web interface, agent reasoning, agent-owned collaborators, reusable workflows and skills, lifecycle hooks, typed actions, HTTP APIs, and persistent app-private databases or files into installable products. Apps are defined with an app.yaml manifest and use the @rome-os/app-runtime and @rome-os/app-web-sdk packages; Rome can generate and iterate on app or workflow source code from a plain-language request. The project can run locally through Docker or be accessed through Rome Cloud, which the repository describes as a private preview environment.
Proliferate is an open-source AI IDE for running coding agents such as Claude Code, Codex, OpenCode, Cursor, and Grok through their native harnesses. It runs agents in parallel within one workspace, giving each task an isolated Git branch and worktree along with its own terminal, conversation, and review state. Agents can delegate scoped work to subagents, while MCPs, skills, Computer Use, Browser Use, and custom tools can be configured once and shared across agents. Its workflow system supports recurring and event-driven runs such as review passes, alert triage, and dependency updates. The desktop application can use a local runtime or connect to a control plane that runs locally or in the cloud. The control plane is self-hostable through Docker, AWS, GCP, Azure, Kubernetes, or air-gapped deployments. The repository lists Rust, Node.js, and pnpm as source-build requirements and is licensed under AGPL-3.0.
Tiger Data is a PostgreSQL data platform for combining transactional and analytical workloads in one database. It uses hybrid row and columnar storage, compression, and continuous aggregates, and is positioned for time-series workloads at scale. Tiger Cloud provides a fully managed PostgreSQL service with separated reads and writes through replica sets, plus tiered SSD and S3 storage.
MiniMind is an open-source tiny large-language-model project and end-to-end training tutorial developed by Jingyao Gong. Its main dense model is approximately 64M parameters, and the repository provides the model architecture, tokenizer, datasets, inference code, and training pipeline. The pipeline covers pretraining, supervised fine-tuning, hand-written LoRA, DPO, PPO, GRPO, CISPO, model distillation, tool calling, adaptive thinking, and agentic reinforcement learning. Core algorithms are implemented directly with native PyTorch rather than relying on high-level abstractions from third-party training libraries. Its agentic reinforcement-learning path performs multi-turn rollouts, executes generated tool calls, appends tool observations to the context, calculates trajectory-level rewards, and updates the policy; rollout can use local PyTorch generation or an SGLang server. MiniMind supports dense and mixture-of-experts variants, single- and multi-GPU training, YaRN-based RoPE length extrapolation, Transformers-format models, and inference through llama.cpp, vLLM, or Ollama. The repository also includes a Streamlit chat interface and a lightweight OpenAI-compatible API server with tool-call and reasoning fields. It is released under the Apache License 2.0.
claudish-to-english is a Claude Code plugin that uses a language model to produce a plain-language rewrite of each assistant message while leaving Claude's reasoning and saved transcript unchanged. Its display hook buffers streamed message chunks, rewrites the completed message through Ollama by default or through the Codex CLI, Anthropic API, or an OpenAI-compatible API, and displays the result in append or replace mode. It follows the input or configured output language and supports configurable rewrite styles, prompts, models, and runtime toggles through the /claudish command and environment variables. An optional PostToolUse hook rewrites Markdown files in a configured directory into a sibling .plain.md file or in place. The hooks fail open: provider errors, missing dependencies or keys, timeouts, unavailable models, and capped outputs leave the original message or file unchanged.
terminal-code is a command-line tool that runs the VS Code interface inside a terminal. It combines terminal-browser, a browser rendered in the terminal, with code-server, which provides VS Code in a browser; the terminal renderer uses the Kitty graphics protocol to draw the interface. The `tode` command opens folders or files, supports workspace, split-pane, diff, extension, theme, source-control review, and timing options, can serve or shut down code-server, and supports importing settings from compatible editors. Its SSH mode keeps the VS Code frontend local while running the backend remotely, proxying network requests over SSH. It can manage code-server installations and upgrades. It is distributed for macOS and Linux through an installation script; Windows is supported only through a compatible terminal and Linux under WSL, without an official Windows build.
Gemma Translator is an open-source, on-device voice-translation application designed to run on a Raspberry Pi 5 with 8GB RAM. It runs the Gemma 4 e2b model locally through LiteRT-LM and uses Moonshine for speech recognition and text-to-speech, providing a two-lane interface for conversations between two people. Each lane records speech, transcribes it, translates it with Gemma, and speaks the result in the other lane’s selected language. The React/Vite frontend communicates with a Python API server and is styled for small displays such as 480×320 kiosk screens. The repository includes model-download, startup, and deployment scripts, production Raspberry Pi OS systemd kiosk service registration, Chromium autostart, and support for development or production modes and landscape active-person or vertical two-operator keyboard layouts. After setup and model download, inference runs without an internet connection.
Scroll Craft is a Claude Code skill for building scroll-driven websites. It treats scroll as the page timeline, supporting frame-by-frame video scrubbing, pinned sections, sideways rails, line-by-line headline assembly, shifting page backgrounds, and pointer-driven interactions. Each build selects one of eight mutually exclusive page grammars and must create a site-specific signature interaction; a fingerprint gate requires it to differ from previous builds across at least four of six dimensions, including grammar, navigation, hero, act shape, ending, and signature move. The skill also applies design and validation rules for typography, spacing, color, depth, motion, feeling curves, and a single engineered peak, while refusing patterns such as identical feature-card grids, gradient text, invented statistics, fake dashboards, and AI-purple gradients.
Utopia is a self-hosted enterprise knowledge platform from DeepLethe, implemented as a Rust binary with a PostgreSQL service using pgvector. It ingests documents and synchronized web, RSS, GitHub, and Jira sources, then combines Tantivy full-text search with pgvector retrieval through reciprocal-rank fusion; its chat answers include inline citations linking to source passages and can use OpenAI-compatible endpoints. Its knowledge model is a bitemporal graph: extracted entities and facts retain validity intervals, evidence, and revision history, so corrections link new facts to closed versions instead of overwriting them. An editable ontology drives three-stage entity resolution, temporal Datalog forward-chaining, provenance-preserving derivation, conflict detection, and ontology-driven queries over mounted PostgreSQL databases. Low-confidence extractions, entity-merge candidates, and cardinality conflicts enter a review queue, while decisions such as confirmations, rejections, merges, and graph rebuilds are recorded in a queryable ledger. The application includes a browser console, graph browser, ontology workbench, and multi-user knowledge-base permissions, and can be deployed offline with Docker. The repository identifies the project as version 0.1 and licenses it under Apache-2.0.
slotstream is an open-source macOS local-LLM inference engine that runs Qwen3.8-Flash-Next, a 125-billion-parameter mixture-of-experts model, on Apple Silicon Macs whose memory cannot hold the model. Implemented as a single Swift binary using MLX, it keeps the dense trunk resident and reads routed expert weights from an SSD into a shared cache of slots, sizing that cache to available memory instead of loading the full model. It supports prefix caching for follow-up turns and exposes Ollama-compatible and OpenAI-compatible chat APIs, as well as a command-line prompt mode; the documented API subset supports streaming, CORS, and sampling options but not tools, images, JSON-schema output, or logprobs. The project requires macOS 14 or later, Apple Silicon, and roughly 110 GB of free disk space for the model weights. Its installer provides signed release binaries, and the repository can also be built with Swift and the macOS Command Line Tools. The source is MIT-licensed; the model weights are distributed under the Qwen community license.
Codewhale is an open-source coding agent for the terminal, built in Rust and independently maintained. It reads repositories, edits files, runs commands, inspects results, and works toward user-defined goals. It connects to hosted providers or local models through Ollama, vLLM, or SGLang, supports switching providers and models during a session, and can run interactively in a terminal UI or non-interactively with the `exec` command. Its control mechanisms include read-only planning, Ask, Auto-Review, and Full Access approval modes; `/undo` reverts the last turn and `/restore` returns the workspace to an earlier snapshot. It also supports saved sessions, durable goals, reviewable workflows, agent coordination, MCP servers, skills, hooks, and configurable agent roles. The project runs on the user's machine with the access granted to it, with optional operating-system sandboxing where supported, and is distributed under the MIT license.
TokensBurned is a privacy-first tool that publishes AI coding activity on a GitHub profile through a live SVG card. It collects token counts and model metadata from AI coding harnesses, reduces raw sessions locally, aggregates usage into 15-minute buckets, and serves totals, activity heatmaps, harness/provider/model comparisons, and an optional site-wide rank. It supports native or official integrations for Claude Code, Codex, Gemini CLI, Cline CLI or SDK, and GitHub Copilot CLI, with OTLP or standalone CLI fallbacks for other harnesses. The tool uploads aggregate token counts, harness/provider/model metadata, hashed session identifiers, time buckets, and request counts, while its stated privacy boundary excludes prompts, responses, source code, tool payloads, repository names and paths, transcript files and paths, and API keys. Public display is a separate explicit opt-in tied to a verified GitHub account; users can build cards with selectable layouts, themes, heatmaps, comparisons, and rankings. It is distributed through harness plugins and an npm CLI, installs no cron job, daemon, proxy, or Git synchronization task, and is licensed under the MIT License.
LightNav-0 is an open-source generalist embodied-navigation model from the Light Origins Team. It uses a pretrained Qwen3-VL backbone with a single egocentric RGB stream and natural-language instructions to control humanoid, quadruped, wheeled, and aerial robots, transferring across tasks, embodiments, and scenes without per-task or per-benchmark fine-tuning. The model extends the vocabulary with dual-channel pointing tokens and residual vector-quantized action tokens rather than adding navigation-specific prediction heads. At each step, it emits an affordance point representing a feasible local direction or waypoint and an object point representing the goal, followed by three action tokens that decode to ten future SE(2) waypoints for an embodiment-specific low-level controller. Its temporally aware history compressor samples older frames less frequently, pools them more coarsely, and preserves ordering with timestamp tokens while bounding the visual context. The repository includes a released checkpoint and action decoder, command-line prediction, a vLLM-based server with WebSocket streaming, a MuJoCo simulation demo, evaluation harnesses, and a ROS 2 deployment stack with adapters for the Unitree Go2 and LimX TRON 1. The project is released under the Apache License 2.0; the included EVT-Bench is separately licensed CC BY-NC-SA 4.0.
TrustMeBro is a Go command-interception tool for controlled red-team testing of coding agents such as Codex, Claude Code, and pi. It uses PATH shims to intercept configured command-line tools without requiring a plugin, hook, or MCP integration; rules can spoof generated or fixed output, rewrite stdout from a real command while preserving stderr and its exit status, pass through to the real binary, or reject the call. Each decision is recorded in a timestamped JSONL audit log. Rules match command names, domains, DNS record types, argument globs, and regular expressions. The bundled DNS generators support dig, nslookup, and host, while custom shims can use fixed output, exit codes, and standard streams. On Linux, lab mode uses Bubblewrap to shadow PATH lookups and discovered absolute paths so an agent cannot bypass interception merely by invoking a resolved system binary. TrustMeBro is distributed as a single MIT-licensed Go binary with installation, configuration validation, status, rule-listing, and uninstall commands. Lab mode requires Linux and Bubblewrap and is explicitly an interception namespace rather than a security sandbox: it reuses the host filesystem, workspace, network, environment, and agent credentials. Outside lab mode, absolute paths, changed PATH environments, and in-process DNS clients can bypass command shims.
HexStellar is an agent-first computational platform whose Cortex service executes structured optimization, decision, scientific-computing, and verification requests through a Python CLI and API. An agent submits a JSON formulation for problems such as QUBO/Ising optimization, maximum cut, traveling-salesperson routing, facility assignment, mixed-integer optimization, selection, ranking, scheduling, coloring, and business-rule feasibility; the managed service returns a structured result with execution metadata, a receipt, and an assurance label distinguishing certified optima, heuristics, operations, and abstentions. The CLI also provides free validation, estimation, service re-checks, local witness recomputation for supported families, reproducible seeds and model versions, batch and compressed-binary transport, and an MCP server over standard input/output, with read-only mode for free analysis and verification tools. The public package is a zero-dependency Python thin client: the proprietary solver runs on HexStellar-managed infrastructure rather than inside the package. A separate enterprise runtime is licensed for compatible customer-controlled compute paths under NDA; it has no public runtime download in version 1.0.
Editable Visual Design is an open-source, coding-agent-driven toolkit for creating editable visual artifacts from prompts. A persistent coding agent interprets a brief, plans a composition, generates assets, implements the design in semantic HTML, renders and observes the result, and applies repairs. An image model supplies visual direction for composition, hierarchy, color, and spatial relationships, while the delivered artifact rebuilds typography and layout in HTML rather than shipping reference pixels. The resulting artifact contains real text, independent assets, selectable and movable layers, a visual editor, rendered PNG output, an animated layer breakdown, and a replayable creation process. Deterministic checks cover the canvas, fonts, layer contracts, rendering, and editor round trips. The repository distributes the workflow as two Codex skills, including editable-design for fixed-canvas designs and a separate html-to-pptx skill for converting compatible HTML designs into editable presentations.
Magnitude is an open-source local inference server and CLI for running language models with AI agent harnesses. It profiles a machine's chip, memory, and bandwidth, recommends compatible models, downloads the selected models, tunes inference settings such as speculative decoding and concurrency, and runs models on demand. Models are loaded when needed and unloaded when idle or when memory is constrained; prompts, files, and models remain on the local machine, allowing offline operation after setup. It integrates with Pi, OpenCode, Hermes, OpenClaw, Codex, Claude Code, Oh My Pi, and Cline, and also provides a built-in harness. Magnitude supports macOS and Linux, with Windows support through WSL, and is licensed under Apache 2.0.
Chrome DevTools for agents (chrome-devtools-mcp) is an MCP server from the Chrome DevTools project that lets coding agents control and inspect a live Google Chrome or Chrome for Testing browser. It uses Puppeteer for browser automation, including navigation and actions that wait for results, and exposes DevTools capabilities for recording and inspecting performance traces and insights, analyzing network requests, taking screenshots, and inspecting browser console messages with source-mapped stack traces. A CLI is also provided for use without MCP, with configuration options including headless, isolated, slim, concurrent-session, persistent-profile, WebSocket, and Android-debugging modes. It requires Node.js LTS, a current stable Chrome version or newer, and npm. Browser contents are exposed to MCP clients; performance tools may query the Google CrUX API, and usage statistics and update checks are enabled by default with an opt-out flag.
Monocolor Editorial Print is an agent skill for generating one-ink or controlled two-ink editorial images for posters, zines, portraits, packaging, and visual field notes. It accepts a theme, phrase, object, article idea, or supplied photograph and produces a raster image when image-generation tools are available, the exact production prompt, and a short recipe describing the print mode, palette, layout, typography, process, and originality changes. The skill identifies the subject and intent, selects a layout family, assigns one or two ink plates, composes the page with visible paper and an asymmetric grid, and generates and inspects the result. Its visual system uses adaptive white, gray, or beige substrates; halftone, risograph grain, cyanotype exposure, or photocopy breakup; responsive typography; and controlled negative space. Without image generation, it returns the prompt and states the limitation.
cn is a JavaScript engine for conditionally joining Tailwind CSS class names and resolving conflicting utilities. It is a drop-in replacement for clsx and tailwind-merge, preserving their APIs and output semantics while learning repeated variadic call sequences so stable calls can skip redundant work. The package has no dependencies and is framework-agnostic, running in browsers, Node, Bun, Deno, edge runtimes, and server templates for frameworks including React, Vue, Svelte, Solid, and Astro. It supports configurable Tailwind class groups, prefixes, themes, and validators through cn/config, and includes a CLI for build-time processing and a lighter strings-only variant. It is maintained by aidenybai and shadcn.
Google's cloud computing platform, providing AI and cloud computing services for security, data management, and hybrid and multicloud environments. In the cited video, GCP is discussed as a source of rented GPU spot instances when a company’s owned or contracted capacity is insufficient.
Sierra is an AI agent platform for businesses. Its agents use tool calling to handle customer support, sales, and other pre-sale and post-sale workflows, with an outcome-based pricing model.
ABYSSAL is an open-source, browser-based real-time ocean and extreme-weather simulation built with Three.js, WebGL2, GLSL3, and Vite. It procedurally generates the ocean, clouds, rain, spray, foam, lightning, hurricanes, waterspouts, whirlpools, rogue waves, and tsunamis on the GPU without external textures, meshes, HDRIs, or sound files. Its ocean uses three FFT wave cascades driven by a JONSWAP spectrum and a per-frame butterfly IFFT; a projected screen-space grid renders the water, while analytic disaster height fields deform its surface. Volumetric clouds are raymarched from Perlin-Worley noise and a procedural weather map, with atmospheric-scattering lookup tables providing shared lighting for the sky, clouds, ocean, and spray. Temporal anti-aliasing, adaptive quality scaling, bloom, depth of field, motion blur, and tone mapping form the post-processing pipeline. The browser experience includes an automatic cinematic storm sequence and a sandbox with free-flight controls.
Pictaria Server is a self-hosted photo-intelligence and curation server for Immich libraries, developed by Pictaria. It provides a private web dashboard for collection insights, optional AI enrichment, human photo review, and Smart Albums that synchronize saved Immich searches on a schedule. Its curation workflow groups photos from the same moment into similarity stacks and suggests which photos to keep using frame-worthiness scores; an optional AI referee can compare grouped photos, while human decisions take precedence. Enrichment classifies selected images with operator-hosted Ollama, LM Studio, llama.cpp, or other OpenAI-compatible endpoints, as well as supported cloud providers. Smart Albums can apply saved searches and a Best of mode that ranks results using enrichment data, curation decisions, and photo scores. The server can run as a Docker container or directly on Node.
HyperFrames is an open-source framework for turning HTML, CSS, media, and seekable animations into deterministic MP4 videos. It can be used locally through a CLI, by AI coding agents through skills, or as a rendering core for hosted authoring workflows. Compositions are defined as HTML with data attributes for timing, tracks, media, and sub-compositions. The renderer seeks each frame in headless Chrome, uses animation adapters such as GSAP, CSS, Lottie, Three.js, Anime.js, or WAAPI, and encodes the result with FFmpeg; its toolchain includes preview, linting, inspection, local rendering, browser-based editing, reusable catalog components, and AWS Lambda rendering. The project requires Node.js 22 or newer and FFmpeg, does not require React or a build step, and is licensed under Apache 2.0. Its agent skills support workflows including product videos, explainers, pull-request videos, captions, motion graphics, slideshows, music-driven videos, and general compositions.
Ponytail is a rule set and plugin for AI coding agents that directs them to implement the smallest working solution. Its decision ladder skips speculative requirements, reuses code already in the repository, prefers the standard library and native platform features, avoids adding dependencies, and then chooses the shortest implementation that works. The project says that validation, error handling, security, and accessibility are not simplified away. It provides chat commands for changing enforcement intensity, reviewing the current diff for over-engineering, auditing an entire repository for bloat, recording deferred shortcuts in a debt ledger, and viewing benchmark results. It can be installed across multiple coding-agent environments, including Claude Code, Codex, Copilot CLI, Gemini CLI, and Pi; the page also lists support for other agents. The project reports lower code volume, token use, cost, and execution time in benchmarked Claude Code sessions editing a FastAPI and React repository, while noting that results vary by task and model.
Nginx Proxy Manager is a Docker-based web application with a graphical interface for managing Nginx reverse-proxy hosts, routing requests to services on a private network, supporting WebSocket connections, and managing free Let's Encrypt SSL certificates with automatic renewal. It requires a database and supports multiple users with permissions to view or manage assigned hosts.
DeskcommCRM is an open-source, self-hosted CRM and AI sales operating system for businesses that sell through WhatsApp and other chat channels. It combines a multi-tenant CRM and sales pipeline with inbox, Kanban pipelines, contacts, team governance, audit logs, LGPD features, webhooks, automations, scheduling, and human takeover capabilities, and supports WhatsApp through WAHA QR connections or the official Meta Cloud API. Its AI agents use tenant-specific RAG, organizational memory, executable skills, intent routing, sentiment analysis, follow-ups, and audited handoff to human attendants. Agents can operate CRM records such as leads and funnel stages, while the CRM exposes an internal MCP interface for agent operations. Tenants can receive leads through public webhook endpoints and process event-driven QUANDO/SE/ENTÃO rules through an event-log queue drained by scheduled workers. The application is built with Next.js, TypeScript, and Supabase/Postgres.
PR Lens is a pull-request visualization tool by Coldtea that generates animated architecture and data-flow diagrams and posts them inside GitHub pull requests. It maps a change's components and call paths, uses green for new elements, amber for changed elements, and red for removed elements, and provides step-by-step walkthroughs, nested diagrams, and interactive canvases with pan, zoom, and light or dark themes. It is available as a GitHub App, GitHub Action, CLI, or coding-agent skill. The CLI analyzes a diff against a merge base, produces a graph document, renders theme-paired SVGs, validates documents against a schema, and can compose the pull-request comment. The GitHub Action can use Gemini, OpenAI, or an OpenAI-compatible endpoint; the agent skill lets a coding agent generate and attach the diagrams. Repository configuration in .github/pr-lens.yml supports renames, exclusions, lane pins, and groupings. The repository documents Node 20.11+ and pnpm 10 requirements and is licensed under the MIT License.
BankMCP is a self-hosted, read-only Model Context Protocol server that lets AI assistants query a user's bank accounts through Enable Banking's PSD2 open-banking API. The user authenticates at the bank through a consent flow; BankMCP stores consent and account metadata but does not store balances or transactions, make payments, or send telemetry. It exposes MCP tools for managing bank connections and accounts, retrieving balances and paginated transactions, and creating rules that poll accounts for conditions such as low balances or matching payments and send webhook notifications. It can run locally or on a server, is distributed as the npm package `bankmcp`, supports standard MCP clients including Claude and Ollama, and is licensed under the MIT License.
NiubiGEO is an open-source, self-hosted tool for tracking how AI models describe products and brands, which competitors they name, and which sources they cite. It accepts a domain, queries selected OpenRouter models with optional web search, and stores each model's original answer, descriptions, competitor associations, keywords, citations, ordinary URLs, errors, and execution conditions. Users can confirm model-derived competitors and keywords for follow-up tests that omit the target brand's name, repeat measurements, and schedule monitoring runs. Results are organized into separate projects for different domains and retain historical records so findings can be traced back to the underlying answers and evidence. The project states that it does not track traditional search-engine rankings and that offline and web-enabled tests should be interpreted separately. The Community Edition is free, open source, and licensed under Apache-2.0. It runs locally with Node.js 22 or later, requires an OpenRouter API key for testing, and can also be deployed with Docker; users pay the model, search-service, and hosting costs they incur. The project also offers paid AI testing by real people and GEO optimization services.
OKF Agent Memory is a Git-native persistent memory layer for AI coding agents, developed as a pure Go CLI and library. It stores project knowledge as plain Markdown files with YAML front matter in a repository, following the vendor-neutral Open Knowledge Format (OKF) v0.2 rather than relying on an external vector database. The tooling parses and validates OKF bundles, builds an in-memory BM25 index for local lexical search, and provides progressive disclosure through hierarchical index files and link graphs. Its search-before-write workflow queries existing concepts before creating new ones; concepts can include provenance, trust tiers, lifecycle metadata, and code references for linking knowledge to source files. The standalone executable also supports creating and updating concepts, bootstrapping the memory structure into a project, and running an embedded Model Context Protocol (MCP) server over stdio for agent platforms. The repository reports sub-300-microsecond concept search and approximately 4-millisecond graph validation for its benchmark cases, with no external databases or API costs for retrieval. It is distributed under the MIT License.
Bot Crossing is a local web application that visualizes coding-agent threads as an astronaut colony. It reads session records from installed agent harnesses on the user's machine, maps repositories to persistent hex zones, and represents sessions as astronauts whose behavior reflects states such as running, waiting, errored, merged, archived, or inactive. Selecting an astronaut or repository opens its associated thread through the owning harness, while the application can also start sessions, reveal folders, mark threads viewed, and archive them. The application uses per-harness adapters to scan local session files and merge them into a common thread model. It currently documents support for Claude Code, Codex, and Cursor, with adapters for other harnesses able to be added under `server/harnesses/`. Astronaut navigation uses a rasterized navigation grid, A* routing, string-pulling, collision handling, and crowd separation. The browser renders the colony with Three.js techniques including instanced GPU-skinned astronauts, shader-based construction progress, procedural terrain and surfaces, and configurable planets, lighting, and quality settings. Bot Crossing runs locally through a Vite development server or a built Node server, binds to loopback by default, and stores the colony layout in `data/colony.json`. It does not upload session data or use an account; it reads harness records and writes only the colony state plus the documented archive field. The repository requires Node 22.13 or newer, supports macOS, Linux, and Windows, and is licensed under MIT. Bundled Kay Lousberg assets are separately covered by CC0, while bundled Material Design Icons use Apache-2.0.
anything2explainer is a Claude Code and Codex skill for producing narrated explainer videos from a topic or document. It outputs a 1280×720 H.264 MP4 in Chinese or English, with synchronized TTS voiceover, word-aligned subtitles, chapter cards, a top HUD, and a chapter progress bar. Every frame is drawn as code with Remotion, React, and TypeScript rather than generated video or stock footage. Its pipeline researches the subject and records sources, writes the narration, generates voiceover and a frame-accurate timeline, storyboards each shot, dispatches parallel agents to build Remotion components, renders the film, and runs quantitative frame metrics plus chapter-level quality-control and fix passes. The repository includes a compilable Remotion template, visual primitives, style and motion specifications, scripts for voiceover, storyboarding, rendering and QC, and a complete reference film with its research and production records. It is not a standalone CLI; it is a skill and production method intended for AI coding agents. The toolkit is licensed under PolyForm Noncommercial 1.0.0: noncommercial use is free, while commercial use requires prior authorization. The videos produced with it are described as belonging to their creators. It renders through headless Chromium on the CPU, supports landscape output rather than 9:16 video, and provides two backdrops within its specified visual style.
Whiteboard Animator is a CPU-only Python package and command-line tool that converts a finished whiteboard-style image into a hand-drawn reveal video. It can write detected text word by word, trace outlines, fill shapes with brush strokes, and decompose branched line art into sequential pen paths; narration audio can pace the drawing, and multiple scenes can be concatenated into one MP4. It is the render engine behind the Whiteboard format at Kinoslide. The engine identifies connected non-white pixel components and uses a bundled CRAFT text detector to distinguish text from drawn shapes. It orders components using relationships such as containers before contents, shapes before labels, reading order for text, and attachment of small dots to glyphs. It assigns reveal times based on component area or a region plan, then follows skeletons and outlines with reveal fronts, brush strokes, or sequential paths before streaming frames to FFmpeg. Rendering requires no GPU, API key, training, or model beyond the bundled text detector; an optional Gemini-based region-detection feature requires a Gemini extra and GOOGLE_API_KEY. The package provides the `whiteboard-animate` CLI and a Python API, requires Python 3.10 or later plus FFmpeg and ffprobe, and is distributed under the MIT license. It is intended for clean marker-style drawings on white backgrounds rather than photos, gradients, or textured backgrounds.
DeepSelect is a high-performance CUDA and PyTorch implementation of TopK kernels for DeepSeek Sparse Attention and sampling workloads. It replaces vanilla `torch.topk` for supported inputs, covering bfloat16 lightning-indexer workloads and float32 sampling workloads, with configurable sorting, index type, value output, variable-length rows, and NaN handling. The repository reports 2–20× speedups against `torch.topk` on its benchmarks. It is distributed as an installable Python package from the `deepseek-ai` GitHub repository. Supported TopK values are limited to 4096 or less, and input tensors must meet specified stride and contiguity requirements; the optimized sampling case targets vocabularies of approximately 128K.
MediaPipe is a computer-vision tool that detects facial expressions and gestures, including for video-call reactions.