1,038 tools and products — trending open source, and what gets used in AI and other work.
A personal AI learning system distributed as a pi configuration. It encodes a teaching philosophy and learning process in skills, including teaching and diagram-based visualization, and adds extensions for question popups, graded quizzes, Markdown session logs, and visualization tools. The configuration delegates research and visual creation to researcher, SVG-maker, and Mermaid-maker subagents; it can also run without subagents, with those delegation-based capabilities omitted. It is intended for one learner and is shared as-is under the repository's stated installation instructions.
An OpenAI API for real-time interactions, powering voice interaction with a 3D globe.
LiveKit is an open-source framework and developer platform for building, testing, deploying, scaling, and observing real-time voice, video, and physical AI agents. Its agent pipeline streams user speech from an app, browser, or phone call to an agent, which applies custom business logic and returns a response; the platform supports automatic turn detection and interruption handling, speech-to-text, language-model, and text-to-speech providers, web and mobile applications, and telephony through phone numbers and SIP integrations. LiveKit Cloud provides deployment and scaling on LiveKit's real-time infrastructure, alongside an inference gateway and full-stack observability for agent sessions.
An AI framework for defining model architectures and supporting model serving.
A downloadable web-development project bundle from GreatStack for building a full-stack AI website builder with MongoDB, Express.js, React.js, and Node.js. The project includes starter assets and source files for a React website generator that accepts text prompts, builds websites step by step, exposes generation progress, and supports manual source-code editing, follow-up AI prompts, exporting, and publishing. Its tutorial project includes user authentication, REST APIs, an OpenRouter model integration, and an agent chat API for updating generated projects.
Heretic is a command-line tool for automatically removing safety alignment from transformer-based language models without post-training. It combines directional ablation (abliteration) with a Tree-structured Parzen Estimator optimizer powered by Optuna, co-minimizing refusal counts and KL divergence from the original model to select ablation parameters automatically. For supported transformer components, currently attention output projections and MLP down-projections, Heretic computes per-layer residual directions from the difference between first-token hidden states for harmful and harmless prompts, then orthogonalizes the associated weight matrices against those directions. Its optimizer can interpolate between residual directions and select separate, flexible layer-weight kernels for different components. The tool supports most dense models, many multimodal models, several mixture-of-experts architectures, and some hybrid architectures; pure state-space models and certain research architectures are not supported out of the box. Heretic runs in a Python 3.10+ environment with PyTorch 2.2 or later, supports optional bitsandbytes 4-bit quantization, and can save or upload generated models, launch a chat evaluation, and run standard benchmarks. An optional research installation provides residual-vector visualization using PaCMAP and residual-geometry analysis. The project is distributed under the GNU Affero General Public License version 3 or later.
Crawl4AI is an open-source Python web crawler and scraper that converts web pages into structured, LLM-ready Markdown for retrieval-augmented generation, agents, and data pipelines. Its asynchronous Playwright-based crawler supports Chromium, Firefox, and WebKit, dynamic JavaScript pages, sessions, persistent browser profiles, cookies, headers, proxies, screenshots, media, iframes, lazy loading, full-page scanning, caching, and deep crawling with BFS, DFS, and best-first strategies. For extraction, it provides heuristic Markdown filtering including BM25-based relevance filtering, CSS- and XPath-based schema extraction, chunking and cosine-similarity strategies, and optional LLM-driven structured JSON extraction. It also includes adaptive crawling, link analysis, URL seeding, virtual-scroll handling, anti-bot and proxy escalation features, and customizable hooks. Crawl4AI can be installed with pip and used through Python or its command-line interface. It is also distributed as a Dockerized FastAPI server with JWT authentication, browser pooling, monitoring dashboards, a playground, and endpoints for crawling, HTML extraction, screenshots, PDF generation, and JavaScript execution. The repository states that it is licensed under Apache License 2.0.
GitNexus, developed by Akon Labs, is a code-intelligence engine that indexes repositories into a knowledge graph for code exploration and AI-agent context. Its indexing pipeline walks the file tree, parses source with Tree-sitter, resolves imports, calls, inheritance, constructor-inferred receiver types and other relationships, groups symbols into functional communities, traces execution processes, and builds BM25-plus-semantic hybrid search indexes backed by LadybugDB. The resulting graph supports MCP tools and CLI commands for process-grouped search, symbol context, call-path tracing, blast-radius and Git-diff impact analysis, structural checks, coordinated renaming, API and route mapping, taint and dependence queries, and Cypher access; repository groups can link contracts and impacts across multiple repositories. The CLI runs locally and can connect editors such as Claude Code, Cursor, Codex and others through MCP, skills and selected hooks. GitNexus also provides a browser-based WebAssembly UI with an interactive graph explorer and AI chat, plus a local HTTP server and Docker deployment mode for accessing indexed repositories through a backend. The web-only mode keeps repository processing in the browser and is constrained by browser memory, while the native CLI stores indexes locally in each repository's .gitnexus directory and uses a global registry for multi-repository access.
Co-Invest is Liquid's AI trading assistant for researching markets, sizing positions, and placing trades through Claude, ChatGPT, iMessage, a Chrome extension, or the Liquid account. It uses market data such as positioning, funding, liquidation maps, on-chain flows, news, macroeconomic events, earnings, and ETF flows to produce trade ideas with a direction, size, named catalyst, and source links. It supports crypto, stocks, ETFs, indices, commodities, foreign exchange, and other markets available through Liquid, including long and short positions, multipliers, stop-losses, and take-profits. Trades require explicit user confirmation: the assistant returns a confirmation card containing the symbol, direction, size, multiplier, and rationale, and the user taps to confirm or cancel. According to the product page, Co-Invest can read a portfolio and propose trades but cannot transfer funds, change account settings, or submit an order without confirmation. The page distinguishes this assistant from Co-Invest Computer, which it identifies as Liquid's option for scheduled, hands-off execution; the video description presents scheduled trading routines as part of Liquid's AI trading offering. Co-Invest itself is described as free, with Liquid's standard trading fees and applicable perpetual-futures funding rates. It also provides a paper-trading mode using simulated balances against live market data.
Suno is an AI music generator and web-based generative audio workstation from Suno, Inc. It creates complete songs from text prompts or user-supplied melodies, lyrics, audio, moods, genres, and themes, including vocals, instrumentation, and production. The service can be used to create customized local-business jingles, including versions tailored to particular markets. Users can regenerate, extend, edit, remix, reorder sections, rewrite lyrics, record or upload audio, and control aspects such as voices, vocal gender, exclusions, weirdness, and style. Suno Studio combines AI music creation with digital-audio-workstation functionality; generated tracks can be separated into time-aligned WAV stems for use in other audio software. Free accounts provide limited daily song creation, while paid plans add commercial rights and expanded creation, editing, stem-separation, and Studio features.
FuXi is a self-contained terminal AI coding agent developed by FUXI. Built in Go and distributed as a static binary, it uses a Think → Act → Verify loop to read and edit code, run shell commands, drive tools, connect to MCP servers, and route requests across multiple LLM providers with automatic failover and cost-aware settings. Its built-in capabilities include file operations, shell execution, code search, web fetching, LSP diagnostics, Jupyter, browser use, background tasks, and parallel sub-agents. Shell commands pass an AST-based safety classifier, while permissions and audit logs govern autonomous actions. Sessions persist to disk, with checkpoints for resuming, rolling back, or forking; the TUI also supports memory consolidation and automatic context compaction. FuXi supports provider API keys or FuXi OAuth, configurable OpenAI-compatible endpoints, MCP clients, hooks, skills, plugins, and slash commands. The repository contains documentation, installers, release information, and issue-tracking materials; it states that the product source is proprietary and not published.
MyContext is a local-first desktop app from openTrinity that builds a private personal work-context layer from sources such as instant-messaging conversations, documents, and meeting records. It stores local copies, indexes, source references, and derived context in an on-disk SQLite vault, then organizes them into a personal context graph linking people, projects, topics, events, conversations, and supporting facts. Its search and answer workflow combines local full-text search, semantic retrieval, and graph queries, with agents assembling answers from traceable source material and falling back to ranked local results when the agent runtime is unavailable. A digital-self workflow recalls relationship-specific context and communication history to draft replies, while sending, deletion, and other consequential actions require explicit user confirmation. The repository describes an Electron and React desktop architecture with source connectors, incremental ingestion, context processing, retrieval, knowledge-graph, persona, and isolated agent-runtime layers. MyContext is in developer preview and under active development; its README warns of compatibility-breaking changes and migrations that may require recollection. It is licensed under the Elastic License 2.0, which permits use, modification, and self-hosting but restricts offering it as a hosted or managed service to third parties.
JoyAI-Video-Edit is an open-source, instruction-guided system for editing live camera streams or uploaded videos as frames arrive. It processes frames causally without waiting for the complete video, requiring a predefined sequence length, or revisiting future frames, and supports subject and local edits, background replacement, style and motion changes, and reference-guided editing. Its autoregressive diffusion architecture combines an MLLM-based condition encoder, a causal video VAE, and a 16-billion-parameter multimodal diffusion transformer. The deployment uses aligned autoregressive distribution-matching distillation, long-horizon optimization, bounded KV-state inference, and deployment-oriented scheduling to reduce train–inference mismatch and temporal drift during streaming generation. The repository reports 30 FPS end-to-end throughput at 720 × 1248 in its deployment benchmark. The repository provides deployment code and model checkpoints, a local server with a browser interface, and instructions for CUDA-based inference. It is licensed under Apache 2.0.
screenpipe is a local AI-agent memory layer that continuously captures computer history on macOS, Windows, and Linux. It records screen frames with OCR, accessibility data, microphone and system audio with transcripts, and application activity, storing the underlying history locally. Agents can search the history through a local REST API, database, or MCP server and use it for tasks such as meeting summaries, follow-ups, and workflow automations. Users can exclude apps, windows, URLs, or time periods and redact sensitive fields on the device; the page describes the project as source-available.
Gemini Deep Think is an AI reasoning system used in mathematical research. In the cited example, a mathematician used it to prove lemmas after human experimentation led to a stronger problem statement.
OpenClaude is an open-source, terminal-first coding-agent CLI for cloud and local model providers. It connects to OpenAI-compatible APIs, Gemini, GitHub Models, Codex, Ollama, Atomic Chat, and other supported backends, providing prompts, streaming output, Bash and file tools, grep, glob, agents, tasks, MCP, slash commands, web search and fetch, and image inputs for compatible providers. The CLI supports guided provider setup with saved profiles, conversation continuation and forking, detached local background sessions, model-specific agent routing, repository maps based on PageRank-ranked code structure, and a headless bidirectional-streaming gRPC server for integrations, CI/CD pipelines, and custom interfaces. A bundled VS Code extension provides launch integration, in-editor chat, provider-aware controls, and theme support. OpenClaude runs on Node.js 22 or newer, is distributed through npm and an Arch Linux AUR package, and can use local inference or remote APIs. Its repository is licensed MIT for the project's modifications and states that it is an independent community project derived from and substantially modified from the Claude Code codebase, without Anthropic affiliation.
Skill Cabinet is a local catalog for agent skills installed on a machine. It scans user-level skill locations such as .agents, .claude, .codex, .cursor plugins, Hermes profiles, and other ~/.* /skills folders, then lets users filter skills by drawer, metadata, risk, invocation mode, and status; inspect rendered or source bodies, YAML frontmatter, extra files, symlinks, duplicates, origins, and broken links; and review disk usage. It runs with Node 20 or later through npx skill-cabinet, starting a server bound to 127.0.0.1 and opening the catalog in a browser. Users can quarantine skills to ~/.skill-cabinet/quarantine and restore them, or delete individual skills or groups; deletion removes folders, files, or symlinks from the scanned locations, while a symlink's target is retained. The project is distributed under the MIT license.
SkillRadar is open-source discovery, security, ranking, and routing infrastructure for Agent Skills and Codex. It discovers public SKILL.md files, parses their contents, performs conservative static checks for capabilities such as shell commands, dynamic execution, secret access, networking, package installation, filesystem writes, and deployment tooling, then classifies and ranks candidates in a safety-gated registry. D and Blocked candidates are kept audit-only and excluded from automatic routing. Its Codex plugin provides task-to-Top-3 routing, skill search, provenance and safety inspection, and a read-only Skill Budget Doctor. Routing can use a bundled offline registry and returns relevance, SkillRadar score, security grade, provenance, reasons, and match details without executing candidate repositories or depending on live GitHub discovery. The repository includes radar data, a matching system, a router-quality benchmark, a local registry UI, and daily bot-refreshed generated data; it is released under the MIT license.
genart-skill is a Claude Code plugin that provides an AI coding skill for generative art. It teaches deterministic, hash-seeded compositions, including seeding a pseudorandom number generator from a token hash, using named sub-streams, rendering the same composition at different resolutions, and designing traits and rarity tables for editions. It covers Canvas 2D, p5.js, Three.js/WebGL, and SVG, with guidance on ethics, platform-specific workflows, debugging, keyboard shortcuts, PNG and video export, and print and pen-plotter output. The plugin includes scripts for checking same-machine repeatability, distinctness, global-state isolation, and feature stability; rendering a hash-specific PNG or contact sheet; measuring rarity across a census; and exporting batches of images. The scripts can be run in a project and use Playwright when browser rendering is required. The repository explicitly limits its determinism claims, noting that GPU shader compilers, floating-point behavior, rasterizers, MSAA, and JavaScript engines can prevent cross-machine equivalence. It is installed through the Claude Code plugin marketplace and is maintained by Camille Roux. The repository says that CI tests the scripts against a known-good fixture and a deliberately broken variant, while a monthly check verifies that platform-documentation URLs still resolve. It is licensed under the MIT License.
A Claude Code skill for converting scripts into MiniMax H3 shot lists and directing character performance. It breaks scenes into shots based on emotional-beat density, using controlled comparisons and documented inferences about MiniMax H3 output; its central rule is to split shots that contain too many facial-expression beats because H3 may otherwise produce a frozen face. The skill also covers dialogue-tag effects on frame allocation, camera placement, observable body actions for silent characters, reference-image use for shapes, edit-based handling of pauses and emotional transitions, PSNR-based verification, dubbing constraints, tail degradation, and a six-step breakdown workflow. It is installed with `npx skills add https://github.com/phileiny/h3-storyboard-skill --skill h3-storyboard` or by copying the skill directory into `.claude/skills/`. The repository distinguishes controlled findings from partly verified and inferred rules, and describes its scope as serialized short-form drama produced with local ComfyUI and MiniMax H3.
FixAnything is an open-source video-refinement tool that repairs rendering artifacts from 3D representations, including 3D Gaussian splats, NeRFs, meshes, and sparse point clouds. It repurposes the pretrained Wan2.1 video diffusion model with minimal modification and fine-tuning, using a FixAnything LoRA to turn a rendered camera-path video into a refined video. The inference pipeline accepts a video file or folder of frames; its documented workflow uses 61-frame renderings resized to 832×480 and writes a generated video, the resized input, and a side-by-side comparison. An optional pipeline reconstructs scenes from a small set of photographs with MapAnything, renders the reconstruction along an interpolated camera path, and then refines that rendering. The repository contains inference code and downloads the underlying Wan2.1 model and FixAnything weights; the code and weights are released under the Apache 2.0 License.
Procedura is an open-source agentic 3D-modeling tool from SpatiaOS that converts text prompts into editable procedural assemblies rather than point clouds or triangle meshes. It generates OpenSCAD source with named parts and typed mates, plans and builds parts incrementally, and can use reference images and optional Blender render feedback during refinement. Optional passes assign per-part PBR materials and plan articulation, exporting motion to OpenUSD and URDF with Isaac-based validation. It runs locally as a Bun/TypeScript pipeline using a configurable OpenAI-compatible, Gemini, or local model endpoint; it does not provide hosted inference or API keys. The pipeline uses a Manifold-capable OpenSCAD build to compile the generated programs and Blender for renders, and includes a web Studio for composing runs and inspecting intermediate artifacts. The repository is MIT-licensed.
CDAF is an open sidecar format and toolkit for video that stores a timestamped plain-text description beside the corresponding video file. Its Python library and CLI can generate, parse, validate, read, and report the status of `.cdaf` files, while an agent skill teaches video agents to check for a matching sidecar before processing footage. Each sidecar contains a minimal versioned header with the video filename, SHA-256 hash, byte size, duration, generator, and creation time, followed by sections such as summary, timestamped segments, transcript, on-screen text, and tags. Conforming tools verify the video's freshness and refuse to use a stale sidecar after the video changes; the format is model-agnostic even though the included generator uses the Gemini Files API, with optional local-model support. The repository includes a normative specification, reproducible sidecar-versus-direct-video benchmarks, an agent skill installable with `npx cdaf-skill`, and CLI commands for generation, validation, reading, and status checks. The core validation functions require only the Python standard library; generation requires Python 3.10 or later and a user-supplied Gemini API key. The project is licensed under MIT.
An Apache-2.0-licensed serving setup for Qwen3.8-27B on a single 24 GB NVIDIA GPU, built around vLLM and exposing an OpenAI-compatible API with optional key authentication. Its preparation pipeline starts from a W4A16 AutoRound model, requantizes the language-model head and embedding matrices to int8, requantizes the MTP components, and applies vLLM patches for the serving stack. The runtime combines int8 tensor-core GEMMs with an fp16 recurrent state, continuous batching, split-KV verification attention, and speculative decoding through either Qwen's MTP path or the optional DFlash2 block drafter; DFlash2 can also draft from the request's cached context for document-reproduction workloads. The repository provides Docker Compose profiles for batch throughput and single-user latency, prebuilt container images, model-download and requantization scripts, benchmark and quality-test tools, and launchers for fast, long-context, and experimental KVarN cache modes. The standard configurations target roughly 64k to 150k tokens of context, while KVarN and alternative int4 KV-cache paths extend the claimed capacity to the 256k-token range with lossy KV quantization. The setup requires a recent NVIDIA driver and a compatible Ampere-or-newer GPU; the repository notes that its benchmark figures were measured on an RTX 3090 subject to a 250 W power limit.
cc-prune is a Python command-line tool for surgical context recovery in Claude Code sessions. It removes already-processed tool output from a Claude transcript while preserving conversation text, reasoning, thinking blocks, transcript structure, and tool-call relationships; it does not summarize or truncate the retained content. The workflow measures actual API usage, inspects transcript byte buckets, optionally audits individual results, and clears selected buckets such as tool results, inputs, attachments, or orphaned command output. Clear operations use a character-length threshold, can protect a tail of recent records, create numbered snapshots, validate invariants, and support restoration. Its validation checks include unchanged record and UUID/parent-UUID chains, byte-identical thinking signatures, and matching tool-result and tool-use identifiers. The tool operates on Claude Code JSONL transcripts under the user's .claude/projects directory and requires the session to be stopped before editing. It supports usage verification after pruning, snapshot listing and restoration, and density calculations based on measured token counts. The project states that the Claude Code transcript format is undocumented and version-sensitive, and that cleared results may need to be regenerated; it is distributed under the MIT license and requires Python 3.9 or later with no dependencies.
An Apache-2.0 C++23/GGML port of SkinTokens and TokenRig for automatic 3D mesh rigging on CPU or Vulkan. It takes a static GLB mesh, uses a Michelangelo point encoder, Qwen3-based TokenRig policy, FSQ expansion, condition encoder, and SkinVAE decoder to predict a skeleton and per-vertex skin weights, then exports a rigged GLB. It can also generate weights for an existing static or animated skeleton, inspect GLB assets, and retarget recognized SOMA30 motion to the 52-joint Mixamo order or to a generated humanoid rig. The project provides a command-line interface, shared library, C11-compatible API, and a Go/WebGL demo; converted GGUF model weights are available separately, with optional Vulkan support.
Claude 5.1 is described in the supplied video evidence as an Anthropic AI model for programming, scientific research, and agentic tasks. The video claims that it can work continuously for dozens of hours on codebases and multistep research, with separate lower pricing for cached, ordinary, and complex agent tasks.
Facebook Ads MCP is an MCP server that connects Claude to Facebook advertising workflows. It can be used to create and manage campaigns, ad sets, targeting, and ads.
Lakebed is a cloud product whose codebase was audited, cleaned up, optimized, and improved through parallel pull requests generated by AI.
DBOS is a database-oriented operating-system project and application environment that stores important system state in a database. Its practical focus includes durable, recoverable workflows, particularly workflows used by agentic AI systems.
Academic Research Skills for Claude Code is an open-source suite of Claude Code skills for academic research and publication workflows, maintained by Cheng-I Wu. It provides separate deep-research, academic-paper, academic-paper-reviewer, and academic-pipeline skills for literature reviews, systematic reviews, guided research, drafting, citation conversion, revision, peer review, rebuttal auditing, methodology review, and re-review. The suite uses staged, human-in-the-loop workflows with multi-agent orchestration, Socratic checkpoints, style calibration, writing-quality checks, citation formatting, and outputs in Markdown, DOCX when Pandoc is available, and LaTeX/PDF through tectonic. Its pipeline orchestrator connects activities through ten stages, user-confirmation checkpoints, Material Passport handoffs, claim and citation verification, integrity gates, and final process summaries. Citation verification can cross-check references against Semantic Scholar, OpenAlex, Crossref, and arXiv when available.
Claude-Mem is an open-source persistent-memory plugin and service for coding agents. It captures agent activity through lifecycle hooks, stores sessions, observations, and summaries in SQLite, and uses hybrid full-text and Chroma vector search to retrieve relevant context across sessions. Its MCP search workflow uses progressive disclosure: the agent first searches a compact index, then reviews a timeline, and finally fetches full observations for selected result IDs. A local worker service, managed by Bun, provides the HTTP API, search endpoints, and web viewer; the project also supports integrations with Claude Code and other listed agent environments, configurable context injection, private-content exclusion tags, and optional cloud synchronization. The repository states that it is distributed under the Apache License 2.0 and requires Node.js 20 or later, with Bun, uv, and SQLite used by the runtime. It can be installed through its npx installer or Claude Code's plugin marketplace.
CodeBurn is a free, open-source, local-first tool by AgentSeal that reads session files written by AI coding tools and reports token usage and estimated cost by provider, model, project, task, and activity. It provides a terminal dashboard and reports, a localhost web dashboard, desktop and tray or menubar views, exports, model comparisons, subscription-plan tracking, and optional cross-device aggregation. Its deterministic analyzers classify work into task categories from tool usage and message keywords, calculate token costs using LiteLLM pricing cached locally, and scan coding-agent sessions for patterns such as repeated file reads, low read-to-edit ratios, uncapped shell output, unused MCP servers, bloated configuration, and retry-heavy work. The optimize command produces estimated savings and fixes; applicable configuration changes are backed up and journaled so they can be undone and later compared with observed usage. The yield command heuristically correlates sessions with Git commits to classify spend as productive, reverted, abandoned, or ambiguous. The guard feature installs opt-in Claude Code hooks for soft and hard session-spending caps, checkpoints, and status-line reporting. CodeBurn also exposes local usage and savings through an MCP server over stdio. The CLI reads data from the local machine without wrappers, proxies, API keys, or uploads; the optional desktop applications can send anonymous bucketed telemetry after consent. It requires Node.js 22.13 or newer, and the repository is licensed under MIT.
here.now is an agent-oriented hosting and publishing service for publishing files and folders—including websites, documents, dashboards, presentations, prototypes, games, and media—to the web and receiving a live URL. Any AI agent that can make HTTP requests can publish to it; no account is required, but unauthenticated sites expire after 24 hours, while registered accounts can keep sites permanently. Sites are public by default with randomly generated URLs, and can be protected with passwords or restricted to invited email addresses or domains. The service also supports custom domains and team workspaces with member-only visibility and workspace subdomains.
C-Dance 2.0 is described in the videos as an AI video-generation system that creates scenes using character sheets, image frames, environment references, and voice-reference videos to maintain visual and voice consistency across scenes.
Kling 3.0 is an AI video-generation product or model described as generating comparison video scenes from environment and character reference images.
Google Flow is a visual-creation platform for generating and upscaling character sheets, creating reference-based images, and producing image-to-video clips. Its workflow uses a character reference sheet that can be reused across image and video generations to maintain a character's facial identity, including across scenes with multiple characters, different outfits or lighting, scars, tattoos, and age changes.
Miles is an enterprise-facing reinforcement learning framework for large-scale post-training of large language and vision-language models. It pairs SGLang for high-throughput, agentic rollouts with Megatron-LM for scalable training, and also provides a PyTorch FSDP2 backend for Hugging Face implementations. Its asynchronous architecture decouples rollout and training workers, supports configurable on- and off-policy schedules, and updates rollout engines in-loop through peer-to-peer RDMA weight transfer. The framework includes token-in-token-out data flow, Rollout Routing Replay for replaying mixture-of-experts routing decisions during training, fault-tolerant recovery of failed SGLang engines, low-precision training with MXFP8 and NVFP4 alongside FP8, INT4 QAT, BF16, and FP16, and LoRA or multi-LoRA training. It supports reinforcement-learning recipes including GRPO, GSPO, PPO, and REINFORCE++, as well as supervised fine-tuning, on-policy distillation, agentic environments, and diffusion-model training. Miles was forked from slime and integrates SGLang, Megatron-LM, and torch_memory_saver; the repository is released as an open-source project, though the provided page text does not state its license.
ComfyUI-Easy-Install is a portable installer and desktop management environment for ComfyUI, supporting Windows, macOS, and Linux, with the documented Windows distribution focused on NVIDIA GPUs. It bundles Git, embedded Python, ComfyUI, ComfyUI Manager, and custom nodes used in Pixaroma tutorials, avoiding separate manual setup of Python or Git. Its EZi Desktop app manages ComfyUI and frontend versions, models, packages, PyTorch/CUDA versions, dynamic VRAM settings, UV/PIP caches, GGUF conversion, and optional components such as Nunchaku, SageAttention, FlashAttention, InsightFace, and Trellis 2.0. The installation is kept in a portable folder that can be moved or backed up, and the repository documents add-ons for model-folder linking, hardware checks, version rollback, package pinning, and custom input, output, and user folders.
Easy Use is a custom node pack for ComfyUI that provides nodes used for mathematical and workflow operations.
RG3 is a custom node package for ComfyUI that adds group settings and controls to workflows. The videos identify it as an optional package installed through the ComfyUI Easy Installer.
Marble is World Labs' product for generating spatially consistent, high-fidelity, persistent 3D worlds from text, images, videos, 360-degree panoramas, or 3D layouts. It uses Gaussian splats to represent generated scenes and supports interactive editing, expansion, and combination of worlds. Users can move through and inhabit the resulting environments, then download or export them in 2D and 3D formats for use in other workflows and pipelines.
An AI filmmaking tool from Higgsfield for creating cinematic scenes and short films. It builds reusable character, location, and prop references, then generates connected video clips using those references to maintain continuity. The workflow supports emotional control, dialogue, background sound, video references, and editing of the resulting scenes.
MTPLX is an open-source native Mac app and command-line tool for running local language models on Apple Silicon. It uses a model's built-in multi-token prediction (MTP) heads to draft several tokens, verifies the block in one batched forward pass, and commits tokens with exact rejection sampling and residual correction, without requiring a separate draft model or changing the model's sampling distribution. It provides local OpenAI-compatible and Anthropic-compatible APIs, native chat, model management, hardware-specific draft-depth tuning, and tooling for building and verifying MTP models. The server can also expose MLX embedding and reranking models. MTPLX requires an M1-or-newer Mac running macOS 14 or later, and the repository is licensed under Apache-2.0.
Speechify is an AI voice and text-to-speech platform that reads written material aloud and also provides speech-to-text, conversational, and business-focused AI tools. Its SpeechifyAI service offers text-to-speech and realtime voice-agent APIs through a single interface, with streaming synthesis, zero-shot voice cloning from a consented reference clip, SSML-based emotion control, and multilingual output for English, German, Mexican Spanish, French, Italian, and Brazilian Portuguese. The service includes the Simba model family, which the site describes as streaming-native speech models designed to model voice identity, expression, and language for realtime conversation. Speechify was co-founded by Cliff Weitzman, who initially built it to consume written material through audio.
Simba 3.2 is Speechify's text-to-speech model. The videos describe it as highly ranked for voice quality and substantially more affordable than competing models. Speechify also offers the model through an API for business customers.
Speechify Work is Speechify’s work-focused AI assistant. It is described as helping users read, dictate, and work with information, with a role comparable to a general-purpose assistant such as Jarvis.
Gradium is a Brazilian text-to-speech company mentioned as a competitor to ElevenLabs and other speech products.
Chinese Patent Skill is an MIT-licensed open-source agent skill for Chinese patent workflows, covering invention, utility-model, and industrial-design patents. It mines patentable points from project materials such as Markdown, code, DOCX, PPTX, and optionally STEP/CAD files; performs novelty, bibliographic, and prior-art searches with China National Intellectual Property Administration sources preferred; drafts patent disclosure documents; rewrites disclosures into claims, specifications, and abstracts; produces Mermaid diagrams or planned patent figures, black-and-white drawings, and optional editable DOCX exports; explains published patents in plain language; and assists with patent-office examination responses and policy briefs. It can read technical and product drawings, extract outlines and component references, derive multiple views from CAD models, and organize utility-model and design workflows with schemas, figure plans, views, and component numbering. Patent explanations can be stored in Obsidian as linked notes, graphs, and canvases, while disclosure work supports self-checking, correction, iterative updates, timestamped drafts, multiple saved versions, and conversation records.
OpenWhispr is an open-source, cross-platform desktop voice-to-text application for macOS, Windows, and Linux. A global hotkey captures speech and inserts the resulting text at the cursor in another application; dictation can use local Whisper or NVIDIA Parakeet speech-to-text engines, where audio remains on the device, or cloud providers through user-supplied keys. It also supports dictation translation, AI-agent commands, meeting transcription with speaker diarization and voice fingerprinting, audio and video transcription, and searchable notes with semantic search and optional cloud sync. The application is built with Electron, React, TypeScript, SQLite, whisper.cpp, and sherpa-onnx, and provides an API and MCP server for programmatic access to notes and transcriptions. The repository describes it as having no data collection or telemetry and distributes installers for the three supported desktop platforms. It is licensed under the MIT license.
Reverify is an AI-agent verification toolkit for reverse engineering and related code analysis. It places deterministic tools between an agent and its claims: the model proposes hypotheses about a binary or candidate implementation, while parsers, disassemblers, pattern scanners, emulators, equivalence checks, and optional analysis engines compare them with ground truth and return evidence-backed VERIFIED, REFUTED, INCONCLUSIVE, OBSERVED, or INVALIDATED results. Its pure-Python core handles PE, ELF, and Mach-O parsing, x86/x64 and ARM/ARM64 disassembly, byte and pattern inspection, CPU micro-emulation, Protobuf/TLV dissection, and Frida hook generation; optional installations add Capstone, Unicorn, LIEF, Z3, or angr for deeper analysis. The project exposes the verification loop through a command-line interface and an MCP server that agents such as Claude Code and Cursor can call. It also records verified, observed, proved, and refuted results in a content-keyed local ledger, allowing grounded facts and negative findings to survive context resets, compaction, and new sessions. A rollover and orchestration system hands sessions off through files rather than model-written summaries, keeping verified facts separate from unverified notes. Reverify includes claim types for bytes, typed reads, instructions, patterns, emulation, behavioral or formal equivalence, imports, exports, sections, and—when angr is installed—functions, calls, references, and reachability. It is distributed from PyPI and licensed under the MIT License. The repository states that it is intended for authorized reverse engineering, malware analysis, CTF work, interoperability research, and software the user owns or is permitted to analyze.
ffmpeg-skill is an agent skill that gives Claude Code, Cursor, Codex, and other agents a local video- and audio-editing workflow built on FFmpeg and ffprobe. It uses a probe → edit, preferring stream copying where possible → check → verify process, with typed Python scripts rather than shell strings. Its 21 tools cover cutting, joining, silence removal, duration and aspect-ratio fitting, captions, overlays, graphics, HDR-to-SDR conversion, LUTs, audio cleanup, loudness, synchronization with drift correction, multicamera editing, delivery checks, project rendering, and batch processing. Each tool supports structured results and dry runs, and the machine-readable contract generates an MCP interface whose tool definitions are derived from the scripts. The skill runs locally without cloud services, API keys, or Python dependencies beyond the standard library; it requires FFmpeg 5.0 or later. It is a tool for local, agent-controlled workflows that probe, edit, verify, and render video.
Commerce Agents is Anthropic's reference implementation for building shopping agents for customers and merchant agents for back-office staff with Claude. It defines each agent through prompts, skills, tool contracts, grounding, memory, approval gates, and backend interfaces, and provides runnable examples for retail, travel, telecom, and entertainment. The shopping agent searches and compares products, plans purchases, fills carts, answers order and policy questions, and maintains customer memory. The merchant agent analyzes performance, manages listings and inventory alerts, handles pricing and promotions, and drafts campaigns; its writes are staged until a host approves them. The agents can run through the Anthropic Messages API, Claude Agent SDK, or Managed Agents, and the repository provides a Claude Code plugin for scaffolding or reviewing commerce agents.
M3E Canvas is a browser-based Material 3 Expressive design tool that turns interface sketches into prompts for AI coding tools. It provides drag-and-drop components, multiple phone and desktop screens, connected tap and swipe navigation, layers and groups, theme controls, previews, and PNG export. Designs can be converted into prompts in English, Japanese, Chinese, or Korean for Android or web targets, including behavior notes for individual components. An optional bring-your-own-key AI helper can draft behavior notes and screen descriptions by sending requests directly from the browser to OpenAI, Claude, Gemini, or DeepSeek. The app is a static Next.js export, saves work in browser localStorage, and is free and MIT-licensed.
Reef is open-source continual-learning infrastructure for self-improving AI agents, developed by Human-Agent-Society. It connects agent inference, interaction records, feedback, learning jobs, evaluation, and versioned delivery, supporting updates to model weights as well as agent harnesses such as prompts, rules, and skills. Each learning cycle has four stages: Reef serves requests and records interactions; matches later scores or structured feedback to those records; produces a candidate update from eligible records; and evaluates the candidate against a configured selection policy before publishing it. Rejected candidates leave the previous release serving, while accepted updates are committed and delivered without restarting the service. The project provides an HTTP API with OpenAI- and Anthropic-compatible inference endpoints, feedback reporting, scenario-based releases, and integrations for training or inference components including Slime and SGLang. It can be installed from PyPI as `reef-infra` or from source; artifact and checkpoint management requires Git LFS.
SlopMonster is a Python linter for detecting formulaic or suspicious AI-generated writing in HTML pages, Markdown, plain text, landing pages, READMEs, emails, and scripts. Its scorer checks AI-associated vocabulary, sentence constructions, punctuation cadence, rule-of-three rhythms, and unsupported sales claims, assigns a score out of 5, and can fail a build when copy scores below 5/5. Its workflow scores the draft, performs a three-pass rewrite that removes suspect vocabulary and sentence shapes and replaces them with specific copy, sends the draft to a different model family for cleansing, and scores the result again. The cleansing script can use a Codex or Claude CLI, refuses to route a draft back to its own model family, and prints a prompt when no supported rival CLI is available. A GitHub Actions workflow is included as a reusable build gate; the scorer itself uses only Python's standard library, while the cleansing step requires an external AI CLI or manual prompt handling. The repository also includes regression tests, a catalogue of detection rules, rewrite principles, and worked before-and-after examples. It is distributed under the MIT license and can also be installed as an agent skill for Claude Code or used by other agents through its plain-Markdown instructions.
UNREEL is a personal AI video streaming service that generates shows while they are watched instead of serving prerecorded episodes. A showrunner LLM writes shots incrementally, and MiniMax H3 Max Turbo renders them through fal; each story shot uses the previous shot’s last frame as its image-to-video input, preserving visual continuity as the stream advances. The player maintains a render buffer and swaps to the next clip when the current one ends, while captions display dialogue written verbatim by the showrunner. The application is built with Next.js, React, and TypeScript, and uses Nano Banana 2 for key art and Gemini 2.5 Flash through fal’s LLM router for the showrunner. It includes chained story films and hard-cut “chaos” channels, and requires a fal API key for generation.
Codenotch is a macOS app that pins a small notch to a screen edge and displays usage limits, reset windows, and activity states for coding assistants including Claude Code, Cursor, Codex, Antigravity, and GLM. It can show whether a provider is working, finished, or waiting for user input, with separate rings for separate Claude Code profiles and live session details on hover. Each provider adapter reads credentials, sessions, logs, databases, language-server RPCs, or provider endpoints already used by the corresponding local tool; Codenotch does not sign in to services itself. A UsageStore polls the adapters, preserves the last good reading across launches, labels each reading by fidelity, and exposes failures as statuses such as stale, needsAuth, or error rather than inventing a percentage. The notch can be placed on any screen edge, expands when approached, and can be configured to show or hide its Dock and menu-bar icons. The app is written for macOS, supports a demo mode with fixed sample data, and uses Sparkle for signed background updates. Its README notes that the provider interfaces may change because some readings come from internal endpoints or local application state. The project is licensed under the MIT License.
ai-evaluation-framework is an open-source framework by Dreamers Inc. for benchmarking model-based solutions by accuracy, p95 latency, and estimated cost against user-defined ground-truth cases. Each task specifies fields and a scoring rule per field, including exact identifiers, normalized dates, numeric tolerances, containment, token overlap, and expected-missing values; cases can also provide alternate accepted renderings. It runs models from OpenAI, Anthropic, or an OpenAI-compatible gateway, compares results in a table, emits detailed JSON reports, and caches responses by model, task, and case so scoring rules can be changed without another model call. The repository uses document-field extraction as its worked example but accepts arbitrary text inputs such as OCR output, transcripts, and PDF text. A separate feedback server collects thumbs-up or thumbs-down responses with required comments, stores append-only JSONL data, and exposes statistics, exports, and candidate cases for human review; it does not train models. The repository includes tests using scripted mock models and is licensed under Apache-2.0.
Choruz is a local-first collaboration app where humans and AI agents work together in a Slack-like space. Each agent runs a real CLI in its own workspace, preserving that CLI's models, tools, and session capabilities while allowing work to be handed to people or other agents through direct messages, groups, mentions, threads, tasks, and files. It provides isolated company and agent workspaces, dedicated directories or Git worktrees, an integrated terminal, file browser and editor, SSH runtime hosts, and browser-based remote control. It supports Claude Code, Codex, Pi, Grok, OpenCode, and webhook-driven external agents, with REST APIs, WebSocket synchronization, webhook agents, Slack and Telegram bridges, and optional plugins. Choruz is in pre-release development and requires Rust, Node.js, pnpm, PostgreSQL, and at least one supported agent CLI when run from source. Its source code and software documentation are licensed under the MIT License; visual assets have separate licensing records.