1,038 tools and products — trending open source, and what gets used in AI and other work.
Higgsfield is an open-source GPU workload manager and machine-learning framework for training models ranging from billions to trillions of parameters, including large language models. It allocates exclusive or non-exclusive access to nodes, queues experiments to manage resource contention, and provides interfaces for distributed training with DeepSpeed ZeRO-3 and PyTorch fully sharded data parallelism. Its framework supports initiating, running, monitoring, and checkpointing experiments on allocated nodes, while its LLM training API includes model, data-loader, and experiment components. Higgsfield integrates with GitHub and GitHub Actions to generate deployment and run workflows: required tools and project deployment keys are installed on servers, code is deployed to configured nodes, and experiments are launched and monitored through GitHub. The repository describes compatibility with Ubuntu nodes offering SSH access and passwordless sudo, and reports testing with Azure, LambdaLabs, and FluidStack.
Codex-X is a cross-platform desktop management tool for OpenAI Codex Desktop and Codex CLI. It provides a visual interface for managing prompt templates, switching between official login profiles and third-party API providers, organizing and repairing local sessions, managing Skills and MCP servers, and inspecting or editing Codex TOML and JSON configuration files. The application can import or edit Markdown prompts, append or replace existing prompt content with automatic backups, test provider endpoints and discover models, import providers from cc-switch, group sessions by project, detect or repair provider inconsistencies, and permanently delete selected sessions from Codex storage. It also previews MCP servers, installs Skills from ZIP files, enables or disables individual extensions, tracks local token usage by date and model, and attributes subagent usage to its main conversation. Codex-X is built as a Tauri 2 desktop application with a React, TypeScript, and Vite frontend, a Rust backend, and SQLite local storage. It provides macOS, Windows, and Linux packages, including Windows portable distribution, and is released under the MIT License.
JimuReport is a web-based data reporting and visualization tool developed by JeecgBoot. It combines three modules: JimuReport for complex reports and printing, JimuBI for dashboards and big screens, and JimuChatBI for conversational data analysis. The system supports Excel-like drag-and-drop report design, grouped and cross-tab reports, master-detail layouts, multi-sheet reports, drill-downs, form filling with database write-back, QR and barcode reports, and exports to Excel, PDF, Word, and images. Its AI features use Claude Code skills to generate editable reports, dashboards, big screens, and conversational BI pages from a natural-language request or screenshot. JimuChatBI converts conversational requests into tables, charts, and grouped summaries, while allowing follow-up refinement. Data can be bound through SQL, APIs, JSON, WebSocket, stored procedures, files, and supported relational, NoSQL, and localized database systems. JimuBI provides draggable visualization components, ECharts-based charts, maps, dashboards, portals, and mobile layouts. The repository provides Spring Boot starters for JimuReport, JimuBI, and JimuChatBI, with integration guidance for Spring Boot 2, 3, and 4. It offers an online service, source-code deployment, Docker deployment, and an installation-free version. The repository describes the software as free for use and documents open-source licensing and supplementary terms, including retention of JimuReport attribution; commercial licensing is available for uses that do not retain the stated notices.
Agent-Native is an open-source TypeScript framework for building agentic applications with purpose-built user interfaces. Developers define each capability once as an action; the same action is used as an agent tool and called by the UI, with shared validation, permissions, implementation, data, and application state. Agents work through this shared action layer rather than clicking through the interface. Actions can also be exposed through HTTP, MCP, A2A, and a CLI. The framework includes agent chat, authentication and permissions, reusable skills and memory, scheduled or event-driven automations, multi-agent teams, and PostgreSQL support with PGlite for local development. It is distributed under the MIT license.
AutoClip is an open-source AI video clipping tool that analyzes transcripts to identify highlights, generate titles, create clips and organize them into collections. It supports local videos, YouTube and Bilibili links, optional SRT subtitles, and transcription through local Whisper components when subtitles are unavailable. The processing pipeline imports video, obtains subtitles or a transcription, performs AI analysis and scoring, generates clips and collections, and exports them using presets for Douyin, Xiaohongshu, YouTube Shorts and Bilibili. Users can select Qwen, OpenAI-compatible APIs, Gemini, SiliconFlow, Ollama or LM Studio models; cloud analysis sends transcript text to the selected provider while video cutting runs locally. AutoClip is distributed as a macOS or Windows desktop application and can also run through a Docker web interface, a Python CLI or an MCP server. It is free software licensed under the MIT License.
AX is Google's open agentic orchestration runtime for declaring and running autonomous agent workloads in Kubernetes clusters. It defines Tasks, Workspaces, Gateways, and Models as ax.io/v1alpha1 manifests: Tasks run agent code in isolated sandboxes, Workspaces preconfigure Git repositories, MCP servers, and skill packages, Gateways restrict outbound traffic with host allowlists, and Models configure the platform's LLM access through Kubernetes secrets. The control plane runs on Agent Substrate and is operated through a kubectl-shaped CLI that applies manifests, watches task status, opens SSH sessions into sandboxes, and suspends or resumes checkpointed agents. AX is described as being designed for high-throughput cluster operation and is still subject to breaking changes before a stable release. It is distributed under the Apache License 2.0.
CLI-Anything is an open-source toolkit and agent plugin for generating command-line harnesses that make existing software accessible to AI agents. It includes integrations for agent platforms such as Claude Code, Cursor, Pi, OpenCode, Codex, and others, plus CLI-Hub for browsing, installing, updating, and launching published harnesses. Its generator follows a seven-phase workflow: it analyzes a target codebase or software, designs command groups and state handling, implements a Click-based CLI with a REPL and JSON output, plans and writes unit and end-to-end tests, documents the result, and packages it for installation. Generated harnesses are intended to call the real application's backend, create or modify its native project files, and invoke the application for operations such as rendering or export rather than replacing it with a simplified implementation. They provide stateful sessions with undo/redo, command-line help for agent discovery, structured JSON responses, and human-readable output. CLI-Anything is distributed under the Apache License 2.0. Its repository also contains agent skill definitions, platform-specific plugins and installers, generated harnesses for numerous applications, and the CLI-Hub package manager installable with pip.
Spirula Studio is a cross-vendor 3D Gaussian Splatting trainer that processes photos or video into Gaussian splats and textured meshes in a self-contained binary. It combines video frame extraction, AI masking, structure-from-motion, Gaussian Splatting training, meshing, depth and normal generation, and skybox processing, with native support for fisheye and 360-degree equirectangular datasets. Its Vulkan backend runs on NVIDIA, AMD, Intel, and Apple GPUs, while a CUDA backend supports CUDA-capable NVIDIA GPUs; the application provides GUI and CLI workflows for Windows, Linux, and macOS. Quantized training is designed to support up to 10 million SH3 Gaussians in 8 GB of VRAM, and masking requires a separately downloaded SAM checkpoint.
PanWatch is a self-hosted AI stock-monitoring assistant for A-shares, Hong Kong stocks, and U.S. stocks. It provides multi-account portfolio management, market and opportunity views, simulated trading, technical indicators, price alerts, stock research, and notifications through Telegram, WeCom, DingTalk, Lark, Bark, or custom webhooks. Its deep-analysis workflow integrates the TradingAgents framework: technical, sentiment, news, and fundamental analyst agents produce analyses, which proceed through bullish/bearish debate, risk review, and portfolio-manager decision synthesis. The system also includes scheduled pre-market, intraday, and post-market agents, plus AI-filtered financial-news collection. Alerts can combine price, percentage-change, turnover, and volume-ratio conditions with AND/OR rules and scheduling limits. PanWatch runs through Docker or local Python and supports a progressive web app for mobile use. It uses FastAPI, SQLAlchemy, APScheduler, the OpenAI SDK, React, TypeScript, and Tailwind CSS; it supports OpenAI-compatible providers including OpenAI, Zhipu, DeepSeek, and Ollama. Optional OpenTelemetry export covers agent runs, LLM calls, and TradingAgents nodes. The project is released under the MIT license.
Umi-OCR is free, open-source, offline OCR software for Windows 7 x64 and later. It recognizes text from screenshots, pasted images, batches of local images, and PDF documents, and can save batch results as TXT, JSONL, Markdown, or CSV files. Its OCR post-processing can merge lines into natural paragraphs, preserve code-block indentation, handle vertical text, and ignore configured rectangular regions such as watermarks or logos. It also includes screenshot, image, and drag-and-drop scanning for QR codes and barcodes, plus QR-code generation. The application can switch between OCR engines through plugins and provides command-line and HTTP API interfaces. The release is distributed as a portable archive or self-extracting archive, requiring no installation or network connection to run. The project is hosted on GitHub and supports multiple interface languages and themes.
stable-diffusion.cpp is a pure C/C++ inference implementation for diffusion image and video models, based on ggml and designed to work similarly to llama.cpp. It provides a command-line executable for generating and editing images from text or image inputs, with support for model families including Stable Diffusion, SDXL, SD3, FLUX, Qwen Image, Wan, LTX, and other listed image and video models. The project supports PyTorch checkpoints, Safetensors, and GGUF weights, and can convert weights to GGUF or Safetensors. It runs on CPU, CUDA, Vulkan, Metal, OpenCL, and SYCL backends across Linux, macOS, Windows, and Android via Termux; features include LoRA, ControlNet, IP-Adapter, latent-consistency models, VAE tiling, ESRGAN upscaling, quantization, and an embedded web UI. The repository is under active development, so its API and command-line options may change frequently.
StarNet is a local-first desktop harness for creating and running AI agents inside a pixel-art space station. The station layout represents the workflow: rooms define capability-scoped teams, hallways authorize handoffs, and placed objects grant capabilities. Each agent run uses its own workspace, transcript, memory, and bounded permissions, while the station reflects runtime state rather than simulating activity. The harness streams model calls through a local Node sidecar and applies explicit capability and consent checks to tool use. It supports OpenRouter and provider sign-ins for Anthropic, OpenAI, and Google, as well as local models through Ollama. Agents can run concurrently, connect to Telegram, Discord, Slack, Signal, and Matrix, use MCP servers and other connectors, launch recipes and reusable skills, run on schedules, and deliver files through an OUTBOX. Spend, budgets, tasks, schedules, transcripts, and memory persist in the local StarNet workspace; provider requests leave the machine when an agent runs, while secrets are held by the sidecar or operating-system keychain rather than the frontend. StarNet is distributed as a desktop application for Windows and macOS, with Windows identified as the most-tested target. It is also runnable from source with Node.js, and its code is available under the MIT License. The StarNet name, logo, artwork, sprites, and other brand identity are not covered by that license.
Claude Code Action is a GitHub Actions integration from Anthropic that uses Claude to answer questions, review pull requests, and implement code changes in GitHub repositories. It detects its execution mode from workflow context, including @claude mentions, issue assignments, and explicitly prompted automation tasks, then interacts with pull requests and issues through GitHub APIs and file operations. It can produce structured JSON outputs for automation workflows and display progress through updating checklists. The action runs on the repository's own GitHub runner, while model requests are sent through Anthropic's API, Amazon Bedrock, Google Vertex AI, or Microsoft Foundry. It is licensed under the MIT License.
OpenShell is an open-source runtime from NVIDIA for running fleets of autonomous AI agents in isolated sandboxes. It lets agents read files, install packages, call APIs, and use credentials while enforcing declared filesystem, process, and network policies. Each agent runs in a sandbox where kernel-level controls govern file access, system calls, and network connections. OpenShell also uses formal verification to review policy changes before approval, flagging newly permitted access such as connections to additional hosts with credentials or calls to new API methods. Credentials are withheld from agents and added only to requests bound for approved endpoints. The project provides a CLI, local gateway, sandbox lifecycle management, policy tooling, inference-provider routing, Kubernetes deployment, extensibility APIs, and Python, TypeScript, Go, and Rust SDKs. OpenShell supports Linux, macOS on Apple Silicon, and experimentally Windows through WSL 2, with Docker, Podman, or host virtualization. It is licensed under the Apache License 2.0 and collects configurable anonymous operational telemetry.
PageIndex is a vectorless, reasoning-based retrieval-augmented generation platform from VectifyAI. It replaces vector indexes and chunking with a hierarchical tree index for each document, then uses an LLM to search that tree agentically, following document structure and context to retrieve relevant sections with traceable references. The PageIndex SDK supports local indexing, retrieval, and chat with a user's own LLM key, as well as PageIndex Cloud for managed parsing, OCR, image understanding, tree-index construction, and storage. It can process text-based PDFs locally and supports cloud-hosted indexes for scanned and image-rich documents; the platform also provides an application for long professional documents, an agent integration interface, and a cloud-only file-level tree index for reasoning across document collections.
CodeGraph is a local code-intelligence tool for AI coding agents, distributed as a CLI, MCP server, and embeddable npm library by Colby McHenry. It parses source code into a SQLite knowledge graph of symbols, files, calls, imports, inheritance, dependencies, framework routes, and cross-language bridges, then resolves references and provides full-text search, source context, call paths, callers, callees, impact analysis, and affected-test discovery. Its primary MCP operation, codegraph_explore, returns relevant verbatim source together with relationships and blast-radius information so agents can query structure without file-by-file exploration. A native Rust parsing kernel supports more than 20 languages, with a portable fallback for unsupported platforms or files that fail native parsing. A native file watcher incrementally updates the local graph after changes, while connect-time reconciliation catches edits made between agent sessions. The tool installs and configures MCP integrations for agents including Claude Code, Cursor, Codex CLI, OpenCode, Gemini CLI, and GitHub Copilot; projects are initialized with codegraph init and stored locally under .codegraph. The repository states that the software is MIT-licensed and that telemetry can be disabled.
TileLang is a Pythonic domain-specific language and compiler for developing high-performance GPU, CPU, and NPU kernels such as GEMM, quantized matrix multiplication, FlashAttention, and linear attention. Built on TVM, it uses tiled programming primitives for memory transfers, pipelining, matrix operations, parallel elementwise work, layouts, and target-specific code generation; its JIT specializes kernels for input shapes and compile-time arguments on first use. The project supports CUDA, ROCm/HIP, Apple Metal, Huawei Ascend, and experimental LLVM, CuTe DSL, and WebGPU backends, with additional ecosystem adapters for other accelerator platforms. It is distributed as a Python package through PyPI, with source builds and nightly wheels available.
UniMate is an AI research system for text-to-motion generation across diverse 3D skeletons, developed by researchers at Princeton, UC Berkeley, MIT, and NTU. Given a rigged asset and a text prompt, it generates articulated motion without per-skeleton retraining or test-time optimization. The system is trained on the UniML3D dataset, which canonicalizes text-paired motion for bipedal, quadrupedal, avian, marine, insectoid, serpentine, and articulated rigid-object skeletons. Its training pipeline uses flow matching: a network predicts the velocity of a linear interpolation between noise and motion data, with masked L2, geodesic rotation, and velocity-smoothness losses. The model supports graph-based spatial and temporal attention or full attention over joint-time tokens, with text conditioning through adaptive layer normalization or cross-attention. At inference, classifier-free guidance generates motion features and can render skeleton animations or export animated GLB and FBX files. The same trained model supports text-to-motion generation, motion in-betweening, text-guided joint-preserving editing, and expansion of sequences from multiple prompts. The repository includes dataset processing, training, inference, and preprocessing code, and is released under the MIT License. The included datasets and source assets remain subject to their original licenses.