1,037 tools and products — trending open source, and what gets used in AI and other work.
PenEcho is an open-source shared canvas for handwriting, equations, diagrams, and spatial context. It reads marks and their spatial relationships on a 20,000 × 20,000 canvas, then places hints, explanations, formulas, plots, diagrams, and other AI-generated drafts beside the original content. Drafts can be moved, resized, copied, accepted, or discarded before becoming part of the canvas ink; users can draw with a stylus or mouse, lasso confirmed ink to move, resize, recolor, or delete it, and send selected content to Typeset without triggering an AI request when editing ordinary ink. The application supports editable sandboxed HTML widgets, professional diagrams, animations, and live-data plugins, which can be refined in place through incremental unified-diff edits rather than full regeneration. It supports up to ten API or CLI connections, including Kimi, MiniMax, Codex, and Claude Code configurations, with a selectable active connection per client. Canvases can be organized into projects, shared across authorized devices, and exported as cropped PNG files. PenEcho provides a desktop app and an npm installation requiring Node.js 20.3 or newer. It can run offline with a user's own API or authenticated CLI setup; API credentials remain on the host device and are not sent to browser code. PenEcho Cloud is an optional companion service for private versioned projects, synchronized favorites, linked-device remote access through a relay, and public sharing of Canvases and Widgets as Echoes or Crafts. The project is supported through Moonshot AI's Kimi Open Source Friends program.
Rakazo is an open-source platform for running persistent AI teammates. Each bot has conversations, memory, routines, and history, and can access a browser, terminal, files, and a graphical desktop. Users can access bots through the web app, Electron desktop app, or Expo mobile app, and can take control of the graphical desktop when needed. The platform supports shared team computers and isolated private computers, bot delegation to peer bots or short-lived subagents, voice interaction, and user-provided model credentials through Pi. It can use Docker, E2B, Daytona, Box, or trusted local computers as computer providers, and supports integrations through Composio, Pipedream Connect, MCP, OpenAPI, and Treg sources. Rakazo can be self-hosted with Docker and is also available as a managed application. The repository describes it as being in beta and identifies a TypeScript stack using React, Vite, Tailwind CSS, Electron, Expo, Hono, oRPC, PostgreSQL, Prisma, Better Auth, and Graphile Worker.
pgbot is a read-only PostgreSQL observability and health-diagnosis tool for AI agents, applications, and operators. A static binary connects to PostgreSQL, reads the database's own statistics views, and produces deterministic, findings-first health reports, including a health score, critical/warning/note findings, verified healthy subsystems, and change information from local baselines. It provides focused commands for indexes, queries, tables, and vacuum, along with versioned JSON output for agents and scripts and an MCP mode. Its optional ask and explain commands add an AI interpretation of the same findings; inspection and database queries otherwise require no model key or external service. The project is in beta and uses a read-only pg_monitor role with session-level read-only, statement, and lock timeouts as additional safeguards.
Colibrì is a pure-C inference engine and open research platform for running large mixture-of-experts models on consumer and heterogeneous hardware without engine dependencies. It treats VRAM, system RAM, and storage as a unified multitier inference hierarchy, streaming expert weights from disk and coordinating placement, storage I/O, scheduling, kernels, speculation, and CPU/GPU overlap. The project provides chat, server, and web front ends, with support for CPU, CUDA, Metal, and dual-SSD streaming; its design rules prohibit silently changing model precision or router semantics when fast memory is insufficient.
img2threejs is an agent skill that reconstructs an object or character from a reference image as a code-only, procedural Three.js model. It generates a TypeScript THREE.Group factory using primitives, procedural shaders, and generated geometry rather than photogrammetry, mesh extraction, or downloaded art assets. The generated scene includes runtime structures such as pivots, sockets, and colliders for animation, and the project documents separate hard-surface and anatomy-aware reconstruction paths. It runs under Claude Code, Codex, or OpenCode and provides live browser demos whose models can be inspected as generated source.
TurboFieldfare is a model-specific Swift and Metal runtime for running the instruction-tuned Gemma 4 26B-A4B model on Apple Silicon Macs, including systems with 8 GB of RAM. It keeps the model’s shared 1.35 GB core and FP16 key-value cache in memory while streaming the expert weights selected for each token from an SSD, allowing inference with about 2 GB of resident weights instead of loading the approximately 14.3 GB model. The model uses MLX affine 4-bit weights with group size 64, an 8-bit router, 4-bit shared and routed experts, and about 3.88 billion active parameters per token. The project includes a native macOS app, command-line interface, Swift runtime library with Metal kernels, decode service, streaming model installer and verifier, and experimental loopback OpenAI-compatible Chat Completions server. These products share a repacked `.gturbo` model directory; the installer streams byte ranges from a pinned Hugging Face revision, repacks them without materializing the source checkpoint, and validates the resulting manifest and file hashes. The app and CLI support instruction and completion generation with optional system guidance but do not execute tools; the server accepts function-tool declarations and returns model-produced tool calls for client authorization and execution. An optional companion vision tower adds image input on M2 or newer Macs, while text-only inference remains available on M1. The package is arm64-only and targets macOS 26, Metal 4, and Swift 6.2 or newer, with approximately 14.3 GB of storage for the text model plus about 1.1 GB for the image pack.
Personal Model is an open-source, local-first AI memory runtime from Intuition Lab for giving coding agents evidence-linked context about a user. After macOS permissions are granted, it captures focused activity from applications on the user's Mac and builds a portable HUMAN.md-style personal context model, storing the data locally and allowing it to be inspected, corrected, exported, or deleted. The runtime progressively organizes sourced observations and events into relationships and changes, supported patterns, higher-order structures, and an integrated current model. It retains source receipts for important claims so later evidence can strengthen, revise, or overturn inferences, then exposes this context through MCP to trusted clients including Claude Code, Codex, Cursor Agent, Claude Desktop, and other compatible tools.
Finger Frame AI is a browser app and Python CLI pipeline by Sophia Yang that places an AI-generated world inside a two-hand finger-frame gesture. It sends an uploaded video to Gemini Omni Flash for video-to-video restyling using a selected style or custom prompt, with alignment instructions intended to preserve framing, facial positions, expressions, blinks, gaze, and motion. MediaPipe Hand Landmarker then tracks both hands and the finger-frame quadrilateral frame by frame; the tracking pipeline uses anatomical corner ordering, spread and area gates with hysteresis, teleport rejection, velocity-adaptive smoothing, dropout holding, and presence fading. The compositor reveals the restyled video through the tracked quadrilateral and adds a dashed outline with pulsing corner markers. The browser app exports MP4 where browser recording supports it and otherwise WebM. The CLI separates restyling and compositing into `stylize.py` and `composite.py`, accepts formats readable by FFmpeg or OpenCV, and can produce H.264 MP4 while retaining the original audio track when present. Users provide their own Gemini API key; the app keeps it in the browser and sends it only to Google's API. Generation is billed per video, and the app recommends keeping uploads under about 15 MB. A no-key placeholder mode runs the tracking, compositing, and export pipeline with a hue-shifted stand-in.
AGENTS.md is an open format for guiding AI coding agents through project-specific context and instructions. It provides a dedicated, predictable file—similar to a README for agents—where projects can document development-environment commands, testing and linting procedures, package or project conventions, and pull-request requirements. The format is documented by the agentsmd project, which also maintains a basic Next.js website with examples.
AlphaFold is an open-source implementation of the AlphaFold v2 protein-structure prediction inference pipeline, developed by Google DeepMind. The repository also includes an implementation of AlphaFold-Multimer for predicting protein complexes and documents the updated AlphaFold v2.3.0 models and inference procedure. It runs on Linux using Docker and a modern NVIDIA GPU, with model parameters and genetic databases supplied separately; the full database setup requires substantial storage.
An Apache-2.0-licensed open-source Go SDK from Grafana for building AI-powered backends and agents. It provides GenerateText and StreamText APIs for calling language models, returning complete or streamed responses, executing Go functions as tools, generating schema-validated structured output, and running multi-step tool workflows, with approval for consequential tool actions. The SDK supports Anthropic, Amazon Bedrock, OpenAI, OpenAI-compatible APIs, and Grafana's hosted endpoint, along with retries, configurable timeouts, fallbacks, logging, Prometheus metrics, and Agent Observability. It follows the design of Vercel's AI SDK and is wire-compatible with its TypeScript React frontend hooks, including useChat, useCompletion, and useObject, allowing Go backends to stream Server-Sent Events directly to an existing React frontend without a protocol adapter.
Zeron is a local control layer for coding agents, including Claude Code, Codex, Cursor, Grok, Hermes, and Pi. Each device runs a small engine that stores agent sessions locally by default, without requiring an account or network connection. Optional synchronization lets users start an agent on one device and follow or control it from another; an always-on machine such as a VPS can keep agents running after a laptop is closed. The Linux installer starts a daemon that persists across reboots, while macOS users can install the desktop release or build from source. Zeron is licensed under the MIT License. The videos refer to the project as Comet.
icm-architect is a Claude skill that designs processes, ideas, and problems as ICM (Interpretable Context Methodology) workspaces, using folder structure as agent architecture, or restructures an existing folder, repository, or vault. Numbered folders encode sequencing, hierarchy scopes context, and plain Markdown files store state, allowing an agent to orient and act by reading the relevant files. In build mode, it extracts stages, human approval points, and stable versus per-run elements, then scaffolds a workspace using one of six forms: Pipeline, Umbrella, Record library, Knowledge bundle, Context map, or System map. In restructure mode, it audits files as catalog, contract, factory, product, or dead, proposes a migration map for approval, then migrates and validates the result. The walk test checks whether an agent with no prior memory can orient, act, and report status from the workspace alone. It can be installed as a Claude Code skill or uploaded to Claude applications, and is MIT licensed.
barehands is an open-source, webcam-powered hand-tracked interface for AI assistants. It runs in Chrome with a webcam and overlays notes, images, and 3D models as cards over the camera feed; users manipulate them with gestures such as pinching, throwing, stretching, force-pulling, and two-finger dragging. The interface uses a local Python server and loads Google MediaPipe for hand tracking and three.js for 3D rendering. Markdown folders, including Obsidian vaults, can serve as note sources, while configured media folders provide images, transparent props, and 3D models. An AI assistant can control the interface through local files and shell scripts: state files drive the on-screen ring's status, and the board protocol presents content on the stage. The repository states that it can be used, modified, and built on commercially within a business, with redistributed versions remaining under the same license.
unlazy is an agent skill for enforcing completion discipline in substantial AI-agent tasks. It requires an acceptance ledger to be written first, with runnable CHECK commands and EXPECT conditions; its gate checker reviews commands, executes approved gates, records evidence, and can reverify completed work. The skill also supports a Depth Tree method that splits tasks into fresh-context subtasks with integration checks, giving each leaf the full task time budget. It can be installed for supported agents with the skills CLI or manually for Claude Code and Codex CLI; its checker and optional hook require Node.js 16 or newer and no third-party runtime packages.
Endoplexity is a Chrome MV3 side panel and local bridge for agentic control of the browser a person is already using. It lets Claude or Cursor agent CLIs operate existing tabs and logged-in sessions by clicking, typing, filling forms, navigating, switching tabs, uploading files, and reading pages, without sending requests directly to a model or requiring a separate metered API key. The side panel owns the Chrome DevTools Protocol connection and communicates over an origin-pinned WebSocket with a bridge on localhost; the bridge exposes browser tools over token-gated MCP/HTTP and enforces the safety policy, including approval gates and autonomy modes. Pages are provided to the agent as accessibility-tree snapshots rather than raw HTML, and actions return the resulting page state so subsequent turns can resend only changed lines. Its file-reading tool extracts text from PDFs, DOCX, XLSX, PPTX, CSV, JSON, Markdown, and other text files.
NorthCinder is an open-source, local-first MCP server for AI shopping agents. It compares products from user-selected sources against a buyer's criteria, explains rankings and exclusions, shows source facts and unverified details, and normally returns up to three useful options rather than claiming one universal best product. Its research workflow provides separate product and seller guides, creates a research plan, and uses controlled research tools; results remain provisional when sources conflict or do not identify the exact item or seller. Purchasing is a separate human-in-the-loop decision: each checkout requires fresh approval for one exact offer and unit, with the merchant, variant, price, known total, and spending cap signed into a single-use approval. It rejects raw card details, using an opaque payment token for supported automated checkout or handing the buyer a cart link. NorthCinder runs alongside an MCP-capable AI application. Its local setup requires Node.js 20 or later; the init command stores configuration locally and starts the MCP server and search engine in one process. The repository states that there is no hosted NorthCinder account or cloud service, and that order outcomes remain local.
TrueForge is an open-source, vendor-neutral agent harness developed by TrueFoundry. It runs the agent execution loop, including model calls, MCP tool use, skills, sandboxed code and file execution, approvals, context management, and session state. Agents can use configured model providers, remote MCP servers, git-backed SKILL.md instruction packs, and on-demand sandboxes; context features include subagents, deferred tool loading, Code Mode, large-result offloading, and compaction. TrueForge exposes agents through a bundled chat UI, an HTTP API with a TypeScript SDK, and an embeddable UI SDK. It supports local mode with SQLite and hosted deployments using Postgres and Redis through Docker Compose or Helm; the repository warns that local mode is intended for personal use on localhost rather than production or internet-facing deployment.
sloptrim is a local prose linter and coding-agent plugin that detects patterns associated with AI-generated writing when an agent saves a document. It extracts prose from supported text and office formats, scores it on a 0–100 scale, and reports suspicious spans for revision; it analyzes writing patterns rather than determining authorship and does not process code. The command-line tool uses Python's standard library, requires no model or network connection, and keeps text on the local machine. It integrates with Claude Code and can be initialized for other agents through an agent contract and editor-rule files.
APEX is an open, verification-first RTL design for an LLM-inference tile developed with Sigmantic AI. It implements one transformer decoder layer—including attention, softmax, RMSNorm, RoPE, SwiGLU, residual processing, and KV-cache compression—and runs Qwen2.5-0.5B through the verified pipeline on FPGA hardware. The KV-cache codec sits in the datapath: cached keys and values are compressed as they are produced and decompressed when consumed, while an importance unit allocates precision according to which parts of the context matter. Each RTL block is checked against an executable NumPy golden model with bit-exact comparisons and mutation-tested testbenches. The repository covers the inference tile rather than a complete system; DRAM control, PCIe, and NoC are out of scope.
An open-source Python tool that extracts local chat histories from AI coding assistants and converts them into a normalized JSONL format for backup, personal analytics, or fine-tuning. It auto-discovers conversation data in the storage formats used by Claude Code, Codex CLI, Cursor, Windsurf, Trae, Continue, Gemini CLI, OpenCode, Cline/Roo Code, and Aider. Extracted records can include user and assistant messages, code context, diffs or suggested edits, tool calls and results, timestamps, session IDs, project paths, and model names when those sources store them. It handles JSONL, JSON, SQLite, Markdown, and heuristic undocumented schemas, searches macOS, Linux, and Windows storage conventions, and runs from the command line using only Python's standard library, requiring Python 3.9 or newer.
The X For You Feed Algorithm is the open-source codebase powering the For You feed on X. It assembles each request from in-network posts kept in memory by Thunder from accounts a viewer follows and out-of-network posts discovered through Phoenix retrieval and SimClusters, then hydrates candidates with post, media, author, account-label, language, and engagement data. A separate blending pipeline can add ads, Who to Follow recommendations, and prompts. The post pipeline applies visibility filtering based on viewer actions such as blocks, mutes, and muted keywords, as well as labels produced by rule-based systems, account-scoring models, media models, and enforcement components. A transformer-based Phoenix model predicts the viewer’s probability of taking actions on each post and other continuous values; configured weights combine those predictions into a score, which ranks posts after filtering. The repository includes retrieval, labeling, filtering, scoring, ranking, Phoenix training and inference code, synthetic data generation for a proof-of-concept training run, and an Under the Hood tool that displays aggregate statistics about labels affecting account and post visibility.
oMLX is an open-source LLM inference server for Apple Silicon Macs, managed through a macOS menu-bar application. It serves text, vision, OCR, embedding, and reranking models through compatible API endpoints, with model discovery and management, monitoring, continuous batching, and caching. Its caching system keeps KV cache in a hot in-memory tier and a cold SSD tier, allowing cached context to remain reusable across requests even when the conversation context changes. The project can be installed as a macOS app, through Homebrew, or from source, and includes a CLI for controlling the app-managed server. The repository states that it requires macOS 15 or later, Python 3.11–3.13, and Apple Silicon hardware.
llmfit is a terminal tool that evaluates which local language models fit and run well on a user's hardware. It detects system specifications including RAM, CPU, GPU, and multi-GPU configurations, then scores models across fit, estimated speed, quality, and context dimensions while accounting for quantization and mixture-of-experts architectures. It provides an interactive terminal UI and a classic CLI, supports local runtime providers including Ollama, llama.cpp, MLX, Docker Model Runner, and LM Studio, and can output recommendations as JSON. Its benchmarking workflow can measure real token-per-second performance on a machine and contribute those results to the project's fit data.
SWE-bench is an execution-based benchmark for evaluating large language models on real-world software issues collected from GitHub. Given a repository codebase and an issue, a model must generate a patch intended to resolve the problem; the patch is evaluated by running repository tests in Docker-based environments. The project provides datasets and an evaluation command-line tool, with dataset variants including full, verified, multimodal, multilingual, and locally supplied datasets. It is distributed as open-source code through its GitHub repository and can also be loaded from Hugging Face.
whisper.cpp is a C/C++ implementation of OpenAI's Whisper automatic speech-recognition model for speech-to-text inference. It is designed for offline, on-device use and supports CPU-only inference, GPU and NPU acceleration, integer quantization, mixed F16/F32 precision, voice activity detection, and a C-style API. The implementation targets macOS, iOS, Android, Linux, Windows, WebAssembly, Raspberry Pi, Docker, and other platforms, with optimizations including ARM NEON, Apple's Accelerate and Metal frameworks, x86 AVX, POWER VSX, Vulkan, CUDA-related NVIDIA support, ROCm, OpenVINO, and other backends. The repository includes the whisper-cli example, which transcribes 16-bit WAV audio using Whisper models converted to ggml format; its zero-dependency implementation is intended for integration into applications such as on-device captioning tools.
Kimi K3 in C is a portable C99 inference engine for running the 2.78-trillion-parameter Kimi K3 base model on CPUs without BLAS, an inference framework, or a GPU. It reads the 1.56 TB checkpoint from disk, keeps a selected depth of the 93-layer dense trunk in memory, and streams the remainder during generation; its routed experts are packed in 4-bit form, multiplied directly from that representation, and managed with an LRU cache. The implementation includes KDA attention with a non-growing memory, MLA attention with a latent representation, and selection of 16 experts from 896. The repository provides a command-line interface, weightless and full-checkpoint validation gates, and measurement data; it reports operation from 8 GB of RAM upward, with additional memory changing speed rather than the generated output. The base model has no chat template, so prompts produce continuations rather than chat replies. Reference requirements include an AVX2-and-FMA CPU, about 1.7 TB of free storage, and a C compiler; Linux/x86-64 is the reference platform, while macOS/arm64 and Windows/x86-64 are also documented as building and passing the tests.
watsonx Orchestrate is an IBM enterprise AI orchestration service in the watsonx family that composes AI assistants and no-code workflow automations to perform knowledge-worker tasks. It accepts conversational input and business data through connectors and executes automated actions across enterprise applications as a commercial SaaS offering on IBM Cloud.
J-Space Cognition Suite is a model-agnostic, inference-time control system packaged as a cross-platform AI-agent Skill for deep reasoning, long-horizon work, tool use, verification, and recovery. It leaves model weights and training unchanged, organizing an agent's working representations through one entry point and nine selectively loaded modules supported by four references. Its fast, full, and loop operating modes use selective workspace loading, a shared broadcast hub for constraints and values, dense internal reasoning traces, explicit intermediate steps before conclusions, metacognitive routing of confidence and failure signals, and bounded empirical verification. An optional standard-library controller records durable task state, including goals, next actions, checkpoints, open questions, seams, and recovery state. The suite is intended for low-friction integration with AI hosts that support Skills or equivalent system/developer-instruction mechanisms.
Docker Sandboxes is a Docker security feature for running AI coding agents and tools inside isolated micro virtual machines. Each sandbox has its own kernel, filesystem, and network, allowing agents such as Claude Code to run with reduced access to the host machine.
Legora is an AI legal technology company that provides a platform for legal professionals to research, analyze, draft, and work with legal documents. It is positioned as a specialized alternative to general-purpose AI tools and competes with products such as Harvey.
Iris is a Rust command-line tool and local stdio MCP server for capturing screenshots of live websites with an installed Chrome-family browser such as Chrome, Chromium, Edge, or Brave. It can capture viewport or full-page images, emulate mobile sizes and user agents, apply dark color schemes, frame the first matching element with optional padding, wait for selectors or page content, and process URL batches concurrently or from standard input. The same binary exposes an MCP `capture` tool that returns an inline image with structured metadata and can optionally write an output file. It is distributed as the `iris-screenshot` crates.io package or through an install script; building from source requires Rust 1.88 or newer.
Wake is a native desktop application that gathers coding-agent sessions from local data directories into a single searchable library. Built with Rust and GPUI, it reads supported agent histories read-only, groups them by agent and project, watches files for incremental updates, and indexes transcripts with SQLite FTS5 trigram search, including code substrings and CJK text. Its transcript view renders messages, collapsible tool-call clusters, thinking summaries, and syntax-highlighted code. Wake can resume supported conversations in a terminal at the original project directory, and also provides session management, Markdown export, deletion with tombstones, and library statistics. Data remains local; the application does not make background network requests and only contacts GitHub when an update is explicitly checked. It is macOS-first, with experimental Linux support and experimental Windows support.
Amagine3D is an open-source 3D capability layer for hardware creation, developed by Amagine. Its parametric CAD workflow turns natural-language hardware requirements, reference images, and key dimensions into editable enclosures and assembly structures around internal components, producing Python and build123d source code alongside STEP, STL, or color-aware 3MF exports. A 3D-native agent first organizes the requirements into a design brief, then runs the generated source in a browser geometry runtime. It uses the resulting model state and check results for dimensions, part connectivity, interference, and motion clearance to revise or accept the design. The workflow supports multipart structures such as covers, hinges, and latches, including assembly clearances and printing tolerances. Generated parameters can be adjusted in the workbench and written back to the source without regenerating the model. The system can preview, measure, modify, and save complete designs and their check reports.
HFlow is an open-source SDK from Hebbian Robotics for multimodal data pipelines in robotics and physical AI. It lets teams implement transformations, quality checks, labels, and enrichments as plain Python functions while HFlow handles orchestration, provenance, storage, versioning, and dataset curation. The pipeline accepts supported MCAP episodes and can import LeRobot Dataset v3 repositories. It runs code in-process for development or generates Airflow 3 DAGs for scheduled execution, producing canonical MCAP episodes, artifacts, provenance records, and a Parquet catalog. DuckDB SQL creates version-pinned manifests for curation, while a queryable catalog records metadata and quality evidence without requiring the underlying recordings to be loaded.
macOS Harness is an experimental, MIT-licensed, macOS-only Python harness from browser-use that gives an LLM or coding agent raw primitives for controlling a Mac in one persistent Python process, without app-specific tools, recipes, or framework rails. Its six core primitives—`see`, `key`, `type`, `click`, `ax`, and `script`—capture native and Electron app windows, send keyboard and coordinate input directly to an application process, expose Apple Accessibility and Apple Events, execute AppleScript, and draw a click-through pointer without moving the physical cursor. The agent can write missing task logic in ordinary Python during execution. The same process provides access to a real, logged-in Chrome browser through Browser Harness, as well as the local filesystem and shell commands through Python, `Path`, and `subprocess`. It can capture background application windows without bringing them to the foreground and does not activate or raise the target app. The project includes installation and skill-registration workflows, a `doctor` command for checking required macOS permissions, and verification instructions. Anonymous telemetry is enabled by default and records the CLI command category, success, duration, package version, operating system and architecture, and detected agent client; the project states that it does not record prompts, app names, screenshots, UI text, scripts, paths, or window titles, and provides `macos-harness telemetry disable` to turn telemetry off.
OpenSpec is an open-source, tool-agnostic spec-driven development framework from Fission AI for AI coding assistants. It stores requirements and development artifacts as plain Markdown files in a repository, using an `openspec/` structure containing specifications and proposed changes. Its workflow supports exploratory discussions, `/opsx:propose` for generating a proposal, requirements and scenarios, a technical design, and implementation tasks, followed by `/opsx:apply` to implement the tasks and `/opsx:archive` to update the specifications. Coding agents can interact with it through slash commands or an MCP server. The project is designed for iterative, brownfield development and supports shared or cross-repository requirements through beta Stores. The CLI requires Node.js 20.19.0 or higher and is installed with npm.
Modelplane is an open-source control-plane product for declaratively deploying and governing self-hosted machine-learning model inference across an organization's GPU fleet, with provider support and networking capabilities for reliable operation across clusters and more complex algorithms.
CarWatch is an open-source, offline in-car AI agent from ThinkOffApp that runs on a Raspberry Pi 5 with a locally hosted language model. It joins chat rooms as a vehicle agent, answers questions from the car's owner's manual using lexical retrieval-augmented generation and page citations, and reads live vehicle and system data while stating when information cannot be sensed. A continuous energy-based voice listener and whisper.cpp provide local speech input without a wake word, with responses played through the car's speakers. The system can monitor OBD and engine data, send departure and arrival messages, trip summaries, and dashcam clips, and expose a phone dashboard for approvals, replies, model selection, and maintenance. The repository describes systemd-managed services for the model server, room agent, voice listener, dashboard, and engine watcher, with Python standard-library code and no cloud subscription required.
Factory is a reference software factory for Claude Code and Codex that installs a repeatable, version-controlled software-delivery workflow into an existing GitHub project. GitHub Issues serve as the work queue, while committed policies, skills, labels, handoff comments, pull requests, and run records preserve state between fresh agent sessions. Scheduled agents triage issues, route them to implementation, specification, questions, or blockers, claim bounded work, implement it on a branch, run configured type, lint, test, build, audit, and architecture checks, obtain independent verification, and open draft pull requests. An independent verifier reads the diff and checks that the new test fails without the implementation. Humans retain responsibility for ambiguous requirements, system design, significant changes, and merging pull requests. The repository has no custom orchestrator or queue service. Claude Code routines provide the default scheduling and compute, with a thin Codex adapter using the same policies, gates, and evidence files; GitHub events or an optional API-triggered Action can also start runs. It is configured through files such as a human-owned charter and gates script, and includes a /factory control room backed by live issues, pull requests, and run records.
FrontierAgent is an open-source agent runtime, terminal product, and evaluation suite for long-horizon research and file-based work. Its native terminal TUI provides a ReAct workflow in which one stateful agent researches, reads files, writes deliverables, runs commands, and iterates in a task-scoped sandbox, and an Agent Team workflow in which a coordinator maintains a task board, delegates bounded independent assignments to parallel sub-agents, collects structured reports, and synthesizes the result. Shell and file tools use a shared sandbox with read-only /inputs, working-state /workspace, and persistent /outputs directories. The TUI supports queued instructions during execution, approval-required diffs for mutating operations, local action traces, checkpointed sessions, resumption, and reversion. The same workflow engine powers a subprocess benchmark runner with deterministic artifact collection, concurrency, progress inspection, and reruns of individual failures. The repository also documents connection to the OpenAI-compatible Apodex-1.1 endpoint through the Apodex API.
hayamimi is a local, real-time multilingual speech-to-text tool that produces live subtitles, a browser dashboard, speaker labels, and optional translation without a GPU or cloud API. It detects the language of each utterance and routes it to a dedicated specialist model for Japanese, Chinese, Korean, Cantonese, English, and European languages, with a broad Omnilingual ASR fallback for other languages. The models run as quantized INT8 ONNX models through sherpa-onnx, without PyTorch or CUDA. The pipeline updates partial subtitles while speech is in progress, emits finalized lines shortly after speech stops, and can perform a second batch-decoding pass after silence to produce a refined transcript. Speaker labeling uses CAM++ embeddings and is intended for turn-taking rather than full diarization. The `--serve` option starts a local HTTP server with a live dashboard and an OBS browser-source overlay. Audio can also be supplied over WebSocket, and model eviction keeps resident models below a configurable memory cap, with the default under 2 GB of RAM. The project reports 5.8% character error rate on its Japanese broadcast-audio scorecard and 10–50× real-time processing on a six-core desktop CPU.
XLA (Accelerated Linear Algebra) is an open-source machine-learning compiler for GPUs, CPUs, and other ML accelerators. It takes models from frameworks such as JAX, PyTorch, and TensorFlow, traces and compiles them into optimized executable code for the target hardware. The project is developed under OpenXLA and is primarily used through the documentation and integration of the corresponding ML framework; its repository is intended mainly for compiler contributors and frontend or hardware-backend integrators.
FreeLLMAPI is an open-source proxy that aggregates free tiers from multiple LLM providers behind a single OpenAI-compatible `/v1` endpoint. It can also connect custom OpenAI-compatible chat, embedding, image, and audio endpoints, including local servers and remote gateways. Its router selects an available model, fails over when a provider is rate-limited, stores provider keys encrypted, and tracks per-key usage against free-tier limits. The router refreshes its model catalog from a signed feed; free installations receive periodic snapshots, while the hosted premium offering provides the live catalog. The project is intended for personal experimentation and is available as a self-hosted install with a hosted service at freellmapi.co.
ai-memory is an open-source local server for long-term memory and session handoff among AI coding agents. It captures prompts, tool calls, and agent context through MCP integrations and lifecycle hooks, storing the material in a searchable, Git-backed Markdown wiki. When a session ends or is explicitly finalized, it generates handoff context for a subsequent session, including the project architecture, failed approaches, and open questions, so work can continue across agent vendors without restating the context; clients without a true session-end hook use explicit finalization commands. The project documents integrations for Claude Code, Codex, Command Code, Devin CLI, OpenCode, Cursor, Gemini CLI, Oh My Pi, Pi, and Crush, with agent-specific configuration, generated plugins or extensions, and capture exclusions. It runs on Linux, macOS, and Windows through WSL2, with experimental native Windows support, and is distributed through Docker images and native release binaries.
Mercor is a data and talent service discussed as providing training information for frontier AI models. The video emphasizes the company’s rapid growth and high valuation.
BetterVoice is an experimental, open-source macOS menu-bar app for local voice dictation with screen context. It transcribes speech locally and captures the full display whenever the user circles an area with the pointer, marking each capture with a blue pointer trail and pulse while preserving multiple references in order. The app inserts the transcript into the selected text field when macOS permits it and pastes captured images when both speech and screen context are available. It supports quick notes and long explanations, configurable microphones, dictation languages, circle-detection thresholds, and keyboard shortcuts. English uses a local Parakeet model by default, while other languages use a separately downloaded multilingual model; an optional English-only grammar-cleanup feature runs locally through quantized ONNX weights.
Microduck is the software brain for Pollen Robotics' approximately 25 cm biped duck robot. It runs reinforcement-learning policies on a Rockchip RK3566, with a 50 Hz control loop for fifteen servos, radios, the camera, and behaviors such as walking, rolling, grasping, kicking, quacking, and self-righting; it also accepts gamepad control. The Rust implementation is organized as daemons communicating through a shared JSON-RPC contract over Unix sockets. `robotd` owns the control loop and motor bus, while other services handle updates, configuration and Wi-Fi, Bluetooth, gamepad input, camera streaming over WebRTC, and the depth sensor. Signed updates are health-gated, reversible, and roll-backable. The policies are trained in the companion `microduck_rl` project with MuJoCo and PPO, then exported to ONNX for loading by this repository.
Experiential is an open-source gateway and router for agent workflows from Experiential Labs. Its `exp` CLI provides hosted, bring-your-own-key, local, and custom models through OpenAI-compatible and Anthropic Messages APIs, with model aliases, access controls, use-case restrictions, and spending limits for users and agents. It can persist provider connections and configuration locally, expose API routes on loopback, and load a fitted project router as an OpenAI client through its Python interface. It collects OpenTelemetry traces from agent traffic to build simulations and optimize routing for quality, speed, and cost, and supports harness optimization, endpoint serving, model distillation, and fine-tuning an owned open-source model through Tinker. The hosted gateway is available at `api.experientiallabs.ai`; anonymous aggregate PostHog telemetry is enabled by default locally and can be disabled.
Doop is an open-source multiplayer design canvas where people and AI agents create and edit designs together. Each canvas contains frames that render real HTML in sandboxed iframes; humans edit in the browser, while agents connect through the built-in Model Context Protocol (MCP) server and can create frames, stream HTML in chunks, inspect screenshots, and revise designs. The app synchronizes cursors, presence, frame edits, agent status, comments, tasks, and activity over WebSocket rooms, and includes a server-side Doop Agent that can process queued cards, mentions, and feedback through specialist roles. Canvases are private by default, with email invitations or optional link sharing, and agents inherit the access of the user who authenticates them. Doop also provides design-memory features for exemplar frames, decisions, and proposed style rules, and can be self-hosted with Docker Compose or Bun using embedded Postgres through PGlite.
ego lite is a macOS browser from Citro Labs designed for AI-agent browser automation alongside a user's normal browsing. It gives each agent or task an isolated Space within the same browser, allowing multiple tasks to run in parallel without taking over the user's tabs; users can observe, take over, or stop an agent's Space. Through the ego-browser skill, agents such as Claude Code, Codex, Cursor, or custom agents can call in-page JavaScript tools including snapshot, fill, click, wait, navigate, and capture; an agent can compose several operations into one JavaScript execution rather than repeatedly issuing CLI commands. During setup, optional Chrome-data migration can transfer existing logins, cookies, extensions, and bookmarks, while browsing data remains on the user's device. The browser is distributed as a separate free download, the repository is MIT-licensed, and Windows and Linux support are listed as planned.
SwarmForge is a local, tmux-based orchestration platform for coordinating multiple AI agents across Git worktrees. It reads a project-local configuration that assigns roles, agent backends, worktrees, task or batch handling, and handoff propagation; launches each role in an isolated tmux session; and provides a browser-based pack cockpit for projects, tasks, approvals, clarifications, live status, agent-pane inspection, and teardown. Agents communicate through a daemon-delivered handoff protocol and helper scripts rather than direct tmux messages. The repository supplies two-pack, four-pack, and six-pack workflow templates with role prompts and constitution articles, while projects can define their own swarm topology. It runs locally with zsh, Git, tmux, Babashka, and at least one supported agent backend such as Claude, Codex, Copilot, or Grok; swarm state is kept in the working director
Rome is an agentic operating environment from Rome OS for collaboration between humans and AI agents. It provides manifests, typed actions, agents, skills, hooks, tools, workflows, memory, interfaces, policies, and persistent database-backed state in a guardrailed environment where agents can build applications, define standard operating procedures, and orchestrate workflows under human guidance. Rome supports one-off tasks, scheduled work, and long-running follow-through. Rome Apps combine a purpose-built web interface, agent reasoning, agent-owned collaborators, reusable workflows and skills, lifecycle hooks, typed actions, HTTP APIs, and persistent app-private databases or files into installable products. Apps are defined with an app.yaml manifest and use the @rome-os/app-runtime and @rome-os/app-web-sdk packages; Rome can generate and iterate on app or workflow source code from a plain-language request. The project can run locally through Docker or be accessed through Rome Cloud, which the repository describes as a private preview environment.
Proliferate is an open-source AI IDE for running coding agents such as Claude Code, Codex, OpenCode, Cursor, and Grok through their native harnesses. It runs agents in parallel within one workspace, giving each task an isolated Git branch and worktree along with its own terminal, conversation, and review state. Agents can delegate scoped work to subagents, while MCPs, skills, Computer Use, Browser Use, and custom tools can be configured once and shared across agents. Its workflow system supports recurring and event-driven runs such as review passes, alert triage, and dependency updates. The desktop application can use a local runtime or connect to a control plane that runs locally or in the cloud. The control plane is self-hostable through Docker, AWS, GCP, Azure, Kubernetes, or air-gapped deployments. The repository lists Rust, Node.js, and pnpm as source-build requirements and is licensed under AGPL-3.0.
MiniMind is an open-source tiny large-language-model project and end-to-end training tutorial developed by Jingyao Gong. Its main dense model is approximately 64M parameters, and the repository provides the model architecture, tokenizer, datasets, inference code, and training pipeline. The pipeline covers pretraining, supervised fine-tuning, hand-written LoRA, DPO, PPO, GRPO, CISPO, model distillation, tool calling, adaptive thinking, and agentic reinforcement learning. Core algorithms are implemented directly with native PyTorch rather than relying on high-level abstractions from third-party training libraries. Its agentic reinforcement-learning path performs multi-turn rollouts, executes generated tool calls, appends tool observations to the context, calculates trajectory-level rewards, and updates the policy; rollout can use local PyTorch generation or an SGLang server. MiniMind supports dense and mixture-of-experts variants, single- and multi-GPU training, YaRN-based RoPE length extrapolation, Transformers-format models, and inference through llama.cpp, vLLM, or Ollama. The repository also includes a Streamlit chat interface and a lightweight OpenAI-compatible API server with tool-call and reasoning fields. It is released under the Apache License 2.0.
claudish-to-english is a Claude Code plugin that uses a language model to produce a plain-language rewrite of each assistant message while leaving Claude's reasoning and saved transcript unchanged. Its display hook buffers streamed message chunks, rewrites the completed message through Ollama by default or through the Codex CLI, Anthropic API, or an OpenAI-compatible API, and displays the result in append or replace mode. It follows the input or configured output language and supports configurable rewrite styles, prompts, models, and runtime toggles through the /claudish command and environment variables. An optional PostToolUse hook rewrites Markdown files in a configured directory into a sibling .plain.md file or in place. The hooks fail open: provider errors, missing dependencies or keys, timeouts, unavailable models, and capped outputs leave the original message or file unchanged.
Gemma Translator is an open-source, on-device voice-translation application designed to run on a Raspberry Pi 5 with 8GB RAM. It runs the Gemma 4 e2b model locally through LiteRT-LM and uses Moonshine for speech recognition and text-to-speech, providing a two-lane interface for conversations between two people. Each lane records speech, transcribes it, translates it with Gemma, and speaks the result in the other lane’s selected language. The React/Vite frontend communicates with a Python API server and is styled for small displays such as 480×320 kiosk screens. The repository includes model-download, startup, and deployment scripts, production Raspberry Pi OS systemd kiosk service registration, Chromium autostart, and support for development or production modes and landscape active-person or vertical two-operator keyboard layouts. After setup and model download, inference runs without an internet connection.
Scroll Craft is a Claude Code skill for building scroll-driven websites. It treats scroll as the page timeline, supporting frame-by-frame video scrubbing, pinned sections, sideways rails, line-by-line headline assembly, shifting page backgrounds, and pointer-driven interactions. Each build selects one of eight mutually exclusive page grammars and must create a site-specific signature interaction; a fingerprint gate requires it to differ from previous builds across at least four of six dimensions, including grammar, navigation, hero, act shape, ending, and signature move. The skill also applies design and validation rules for typography, spacing, color, depth, motion, feeling curves, and a single engineered peak, while refusing patterns such as identical feature-card grids, gradient text, invented statistics, fake dashboards, and AI-purple gradients.
Utopia is a self-hosted enterprise knowledge platform from DeepLethe, implemented as a Rust binary with a PostgreSQL service using pgvector. It ingests documents and synchronized web, RSS, GitHub, and Jira sources, then combines Tantivy full-text search with pgvector retrieval through reciprocal-rank fusion; its chat answers include inline citations linking to source passages and can use OpenAI-compatible endpoints. Its knowledge model is a bitemporal graph: extracted entities and facts retain validity intervals, evidence, and revision history, so corrections link new facts to closed versions instead of overwriting them. An editable ontology drives three-stage entity resolution, temporal Datalog forward-chaining, provenance-preserving derivation, conflict detection, and ontology-driven queries over mounted PostgreSQL databases. Low-confidence extractions, entity-merge candidates, and cardinality conflicts enter a review queue, while decisions such as confirmations, rejections, merges, and graph rebuilds are recorded in a queryable ledger. The application includes a browser console, graph browser, ontology workbench, and multi-user knowledge-base permissions, and can be deployed offline with Docker. The repository identifies the project as version 0.1 and licenses it under Apache-2.0.
slotstream is an open-source macOS local-LLM inference engine that runs Qwen3.8-Flash-Next, a 125-billion-parameter mixture-of-experts model, on Apple Silicon Macs whose memory cannot hold the model. Implemented as a single Swift binary using MLX, it keeps the dense trunk resident and reads routed expert weights from an SSD into a shared cache of slots, sizing that cache to available memory instead of loading the full model. It supports prefix caching for follow-up turns and exposes Ollama-compatible and OpenAI-compatible chat APIs, as well as a command-line prompt mode; the documented API subset supports streaming, CORS, and sampling options but not tools, images, JSON-schema output, or logprobs. The project requires macOS 14 or later, Apple Silicon, and roughly 110 GB of free disk space for the model weights. Its installer provides signed release binaries, and the repository can also be built with Swift and the macOS Command Line Tools. The source is MIT-licensed; the model weights are distributed under the Qwen community license.