1,038 tools and products — trending open source, and what gets used in AI and other work.
KTransformers is an open-source research framework for efficient large-language-model inference and fine-tuning through heterogeneous CPU-GPU computing. Its user-facing capabilities are provided through the kt-kernel source tree: high-performance inference and supervised fine-tuning (SFT). For mixture-of-experts models, it uses heterogeneous expert scheduling, placing frequently used or “hot” experts on GPUs and other experts on CPUs with optimized AMX, AVX512, or AVX2 kernels and NUMA-aware memory management. It supports CPU-side INT4/INT8 quantized weights, GPU-side GPTQ, native BF16 and FP8 precision, and multiple CPU, GPU, and NPU backends. The SFT capability integrates with LlamaFactory for hybrid CPU-GPU fine-tuning of large MoE models, including LoRA and full-parameter training, with INT4/INT8 quantization options.
Hiring Agent is an open-source Python resume-to-score pipeline. It extracts structured data from resume PDFs, enriches candidate information with GitHub signals, and evaluates projects, work experience, and technical skills with supporting evidence to produce an explainable evaluation. The repository describes it as a ranking aid for prioritizing resume review, not an applicant-tracking system or a product for HackerRank customers. It includes local-model configuration, with Gemma 4 listed as the default demonstration model, and provides command-line setup and usage instructions.
oh-my-pi is an open-source terminal coding agent and fork of Pi by Mario Zechner, with an IDE-oriented interface wired into the agent. It supports context-aware code edits, code search, language-server operations such as renames, debugging through debugger operations, persistent Python and Bun code-execution workers, web browsing, sub-agents, and MCP; the project describes support for dozens of model providers and includes built-in tools. It runs on macOS, Linux, and Windows, with installation options including a shell installer, Homebrew, Bun, Nix, and PowerShell.
Klaat Code is an open-source, terminal-native AI coding agent from KlaatAI. It reads and edits project files, runs shell commands, performs type checks after edits, and asks for permission before risky operations. It supports planning, sub-agents, MCP, hooks, and skills, and indexes projects into a code knowledge graph containing symbols, callers, callees, and semantic-search data for targeted exploration. Its hosted Klaatu-o1 router classifies each request and dispatches it across six model-cost tiers—nano, fast, code, reason, heavy, and titan—rather than using one model for every task. The CLI is a thin client to this service; the repository states that routing decisions, model health tracking, pricing, and the code-graph index run server-side. It supports providers and models including Claude, GPT, Gemini, and DeepSeek, and can be installed with npm.
Frontman is an open-source, browser-based AI coding agent for visual frontend editing. It lets users select an element in a running application, describe a change in plain language, and edit the underlying source files with hot reload. The agent uses the live DOM, component tree, computed CSS, routes, source maps, and server logs to identify the relevant code, then presents the resulting diff for review. It supports frontend projects using frameworks and tools including Next.js, Astro, Vite, React, Svelte, and Vue, and is distributed under Apache 2.0 / AGPL-3.0 as listed in its repository.
Project AIRI is a self-hosted AI companion for creating virtual characters or digital pets, with Live2D or VRM avatars and real-time voice chat. It connects to multiple model providers and can interact through Discord, Telegram, and games including Minecraft and Factorio. The project supports web, macOS, and Windows, and provides installation options through Winget, Scoop, and Homebrew Cask.
Exo is an AI agent harness for recursive self-improvement, providing tools, tasks, integrations, and a layered architecture that separates executive policy, protected state, and sandbox execution. It can inspect its own code and runtime logs, incrementally modify prompts, memory, tools, and harness policy, clone itself, and manage a lineage of clones. An event log remains outside these modifications as a canonical record intended to prevent recursive loops.
CoachAI is an iOS fitness app presented at coachai.tech that uses an iPhone’s sensors and AI-powered, on-device movement analysis to provide guided workouts, exercise coaching, and feedback.
mpai is a small GitHub-hosted project (presented on a GitHub Pages site) that positions itself as a collaborative AI coding tool, describing the concept as "AI coding is multiplayer now" and using the tagline "Open your teammate." The project appears to be created and published by the GitHub user godfaddaai.
MascotAI is a web application that generates animated SVG mascots for web and mobile applications. It offers customizable gestures and themes, and produces download-ready mascot asset packs for direct integration into products.
Airtop is a platform that automates Google Ads campaign management. Its Google Ads offering builds, monitors, and optimizes campaigns from a conversation, providing keyword research, campaign creation, waste audits, and performance reporting.
Human Review is a local visual review tool and agent skill for editing HTML and Markdown files, reviewing localhost pages, and sending feedback to an AI coding agent. It opens a file or URL in a browser, where users can edit text and basic formatting, rearrange or remove page elements, resize or paste images, and attach comments to exact phrases, images, charts, or sections. Direct HTML edits and image changes save automatically; Markdown and localhost reviews are sent to the agent, which applies the changes to the source and refreshes the page. The tool includes CLI commands for review sessions, polling, status, and setup, plus a local server, editing and feedback SDK, browser client, Markdown renderer, and instructions for agent harnesses such as Claude Code and Codex. It runs on the user's computer without an account, cloud service, database, or API key.
ASD-STE100 Skill is a Claude Code skill that rewrites dense or ambiguous English into ASD-STE100 Simplified Technical English for AI-agent outputs, including tool descriptions, error messages, READMEs, and inter-agent instructions. It offers Strict mode for procedures, error messages, and tool descriptions, and STE-flavored mode for explanatory prose. The skill reads the input for meaning, flags sentence-level issues such as ambiguous wording, complex tense, unclear passive voice, multiple instructions, long sentences, noun clusters, phrasal verbs, nominalizations, semicolons, hedge stacks, and marketing adjectives, then rewrites flagged sentences while retaining facts, conditions, and scope qualifiers. It can also provide a before-and-after table naming the rules that each sentence breaks.
tokentab is a local tool that reads Claude Code session logs and converts token counts into cost reports organized by model, project, day, and activity. It provides JSON output and a local dashboard, with the reported data remaining on the user's machine.
An open-source, executable playbook of agent skills for U.S. utility-patent work. It provides SKILL.md instructions, checklists, intake questionnaires, examiner and adversary protocols, and deterministic Python checks for dates, claim counts, and fees. Its three main workflows are a pre-filing audit, a simulated USPTO prosecution examination loop, and an adversarial design-around test. The repository grounds its simulations in 35 U.S.C., 37 CFR, the MPEP, and current USPTO materials that the agent is instructed to refresh, but it does not file applications, contact the USPTO, provide legal advice, or perform acts reserved for registered practitioners. It is licensed under GPL-3.0.
MiniMax Music 3 is a music-generation model from MiniMax AI that creates complete songs of up to five minutes from lyrics and a music description. Lyrics can use section tags such as verse, chorus, bridge, instrumental, solo, and outro; the music description specifies style, emotional progression, vocals, instrumentation, arrangement, and production. A structured-caption format can organize these instructions into global metadata, vocal details, and arrangement sections. The model uses a hierarchical autoregressive hybrid architecture: an 8B Global LLM predicts the first residual-vector-quantization codebook frame by frame to model long-range musical structure, while a 0.6B Local LLM predicts the remaining acoustic codebooks within each frame. Its continuous hidden-state synthesis module combines the two models' final hidden states with Flow Matching and Flow-VAE to produce 32 kHz, 16-bit stereo WAV audio, preserving information for vocal articulation, instrumental texture, and temporal continuity.
Agent-Safe Pipeline is a Decionis TypeScript library and runnable reference implementation for routing AI-agent actions through an independent authorization boundary. It captures an agent’s proposed action, target, and parameters as an immutable intent, sends the intent to a Decionis policy gate for an ALLOW, ESCALATE, or BLOCK decision, and coordinates verified human approval followed by policy re-evaluation for escalated actions. Agents cannot authorize their own actions, access downstream privileged credentials, or select trusted handlers. SafeExecutor accepts only the captured intent and decision, uses a sealed ActionRegistry to map action names to trusted handlers and validate parameters, and consumes a single-use grant bound to the intent before calling an API. The repository includes examples for Shopify refunds, GitHub deployments, procurement, and governed MCP tools, plus canonical-hash conformance vectors and architecture, threat-model, and security-evidence documentation. It is a self-hosted reference implementation rather than a hosted authorization service; its documented safety boundary also depends on provider-side identity, least privilege, network isolation, and incident response. The demos require Node.js 22.14 or later and pnpm 9.
oss-pr-reviewer is an AI-powered command-line tool for reviewing GitHub pull requests and producing structured Markdown or JSON reports for maintainers. It fetches pull-request metadata and changed-file patches with Octokit, normalizes reviewable content, skips binary, patchless, truncated, and oversized files with explanations, and splits large text diffs into deterministic batches without cloning or executing the repository code. The tool sends batched changes to OpenAI for structured findings, validates responses with Zod, then merges, deduplicates, orders, and filters findings by severity. It can apply repository-specific review rules and ignored paths from a trusted base-branch `.oss-pr-reviewer.yml` file, retry transient GitHub and OpenAI failures, and optionally publish bounded advisory reports through GitHub Actions. It requires Node.js 20 LTS or newer and, according to the repository, is run from a checkout because its package-ready release is not yet published to npm.
MCP-Memory is a Model Context Protocol server that gives AI coding agents persistent memory across chat turns and sessions. It stores memory records as human-readable Markdown documents following the Open Knowledge Format (OKF v0.2), with YAML frontmatter, hierarchical index files, and an update log, while maintaining a local SQLite FTS5 index for key lookups, keyword search, tag filtering, and namespace-scoped retrieval. The server exposes MCP tools for storing, retrieving, searching, and deleting memories, and supports project, user, and default namespaces; a setup wizard configures supported MCP clients including Claude Desktop, Cursor, Antigravity, Windsurf, and Codex.
pi-transcribe is a Pi extension for local speech-to-text dictation. It records microphone audio through a configurable Pi terminal shortcut, processes it with a locally downloaded speech model, and inserts the transcription at the editor cursor; streaming-capable models process audio in roughly 500 ms chunks, while other models process the complete recording after capture stops. It also registers a transcribe_file tool for local audio and video files and a /transcribe command for configuring languages, models, microphones, and shortcuts. File transcription requires FFmpeg, shares a loaded model across queued jobs, admits at most two file operations at once, runs one FFmpeg decoder at a time, and limits decoded audio to 128 MiB.
An experimental community preset for DeepSeek Harness, an LLM-agent harness. It starts a session with the real Minimal tool schema containing only a shell tool and a file-reading tool, using that condition to anchor the model's initial trajectory. After the first durable tool call or assistant reply, it promotes the session to the full Standard tool catalog, making heavier tools available on demand. The repository also includes related zero-anchored, self-introduction, prefab, and eternal-Minimal modes, and states that it is not an official DeepSeek preset or affiliated with DeepSeek.
Multi-Agent Workbench is a local-first browser control room for orchestrating coding agents such as Claude Code, Codex CLI, and generic PTY-based command-line agents. Agents work in task rooms under editable YAML role cards that define permissions, speaking style, and decision boundaries; a rule-driven orchestrator pauses execution and risky commands for human approval, while persistent sessions expose terminal activity and status. The system uses a local event-driven architecture with a React interface, Fastify backend, REST API, server-sent events, agent adapters, an approval service, and a SQLite event store. Every action passes through the API and is recorded in an append-only event log, whose events are projected into timelines, decisions, terminal views, artifacts, and diff views. It also provides task templates, Markdown reports and decision logs, real-time room updates, and configurable risk resolutions. Code-writing agents are intended to use isolated Git worktrees, with diff and patch review before changes reach the main workspace; the repository marks this worktree and merge flow as in progress.
Graft is an open-source codebase-context tool for coding agents, developed by NanoNets. It builds a repository graph and stores the result as linked Markdown files, with one node for each system, API, or concept and links describing how parts of the codebase connect. The graph supplies targeted context to agents instead of requiring them to rediscover repository structure during every task. It provides a CLI and MCP server, with integrations for Claude Code, Cursor, Codex, Gemini, and other coding agents. The CLI supports repository initialization, graph building, targeted search and orientation through commands such as `graft grep` and `graft map`, and graph visualization. `graft init` can install the agent wiring and background rebuild hooks; the generated graph is a local, regenerable cache that is added to `.gitignore` by default.
lumabri is a pure-C system for running large mixture-of-experts models across a swarm of peer machines using the colibri engine. A machine hosting a model serves it to other machines; model bytes needed during inference are fetched from peers on first use and retained in a shared content-addressed store. With Segment available, peers retain assigned layer ranges and their state while clients receive the tokenizer, embeddings, final transform or head, and conversation state; the system can instead fall back to the expert/CAS engine or local execution when a complete Segment route is unavailable. It supports CPU- and SSD-first participation as well as GPUs, and provides serving and chat commands, a terminal UI, configurable disk and compute donation, machine and health diagnostics, swarm monitoring, and low-priority peer-to-peer or signed relay transfers. The video also describes routing of missing model blocks or expert activations between peers, with hashes, signatures, and optional encryption used for protection.
mcptoon is a cross-agent Model Context Protocol (MCP) management CLI that gives agents access to configured MCP servers through ordinary command-line calls, rather than loading large tool catalogs into their context. It discovers existing server configurations, maintains ~/.mcptoon/config.json as a single source of truth, synchronizes that configuration across agents, serves a tool-name manifest on demand, and composes a single-entry proxy for MCP access. Tools are launched when invoked, and the project supports native MCP JSON, structured output, optional TOON compression, and envelope passthrough; the videos also describe output checks for prompt injection and credential leaks. The project is distributed as a Python package for Python 3.10+ and is described as using only the standard library, with no third-party dependencies. Its repository documents compatibility with agents including Claude Code, Cursor, and Codex, and provides commands such as `quickstart`, `demo`, and `call`. It is free and open source under the Apache-2.0 license.
Grok2API is a self-hosted, Go-based API gateway with a built-in React administration console. It manages separate account pools for Grok Build, Grok Web, and Grok Console, synchronizes account credentials, quotas, and models, and provides unified OpenAI- and Anthropic-compatible APIs with account and model routing. The gateway exposes Grok capabilities including chat, image generation and editing, and video. The project is intended for technical research and learning and asks users to comply with Grok's terms of use and applicable laws.
An open-source agent skill and standard-library Python service for stripping AI provenance marks from text and files owned by the user. The skill is a thin HTTP client that invokes a separately running service, so the agent host does not need Python. Deterministic processing handles invisible Unicode characters, exotic spaces, bidirectional and tag characters, while an agent rewrite pass and optional rewrite_text.py hook address statistical token-sampling watermarks. File processing targets C2PA, EXIF, XMP, and document properties across image, document, ebook, web, video, and audio formats, including PNG, JPEG, WebP, AVIF, HEIC, BMP, GIF, TIFF, SVG, PDF, DOCX, XLSX, PPTX, EPUB, ODT, HTML, Markdown, MP4, MOV, M4A, M4V, WAV, MP3, and FLAC. The repository describes support for class-level provenance ecosystems including Claude, Gemini/SynthID-Text, OpenAI provenance surfaces, Kirchenbauer-style green-list marks, and keyed-Gumbel/EXP marks. It includes installable skills for Claude Code, Cowork, Cursor, and other supported hosts; the installer uses Python 3.10+ with no external dependencies.
Soloop is a startup offering a “one-person company OS” for solo founders, independent makers, and indie founders. It provides an AI-based founding team and AI-powered tools and workflows intended to help users grow and monetize the products they build.
withoutBG is a Python SDK and command-line tool for removing image backgrounds locally or through a cloud API. Its local mode runs open-weight CPU ONNX models on the user's machine, downloading the model weights once; processing then works offline without an API key or GPU. The same API supports single-image and batch processing, progress callbacks, and returns transparent RGBA PIL Images. The package also provides a cloud mode through the withoutBG API and can be installed with uv or pip.
Agentic Inbox is a self-hosted email client from Cloudflare that runs on Cloudflare Workers through the user's Cloudflare account. It receives and sends mail through Cloudflare Email Routing, isolates each mailbox in a Durable Object with SQLite storage, and stores attachments in R2. The web interface supports threaded replies and forwards, rich-text composition, folders, search, and attachments. Its built-in AI agent, using the Cloudflare Agents SDK and Workers AI, can read inboxes, search conversations, draft replies, and send messages through email tools; automatic drafts for incoming mail require explicit confirmation before sending. Deployment provisions Workers, Durable Objects, R2, and Workers AI, and the setup requires Cloudflare Access, Email Routing, Email Service, and a mailbox on the configured domain.
Langfuse is an open-source LLM engineering platform for developing, monitoring, evaluating, and debugging AI applications. It ingests traces for LLM calls and related operations such as retrieval, embeddings, and agent actions, with integrations including OpenTelemetry, LangChain, the OpenAI SDK, LiteLLM, and other LLM frameworks. Its features include prompt management with version control and collaborative iteration, LLM-as-a-judge and code-based evaluations, user feedback and manual labeling, datasets for test sets and benchmarks, trace replay and debugging, metrics, and a playground for testing prompts and model configurations. Langfuse exposes APIs and typed Python and JavaScript/TypeScript SDKs for custom LLMOps workflows. It is available as a managed cloud service or for self-hosting with Docker Compose, virtual machines, Kubernetes via Helm, or Terraform templates.
fak is a self-hosted agent kernel distributed as a single Go binary. It wraps existing coding-agent sessions without replacing their interface or model, managing shared context, provider cache continuity, session recovery, model routing, tools, policies, and execution receipts. Tool calls pass through explicit policies that can allow, deny, transform, or witness them. Its default-deny capability floor blocks tools outside the configured allow-list, while the system can forward an existing provider subscription credential, route tasks to a configured provider or its native inference engine, and avoid silent fallbacks. The repository also provides an offline proof mode and installation through a shell script or Go.
Superlog is an open-core, self-hosted observability workspace for OpenTelemetry data, developed by Superlog Labs. It ingests traces, logs, and metrics, groups noisy signals into incidents, and provides a local-first interface for investigating production systems. The open-source community edition includes a web app and API, an OTLP intake proxy, worker processes for incident grouping and background jobs, a Postgres schema, ClickHouse-backed telemetry queries, and pluggable agent-runner interfaces with a default runner that records a local incident summary. The repository is licensed under Apache License 2.0; a hosted Superlog Cloud edition is also available with a free tier, pay-as-you-go plan, and monthly credit packs.
Renfield is a self-hosted, voice-first AI household assistant built with FastAPI, React, and Ollama. It uses Raspberry Pi Zero 2 W voice satellites with ReSpeaker HATs for local wake-word detection, Whisper speech recognition, Piper text-to-speech, and SpeechBrain speaker recognition. The system can track room presence through BLE scanning, voice recognition, and web authentication, and integrates with Home Assistant and DLNA for multi-room control. Its ReAct agent chains tools through 10 MCP servers covering weather, web search, news, calendars, Jellyfin, DLNA, n8n workflows, Home Assistant, Paperless, email, and related services. It also provides conversational long-term memory with contradiction detection, a knowledge graph with validated entity-relation triples, and a RAG knowledge base that combines dense pgvector embeddings with BM25 full-text search using reciprocal-rank fusion. Documents can be uploaded from formats including PDF, DOCX, PPTX, XLSX, HTML, Markdown, and TXT, with Tesseract or EasyOCR available for scanned or poor-quality PDFs. The project is designed to run entirely on local hardware without cloud dependencies and includes a web interface, satellite monitoring, proactive notifications, plugin hooks, and Paperless document-audit workflows. The repository describes more than 100 available tools and supports Docker-based self-hosting.
best-claude-hud is a Rust-based statusline HUD for Claude Code terminals. It displays the active Claude model and live reasoning effort, Claude Code's launch directory, workspace, Git branch and status including ahead/behind counts, and context-window usage from Claude Code's statusLine data with an active-transcript fallback. Optional segments provide usage or rate-limit metadata, cost, session, and output-style information. It is distributed through npm with prebuilt native binaries, so Rust is not required for installation. The repository also provides a Nix flake for running or installing the tool in declarative environments. Its setup command writes the required statusLine configuration to Claude Code's settings file while preserving existing settings.
agentacct is a local-first audit and observability tool for coding-agent work. It reads session logs from Claude Code, Codex, OpenCode, and Hermes, combines them with recorded tasks and test evidence, and produces a Work Receipt for each task covering commands run, files touched, tools used, time and token usage, cost, outcome, and evidence strength. Agent claims and machine verification are kept as separate axes, with verified status reserved for tasks whose live checks pass after the latest recorded work. The data is available through a macOS app, the live terminal dashboard launched with `agentacct tui`, and a local JSON API. State is stored in local files; the API listens only on loopback, and the project states that it has no login, cloud sync, or telemetry and does not request provider API keys.
AI Copywriter is a portable, plain-Markdown agent skill for writing marketing copy and removing common patterns associated with AI-generated prose. It supports headlines, short descriptions, microcopy, subject lines, button labels, and LinkedIn posts, and runs in any harness that supports skill-style instructions. Before drafting, it interviews the user about the intended customer, product category, and underlying story, then pressure-tests the story for specific and interesting details. Its copywriting method focuses on the reader's immediate situation and explains the product in simple language. It incorporates the 33 detectable and fixable writing patterns from blader's Humanizer while adding the reverse workflow of generating marketing copy designed around those constraints.
MoonEP is an expert-parallelism communication library from Moonshot AI for mixture-of-experts training and inference. It plans a small number of dynamic redundant experts from the current router outputs, prefetches their weights before expert computation, and reduces the resulting gradients back to the experts' home ranks during the backward pass. Its near-optimal GPU planning kernel, fused permute/unpermute operations, and zero-copy dispatch send tokens directly to expert-grouped positions on remote ranks and return buffer views to computation. The library uses a fixed S × K token buffer per rank, where S is the input-token count and K is the routed top-k, keeping communication and computation shapes static across layers. The repository documents support for NVIDIA GPUs and lists Zhenwu PPU support as under review.
Better Harness is an open-source Harness Engineering platform for coding agents. It analyzes project and, where supported, session evidence around an agent's work rather than only its final code diff, evaluating the Agent Work Loop for issues such as task understanding, validation, delivery control, and retained lessons. It turns supported gaps into prioritized findings linked to their evidence, expected outcomes, repair boundaries, and acceptance checks, and produces host-specific reports in formats including HTML, paired Markdown, or native Canvas reports. The platform also supports defining harnesses as code, running controlled experiments, inspecting evidence, and comparing outcomes across coding-agent hosts such as Claude Code, Codex, Qoder, Cursor, and GitHub Copilot CLI.
humanizer-stack is an open-source writing-humanization toolkit packaged as Claude Code Skills. Its two-pass pipeline first scans and revises surface-level tells such as hype vocabulary, em-dash overuse, formulaic phrasing, antithesis, vague attribution, inflated symbolism, and redundant exposition; a second structural pass audits discourse-level patterns grounded in the StoryScope study. The structural audits cover theme explicitness, structural tidiness, emotional framing, reference specificity, reader engagement, and shape convergence, with deterministic Python scanners for mechanical copy and structural tells. The repository includes the skills, reference materials, genre-specific guidance, and documentation describing how the passes are chained.
Lexicon is a free, open-source, local-first writing assistant and distraction-free rich-text editor for Windows, macOS, and Linux. It provides inline proofreading through the separate LanguageTool engine, while local AI tools handle rewriting, tone adjustment, summarization, and other writing tasks without requiring an account or cloud writing service. Users can review suggestions in a dedicated panel, format drafts with rich text and LaTeX, autosave work, and export it to Markdown, HTML, text, or PDF. The app can download and run a local AI model or use an existing local Ollama server; the editor and proofreading tools work without an AI-model download. It is MIT-licensed, and the repository notes that its Windows builds are unsigned and its macOS build is pre-release and unsigned.
opentax-engine is an open-source deterministic US tax calculator designed for AI agents and other applications. It accepts tax facts such as wages, filing status, and dependents, applies an encoded rule corpus, and returns cent-exact results with assumptions, legal citations, and a proof tree showing how each value was derived. If its rules cannot derive an answer, the project is designed to refuse rather than guess. The calculation can be selected for a date with an --as-of option, and the CLI accepts flags or JSON fact files; the videos also identify CLI, browser, and MCP integrations. The engine is implemented in TypeScript with no platform dependencies and can run in a self-contained browser HTML file containing the engine, rule corpus, and verifier. Proofs use canonical JSON, hashing, and a Merkle construction; the verifier independently re-derives the calculation and reports altered proof data, rule-corpus differences, or steps that do not match. The repository describes the project as AGPL-3.0 licensed, with commercial licenses available.
τ₀-VLA is an open-source hierarchical robot foundation model and reference implementation for long-horizon manipulation. A memory-augmented high-level policy generates the next subtask and uses world-model-guided test-time computation to search over alternatives when additional reasoning is needed; a generalist low-level policy then executes the selected subtask across robot embodiments. The low-level policy combines a Qwen3.5 vision-language backbone with a Mixture-of-Transformers action expert trained through conditional flow matching, using a unified 40-dimensional state/action space and multimodal co-training on tens of thousands of hours of heterogeneous real-world robot data. The repository provides LeRobot-format data loading, prompting, masking and normalization, embodiment-specific adapters, training and post-training recipes, deployment contracts, a policy server, and open-loop evaluation tools. It includes an AgiBot World example subset and templates for other datasets and robots. Public v1 serving supports joint-control checkpoints; native end-effector data can be used for training but end-effector serving is not supported in that release. The reference environment uses Python 3.11, CUDA 12.8, and PyTorch 2.7.1. Code and model weights are released under the Apache License 2.0.
OptMem is a permanent-memory tool for AI agents. It installs as a dependency-free Python 3 script and integrates through a prompt block added to an agent's AGENTS.md or CLAUDE.md file. The agent uses `memo wake` at session startup, `memo note` to append one-line memories, `memo recall` for exact regular-expression searches, and `memo zoom` to navigate summaries. Raw memories are stored in an append-only `LOG.txt`; a binary tree of one-line summaries provides a compact reading view, with summaries rebuilt when needed. `memo nap` processes due merges, while `memo forget` discards a summary so the next nap can rebuild it. Memory can be kept in a configurable directory such as a synced folder or Git repository, and the system is designed to persist across sessions, context compaction, models, and vendors.
Herdr Browser is a deprecated browser-automation experiment that renders a real Chromium view inside a Herdr terminal pane. It captures Chromium frames through the Chrome DevTools Protocol, encodes them as PNG images, and sends them through Herdr's pane graphics stream for display in the terminal; automation clients can drive the browser over CDP while users view or interact with the persistent session. The project is no longer maintained and directs users to Terminal Browser, which uses Electron offscreen rendering and file-backed raw RGBA frames instead of the PNG encode/decode path.
Deltafin is a native binary for running the full, unpruned Kimi K3 mixture-of-experts model on consumer hardware, along with an OpenAI-compatible API server for local chat and coding agents. Its demonstration approach keeps the model's attention spine on local disk and fetches routed experts from Hugging Face's CDN as needed, allowing the full model to run on a 64-gigabyte MacBook rather than loading the entire expert bank into memory. The project emphasizes preserving Kimi K3's original expert weights and having K3 verify every token; it is presented as an experiment in pushing large-model inference on local hardware rather than as a practical chat setup.
OpenWorker is an open-source desktop AI coworker that produces finished deliverables from everyday tasks, including code-security reviews with proposed fixes, cloud-configuration audits, incident reports, documents, spreadsheets, drafted messages, and scheduled briefs. It runs on the user's machine through a native desktop app and a local Python agent server built on aisuite, working across local files, the terminal, connected applications, and more than 25 connectors. Specialist coworkers cover security review, cloud posture, incident triage, everyday work, and recurring automations; security review combines deterministic scanners such as Semgrep with model reasoning, then re-scans and diff-reviews proposed fixes before approval. The agent breaks requested outcomes into steps and uses models from providers such as OpenAI, Anthropic, and Google, open-weight providers, or a local Ollama deployment, with the user supplying model access. Consequential actions—including sending messages, changing calendars, and running commands—are approval-gated. Its governance design includes human-only floors for dangerous or irreversible operations, explicitly granted and revocable autonomy rules, circuit-breaker escalation for uncertain or repeatedly denied actions, and an audit trail recording tool calls, approval provenance, and reviewer reasoning; unattended runs cannot self-approve. The project is in open beta and provides downloads for macOS on Apple Silicon and Windows on x64.
scroll-world is an agent skill for Claude Code, Codex, and other SKILL.md-compatible agents that builds scroll-scrubbed 3D world landing pages for brands and industries. It generates cohesive isometric diorama scenes and image-to-video camera flights, linking scenes through first/last-frame conditioning so the camera appears to move continuously as the visitor scrolls. The skill provides prompt templates, an AI image and video rendering pipeline using Higgsfield, Monid, Seedance, Kling, or Codex image generation, and a framework-agnostic vanilla JavaScript scrubbing engine that can be used with plain HTML, Next.js, Vue, or a Python-served page. It can also render a separate portrait chain for mobile devices and requires external rendering services or CLIs, ffmpeg/ffprobe, and Python with Pillow.
No AI Slop is an AI-writing editing skill that removes more than 20 canned machine-writing patterns while preserving the author's vocabulary, cadence, humor, and imperfections. Its rules target binary contrasts, throat-clearing openers, faux-insight setups, colon reveals, dramatic fragments, superficial analysis, importance puffery, weasel attribution, synonym cycling, and fake-profound endings; they also cover active voice, leading with the point, untangling difficult sentences, and preferring concrete details over abstractions. The skill supports editing text, detecting and quoting suspected patterns without judging whether AI produced the writing, and generating deliberately exaggerated AI writing for satire. It can be invoked as `/no-ai-slop` in ChatGPT, Claude Code, Codex, or another coding agent, and is distributed through the Skills package manager or an npx command. The repository contains the editing rules in SKILL.md, evaluation checks in eval.md, ChatGPT and Codex plugin metadata, and a plugin build script; it is licensed under MIT.
OpenScience is an open-source AI workbench for scientific research developed by Synthetic Sciences. It runs as a browser-based workspace where a single research agent takes a goal through literature review, hypothesis formation, code writing and execution, experiments, analysis, and a written report. The agent can load domain skills, delegate bounded exploratory or execution work, query scientific databases, and preserve sessions as observable research traces. The local server hosts the workspace UI, agent runtime, skill library, and tool layer. The agent plans with a research harness and calls shell, editor, LSP, MCP, scientific database, and skill tools; sessions, skills, artifacts, and provenance are stored on disk. The workbench includes a file tree, code editor, terminal, session history, and inline rendering for molecules, structures, genomes, and plots. Its bundled skills cover training, evaluation, datasets, molecular and clinical biology, cheminformatics, papers and LaTeX, figures, and scientific runtimes, while database connectors provide direct access to services including UniProt, PDB, Ensembl, ChEMBL, PubChem, arXiv, OpenAlex, and Semantic Scholar. It supports frontier and open-weight models through user-provided keys, eligible ChatGPT or Codex access, local models, and optional Ace-managed models. Models are routed per request, allowing providers to be switched without changing the workspace. Extensibility includes LSP integration, MCP servers, plugins, custom agents and commands, experimental Python environments and BioNeMo adapters, and a TypeScript SDK. It is installed with the @synsci/openscience npm package or launched through npx; platform binaries and desktop installers are distributed through GitHub Releases. A free Synthetic Sciences account links an installation and can provide credential synchronization, private research graphs, enhanced search, and optional credit-backed models, while BYOK and local-model usage remain separate from Ace.
An open-source project that runs a 28.9-million-parameter language model locally on an ESP32-S3 microcontroller, without network connectivity. It uses a tiered memory layout based on Per-Layer Embeddings: frequently accessed activations and normalization weights stay in SRAM, the model core and output head reside in PSRAM, and a roughly 25-million-parameter embedding table is stored in flash; only the few rows needed for each token are read into faster memory. The included TinyStories model generates short stories, while a Barista model provides espresso question answering. The repository notes that the TinyStories model is not designed for general question answering, instruction following, code generation, or factual knowledge, and provides separate scripts for downloading and verifying model assets and deploying a selected model to the board.
The Fable Method is an agent-workflow repository that distills the Fable Workflow into skills that models can run. Its process classifies the request, defines completion with named verification, gathers evidence from primary sources in parallel, commits to one recommendation, changes the smallest correct thing, verifies the result by observation, and reports the outcome with caveats. The repository provides four related skills—fable-method for thinking, fable-loop for acting, fable-judge for evaluating results, and fable-domain for generating domain adapters—along with evaluation cases, transcripts, judge outputs, and logs. The project reports testing the method across fifteen evaluation rounds and more than 260 agent runs, with judges checking diffs and execution rather than relying on agent reports.
riddle is a Rust application that turns a reMarkable Paper Pro into a diary-style AI interface. It reads pen strokes from the tablet, waits for an idle pause, commits the page as a PNG to a resident oracle LLM process, and streams the response back as animated handwriting. The handwriting pipeline rasterizes the response, applies Zhang-Suen thinning, traces the result into single-pixel pen paths, and replays those paths stroke by stroke. The app supports a qtfb display backend for running inside xochitl and a quill-based takeover backend that drives the vendor e-ink engine directly; the prebuilt bundle uses takeover mode. It can be configured with an OpenAI-compatible API key or used with pi, according to the repository instructions. Installation requires a reMarkable Paper Pro in developer mode with xovi and AppLoad, or the remagic installer. Takeover mode stops the normal reMarkable interface, runs as root, and takes control of the display; the repository says it has been tested on the Paper Pro ferrari model with OS 3.26–3.27 and warns that installation modifies the device and may not work on other models or versions.
Jacobian Lens (jlens) is a research and interpretability tool from Anthropic for examining what an internal language-model activation is disposed to make the model say. It linearly transports a residual-stream vector from any layer and position into the final-layer basis using an average input–output Jacobian computed over a text corpus, then applies the model’s own unembedding to produce ranked vocabulary tokens. The reference implementation fits lenses on open-weight decoder transformers, applies pretrained or newly fitted lenses, and renders an interactive layer-by-position view with top-token ranks, pinned-token tracking charts, and rank heatmaps. It is distributed as an Apache-2.0-licensed Python package and repository; the code is described as not maintained, and model weights and text corpora are not bundled.
img2obj is a Codex plugin that reconstructs an object from an attached image, screenshot, or local image path as a code-only, procedural Three.js model. It first validates the reference, then plans an ObjectSculptSpec covering the component hierarchy, geometry, materials, pivots, sockets, motion requirements, and visual priorities. The model is built through blockout, form, look-development, and interaction phases; browser renders are compared with the source image, with quality evidence and approval required before advancing. The plugin generates procedural Three.js geometry rather than extracting or downloading a mesh. It supports hard-surface and organic assets, vegetation, fabric, glass, hair or fur, emissive elements, and decals; can define animation-ready parent-child relationships, detachable parts, pivots, and sockets; and can produce reference-derived PBR evidence such as palette, roughness, height, normal, and ambient-occlusion maps. It is intended for stylized props, mechanical objects, plants, scene assets, and interactive models, not photogrammetry, exact mesh extraction, or guaranteed production-ready geometry from a single image. The repository requires Codex with local plugin support, Python 3.10 or newer, and a browser project using Three.js.
thinking-orbs is a React component package providing dotted animated loading indicators for AI and agent interfaces. It offers nine hand-tuned states— including working, searching, solving, listening, connecting, weaving, composing, breathing, and shaping—rendered with plain 2D canvas rather than WebGL or filters. The indicators have separate 64-pixel chat-avatar and 20-pixel inline-text designs, automatic or pinned light/dark themes, adjustable speed, pause control, and pass-through canvas properties. It includes per-state accessible labels, renders a static frame for prefers-reduced-motion, pauses offscreen instances through IntersectionObserver or when the tab is hidden, and resumes them in phase using a shared clock. The package is distributed through npm under the MIT license.
GC Minimal Zine Poster is a Codex skill for turning a theme, sentence, article idea, object, mood, photograph, or reference set into a sparse editorial poster, production-ready image prompt, or reusable visual system. It compiles requests into a vertical aged-paper composition with 70–90% negative space, one small visual subject or event, restrained typography, a high-chroma color accent, and xerox, risograph, halftone, letterpress, or scanned-paper textures. The skill supports Generate, Reference Analysis, Prompt-only, Analyze + Generate, and Photo Input routes; its reference-analysis and quality-review workflows require image inspection, while image generation depends on the host runtime's available model. It is distributed as a Git repository for use with Codex or another compatible Skill runtime and contains no scripts, external fonts, API keys, private paths, or downloaded runtime assets.
QuerySplat is an open-source implementation for predicting and rendering 3D Gaussian splats from multiple images of a scene while decoupling geometry and appearance representations. Its inference pipeline preprocesses custom images, predicts 3D Gaussians, and uses VGGT-Omega components for camera and depth prediction; optional test-time optimization can refine the result. It can export Gaussian PLY files, predicted input cameras, per-view depth and confidence data, and depth-derived point clouds. The repository requires Linux, an NVIDIA CUDA-capable GPU, and CUDA-enabled PyTorch.
doc7 is an open-source document-to-Markdown tool that converts PDFs, Office files, scans, screenshots, charts, formulas, and diagrams into AI-ready Markdown using an OpenAI-compatible multimodal vision model. It renders and understands complete pages rather than relying on character extraction, preserving text, tables, formulas, chart and diagram relationships, image meaning, and visible UI state; the video also describes page-level retry and resume. The model can run locally through LM Studio or Ollama or through a remote endpoint, with no required OCR stack or document-processing service. The same binary provides an interactive CLI, batch processing, model checks, MCP support, a Go SDK, and an asynchronous HTTP service, with installers for macOS, Linux, and Windows.
A ComfyUI acceleration extension for the native MiniMax H3 audio-video model. It reduces expensive transformer evaluations during sampling by retaining the packed hidden state after actual transformer steps, then using Chebyshev ridge regression to forecast that state on selected forecast steps. Forecast steps skip the H3 transformer blocks and continue through the native output and sampler path, with adaptive scheduling, sampler-aware safeguards, CPU/VRAM history storage, and native fallbacks. Spectrum is a training-free approximate accelerator: forecasted states alter the denoising trajectory, so outputs can differ from native H3 even with the same seed and workflow.