1,038 tools and products — trending open source, and what gets used in AI and other work.
Hallmark is an open-source design skill for Claude Code, Cursor, and Codex that applies anti-AI-slop rules to generated interfaces. It selects a macrostructure and one of twenty-one themes for a brief, applies its design rule set, runs fifty-seven anti-pattern gates and a pre-emit self-critique, and returns self-contained HTML and CSS. Its four operations build new UI, audit existing code without editing it, redesign an interface while retaining its copy, information architecture, and brand, or study a screenshot or URL to extract macrostructure, type pairing, and a color anchor; the study operation can also emit a portable design.md file. Custom briefs can receive a made-to-measure palette, typography, and layout rather than a catalog theme. Hallmark is made by Together AI and distributed under the MIT license.
Agent Substrate is an open-source runtime and control plane for running agent-like workloads at high density. It manages the lifecycle of sandboxed actors, including creation, destruction, suspension, and resumption, and assigns them to a smaller pool of worker pods while routing traffic to the active workers. The system uses Kubernetes for infrastructure provisioning and pod lifecycle management, with support for sandbox technologies including microVMs and gVisor. It preserves actors' volatile memory and filesystem state through full-state snapshots, enabling sub-second suspend and resume operations and multiplexing many stateful actors across shared infrastructure. It manages standard OCI containers at the kernel level and is designed to support different agent frameworks and harnesses, including ADK-compatible actors, LangChain agents, coding environments, and MCP servers. Agent Substrate is intended for running agents at scale rather than building them, and its workloads do not have to be literal AI agents. The repository notes that the project is not an officially supported Google product.
A desktop agent operating system for creative work that coordinates specialized agents to plan, execute, revise, and deliver projects across film, games, music, and social media.
AgentConnect is an open-source, self-hostable platform that integrates autonomous AI agents (ACP agents such as Claude Code, Codex, and Gemini CLI) into communication platforms including Slack, Telegram, Discord, and GitHub. It provides tagging of agents and facilities to assign roles, tools, and memory so agents can act within those channels.
Gemini CLI is an open-source, Apache 2.0-licensed AI agent that provides terminal access to Google's Gemini models. It can understand and edit codebases, generate applications from PDFs, images, and sketches, debug through natural-language prompts, and run non-interactively for scripted automation. Built-in tools include Google Search grounding, file operations, shell commands, and web fetching; Model Context Protocol (MCP) support allows custom integrations and additional capabilities. The project also supports conversation checkpointing, project-specific GEMINI.md context files, and a GitHub Action for pull-request reviews, issue triage, and on-demand assistance. It can be run with npx or installed through npm, Homebrew, MacPorts, or Anaconda.
Argos is a Chrome extension that allows users to control browser actions (open and close tabs, scroll, click, fill forms) and interact with DOM elements. It integrates with Google Docs, Gmail, and Google Sheets to support research, task automation, and workflow organization inside Chrome. The extension exposes browser automation and scripting capabilities directly in the browser environment.
DocsAlot is a documentation platform for publishing AI-readable help centers and developer documentation. The service provides hosted MCP and llms.txt support and tools to keep documentation current.
Soup CLI is a command-line toolkit for training, fine-tuning, evaluating and managing large language models (notably Llama-3.1-8B) with emphasis on running on consumer GPUs. It implements streaming/quantization (NF4), multiple training algorithms and evaluation suites, VRAM probing and model shipping/verification features. The project is published at trysoup.dev and installs without requiring PyTorch, supporting Python 3.10–3.12.
Bevel is a git-backed control plane for enterprise AI agents, providing a platform to manage agent configuration and lifecycle using Git workflows. It is presented by Bevel on its official site as a tool for organizations deploying AI agents.
Toolport is a free, open-source local MCP gateway that allows multiple AI agents to share a single server for accessing tools and managing credentials. It integrates with agents and clients such as Claude, Cursor, VS Code, and Codex, and the project claims it can reduce tool tokens and secrets stored in the OS keychain by up to 91%.
Coldtea is a software development lifecycle platform that uses autonomous agents to test code and monitor production. It runs tests on real devices for every pull request, observes production behavior, and files issues for failures before end users encounter them.
Rindler is a web automation platform from Rindler that signs into websites and performs repetitive team tasks such as pulling records, downloading statements, checking status, and submitting forms on sites that lack APIs. It is positioned as a tool to automate manual web-based busywork for teams.
BrowserOS is an open-source web browser that integrates agentic AI capabilities and runs AI agents locally on users' computers. It is presented as a privacy-first alternative to Chrome.
Troopr is an AI project manager platform for engineering teams, offered at troopr.ai. The product provides AI-driven project management capabilities and includes an AI Scrum Master capability to assist with agile workflows and team coordination.
Prompt Bridge is a Chrome extension that carries full AI conversation context across major AI platforms, enabling users to continue threads when switching between services. It is offered as a free browser extension.
GPTs is a feature from OpenAI integrated in ChatGPT that lets users create, customize, and share specialized chat models called GPTs. It includes a marketplace (Explore GPTs) for discovering and using user-created GPTs and supports public sharing and invite links.
An agentic retrieval-augmented generation platform for ingesting timestamped Markdown, filtering a corpus, and searching it conversationally.
Figure is an AI robotics company that develops a general-purpose humanoid robot. The company states its goal is to build an AI-driven humanoid capable of performing a range of physical tasks.
Cover is a company that offers technology for protection against concealed weapons. Its website (cover.ai) describes the offering as "invisible protection against concealed weapons," indicating an AI-focused approach to detection and protection.
An end-to-end model-development system created by Poolside for data pipelines, distributed training, post-training, reinforcement learning, experiment tracking, reliability, and rapid model launches.
Kilo Code is an open-source AI coding agent and agentic engineering platform from Kilo, available as extensions for VS Code and JetBrains, a command-line interface, and cloud services. It can generate and edit code across multiple files, provide inline autocomplete, run terminal commands and control a browser, and act as an external coding harness for cross-harness model testing. The platform includes Code, Plan, Ask, Debug, and Review agents, supports custom agents, allows model switching during a task, and uses self-checking to review and correct its work. It also provides MCP server discovery, autonomous CI/CD execution, cloud agents, and automated pull-request code reviews.
Wardrobe is a local clothing-library and outfit-generation application that uses OpenAI image and vision APIs. It detects garments in photos with the OpenAI Responses API, creates clean clothing cutouts with the OpenAI Images API, and can generate modeled editorial previews using a local reference photo. The application stores original images, generated images, processing jobs, and a JSON database in its local data directory, and provides drag-and-drop, paste, editing, review, regeneration, and approval workflows. The repository includes Codex skills for importing clothes from a folder and for generating modeled outfit lookbooks. The importer reviews cutouts and modeled images before writing them to the local library; the outfit skill curates, generates, verifies, and saves complete looks. It runs as a local npm application, requires an OpenAI API key and a PNG model-reference image, and is licensed under MIT.
Bonsai Demo is an open-source repository from Prism ML for running the Bonsai and Ternary-Bonsai language models locally. It supports 1-bit and ternary model families in several sizes, using llama.cpp or MLX across macOS Metal, Linux and Windows CUDA, Vulkan, ROCm, and CPU backends. Setup scripts download the models and binaries, while the included server and web interfaces provide local chat; the 27B models additionally support image, screenshot, and PDF input, OpenAI-style tool calls, MCP servers, configurable reasoning effort, and long-context conversations. The repository provides command-line launch scripts, model-family and size selection, and links to the corresponding Hugging Face collections and technical whitepapers.
LobeHub is a self-hostable platform for organizing teams of AI agents, presented as a chief agent operator. It treats agents as units of work and provides tools to create and configure agents, schedule and report on their operations, coordinate them in agent groups, and monitor shared workspaces and memory. The platform provides access to multiple models and modalities, an agent builder, a library of skills and MCP-compatible plugins, and an IM gateway for interacting with agents through existing chat services. It can be deployed using Docker or listed cloud deployment options, and is described as being under active development.
Cangjie Skill is an open-source pipeline that distills methodologies from books, long-form videos, podcasts, interviews, and courses into callable AI Agent Skills rather than summaries. The resulting skills package evidence, examples, executable steps, and usage boundaries so coding agents can invoke the extracted methodology. Its capability-bundle workflow extracts stable capability cards and metadata before compiling installable output. It supports a single router-style Skill or a compact pack containing a router and promoted standalone Skills. The unified local toolchain includes diagnostics, compilation, output replanning, incremental updates, repair, rollback, evaluation, and benchmarking, while content-addressed preprocessing, source diffs, impact analysis, transactional patches, snapshots, and edit detection support controlled evolution. The repository also provides a standalone installation package for DeepSeek Harness.
An extension for the Pi coding agent that delegates work to focused child-agent sessions. Pi acts as the parent session, assigning tasks to subagents and bringing their results back; foreground runs stream in the conversation, while background runs continue asynchronously and can be checked later. Built-in roles include scout for codebase reconnaissance, researcher for web and documentation research, worker for implementation, reviewer for code review and fixes, oracle for challenging assumptions, and delegate for general delegation. It supports parallel and chained workflows, saved workflows, council-style model-based reviews, artifacts, truncation, session sharing, and optional isolated worktrees. The extension is installed with `pi install npm:pi-subagents` and does not start automatic reviews unless requested in a prompt or project instruction.
clawk is an open-source command-line tool that runs AI coding agents in disposable Linux virtual machines instead of directly on the host. It mounts a repository from the host into the guest, keeps host files and keychains outside the VM, forwards the host's SSH agent, and restricts outbound connections with a per-sandbox network allowlist. It supports agents and shells including Claude Code, Codex, pi, and a standard shell; sandboxes can be destroyed and recreated while host-side code and agent conversations remain available, and sessions can be resumed. The repository describes the project as pre-1.0 and subject to breaking changes; allowed destinations and forwarded credentials remain accessible to the agent and can be used to publish data.
waggle is an open-source reference layer for handing off artifacts between AI agents. It replaces pasted files with compact, approximately 30-byte tokens that resolve into an agent-specific view while preserving attribution and recording which parts were read. The token travels without automatically expanding the underlying artifact. Resolve, read, and search operations return only the requested projection or slice within byte budgets, and corrections can propagate to token holders. waggle is MCP-native and runs as an MCP server for agent tools.
mindwalk is a local visualization tool that replays coding-agent sessions on a 3D map of a codebase. Its Go binary reads Claude Code, Codex, and pi session logs, showing searched, read, and edited files as changing glow and touch states on radial-tree or treemap views; files no longer present remain as wireframe ghosts. A playback deck supports timeline scrubbing, event-bucketed observation and mutation phases, speed control, and client-side WebM export, while timeline marks, subagent lenses, and a file inspector connect session events to individual file histories. It runs locally and sends nothing elsewhere during viewing; optional session evaluation sends a session summary to the model used by the user's local Claude or Codex CLI when explicitly invoked.
CLODEx is an open-source, local-first, zero-trust agentic IDE for long-running software engineering work. It maintains durable tasks with searchable history, workspace-aware context, restart recovery, and continued work across sessions, combining code editing, pending edits, line-level diffs, Git and worktrees, persistent terminal sessions, embedded browsing, screenshots, models, and MCP tools in one desktop workspace. Sensitive actions can require explicit approval and remain reviewable; its governing principle is that model output is input, not authority. CLODEx supports account-backed models, bring-your-own provider keys, compatible endpoints, and local Ollama models, and is available for macOS, Windows, and Linux.
Flawless, also named CISRE (Cloud Infrastructure Site Reliability Engine), is an AI SRE and AgenticOps control plane for Kubernetes and cloud infrastructure. It connects risk discovery, evidence collection, diagnosis, Skill-based remediation, human approval, controlled changes, recovery verification, and auditable records into a remediation loop. Its architecture separates model planning and explanation from domain Skills, composable plugins, a Harness for state, permissions, and orchestration, controlled executors for real changes, and Verifiers that determine whether the target has actually recovered. The documented flow is discovery → evidence collection → diagnosis → Skill routing → change preview → human approval → execution → same-target readback → stability verification → records; if recovery fails, the system retains evidence and can continue with another strategy. Kubernetes is documented as having a complete loop, while database, VM/host, storage, middleware, cloud-resource, and network integrations are described as contract-ready extension areas. The platform includes SRE Run, scheduled or manual AI inspections, topology impact analysis, a Skill library, a plugin center, operational dashboards, and Agent Trace. Real changes must pass through a typed action, policy and blast-radius checks, human approval, an executor, same-target readback, a recovery verifier, and a record. Its plugin-first architecture supports declared service dependencies, event-driven orchestration, reversible loading and hot reload, event-sourced audit records, replay, fork, resume, and resource-domain agents.
River is a sales automation company that offers a "digital account executive" product designed to deliver on-demand product demos to prospects and assist in closing SMB deals. The service is positioned to help go-to-market teams scale outreach and conversions without additional headcount.
Nitrosend is an AI-native email automation platform and email layer for agencies, SaaS teams, developers, and autonomous agents. It provides agent onboarding, agent-owned inboxes, marketing and transactional email, customer replies, full-stack workflows, and multi-brand campaigns, with human approval gates and a free-to-start model. Agents can be integrated by pointing them to nitrosend.com/SKILL.md.
In Parallel is a platform that turns the conversations, documents and decisions around an initiative into persistent business context that teams and their AI tools and agents can use. It is positioned to provide living context for collaboration and for AI-driven workflows.
Graft AI is an enterprise AI infrastructure product from Axcelner that enables AI agents to interact with legacy, internal, desktop, and browser-only software workflows. It provides integration layers and connectors to expose those systems to agent-based automation while emphasising secure and stable operation.
Verse is an AI platform that lets teams create autonomous AI “employees” from a single prompt without writing code. The agents can access connected customer tools, retain context across interactions, and run automated tasks continuously.
Weave is a meeting intelligence tool that listens to spoken conversations and generates a live, editable map of the thinking during a discussion (decisions, open questions, and who said what). Maps are saved, replayable, and shareable for review and asynchronous collaboration.
SonOf is a software delivery service that connects to code repositories and project-management tools to automate ticket creation and work estimation. It can deliver approved changes with senior engineering review, using a billing model in which customers pay only for work that ships.
Kit for AI is a service that provides a persistent memory and grounded-knowledge layer for AI agents. It exposes APIs and native memory tools that let agents call stored knowledge directly, supports ingesting files and URLs, and is intended to avoid building a separate retrieval-augmented generation stack. The site states the product works with any model and offers a free-to-start tier.
A browser-based AI 3D creation platform that generates 3D assets from text prompts or images. It provides material generation, retopology, rigging, and motion-capture features, with exports for game engines and 3D printing.
Agently is an AI Work OS that integrates with a company's software stack to build a centralized "company brain" and deploy autonomous agents to perform tasks and workflows. The platform describes an orchestration component named Jarvis that coordinates those agents across the stack.
YAGNI is a coding-agent platform for engineering teams, offered as a CLI and desktop client. Its agents are grounded in a company’s data and use a shared decision ledger, while per-developer budget controls govern usage. The platform routes requests across vetted, US-hosted open-weight models and supports developer skills, plugins, and custom components; it advertises lower per-call inference costs than frontier APIs.
Open Interpreter is a command-line coding agent optimized for open and low-cost models. It is a fork of OpenAI's Codex focused on emulating provider-specific agent harnesses; its Rust-native harnesses can be switched with `/harness`, while providers and models can be changed from its terminal interface. It supports native sandboxing on macOS, Linux, and Windows, MCP, skills, hooks, permissions, AGENTS.md instructions, and local configuration and session state under `~/.openinterpreter`. The agent can run as an Agent Client Protocol agent for compatible editors, speak the Codex exec protocol, and provide a one-line binary override for applications using the Codex SDK. Its built-in QA skill can test web applications through a real browser and native applications through trycua. The project is distributed for macOS, Linux, and Windows through shell or PowerShell installation commands and uses shared agent directories and protocols where available.
UI Skills is a command-line tool and collection of skills for design engineers working with AI coding agents. Its registry can be browsed from a terminal with commands such as `npx ui-skills start`, `categories`, and `list --category motion`, while individual skills such as `baseline-ui` can be fetched with `get`. It also exposes the registry to agents over the Model Context Protocol through `list_skills` and `get_skill`, and provides a playbook of distilled design-engineering lessons.
Graphify is a command-line tool and agent skill that turns a codebase and related project materials—including documentation, SQL schemas, configuration files, PDFs, images, and videos—into a locally queryable knowledge graph. It uses deterministic, local tree-sitter AST parsing for code, while non-code materials can receive a semantic pass from the configured AI assistant or API key; it does not use embeddings or a vector store. Graphify labels graph edges as EXTRACTED when they are explicit in the source or INFERRED when resolved by the tool, and can generate a clickable graph.html visualization, a GRAPH_REPORT.md summary, and a graph.json representation for later queries. The CLI is installed as the graphifyy package, registered with `graphify install`, and invoked through `/graphify` in supported AI coding assistants such as Claude Code, Cursor, Codex, Gemini CLI, and GitHub Copilot. The project is developed by Graphify Labs and provides free, fully local code mapping; its homepage describes a separate Graphify platform available through early access.
OpenMontage is an open-source, agent-driven video production system that turns plain-language instructions into video projects through an AI coding assistant. Its workflow covers research, scripting, production planning, asset generation, editing, and final composition, with production pipelines, approval gates, cost estimates, and post-render checks. The system can produce image-based videos as well as videos assembled from real motion footage: its agent can build a corpus from free stock footage and open archives, retrieve clips, edit them into a timeline, and render the result. It runs locally with FFmpeg and Remotion and includes 12 production pipelines, more than 100 tools, and hundreds of agent-skill and production-knowledge files.
codebase-memory-mcp is an MCP server from DeusData that indexes a codebase into a persistent local knowledge graph for AI coding agents. It parses source with tree-sitter AST grammars and uses Hybrid LSP semantic type resolution for selected languages to model functions, classes, call chains, HTTP routes, and cross-service links. Its MCP tools support code search, structural and Cypher-style queries, call tracing, architecture overviews, impact analysis, dead-code detection, and graph visualization. The project is distributed as native executables for macOS, Linux, and Windows, with no language runtime, hosted service, or API key required. Processing is local; the repository describes a RAM-first pipeline using compressed data, in-memory SQLite, and fused Aho-Corasick pattern matching, and reports sub-millisecond structural queries. A local visualization interface is available at localhost:9749.
Page Agent is a JavaScript in-page GUI agent that lets users control web interfaces with natural-language instructions. It operates through text-based DOM manipulation inside the webpage, without requiring a browser extension, Python, screenshots, multimodal models, or a headless browser. It can be integrated with a script tag or installed from NPM, and supports bring-your-own LLM configurations, including locally deployed models. An optional Chrome extension supports multi-page tasks, while a beta MCP server allows external agent clients to control the browser. The project is distributed under the MIT License and is designed for client-side web enhancement rather than server-side automation.
Cognee is an open-source AI memory platform for agents. It ingests data in varied formats and builds a self-hosted knowledge graph that combines vector embeddings, graph relationships, and ontology generation, allowing agents to search by meaning and connect related information across sessions. Its API provides remember, recall, forget, and improve operations, with session memory synchronized to the graph. The project provides Python, Rust, and TypeScript clients, a CLI and UI, MCP support, and integrations for Claude Code and OpenClaw.
omg.dev is an open-source, self-hosted parallel coding-agent harness and control plane developed by BennyKok. It runs coding-agent sessions on a local computer or hosted Computer, keeps them running when the web interface disconnects, and provides a single web UI for parallel sessions, transcript reading, follow-up instructions, chats, bots, schedules, and notifications. The project supports agents including Claude Code, Codex, Grok, Cursor, OpenCode, Jcode, GitHub Copilot, and Pi, using the user's existing agent subscriptions or API keys. Local installations run through a Bun CLI and expose the interface on localhost; remote phone access can be provided through Tailscale. The repository is licensed under the MIT License.
loop.js is a TypeScript framework and CLI for loop engineering: it runs an agent against a declared goal and verification criterion until a separate, skeptical Verify agent judges the work complete. Each round starts with fresh context and reads state and handoff notes from disk; the worker agent cannot approve its own work, and a failed verification supplies a reason for the next round. The framework provides round, spending, and per-round timeout guards; typed exits; crash recovery; idempotent resumption from a stored cursor; and a compare-and-set lock with heartbeat to prevent overlapping live runs and take over dead ones. It can run from a terminal, be embedded in a product, or be scheduled through crontab, launchd, Task Scheduler, or Modal's cloud.
Codex Orchestration is a Codex plugin that coordinates multiple models and role-based stages within a single task. It can assign optional Planner, Advisor, and Designer roles, plus a required Executor; the selected Codex model passes work between them, checks each result, and returns the final answer. The Planner creates and revises a plan, the Advisor reviews it and can return PLAN_APPROVED, the Designer produces an optional visual or UX handoff, and Executors implement the approved plan, potentially in parallel. The workflow stops after approval or after a safety limit of eight reviews, showing unresolved issues if approval is not reached. It is installed through the Codex plugin marketplace and requires Python 3.11 or newer.
self-learning-skills is a meta-skill for AI coding agents, including Claude Code, Cursor, Codex, and agents that read AGENTS.md or similar standing-instruction files. It recognizes hard-won, reusable procedures such as non-obvious commands and recurring operational workflows, then captures them as skills, rules, or project instructions for later sessions. The capture includes known dead ends and excludes one-off details and secret values. The project stores learned procedures as new skills/<name>/SKILL.md files for Claude Code, Codex, and Agent Skills clients; as .cursor/rules/learned/<name>.mdc files for Cursor; or in AGENTS.md or project notes for other agents. It can be installed with the community skills CLI, as a Claude Code plugin, or manually by copying the relevant files.
Three.js Game Skills is a set of self-contained Codex and Claude Code agent skills for building playable Three.js browser games. Its threejs-game-director skill routes work across gameplay, graphics, UI, asset generation, audio, debugging, and release verification, so users can request an outcome without manually selecting specialist skills. The package includes SKILL.md files, references, checklists, prompt templates, helper scripts, and a Vite, TypeScript, and Three.js scaffold. Generated games include deterministic test hooks and seeded randomness, with Playwright templates for smoke tests, visual-regression baselines, and bot playtests; the documented workflow also checks browser behavior, mobile viewports, visual output, UI, performance, and release readiness. It is created by Majid Manzarpour and can be installed for Codex or Claude Code through the skills CLI or the repository's installer.
Gigatoken is an open-source tokenizer for language-model training data, implemented in Rust and distributed as a Python package. It tokenizes large text corpora at gigabytes per second, supports common tokenizer formats across x86 and ARM hardware, and provides compatibility modes for Hugging Face Tokenizers and Tiktoken. Its native API can read text files directly and encode them with parallel processing, while the compatibility modes are designed as drop-in replacements but incur additional overhead. The repository reports throughput of up to roughly 1,000 times that of Hugging Face tokenizers in its benchmarks.
Palmier Pro is a Swift-native macOS video editor built for AI, with generative video and image models integrated into the timeline. It can expose a local MCP server at http://127.0.0.1:19789/mcp for connections from Claude, Codex, and Cursor, and includes an in-app agent that can create and edit within the same project. It requires macOS 26 (Tahoe) or later on Apple Silicon. Releases through v0.7.6 and the corresponding published source are available under GPLv3; later binary releases are proprietary.
deepsec is an open-source, agent-powered vulnerability scanner and security harness from Vercel Labs for on-demand review of large codebases within the user's own infrastructure. It identifies candidate vulnerabilities, uses AI models to investigate them, supports triage and optional revalidation to reduce false positives, and can run work across parallel worker machines. The CLI initializes a repository with `npx deepsec init`, stores state and findings in a `.deepsec/` directory, and resumes interrupted runs while skipping files already analyzed. Subsequent commands separate fast pattern scanning from AI processing and revalidation, and findings can be exported as Markdown directories or JSON. Model access can use Vercel AI Gateway, direct OpenAI or Anthropic credentials, or a custom HTTPS provider; scan cost and duration can be capped with command-line options.
MuScriptor is an open multi-instrument music transcription model developed by Kyutai and Mirelo. It converts audio recordings into note events, MIDI, or engraved sheet music with instrument information, and provides a locally hosted web UI, command-line interface, and Python package. Its architecture is a transformer decoder-only model; small, medium, and large variants are distributed through Hugging Face and downloaded and cached automatically after authentication. Sheet-music output uses MuseScore 4 or newer to generate MusicXML and PDF scores, including instrument-specific and tablature PDFs. The project can be run with uvx or installed from PyPI, and the model weights require acceptance of the CC BY-NC 4.0 license. The README notes that transcription works best with a steady tempo and performs significantly worse on rubato recordings.
LLM Space v4 is a local-first desktop workbench for prototyping and developing AI agents. It lets users version prompts, system messages, tools, and model settings; trace each model call and tool run in an agent loop; replay historical runs to debug failures; and evaluate agent performance across runs. It can generate prompts and tools with AI and turn a thread into a runnable LangGraph agent. Threads, project files, and API keys remain on the user's computer, while the project describes the app as cloud-ready for managed agents. The desktop application uses Electrobun with a React, Tailwind CSS, and shadcn/ui interface, and is built as a Bun monorepo around Pi Agent Core.
OpenAI Evals is an open-source framework for evaluating large language models and systems built with them, including tool-using agents and prompt chains. It includes a registry of benchmark evals, supports custom model-graded and private evals based on a user's data, and uses a completion-function protocol for advanced workflows. The package can be installed with pip and run locally with an OpenAI API key; the repository also documents optional result logging to Snowflake and configuration through the OpenAI Dashboard.