AI Manager is a local-first desktop application for installing, updating, configuring, and repairing AI coding command-line tools on a user's computer. It provides a single interface for checking installed tools, managing versions, connecting tools to reviewed or custom HTTPS AI-service endpoints, and viewing the effective endpoint and environment-variable overrides without editing shell profiles or PATH files. Its sidebar includes software management, API endpoints, skills, MCP connections, global prompts, sessions, and settings. It installs skills from trusted GitHub catalogues or local ZIP files, supports stdio, HTTP, and SSE MCP servers, and searches local coding-tool sessions that can be resumed in a terminal. Tool behavior is controlled by a capability registry, and the application uses local signals for checks rather than claiming a process is healthy without evidence. The application is built with Tauri 2 and React 18 and is released under the GNU AGPL-3.0-or-later. It has no analytics or telemetry SDK, stores product data locally, and offers signed update verification and platform-specific installers. Provider credentials are currently stored in plaintext in the local database, backups, and SQL configuration exports. AI Manager began as a fork of CC Switch, with inherited portions retaining their original MIT licensing.
A personal Codex skill and orchestration workflow for splitting software-development tasks between GPT-6 Astra and DeepSeek V4.1 Flash. Astra handles planning, architecture, high-stakes decisions, task briefs, and final review; Flash handles scoped repository discovery, implementation, testing, debugging, and routine verification, then returns a patch and evidence for Astra's acceptance pass. The workflow supports existing plans and phased implementation bundles, normally uses one writer, and includes verification, checkpointing, dry-run installation, backups, and guarded undo receipts. It is an early release requiring a Codex client with native subagents, Python 3.11 or newer, GPT-6 Astra as the root model, and an already configured Codex Router route for DeepSeek V4.1 Flash. It installs a skill, a named builder agent, and a scoped workflow policy without changing credentials, the root model, or Router configuration; no third-party Python dependencies are required.
Bespoke Nimble is an open-source model, training recipe, and local scoring library from Bespoke Labs for making typed decisions about text. Given a context and a flat schema of boolean or fixed-choice fields, it selects an allowed answer and returns candidate logits and probabilities without generating a reasoning explanation or free-form answer. Nimble fine-tunes Qwen3.5-9B with LoRA using contrastive training pairs: two nearly identical examples differ in one relevant fact, which changes the correct label. At inference, the scorer reads the logits for one-token candidate codes and converts them to probabilities with softmax; the Python library constructs the typed result rather than parsing generated JSON. The MLX ParallelScorer processes shared context once and scores fields in parallel, while the CUDA scorer scores each field separately. The repository includes the model recipe, curation and evaluation code, datasets, examples, tests, and deployment documentation. Bespoke-Nimble-9B can run locally on Apple Silicon or an NVIDIA BF16-capable GPU. It accepts text only, does not support nested schemas or answers outside the supplied choices, and its probabilities are not guarantees of correctness.
BongoCat is a native animated desktop companion that reacts to keyboard, pointer, and gamepad input. It is built in C and C++ with SDL3 and OpenGL, supports Windows, macOS, and Linux, and processes keyboard and mouse input locally rather than recording or uploading it. The runtime uses platform-specific input listeners, an atomic input queue, an SDL3 main loop, model-parameter updates, and OpenGL composition for the pet window and overlays. It includes built-in model assets and supports external model sources, including Mver packages, Tauri sources, `.model3.json` files, and image patches. Live2D rendering uses the Cubism SDK when available; without it, the project builds a diagnostic backend. The source code and native runtime are licensed under AGPL-3.0-only, while the default model mode and bundled model assets have separate MIT licensing.
ddc is a Rust command-line Android decompiler that converts DEX bytecode from APKs and related containers into readable Java. It supports multi-dex applications, XAPK/APKS/APKM packages, DEX versions 035–041, lambdas, string concatenations, and selected Android and Kotlin-specific output transformations. Its progressive-analysis interface provides subcommands for inspecting application metadata, manifests, resources, strings, cross-references, classes, methods, disassembly, and class hierarchies without requiring a complete decompilation first. The repository reports that its generated Java output passes a javac syntax check across its validation corpus and uses deadline-bounded processing for pathological classes. The CLI can be installed through Homebrew, Scoop, cargo-binstall, or built from Rust source. It provides bilingual English and Simplified Chinese messages, emits reproducible output without timestamps, and is released under the MIT license.
DocJev is an open-source Python library, CLI, and local app for classifying and splitting PDF, DOCX, and PPTX files with user-written natural-language rules. LiteParse extracts page text locally, then the hosted Jev decision engine assigns document categories or identifies boundaries between component documents; optional LlamaParse tiers provide cloud OCR for difficult inputs. Classification returns a category, probabilities, and review flags. Splitting returns ordered categories and contiguous page ranges, with optional export of each segment as a PDF and review reasons for uncertain boundaries. DOCX and PPTX processing requires LibreOffice, and inference is not offline because Jev is a hosted service. The project is distributed under the Apache-2.0 license.
Fast Browser Use is a local-first browser automation engine and agent skill for Claude Code, Codex, Cursor, and Python workflows. It runs Qwen3.5 models locally through MLX or PyTorch on Apple Silicon, CUDA GPUs, or CPUs, without cloud inference. The engine scans the rendered page for visible, interactable controls and converts them into bounded action candidates such as click, select, type, and done. It maps those candidates to single vocabulary tokens, scores them with one forward pass, and dispatches the selected browser action instead of generating CSS selectors, XPath expressions, or Playwright code. Text generation is used specifically for typing into fields, while structural actions use discrete logits scoring. The runtime also includes page-settling waits, visibility and occlusion checks, DOM-freshness checks, read-only protection, immutable execution traces, and external URL, title, and text assertions for outcome verification. The project provides a command-line interface, a Python API, and installable agent-skill integration, using Playwright Chromium for browser execution. It supports offline model operation after local weights are downloaded and is released under the MIT License.
FluidUse is an open-source macOS library and demo application for local computer use on Apple silicon. It reads forms in running Mac applications or browsers through the macOS Accessibility API, uses an on-device decision model to match visible fields with supplied profile data, and types or selects the resulting answers in the real application. A WebFormDriver supports embedded WKWebView forms. The project converts CUA-S1-FORMS to Core ML and runs it through FluidAudio; it also includes laya-based typed decision models for choice, score, and yes/no-style questions. DocumentEntities can turn PDF or text files containing label-and-value lines into a profile, while a predetermined-answer sheet handles question-style fields outside the model's decisions. The repository states that processing stays on the machine and that the tool does not read résumés, reason about dropdown options, write free text, upload files, or click Submit unless enabled. FluidUse is licensed under Apache 2.0 and requires Accessibility access for its launching terminal.
Flute is an open-source React tool by Web Prodigies for creating cinematic 3D scenes from an application's existing live UI. It integrates with Next.js and Vite projects, and provides a portable React wrapper for other browser-based hosts. Scenes use JSON definitions paired with React components that import the host application's real components and return Flute Surface elements. The shared studio handles camera movement, depth of field, playback, scene selection, snapshots, and deterministic video capture; depth of field uses a DOM-blur approximation rather than ray-traced bokeh. The renderer preserves the app's existing providers, styles, and components instead of duplicating the UI. Flute runs locally and is distributed as the `@webprodigies/flute` npm package under the MIT license. Its CLI supports scene initialization, synchronization, preview, snapshots, and MP4 export at 30, 60, or 120 FPS; Chromium and FFmpeg are required for export. It supports React DOM 18.2 or 19 and requires Node 22.12 or newer for the CLI. A React Native application without browser DOM is not supported.
Foremerge is an open-source coordination protocol and CLI for parallel coding agents, developed by naw103 and built above Git. It gives agents isolated Git worktrees while recording shared intents, semantic scopes, claims, dependencies, ChangeSets, decisions, validation results, and provenance in a local SQLite store under the repository's Git common directory. Agents publish what they intend to change before editing. Foremerge compares typed scopes such as symbols, APIs, schemas, configuration, infrastructure, tests, migrations, files, components, contracts, and domains, producing deterministic, explainable conflict advisories and related-work information. Claims are leased and advisory rather than exclusive: the system does not lock files, block agents, or ask a model to judge conflicts. Its workflow can publish and validate a ChangeSet against a recorded Git fingerprint, require a clean worktree and resolved high-severity conflicts for acceptance, and record the eventual integration commit without merging, rebasing, cherry-picking, or pushing code. The project provides a Rust CLI, JSON API, MCP server, SQLite coordination store, Git worktree support, and integrations for Claude Code, Codex, and Cursor. It is a local-first pre-1.0 MVP; coordination across machines is outside its scope, validation commands run with the local operating system's permissions, and the project reports no published benchmark results. Foremerge is distributed under the Apache License 2.0.
Glyd is a Rust lossless-compression library and command-line tool with a C ABI, streaming interfaces, and Python and Go bindings. It targets object storage, data lakes, logs, telemetry, backups, versioned exports, RPC payloads, and caches as an alternative to LZ4, Snappy, and zstd. Its record mode detects delimited records, SQL dumps, JSON lines, and varying-shape logs, converts fields or template variables into typed streams, and compresses them in parallel. Base mode compresses a new version against an older version by using regions of the old data as history. The format uses independently decodable blocks and parallel paths; higher levels use separated token, offset, length, and literal streams with SIMD-oriented decoding, while long-distance matching can find repeated content up to 128 MB back. The repository also provides shape dictionaries for small objects, packs for many small objects, container and JPEG handling at selected levels, and a cold context-mixing level for infrequently read data. The separate glyd-store crate and CLI select similar stored objects as delta bases using fingerprints, cap delta chains, and support local directories and S3-compatible backends. The codec, CLI, C ABI, and bindings use the BSD 3-Clause License or GPL-2.0; glyd-store uses the Business Source License 1.1, with commercial production use requiring a license and conversion to Apache-2.0 four years after each release.
Herdr GPUI is a native Rust/GPUI desktop client for a separately installed Herdr daemon. It displays the daemon's terminal sessions, split panes, workspaces, Git worktrees, and agent activity without running another terminal emulator or wrapping the TUI. The daemon owns terminal processes and session state. Herdr GPUI connects to its local or saved remote daemon through the Herdr client socket, receives binary protocol frames and surface patches, renders the terminal surfaces, and sends semantic input back; closing the GUI leaves the daemon and its sessions running. The project is distributed from the repository as a signed, notarized macOS app through Homebrew, with additional Linux packages, experimental Windows builds, Nix support, and source-build instructions. The daemon must be installed separately. It is licensed under Apache-2.0.
jeff is a self-hosted, MIT-licensed drop-in implementation of TypeSafe's jev System One API, powered by the GLiFormer encoder model. It accepts classification questions through a compatible SDK or the POST /v1/systemone endpoint and returns choice selections, ordered scores, and noul yes/no probabilities. The server supports CUDA, Apple MPS, CPU, and ONNX Runtime deployments, with batching, bearer-key authentication, rate limits, request limits, health and statistics endpoints, and deployment configurations for Modal. Its repository reports lower hosting costs than jev but weaker results on reasoning-heavy tasks.
Jevmind is a Python decision layer for software agents. It turns agent decisions into typed answers such as yes/no results, choices, and scores with confidence values, then applies a code-defined gate: uncertain answers escalate, trusted positive answers may act, and trusted negative answers are held. Its decisions are recorded in an append-only, hash-chained ledger, graded against later outcomes, and optionally recalibrated with Platt scaling. It includes nine skills for compacting tool output, navigating repositories, reviewing diffs, routing tasks to model tiers, checking completion claims, curating training data, walking linked markdown vaults, guarding commands, and running a structured-state control-loop demo. The default local brain uses readable rules, reflexes, and BM25 retrieval without a language model; an optional Jev HTTP brain can answer the same typed questions, while replay mode uses recorded answers. Jevmind runs as a command-line tool and MCP server, includes a Claude Code hook for shell-command screening, supports Python 3.10+, has no runtime dependencies, and is distributed under the MIT license.
JevRouter is a local-first agent capability router for models, subagents, skills, MCP tools, CLIs, and DSH plugins. It presents these capabilities as one candidate set and uses Jev for typed selection while enforcing availability, permissions, risk, and confirmation policies. It is decision-only by default and does not execute selected tools implicitly. The `route` operation answers one capability-selection question; `plan` selects capabilities for multiple steps using serial, batch, or decomposed strategies, with optional beam sequence search and contextual routing. Candidates can be supplied inline or discovered from Skill directories, MCP configurations, CLI help output, and DSH manifests. The router preserves Jev probabilities and provider responses, validates supplied permissions and input schemas, filters unavailable or disallowed candidates, and can select a safe fallback while retaining the original Jev choice. The CLI, Node.js SDK, HTTP server, and MCP adapter expose the routing contract. Decisions and plans are written as append-only receipts with provenance hashes, and a local dashboard reads those files without making provider requests or uploading data. It supports the official Jev API and OpenRouter, has an offline demo provider, requires Node.js 20 or later, and is distributed under the MIT license.
jev-voice-browser is a Node application that controls a headed Chromium browser by voice or typed commands. It streams partial speech transcripts from the browser's Web Speech API to a Node server, which uses TypeSafe's Jev model to classify the intent, select a page element or site, extract candidate text and URL spans, determine whether the command is complete or addressed to the browser, and assess whether an action is destructive. Playwright then executes navigation, searches, typing, clicking, scrolling, history, and tab actions. The application sends the page URL, title, detected site, a compact snapshot of visible elements, and recent actions to Jev for each transcript update. Code—not the model—constructs URLs, copies text verbatim, and applies policy thresholds to decide whether to act, wait, ignore, request confirmation, or show numbered alternatives. Recent action context supports corrections such as reversing an action, rejecting a selected result, or choosing another candidate. It requires Node.js, Chrome or Edge for microphone input, Playwright's Chromium installation, and a TypeSafe API key. The server can launch a persistent-profile Chromium window or attach to an existing browser through CDP; it also supports headless operation and typed commands when no microphone is available.
Laya is an open-source, Jev-compatible System-1 decision model package by Convai Innovations for running typed decision questions from Node.js and TypeScript through ONNX Runtime. It accepts a state such as a ticket, email, or JSON object and returns calibrated answers for choice questions with per-option probabilities, ordered score questions with a probability distribution and expected level, or yes/no statements with a probability of true. Questions in one call are batched into a single model run, and the package uses the same request and response shape as the Python reference implementation without requiring Python or PyTorch at runtime. The package downloads and caches ONNX model weights from Hugging Face by default, or can load a local ONNX bundle; it requires Node.js 20 or newer. The JavaScript package is MIT-licensed, while the Laya model weights are published under Apache 2.0.
Laya-CoreML is an open-weight on-device decision-model runtime and port of Laya for Apple Silicon, using Core ML and Apple's Neural Engine. It answers typed choice, ordinal-score, and boolean questions with probabilities or scores rather than generating text or JSON; its Python API loads a model and returns structured answers. The repository includes Core ML checkpoints, tokenizer and configuration files, an offline Python package, a command-line Snake demo, and conversion and benchmarking tools. The Neural Engine bundle uses a short fixed-capacity input path, while longer or general-purpose requests can use CPU-and-GPU checkpoints. The project documents a separate Neural Engine graph using BC1L activations, 1×1 projections, and per-head attention, with CPU handling input and output boundaries. It supports macOS 15 or later on Apple Silicon and can run without PyTorch, Transformers, or MLX at inference time. The project is licensed under Apache-2.0 and describes itself as an independent port by Convai Innovations and contributors, not an official Convai Innovations or Apple release.
magpie is a cross-platform menu-bar and command-line application for selecting models used by local coding agents. It detects configured agents, lets users change their provider, model, and related settings, edits only the changed keys in configuration files while preserving formatting, and saves profiles for switching groups of agent configurations. It also runs a local gateway at 127.0.0.1:3425 that exposes OpenAI Chat Completions, OpenAI Responses, Anthropic Messages, and Google Gemini-compatible endpoints. The gateway translates requests between these APIs, forwards them to configured vendors or local servers, and presents provider models through a shared provider/model catalog; signed-in Claude Code, Codex, and Copilot accounts can also be used as providers. A terminal UI, plain CLI, provider management, model-list refreshing, import links, and encrypted backup and restore are included. The application is distributed for macOS, Linux, and Windows, with a smaller terminal-only build, and is licensed under the MIT License. Its official homepage is usemagpie.ai and its source repository is maintained at yetone/magpie.
mini-AGI is an experimental continual-learning, byte-level language model that trains from scratch on a CUDA-capable GPU with at least 8 GB of VRAM. It reads 256 byte values directly rather than using a conventional tokenizer, processes data one chunk at a time, and uses the same forward path for training and generation. Its architecture combines dense prelude blocks with a recurrent block applied up to 24 times, adaptive PonderNet-style halting, and per-application top-8 routing through a dynamically growing and pruning expert pool. Expert weights and Adam moments are stored as files on disk; a working set is paged through RAM and VRAM as needed, allowing the pool size to exceed available VRAM. The project includes Python commands for building corpora, reading files, streaming training, serving a local web interface, and inspecting or replicating continual-learning experiments. The repository describes the current model as toy-level rather than frontier-capable, and its weights are generated in the local weights directory rather than distributed with the source.
Mr. Mak Workspace is a local Windows desktop workspace for Codex CLI and Claude Code CLI. It combines real CLI chats with a project workspace for folders, research, images, files, Markdown documents, project reports, and saved project materials. The application provides chat history and tabs, agent activity indicators, file and folder insertion, project cards, previews for media files, and a right rail for project skills, knowledge, workflows, MCP connections, settings, and help. The repository includes fourteen project skills covering areas such as planning, handoffs, image references, media generation, Three.js, Blender, video inspection, and dictation. It can also provide an optional voice coordinator using Codex CLI and an OpenAI API key; external providers and MCP connections are optional and use the user's own accounts. The Windows installer includes the local Node service but does not bundle the supported CLIs or provider accounts. The native application targets Windows x64, while browser report previews can run elsewhere. Application code and original workflow documentation are MIT-licensed, with third-party skill materials retaining their original licenses.
Null Motion is a local browser-based motion-design preview and export tool that presents a finished advertisement alongside the black-and-white HyperFrames drafts used to plan it. It includes an earlier launch editor and ships with example references and authored draft scenes. The project uses FFmpeg scene detection to divide reference videos into sections while preserving their cuts. Each section is represented as a 640×360 HTML scene on a paused GSAP timeline, with beats such as text, logos, phones, windows, chats, cards, lists, charts, and clouds. During playback, the active draft is driven by the film’s clock and stays synchronized frame by frame with the reference video. For export, a hidden video plays at half speed; requestVideoFrameCallback captures decoded frames, the drafts are rendered at each frame’s timestamp, WebCodecs encodes the video, and mp4-muxer produces an MP4 with copied reference audio. The tool runs with Node.js 22 or newer and requires Chrome or Edge for WebCodecs export. Reference media remains local, and the repository states that the drafts are authored rather than AI-generated. No project-wide open-source license has been selected.
OpenMuse is a personal-agent application built by CopilotKit for iOS, Android, and web. It combines a persistent Chromium browser with an optional isolated Linux container, file and PDF handling, visible task plans, action reviews, and browser or terminal takeover so a person can inspect or continue the agent's work. It also provides Gmail and Google Calendar adapters with review required for sends and calendar changes. The application runs an API, durable task worker, and browser worker. Tasks use stored plans, progress, approvals, pauses, retries, and SQL leases; the browser worker maintains persistent Chromium profiles, while the optional non-root Linux container provides a retained workspace volume with bounded commands and no host-directory mounts or credentials. The interface is built with CopilotKit headless chat and AG-UI events, with inline email, browser, PDF, plan, and finance results. The repository describes OpenMuse as an alpha for self-hosting and building on, with model and Google-account configuration required for live agent and mail/calendar workflows. It is MIT licensed; CopilotKit Intelligence is a separate required service for rich conversation persistence and is not included under the repository's license.
reflex is an open-source local decision model for classifying tickets, documents, photos, and other state against typed questions with fixed answer choices. It runs the state through an open-weights model in one forward pass, branches each question independently, reads only the answer-label scores rather than generating free text, and returns probabilities for yes/no, choice, or ordered-score questions. It can read each question in multiple option orders and average the results to reduce position bias; optional temperature calibration adjusts confidence estimates for a target workload. The project provides a Python engine, an HTTP endpoint at /v1/systemone, GPU and Apple Silicon execution, an SGLang backend, a WebGPU browser demonstration, and optional LoRA training and distillation tools. The default serving configuration uses a frozen Qwen checkpoint, and the repository is licensed under MIT; model weights have their own licenses.
Search is a small, fast WebKit browser for macOS developed by Office Commun. It provides a single address-and-search field, tabs, reading mode, picture-in-picture video, bookmarks, history, downloads, built-in ad and tracker blocking, per-site clutter hiding, and password storage in the macOS keychain. It has no account, sync, cloud storage, telemetry, or analytics; browsing data remains on the Mac. The browser uses WebKit and lazily creates each tab's web view, while its ad blocker runs as a compiled WKContentRuleList before network requests are made and its hidden-element rules are injected at document start. It can run Chrome extensions through WebKit's extension engine, adding shims for Chrome APIs that WebKit lacks; extension support requires macOS 15.4 or later. Search supports macOS 14 or later, is distributed as a free download of about 3 MB or through Homebrew, and is licensed under the MIT License. The repository contains a Swift and SwiftUI/AppKit implementation with no dependencies beyond Apple's macOS frameworks.
Shapeshift is a browser-based natural-language interface that turns one text input into structured cards such as events, checklists, timers, color pickers, bill splitters, converters, polls, contacts, and notes. Its intent layer can use TypeSafe AI's Jev model to answer typed questions in parallel, while deterministic parsers handle dates, amounts, units, and calculations; a built-in keyword classifier provides offline operation without an account. A state machine stabilizes changing classifications as text is entered, and saved cards remain in the browser's local storage until deleted. The open-source project is built with Next.js, React, TypeScript, and Tailwind CSS and is licensed under MIT.
ThinkingOrbs is a SwiftUI package by Haplo LLC that provides dotted 3D loading indicators for AI and agent interfaces. It includes nine hand-tuned designs representing states such as searching, solving, listening, connecting, composing, and waiting; eight use rotated, depth-shaded, z-sorted 3D forms and one uses a morphing outline. The package provides regular and small sizes, status-label views with optional shimmering text, localization and accessibility labels, speed and pause controls, and public frame geometry for custom renderers. Each orb is drawn with grayscale dots in a Canvas inside a TimelineView. Timelines pause when an orb is offscreen or the app is backgrounded, while Reduce Motion displays a representative static frame. The package supports iOS 17+, macOS 14+, tvOS 17+, watchOS 10+, and visionOS 1+, and is distributed as a Swift Package under the MIT license.
Unreal Agent is an async-first agent harness from Unreal Labs. Its library and command-line executables coordinate persisted agent sessions, LLM turns, tool calls, and asynchronous operations. Inputs carry caller-supplied globally unique IDs for deduplication; sessions store append-only history that can be recovered or forked; and tool translators validate model-produced calls and convert them into serializable operations whose execution state is tracked separately. The harness includes a session inbox, coordinator, session store, context builder, LLM adapter, tool registry, tool translators, and an operation manager. Session and operation data are versioned and serializable, allowing operations to be dispatched to alternative or remote runtimes; unsupported session versions produce an explicit resume error. The repository contains the harness library, executables, and benchmark runners.
Z.ai's coding agent harness and AI programming workspace, offered through a desktop application, browser interface, and terminal agent. Its repository contains the client, backend services, shared React UI, Agent CLI, and runtime; the desktop client uses Electron, while the web and server components communicate through HTTP and WebSocket services. It supports remote project connections over SSH and WSL. The standalone command-line distribution combines a terminal UI, web server, and agent, and can run locally without Electron; the same command starts either the terminal interface or a browser-accessible web interface.
Searchable transcript of GitHub Trending Weekly #50: openmuse, open-glean, unreal-agent, BongoCat, laya-mlx, mini-AGI, flute — Github Awesome (15:11). Search for a phrase, then click its timestamp to jump straight to that moment in the video.
Captions sourced from the original video on YouTube, published by Github Awesome. The video, its captions and all related intellectual property remain the property of their respective owners; AINotes claims no ownership. Provided for research, accessibility and search — see the Transcript Notice and Copyright Policy.
00:00 Welcome back to GitHub awesome. This is GitHub trending weekly number 50. 35 trending open- source projects on GitHub right now. Let's go. Open Muse is a personal AI agent from Copilot Kit that does its work where you can see it. It gets a persistent browser, an optional Linux terminal in an isolated container, and read access to your Gmail and Google Calendar.
00:22 And you follow the plan as it runs. If it's heading the wrong way, you take over the browser or terminal mid task. It runs on iOS, Android, and the web from one expo app. Search is a Mac browser built to stay out of your way. It uses the web kit already on your machine, keeping the app around 3 megabytes. You get tabs, one field addresses, an ad blocker, and a way to hide clutter that keeps coming back.
00:50 Passwords live in the Mac OS keychain. There's no account or sync, so your browsing history stays on that Mac. Open Glean searches your files, notes, and connected apps through HydroDB, then uses your chosen model to write answers with source citations. Its deep research mode splits a question into smaller searches, runs independent branches in parallel, and combines their findings.
01:17 You can host the interface yourself and choose an OpenAI compatible model endpoint. If an agent receives the same event twice, you don't want it doing the job twice. Unreal agent is an async firstgo harness built around that problem. It dduplicates inputs, persists sessions and operation state, and supports recovery and forks. Tool calls become serializable operations that execute outside the coordinator's event loop.
01:44 That separation matters when a run fails halfway through. There's a record to resume from. Zcode puts the same coding agent runtime behind a desktop app, browser workspace, and terminal interface. You get a desktop app, a web workbench, and a terminal agent built from the same codebase. So, the interface you're in doesn't change what the agent knows.
02:04 It works over SSH and WSL for remote projects. And the CLI piece runs standalone if you don't want Electron near your terminal. Bongo Cat puts an animated cat on your desktop that reacts to your keyboard and mouse. Completely unnecessary, pretty charming. This native version uses C, SDL3, and OpenGL with support for custom models. It runs on Windows, Mac OS, and Linux, and the maintainers say your input stays local rather than being recorded or uploaded.
02:35 Your desktop gets a tiny co-orker who takes typing far more seriously than you do. LIA MLX handles small decisions that don't need a chatbot response. Give it a support message and candidate departments and it returns probabilities for each choice without generating text. This MLX port runs locally on Apple silicon without PyTorch or a cloud API. On an M3 Max, the maintainers report a 13.4 mm's median for a short English decision.
03:06 Mini AGI is a bite level language model you train from scratch on an 8 gigabyte GPU. It reads raw bytes with no tokenizer and pages its experts in from disk 32 at a time. It keeps learning as you feed it new data and the author reports 99.84% retention of what it learned earlier. The author calls it a toy model despite the name. Continual learning on a laptop GPU is the interesting part.
03:33 LiaML brings LIA's typed decisions to Apple's neural engine. Give it a short question and it returns probabilities for choices or scores without generating a paragraph. On an M3 Max, the maintainer reports a 4.98 millisecond median for one multilingual decision and a 2.78fold improvement in system energy per decision over compiled MLX. Bespoke Nimble is a Quen 3.5-9B fine-tune that answers typed questions about text with no reasoning text.
04:02 You hand it a schema of multiple choice or yes or no fields and it scores only the candidate answers returning JSON with a probability for each option. On their 324 example test set, the team reports 90.12% agreement, up from 66.36 for the base model. VMO's Swift SDK lets iOS and Mac OS apps search for moments inside video. Upload a clip, then look for a scene, spoken words, or text that appeared on screen.
04:35 It handles streamed uploads, progress, cancellation, and typed search results with timestamps so your app can jump to the matching moment. VMO's backend does the indexing and search, and you'll need an API key. Magpie is a menu bar app that picks the model for all your coding agents from one place. It covers nine of them, including claude code, codecs, Gemini CLI, and cursor, and rewrites each config file in place without wrecking your comments or indentation.
05:01 A local gateway speaks open AIS and Anthropics APIs and forwards to whichever vendor you chose. Profiles snapshot every agent's settings, so you can switch the whole setup at once. Null Motion plays a finished ad film at full width and runs black and white motion drafts underneath it, frame synced, so you can see how each section was built. FFmpeg scene detection splits splits the reference to three to eight sections and each draft is HTML and GSAP on a pause timeline driven by the film's clock.
05:34 Export renders the whole thing to MP4 in the browser with web codecs. It's a tear down view for motion design and it runs entirely local. Shapeshift starts with one text box, then turns what you type into the right little tool. Dinner with Priya Friday becomes an event card. A shopping list becomes a checklist. Split 2400 between three calculates each share.
05:58 Jev can identify the intent while ordinary code handles the dates and math. It also works offline with a built-in classifier and saved cards stay in your browser. Rizo Windows is a collection of procedural films, each living in a single HTML file. One follows a night train, another turns a flock of starings into moving Rizograph ink. Canvas draws the frames.
06:23 Web audio handles the sound and the repo includes clawed code skills and rendering tools for making your own. Every frame depends on time alone, so you can jump to any moment and inspect it exactly. Flute turns the React components you already have into cinematic 3D product shots. It installs into a Nex.js or Vitey app and your coding agent writes JSON scenes that point at your real UI with camera moves and depth of field layered on top.
06:52 A built-in studio previews each scene and exports MP4 at up to 120 frames per second. Launch videos built from your actual interface instead of screenshots are the part that stands out. ARC CUA lets a computer use agent hand off the clickby-click part of a task. A planner sets the goal, allowed text, constraints, and success check. JV chooses actions from controls the Mac desktop exposes.
07:19 It uses accessibility data and local OCR, then waits for the UI to settle before the next move. If it gets stuck or hits its action budget, it hands control back to the planner. Former merge catches the merge conflict before anyone writes the code. When parallel coding agents publish what they plan to touch, it compares those intents. So if one agent plans to replace payment service while another plans to extend it, you hear about it now instead of at merge time.
07:47 A verification gate runs your tests on a clean commit before anything lands. Browser agents sometimes try to click buttons that aren't there. Fast browser use scans the controls actually visible on a page, then asks a local model to choose from those options. It can navigate sites and fill forms as a skill for codeex or clawed code and leaves a trace you can check afterward.
08:12 It runs on your machine, though the model needs substantial memory. Fluid use a surprisingly small AI model for a familiar chore, filling out forms. On a Mac, it reads fields in a real app or browser, matches them to details you provide, and types the answers in. The decisions run on your device, and the repo reports about one millisecond per field.
08:36 You handle anything that needs judgment, like essays, uploads, and clicking submit. Mr. Mac Workspace is a Windows desktop app that puts CodeCli or Claude code in one window and your project files in the other. The right side has folders for your inbox, projects, knowledge, and processes with a markdown editor and preview built in. You can pin chats, color code tabs, and see which agent is still working.
09:01 It ships with 14 project skills, including Blender Workflows. AI manager is a desktop app that keeps your AI coding CLI in one place, forked from CC Switch. It detects 10 tools including Claude Code, Codeex, Gemini CLI, and Open Code, and installs, updates, or repairs them across npm, PNPM, Bun, or Volta. From the same window, you point each tool at an API endpoint and test it, add MCP servers, install skills, and search old sessions to pick them back up.
09:31 Astra Flash Orchestrator splits a coding task between two models. Astra plans the work and reviews the finished patch. A DeepSeek Flash sub agent handles implementation, tests, and routine fixes. The idea is to spend Astra's attention on decisions while Flash does the longer build loop. It includes phase task briefs, verification steps, and reversible setup.
09:56 Jevouter decides which tool your agent should reach for and stops there. It gathers models, sub aents, skills, MCP tools, and CLIs into one candidate list. Asks Typesafe's Jev a type choice question, then filters the answer through permission, availability, and risk checks. It never runs anything itself, and every decision goes into an appendon log.
10:19 ISIS 1 skills pack is three clawed skills for explaining work, and they pick between themselves based on what you want to end up holding. Eli 5 succinct gives a short plain language answer in chat and never writes a file. The summary skill writes one self-contained HTML page with modes you can stack accessible for ADHD or dyslexia, artistic and animated.
10:40 It reads the actual files before it summarizes instead of working from memory. Thinking orbs is a swift UI port of Jako Benalik's thinking orbs animated loading indicators for AI interfaces. You drop in one line like thinking orb searching and get a dotted 3D sphere for that agent state with nine to choose from. They pause offcreen and respect reduce motion.
11:06 The port was checked against the original TypeScript engine across 725,49 frames. Jev voice browser lets you drive a Chromium window by talking to it and it often acts before you finish the sentence. Partial transcripts go to type safeaf's jev with about a dozen typed questions and it picks the intent and the target instead of generating text. So URLs and search terms come through verbatim.
11:31 Say no the other one and it corrects itself. The author reports about 300 milliseconds from your last word to a decision. Reflex is a local decision model for jobs where you need a choice not a paragraph. Give it a support ticket, document, or photo and several fixed answer questions. It returns probabilities in one model pass. It asks each question with the options in two different orders to reduce position bias.
12:00 There's a Jev compatible API, but you should calibrate the confidence numbers on your own data before using them as thresholds. This lia package runs conveys decision model from Node.js and TypeScript with no Python and no PyTorch. It uses ONX runtime and the author says outputs match the Python reference to four decimal places. You pass a state like a ticket or an email plus typed questions and get back per option probabilities, a rubric score or a calibrated yes or no.
12:33 A single PDF can contain several documents and sorting them by hand gets old fast. DocJev takes PDFs, word files, or slide decks, classifies them using rules you write, and splits combined packets into page ranges. It flags uncertain boundaries for review and can export the pieces as PDFs. Text extraction runs locally by default, but the extracted text goes to the hosted Jev service for decisions.
12:59 Jevmin pulls an agent's decisions out of its text and treats them as data. Every call becomes a yes or no with confidence, a pick from a list, or a score on a scale. And code-based rules decide whether the agent gets to act on it. Each decision lands in a hashchained ledger before the outcome is known, gets graded afterward, and feeds back into calibration.
13:21 Logging the guess before the result is the part worth copying. DDC is an Android decompiler written in Rust, going after the job Jadeex does. The author reports a 226 megabyte APK decompiled in 5 seconds and says all 99,689 output files across seven real APKs parse cleanly under Java. You don't always need a full decompile either. Over 20 subcomands answer questions like strings, cross references, and class hierarchies in milliseconds.
13:53 Jeff is a self-hosted standin for Typesafe's Jev API. Point the official SDK at your own server and it answers choice score and yes or no questions using a 400 million parameter gllyformer model. The maintainer estimates about 2.6 per million single question requests on an L4 against 15.6 for Jev. Their benchmark shows the trade-off. Jeff trails Jev on harder reasoning tasks.
14:19 Herder GPUi gives Herder a native Mac window for its terminal sessions, split panes, workspaces, and agent activity. It connects to a Herder Damon you install separately, then draws the terminal itself using Rust and GPUi. The useful bit is that the Damon owns the sessions. Close the window and your work keeps running until you attach again. Glide is a Rust lossless compressor built for logs, data dumps, and version backups.
14:46 Its record mode groups fields from structured data before compressing them. Its base mode stores a new version against an older one, while a companion store can find that older match automatically. The appeal is smaller files and fast parallel reads. The trade-off is slower compression, especially when record mode has to parse the data first.