BridgeClip is an open-source AI desktop app from BridgeMind that turns local videos or podcast, YouTube, and Twitch VOD links into captioned short-form clips. It downloads the source when needed, transcribes speech, uses OpenRouter to identify promising moments and plan the clips, reframes them to 9:16 or 16:9, and renders the results locally with FFmpeg and word-by-word captions. It supports smart face and screen-share framing, multiple caption styles, clip libraries and job tracking, and optional publishing or scheduling through Zernio. The app runs on the user's computer without a BridgeMind backend; provider keys and transcription or planning data are sent to the selected OpenRouter services, while rendering is local. macOS and Windows releases are provided, Linux development is supported, and the project is MIT licensed.
changelog.earth is an environmental news project that presents selected reporting as planetary patch notes. It collects stories from RSS feeds, GDELT, and Spaceflight News, validates and deduplicates them, then uses Groq to select stories and write game-style patch titles while retaining the original headline, publisher summary, publication details, and source link. Scheduled archive jobs save editions that the homepage, API, and RSS feed serve; the RSS feed provides the latest published updates with stable IDs, and saved editions remain available when collection or upstream services fail. The project is an MIT-licensed open-source application built with React, TypeScript, Tailwind CSS, and Next.js, with deployment support for Vercel and Cloudflare.
cli-faq-shortcuts is an agent skill that analyzes a developer's coding-agent history and turns repeated requests into project-specific slash-command shortcuts for Claude Code, Codex, and Cursor. Its local extractor reads prompts from the relevant Claude Code, Codex, and Cursor history files, excluding subagent turns and scripted runs. The agent groups differently worded prompts by intent, counts recurring asks, checks existing shortcuts and project tools, and presents candidates for the user to select. Each selected shortcut is written as a SKILL.md file under the project’s .claude/skills directory, linked through .agents/skills so the supported agents can load it; the repository also records mappings in CLAUDE.md or AGENTS.md. Shortcuts can point to existing scripts and queries, include project-specific cautions, and use a dry run before actions that write or send data. The extractor can be run independently with Python and reads local files without making network calls. The repository supports installation on macOS, Linux, and Windows, and is licensed under Apache-2.0.
CogSend is a self-hosted social media scheduler that runs on the user's Cloudflare account. It lets users write a draft, customize it per platform, add images and alt text, automatically split long drafts into threads, and publish immediately or schedule posts for Mastodon, Bluesky, LinkedIn, Threads, and X. The scheduler reports delivery results per account, retries temporary failures, and supports cancellation, rescheduling, and manual retries; it also provides link previews and publication-failure insights. The application stores data in Cloudflare D1 and media in R2, sends posts through the user's API credentials, encrypts credentials, and protects the administrator account with two-factor authentication. It provides a personal API key and an MCP server for scripts, Shortcuts, Claude Code, Codex, or other MCP clients to draft, schedule, and publish posts. It is built with SvelteKit on Cloudflare Workers, uses Drizzle with D1 SQLite, and is licensed under the MIT License.
Company Brain is an open-source Slack teammate developed by supermemory. It remembers decisions, projects, owners, and other context from the Slack channels and connected tools it can access, answers questions from that information, and can act through tools such as GitHub, Linear, Notion, and Google Workspace via MCP. It can open issues, read pull requests, run repository scripts in a sandbox, return charts, CSVs, and PDFs in Slack, speak up when relevant, and send scheduled digests or perform research. Its memory is organized as a permissions graph: public-channel queries use the shared organizational memory, private-channel queries add that channel's memory, and direct messages can use the asker's personal memory and private channels. Tool writes run through the user's own connection, while access requiring another teammate's connection requires that teammate's approval. The project runs on a user's Cloudflare account using Workers, Durable Objects, D1, KV, and supermemory for memory; it can be deployed through a setup flow or run locally. The repository states that the formerly paid product was discontinued and released under the Apache 2.0 license.
CLM is an AI decision-making model and Python service that selects among candidate actions by embedding a state and each action separately, then scoring their alignment instead of generating a textual response. It is trained with bidirectional InfoNCE: matched state-action pairs are pulled together while incorrect or hard-negative actions are pushed apart. At deployment, cached state and action embeddings are scored with a projection head, enabling typed decisions such as binary judgments, choices, and ordered scores, as well as ranking candidate answers, tools, trajectories, or next moves. The repository provides CLM-8B through a TypeSafe-compatible HTTP API and a local Python engine. It runs a Qwen3-8B pooling encoder through vLLM, exposes endpoints including `/v1/systemone` and `/v1/rank`, and includes a browser playground, embedding and action caches, checkpoint hot-reloading, and scripts for fine-tuning projection heads on typed decisions or agent trajectories. The reference head can run on CPU or GPU, while the encoder is served separately; states longer than 2,048 tokens are truncated by default unless the limits are raised together. The project is released under the Apache 2.0 License, and the repository's CLM-8B weights are also released under Apache 2.0 on Hugging Face.
disktree is a disk-usage treemap for finding and removing files and directories that occupy space, developed in Rust with GPUI. It scans a home directory, selected directory, mounted volume, or whole disk; draws nested blocks sized by actual disk usage; and supports navigation by size, file count, or age, filtering, hidden-file inclusion, apparent-size measurement, and classification of data such as caches, build output, media, and repositories. Users can mark items, review the complete selection, estimate the resulting free space, and move selections to trash or delete them permanently with confirmation and removal safeguards. The scan accounts for hardlinks, avoids following symlinks by default, and aggregates directory sizes in parallel. It runs on Linux, macOS, and Windows, with platform-specific disk-usage and trash behavior, and is licensed under the MIT License.
Flux is an open-source device-integration tool that connects an Omarchy desktop to an Android phone or Mac over a local network, or through Tailscale when away from home. It transfers files, clipboard text, and links; reads phone notifications; sends SMS; controls media; synchronizes Do Not Disturb; runs desktop commands; and uses the phone as a camera or microphone. It can also mirror the phone screen, support fingerprint-based sudo approval, and expose coding-agent output on the phone. The project includes a command-line interface, background daemon, native Qt desktop window, Omarchy shell plugin, Android app, and macOS app. The desktop initiates network connections, so the default Omarchy firewall does not require an additional inbound rule. Flux for Android requires Android 10 or later, while the macOS app requires macOS 14 or later.
HaloBattery is a Windows system-tray utility that displays battery levels for wireless peripherals, including devices connected through 2.4 GHz receivers, USB, and Bluetooth. It gives each device a battery-ring icon with a device pictogram, low-battery and charging states, charging animation, notifications, and an exact percentage on hover. The program reads vendor-specific HID reports and standard controller or headset reports for supported mice, headsets, keyboards, and game controllers, while Bluetooth devices use the battery level reported by Windows. It combines receiver and cable readings into one device icon where supported, retains the last value for sleeping devices, and provides diagnostics that list HID devices and raw protocol replies for troubleshooting and adding device support. The repository distributes a ready-made Windows executable as a release ZIP or supports running from Python 3.10+ source. It requires no Python installation when using the packaged executable, is built through GitHub Actions, stores settings and diagnostics under the user's application-data directory, and is licensed under MIT.
Hindsight is an agent memory system developed by Vectorize.io for storing, retrieving, and synthesizing long-term memories for AI agents. It organizes memories into world facts, experiences, observations, and mental models within isolated memory banks, with support for multilingual data and optional per-bank scanning for secrets and personally identifiable information. Its retain operation uses an LLM to extract facts, entities, relationships, and temporal data, then normalizes them into searchable representations. Recall combines semantic vector search, BM25 keyword matching, entity/temporal/causal graph links, and time-range filtering, merging results with reciprocal-rank fusion and cross-encoder reranking. Reflect performs deeper analysis of stored memories to form connections and answer questions requiring more than retrieval. Mental models and knowledge pages provide continuously updated answers or documents derived from a bank's memories. Hindsight can run as a Docker service, Python package, embedded database, Kubernetes deployment, or hosted Hindsight Cloud service. It provides Python, Node.js, Go, CLI, and REST clients, an OpenAI-compatible LLM wrapper, integrations for agent frameworks and coding agents, and a built-in Model Context Protocol server. Self-hosted storage uses PostgreSQL with pgvector or Oracle AI Database 23ai; the project is MIT-licensed.
INKWAVE is an open-source browser-based 4v4 turf-war shooter built with three.js. Players paint arena surfaces, swim through their team's ink to move quickly and refill their tank, use weapons and special abilities, and compete against bots across three stages; the team with the most painted ground after three minutes wins. It includes seven weapon types, squid movement, map viewing with Super Jumps, character customization, and keyboard, mouse, and gamepad controls. Ink is painted into texture-space regions on a GPU-managed 4K atlas, while a coarse CPU grid tracks turf scores and gameplay queries. Shaders add height, gloss, wetness, drying, and wall-drip effects. Stages are defined as data and mirrored by 180 degrees for symmetrical team layouts; characters use procedurally generated geometry, materials, a 60-bone rig, and code-driven animations. Weapons, actors, match logic, effects, HUD, and audio communicate through typed events, and deterministic freeze-and-step tooling supports reproducible bot simulations and filmstrip generation. The game has no build step for local development and can run from a static file server; its included server starts the game at localhost:8490. Chrome and Edge are the target browsers, with Firefox supported and Safari described as slower. The repository is licensed under MIT and states that INKWAVE is an independent project unaffiliated with Nintendo.
Interference Search is a research project and Python library for reasoning over explicit states rather than a single linear language-model transcript. It expands live branches in parallel, lets the environment execute their moves, merges branches that reach the same state, uses a trained judge to discard states unlikely to reach the goal, and advances the surviving frontier one level at a time. The method is classical and does not claim quantum speedup. The repository implements Countdown arithmetic search and program search, including move generation, exact solving, judge training, sandboxed program tests, behavioral merging, benchmarks, raw results, failed experiments, a research log, and the accompanying paper. Its reported experiments compare the approach with linear language-model reasoning and other search strategies; it also includes a one-line baseline and tools for reproducing the Countdown results. The package requires Python 3.10 or newer and runs with PyTorch. Language-model experiments use MLX on Apple silicon, while the core search, judge, and tests can run on other PyTorch-supported systems. It is distributed under the Apache 2.0 license.
Jevgrep is a command-line tool for coding agents that finds relevant files and source context in unfamiliar repositories from natural-language questions. It uses Jev to judge relevance across folders, files, and declarations, explores qualifying branches of the repository hierarchy, selects files through content previews, identifies useful source units and surrounding context, and returns a summary followed by file locations, reading leads, and verbatim excerpts with line references. Python and TypeScript/JavaScript support declaration parsing, while other text uses a fallback. It also includes an agent skill that teaches supported coding agents when to invoke the CLI and how to use its results. The CLI is distributed through npm as @dzhng/jevgrep, requires Node.js 22 or later on macOS or Linux, and requires authentication with a supported AI provider. Searches send eligible source content to Jev through the selected provider; local filtering excludes ignored, hidden, dependency/build, binary, and obvious credential files but is not a guarantee that sensitive data is removed. The project is MIT-licensed.
Jevry is an open-source, MIT-licensed desktop browser agent created by Michael Swissa. It is an Electron browser that combines a language model for planning and text or visual reasoning with TypeSafe's Jev model for selecting typed actions from controls observed on the current page. Chromium executes the selected operations through guarded native input rather than generated JavaScript, arbitrary selectors, or shell commands. Its loop observes page text and compatible controls, compiles supported operation-and-target choices such as clicking, typing, scrolling, or stopping, asks Jev to choose, validates the choice against the current document and target, executes it, records an input receipt, and checks the resulting evidence before continuing. It supports browser tasks such as forms, page-grounded questions, research and citation, documentation navigation, catalogs, and selected visual or numeric games. The project is an experimental developer preview for macOS and Windows; it requires a TypeSafe Jev connection plus a text-model connection, and some capabilities—including uploads, extensions, password-manager workflows, and certain cross-origin interactions—remain unsupported or unverified.
jev-use is a macOS voice and text computer-use harness built on Jev. It reads the frontmost app's Accessibility tree rather than taking screenshots, sends the command, app and window names, visible targets, and recent actions to Jev, then selects and performs operations such as clicking, typing, pressing keys, scrolling, opening items, and arranging windows. After each operation it reads the Accessibility tree again and continues until the task is complete, blocked, or waiting; low-confidence or destructive actions stop for confirmation. The app supports hold-to-talk, typed commands, hands-free listening, and an optional “Hey Jev” wake phrase, with speech handled by Apple Speech. It requires macOS 14.2 or later and Xcode, has no external dependencies, and stores the TypeSafe API key in the macOS Keychain.
Jive is an open-source terminal coding agent that plans work as executable graphs. It replaces sequential tool calls with graph calls: the planner produces a directed acyclic graph containing tool calls and Jev calls, allowing the agent to perform multi-step profiling, dataset analysis, and repetitive tasks with fast decisions inserted between harder reasoning steps. The project emphasizes a lightweight, token-efficient agent scaffold with eager execution, flexible graphs, and support for system-one and system-two models. It can be installed with the project's shell installer and is distributed under the MIT license.
Keel is a local-first native macOS coding workspace that runs coding agents through a Rust/GPUI application. It combines sessions, a composer, transcripts, a terminal, change inspection, settings, Git history, crash recovery, and an agent graph showing active agents, subagents, and files as shared context; agents can be steered or stopped from the graph. It connects Claude Code, Codex, Cursor, Grok, Hermes, and pi through the Agent Client Protocol, and also includes an embedded DeepSeek agent loop with host-enforced tool focus. For unpinned tasks, a local Laya selector running through Core ML or an optional hosted Jev selector chooses among host-prepared routes or abstains. Keel validates each choice before applying it, falls back to the ordinary route when a choice is rejected, stale, or expired, and records candidates, validation, fallback, and observed outcome in bounded decision receipts. Decision receipts can be exported as JSONL and summarized with CLI reports. Keel also provides headless and daemon modes, while each coding provider retains its own authentication, model configuration, and tool loop. The repository supports Apple Silicon and macOS 15 or later for the packaged Laya model. It can be built from source, but the repository does not currently publish downloadable app releases; packaged builds are ad hoc signed development builds. Automatic training, public installer distribution, cloud sync, and native Codex realtime voice are not implemented or included in this build.
LaunchVideo is a tool that turns a website URL or prompt into a 20–40-second launch video. Opus 5.5 writes a single HTML film, while a serverless OpenComputer agent checks the scene at multiple timestamps and renders each frame in headless Chromium under a controlled virtual clock. The frames are encoded into an MP4 with ffmpeg and uploaded to Vercel Blob. The project includes a Next.js web form, OpenComputer agent configuration, API routes for job creation and status, and a CLI or one-click deployment path. URL inputs are fetched for page metadata, headings, colors, and Google Fonts before generation. Rendering runs in a fresh Amazon Linux arm64 microVM and supports deterministic HTML scenes subject to restrictions such as no video, audio, iframe, CSS transitions, random values, or external images.
An agent skill for Claude that guides logo projects from discovery briefs through concept development, SVG construction, testing, refinement, and delivery of brand assets. Its workflow researches category conventions in a classified library of more than 1,400 reference SVG logos, develops 8–12 concepts, builds selected concepts in SVG, and pauses at a checkpoint for direction selection before producing a full kit. Dependency-free Python tools audit SVG structure and craft rules, generate concept and test sheets, check marks at sizes down to 16 px in one-colour and reversed treatments, compare them in contextual and competitor-shelf tests, render PNGs, create presentation boards, and export favicon, app-icon, web-manifest, lockup, and other variants. It can also produce icon sets and brand-guideline materials. The skill follows the Agent Skills format, supports Claude Code, Claude.ai, Claude Desktop, and compatible agents, requires Python 3.8+ for its tools, and is distributed under the MIT license; the reference logo files are third-party trademarks excluded from that license.
mu (μ) is a coding agent with a judgment kernel, built on pi and AionUi. A small judge handles bounded decisions around context admission, command risk and approval, task framing, memory, completion, drift, browser actions, and multi-agent coordination, while the main model performs the coding work. Each decision can be active, shadow, or off and can use Jev, the local Laya judge, or another LLM; verdicts, probabilities, and timings are recorded in a local ledger that can be inspected from the command line or desktop app. The system admits tool output chunk by chunk, archives or removes stale context, applies rule-based safety checks before judgment, and uses a shared board to route findings among non-editing sub-agents. It provides an npm-installed command-line interface and a native desktop app, supports model-provider connections and importing Claude Code or Codex conversations, and states that it is in early development with no release yet; names, settings, and formats may change.
NeoHorse is TokenRhythm's family of open-weight causal language models for text-based agent workflows, including tool use, coding, and instruction following. NeoHorse-1 uses routing-guided agentic post-training: a harness assigns tasks to a heterogeneous model pool, records tool interactions and outcomes, estimates capability demand, and feeds capability-level feedback into later training mixtures. Updated models can return to the harness in an evaluation-selection-update loop intended as a prototype for recursive self-improvement. NeoHorse-1 is released in approximately 4B and 9B parameter variants, post-trained from Qwen3.5 models, with text input and text output interfaces and deployment examples for vLLM and SGLang. The related NeoHorse-Jev-4B model uses prefill-only inference to return structured Choice, Noul, and Score decisions and probabilities for routing requests, selecting tools, or controlling workflows. The models are released under the Apache License 2.0.
Ollaya is an open-source system for running decision models locally. It accepts a message, email, ticket, or other JSON state together with typed questions and returns calibrated probabilities for choices, scores, or binary decisions rather than generated text. The Ollaya daemon pulls models by name and serves them through a CLI and HTTP APIs, including TypeSafe-compatible `/v1/systemone`, `/v1/decisions`, and `/v1/models` endpoints. Its native API adds routing information and timings, while model definitions can embed reusable question sets. It supports ONNX Runtime on CPU and NVIDIA GPUs, and llama.cpp for GGUF models on CPU, CUDA, and Apple silicon. Ollaya also provides an MCP server for clients such as Claude Code, Claude Desktop, and Cursor. It publishes small ONNX graphs while obtaining the original model weights from their authors' Hugging Face repositories, pinned and verified by commit and SHA-256; it does not re-host those weights. The project is distributed under Apache-2.0, while individual models retain their own licenses, and it is available as a CLI, daemon, desktop application, and Docker image.
OmaPhoto is a layered image editor for Linux, developed first for Omarchy. It provides layers, folders, masks, clipping masks, blend modes, selections, painting and retouching tools, transforms, adjustment layers, layer effects, filters, text, shapes, and image export. It imports JPEG, PNG, TIFF, HEIC, Photoshop files, and camera RAW images. OmaPhoto is a Linux port of Wonder Assembly's macOS Compositor and follows Compositor's tools, menus, shortcuts, and layered project format. It can read and write .comp projects so projects can move between the two applications. Its offline Remove Background feature uses U²-Net through ONNX Runtime; the editor also includes Content-Aware Fill, Gaussian and motion blur, noise, lens correction, and other image adjustments. The project distributes packages for x86_64 Arch, Ubuntu, Fedora, and NixOS systems, and can also be built from source with Docker or a C++20, Qt 6.4-or-later, CMake, Ninja, image-format, RAW-processing, font, and ONNX Runtime toolchain. It is licensed under MIT.
Omarchy Meeting Recorder is a local meeting-recording application for Omarchy on Hyprland, built with Rust, GTK 4 and libadwaita. It captures microphone audio and the computer's output as separate tracks, then uses whisper-rs and a local Whisper model to produce a timestamped transcript with speaker labels. Recordings can also be imported from files; imported single-track audio uses NVIDIA Nemotron 3 Diarization through ONNX Runtime to distinguish speakers locally. The application provides an editable transcript, waveform playback synchronized with transcript lines, speaker renaming, chapter markers, pause and crash recovery, command-line controls, and optional bar-widget integration. Chapters can be generated by the configured Omarchy coding agent; in that case, the transcript text is sent to that agent, while recording and transcription otherwise remain on the user's computer. It stores meetings as folders containing audio, Markdown transcripts, metadata, and separate track files. The project is distributed under the MIT license and can be installed from the Omarchy package repository or built from source.
omnibin is a Linux tool that exposes binaries from Nixpkgs on the PATH without installing all of their packages in advance. It mounts a FUSE-backed filesystem over the Nix store and uses package metadata to resolve executable names and versioned forms; package files are fetched lazily from cache.nixos.org when they are read, then served locally. The omnibin command can locate the newest or all available versions of a binary, including historical Nixpkgs versions. It can run in a shell, as a NixOS module, or in a container; the container requires /dev/fuse and the SYS_ADMIN capability. The project is distributed under the MIT license.
Phantomat is a Hyprland plugin that replaces workspaces with a zoomable, infinite 2D canvas for windows. Windows can be placed freely, while linked monitors show adjacent parts of the same canvas and move together. It supports Wayland and X11 applications, including games running under Wine and Proton. The plugin provides a zoomed-out navigation view with title- and app-name search, camera movement to matching windows, recent-window switching, panning, zooming, window movement and resizing, minimap navigation, grid arrangement, undo and redo, fullscreen and screen-fill modes, and optional pinned windows. It remembers window positions and camera state across restarts and includes a live tuner for lens distortion, blur, grid, parallax, HUD, colors, and related settings. Phantomat is distributed as source code under the BSD 3-Clause license. It requires Hyprland 0.56 or newer, Lua configuration support, GCC 15 or newer, and Hyprland development files; the plugin must be rebuilt after a Hyprland update. It was formerly called Spatial Overview, so its configuration files and commands retain that name. The project began as a fork of hyprland-scroll-overview and is described as less tested outside its Omarchy, Hyprland, NVIDIA, and two-monitor development setup.
product-film is a Claude Code skill for creating product films in Remotion, including landing-page loops, launch videos, promos, and demo reels from a product's real components, design tokens, logo, and stated claims. It first inspects the product's design rules, tokens, components, live site, logo, and claims, then interviews the user about the film's duration, placement, music, features, and visual ingredients. It records the resulting brand kit in videos/BRAND.md, creates a beat sheet, analyzes the music's beat grid, builds time-based Remotion scenes, and supports review through style frames, stills, contact sheets, and a draft render. The final workflow renders a 240 fps master with motion blur and can produce a muted loop, a music version, a WebM file, and a poster; verification scripts decode the outputs and check colors, duration, and loop continuity. It runs in a local Claude Code toolchain rather than the Claude chat apps, requires Node.js, Bun, and uv, and is distributed under the MIT license for its own text and code.
reladraw is a text language and command-line tool for creating diagrams with relative positioning. It sits between automatic layout systems such as Mermaid, Graphviz, and D2 and manually positioned editors such as draw.io: users declare nodes, relationships, labels, and relative placement without specifying absolute coordinates. The parser and layout engine resolve the declarations into a diagram, while the tool reports issues such as overlaps, crossed edges, and text overflow without requiring a rendered preview. Its TypeScript CLI converts `.reladraw` files to SVG, and the repository also provides a browser version and an optional agent skill. The project is early-stage, its syntax is subject to change, and the code is licensed under Apache-2.0.
ritridata is a Rust command-line tool for read-only inspection of disk images and synthetic image generation. It reads MBR and GPT partition metadata, checks GPT header and entry-array CRCs, reports bounds and overlapping partitions, reads sectors and byte ranges, produces hex dumps and hashes, and exposes public file-signature definitions through ID lookup. Its generator creates blank, MBR, and GPT test images, including synthetic invalid-CRC and truncated fixtures, with JSON output for structured commands. The tool operates offline and accepts regular .img, .raw, and .dd files; input images are opened read-only, while generated files use an exclusive new-file operation. It does not scan or carve files, parse filesystems, recover deleted data, repair images, or substitute GPT backup entries for an invalid primary. The repository is licensed under AGPL-3.0-only.
Rooms is a macOS menu-bar window-management app that saves each project as a named room containing selected windows and their layout. Pressing ⌥Space and entering a room name restores the windows on the current screen while hiding or parking unrelated windows; rooms can also be assigned direct keyboard shortcuts. It offers Focus, Columns, Grid, custom grid, and Stack layouts, measures application minimum sizes, and remembers separate arrangements for laptop and external displays. Rooms stores room and parked-window data locally on the Mac and makes no network connections. It requires Accessibility permission, supports macOS 14 or later on Apple Silicon and Intel, and is distributed under the MIT license.
Tidewater is a browser-based island fishing game. Players fish from a pier, beach, or boat, manage line tension while fighting catches, sell fish to a village vendor, and spend the proceeds on equipment such as stronger line, reels, rods, a larger hold, fuel capacity, an engine, a fish finder, and deck lights. The island and surrounding ocean can be explored on foot, by swimming underwater, or by boat, with fishing conditions varying by water depth and time of day. The game runs directly on WebGPU and WGSL using its own rendering engine rather than a framework. Its systems include a four-cascade FFT ocean based on Tessendorf spectra, shallow-water swash simulation, boat wakes, underwater lighting and caustics, a physically based atmosphere, volumetric clouds, dynamic lighting, positional audio, and browser-saved progress. It requires a recent browser with WebGPU and a capable GPU; the first load compiles several hundred shaders. The source code is released under the MIT license and is deployed as a static build through GitHub Pages.
Whiteboard is an open-source desktop app for humans and coding agents to architect software in a shared workspace. It connects to agents such as Claude Code and Codex, giving them an SDK for drawing diagrams and describing their work on an in-app canvas. Visualizations including sequence diagrams, entity-relationship diagrams, and agent-trace quotations can link directly to underlying code, while the code-review interface provides VS Code keybindings and LSP support. Whiteboard includes a semantic, AST-aware diff viewer written in Rust. It summarizes large added functions as pseudocode and can collapse or hide changes such as tests and documentation; a WebAssembly-based plugin system makes these rules customizable. Its decision-log tools let agents query and link their traces so users can inspect requirements, implementation choices, and autonomous decisions. The app runs against local checkouts, is distributed for macOS, Windows, Ubuntu, and Fedora, and is licensed under the MIT License. It is self-hostable, although hosted team functionality is planned. Current limitations include no file editing within Whiteboard and limited support for reviewing multiple repositories together.
YuE2 Studio is a Windows desktop application for local AI song generation. It takes a musical style and lyrics, uses YuE2 to compose a melody and chords as ABC notation, engraves the result as sheet music, and renders the composition as a full song with vocals. The score can be edited and rendered again with different sounds, while saved audio codes support exact replay and variation without recomposing. The studio also supports score-first composition, melody transcription for covers through SheetSage2, lyric-level karaoke timing, six-stem separation, audio-to-MIDI conversion, LoRA loading and training, audio processing with VST3 plugins, and an optional local or OpenRouter writing assistant. It can be controlled by AI agents through a local MCP server. It is distributed as a native Windows installer with auto-update or a portable folder and does not require Python or Node.js at runtime. Generation runs offline through the bundled yue2.cpp engine on NVIDIA, AMD, or Intel GPUs, with Vulkan support for the latter two described as experimental; CPU execution is also available. The studio is MIT-licensed, while the YuE2-3B, YuE2 VAE, and SheetSage2 models are CC BY-NC 4.0 and restrict generated songs to non-commercial use unless other terms are obtained.
Searchable transcript of GitHub Trending Today #51: hindsight, reladraw, Whiteboard, NeoHorse, company-brain, CLM, cogsend — Github Awesome (14:51). Search for a phrase, then click its timestamp to jump straight to that moment in the video.
Captions sourced from the original video on YouTube, published by Github Awesome. The video, its captions and all related intellectual property remain the property of their respective owners; AINotes claims no ownership. Provided for research, accessibility and search — see the Transcript Notice and Copyright Policy.
00:00 Welcome back to GitHub awesome. This is GitHub trending today 51 35 trending open source projects on GitHub right now. Let's get into it. Hindsight is an agent memory system from Vectorize built to help agents learn from what they've done beyond just recalling old chats. It sorts what it keeps into world facts, the agents own experiences, observations backed by several memories, and mental models that update as it learns.
00:26 Recall runs semantic, keyword, graph, and timebased searches in parallel, then reranks the results. Real is a text diagram language that sits between mermaid style auto layout and dragging boxes around in excellent. You write things like above left of the hub or between desktop one and laptop one. And a resolver treats those as minimum distances and finds the tightest layout that fits.
00:54 It also reports overlaps, crossed edges, and text overflow without rendering. Whiteboard is a desktop canvas where you and your coding agent design software together. Clawed code or codecs draws sequence and entity diagrams onto the canvas and clicking a box jumps you to the code behind it. Its diff viewer is aware so big functions collapse into pseudo code and tests fold away.
01:19 Agents also log their decisions so you can trace why something was built the way it was. Neoor is a family of openw weight models trained from Quen 3.5 for coding and tool use. Its training harness routes tasks among models, records their actions and outcomes, then uses that feedback to shape the next training mix. The release includes 4B and 9B language models, plus a 4B Jev variant that scores predefined decisions.
01:47 The repository describes recursive self-improvement as a goal. Company brain is a Slack bot from Super Memory that remembers what your team says and can act on it. It only reads with the asker's own access. So a question in a public channel draws on shared memory while a DM can also pull from your private channels. It opens issues and reads PRs with writes running under your own credentials.
02:13 It self-hosts on Cloudflare workers, but memory still runs through Super Memories API. CLM, short for contrastive language model, makes agent decisions by matching embeddings instead of generating text. It encodes the state and each candidate action separately on a Quen 38B backbone with small trainable heads, then picks the action that lines up best.
02:36 It serves through a Typesafe compatible API and the authors report parody with Jev at up to nine times lower latency. Golive is an agent skill for codecs and claude code that takes the app your agent just built and puts it into production on your own accounts. It figures out what the app needs, shows you a plan for hosting, database, domain, email, and payments, and waits for your approval before touching any provider.
03:05 When it's done, it writes a handover file listing everything it created so you know exactly what you own. Cogsend is a self-hosted social scheduler that posts to Mastadon, Blue Sky, LinkedIn, Threads, and X from one draft. You write once, tweak the copy per platform, and publish now or schedule it. It runs on Cloudflare workers with D1 and R2, and a per minute cron sends what's due.
03:29 Failed posts retry five times, then park in a failed list instead of vanishing. Credentials are encrypted at rest and the admin account has 2FA. Disk tree shows what's filling your drive as a map of nested blocks sized by the space files actually use. Zoom into a folder, spot bulky caches or old builds, and mark what you want to remove. Before anything goes, you review the full list and see how much space it could free.
03:57 It prefers the trash and permanent deletion asks for confirmation. It runs on Linux, Mac OS, and Windows. Inkwave is a browser game where winning means painting more ground than the other team. You play threeinut fouron- matches against bots, switching between seven weapons and a squid form that swims through your ink and climbs painted walls. It's built with 3JS and runs without a build step.
04:22 Even the characters, animations, textures, sound effects, and music are generated in code. decision models locally. Much like Olama lets you run language models, give it a message, email, or ticket, and ask things like, "Is this urgent?" or "What's the customer asking for?" It returns scores and probabilities instead of generating a paragraph. You can pull models by name, serve them through an API, and connect them to coding agents through MCP.
04:54 Jive is a terminal coding agent that turns a plan into an executable graph. Instead of asking a large language model what to do after every tool call, it can run connected steps and use smaller decision models along the way. That makes it an interesting fit for repetitive work like reviewing lots of records or tracing errors across a code base where constant model round trips add up.
05:18 MW is a coding agent with a small judge handling routine decisions while the main model focuses on the code. The judge decides things like which tool output belongs in context, whether an action fits your instructions, and when a task is actually finished. Each verdict goes into a ledger you can inspect. You can even run a judge in shadow mode to see its decisions before letting them affect the agent.
05:44 Clifact Shortcuts looks through your coding agent history for requests you keep typing. It groups similar asks, shows you the repeats, and lets you turn the useful ones into short project commands. So, did the newsletter reach everyone? Can become/sent with the right query and reporting steps saved behind it. The history extractor runs locally, and you choose which shortcuts get created.
06:10 Anadoodle makes handdrawn art with code, so the same source can produce an illustration, a drawing time lapse, or an animated film. It includes 31 styles from charcoal and watercolor to woodcut and pixel art. You can even add music composed in code. The neat part is consistency. Render it again at a different size and the marks stay where you put them.
06:32 Keel brings the coding agents you already use into one native Mac OS workspace. You can follow sessions, inspect changes, use a terminal, and see agents and the files they touch in a live graph. For new tasks, a local lia model can help choose which agent to use. Keel checks that choice and records the decision so you can see how the task was routed.
06:56 Logo design skill guides Claude through a logo project from the first brand brief to SVG concepts you can actually test. It checks how each mark looks at tiny sizes in one color on different backgrounds and beside competitors. There's also a searchable library of more than 1,400 real logos for studying design patterns. Once you pick a direction, it can build the icon set and brand guidelines.
07:22 Ship video turns a website URL or a short prompt into a launch video. An AI agent writes the animation as a single HTML page, checks how it looks at different moments, then renders each frame in a headless browser and encodes an MP4 with ffmpeg. The useful twist is that the video comes from editable web code, so its text, colors, and motion have a clear source.
07:46 Tidewater is a browser fishing game built around a tropical island you can explore on foot, underwater, or by boat. Fish bite in different places and at different times, and you have to manage line tension to land them. Sell your catch, upgrade your gear, and head farther offshore. Its ocean and sky run on a custom web GPU engine, so you'll need a capable browser and GPU.
08:10 UE2 Studio generates full songs with vocals on your own Windows PC. Give it a style and lyrics, and UE2 writes a melody and chords as sheet music before performing the track. You can edit that score, then render the same composition with a different sound. It runs on UA2.cpp, CPP and the author reports a 338 song in about 46 seconds on an RTX 4090. Rooms turns each project on your Mac into a saved set of windows and a layout.
08:40 Press option space, type a room name, and its windows come back in place while everything else hides. It can adjust the layout for your laptop or an external monitor, and you can save your own arrangement. The windows stay open, and your room data stays on your Mac. Jevy is a desktop browser agent that plans with a language model and picks each action with Typesafe's Jev.
09:02 It observes the page, turns the real controls on it into a list of choices and Jev picks one, so it never makes up a selector or writes its own JavaScript. You can redirect it midtask on a 12 task slice of web arena verified. The author reports 10 passes. Phantomat is a hyperland plugin that replaces workspaces with one infinite canvas. Every window sits on the same zoomable plane and multiple monitors show neighboring parts of it moving together like one desk.
09:32 Hit super control G and it zooms out to show everything. Then type part of a window title to fly straight there. It remembers where everything was across restarts. Product film is a clawed code skill that builds launch videos and landing page loops out of your product's real components with Remotion. It reads your design tokens and logos first, then writes a brand kit and a beat sheet.
09:57 A Python script maps the beat grid of your track so cuts land on the music. You approve stills and contact sheets before the final render, which comes out as a 240fps master with motion blur. Omnibin puts tens of thousands of Nyx package commands on your path without installing them all up front. A virtual file system knows which package provides each command and fetches its files from the Nyx cache only when you use it.
10:25 You can even call a specific version like an older Python release. That's handy for disposable Linux environments, though the first run of a tool needs a download. Changelog.ear Earth presents the news like patch notes for the planet. It gathers reporting, picks a few stories, and turns their headlines into short updates about discoveries, conservation wins, and things still unresolved.
10:50 Each note keeps the original publisher, date, and source link, so you can check the story behind the joke. It also saves past editions and publishes an RSS feed. Flux links an Omari desktop with an Android phone over your local network. Once you pair them, you can move files and clipboard text, see phone notifications, control media, and even use the phone as a webcam or microphone.
11:16 There's a desktop app for everyday use and a CLI for scripts. Pairing shows a verification code on both devices, so you can check the right phone. Interference search is a research project that explores reasoning through explicit states. It advances several possible moves at once, merges branches that reach the same state, and uses a small judge to drop dead ends.
11:39 The approach performed well on the project's countdown arithmetic tests, while gains on code problems were much smaller. The repo includes its code, model weights, raw results, and failed experiments. Bridge Clip turns long videos you have permission to use into short captioned clips. Drop in a file and it transcribes the speech, picks promising moments, refframes them for vertical or widescreen video and renders wordbyword captions.
12:09 The final video is cut on your computer with ffmpeg. Transcription and clip selection use your open router account, so audio and transcript data go to that provider. Omari meeting recorder records your calls without a bot joining them, and the audio never leaves your machine. It captures your mic and the computer's audio as two separate tracks, then transcribes locally with whisper.cpp and labels who said what.
12:34 Your coding agent adds chapters, and the only thing it sees is the transcript text. You get a waveform player and a transcript you can fix in line. Retradata is a Rust command line tool for looking inside disk images without changing them. It reads MBR and GPT partition tables, checks header and entry CRC's, flags overlapping partitions, and identifies files by their signatures.
13:01 You can also generate blank or deliberately damaged images for tests. Everything prints as JSON, and it runs fully offline. Halo battery puts the battery level of each wireless peripheral in the Windows system tray, so you can close Synapse. Every device gets its own ring icon that fills clockwise, turns amber when it's low, and breathes green while charging.
13:23 It talks directly to a Razer Black Shark V2 Pro, a WL mouse Beastex Max, and a Game Seir G7 Pro, and reads Bluetooth batteries from Windows. Jev use lets you speak or type a task and have your Mac carry it out. It reads the current app's accessibility controls, asks Jev to choose the next action, then checks the screen again after acting. That means it can work without sending screenshots.
13:48 Though your command and visible control details go to type safe, it needs a type-S safe API key and Mac OS accessibility permission. Omoto brings compositor's layered image editor to Linux with special attention to Omari. It even follows your desktop theme. You can work with masks, selections, retouching tools, adjustment layers, and offline background removal.
14:14 It also opens and saves compositors, comp projects, so a layered edit can move between Linux and Mac without being flattened into a single image. Jevp helps coding agents find the right file in an unfamiliar codebase. Ask where is authentication checked and it walks the repository, judges which files matter, then returns source excerpts with line numbers in one response. A companion skill shows agents when to use it. Relevant source content goes to the provider you configure, so check what you're comfortable sending.