A prediction engine that turns source material into a knowledge graph, generates agents with personalities and memory, and simulates them in a digital world.
An open-source design skill that applies anti-AI-slop rules, selects layouts and themes, critiques generated designs, and returns HTML and CSS.
An agent-native tutoring workspace combining chat, quizzes, research, problem-solving, visualization, and mastery practice while keeping learner context connected.
Agent Substrate is an open-source runtime and control plane for running agent-like workloads at high density. It manages the lifecycle of sandboxed actors, including creation, destruction, suspension, and resumption, and assigns them to a smaller pool of worker pods while routing traffic to the active workers. The system uses Kubernetes for infrastructure provisioning and pod lifecycle management, with support for sandbox technologies including microVMs and gVisor. It preserves actors' volatile memory and filesystem state through full-state snapshots, enabling sub-second suspend and resume operations and multiplexing many stateful actors across shared infrastructure. It manages standard OCI containers at the kernel level and is designed to support different agent frameworks and harnesses, including ADK-compatible actors, LangChain agents, coding environments, and MCP servers. Agent Substrate is intended for running agents at scale rather than building them, and its workloads do not have to be literal AI agents. The repository notes that the project is not an officially supported Google product.
anydoc is an open-source Rust library from Firecrawl that converts Word, PowerPoint, Excel, OpenDocument, RTF, EPUB, CSV, and PDF files into GitHub-Flavored Markdown. It uses a shared document model and provides Node.js, Python, browser WebAssembly, and Rust interfaces, along with a command-line interface that reads files or standard input and writes Markdown to standard output or a file. The library can return either Markdown or its intermediate document model, which carries embedded assets. A browser demo runs the WebAssembly build locally, while hosted OCR through Firecrawl Parse can process scanned PDF pages that anydoc cannot read by itself. It is also distributed as an agent skill compatible with agent tools such as Claude Code, Codex, Cursor, and OpenCode.
AutoGPT is an open-source platform for building, deploying, and running AI agents that carry out complete workflows. Users can describe an outcome in plain English, or use its visual builder to drag, connect, branch, and inspect workflow blocks. Agents can run on demand, on schedules, or from event triggers, and the marketplace provides ready-made agents that can be added and customized. The hosted platform is a paid, usage-based service that manages infrastructure, model access, credentials, and updates; the self-hosted path is free of license fees but requires users to provide infrastructure and model API keys and maintain the deployment.
Backchannel is an open-source userscript and renamed continuation of HNewhere that brings discussions from Hacker News, Bluesky, Reddit, Lemmy, Lobsters, and other selectable sources into a sidebar beside the page or PDF being read. It checks visited pages for matching threads, merges enabled sources into one conversation with source labels, and matches quoted comments to article passages as clickable annotations; enhanced PDF highlighting is supported except in Firefox. Users can vote, reply, and submit through popups on supported source sites, add locally stored notes to pages and PDFs, and use a customizable front page combining content from selected sources. The single-file script runs without a build step, backend, analytics, or telemetry, and is distributed through userscript managers such as Tampermonkey, Violentmonkey, and Userscripts.
Can I Vibecode It? is an open-source catalog that evaluates whether AI coding agents such as Claude Code, Codex, or Cursor can build a usable personal replacement for a paid SaaS application. Each app entry provides a YES, KINDA, or NOT REALLY verdict, the exact prompt for attempting the replacement, and the tradeoffs of leaving the original service, including lost network effects, data, or infrastructure. App entries are contributed as one JSON file per app, with the schema and verdict criteria documented in the repository. The site uses Astro server output with a Node adapter to render fully server-rendered HTML, SQLite via better-sqlite3 for vote counters and waitlist data, and vanilla JavaScript and CSS for interactions. Satori and resvg generate Open Graph images at build time. It runs locally with npm install, npm run dev, npm run build, and npm start, requires no environment variables for local development, and supports optional analytics, sign-in, payments, email, and media-storage configuration. It can be deployed on a VPS behind a reverse proxy, with DATA_DIR used to keep user data outside the repository. The project is released under the MIT License, and its prompts are intended to remain free.
Capptivo is a free, open-source screen recorder and editor for creating demo videos on macOS, Windows, and Linux. It tracks cursor movement and clicks to apply follow-cursor and click-based automatic zooms, and provides editor presets, annotations, facecam, backgrounds, captions, and MP4, WebM, or GIF export. Captions use a system whisper.cpp command-line binary, with model weights downloaded on first use. The application captures macOS screens through ScreenCaptureKit with VideoToolbox H.264 encoding, Windows screens through Windows.Graphics.Capture with hardware-encoding support, and Linux screens through xdg-desktop-portal and PipeWire. It is distributed through platform installers, including DMG or app archives for macOS, MSI or setup executables for Windows, and DEB, AppImage, or RPM packages for Linux. The repository releases it under the MIT license.
Comp AI CRM is an open-source, self-hostable customer relationship management system for AI agents, developed by Comp AI and hosted by Try Comp AI. Its agent runs as an independent durable deployment on its own schedule and database-backed work queue rather than waiting for browser requests: it selects records to investigate, researches contacts and companies, spends a research budget, schedules follow-ups and rechecks, records observed evidence, and sends weak or ambiguous matches to a human instead of writing them directly to records. It supports durable sessions, contact and company records, an Agent tab for viewing work and answering questions, authored tools for reading CRM history, searching records, identifying contacts, researching people, enriching companies, recording facts and scheduling rechecks, and versioned Markdown skills. The agent uses file-based tools and Markdown skills on Vercel's eve durable-agent framework, while its queue leases due tasks with PostgreSQL row locking so concurrent workers claim disjoint work and expired leases release tasks from failed runs. A restricted shell sandbox provides bash, grep, glob and a workspace without network egress, database credentials or direct database access. Optional sources include mailbox history, company brand data, LinkedIn and Perplexity web research.
DeepTutor is an open-source, self-hostable AI tutoring platform developed by HKUDS for lifelong personalized tutoring. It provides a web application and command-line interface for interactive learning, problem solving, quiz generation, deep research, visualization, and mastery practice, with persistent memory and learning state shared across knowledge bases, books, notebooks, tutor personas, and other activities. The platform supports multi-agent problem solving, retrieval-augmented learning with source-traced and page-level citations, math animations, interactive visualizations, and tutor bots built from personal study materials. Its knowledge sources include personal documents, EPUBs and books with annotations, GitHub repositories, web search, and connected libraries; documented retrieval and ingestion options include GraphRAG, PageIndex, LightRAG, linked knowledge bases, Obsidian, and configurable parsing and vector backends. DeepTutor also includes courses, research workflows, an ecosystem of MCP services and tool or capability plugins, and integrations with connected coding agents and other partners. It can run locally or through Docker.
Desktop Commander MCP is an open-source Model Context Protocol server that lets AI assistants search, read, write, move, and edit files; run terminal commands and code; and manage interactive or long-running processes. It is built on the MCP Filesystem Server and adds search-and-replace editing, streamed command output, session management, process controls, configurable timeouts, and background execution. It also supports file previews and Markdown editing, data analysis for CSV, JSON, and Excel files, and native operations on Excel, PDF, and DOCX documents. The repository documents use with Claude Desktop and other MCP clients, as well as remote AI control through ChatGPT, Claude web, and other AI services. Its associated Desktop Commander App adds a graphical interface, live file-change previews, arbitrary model support, and custom MCP configuration.
Hallmark is an open-source design skill for Claude Code, Cursor, and Codex that applies anti-AI-slop rules to generated interfaces. It selects a macrostructure and one of twenty-one themes for a brief, applies its design rule set, runs fifty-seven anti-pattern gates and a pre-emit self-critique, and returns self-contained HTML and CSS. Its four operations build new UI, audit existing code without editing it, redesign an interface while retaining its copy, information architecture, and brand, or study a screenshot or URL to extract macrostructure, type pairing, and a color anchor; the study operation can also emit a portable design.md file. Custom briefs can receive a made-to-measure palette, typography, and layout rather than a catalog theme. Hallmark is made by Together AI and distributed under the MIT license.
MAGI-2 Preview is the inference implementation for SandAI’s unified audio-video generation model. The 114-billion-parameter architecture, built on MagiMoE, activates about 6 billion parameters per token and generates 10-second clips from text prompts (T2V) or a prompt plus a still image (I2V), with sound generated alongside the video and muxed into the output file. Generation runs in two stages: `magi2_preview` denoises the clip at low resolution, after which `magi2_refiner` increases it to 1080p. The current base release is not step-distilled, so denoising steps account for most of the generation time; a distilled release is described as forthcoming. The repository contains inference code, while the model weights are downloaded separately from the `sand-ai/MAGI-2-preview` Hugging Face repository and occupy hundreds of gigabytes in total. An optional prompt-enhancement stage sends the input to an OpenAI-compatible instruction-following LLM, which rewrites it as a structured JSON caption for the 10-second clip, renders that result as Markdown, and passes it to the generation pipeline. It can be disabled to use the raw prompt. The documented runtime requires eight NVIDIA Hopper GPUs, Python 3.12, a recent CUDA toolkit, and FFmpeg for audio-video muxing; Docker images and a source installation are provided.
MiroFish is an open-source multi-agent prediction engine that builds simulated digital worlds from source materials such as news, policy drafts, financial signals, reports, or stories. Its workflow extracts entities and relationships into a knowledge graph, injects individual and collective memory through GraphRAG, generates agent personas and configurations, and runs parallel simulations in which agents with independent personalities, long-term memory, and behavioral logic interact and evolve. Users can describe prediction requirements in natural language, introduce variables during a simulation, chat with agents in the simulated world, and receive a generated prediction report.
morphicons is an open-source, zero-dependency JavaScript library for morphing one stroke-based SVG icon into another with spring physics. It accepts icon data such as IconNode objects or raw path strings, supports icon sets including Lucide, Tabler, Heroicons, and Iconoir, and can use custom paths. Its morphing algorithm solves the optimal similarity between shapes using 2D Procrustes analysis and polar interpolation, allowing rotation and scale to emerge mathematically rather than being declared for each icon pair. The library provides React, Vue, Svelte, React Native, vanilla custom-element, Astro, and core usage options, including uncontrolled prop-driven animation, controlled progress, and imperative morphing. It is distributed as an ESM package with optional framework peer dependencies.
OfficeCLI is an open-source command-line Office suite developed by iOfficeAI for AI agents to create, read, edit, render, and automate Word (.docx), Excel (.xlsx), and PowerPoint (.pptx) files without a Microsoft Office installation or external dependencies. It is distributed as a single binary and provides commands for creating documents, adding, modifying, moving, copying, and removing elements, reading text and structure as plain text or JSON, and saving changes. Its built-in HTML rendering engine converts DOCX, XLSX, and PPTX files to HTML or PNG for a render-and-review workflow. The `watch` command provides a live browser preview that refreshes after document changes, while the tool can also analyze formatting and structural issues, evaluate Excel formulas, and install an agent skill for supported AI coding agents.
Semantica is an open-source Python infrastructure layer for building context graphs and knowledge graphs for AI systems. It ingests enterprise and other multi-source data, extracts entities and relationships, flags conflicting facts, merges duplicates, and supports ontology management, knowledge modeling, graph analytics, and causal reasoning. The system records provenance and execution trails for the context supplied to an AI system and the decisions produced from it, making those relationships and decisions queryable. Its repository describes deterministic graph construction, reasoning, and provenance that do not require an LLM; it explains the data and policies outside an LLM rather than exposing the model's internal reasoning. Semantica supports RDF and labeled-property-graph storage, W3C standards, self-hosted deployment, and installation with pip.
SimpleEnglish is an open-source agent skill that makes large language models write technical documentation in ASD-STE100 Simplified Technical English, a controlled language used in aerospace. Its rules enforce short sentences, active voice, explicit instructions, simple tenses, and consistent terminology when rewriting READMEs, error messages, incident reports, runbooks, and release notes. The skill is distributed as a dependency-free folder for tools that support the Agent Skills standard, including Claude Code, Cursor, VS Code Copilot, OpenAI Codex, Gemini CLI, Goose, and OpenCode. It can also be installed as a Claude Code plugin and output style, or used by adding its prompt to a system prompt, AGENTS.md, or .cursorrules file. The repository is licensed under MIT.
Swiftlet is an open-source Swift and Metal runtime for running Qwen3-Next and Qwen3.5/3.6 mixture-of-experts models locally on Apple devices, including iPhones. It keeps the models' small dense core in memory and streams routed expert weights from storage on demand, allowing 35B and 80B models to run with comparatively low RAM. The project is available as a Swift package, command-line interface, OpenAI-compatible server, and iOS app; its repository describes the runtime as working end to end and notes that only about 3B parameters are active per token in the supported models.
TencentDB Agent Memory is a self-hostable, team-level memory hub for AI agents. It extracts conversations and tasks into reusable Chat Memory and Skills, and converts documents and code into an LLM Wiki and CodeGraph. These assets can be reviewed, versioned, governed, shared, routed, and reused across agents, frameworks, and team members; existing documents, codebases, and agent sessions can also be imported to reduce cold-start work. The system provides a shared memory server and a proxy that preserves the agent protocol, so supported clients can use the same memory by pointing their base URL at the proxy without plugins, hooks, or an MCP server. The repository deploys memory-core, memory-hub, and the proxy together, with a local management panel and configuration for separate memory and proxy LLM parameters. Documented integrations include DeepSeek Harness, Claude Code, Codex, CodeBuddy, WorkBuddy, Hermes, and OpenClaw.
vPhone Workstation is a native macOS SwiftUI application for managing virtual iPhones. It provides a library with live running or stopped state, iOS and cloudOS builds, search, and VM details including variant, CPU, memory, disk, network, device, and UDID. Users can start, stop, foreground, clone, rename, export, configure, and delete virtual machines. Its creation wizard selects a VM variant, firmware pairing, and CPU, memory, and disk resources, while streaming progress from vphone-cli and producing an exportable log. vPhone Workstation is a graphical front end for vphone-cli, which boots iOS research virtual machines through Apple's Virtualization.framework using PCC research VMs. It also checks host readiness, including the required macOS version, Apple Silicon hardware, vphone-cli availability, research-guest permissions, and AMFI configuration. The repository documents macOS 15 or newer as a requirement and provides Homebrew installation and source-build instructions.
Waku Agent is an open-source, local-first personal assistant and AI-agent harness built by seanchen.io. It runs on a laptop through a terminal, local browser dashboard, voice input, or Telegram gateway, and is organized around four components: a harness for tool execution, an approximately 95-line plain-Python agent loop, memory, and evaluation/LLM operations. Its memory uses semantic, episodic, and procedural stores in a local SQLite database, with a gate that decides whether to retrieve or save memory and a pass that determines what to retain. The project includes deterministic tests alongside LLM-as-judge evaluation with a release gate. The dashboard displays message flow through the harness, including gate decisions, tool calls, loop iterations, memory updates, traces, costs, and latency; it also includes views for workflows, tools and MCP connectors, memory, and the SQLite state database. It can be installed with pip and exposes the `waku` command. The dashboard runs a local web server at `127.0.0.1:7777`, and the Telegram gateway starts when `TELEGRAM_BOT_TOKEN` is configured. Model providers are selected through a provider setting and API key, with support stated for Anthropic, OpenAI, Gemini, DeepSeek, MiniMax, Kimi, GLM, OpenRouter, OpenCode Zen, and OpenCode Go.
Searchable transcript of Top Open-Source GitHub Projects : AutoGPT, Capptivo, anydoc, Swiftlet, waku-agent & morphicons #282 — ManuAGI - AutoGPT Tutorials (13:37). Search for a phrase, then click its timestamp to jump straight to that moment in the video.
Captions sourced from the original video on YouTube, published by ManuAGI - AutoGPT Tutorials. The video, its captions and all related intellectual property remain the property of their respective owners; AINotes claims no ownership. Provided for research, accessibility and search — see the Transcript Notice and Copyright Policy.
00:00 Every week developers release new open-source projects on GitHub, and this is our weekly GitHub project update video covering the best of them. Top trending open-source GitHub projects this week. You'll discover useful and trending developer tools from local AI agents to self-hosted platforms. Without wasting time, let's get started. >> Before we jump into today's project updates, here's a quick announcement for everyone.
00:24 We've launched a brand new YouTube channel called AI Agent Studio dedicated entirely to AI agent projects, tutorials, and tools. So, if you're interested in staying up to date with the latest AI agent open-source projects, learning how to build your own agents, or exploring cutting-edge agent frameworks, make sure to check it out. Subscribe now to get weekly videos, in-depth guides, and real-time project breakdowns.
00:50 The link is right there in the description. Don't miss it. All right, let's get into today's video. >> Project number one, Semantica, graph infrastructure for accountable, auditable AI agents. Semantica is an open-source Python infrastructure layer that gives AI agents a knowledge graph with audit trails. It sits under your LLM, vector store, and agent framework, ingesting data from databases, Databricks, and Snowflake, then extracting entities and relations into a context graph.
01:19 Every fact carries W3C PROV-O provenance. Every decision becomes a traceable, queryable node, and reasoning runs deterministically through re-data log and SPARQL without an LLM. Storage spans RDF and property graph backends. Install it and ground your agents. Project number two, Miro Fish, multi-agent simulation engine for predicting the future. Miro Fish is an open-source prediction engine that uses multi-agent simulation to rehearse the future.
01:48 You upload seed material like news, a policy draft, or a novel, and describe your prediction. It builds a knowledge graph, generates thousands of agents with personalities and memory and runs them through a digital world where they evolve. You inject variables mid-run and chat with any agent and it returns a report. It runs on any OpenAI format model.
02:08 Deploy it and rehearse the future. Project number three, AutoGPT, build, deploy, and run continuous AI agents. AutoGPT is a free, self-hostable platform for building, deploying, and running continuous AI agents that automate workflows. You assemble an agent in a low-code builder by connecting blocks, where each block performs one action, then deploy it on the AutoGPT server, where agents run continuously and can be triggered by outside events.
02:37 A marketplace offers ready-made agents. For example, an agent turns Reddit topics into short videos. It installs with one command via Docker. Self-hosted and automate your workflows. Project number four, Tencent DB agent memory, team memory hub of reusable assets for agents. Tencent DB agent memory is an open-source, self-hostable memory hub for teams of AI agents.
02:59 It turns conversations, documents, and code into four reusable assets: chat memory, version skills, a wiki, and a code graph of symbols and call paths. A human-controlled panel governs ownership, versions, and visibility, so you equip agents with only what they need and share safely. New agents import repos, docs, and past sessions. It works with Claude code and open claw.
03:24 Deploy it and give your agents memory. Project number five, Enaun Hallmark, anti-AI slop design skill for Claude code, cursor, codex. Hallmark is an open-source design skill for Claude code, cursor, and codex that refuses to look AI-generated. It encodes anti-AI slop rules for typography, color, layout, and motion. It picks one of 21 structures, dresses it in one of 22 themes, then runs 65 slop test gates in self-critique before returning HTML and CSS.
03:57 Four verbs let you build, audit, redesign, or study a design from a screenshot or URL. Per project memory keeps outputs from repeating. Install it and stop shipping slop. Project number six, Deep Tutor, self-hostable agent-native workspace for personalized tutoring. Deep Tutor is an open-source self-hostable AI tutoring workspace from H K U D S. It runs chat, quizzes, deep research, problem-solving, visualization, and mastery practice on one agent loop, so context follows the learner.
04:29 Knowledge bases, books, notebooks, and personas stay connected with retrieval across llama index, graph rag, light rag, or a linked obsidian vault. You consult a live Claude code or Codex agent mid-turn. Install community skills and inspect a three-layer memory that traces every fact to its source. Install it and start tutoring. Project number seven, Desktop Commander MCP, terminal, file, and process control for AI.
04:56 Desktop Commander is an open-source MCP server that gives an AI assistant control of your terminal, file system, and processes. It runs long-running commands, manages processes, and searches files by name or content with rip grep. It reads and edits text, Excel, PDF, and Word files, applies search and replace edits, and runs Python, Node, or R code in memory for data analysis.
05:20 It installs into Claude desktop, Cursor, VS Code, and Codex with Docker isolation. Install it and hand over your desktop. Project number eight, Office CLI, command-line office suite built for AI agents. Office CLI is an open-source command-line office suite, so AI agents can read, edit, and automate Word, Excel, and PowerPoint files. It ships as a single binary with no office install, giving every element a stable path, so agents navigate without XML.
05:50 A built-in rendering engine turns files into HTML or PNG, so an agent can see its output and fix layout issues. It evaluates Excel formulas and runs as an MCP server. Install it and let your agent build documents. Project number nine, AnyDoc, convert any office document to clean markdown. AnyDoc is a fast open-source Rust library from Firecrawl that converts documents into clean markdown.
06:16 It turns Word, PowerPoint, Excel, OpenDocument, RTF, EPUB, CSV, and PDF files into GitHub-flavored markdown. Parsing every format through one document model, so tables, lists, and footnotes stay consistent. It detects formats from bytes and runs in milliseconds with no ML models or external services. It ships bindings for Node.js, Python, and browser, plus a CLI and agent skill.
06:44 Install it and convert any document. Project number 10, Simple English, make AI write in simplified technical English. Simple English is an open-source agent skill that makes an LLM write documentation in ASD-STE 100 Simplified Technical English, the controlled language aerospace uses, so instructions can't be misread. It applies 53 numbered rules, short sentences, one word per meaning, active voice, simple tenses, no hedging modals, condition before command, and reduces AI slop.
07:16 It adapts to error messages, runbooks, and release notes. It installs into Claude Code, Cursor, Codex, and any agent skills client. Install it and write docs that can't be misread. Project number 11, Agent Substrate, run agent workloads at scale on Kubernetes. Agent Substrate is an open-source system built on Kubernetes that runs agent-like workloads at higher density and lower latency.
07:40 It maps idle actors onto a pool of worker pods, suspending and resuming them with state intact for multiplexing. It keeps the Kubernetes control plane off the request path, isolates workloads with G visor, and stays framework agnostic across ADK, LangChain, Claude Code, and MCP. It runs on any cluster through a kubectl plugin, and it's early. Deploy it and scale your agents.
08:04 Project number 12, Camp AI CRM. Open source agentic first CRM. Built for AI agents, Camp AI CRM is an open source self-hostable CRM built around an AI agent. The agent runs on its schedule, filling records and booking its own follow-ups, and keeps going after you close the browser. Its rule is that nothing about a person is guessed. Tools report only what they observe, and weak evidence becomes a suggestion a human settles.
08:31 It treats your own email as evidence and needs no API keys. Deploy it and let the agent keep your notes. Project number 13, Waku Agent. A readable, local-first personal AI agent. Waku is an open source, local-first AI assistant you run on your laptop in readable code. Its memory lives in one SQLite file, semantic, episodic, and procedural, with a gate that decides whether to retrieve.
08:56 The loop is 95 lines of plain Python, and a local dashboard shows each message flow through the harness. Deterministic tests and an LLM judge run behind a release gate. You bring one API key, reach it by terminal, voice, or Telegram. Clone it and read the agent. Project number 14, Morphicons. Morph any stroke-based icon into any other. Morphicons is an open source JavaScript library that animates any stroke-based icon morphing into any other with spring physics.
09:25 It works with Lucid, Tabler, Heroicons, Iconoir, or your SVG paths, and needs no hand-declared rotation pairs. It solves optimal alignment between shapes and interpolates in polar space, so rotations emerge from the math. Morphs are interruptible, keeping velocity mid-flight. It ships React, Vanilla DOM, and Pure Core builds, has zero dependencies, and weighs 6KB GZipped.
09:51 Install it and morph your icons. Project number 15, Backchannel. Add Hacker News and Reddit threads to articles. Backchannel is an open-source user script that adds Hacker News and Reddit discussions to any article you read. It checks each page against sources you've enabled and lights up a button when a thread exists, then opens a sidebar merging both communities, tagged by source.
10:13 A beta layer matches quotes back to the article, so you can filter the thread by passage. You can vote and reply on Hacker News. It has no back-end or telemetry. Install it and read along. Project number 16, vPhone Workstation. Native macOS app for managing virtual iPhones. vPhone Workstation is an open-source native macOS app for managing virtual iPhones.
10:35 It's a SwiftUI front-end for vPhone CLI, which boots iOS research VMs on Apple's virtualization framework, so you browse, create, and boot them without the terminal. A create wizard walks you through variant, firmware, and resources, while a library shows each VM's state, build, and specs. You can clone, rename, export, and delete VMs, and it checks your host first.
10:58 Install it and boot a virtual iPhone. Project number 17, Swiftlet. Run 35B and 80B Qwen models on iPhones. Swiftlet is an open-source Swift and Metal runtime that runs 35B and 80B Qwen mixture of experts models on Apple devices, including iPhones. It keeps a model's dense core in memory and streams routed experts from storage, so the 80B needs about 4.3 GB of RAM, and the 35B fits on an iPhone.
11:27 Experts are repacked, so each fetch is one read. It ships as a Swift package, CLI, OpenAI compatible server, and iOS app. Build it and run a big model locally. Project number 18, Can I Vibe It Coded, which SaaS you can replace with one prompt. Can I Vibe It Coded, or Can I Vibe It Coded, is an open-source website that tells you which SaaS subscriptions an AI coding agent could replace.
11:51 For each app, it gives a verdict, yes, kind of, or not really, on whether Claude Code, Codex, or Cursor can one-shot a personal version, the exact prompt to build it, and what you'd lose by leaving. Entries are JSON files contributed by pull request. It's built with Astro and SQLite, self-hostable on any VPS. Browse it and cancel a subscription. Project number 19, Captivo, open-source screen recorder with smart cursor zoom.
12:19 Captivo is a free, open-source screen recorder and editor for macOS, Windows, and Linux, and an alternative to Screen Studio. It captures your screen through each OS's native pipeline with hardware H.264 encoding, saving a cursor and click track. From clicks, it auto-suggests zooms, follows the cursor, and lets you add annotations, a facecam, backgrounds, and editor presets.
12:44 Captions run on device through Whisper with no cloud, and you export to MP4, WebM, or GIF. Download it and record your demo. Project number 20, Magai 2 Preview, 114B mixture of experts model for audio-video generation. Magai 2 Preview is an open-source 114 billion parameter model from Sand.ai that generates video with sound from text. It's a mixture of experts model that activates 6 billion parameters per token, and this repository is the inference code from a text prompt or a prompt plus an image.
13:20 It produces a 10-second clip with audio generated in two stages, a low-resolution pass, then a 1080p refiner. Weights are on hugging face, and it needs eight Nvidia Hopper GPUs. Download it and generate a clip. Thanks for watching. See you in the next update.