Tool Open source
Exploitarium is a curated collection of public proof-of-concept exploits and vulnerability-research writeups maintained as a GitHub repository. It organizes former standalone PoC repositories into self-contained folders that preserve their original READMEs and tracked files, while new research entries are added directly to the archive. The repository documents a consolidation process that verifies paths, Git object types, tree modes, and blob IDs for migrated files, and states that the material is intended for good-faith, open-disclosure research rather than malicious use.
claude-red is a curated library of offensive-security skills for Claude, maintained by SnailSploit. Each skill is a structured SKILL.md file covering a particular attack surface, including web vulnerabilities, identity systems, wireless, cloud, mobile, IoT, exploit development, fuzzing, reconnaissance, API security, containers, cryptography, privilege escalation, post-exploitation, and AI security. The files are intended to be installed into a Claude Skills environment, where matching skills load on demand from conversational triggers, or they can be supplied manually as system files or project prompts. The repository includes an installation script with category and target options, a skill index, a roadmap, contribution guidance, and a review-oriented skill template. It is distributed under the MIT license and lists authorized red-team engagements, bug-bounty triage, security research, CTF preparation, and operator training as use cases.
Codex Security is an OpenAI CLI and TypeScript SDK for finding, validating, and fixing security vulnerabilities in authorized code repositories. It scans directories and codebases, including repository changes, uses AI-assisted analysis to trace attack paths and apply custom threat models, validates findings, and can generate reviewable patches and exports in SARIF, JSON, and CSV formats. The CLI supports CI use with an API key and can stop builds at a selected severity threshold. The package also supports configurable deep scans, worker and subagent counts, time limits, containerized bulk scans, and a preview findings service. The service stores findings and embeddings in SQLite, provides paginated listings and a read-only dashboard, identifies potential duplicates through embedding similarity, and allows completed findings to be published and deduplicated through the CLI or SDK. It requires Node.js 22.13.0 or later and Python 3.10 or later, and can use OpenAI or other documented inference providers.
Harness Engineering is Ryan Lopopolo’s anthology, field guide, and agent-context repository for improving coding-agent output without changing the underlying model or coding agent. It shapes the surrounding environment through context and tools so an agent can recover intent, operate the real system, respect authority, prove its outcome, and incorporate lessons from accepted work, corrections, failures, and user responses. The repository uses AGENTS.md to route agents to relevant arguments, cases, and proof, and provides a thesis index, playbooks, source material, and executable constraints for encoding organizational requirements and decisions. Repository-authored material is licensed under CC BY 4.0.
OpenResearch is a local-first workspace for research agents and autoresearch, available as a desktop application and a macOS/Linux CLI. It turns Claude Code, Codex, or OpenCode into agents that can review literature, develop hypotheses, run experiments, and produce research artifacts. The workspace assigns independent agent sessions and isolated Git worktrees to parallel research directions. It tracks experiment variants in a Git-native experiment tree, archives each run against its recorded commit, and keeps logs, diffs, files, results, and artifacts tied to the work that produced them. Its autoresearch loop can propose an idea, modify code, launch an experiment, inspect evidence, and choose the next direction. The same committed source snapshot can run locally or through SSH, Slurm, Kubernetes, Ray, Hugging Face Jobs, Modal, Tinker, or managed OpenResearch compute. By default, projects and run data remain on the user's machine in a local SQLite store, with a browser dashboard served on localhost. An OpenResearch account is used for service-owned capabilities such as organizations and managed compute. The CLI also installs an OpenResearch skill into supported coding agents and provides commands for projects, runs, logs, experiments, discovery, and paper lookup.
Better Harness is an open-source Harness Engineering platform for coding agents. It analyzes project and, where supported, session evidence around an agent's work rather than only its final code diff, evaluating the Agent Work Loop for issues such as task understanding, validation, delivery control, and retained lessons. It turns supported gaps into prioritized findings linked to their evidence, expected outcomes, repair boundaries, and acceptance checks, and produces host-specific reports in formats including HTML, paired Markdown, or native Canvas reports. The platform also supports defining harnesses as code, running controlled experiments, inspecting evidence, and comparing outcomes across coding-agent hosts such as Claude Code, Codex, Qoder, Cursor, and GitHub Copilot CLI.