A first, low-risk slice of the software factory: an agent that reads monitoring alerts (e.g., Sentry), pulls traces and surrounding context, and posts a diagnosis to Slack for engineers to review and act on. Builds trust incrementally before any broader automation.
Factory is a reference software factory for Claude Code and Codex that installs a repeatable, version-controlled software-delivery workflow into an existing GitHub project. GitHub Issues serve as the work queue, while committed policies, skills, labels, handoff comments, pull requests, and run records preserve state between fresh agent sessions. Scheduled agents triage issues, route them to implementation, specification, questions, or blockers, claim bounded work, implement it on a branch, run configured type, lint, test, build, audit, and architecture checks, obtain independent verification, and open draft pull requests. An independent verifier reads the diff and checks that the new test fails without the implementation. Humans retain responsibility for ambiguous requirements, system design, significant changes, and merging pull requests. The repository has no custom orchestrator or queue service. Claude Code routines provide the default scheduling and compute, with a thin Codex adapter using the same policies, gates, and evidence files; GitHub events or an optional API-triggered Action can also start runs. It is configured through files such as a human-owned charter and gates script, and includes a /factory control room backed by live issues, pull requests, and run records.
Sentry is a developer-focused application performance monitoring, error-tracking, and debugging platform. It helps identify and trace issues in real-world applications through official SDKs for JavaScript, Python, Ruby, PHP, Go, Rust, Java/Kotlin, C#/F#, C/C++, Dart/Flutter, and other platforms and languages. Its agent-focused features and Sentry MCP expose traces, request and token costs, timelines, chat transcripts, errors, and the full request pipeline so large-language-model agents can investigate and debug production issues. The project is developed by Sentry and is available as an open-source repository with a fair-source topic designation.
Searchable transcript of The Software Factory: From Bug Report to Production Code — Davis Palmie, Factory — AI Engineer (18:32). Search for a phrase, then click its timestamp to jump straight to that moment in the video.
Captions sourced from the original video on YouTube, published by AI Engineer. The video, its captions and all related intellectual property remain the property of their respective owners; AINotes claims no ownership. Provided for research, accessibility and search — see the Transcript Notice and Copyright Policy.
00:01 [music] Hey folks, I'm Davis Palmy, a member of technical staff here at factory. Uh before this I co-founded Lummetric which was an agent platform for investment teams that we took through Y Combinator and was acquired by factory and prior to that I was a engineering lead for Slalom's innovation lab. So I've seen AI deployments both across large enterprises and more AI native orgs.
00:41 With that in mind, today I'm going to be talking to you about the software factory and how you can move towards an autonomous system that translates input signals into production code. So, first we've moved through three distinct eras of AI engineering. First, we had tab autocomplete, next token prediction, and the engineer was still firmly in the driver's seat.
01:11 Then AI could generate entire files and the software engineer became a bit of a discerning copy pastaster. Um, but they still heavily had control over the output of the AI. And then you had agents that were more tightly coupled with the code. They could gather their own context, call tools, and debug their own output. Now, autocomplete used to feel magical, and now at best it's it's a quaint thing we remember.
01:37 You want to look at the speed at which AI engineering is climbing the abstraction ladder from tokens to files to context now to entire systems. Every shift you have here is an adjustment. But each move up the abstraction ladder has made engineers more valuable, not less. And so now we're at a transformational point. Engineers are going from typing code to governing agents and eventually you'll organize those agents into a system.
02:16 That's your software factory. So with this the engineer's role shifts to keeping guard rails tight, catching code drift and setting the highle priorities and direction. But meanwhile, the system is ingesting input signals like bug reports or user feedback and outputting production code. So if you have AI systems that are producing more code than humans, there are some big questions that everyone is rightly asking.
02:50 First, how can I prevent massive cost overruns? We've all seen the headlines about Uber's CTO blowing through the years, AI budget by April, or Microsoft clawing back clawed code licenses to try and cut down on spend. At Factory, you're not locked into a single model provider. And this means you unlock the entire spectrum of cost performance with LLMs.
03:21 You want the maximum uh cost efficiency and the minimum effectiveness necessary for each task. As our CEO likes to say, uh if your kid needs a algebra tutor, you can definitely find someone cheaper than Albert Einstein. Second point, how can I ensure the best access to whatever model is top right now? Well, we think that having to back the right horse is a false choice.
03:49 Um, especially with the increasing effectiveness of open source models. We're strong proponents of being able to choose the right tool for the right job or even now have it dynamically routed. Number three, what happens to software engineers? What does the future of software engineering look like? Well, the role is obviously evolving and the line between product and engineering is blurring.
04:19 Engineers will now build and maintain the system that's responsible for building your production code and they'll set the highle direction for the software factory. And lastly, how is this going to change the organization? Well, team boundaries are also starting to blur and context needs to flow seamlessly across these boundaries. Information can't be siloed.
04:45 That was true for humans, but it's especially true for AI agents. Everyone across all your teams will have a shared responsibility for the pillars of the software factory. Now, coding was the easy part. This actually was never the bottleneck. Your context is laid out in natural language. I bet if you look at your real engineering metrics, what's the median time to review a PR?
05:15 How long does it take you to reproduce user bug reports or go through and debug code? Who keeps your docs current with every PR and who keeps it current with every strategy meeting? Who maintains your test suites? oftentimes the testing teams and the engineering teams and the product teams are each on their own continent. If you pull your dora metric, the time from code written to time in production, I would bet dwarfs the time actually making those engineering changes.
05:50 The real need you have today is cohesion across all of these surfaces. Right now, docs will go stale, dead code lingers, and there are so many layers of communications between your different teams. Again, this is hard for human engineers, but it is especially more pronounced for agents. The necessary undertaking to build your factory is to make sure the agents can bridge all of these services and understand the nuance within your organization.
06:24 So this brings us to the software factory which is the system of agents that will take you from input signal to production deployment. Your agent's not just coding anymore. It's taking on customer support, product engineering, deployment ops, and many other roles. Across these pillars, your agent should be triaging incidents, making plans, creating and updating docs, actually executing the code, reviewing it, testing it, rolling it out, and then monitoring that release.
07:03 Now, that's obviously a ton of work, and you don't want to jump into that all at once. For one, it's too heavy of an undertaking, but for two, your teams will not have time to build trust in this system. You want to start with something narrow and verifiable like incident triage. Let your agent read a sentry alert, pull the traces in other context, and then post in Slack.
07:29 But keep it so your engineers can review this and decide what to do. And as more correct diagnoses roll in and in from the agent, your team can start to build trust and an understanding of this pillar. So each step in the software factory needs careful ownership and monitoring before you even think about letting the the pipeline loop. But once you can tie the output of this system back to the input signal, that's when you unlock organizational self-improvement.
08:07 Now we feel pretty strongly that your software factory should have a few key principles. First, it should be model agnostic. Obviously, different models excel at different tasks and with very different cost efficiencies. We found, for instance, that you can achieve the same performance on code review with GPT 5.2 as with the latest Opus models, but at half the price.
08:32 And now if you look at the open- source alternatives that can slash your cost down to 10 to 30x cheaper. So at factory again you have the entire interplay of speed to task efficacy to cost. Not being locked into a single lab means you unlock the entire PTO frontier of LLMs. Now number two your deployment model should be sovereign. You shouldn't have to compromise on your security architecture or your policies to try and fit your AI deployment model in.
09:13 Now at factory you can choose your workspace and your deployment model. We support everything from fully managed to bring your own machine to entirely airgapped. And number three, your software factory needs to be integrated across the SDLC. These are no longer separate uncoupled concerns. Docs changing should inform your code and vice versa. Planning goals should inform what you test for, not however many communication layers there are between strategy, product, engineering, and your testing teams.
09:50 Each one of these pillars supports the software factory. And weakness in one means a weakness in the entire system. So tactically in an AI native or what does this look like? Well, we typically see it as the fusion and the incremental automation of many sub teams under the umbrella of one software factory. Some really common patterns we see are triaging incoming signals like sentry alerts, having the agent make diagnosis before the human can even answer the page.
10:30 Proactively testing security, scanning for secrets or vulnerabilities. Creating, maintaining your tests, monitoring the ones that are becoming brittle or unneeded. Again, keeping docs in sync with code. keeping docs in sync with your different meeting channels and then the classic reviewing code, commenting on PRs, rollouts, monitoring releases and with the agent doing all of this, the human engineer's role shifts to responding to exercising judgment and correcting the AI in such a way that it can learn from these
11:10 mistakes. Now everything is transitive across sub teams. An improvement to one group means an improvement to the entire system. So why should your enterprise care about a software factory? Well, the larger and the more fragmented you are, the more you feel the linear pain of managing teams. In tech terms, your software factory should take this pain from O of N to O of one conservatively.
11:45 You can define your guardrails, your agents, your docs, testing policies, really governance just one time and then standardize it across the organization. And this standard should make sure that everyone is pulled to the highest bar, not the lowest common denominator. Now also if agents can move across these surfaces then they can hold that quality bar across teams without you having to linearly scale your time to manage each one.
12:19 And if agents are moving across all these services then you have one place where context can be unified and you can have executive visibility into how the entire SDLC is performing. Now, just like model agnosticism removes lock in, your agents should also be surface agnostic. They need to be highly available to your teams. Regardless of what tools they use, your teams should not have to compromise on their processes or their tooling to try and fit your AI provider in.
12:56 At Factory, we have agents available everywhere from remote machines to the CLI to Slack, any tool your teams use. And even though humans and agents may use different services, everything should share the same underlying harness and context. So the undertaking of building the factory will expose tribal knowledge, manual processes, and outdated info.
13:28 Unblocking these is not only going to help your agents, but it's also going to help your engineering teams. At factory, we have the concept of agent readiness and we evaluate this based on eight pillars. The first is validation. Do you have guard rails like llinters and formatterers in place? Then your build system. Are your build commands and CI well documented?
13:56 Feedback loops. You have unit and integration tests tight and you have docs like readmes, agents.mmd files. Are your dev environments cleanly reproducible such that an agent can use them? Is your code modular with clear boundaries and rules against sprawl? Do you have observability? How long does it take you to find out why something went wrong? and are you proactively security scanning so that you don't have to worry about leaked secrets or vulnerabilities.
14:30 Basically, the key takeaway here is that if your agents and your human engineers are going to walk the same roads, paving them well is doubly useful. Now, you can't govern what you can't see. And that's true for agents and human engineers. You can start with this agent readiness assessment, but no matter what, your engineers must monitor agent performance and control behavior.
15:03 At factory, every action is auditable. It operates under your standards with role-based access control and least privilege. For example, your teams that are building this software factory need to first build trust and understanding. That's why you should roll this out incrementally. You want to automate the individual pillars first before you turn this system into a loop.
15:31 And again, precise execution does not prevent poor design. It never has. You still heavily need humans in the loop for architecture, for strategy, and for setting direction. Engineers will not be replaced, but judgment and prioritization now beat mechanical implementation skill. Agents won't be able to handle everything. There will be guardrails to tighten, code that drifts.
16:01 Humans need to stand as validation gates in this factory. But just like we don't have to feed punch cards into the machine anymore, code does not have to be written by hand. Your team should be able to set the strategy and their risk parameters and then let agents execute. And knowing if your agents are executing well comes down to measuring the right things.
16:30 Basically, you want to measure outcomes, not tokens. Lines of code and tokens generated are gameable metrics. These are not actually associated with outcomes. Token leaderboards at big companies just encourage proflegate spending. It's goodart's law in real time. If you decide to measure token usage, engineers are going to optimize for token usage and then it's no longer a good metric.
17:03 It's not about who spends the most on tokens. It's about who can get the highest leverage outcomes from that spend. Some metrics that we like at Factory are signal to production time, human intervention count, median time to repair, shelf life of code, and cost per PR. These are real tangible outcomes. So agent readiness is really a precursor to effective engineering.
17:35 Whether it's an agent or a human that ends up in that system, your agents need the exact same careful governance as your engineers, tooling, access levels, systems for review. Don't just turn them loose in your codebase. And don't token max. Measure what actually matters. Your end users don't care about how much you're spending on AI. They care about feeling the dramatic improvements that will come from your software factory. Thank you. >> [music]