Building Reliable, Affordable AI Workflows

Useful AI depends on deliberate training, customer feedback, automated validation, and cost-aware model routing.

Summary

Useful AI work requires more than handing people a model: Automattic gives employees dedicated learning time, while Sonar turns validation into an agentic pipeline. Production reliability comes from customer-defined success, instrumentation, direct conversations, and repeated diagnosis rather than static benchmarks, as Nova Act’s development illustrates. Model economics add another lever: Haiku 5.5 offers low-cost, high-volume work and performs especially well as a sub-agent. The emphasis differs across workforce fluency, product reliability, software quality, and model cost, but each area makes the surrounding system decisive.

The videos

Automattic builds AI fluency through two-week programs split between workshops and real projects, with seven cohorts of about 45 people completed and a goal of reaching roughly 600 employees, including about two-thirds of its engineers.

Nova Act moved from a March research preview to an AWS service in December by replacing static benchmark confidence with an eval flywheel based on customer-defined success, instrumentation, diagnosis, and repeated prioritization.

Sonar’s Gitar agent automates pull-request review, CI diagnosis, flaky-test handling, autofixes, and merging, while SonarQube’s taint, control-flow, data-flow, and software-composition analysis strengthens quality and security checks.

Haiku 5.5 costs $0.10 per million input tokens and $0.50 per million output tokens, more than doubles Haiku 4.5 on most benchmarks, and works best as a cheap sub-agent for searching, categorizing, auditing, and testing.