
An agent can finish a change in minutes and leave it waiting for review for two days. The generation was fast. The delivery wasn’t.
That’s the problem I keep coming back to as I read this month’s AI announcements. Models are getting cheaper. Platforms are taking over more of the work required to run agents. But someone still has to establish whether the result is correct and belongs in the system.
For a team already producing more changes than it can review, cheaper generation adds to the backlog. The constraint has moved, and the way we measure progress needs to move with it.
On September 22, Anthropic released Claude Opus 5.5 at $4 per million input tokens and $20 per million output tokens. Anthropic reports roughly 40% lower cost on typical workloads than Opus 5, accounting for efficiency improvements. That’s the vendor’s finding; your workload still needs its own evaluation. Anthropic’s announcement.
OpenAI released GPT-6 Sol and Luna the same day. Their standard input/output prices are $2/$10 and $0.10/$0.50 per million tokens, respectively, for prompts up to 272,000 input tokens. That puts a 40-fold spread between Opus 5.5 and Luna on input pricing alone. OpenAI’s changelog.
That spread makes model routing worth doing. Give routine work to a cheaper model when it can meet the acceptance criteria. Use a stronger model where the task requires it, and revisit the choice with each release.
The mistake is stopping at the token bill. A cheaper run that needs another hour of an engineer’s attention may cost more overall. The rework still happened. It just landed in a different budget.
I’d track cost per accepted change: model and execution costs, plus the human time spent reviewing and correcting the result. Track elapsed time alongside it, including time spent waiting. Keeping those measures separate makes it easier to see whether the problem is expensive work or work that isn’t moving.
This is why Sinfonia, the runtime at the core of CTOsphere, treats the agent backend as configuration. A change in model pricing should lead to a policy update. It shouldn’t require rebuilding the platform.
The software around the model is changing, too. On September 10, OpenAI released its Agents API in public beta, with durable sessions, context compaction and recovery. It supports tool connections and either hosted or customer-supplied sandboxes. More of the machinery needed to run an agent is becoming a platform feature. OpenAI’s changelog.
Atlassian’s September announcements point in the same direction: agents working from Jira, shared code context, organizational standards and automated review. Availability matters here. Agent loops, Standards and AI Review are still in private early access. Atlassian’s announcement.
I’m glad the basic execution loop is becoming easier to buy. There is limited value in every team rebuilding session recovery. The part I’d examine carefully is where the team’s knowledge ends up.
Engineering standards and approval rules accumulate over time. If they only work inside one vendor’s configuration, moving away means reconstructing how the organization operates. That cost won’t appear on the pricing page. Before adopting a platform, I’d want to see how those rules can be exported and used elsewhere.
CTOsphere is being built around that concern. Sinfonia is a Rust implementation of OpenAI’s open Symphony specification, extended for multiple projects and execution across a cluster. Conformance gives it a common contract with the reference implementation. The extensions still need to be considered when moving between implementations.
Partitura, the knowledge plane we’re developing, takes the same approach to standards and decisions. The intent is to declare them once, then use that source both to guide agents and to check their work. Keeping a document for people and a separate set of rules for machines creates two things to maintain. Eventually, they disagree.
Governance has a similar gap between what a tool supports and what a team actually operates. On September 22, GitHub added centrally managed OpenTelemetry configuration for Copilot app sessions. The next day, it previewed local sandbox controls for file access, networks and credentials. GitHub’s telemetry announcement and sandbox announcement.
The sandbox is off by default. Turning it on for a project applies to new sessions; an active session can be enabled separately. Those details matter when someone assumes the rollout is protected because the feature exists. An available control only helps once it’s configured and enforced.
For an agent rollout, I’d expect a named owner for permissions, a versioned record of those permissions, and evidence that the restrictions hold. Test whether the agent can reach something it shouldn’t. Keep the telemetry needed to reconstruct its actions when something goes wrong. Written restrictions need evidence of enforcement.
That’s the premise of the control plane we’re designing for CTOsphere. Fleet configuration and permissions become declarations that can be reviewed and versioned. We’re planning the same for telemetry configuration. The record needs to connect what was allowed with what actually ran.
The September announcement I find most useful is less dramatic than a new model. GitHub’s Copilot metrics API now separates review time into three stages: ready for review to first review, first to final review, and final review to merge. It reports the median and 90th percentile for each. GitHub’s review metrics announcement.
There is a limitation: this release covers human-authored PRs reviewed by another person. Bot reviews aren’t timed, and the data isn’t backfilled. It won’t measure the entire agent workflow for you. It does make the places where human review waits easier to distinguish.
That distinction changes what you fix. Work waiting for a first review may need clearer ownership or protected review time. Repeated review cycles may point to weak requirements or missing context. Work sitting after final review needs an explanation of its own.
If review is already at capacity, increasing the rate of incoming PRs grows the queue. Counting generated PRs can make an AI rollout look productive while the time to production gets worse. The engineers doing the review then inherit the pressure to keep up.
I’d start by reducing avoidable review work. Check architectural boundaries and dependency rules while the agent is still working. Give it clear acceptance criteria. Keep changes small enough to understand, and limit work in progress when the review queue is full.
Human judgment still matters. A passing check can’t tell you whether every product decision was sensible or whether a tradeoff is acceptable. It can keep reviewers from spending their attention on the same mechanical violations in every change.
This is the direction behind CTOsphere at O’Side Systems: a portable execution runtime, with knowledge, control and observability capabilities being developed around it. It’s also the kind of operating problem I help teams work through in fractional CTO engagements. If you’re working through it, reach out.
You can start without adopting another platform. Pick five to ten representative tasks from your backlog and run them through two or three models under the same acceptance criteria. Record acceptance, review time, rework, total cost and elapsed time. Include the failed runs in the accounting.
Use that small sample to decide what to test next. If generation gets faster while accepted work takes longer, invest in the part of the system where work is waiting. The value of faster code generation depends on the team’s ability to turn it into software people can rely on.