
Last week, I wrote about code getting cheaper to generate while the work of verifying it remains. The releases from September 29 through October 1 extend that argument. Orchestration itself is becoming part of the platforms engineering teams already use.
That’s useful. It also changes where I’d spend engineering time. As vendors take on more of the machinery for running agents, the organization still has to define what those agents may do and how anyone will know the work was good.
Getting agents to work together is becoming easier. Being accountable for their work remains a leadership responsibility.
The workflow is becoming part of the platform
On October 1, GitHub introduced dynamic workflows in public preview across the Copilot app, CLI, and SDK. These are code-defined processes that combine deterministic steps with agent work. Stages can run in parallel and pass structured results between them, with checkpoints for human review.
This is a useful direction for agentic engineering. A defined process gives you somewhere to put checks and handle failure. You can decide where a person needs to step in. It also makes reusable agent workflows a harder thing to sell as a standalone capability when they’re included in an existing subscription.
The complication is that an organization’s work rarely fits inside one tool. Consider teams using Copilot and Claude Code while the platform group runs Codex alongside internal automation. Each may have a sensible local workflow. Someone still has to own the rules across all of them.
I’d treat each vendor’s agent runtime as an executor within the organization’s process. That leaves room to change tools without rebuilding the approval model or losing the record of what happened. It costs more integration work upfront, so the boundary needs to earn its place. Start with the policies and records that actually need to survive a vendor change.
That’s the direction behind CTOsphere and its execution runtime, Sinfonia. Sinfonia implements OpenAI’s open Symphony specification, with the agent backend kept replaceable. The organization’s operating rules should hold even when its preferred coding tool changes.
Desktop access expands the responsibility
Between September 29 and October 1, OpenAI added capabilities including computer use and multi-agent execution to its managed Agents API. GitHub also added desktop computer use to Copilot, extending its reach into applications without an API or MCP integration.
That opens up work that has been difficult to automate. An agent can potentially handle an internal admin screen or a manual step in an older application. It also brings the permissions of those systems into the engineering workflow.
An agent working through a person’s desktop may inherit access far beyond what the task needs. A routine change can become an operational incident if the agent can reach production controls with the same credentials. The useful boundary is the task’s required access.
Before expanding that access, I’d want the organization to be able to explain:
- Which systems an agent can read and which it can change.
- Which actions require approval, and who can give it.
- How credentials are scoped to the agent and the task.
- What record is retained when an action succeeds, fails, or is stopped.
- How the team tests workflows that interact with changing screens.
GitHub’s application approval prompt and organization controls are useful starting points. The organization still needs to decide what approval means. Permission to use an application doesn’t settle whether an agent may change a production setting inside it.
A policy needs to be specific enough to enforce: a production configuration change requires approval from the designated owner, and the change and approval are recorded together. That rule can be reviewed and tested. It gives the person approving the work a clear responsibility instead of leaving them to interpret a dialog at 6 p.m. on Friday.
This is why human checkpoints and action-level policy are part of the CTOsphere control plane we’re designing. They affect how work runs from the beginning.
Access needs a supported foundation
On September 29, OpenAI expanded Sign in with ChatGPT so eligible Plus and Pro users could allocate plan usage across sixteen participating products. The broader direction matters: providers are creating explicit ways for other tools to act on a user’s behalf.
I’d build around those supported mechanisms where they’re available. Customer-owned API or enterprise credentials remain the practical option elsewhere. Either way, access should be replaceable without requiring a redesign of the system that depends on it.
That’s the posture behind Sinfonia’s community edition: customer-owned API keys, with provider support behind an abstraction. New authorized access mechanisms can be added there. Depending on an unofficial credential workaround creates a failure point outside the team’s control.
Model choice needs evidence from your own work
OpenAI’s September 29 release of GPT-6.1 Sol put its reported near-frontier coding performance at $2 per million input tokens, $0.10 for cached input, and $10 for output. Those benchmarks are a reason to evaluate it. Your own work determines whether it’s a better choice.
I still care more about cost per accepted change than cost per token. A cheaper run that leaves an engineer with two hours of cleanup can cost more than a run that gets the work right. Review time belongs in the calculation.
As models get closer in price and capability, choosing between them depends more on the record you’re keeping:
- Success on similar tasks and repositories.
- Human rework needed before acceptance.
- Review findings and defects that reached production.
- Total execution cost and time to an accepted result.
A benchmark can’t tell you which model handles your migration conventions or consistently misses your authorization rules. That knowledge comes from observing work in your own environment. If you aren’t recording the outcome, every model change starts with another round of opinions.
This is the reason outcome evidence has its own place in CTOsphere’s design. It should help a team make the next decision, including whether using an agent for that task is worthwhile at all.
Your engineering knowledge still needs engineering
AMD’s September 30 introduction of Ross, an assistant for embedded engineering, illustrates another part of this. The design connects agents to AMD toolchains through MCP. It also supplies expert-authored skills and trusted reference material, with tested examples to work from. Developers can use their choice of model and MCP-compatible client.
The part I’d pay attention to is the expertise around the model. A timing-closure workflow is useful because someone who understands the work has defined it and provided the tools and reference material to carry it out.
Engineering organizations have their own version of that knowledge. Architecture decisions, release procedures, migration rules, and the reasons an old service needs careful handling all affect whether a change is safe. Too much of that context depends on knowing which senior engineer to ask.
Putting it in a Markdown file is a start. Then you need to separate context from enforceable rules and decide which judgments still belong to a person. A rule about tenant isolation needs a way to detect a violation. A release procedure needs stopping conditions and a recovery path.
That’s work I’m building into CTOsphere’s knowledge layer and into how I scope agent adoption through O’Side Systems. An engineering agent pack should reflect the organization using it: its standards, repeatable workflows, internal integrations, reference tasks, and permission policies. It also needs an owner who keeps it current.
A model provider can improve the model. The organization has to maintain the knowledge that makes it useful there.
Start with what your agents can touch
I’d begin with a simple inventory. For every agent tool in use, write down 4 things: what it can read, what it can change, where its rules live, and what evidence it leaves behind. Include the scripts and unofficial workflows people rely on to get their work done.
Then choose one consequential workflow and follow it through. Check whether the permissions match the task and the approval reaches the right person. Someone should also be able to reconstruct the result afterward. Fix the gaps you find before expanding its access.
This is the work we’re pursuing with CTOsphere and Sinfonia, alongside my fractional CTO practice at O’Side Systems. If you’re working through it in your own organization, I’m happy to connect to see how we can help.
The tools will keep changing. Accountability for what they do stays with the organization.