Upcoming Workshop: Citizen Developer Agentic Software Factory on AWS
Register Here!Agentic SDLC and Software Factories: The Path to Production
Agentic SDLC and software factory guides cover planning through code review. Here is what agents need to release and operate infrastructure safely.
I've read a lot of agentic SDLC and software factory writeups over the last year. AWS has the AI-Driven Development Life Cycle. Microsoft has Agentic DevOps. Dan Shapiro's five levels run from autocomplete up to a "dark factory." StrongDM published how its software factory works, under two rules: "Code must not be written by humans" and "Code must not be reviewed by humans." Factory rebuilt its product around the idea, and Port published a lifecycle rebuilt around agents.
The SDLC posts describe the stages, and the factory posts describe the system that runs them. Most of them draw the same loop: plan, build, test, review, release, operate, with signals from production feeding the next plan. They mostly agree on the human's new job too. People set intent, write the specs and scenarios, and decide what ships.
They also share a gap. The stages up to the merge get detailed mechanics: how specs get written, how agents pick up issues, how test scenarios are held out so the agent can't game them, how review gets scaled or removed. The stages after the merge get less, and what they get is mostly about shipping application code through a pipeline that already exists. Port's release agent scores the risk of a change and manages the rollout, and its operations agent ties an alert to the last deploy and drafts a rollback. How an agent changes the infrastructure underneath, and what credentials it holds while it does, gets about a sentence. In AWS's version, operations is "AI applies the accumulated context from previous phases to manage infrastructure as code and deployments, with team oversight." Factory's loop has changes that are "built, tested, reviewed, secured, shipped, and monitored," and shipping gets no more detail than that. StrongDM's Digital Twin Universe clones third-party services like Okta and Slack so its agents can "test failure modes that would be dangerous or impossible against live services." That's a good idea for SaaS dependencies. It doesn't cover the AWS account the software runs in.
Port's writeup is the one that names the credential problem: "When a person needs production access, there's a scoped, logged process. An agent built locally just inherits whatever credentials its creator wired in, usually with no review."
That describes most of the agent setups I see. The agent gets whatever the developer's shell has. Sometimes that's a read-only role. Often it's an admin profile left over from last quarter's incident.
I've spent my career on the infrastructure under that loop, so that's what this post covers. Two parts of a software factory make it workable for agents: a context engine that tells the agent what exists and what it's allowed to do, and a platform orchestrator that gives it one governed way to change the cloud.
The repo does two jobs for code agents
Code agents got good fast partly because the repository hands them almost everything they need. The repo is the context: source, tests, specs, history, and conventions, all readable and searchable. The repo is also the action surface. The agent writes to a branch and opens a pull request, and nothing it does is real until someone merges. CI runs in a sandbox. A bad commit costs a revert.
Infrastructure has neither. The context is spread across Terraform state in a dozen repos, cloud APIs, Kubernetes, a CMDB that was accurate in 2023, and the heads of three people. The action surface is a credential. When an agent runs terraform apply or aws rds modify-db-instance, the change is real as soon as the API call returns, and the only review was whatever the agent decided before it made the call.
That's why infrastructure stays thin in these writeups. It's the part of the stack where the agent has no repo to lean on. A context engine supplies the context half. A platform orchestrator supplies the action half.
A context engine answers questions about the running system
Anthropic describes context engineering as "the set of strategies for curating and maintaining the optimal set of tokens (information) during LLM inference." For infrastructure, the tokens you want in front of the agent are answers. What's running in this environment, what depends on it, which module version produced it, what changed last, and what the rules say about changing it. An agent working against raw cloud APIs has to assemble those answers itself.
Take a simple request: resize orders-db, because checkout latency is bad. Before it touches anything, a careful agent wants to know which services connect to that database, whether any of them belong to another team's project, which instance sizes are allowed in production, and whether a resize needs someone to sign off. Against raw APIs, that's a describe call per service, a pull of every state file that might reference the database, a crawl of Kubernetes manifests for connection strings, and then a guess. A consumer that reads the connection string from a secret another team's pipeline injected won't show up in any of those. The approval rule doesn't show up anywhere the agent can reach.
Massdriver keeps that model current as a side effect of how deploys happen. Every piece of infrastructure is an instance of a bundle, which is a versioned IaC module with a JSON Schema for its inputs. Instances connect through resource types, the versioned contracts for the data one component passes to another. The graph knows orders-api consumes orders-db because that connection is how orders-api got its database credentials in the first place. Environments, deployment history, policy results, alarms, and cost data hang off the same graph, and agents reach it through the MCP server, the GraphQL API, or the CLI.
Some of the questions it answers in one call:
- Which instances depend on this one, across projects
- What differs between staging and production, down to bundle versions and parameters
- Which inputs the agent may set on this bundle, and the valid values for each
- Why a policy check failed, in terms the agent can act on
- Which component an alarm belongs to, and what sits downstream of it
- What the last deployment changed, and what its logs said
The schema one matters more than it looks. When the agent knows the parameter space ahead of time, it doesn't propose an instance class your platform team never published or an engine version you don't support. If it tries anyway, the proposal fails validation before any plan runs. That's a lot cheaper than catching the mistake in review or in a failed apply at 2am.
A platform orchestrator is the only way in
Context tells the agent what it could do. The agent still needs a way to act that doesn't involve handing it cloud credentials.
A platform orchestrator sits between intent and the cloud. Developers, pipelines, and agents all ask it for changes, and it decides which IaC runs, with which inputs, under which policies, and with whose approval. In Massdriver, cloud credentials are attached to environments inside the platform, and each deployment runs in an ephemeral pipeline Massdriver generates for it. The agent authenticates with a scoped Massdriver token. It never holds an AWS access key, a GCP service account file, or an Azure client secret.
Here's what happens to the orders-db resize.
The agent proposes a deployment: this instance, this bundle version, with instance_class changed from db.r6g.large to db.r6g.xlarge. The orchestrator validates the parameters against the bundle's schema and runs the policy checks your team configured, such as Checkov and OPA. It checks whether the agent's token may act on this kind of resource in this environment. Massdriver's access control is attribute-based, so the rule reads like "agents may manage databases in non-production environments and propose changes in production," with no list of resource IDs to maintain. Then it checks the environment's approval rules.
In dev, no approval is required. The change applies, and the agent can watch the database come up and run its load test against it. In production, the proposal waits. With separation of duty turned on for that environment, the account that proposed a change can't approve it, which matters when the proposer is a service account a dozen agents share. A person reviews a two-parameter diff against a module the platform team already vetted and approves it, and the same kind of pipeline that ran in dev applies it to production. Every step lands in the audit log with the identity that took it.
The agent has no way around this. It can't run terraform apply against production because it has nothing to authenticate with. There's also no separate agent path to secure and keep in sync with the human one. The portal, the CLI, the API, and the MCP server all go through the same catalog and the same policy.
Several of the vendors above describe an orchestration layer for agents too. For infrastructure, the questions to ask of any of them are where the cloud credentials live and whether an agent's change can reach the cloud any other way.
What changes after the merge
Agents test against real infrastructure
StrongDM built behavioral clones of SaaS APIs because testing against the live services was too risky. For infrastructure you can give the agent the real thing. Massdriver can fork an environment from production's blueprint, so the agent gets the same topology and the same bundles at lower scale, with credentials that only reach a sandbox account. The agent deploys, fails a compliance check, reads the finding, fixes it, and redeploys. I covered that pattern in more detail in Disposable Environments.
Production review gets smaller
A production change arrives as a proposal against a module your platform team already reviewed. The reviewer looks at the parameter change, the policy results, and the dev deployment that already ran. That's a much smaller review than seventy files of HCL an agent wrote from scratch, and review size decides how much agent output a team can actually ship. Harness puts delivery at 60 to 70 percent of engineering time and says "AI is producing more code than teams can reasonably ship using manual workflows." Shrinking each production review is how you move that number.
Credentials and blast radius live in one place
When every cloud change goes through the orchestrator, an agent has no reason to hold cloud keys, so they never land in a context window, a transcript, a log line, or a prompt injection's payload. Revoking an agent's access means revoking one Massdriver token. What the agent may do is a policy on classes of resources in classes of environments, and it sits above every cloud and tool you run. I made the case for why that has to be attribute-based, and not per resource, in Your Golden Paths Weren't Built for Agents.
Operations covers more than incidents
Most lifecycle writeups treat operations as incident response: an agent reads the alert, correlates it with the last deploy, and pages someone. A graph helps with that, since the alarm is attached to a component with known dependents and a deployment history, and the fix is often a rollback to a known-good deployment the agent can propose. The larger gain is day-two work that never reaches the incident channel, like bundle upgrades, engine end-of-life deadlines, and resource type migrations. An agent that can read the graph sees which instances are behind, tries the upgrade in a forked environment, and hands an operator a tested proposal. Without that, the upgrade sits in a backlog for another two quarters.
Autonomy levels depend on the platform
Shapiro describes level 3 as the stage where you're "a manager" of the AI's work, and says "almost everyone tops out here." For application code, the way past level 3 runs through better tests and holdout scenarios, so the agent's output can be judged without a person reading every diff.
Infrastructure doesn't get past level 3 that way. The risk sits in credentials and in changes you can't take back, and a better test suite doesn't change who can call the cloud API. What moves infrastructure up the levels is the platform around the agent. If every change is a schema-validated proposal against a vetted module, checked by policy, gated by per-environment rules, and recorded, you can let agents run unattended in dev and staging now and spend human attention on production approvals. As confidence grows, you drop the approval requirement where the risk is low and tighten access policy where it isn't. Your platform team sets that pace by writing policy, and the model's capability is only one input.
Massdriver is a context engine and platform orchestrator for this part of the lifecycle. Your platform team publishes the Terraform, OpenTofu, and Helm it already trusts, and agents deploy from that catalog through the same policy and approvals as your engineers, without ever holding cloud credentials. Book a demo and we'll have a production-ready self-service platform for you to pilot in less than 24 hours.
Build one in a free workshop
In Citizen Developer Agentic Software Factory on AWS, you'll set up this pattern in your own AWS account from a Massdriver canvas: an open-weight model on Amazon Bedrock, a locked-down VM sandbox running the opencode coding agent, and an app platform with a GitHub repo that builds into ECR plus DynamoDB, S3, and Lambda. Then you'll write an app with the agent and have it set up the app's infrastructure through the context engine. The agent never holds AWS credentials, and every change it makes is policy-checked and limited to dev.
You'll need an AWS account (expect about $0.50 in usage), a Massdriver account with an AWS role connected, and a GitHub account. You don't need Terraform experience. Save your spot.


