Infraze · agentic DevOps · under development
The engine is a dial. Not a dependency.
An infrastructure agent that runs in your terminal, routes each step to a cheap or an expensive model, and remembers what it built. It owns neither your cloud, nor your weights, nor your data.
In active development · early access by conversation · architecture published in full
- Status
- in development
- Form
- a command line interface
- Engines
- Claude Code, or Kestrel
- Memory
- Trove, in your repository
- Targets
- cloud, CI/CD, Grafana
- Hosted control plane
- none
- Your credentials
- stay yours
- Early access
- by conversation
The problem
The agent that provisions your cloud should not be selling it to you.
There is no shortage of agents that provision infrastructure. Amazon shipped the AWS DevOps Agent in March 2026. Datadog has shipped one. StackGen sells provisioning from intent. Clanker is an open command line agent across every major cloud, and a curated list counted 474 tools in this category by July 2026.
They are competent inside the business model that produced them, and that model has a consequence. An agent built by a cloud vendor optimises for that cloud. An agent built by an observability vendor routes you through that product. An agent sold as a service keeps your infrastructure state on someone else's servers.
Infraze is proposed for the case where you want none of that: a local binary, your credentials, your repository, and no opinion about which cloud wins.
The argument
Infrastructure work is not uniformly hard, so why is the model fixed?
Designing a multi-region failover topology is a judgement problem. Bumping a chart version, adding a label, wiring a panel to a datasource that already exists: those are mechanical. Today both go to the same expensive model, because the model is chosen once at the start of the run.
Claude Code
The expensive engine. Fast and reliable, and the right place to send anything that needs judgement or carries risk.
- Network design
- topology, failover, identity and access boundaries
- First provisioning
- the plan nobody has written yet
- Anything destructive
- and only after a human approves it
Kestrel, on open weights
The cheap engine. Suited to the steps that are repetitive, declarative and easy to verify by running them.
- Dashboards and alerts
- declarative, small blast radius
- Pipeline configuration
- file generation you can test
- Mechanical edits
- versions, labels, wiring
The routing is explicit, logged and set per step class, not guessed per token. Model confidence is used only to escalate a step upward, never to send one down, because a mis-routed step that drops a database is not recoverable by apologising. Anything destructive stops at a human. We have not finished benchmarking the open-weight model behind Kestrel, so it is not named here: it will be published together with its measured cost per completed step, not before.
Memory
The second visit should be cheaper than the first.
Most agents in this space are stateless between sessions, or keep their state in the vendor's store. Infraze is proposed with Trove underneath: what was built, why, which engine built it, and what broke afterwards, written as files inside your repository where you can read them and so can the next engineer.
What gets written
- Plans
- the intent, and the step graph it became
- Runs
- which engine, what it did, what it cost
- Approvals
- what stopped, and who said yes
- Failures
- what broke, and what fixed it
Why files
A memory you cannot read is a memory you cannot audit. Plain files in the repository survive the tool, travel in the pull request, and work when Infraze is not running. They also make it possible to answer the question that matters after an incident, which is not what happened but why someone decided that.
Scope · in order of confidence
Monitoring first, cloud last.
starts hereMonitoring
- Grafana dashboards, datasources and alert rules.
- The artefacts are declarative and the blast radius of a mistake is a wrong graph.
- Repetitive enough that the cheap engine should carry most of it, which also makes it the first real test of the dial.
secondContinuous integration and deployment
- Pipeline configuration, build and deploy steps, promotion between environments.
- Moderate risk, mostly file generation, and verifiable by running the pipeline.
- Approval gates arrive here, measured by how many were hit rather than by impression.
last, and gatedCloud provisioning
- Networks, compute, managed databases, identity and access.
- Highest value and highest risk, which is why it goes last.
- Every destructive operation stops at a human, and every path ships with a written rollback.
Where it stands
What is running, and what we are building.
Three of the pieces are already running in production work. The layer that composes them is what we are building now. We would rather show you the seams than imply a finished product.
Running today
- Kestrel
- the cheap engine. kestrel.velofy.co
- Trove
- the memory layer, open source and in use
- Claude Code
- the reliable engine, third party
In build
- The planner
- intent to step graph
- The router
- the dial itself
- The tool layer
- Terraform, kubectl, Grafana
- Published benchmarks
- when the numbers are in
The build plan · written up front, so it can be checked against the result
Four milestones, each with a number attached.
Days 1 to 21 · monitoring only
The CLI skeleton, the Trove schema and a tool layer for the Grafana API. Exit condition: a dashboard and an alert rule stood up from an intent, twice, in a clean project and then an existing one, with the second run demonstrably reading the first run's memory.
Days 22 to 49 · the dial
The second engine and the step router. Exit condition: a published table of cost and wall-clock time per completed step for the same plan run entirely on Claude Code, entirely on Kestrel, and routed. If routing does not come out cheaper at equal completion, we publish that table too.
Days 50 to 70 · pipelines and gates
Exit condition: a pipeline generated and green, with every destructive operation having stopped at a gate, counted rather than estimated.
Days 71 to 90 · one cloud path
Exit condition: an environment that comes up and gets torn down in a disposable account, with a Trove record sufficient for a second engineer to understand what happened without reading the Terraform.
Day 49 is the one that decides the product. Whichever way the cost table falls, it gets published with the numbers in it.
Questions
Reasonable objections.
Can I use it today?
Not yet, and we will not pretend otherwise. Infraze is in active development and early access goes out by conversation rather than a signup form, so we can shape it around real stacks. Book a demo and we will show you where it actually is.
Is this the first DevOps AI agent?
No, and our architecture write-up names the prior art. AWS DevOps Agent reached general availability on 31 March 2026, Datadog's Bits AI for site reliability engineering in December 2025, and Clanker is an open-source command line agent across every major cloud that you should look at if you want one today. Infraze is a different set of trade-offs, not a first.
Why not just use the expensive model for everything?
You can, and for a while you should. The dial is a claim that most steps in a provisioning run do not need the best model available, and that paying for it anyway is simply the default behaviour of every tool in the category. Day 49 exists to find out whether that claim survives contact with a real cost table.
What happens when the router is wrong?
It routes upward. Confidence is used only to escalate a step to the stronger engine, never to move one down, because the saving is small and the downside is not. Everything destructive stops at a human regardless of which engine produced it.
Which open-weight model is behind Kestrel?
Not named yet. The benchmarking is unfinished, and naming it now would be the kind of claim this lab tries not to make. The model and its measured cost per completed step will be published together.
Who is this for?
Teams who want the provisioning agent and the cloud bill to come from different companies, and who would rather read a memory file than open a vendor console. If you are happy inside one cloud, that cloud's own agent has distribution and integration that Infraze will not have for a long time.
Tell us what you would point it at. While it is still being built.
The scope above is our best read on what matters first. If your stack says otherwise, that shapes what we build next, and it is worth more to us than a signup.