Skip to content
Bricks Footer
Blog

Why RBFS Treats Config as a Transaction

Date Published

By Hannes Gredler, CTO of RtBrick

Configuration is not a script

Get ready for Agentic AI Ops - why RBFS treats config as a transaction

If you've operated a wide-area network for more than a few years, you already know the moment: it's 2 a.m., you're pushing a change to a router during a maintenance window, and three lines into the paste something goes sideways. Maybe the session hiccuped. Maybe a dependent object hadn't been created yet. Now the device is in a state that was never in your intended configuration and never in your previous configuration either. It's stuck in between — and "in between" is not a valid operating state.

That failure mode isn't an accident of fat fingers. It's a direct consequence of a design decision made decades ago and copied, largely unexamined, into most of the network operating systems that followed it.

The first generation: configuration as a sequence of commands

Classic router CLIs — the Cisco IOS model that an entire generation of NOSes imitated — treat configuration as a stream of imperative commands, applied one at a time, in order, against a single running state. You type a line, it takes effect immediately. There is no notion of atomicity of a set of instructions i.e either "all 40 lines all apply or none of them do." There is no built-in notion of two units of comparison: a version A configuration versus version B configuration, that  you could show, compare, or choose between. There is only "the running configuration, as it currently stands after whatever has been typed into it so far." There are also little traces for recreating the root cause during a post-mortem analysis of what went wrong 

This model was a reasonable fit for the operating conditions of routers in the 1990s: one engineer, one terminal, one device, small and infrequent changes. It is a poor fit for operating a wide-area network today, for reasons that have nothing to do with CLI syntax and everything to do with what the model fundamentally cannot express:

  • No atomicity. A multi-line change is not a unit. If it's interrupted halfway — a dropped SSH session, a typo three lines in, a script that crashes — the device is left in a state that was never validated as a whole. Tools and operators have to work around this with external discipline (config generation pipelines, pre-change snapshots, careful ordering of commands) rather than relying on the device itself to guarantee it.
  • No native history. "Running config" and "startup config" are two text blobs, not a timeline. If you want to know what the configuration looked like six weeks ago, or what specifically changed between last Tuesday's maintenance and this one, the NOS itself has no answer — you reach for RANCID or a git repo of scraped show run output that somebody remembered to commit. History lives beside the system, maintained by convention, not in the system, guaranteed by design.
  • No safe, instant rollback. Reverting a bad change usually means re-applying a previous config wholesale, re-running the same imperative apply process that got you into trouble in the first place — with all the same atomicity problems, under time pressure, during an outage.
  • No structured diff. Comparing two configurations means comparing two flat text files, line by line, hoping that whitespace and ordering don't get in the way. There's no notion of "this is a change to the BGP neighbor object with this key" that survives reordering or reformatting.

On a lab device, these gaps are annoying. Across a production network with dozens or hundreds of edge and core routers, where configuration changes happen daily through automation, where multiple teams touch overlapping parts of the config, and where a bad change needs to be identified and reverted in minutes, not hours — these gaps are catastrophic leading to network outages.

Needs of modern networks 

Operators actually need a system closer to that which every production software system settled on decades ago for managing state: configuration as versioned, transactional data, not configuration as a diary of commands.

Concretely:

  • Every change is proposed as a candidate, validated as a whole, and only takes effect via an explicit, atomic commit. Either the entire change applies, or none of it does.
  • Every commit is a first-class, addressable object — not a vague "the config, currently" but a specific, identifiable version with a timestamp and an identity you can refer to later.
  • Rollback is a native operation, not a disaster-recovery procedure. Reverting to a known-good state should be as fast and as safe as the commit that got you into trouble.
  • The history of commits persists across restarts and is queryable by the system itself — "what changed, and when, and by whom" is a question the device can answer, not a question you have to answer with an external change-tracking tool bolted on after the fact.

This is exactly the model RBFS builds in at the architecture level, not as an afterthought bolted onto a CLI. Configuration in RBFS is owned by a dedicated configuration daemon (confd), one of a set of independent microservices — "brick daemons" — that make up the system (alongside daemons like bgpd, isisd, ribd, subscriberd, etc. for routing and fibd for forwarding). Changes are staged, then applied with an explicit commit. RBFS keeps a genuine commit history: show commit log reviews it, and reverting is a native command — rollback 1 for the previous configuration, rollback 2 for the one prior to  that, or rollback commit-id <hash> to jump straight to a specific commit by its identity. That's a deliberately git-shaped idea: configuration is a sequence of addressable, identified versions, not an undifferentiated blob that happens to be whatever was last typed in.

Under the hood, that pipeline is concrete, not just philosophy. The CLI and a RESTCONF API both front the same YANG-based configuration manager, Clixon; Clixon's backend integrates into confd, which owns two separate databases — a candidate configuration and the running one — and creates the configuration tables every other daemon subscribes to. A CLI session edits the candidate: show diff to see what's staged, discard to throw it away, commit to apply it. Nothing takes effect by merely being typed — only commit moves it into the running configuration and onto every subscribed daemon's plate. show commit log makes that version history literally visible:

1Rollback_ID Commit_ID Timestamp
20 9c9a3df2b3a6bacb5709a647b907b111 2026-10-01T12:20:13.353259+0000
31 b2bb31a2b16a7a1562dba165a13f18a0 2026-10-01T12:07:54.207065+0000
42 2ffe8c8d867fb90f3a7143b3c0a5e19d 2026-10-01T12:07:49.229863+0000


Each of the commit IDs is a real, content-addressed object on disk, not just a line in a log — each past commit gets its own directory, named by its hash, under /var/rtbrick/commit_rollback/. And rollback doesn't quietly bypass the discipline the rest of the system runs on: rollback 1 loads that previous configuration back into the candidate, show diff shows you exactly what reverting it would change, and it still takes an explicit commit to actually apply it. Undo goes through the same gate as every other change — there's no separate, less-supervised "emergency revert" path for a bad decision made under pressure to slip through.

The harder half: making the transition graceful

Storing configuration as versioned data is necessary, but not sufficient as  it only solves the bookkeeping problem. It doesn't solve the problem that actually breaks routers during a change: what happens to everything downstream of the configuration when it moves from version A to version B?

A line-oriented CLI has only one answer to that question: re-execute the commands, in order, against the live system, and hope the protocol stacks and forwarding state underneath react sensibly to whatever sequence of individual mutations that produces. There's no single point where the system reasons about "configuration A becomes configuration B" as a whole — only a string of individual, uncoordinated pokes.

A transactional configuration manager has a fundamentally different job at commit time: it has to compute the actual delta between the two structured configurations — not a text diff, but a diff over the real configuration objects and their relationships — and then drive every subsystem that depends on the changed objects through a correct transition, not a replay of commands. In RBFS, this is possible precisely because the architecture is already decomposed into independent daemons sharing configuration and operational state through a common, schema-driven, in-memory datastore (the Brick Data Store). A commit doesn't get broadcast as a list of CLI lines for every daemon to reinterpret rather it changes the shared state, and the protocol and forwarding daemons that actually care about those specific objects — BGP, IS-IS/OSPF, interface management, the RIB, the FIB — update and reinitialize only the parts of their own state that the delta actually touches.

That distinction matters enormously in practice. Changing an interface's MTU shouldn't require restarting every protocol that happens to run over that interface — it should prompt exactly the IGP and BGP sessions that advertise or depend on that MTU to re-adjust, and nothing else. Adding a BGP neighbor shouldn't perturb the sessions that already exist. Removing a routing policy shouldn't require a full RIB recompute if nothing downstream of that policy actually changed. A commit-based manager that only stores and diffs text can't make these distinctions — it doesn't know what an "object" is, only what a line of text is. A transaction aware commit-based manager built on structured, schema-driven data can, because the delta it computes is a set of objects, with identity and relationships, not a set of lines.

This isn't just an abstraction, either. RBFS's running configuration is literally YANG-modeled JSON — show config and a RESTCONF GET return the same namespace-qualified object tree, built from the same YANG models the configuration is validated against. A "delta" is a structural diff over named objects like rtbrick-config:instance or rtbrick-config:system, not a diff over indentation and keywords. Those same YANG models even let a candidate configuration's syntax be verified offline, before it's staged on a device at all — a check a line-oriented CLI has no equivalent of, because it has no notion of "configuration" independent of the device it's currently being typed into.

Prepare for a world with AI agents as operators

Everything so far has been framed around human operators — engineers pushing changes, automation scripts written by people, pipelines reviewed by people. That framing is already incomplete. Network operations are rapidly moving toward AI agents that observe telemetry, reason about intent, and propose or directly execute configuration changes with limited human intervention in the loop than a change-managed script would have had.

That shift moves configuration management from "good practice" to "precondition." An agent acting against a first-generation, line-oriented CLI is doing exactly what a tired engineer does at 2 a.m. — pasting a sequence of commands into a running system and hoping nothing in between goes wrong — except it can do that across hundreds of devices, continuously, at machine speed, with no one watching the lines scroll by. A partially-applied change that would have been one bad maintenance window for one engineer becomes a fleet-wide incident when the actor never pauses to notice something looks wrong.

A commit-based, transactional config store is what makes it defensible to let an agent act at all, rather than merely draft commands for a human to type:

Atomicity turns an agent's uncertainty into a safe failure, not a stuck system. An agent's proposed change is, structurally, a hypothesis — informed by telemetry and a model of intent, but still a guess about what the network needs next. If a commit either fully applies or is cleanly rejected, a wrong hypothesis costs nothing. If it can half-apply, a wrong hypothesis becomes a new incident the agent itself caused.

  • A real commit history is the audit trail that makes an agent's actions reviewable, not just reversible. "What changed, when, and why" have to be answerable after the fact for a human-driven change; they are non-negotiable for a machine-initiated change, where engineers, auditors, and the agent's own operators all need to reconstruct exactly what it did, tied to a specific, addressable commit rather than a vague window of time.
  • Native, instant rollback is the safety net that makes closing the loop tolerable. Giving an agent write access to production infrastructure is only defensible if a wrong decision can be undone as fast and as deterministically as it was made — rollback 1, not a frantic manual reconstruction of what "prior" looked like.
  • Structured, schema-driven configuration is also simply the interface an agent reasons about more reliably. An agent working against modeled objects in a datastore can validate a proposed change, express it as a diff, and reason about its blast radius before committing. An agent working against a stream of CLI text is reduced to pattern-matching strings and hoping the syntax holds together — the same fragility a human has, with none of a human's instinct that something looks wrong.

That last point comes with a sharp edge worth naming, not glossing over. RBFS also exposes a RESTCONF API — standard HTTP GET/POST/PUT/PATCH/DELETE against the same YANG-modeled objects, in principle a far more natural interface for an agent than a CLI session to scrape. But RESTCONF requests apply to the running configuration immediately, in keeping with REST's own stateless principle — they don't land in the candidate database the way a CLI session's edits do. A human operator reasons about that distinction once, deliberately, every time they open a config session versus script against the API. An agent calling the API reflects its designer’s discipline, and not the one that the protocol gives it for free — and if that discipline isn't there, the agent has just rediscovered the no-atomicity problem this piece started with, over HTTP instead of a terminal. A transactional config store only protects an agent that actually routes its changes through the transaction, not around it.

This isn't a hypothetical for some distant future; it's the automation-at-scale argument above, carried to its logical conclusion. The less a human verifies each individual step, the more the guarantees have to live in the system itself. A transaction aware configuration manager isn't only safer for the humans scripting against it today — it's the precondition for letting an autonomous agent touch the network safely tomorrow.

Why specifically this is a network problem

None of this is about CLI aesthetics. It's about what happens when the blast radius of a mistake and the cost of a slow, uncertain recovery both scale up. That is exactly what happens when you move from managing one box in a lab to running a fleet of routers carrying production traffic across a large network, under constant changes from automation, multiple teams, and compressed maintenance windows, and increasingly from agents acting with less and less human oversight between a decision and its effect on the network.

A first-generation, IOS-shaped CLI treats every change as a leap of faith. You type the commands, hope nothing fails mid-sequence, and if something does, you start troubleshooting from scratch, with no structured record of what the system believed its own configuration to be five minutes earlier.

A transaction aware, commit-based configuration manager treats every change as what it actually is: a transition from one well-defined, versioned state to another. It takes architectural responsibility for making that transition safe, atomic, auditable, and reversible, whether the change comes from an engineer, a script, or an agent. Extend that model across the fleet, and the unit of change is no longer a single device but the network itself, committed or rolled back as a whole. Couple that with AI Agents and the speed and impact of change is profound.

For a network where downtime is measured in customer impact and dollars, we argue that this concept isn't a nicety, it is the difference between an operational model and a hope. As my co-founder and CEO Pravin Bhandarkar often says, "Hope is not a strategy."

/Hannes Gredler
CTO RtBrick Inc.