Operations Autopilot
Let software handle the routine. Your team handles the exceptions.
We design and implement governed operations autopilots. That's always-on software that watches your systems, mailboxes and feeds and applies the rules your team would follow by hand. It takes safe actions on its own and asks a person before anything risky.
- Rules you can read
- Human approval where it matters
- Every action audited
Operations by hand
- Someone has to notice, every time, at any hour
- The same ticket, email and follow-up written for the hundredth time
- Ten alerts for one outage, and the important one gets buried
- Follow-ups slip when the person who owns them is away
- Nobody can reconstruct who did what, and why
With a governed autopilot
- Continuous watching, with noise suppressed and duplicates merged
- Routine responses handled the same way every time, from templates
- Related alerts grouped into one incident, one ticket, one message
- Follow-ups and recovery checks on schedule, never forgotten
- An audit record written before every action, with the reason for it
What's inside
The building blocks of every autopilot we deliver
Watch and correlate
- Connectors for monitoring, email, ticketing, chat and APIs
- Stale-data detection that pauses decisions
- Grouping, de-duplication and maintenance windows
- Circuit breakers when an integration is unhealthy
Decide by rules
- A deterministic rule engine with no hidden logic
- Thresholds and templates in reviewed configuration
- A clear state for every incident, from open to resolved
- Tests built from real past incidents
Act safely
- One gate for every side effect
- Built-in protection so nothing happens twice
- Rate and content limits on outbound messages
- A never-automate list enforced at startup
Keep people in charge
- Approve or reject with a one-line reply
- Commands to take over, pause or resume any incident
- Automation steps back as soon as a person owns an incident
- Daily digest and health alerts with fallback channels
How it works
One pipeline, and one gate that every action has to pass
- Deterministic code
- AI, with a narrow job
- People
Code
Watch
Monitoring tools, mailboxes, ticket systems, APIs and data feeds. If a source goes stale, the cycle stops instead of acting on bad data.
Code
Correlate
Related alerts are grouped, maintenance windows honoured, test systems ignored and duplicates dropped. A mass outage becomes one incident, not a hundred.
AI, optional
Understand free-text messages
When a reply or notice is free text, a language model can label it from a fixed list. It never decides what happens next. Low confidence goes to a person.
Code
Decide
A deterministic rule engine maps each incident to its next actions. Every threshold lives in configuration your team can read and review.
Code
Pass the gate
Every action goes through one dispatcher. It checks for duplicates, applies rate limits and allow-lists, picks the approval tier and writes the audit record before anything runs.
Your team
Act, or ask first
Routine actions run on their own. Risky ones wait for a one-line approval by email, chat or command. A daily digest shows everything that happened.
Engineering rules we don't bend
Automation you can trust has to be boring in the best way
Separate thinking from acting
Deciding is done by pure functions with no side effects. Acting goes through one governed gate. That split is what makes the system testable.
Safety in code, not in prompts
Limits, allow-lists and approval rules are enforced by code and checked at startup. No instruction to a model can switch them off.
Autonomy is earned
New capabilities start switched off, then run in shadow mode recording what they would have done. They go live only after that. Actions move from ask-first to automatic once approval history shows they are routine.
Fail closed
Missing or stale evidence means no action and an alert. The autopilot never fills a gap with a guess.
Nothing happens twice
Layered duplicate protection means a restart, a retry or a replayed message can’t send the same email or open the same ticket again.
When a person takes over, the machine steps back
Once someone owns an incident, automated messages for it stop. No awkward duplicate follow-ups.
Where the pattern fits
Any loop with a trigger, a check, a routine action and a clear point where a person should decide
Vendor and supplier outages
- When
- A monitor reports a service down
- Then
- Open a ticket, email the vendor from a template, follow up on schedule, confirm recovery and close
- Ask a person
- No vendor reply by the deadline
Certificate expiry
- When
- Daily scan of every certificate
- Then
- Renew, deploy and verify the new certificate
- Ask a person
- Renewal fails or the certificate can't be automated
Backup verification
- When
- A nightly backup job finishes
- Then
- Run a restore test, compare size and checksums, re-run if needed
- Ask a person
- Second failure or an unexplained size change
Disk, queue and capacity pressure
- When
- A threshold is breached for several checks in a row
- Then
- Rotate logs, prune safely or add workers
- Ask a person
- Pressure persists after the action
Configuration and DNS drift
- When
- Scheduled comparison of live state against declared state
- Then
- Open a ticket with the diff and revert pre-approved settings
- Ask a person
- Drift in anything not on the safe list
Data pipeline freshness
- When
- A key table falls behind its agreed freshness
- Then
- Re-trigger the upstream job and re-check
- Ask a person
- A second miss or a reporting deadline at risk
Credential rotation
- When
- A secret reaches its maximum age
- Then
- Rotate it and verify every service that depends on it
- Ask a person
- A dependent service fails after rotation
SLA breach watch
- When
- Open tickets age against contract terms
- Then
- Nudge the assignee and raise the priority
- Ask a person
- A breach is imminent
Payment reconciliation
- When
- A bank or payment provider export arrives
- Then
- Match payments to invoices and flag the remainder
- Ask a person
- Unmatched amounts above a threshold
Stuck orders
- When
- An order sits in one state longer than expected
- Then
- Retry the fulfilment call and update the customer
- Ask a person
- Retries exhausted
Production line signals
- When
- A sensor crosses its limit, confirmed by a second signal
- Then
- Open a work order and notify the shift lead
- Ask a person
- Safety-related signals always go straight to a person
Compliance evidence
- When
- A weekly control check
- Then
- Collect and store the evidence for auditors
- Ask a person
- A control is failing
How we build yours
We start with one loop and earn trust before adding autonomy
Map the loop
We sit with the people who run it by hand and write down the triggers, decisions, actions and exceptions. That becomes the rulebook, reviewed by your team.
Run in shadow mode
The autopilot runs against live data but only records what it would have done. You compare that with what your team actually did, and we fix the gaps.
Go live with approvals
Real actions, with the risky ones waiting for a one-line approval. The audit log and daily digest are there from day one.
Earn autonomy
When approval history shows an action is routine, it moves to automatic. Rolling that back is a configuration change. Your team keeps the exceptions.
Is this right for you?
A good fit if
- The same operational loop runs many times a week, by hand
- The rules exist, even if only in someone's head
- Mistakes are costly enough that you want approvals and an audit trail
- You want automation you can read, test and switch off
Probably not a fit if
- The process changes every week. Let it settle first
- It happens twice a year. A good checklist is cheaper
- You want an AI agent improvising actions on its own. We build the opposite
Proven in production
We first built this pattern for an infrastructure operator running hundreds of servers across dozens of hosting providers. Every outage used to mean finding the provider, opening a ticket, emailing, chasing and confirming recovery by hand. The autopilot now runs that loop around the clock. Operators approve the exceptions with a one-line reply and read one daily digest.
Frequently asked questions
Is this an AI agent?
No. The core is a deterministic rule engine. A language model is optional and only labels free-text messages, such as a vendor reply or a maintenance notice, using a fixed list. It never proposes or performs actions. Its output has to pass deterministic checks before anything changes, and low-confidence results go to a person. Many operational loops need no AI at all.
What can it connect to?
Anything with an API, a mailbox or a log. That covers monitoring tools such as Nagios, Zabbix or Prometheus, Microsoft 365 and other email systems, helpdesk and ticketing tools, Teams or Slack, and your internal APIs and databases. Each integration is a small, separately tested adapter, so adding a new one doesn’t touch the decision logic.
What if it does something wrong?
Several layers make that unlikely and limit the damage if it happens. New behaviour runs in shadow mode first. Risky actions need approval. Duplicate protection and rate limits prevent runaway loops. Every action is audited before it runs. Your team can take over any incident at any time, and automation for that incident stops immediately. There’s also a global pause.
How do people approve actions?
However your team already works. That can be a one-line reply to an email with an approval code, a command in chat or a ticket note, or a command-line tool. Approvals expire, so nothing old gets executed by surprise.
Where does it run?
On your server or in your cloud account, as a single hardened service. It’s not a SaaS product. You own the code, the configuration and the data, and it keeps working whether or not we’re involved.
How long does it take?
It depends mostly on how many systems the loop needs to talk to and how clear the rules already are. We agree a realistic plan once we’ve mapped the loop with your team. Going live is then a decision based on shadow-mode evidence, not a date on a calendar. Once the foundation exists, each additional loop is much faster to add.
What about AI costs?
Where we use a language model, it has hard caps per call, per incident and per day, enforced in code. Most of the pipeline uses no AI at all, so costs stay small and predictable.
Which loop does your team run by hand every week?
Tell us about it. We'll tell you whether an autopilot makes sense, what it could handle on day one and what should stay with a person.
Automate an operational loopFree 30-minute call · No sales pitch · Honest assessment