Operations Autopilot

Let software handle the routine. Your team handles the exceptions.

We design and implement governed operations autopilots. That's always-on software that watches your systems, mailboxes and feeds and applies the rules your team would follow by hand. It takes safe actions on its own and asks a person before anything risky.

  • Rules you can read
  • Human approval where it matters
  • Every action audited

Operations by hand

  • Someone has to notice, every time, at any hour
  • The same ticket, email and follow-up written for the hundredth time
  • Ten alerts for one outage, and the important one gets buried
  • Follow-ups slip when the person who owns them is away
  • Nobody can reconstruct who did what, and why

With a governed autopilot

  • Continuous watching, with noise suppressed and duplicates merged
  • Routine responses handled the same way every time, from templates
  • Related alerts grouped into one incident, one ticket, one message
  • Follow-ups and recovery checks on schedule, never forgotten
  • An audit record written before every action, with the reason for it

What's inside

The building blocks of every autopilot we deliver

Watch and correlate

  • Connectors for monitoring, email, ticketing, chat and APIs
  • Stale-data detection that pauses decisions
  • Grouping, de-duplication and maintenance windows
  • Circuit breakers when an integration is unhealthy

Decide by rules

  • A deterministic rule engine with no hidden logic
  • Thresholds and templates in reviewed configuration
  • A clear state for every incident, from open to resolved
  • Tests built from real past incidents

Act safely

  • One gate for every side effect
  • Built-in protection so nothing happens twice
  • Rate and content limits on outbound messages
  • A never-automate list enforced at startup

Keep people in charge

  • Approve or reject with a one-line reply
  • Commands to take over, pause or resume any incident
  • Automation steps back as soon as a person owns an incident
  • Daily digest and health alerts with fallback channels

How it works

One pipeline, and one gate that every action has to pass

  • Deterministic code
  • AI, with a narrow job
  • People
  1. Code

    Watch

    Monitoring tools, mailboxes, ticket systems, APIs and data feeds. If a source goes stale, the cycle stops instead of acting on bad data.

  2. Code

    Correlate

    Related alerts are grouped, maintenance windows honoured, test systems ignored and duplicates dropped. A mass outage becomes one incident, not a hundred.

  3. AI, optional

    Understand free-text messages

    When a reply or notice is free text, a language model can label it from a fixed list. It never decides what happens next. Low confidence goes to a person.

  4. Code

    Decide

    A deterministic rule engine maps each incident to its next actions. Every threshold lives in configuration your team can read and review.

  5. Code

    Pass the gate

    Every action goes through one dispatcher. It checks for duplicates, applies rate limits and allow-lists, picks the approval tier and writes the audit record before anything runs.

  6. Your team

    Act, or ask first

    Routine actions run on their own. Risky ones wait for a one-line approval by email, chat or command. A daily digest shows everything that happened.

Some actions can never be automated, by design. The autopilot refuses to start if a configuration change would allow it to run them without a person.

Engineering rules we don't bend

Automation you can trust has to be boring in the best way

Separate thinking from acting

Deciding is done by pure functions with no side effects. Acting goes through one governed gate. That split is what makes the system testable.

Safety in code, not in prompts

Limits, allow-lists and approval rules are enforced by code and checked at startup. No instruction to a model can switch them off.

Autonomy is earned

New capabilities start switched off, then run in shadow mode recording what they would have done. They go live only after that. Actions move from ask-first to automatic once approval history shows they are routine.

Fail closed

Missing or stale evidence means no action and an alert. The autopilot never fills a gap with a guess.

Nothing happens twice

Layered duplicate protection means a restart, a retry or a replayed message can’t send the same email or open the same ticket again.

When a person takes over, the machine steps back

Once someone owns an incident, automated messages for it stop. No awkward duplicate follow-ups.

Where the pattern fits

Any loop with a trigger, a check, a routine action and a clear point where a person should decide

Vendor and supplier outages

When
A monitor reports a service down
Then
Open a ticket, email the vendor from a template, follow up on schedule, confirm recovery and close
Ask a person
No vendor reply by the deadline

Certificate expiry

When
Daily scan of every certificate
Then
Renew, deploy and verify the new certificate
Ask a person
Renewal fails or the certificate can't be automated

Backup verification

When
A nightly backup job finishes
Then
Run a restore test, compare size and checksums, re-run if needed
Ask a person
Second failure or an unexplained size change

Disk, queue and capacity pressure

When
A threshold is breached for several checks in a row
Then
Rotate logs, prune safely or add workers
Ask a person
Pressure persists after the action

Configuration and DNS drift

When
Scheduled comparison of live state against declared state
Then
Open a ticket with the diff and revert pre-approved settings
Ask a person
Drift in anything not on the safe list

Data pipeline freshness

When
A key table falls behind its agreed freshness
Then
Re-trigger the upstream job and re-check
Ask a person
A second miss or a reporting deadline at risk

Credential rotation

When
A secret reaches its maximum age
Then
Rotate it and verify every service that depends on it
Ask a person
A dependent service fails after rotation

SLA breach watch

When
Open tickets age against contract terms
Then
Nudge the assignee and raise the priority
Ask a person
A breach is imminent

Payment reconciliation

When
A bank or payment provider export arrives
Then
Match payments to invoices and flag the remainder
Ask a person
Unmatched amounts above a threshold

Stuck orders

When
An order sits in one state longer than expected
Then
Retry the fulfilment call and update the customer
Ask a person
Retries exhausted

Production line signals

When
A sensor crosses its limit, confirmed by a second signal
Then
Open a work order and notify the shift lead
Ask a person
Safety-related signals always go straight to a person

Compliance evidence

When
A weekly control check
Then
Collect and store the evidence for auditors
Ask a person
A control is failing

How we build yours

We start with one loop and earn trust before adding autonomy

  1. Map the loop

    We sit with the people who run it by hand and write down the triggers, decisions, actions and exceptions. That becomes the rulebook, reviewed by your team.

  2. Run in shadow mode

    The autopilot runs against live data but only records what it would have done. You compare that with what your team actually did, and we fix the gaps.

  3. Go live with approvals

    Real actions, with the risky ones waiting for a one-line approval. The audit log and daily digest are there from day one.

  4. Earn autonomy

    When approval history shows an action is routine, it moves to automatic. Rolling that back is a configuration change. Your team keeps the exceptions.

Is this right for you?

A good fit if

  • The same operational loop runs many times a week, by hand
  • The rules exist, even if only in someone's head
  • Mistakes are costly enough that you want approvals and an audit trail
  • You want automation you can read, test and switch off

Probably not a fit if

  • The process changes every week. Let it settle first
  • It happens twice a year. A good checklist is cheaper
  • You want an AI agent improvising actions on its own. We build the opposite

Proven in production

We first built this pattern for an infrastructure operator running hundreds of servers across dozens of hosting providers. Every outage used to mean finding the provider, opening a ticket, emailing, chasing and confirming recovery by hand. The autopilot now runs that loop around the clock. Operators approve the exceptions with a one-line reply and read one daily digest.

Always on
Software runs the loop, people handle exceptions
6,000+
Automated tests guarding behaviour
100%
Actions audited before they run

Frequently asked questions

Is this an AI agent?

No. The core is a deterministic rule engine. A language model is optional and only labels free-text messages, such as a vendor reply or a maintenance notice, using a fixed list. It never proposes or performs actions. Its output has to pass deterministic checks before anything changes, and low-confidence results go to a person. Many operational loops need no AI at all.

What can it connect to?

Anything with an API, a mailbox or a log. That covers monitoring tools such as Nagios, Zabbix or Prometheus, Microsoft 365 and other email systems, helpdesk and ticketing tools, Teams or Slack, and your internal APIs and databases. Each integration is a small, separately tested adapter, so adding a new one doesn’t touch the decision logic.

What if it does something wrong?

Several layers make that unlikely and limit the damage if it happens. New behaviour runs in shadow mode first. Risky actions need approval. Duplicate protection and rate limits prevent runaway loops. Every action is audited before it runs. Your team can take over any incident at any time, and automation for that incident stops immediately. There’s also a global pause.

How do people approve actions?

However your team already works. That can be a one-line reply to an email with an approval code, a command in chat or a ticket note, or a command-line tool. Approvals expire, so nothing old gets executed by surprise.

Where does it run?

On your server or in your cloud account, as a single hardened service. It’s not a SaaS product. You own the code, the configuration and the data, and it keeps working whether or not we’re involved.

How long does it take?

It depends mostly on how many systems the loop needs to talk to and how clear the rules already are. We agree a realistic plan once we’ve mapped the loop with your team. Going live is then a decision based on shadow-mode evidence, not a date on a calendar. Once the foundation exists, each additional loop is much faster to add.

What about AI costs?

Where we use a language model, it has hard caps per call, per incident and per day, enforced in code. Most of the pipeline uses no AI at all, so costs stay small and predictable.

Which loop does your team run by hand every week?

Tell us about it. We'll tell you whether an autopilot makes sense, what it could handle on day one and what should stay with a person.

Automate an operational loop

Free 30-minute call · No sales pitch · Honest assessment