---
title: "Automation Does Not Call In Sick. That Is the Problem."
url: "https://cooinsider.com/insight/automation-does-not-call-in-sick-that-is-the-problem/"
author: "Ankush Gupta"
published: "2026-10-01"
updated: "2026-10-01"
---

# Automation Does Not Call In Sick. That Is the Problem.

A client asked where their weekly coverage report was. We checked. The workflow that produced it had stopped some time before, and nothing in our stack had mentioned it. No error alert, no failed run sitting in anyone's inbox. Executions were piling up behind one stalled step, and because nothing had crashed, nothing asked for attention. The automation was not broken in any way a monitoring tool would recognize. It had gone quiet, and quiet reads as fine.

Around the same period a messaging integration dropped out in production during a deploy. We knew inside minutes, because messages stopped arriving and people said so loudly. On paper that was the worse incident. It cost us far less.

That gap is what I want to argue about. At FameNinja we run reputation work, PR production and a publishing network on automated pipelines, and the failures that have hurt us most were never the loud ones.

A person who stops working tells you. A workflow will not.

When someone on an ops team stops doing a task, the organization finds out through ordinary human noise. They mention it in standup. Someone covers for them. A client chases. The absence generates signal because people are connected to other people.

Software has no such property. A workflow that stops producing output produces nothing, and nothing looks identical to a system that had no work to do. A queue with no items and a queue that has stopped draining sit side by side on a dashboard built to report errors.

This is why careful teams still get caught. They test that the build works. They never build anything that notices absence.

There are two silent failures and they need different checks. One is the stall, where work stops and the backlog grows behind it. The other is the wrong answer, where every step completes and what comes out the far end is nonsense. The second is the dangerous one in our line of work, because de-indexing a page runs two to six weeks in the cases we handle, not minutes.

Monitor the output, not the automation.

Most automation monitoring watches the wrong object. Uptime, execution success, API status, last-run timestamp. All of that tells you the machine is running. None of it tells you the work got done.

The check that has actually saved us is duller. Did the expected artifact appear, in the expected place, inside the window it was meant to appear in? A report in a folder. A row in a sheet. If it did not, something asks a human, and the alert names the business outcome rather than the failing node, because whoever reads it over a weekend needs to know what a client is missing.

We now write that check before the workflow it watches. It feels backwards. It takes a small fraction of the time the pipeline takes to build, and it is the only part of the system that behaves like an employee.

One warning about where the alert lands. A channel that receives every automated message becomes a channel nobody reads, and a muted alert is worse than no alert, because it buys false comfort. We keep output checks on a separate route from routine notifications and keep the volume low enough that anything arriving there still means something.

Put a name against every workflow.

The second failure is ownership, and it is a leadership problem rather than a technical one.

Automation gets built by whoever was closest to the problem, then recedes into the background. Months later nobody is quite sure who owns it or whether it is still needed. It keeps running. Everyone assumes someone is watching.

Our rule is that every pipeline carries one named owner, and that owner is an operations person rather than a developer. That cuts against the instinct to hand anything automated to engineering.

In the work we run, the ops people who learned the automation tool have outperformed the developers we brought in for the same job. They know what the output is supposed to look like when it is right, so they catch the workflow that is executing perfectly and producing garbage. A developer sees green. An operator sees that the third field is wrong.

The fallback only counts if someone has run it.

Every automated process needs a manual version that a named human has performed recently, start to finish. Not documented. Performed.

We learned this after a security incident on our publishing network, where an admin account was used maliciously and we had to lock down and rebuild access across a large number of sites at once. Automation did not make that survivable. What made it survivable was that the manual path still existed in people's hands.

Writing the fallback procedure is easy and mostly useless on its own. The first time anyone runs it, they find the credential nobody has and the approval that used to come from someone who left the company. Far better to find those on a slow Tuesday.

### **Where this gets uncomfortable**

All of this costs time that produces nothing visible.

Every one of these steps adds work to the automation you are building, and none of it shows up as output. The pipeline would have run fine without them, right up until the morning it did not. So they get cut. They get cut most often by the teams most confident in their systems.

I do not have a clean answer to that incentive. What we do is treat the watcher as part of the build rather than as a separate task that can be pushed to next month. If the check does not exist, the workflow is not finished, and it does not go to production.

The odd result is that this makes us automate more, not less. We hand a process to a machine far more readily once we know the machine will admit it stopped.

---

Ankush Gupta is a Fractional CMO at [FameNinja](https://fameninja.com), where he works on online reputation management, digital PR and marketing automation.
