The question nobody asks after launch
You automated something. Six weeks later, does anyone actually know if it worked?
Most small businesses skip this step. They set up the Zapier flow, roll out the new tool, feel the initial relief — then move on to the next fire. Nobody circles back with numbers. The automation quietly earns its keep or quietly rots, and either way, no one can say for sure which.
This is the Check stage — where you stop assuming and start verifying. Automation isn't free. It costs setup time, a subscription fee, and some process disruption while people adjust. If you don't measure the payoff, you can't tell a real win from a placebo.
Start with the leverage type, not the tool
Before picking metrics, remember why you automated the thing in the first place. Every automation should map to one of four kinds of leverage: Efficiency, Scale, Insight, or Risk Reduction. The metric you track should match the leverage you were after.
- If you automated to save time — that's Efficiency. Measure hours and error rates.
- If you automated to handle more volume without more headcount — that's Scale. Measure capacity and response time under load.
- If you automated to see something you couldn't before — that's Insight. Measure whether the data actually changed a decision.
- If you automated to prevent mistakes or disasters — that's Risk Reduction. Measure incident frequency and recovery time.
Mixing these up leads to fuzzy audits. If you built a chatbot to handle more support tickets (Scale) but only measure "customer satisfaction" (a soft Insight-style metric), you'll never know if it actually did its job.
Here's a starting set, organized by leverage type. Pick the two or three that map to what you were actually trying to fix.
| Leverage Type | Metric | How to Instrument It | What "Good" Looks Like |
|---|---|---|---|
| Efficiency | Time spent on task per week | Time-track before automating; compare tool logs or calendar blocks after | 50%+ reduction, or task disappears entirely |
| Efficiency | Error/rework rate | Count mistakes caught in review, or customer complaints tied to the task | Errors drop toward zero |
| Scale | Volume handled per person | Orders, tickets, or leads processed ÷ headcount | Volume rises without proportional headcount growth |
| Scale | Response/turnaround time | Timestamp logs in your tool (first response, resolution time) | Stays flat or improves as volume grows |
| Insight | Decisions made using the data | Log instances where a dashboard or report changed an action | At least one concrete decision per month traced to the data |
| Insight | Time to answer a business question | Clock how long it takes to get an answer today vs. before | Minutes instead of hours or "I don't know" |
| Risk Reduction | Incident frequency | Track near-misses, downtime events, or compliance flags | Fewer incidents, or same incidents with less damage |
| Risk Reduction | Recovery time | Time from failure to full restoration (e.g., data restore test) | Recovery measured in minutes/hours, not days |
Every row has an instrumentation method, not just a vague goal. If you can't say how you'll measure it, you haven't finished designing the automation — you've just installed software.
Instrumenting metrics without overbuilding
You don't need a business intelligence platform for this. Most of these numbers live in tools you already have:
- Time logs. A simple before/after timer — even a stopwatch and a notebook for a week — beats guessing. Ask "how long did this actually take me" before you touch a tool.
- Tool-native reports. Most SaaS products (scheduling apps, CRMs, help desks) log volume and response time automatically. Check the reporting tab before building your own tracker.
- A single tracking sheet. One spreadsheet with a date, a task, and a number is enough. Don't build elaborate dashboards to measure a small automation — that's leverage mismatch in the other direction.
- Baseline first, always. If you didn't capture "before" numbers, estimate them honestly from memory or old records rather than skipping the comparison. An imperfect baseline beats none.
A before/after example
A small accounting shop automated recurring invoice generation and reminder emails — previously done by hand every month. Here's what an honest audit looked like:
| Metric | Before | After (60 days) | Verdict |
|---|---|---|---|
| Hours/month on invoicing | 9 hours | 2 hours | Real win — 7 hours/month freed |
| Late payments | 14% of invoices | 9% of invoices | Modest win — reminders help |
| Errors (wrong amount/date) | 3 per month | 0 per month | Strong win — consistency effect |
| Staff sentiment | "Dread the first of the month" | "Barely notice it happens" | Qualitative confirmation |
Nothing exotic here — just numbers pulled from time logs and the accounting software's own reports. But now the business owner has proof, not a feeling, that the automation earned its subscription cost many times over. That proof also makes it easy to defend the tool when budgets get scrutinized later, and it's the same kind of evidence you'd want on hand during a 90-day automation audit.
How to audit yourself honestly
The biggest risk in any Check-stage review is grading your own homework generously. A few guardrails:
- Compare to a real baseline, not a guess. "It feels faster" isn't a metric. If you skipped the baseline, reconstruct it from old invoices, timestamps, or team recollection — then be conservative.
- Watch for hidden costs. New errors introduced by the automation, time spent troubleshooting, or a subscription nobody uses fully all count against the win.
- Separate adoption from impact. People logging into the new tool doesn't mean the underlying number moved. Usage metrics are a leading indicator, not proof.
- Give it enough time. Thirty days is often too short to judge Scale or Risk Reduction wins — some benefits only show up under real stress (a busy season, an actual outage attempt). Note the measurement date and revisit later if needed.
- Write down the verdict. "Worked," "partially worked," or "didn't move the needle" — pick one, in writing. Ambiguity is how failed automations linger for years unquestioned.
If your audit turns up a genuine win, don't just file it away — lock in the win with documentation so it survives staff turnover and tool changes. If the number didn't move, that's useful information too: maybe the process underneath was never fixed, which is a case for going back to an automate/improve/ignore triage rather than tweaking the automation further.
What "good enough" looks like
You don't need scientific rigor. You need enough evidence to make a confident call: keep it, tune it, or kill it. A single clear metric moving in the right direction, backed by a rough before/after comparison, is plenty for a small business decision.
The habit that matters more than any specific number is this: every automation gets checked. Not celebrated and forgotten, not quietly resented — checked, on a date you set in advance, against a number you defined before you started. That's what separates automation as a discipline from automation as a hobby.