The goal: decide fast, not perfectly
You've spotted a place where AI could help — maybe flagging expense anomalies, drafting customer replies, or predicting which clients might churn. The question isn't whether AI is useful. It's whether you should buy an off-the-shelf tool, build something custom, or wait.
Most owners either freeze (endless research) or overcommit (a six-month build for a problem they haven't tested). Here's how to answer the question in one week, using a real pilot instead of a spreadsheet of pros and cons.
Step 1: Name the specific job, not the category
"We should use AI for customer service" is too vague to act on. Narrow it to one task with a clear output.
Example: "Draft a first-response reply to incoming support emails so my team edits instead of writes from scratch."
Write it as a single sentence: what goes in, what comes out, who uses it. If you can't state it that simply, you're not ready to pilot anything yet — go map the process first. Our guide on mapping a process before you automate it is a good next stop if that's where you are.
Step 2: Run a one-week pilot with an existing tool
Don't build anything yet. Find the closest off-the-shelf option and use it as-is for five business days.
For the support-reply example, that might mean:
- Turning on the AI draft-reply feature already built into your helpdesk (Zendesk, Intercom, and Help Scout all have one now).
- Or feeding a handful of real incoming emails into a general tool like ChatGPT or Claude with a simple prompt template, with a human reviewing each draft before it's sent.
Keep the setup dead simple. One person, one inbox, one week. The point isn't a polished workflow — it's real data on whether AI adds value here at all. This is the same logic behind piloting an AI insight tool: small, fast, and reversible.
While you run the pilot, track three things daily:
- Time saved or lost — did drafting take less time than writing from scratch, including review?
- Quality — how many drafts needed heavy edits versus light touch-ups?
- Team reaction — did people actually use it, or quietly go back to the old way?
Step 3: Score the pilot, then make the call
At the end of the week, sit down with your notes and answer three questions honestly.
Is this a common need? If dozens of other businesses have the same problem — drafting replies, summarizing calls, flagging odd transactions — a mature product almost certainly exists. Lean toward buy. Most AI needs for a small business fall here: the tool is common, the vendor keeps improving the model, and you get updates for free.
Is it tied to something genuinely unique about how you operate? If your pilot revealed that the generic tool almost works but keeps missing something specific to your business — a scheduling quirk, a proprietary pricing rule — that's a signal for build, but only if that quirk is core to your edge. A one-person consultancy customizing a chatbot script is not the same as a logistics company with a scheduling algorithm nobody else has. Be honest about which one you are.
Did the pilot actually move the needle? If time saved was marginal, quality was inconsistent, or your team ignored the tool, that's not necessarily a failure — it's information. It might mean defer: the task doesn't have enough volume yet, the tool wasn't the right fit, or the underlying process needs cleanup before any AI layer will help.
Once you land on one of the three, write a single sentence explaining why. This matters more than people expect — it turns a vague feeling into a decision you can revisit.
- Buy: "The helpdesk's built-in draft feature cut reply time by 40% with light edits needed. Subscribing to the upgraded tier costs $30/month more — approved."
- Build: "Off-the-shelf tools can't handle our multi-location inventory logic. We're scoping a small custom integration with a freelance developer, budgeted at $4,000, because this directly affects fulfillment speed — our main bottleneck."
- Defer: "Draft quality was inconsistent and the team reverted to manual replies by day three. We'll revisit in Q3 once ticket volume grows past 200/week, when the case for automation gets stronger."
If you land on defer, put a trigger and a date on the calendar. "Revisit when X happens" beats an open-ended "maybe someday."
Common pitfalls to avoid
Skipping the pilot and going straight to a contract. A one-week trial costs you almost nothing. A one-year subscription or a custom build you didn't test costs real money and momentum if it flops.
Confusing "impressive" with "needed." A slick AI feature that isn't tied to your actual bottleneck is a distraction. If your pilot didn't touch a real pain point from your own list, you're chasing hype, not solving a problem — see the ignore criteria in the efficiency, scale, insight, or risk framework for how to tell the difference.
Building because it feels more "serious." Custom development feels like progress, but it's the highest-cost, highest-risk option. Reserve it for cases where the need is both real and tied to something competitors can't easily copy.
Treating one good pilot week as proof it'll scale. A single week with one person testing drafts is a signal, not a guarantee. Before you roll a winning pilot out to the whole team, run it through a proper scale check — our pilot-to-scale checklist covers what to verify before you commit further budget.
Forgetting to track the numbers. "It felt faster" isn't a decision basis. Even rough numbers — minutes per reply, percentage needing edits — turn a gut call into something you can defend later, and something you can compare against the next tool you consider.
The bottom line
You don't need a strategy offsite to decide on your next AI tool. You need one narrow task, one week with an existing tool, and three honest questions at the end. Buy when the need is common, build when it's core to your edge, defer when the case isn't there yet — and write down why, so future-you isn't relitigating the same decision from scratch.