Guide

How to scope your first AI project so it actually works

Scoping your first AI project is less about chasing the single highest-value process in the business and more about picking the one most likely to actually work, end to end, in a way people notice. The job of a first project isn't to deliver the biggest return — it's to prove the approach works inside your business and build the internal confidence that makes a second, more ambitious project possible. The second project is usually where the real money is, and it only happens if the first one lands.

· Reviewed by Artur Horimoto, Founder & CEO

The short answer

A good first AI project is small enough to finish in weeks, narrow enough that everyone agrees on what "done" looks like, and attached to work someone is already doing today — not a greenfield idea nobody actually asked for. Screen candidates against five factors before you screen them on potential return: how often the task happens, how clearly it can be bounded, how little damage a mistake causes, whether someone genuinely wants it built, and whether the result is visible to someone whose opinion carries weight. Run a real candidate list through those five questions and the shortlist gets shorter fast — usually down to something considerably less exciting than whatever started the conversation, and considerably more likely to ship.

What makes a good candidate for a first project

None of these factors decide it alone, but a candidate that's weak on more than one is a warning sign, not a detail to fix in the build.

  • High frequency. A task that happens dozens of times a week teaches you more, faster, than one that happens twice a month. Frequency is also what makes a small win visible — nobody notices a system that ran once, quietly, three weeks ago.
  • Clearly bounded. You can describe the task in a sentence, name its inputs and outputs, and point to where it starts and ends. If explaining the task takes a paragraph full of "well, it depends," it isn't bounded yet — narrow it further before it becomes a project.
  • Low blast radius if it fails. The first system you ship will get something wrong. A good first project is one where that mistake is annoying, not dangerous — a slightly wrong draft someone reviews, not a payment sent to the wrong account or a promise made to a customer that can't be walked back.
  • An owner who wants it. Every first project needs someone on the business side who is glad it's happening, not someone who inherited it. An owner who wants the result will answer questions quickly, flag when something looks wrong, and defend the project when it's inconvenient — all things a reluctant owner won't do.
  • A result visible to someone who matters. The point of a first project is to earn the case for a second one. That case gets made in front of whoever controls the next budget decision, so the result needs to be something they can see and understand, not a backend improvement only the original team will ever notice.

Why your most valuable process is usually the wrong first choice

The process everyone wants to fix first is usually the one with the biggest number attached to it — and that's exactly why it's rarely the right place to start. High-value processes are high-value because they're complicated: several systems, several stakeholders, judgment calls that took years to develop, and enough political weight that everyone has an opinion about how it should work. All of that makes a process worth fixing eventually. None of it makes a process easy to prove an approach on for the first time.

Starting there stacks every risk in one project: technical risk (untested integrations), adoption risk (multiple teams who all have to change how they work), and reputational risk (if it goes badly, everyone who mattered was watching). If that project stalls, the organisation doesn't conclude "we picked the wrong first project" — it concludes "AI doesn't work here," and the second project never gets funded. Ranking the whole list of candidates against each other, rather than defaulting to whichever one is loudest, is exactly the gap an AI strategy engagement is built to close before anyone commits to building the wrong one first.

The fix isn't to ignore the valuable process — it's to sequence it. Prove the approach on something smaller and more forgiving first, then bring the evidence, the internal trust, and the lessons about what your data and systems actually look like to the harder project once there's a track record behind you.

Writing a scope that can actually be finished

A scope document earns its keep by saying what the system will do — and, just as explicitly, what it will not do. Most first projects that drag on for months didn't fail technically; they never finished being scoped, because every new conversation added one more case the system was now supposed to handle. Writing the exclusions down converts "we'll figure that out as we go" into a decision made once, in writing, before anyone starts building.

In practice that means naming the specific inputs the system will handle and stating plainly that anything outside that list gets routed to a person instead of guessed at. It means naming the systems it will read from and write to — and the systems it explicitly won't touch in this version. It means deciding, before the build starts, what happens on a case the system isn't confident about: does it guess, or does it hand off? A scope that only says what the system does, without saying what it deliberately leaves out, isn't really a scope — it's a wish list with a deadline attached. This is the same discipline behind how we structure every engagement: a written scope, agreed before build work starts, that anyone in the business could read and understand.

Agreeing on what "done" means before you build

"Done" needs a definition the same week the scope gets written, not a debate after the system is already live and someone is unhappy with it. Without that definition in writing, every stakeholder quietly carries their own version of success into the project, and the arguments after launch are really just those unstated definitions colliding for the first time.

Write the success criteria down before the build starts: what the system needs to handle correctly, how errors get measured, and what the acceptable range looks like for a first version rather than a mature one. Put a number or a clear condition on each one wherever you can, so "it's working" is a fact you can check rather than an opinion someone argues for. If you're not sure yet how to define success for a specific system, working through how to calculate the return on an AI project before the build starts is worth the hour — it forces the success criteria into a form you can actually measure once the system is live, instead of guessing after the fact.

Bring in the people whose work changes, early

A system that technically works and gets routed around by the team that was supposed to use it is still a failure — arguably a worse one, because it looks like a success on paper while doing nothing. The people whose day-to-day work the system touches need to be in the room while it's being scoped, not introduced to a finished tool in a training session the week before launch.

That means asking, before anything gets built, what the system should never do without a human checking first, what would make someone trust it less, and what part of the current process they'd actually be glad to hand off. People are far more willing to work alongside a system they helped shape than one that arrived as an announcement. This matters just as much for a straightforward automation as it does for a more autonomous AI agent making decisions inside a workflow — the more a system acts on its own, the more the people around it need to trust how it behaves before they'll actually rely on it.

A worked example: from vague ambition to a scopeable project

"We want AI to help with customer service" is an ambition, not a project — it's too large to build, too vague to finish, and nobody could tell you what "done" looks like. Narrowing it into something scopeable usually takes a few rounds of the same question: what, specifically, happens most often, costs the most time, and is safe enough to hand to a first version of a system?

  1. Start broad. "AI for customer service" could mean live chat, phone calls, email triage, refund decisions, or an internal knowledge base for the support team — five different projects wearing one sentence.
  2. Narrow to a workflow. The team says the phones ring constantly after hours, and every call that goes to voicemail is a lead or a frustrated customer lost until morning. That's a real workflow, not just a category.
  3. Bound it. Not "handle every after-hours call" — start with the two or three call reasons that come up most often (hours, availability, simple order status), and route everything else straight to voicemail or a human the next morning, exactly as it works today.
  4. Check the five factors. It happens every night (frequent), the boundary is a short, specific list of call types (bounded), a missed edge case just falls back to today's voicemail (low blast radius), the office manager who deals with the voicemail backlog every morning wants it gone (an owner), and fewer missed calls is something the owner will notice inside a week (visible).
  5. Write what it won't do. No billing changes, no account access, no promises about pricing or availability the system can't confirm — anything outside the short list gets handed to a person, exactly as an unanswered call does today.

That's a scopeable first project. "Handle all customer service with AI" never was one.

Plan the second project while the first one is running

The best time to think seriously about the second project is while the first one is still in progress, not after it ships. A live first project generates evidence a planning document never could: whether the data was as clean as everyone assumed, whether the team actually adopted the system or worked around it, and whether the integration was as simple as it looked on paper. Every one of those answers changes what the second project should be and how confidently you can scope it.

Use that evidence deliberately. Revisit the wider list of candidate processes — the readiness checklist is a reasonable place to work through which one is next in line — and re-rank it with what you now know instead of the assumptions you started with. Some candidates will look more promising once you've seen a system run against real data; others will look harder than they did on a whiteboard. That's exactly the information the first project was supposed to produce, and it's also usually the point where the build-versus-buy question is worth revisiting for whatever comes next — a bigger or more judgment-heavy workflow often tips that decision differently than the first one did.

Frequently asked questions

How do I know if our first AI project is scoped correctly?

If you can describe what it will do, what it explicitly won't do, who owns it, and what "done" looks like in a single page everyone in the room agrees with, it's scoped correctly. If any of those four things is still fuzzy, or the description needs several caveats to explain, narrow it further before you start building.

What if the process we most want to fix doesn't meet these criteria?

Keep it on the list — just not first. Sequence it behind a smaller, more forgiving project that proves the approach and builds the internal evidence and trust that make the harder one easier to scope well. Most teams get to the ambitious project faster this way than by attempting it cold.

Who should own a first AI project internally?

Someone close enough to the workflow to answer questions quickly and make small decisions without a committee, and who actually wants the result — not whoever happened to be free. A reluctant owner slows every stage of the project down and is the first person to let the team quietly route around a system once it's live.

How long should a first AI project take?

Short enough that the organisation doesn't lose interest before it ships. Most well-scoped first projects are live in weeks, largely because a narrow, well-bounded scope is what makes that speed possible in the first place — size and clarity of scope matter far more than the underlying technology.

What happens if the first project doesn't work as well as hoped?

That's useful information, not a verdict on AI generally, provided the success criteria were written down in advance so everyone can look at the same facts. Because the blast radius was kept small by design, a first project that underperforms is a lesson you adjust from — not a setback the whole initiative has to recover from.

Bring the workflow you're considering for a first AI project to a free 30-minute strategy call. We'll tell you honestly whether it's scoped to succeed, what to cut if it isn't, and what the second project should probably be once this one is live.

Get a straight answer on your project

Guides only go so far. Bring your numbers and we will scope it properly.

Free 30 minutes. No pitch deck. You leave with a plan either way.