The pain is addition, not absence
Retail is the most AI-saturated category we work in, and that is exactly the problem. A typical mid-size store's helpdesk vendor sells an AI reply suggester. The email platform sells an AI send-time optimizer and subject-line generator. The reviews tool sells AI summarization. Search sells AI-ranked results. The ad platform sells automated bidding that calls itself AI because that sells better than "algorithm." Each one arrived bundled into a plan the team was already paying for, or pitched in a renewal call as the reason for a price increase.
None of that is fraud — most of these features do something. The problem is nobody added them up. A brand can be carrying six or seven AI line items across its stack, unable to say which ones are actually turned on, which ones changed a number when they were enabled, and which ones are dead weight the account manager talks about on every renewal call. Before Calfy scopes a single new build, the engagement answers that question first.
What the audit actually checks
The first working session is not a whiteboard of AI ideas. It is a walk through the stack you already run — helpdesk, email and SMS, reviews, search, ad buying, and whatever else carries an "AI-powered" badge in its marketing — and for each one, three questions: is this switched on, is anyone using what it produces, and does turning it off change anything a person would notice.
That third question does most of the work. A send-time optimization feature nobody has compared against a fixed schedule is not proven, it is assumed. An AI reply suggester sitting inside a helpdesk that agents click past because the tone is wrong is a subscription line, not a capability. The audit is not there to shame anyone for the accumulation — it happens to every fast-growing store — it is there to produce an honest starting inventory before anyone recommends spending more.
What the engagement produces
The output is not a list of AI ideas. It is three things you can act on directly. First, an inventory of what you already pay for, what it does, and a plain recommendation for each line: keep it configured properly, reconfigure it, or cut it. Second, a build-versus-buy call on the handful of gaps that survive the audit — the workflows no existing tool in your stack actually covers. Third, a sequencing plan that says what to touch first, and what to deliberately leave alone until after your next peak season.
For most stores, the audit alone pays for a meaningful share of the engagement, because cutting or reconfiguring unused AI features shows up on a bill immediately. What is left after that — the genuine gaps — is usually a shorter and more specific list than the one the brand walked in with, which is what makes the roadmap that follows fundable rather than aspirational. This is the same discipline behind the AI strategy work we run across every industry, applied to a market where the starting point is subtraction rather than a blank slate.
The build-versus-buy line sits closer to buy than in most industries
We will say this plainly even though Calfy is the build side of that decision: ecommerce is the industry where buying an existing tool is right more often than building custom. Order lookups, review replies, subject-line generation, basic search ranking, ad bid optimization — these are workflows thousands of stores share, and vendors who serve thousands of stores have had years to refine them. A narrow, well-understood problem with a forgiving failure mode is close to the textbook case for buying, and in ecommerce that describes most of the stack.
That is not a hedge before the sales pitch. It is the honest read of where the market actually sits, and our build-versus-buy framework applies the same test everywhere we work: how standard is the workflow, how deep does it need to reach into systems only you run, and how much of the logic is actually specific to how your business operates versus common retail practice. In ecommerce, that test resolves toward buy for the majority of what a vendor already offers. What is left over is smaller than most brands expect — and worth building precisely because it is left over.
Where custom genuinely wins
Three patterns keep showing up in the minority of cases where a vendor tool cannot do the job.
Deep integration across storefront, 3PL, and ERP. A vendor's AI feature typically reasons over the one system it lives inside — the helpdesk sees tickets, the email tool sees campaigns. The moment an answer depends on stitching your storefront, your third-party logistics provider's shipment data, and your ERP's inventory truth together in real time, no bundled feature reaches far enough. That is custom agent territory, not because the individual lookups are hard, but because no vendor has a reason to build the specific connective logic across systems only you combine that way.
Product knowledge specific to your catalogue. A generic AI answer engine knows general facts about general products. It does not know that your size 10 runs narrow, that SKU 4471 is a seasonal color that will not restock, or that two products in your own catalogue are functionally interchangeable substitutes a shopper should be offered when one is out of stock. That knowledge lives in your data and your team's heads, not in a vendor's training set, and building it in is what makes a pre-purchase Q&A agent trustworthy instead of generically plausible.
Operational logic no vendor model carries. Bundle and kitting rules, subscription-cadence exceptions, loyalty-tier treatment, which supplier gets the reorder for which SKU — this is the accumulated judgment of how your specific operation runs, and it is not something an off-the-shelf AI feature was built to encode, because every store's version of it is different.
Data readiness: what an inconsistent catalogue needs first
Almost every catalogue we look at has the same underlying problem: attributes that were filled in by three different people over four years, migrated across at least one platform change, and never fully reconciled against what marketplaces expect. Size fields that are sometimes a dropdown and sometimes free text. Material and care fields populated for a third of SKUs. Stock-status flags that lag the warehouse system by a day.
None of that stops a strategy engagement from happening, but it changes what gets sequenced first. A pre-purchase Q&A agent trained against a catalogue with unreliable attributes will answer confidently and sometimes wrong, and a confidently wrong answer about fit or stock costs more than a slow honest one. Where the audit finds catalogue data this unreliable, the roadmap puts a scoped data-cleanup pass ahead of the agent that depends on it, rather than building on a foundation everyone already knows is shaky.
Sequencing: nothing launches in Q4
Peak season sets the calendar for everything else in this roadmap, and the rule is simple: nobody launches a new AI system in the run-up to Black Friday, Cyber Monday, or the holiday peak that follows. That window is when contact volume is highest, when a wrong answer costs the most in refunds and lost trust, and when your team has the least attention to spare for teaching a new system its edge cases.
New builds get sequenced to go live and stabilize well before peak, or deliberately scheduled for January once the season's data is in and the team has recovered. What is safe to change during peak itself is narrow: configuration tweaks to AI features you already run and already trust, not new agents learning your operation for the first time under the worst possible load.
Walkthrough: the audit that paid for itself before anything got built
A skincare brand selling through its own storefront and two marketplaces came to us wanting "an AI agent for customer support." The first working session mapped the stack instead: a helpdesk plan that included an AI reply suggester nobody had configured past the default, an email platform's send-time optimizer running untested against a fixed schedule, and a reviews tool's AI summary feature the team had never opened. None of it was doing anything measurable.
Reconfiguring the helpdesk feature properly — feeding it the brand's actual return policy and tone guide instead of the vendor default — closed a real share of the gap the brand had assumed needed a custom build. What was left after that was narrower and specific: order-status lookups that needed to reach into a 3PL the helpdesk vendor had no visibility into. That became the one thing worth building, and it shipped as a scoped agent instead of a rebuild of the whole support function.
Walkthrough: the catalogue question that changed the sequence
A multi-channel outdoor-gear retailer shortlisted a pre-purchase Q&A agent during the opportunity audit — shoppers keep asking sizing and compatibility questions the support team answers the same way, dozens of times a day. Before scoping the build, the data-readiness check pulled a sample of the catalogue against what the agent would need to answer confidently: sizing, material, and compatibility fields.
About a third of the SKUs had those fields filled in fully; the rest were partial or blank, inherited from a platform migration two years earlier. Building the agent against that catalogue as-is would have meant confident wrong answers on a third of the store's inventory — worse than the status quo. The roadmap moved a scoped attribute-cleanup pass ahead of the agent, sequenced to finish before the busy season rather than during it, so the agent launched against data it could actually trust.
What this becomes: connecting the audit to a build
The roadmap that comes out of this engagement points in one of a few directions, and which one depends entirely on what the audit found — not on which build looks more impressive. Where a gap needs judgment applied across systems no single vendor tool reaches, it becomes a custom AI agent. Where the gap is a fixed, repeatable process — a reorder alert, a supplier follow-up, a catalogue consistency check — it is usually simpler and cheaper as a workflow automation with no conversational layer at all. And where the honest finding is that an existing tool already covers it, the roadmap says so and the recommendation is to configure what you already pay for properly, not buy again.
Whatever gets built inherits the audit's findings directly: which catalogue fields needed cleanup first, which systems the build has to reach, and where the sequence sits relative to your next peak season. None of that gets redone once build work starts. Some brands take the roadmap and build with an internal team instead of Calfy, and the audit and sequencing work stand on their own either way. The wider picture of how this fits day-to-day retail operations is covered on our ecommerce industry page, for teams weighing this against everything else competing for budget.