The constraint is defensibility, not capability
Most AI strategy conversations start with a demo. Someone shows what a model can pull out of a document or draft in a chat window, and the question that follows is how fast the firm can get some version of that live. For a regulated financial firm, that is the wrong question to start with.
The question a risk committee actually asks is narrower and harder to answer: if this system gets something wrong, can the firm show exactly what it saw, what it did, and who signed off before it reached a client? A capability that cannot answer that question is not ready for this firm, no matter how good the demo looked. This engagement is not primarily about finding what AI can do — plenty is possible. It is about finding what is defensible, and building the case for it before anyone writes a line of code.
What the engagement actually produces
The engagement produces three things a risk committee, a compliance officer, and an operations lead can each read and act on. First, a ranked list of opportunities scored against both business value and risk exposure — not just how much time a change would save, but what happens if it is wrong and how the firm would find out. Second, a written decision for each opportunity about where a human stays the final decision-maker and where a system can act inside limits set in advance. Third, a sequencing plan that starts with the project least likely to create a problem nobody can explain later, so the firm builds confidence in how it oversees AI before it takes on anything more ambitious.
None of that is a compliance sign-off, and it is not offered as one. Calfy works alongside your firm's compliance function; it does not replace it, and this engagement is not legal or regulatory advice. What it produces is the material your own compliance and risk people need to make that call, organized the way they would actually need to see it — which sits alongside the wider AI strategy work we run for businesses outside financial services too.
Opportunity assessment filtered through risk exposure
The audit starts the way any AI opportunity audit does — walking the workflows that eat the most hours or generate the most rework. What changes for a financial firm is the second filter every candidate has to pass: not just whether it is worth automating, but what the exposure is if the automated version gets it wrong, and how the firm would know before a client or an examiner does.
A workflow that drafts internal notes from data the firm already holds carries a different exposure profile than one that touches what a client is told about their own money. Scoring separates those cases early, rather than letting the difference surface mid-build. What survives is usually a shorter list than the one the firm walked in with. Some ideas get parked — not because the technology cannot do them, but because the firm is not yet the right shape to run them safely, and saying that up front is cheaper than finding out after launch.
Where a human decision-maker has to stay in the loop
Every opportunity gets a specific answer to one question: does a person make the final call, or does the system act inside limits a person set in advance? This is not a blanket rule applied to the whole firm — it is decided workflow by workflow, based on what the output actually does. A system that assembles the facts behind a recommendation is a different case from a system that decides what the recommendation is. A draft a person reviews before it reaches a client is a different case from a message that goes out on its own.
Where advice, suitability, credit, or a similar judgment call is involved, the decision-maker stays human. Where the work is gathering, checking, or drafting from records the firm already holds, a system can carry more of the load, provided the boundary between "prepared this" and "approved this" is written down rather than assumed. That human-in-the-loop boundary goes into the strategy document itself, workflow by workflow, so it is not reinvented — or quietly skipped — the first time a live deadline puts pressure on it.
Governance and model oversight sized to the team you actually have
A lot of AI governance advice reads as though it were written for a firm with a dedicated model-risk function and the headcount to match. Most financial services firms we work with do not have that, and a governance structure nobody can actually staff is worse than no structure at all — it creates a paper process that stops matching reality within a quarter.
The governance work in this engagement is sized to who you actually have: who reviews what a system produced before it goes anywhere, what gets logged and for how long, who is accountable when something goes wrong, and how often the whole arrangement gets checked against what the system is actually doing versus what it was approved to do. For a small advisory practice that might be one named reviewer and a short log. For a larger firm it might be a standing review against a documented policy. Either way, the guardrails we design are meant to survive a normal week under real workload, not just a pilot demo in a quiet month.
Vendor due diligence and concentration risk
Most firms building an AI capability are not writing every layer themselves. They are combining a model provider, a document-processing tool, perhaps a CRM's built-in AI feature, and custom logic that ties it together. Each of those is a vendor relationship the firm now depends on, and depending on several of them inside the same critical workflow creates concentration risk that is easy to miss until one of them has an outage or changes its terms.
Part of the strategy engagement is mapping that dependency honestly: which vendors sit in the path of a client-facing decision, what happens if one of them is unavailable or changes its product without warning, and where the firm would be exposed if a single vendor problem took out more than one workflow at once. That map becomes part of what your own procurement and risk review needs to assess it properly. We hand over the dependency picture; approving any given vendor stays your firm's own call.
Data residency and third-party processing
Where client data physically sits, and which third parties process it along the way, is a question every AI tool answers differently, and it matters more for a financial firm than for almost any other kind of business. A strategy engagement surfaces this for every candidate system before it is built, not after: what data the workflow needs to touch, whether it leaves the firm's own systems at any point, and if it passes through a third-party processor, what that processor does with it and where.
We do not tell a firm what its own data residency obligations are — that is a question for the firm's own legal and compliance advice, and it varies by firm and jurisdiction. What we do is make sure the question gets asked and answered for every system on the roadmap before it is built, so it becomes a design input rather than a surprise raised during a later review.
Staff training and an acceptable-use policy
A governance document that lives in a shared folder does not change what a coordinator does under deadline pressure. The people actually using a system need two things: training specific to what they will be doing with it, and a short, plain-language acceptable-use policy that tells them what the system is for, what it is not for, and what to do when its output looks wrong.
We write that policy as part of the engagement — not a generic responsible-AI document, but one scoped to the specific systems in the roadmap: what a person checks before approving a draft, when to escalate instead of accepting an output, and what never gets pasted into an external tool regardless of how convenient it would be. The measure of whether it worked is not whether people signed it. It is whether they can explain it back a month later, unprompted, and still follow it when nobody is watching.
Sequencing: the first project earns confidence, not headlines
Ranked opportunities do not get built in order of ambition. They get built in the order that teaches the firm the most about its own oversight while the stakes are still low. The first project should be small enough to see running within weeks, low-stakes enough that a mistake is cheap and visible rather than hidden, and useful enough that the team notices it working.
That first project also proves things a planning document cannot: whether the review step actually gets followed under real workload, whether the logging is good enough to answer a question six months later, whether staff trust the output enough to use it properly instead of rubber-stamping it. Everything after it gets sequenced with that evidence in hand. The most ambitious idea on the list — often the one that started the whole conversation — usually belongs later, once the firm has practice running an AI system inside its own oversight rather than learning that oversight for the first time on the hardest case.
Walkthrough: scoping a client-reporting assistant against risk exposure
A mid-sized wealth management practice comes to us wanting "AI for client reporting." The first working session does not start with a build. It starts with mapping exactly what the current reporting pack contains, who signs it before it goes to a client, and what happens if a figure in it is wrong. That mapping produces the risk score: this workflow touches client-facing statements about their own money, so a drafted pack always needs a named reviewer's sign-off before release, with no exception carved out for busy weeks.
From there we map what would actually make the reviewer's job faster rather than replace it: data pulled and pre-populated from the firm's source systems, narrative sections drafted in the firm's own voice, and a clear flag on anything the system could not reconcile cleanly against the underlying records. The strategy document that comes out of the session specifies exactly this — what the system drafts, who reviews it, what triggers escalation instead of drafting, and what gets logged so a reporting error found later can be traced back to what the system saw and what the reviewer changed.
Walkthrough: the vendor question that reshaped a shortlisted project
An accountancy practice shortlists a document-extraction tool for onboarding files during the opportunity assessment. It looks like an easy win — quick to set up, a clean demo, a reasonable price. The vendor due diligence step in the engagement asks a few more questions before anything gets approved: where does the tool actually process documents, does it route through a subprocessor the firm has never vetted, and what else in the firm's stack would be affected if that one vendor had an outage or a security incident.
The answers matter. The tool's extraction step ran through a subprocessor the firm had no visibility into, and that same subprocessor sat behind two other tools already in the firm's stack — meaning a single incident there could take out three workflows the firm had assumed were independent of each other. That is concentration risk the firm would not have seen without asking the question at the strategy stage instead of after the contract was signed. We did not rule the tool out ourselves. We put the finding in front of the firm's own procurement and compliance review, with the dependency map attached, so the firm could decide with full visibility rather than a vendor's own assurances.
What this becomes: connecting the roadmap to a build
A strategy engagement is only worth doing if it leads somewhere. For most of the firms we work with, it leads directly into a build — often a custom AI agent for a workflow that genuinely needs judgment inside defined limits, and just as often a simpler workflow automation for the parts of the roadmap that follow a fixed set of steps and do not need an agent making decisions at all. Part of an honest strategy engagement is telling you which is which, rather than defaulting to the more impressive-sounding build.
Whatever gets built inherits everything the strategy stage already established — the risk scoring, the human-in-the-loop boundaries, the governance structure, the vendor and data-residency answers, the acceptable-use policy. None of that gets redone once build work starts; it is the specification the build works from. Some firms take the roadmap and build with their own team or another vendor entirely, and the strategy work stands on its own either way. The wider picture of how this fits a regulated firm's day-to-day operations is covered on our financial services industry page, for firms weighing this against other priorities.