What a Retail Security Review Should Ask an AI Vendor

What a Retail Security Review Should Ask an AI Vendor

Your questionnaire was written for software that stores data and shows it back. An agent asks for the ability to act on your systems of record. Six questions that cover the difference, in the order the answers matter.

See how Ward detects read-only architecture and gated write-back

Get a demo → Take the 3-minute assessment
Contents

Your questionnaire was written for a SaaS tool, not an agent

Most retail security questionnaires were built for software that stores your data and shows it back to you. Where does it live, who can see it, is it encrypted, what happens when we leave. Those questions still matter, but they no longer cover the risk, because an AI vendor is asking for something a reporting tool never did: the ability to act on your systems of record.

That changes the shape of the review. The question is no longer only "what can you see" but "what can you change, who approved it, and can we prove that afterwards." A vendor who answers the first set fluently and goes vague on the second is telling you something.

What follows is the list we think a retail security review should work through, in the order the answers matter. It is written to be used against any vendor, including us.

1. Does our data move, and can you prove it does not?

The cheapest way to reduce risk is not to copy the data. Ask whether the vendor runs federated queries against your warehouse or ingests into their own store. If they ingest, every subsequent answer about residency, retention, deletion and breach exposure gets harder, and you now have two copies of your customer data to govern instead of one.

Ask for the proof, not the claim. A read-only service account you issue, locked to SELECT on named schemas, is verifiable in your own warehouse logs on day one. You should be able to see every query the vendor ran, when, and against what, without asking them for a report.

What a good answer sounds like

"You issue the account. It has SELECT on the schemas you name and nothing else. You can revoke it in your console without telling us, and every query we run appears in your warehouse monitoring under that account, because it runs as that account."

A bad answer describes an ingestion pipeline and then explains why the pipeline is secure. That may be true, and it is still a larger surface than not copying the data at all.

2. When it writes to SAP, who approved it?

This is the question that separates an analytics tool from an agent, and it is where most reviews stop too early. If a vendor can write to your system of record, the review has to establish four things, in writing:

  • What can it write. Scoped per playbook and per resource, or a general grant?
  • Who approves each write. A named role you defined, or a vendor-side setting?
  • What is recorded. State before and after, or an event saying something happened?
  • How you turn it off. A support ticket, or a change you make yourself?

The strongest version of this we know of is policy as code in your own repository. Agent scope lives in your Git, changes by pull request through the review process you already run for everything else, and rolls back the way the rest of your code does. That turns a vendor configuration question into a change management question your organisation already knows how to answer.

3. What happens when the model gets a number wrong?

Ask which numbers the language model produces. The right answer is none of them.

Forecasts and aggregates should come from classical statistical models: ARIMA, Holt-Winters, Bayesian hierarchical, gradient-boosted residuals. The language model's job is to frame results it did not compute. If the vendor's architecture has an LLM generating figures, no amount of prompt engineering makes that auditable, and your merchants will find the first wrong number before your auditors do.

Then ask for the error rate. A forecast without a published mean absolute percentage error, backtested over a couple of years of your kind of seasonality, is an opinion. You are entitled to see how wrong it has been before anyone acts on it.

4. Can we re-derive any number without you?

Every figure should be one click from the SQL that produced it, the source tables it hit, and the parameters used. Not a description of the query. The query.

This matters for two reasons that have nothing to do with trust. It is how your data team debugs a disagreement between the vendor's number and yours, which will happen in week two. And it is how a number survives an audit eighteen months later when the person who accepted it has left.

A number nobody can reproduce is a rumour, however good the tool that produced it.

5. Whose identity system governs access?

If the vendor maintains its own user directory, you have a second access surface that drifts from your first one the moment someone leaves. Ask for SSO/SAML and SCIM as table stakes, and ask specifically whether deprovisioning is automatic. "We support SSO" and "removing a user in your IdP removes their access here within minutes" are different claims.

Ask where the audit stream goes. It should land in your SIEM, in a format you can query, from the first day rather than at go-live. If the audit trail lives only in the vendor's console, your investigation capability depends on their availability at the exact moment you need it most.

6. What does leaving actually involve?

Run the exit conversation before signature, when you still have leverage. The useful question is not "can we export our data" but "what did you build that we would have to rebuild."

If the vendor never held a copy of your data, exit is not a migration. The warehouse is yours and unchanged, the policies are in your repository, and what you lose is the findings rather than the infrastructure. If instead there is a vendor-side model trained on your history, a semantic layer only they can read, or a second warehouse, price that rebuild now and put the number in the business case.

Ask for all of it before the contract, not after

Everything above should be answerable from documents that exist before there is anything to sign: a data flow diagram, network topology, the policy bundle as deployed, the sub-processor list, SOC 2 status, a pre-answered questionnaire, the MSA and DPA, and the certificate of insurance.

A vendor who can only produce those after a signed order form is telling you the documents do not exist yet. That is survivable in an early company, and it is not the same as a vendor who has them and would rather you did not read them closely. The way to tell the difference is to ask for the specific artifact rather than the category, and to notice which questions produce a document and which produce a meeting.

See how Ward detects read-only architecture and gated write-back

Get a demo →

Using this against us

We publish our answers to all six on the evaluation packet and the security posture, and the document room opens on a click-through NDA without a sales call in between. Where our answer is "not yet", it says not yet. SOC 2 Type II is underway rather than complete, and the page says so rather than implying otherwise.

Run the list against whoever is on your shortlist. The questions are not ours and the answers should not have to be taken on faith from anyone.

See how Ward detects read-only architecture and gated write-back

Ward monitors your stores 24/7 and delivers insight cards, not dashboards. First cards in 48 hours.

security review CIO vendor evaluation AI governance SOC 2

Not sure where AI fits in your operation? Ten questions, about three minutes. Your score out of 100 appears on screen when you finish, with no email required.

Take the 3-minute assessment

From the article to the product.

How this topic maps to what Ward does, who it’s for, and the alternatives buyers benchmark against.

Your stores are generating data right now.

Ward turns it into decisions. First insight cards in 48 hours.

Read-only to start · your LLM keys · SOC 2 Type II underway · or book a call directly

Find out what your data has been hiding.

Tell us about your operation. We’ll show you the problems Ward catches, and the ones your current tools miss.

Step 1 of 3
What are your goals?
Step 2 of 3
About your operation
Step 3 of 3
Your contact info