How to Vet an AI Tool Before Rollout

Photo: Mikhail Nilov / Pexels
Key takeaways
- Vetting an AI tool starts with naming the data that will actually go through it, not the vendor's reputation.
- Check training posture, DPA, BAA (if health data), SOC 2/ISO 27001 and data residency, sourced from the vendor's own documentation.
- Almost every finding applies to one plan tier, not the product as a whole - pin it down and document it.
- Write a one-paragraph verdict so the same tool doesn't get re-debated in six months.
- For a vendor with a public trust centre, the whole process realistically takes 15-30 minutes.
The moment an AI tool goes from 'one person trying it out' to 'the whole team relies on it', the risk changes completely: more data, more people, more assumption that someone already checked it. Knowing how to vet an AI tool properly is the difference between that assumption being true and it being a guess. Most teams either skip the step entirely, which is how shadow AI starts, or turn it into a weeks-long review that people quietly route around. Here's a middle path fast enough to actually get used, built around five short steps.
Step 1: name the data before anything else
Before checking a single vendor policy, get specific about what will actually go through the tool: customer details, health information, source code, financial records, or just general drafting like blog outlines and internal memos. The data type decides which questions matter next. Health data needs a HIPAA question. EU personal data needs a GDPR question. Proprietary code needs a training-and-retention question, since a code-completion tool with the wrong default can quietly turn your codebase into someone else's training set. Skip this step and you end up checking the wrong things thoroughly, which feels productive and isn't.
Step 2: check sourced facts, not reputation
Training posture: does it train on inputs, and does that depend on the plan? A BAA, required only if PHI is in scope, per HHS guidance on business associates. A GDPR DPA, required if EU or UK personal data is in scope - the ICO's guidance on AI and data protection is a useful reference for what a DPA needs to cover. SOC 2 or ISO 27001, a baseline security-hygiene signal assessed against the AICPA's Trust Services Criteria, not proof of the above. Data residency and subprocessors, relevant if you have jurisdiction-specific rules about where data physically sits. Get every answer from the vendor's own trust centre or privacy page, not a colleague's impression of the product from a demo six months ago. ModelCharter's AI Tool Risk Directory has already done this lookup for the most common tools, so check there first; it may save this whole step.
How long should vetting actually take?
For a vendor with a public trust centre and a documented DPA, the honest answer is fifteen to thirty minutes once you know what to check. The weeks-long version most teams dread happens when nobody's agreed what 'checked' means, so the same questions get re-asked by three different people in three different tools that never talk to each other. Standardise the questions once, in a template, and the time collapses.
Step 3: pin the tier, not just the tool
Almost every finding from step 2 applies to a specific plan, not the product as a whole. Decide and document which tier is approved, 'Business or Enterprise only', for instance, and check whether anyone already using the tool is on a different one: a free or personal account someone signed up for before the review even started, back when it was just 'that thing Dave uses'. Write the tier into the approval itself - see consumer vs business AI tiers for why 'we approved ChatGPT' isn't a specific enough answer on its own.
The single biggest mistake
The most common failure isn't skipping the check. It's checking the tool once, at a moment in time, and never again. A vendor's privacy policy from the sign-up email is not a permanent contract; policies get updated, sometimes to your advantage and sometimes not, and a tool that passed a review in January can silently change its training defaults by June without a single notification landing in anyone's inbox. Step 5's recheck trigger exists specifically to catch this, but it only works if someone actually owns it.
Why 'everyone uses it' isn't a verdict
A tool being popular, well-reviewed, or already used by a company much bigger than yours tells you almost nothing about its data policy for your plan. Large customers frequently negotiate custom contracts that a self-serve sign-up form never sees, and a tool's mainstream reputation is built on product quality, not on how it handles your specific type of data. A glowing review from a friend at another company usually describes a different plan, a different data type, or both. Treat 'everyone uses it' as a reason to check faster, since a sourced trust centre probably already exists, not as a reason to skip checking altogether.
Step 4: write the verdict down
A one-paragraph record - what was checked, what the source said, which tier it applies to, and the resulting verdict (approve, conditional, reject, or needs more information) - turns a one-off conversation into something an auditor, a customer security questionnaire, or your own future self can actually rely on. A support lead at a 30-person SaaS startup we spoke to had vetted the same AI notetaker twice in a year, eleven months apart, because the first decision only ever existed in a Slack thread nobody could search, and the person who remembered it had since left the company. The second review took the same fifteen minutes as the first - fifteen minutes a two-line written verdict would have saved entirely.
What if the vendor doesn't publish a DPA or BAA?
Ask directly, and keep the written reply as your source, whether it's an email or a support-ticket transcript. If there's no answer, or the answer is vague ('we take security seriously'), that's not a technicality to wave through; treat it as an open risk and record it as 'needs more information' rather than a silent approval that quietly becomes permanent because nobody revisits it.
Step 5: set a recheck trigger
Rechecking is triggered by whichever comes first: a fixed cadence (six or twelve months is typical for a full re-review), a vendor terms-of-service or privacy-policy update, a publicly disclosed security incident, or a plan-tier change on your side - upgrading, downgrading, or simply adding new seats. None of this needs weeks. A free, structured template at the AI vendor risk assessment tool walks through exactly these steps, and it feeds a living tool register rather than a one-off document that ages the moment it's written. Cross-check against our AI tool security checklist for the deeper technical questions worth asking any vendor handling sensitive data.
| If the tool will process... | Ask specifically about... | Where to find the answer |
|---|---|---|
| Customer or EU personal data | GDPR lawful basis and a signed DPA | Vendor's DPA or enterprise-privacy page |
| Health information (PHI) | A signed Business Associate Agreement | Vendor's business-associate or trust page |
| Source code or proprietary text | Training and retention posture, and any opt-out | Vendor's enterprise or API privacy documentation |
| General internal drafting | SOC 2/ISO 27001 baseline and retention window | Vendor's trust centre |
“A fact you can't source isn't a fact yet. It's a rumour with good marketing.”