By
August 1, 2026
10 min read
What to Ask an AI Vendor Before You Sign the Contract



What AI vendor risk actually means (and why it's not the same review you already run)
An AI vendor risk assessment is the process of evaluating whether a third-party AI product handles your data, your model outputs, and your compliance obligations safely enough to sign a contract. It differs from a standard vendor security review in one core way: the thing you're buying keeps changing after you buy it. A normal SaaS tool has a fixed feature set and a static data flow you can diagram once. An AI vendor's model might retrain on customer inputs, route requests through a subprocessor chain that shifts quarterly, or update its behavior with no changelog you'd ever see.
That distinction is why the checklist your procurement team has used for the last decade doesn't cover it. You can ask a normal vendor "where is our data stored" and get a durable answer. Ask an AI vendor the same question and the honest answer is often "it depends on which model we're routing to this week." OneTrust's framing of this gap is useful: AI governance has to be folded into vendor assessment as its own category, not bolted onto the existing security questionnaire as three extra rows.
If you've been through governing AI agents you build in-house, this is the mirror image of that problem. You control the architecture of what you build. You control almost nothing about what you buy, except the terms you negotiate before you sign.
Why this became a board-level conversation in 2025 and 2026
Three things converged at once: certification standards matured, state law caught up, and AI features got embedded into tools nobody formally re-vetted. Search interest in "ai vendor risk assessment" is up roughly 600% quarter-over-quarter and year-over-year as of a July 2026 pull from DataForSEO Labs, and the related term "ai vendor risk" is up about 400% YoY over the same window. Volume is still small in absolute terms, but the trajectory tells you this went from a niche compliance concern to a live procurement question inside two years.
The certification landscape moved first. NIST published AI RMF 1.0 in January 2023, and it's now the default US reference for categorizing AI risk in a vendor context. ISO/IEC 42001, the first international AI management system standard, followed and is increasingly what enterprise buyers ask AI vendors to certify against, per isms.online's breakdown of the standard's vendor-risk provisions. Financial institutions got a concrete workbook to use: FS-ISAC's Generative AI Vendor Evaluation & Qualitative Risk Assessment guide, published for banks and credit unions in early 2024, structures exactly the kind of questionnaire most mid-market companies are now improvising from scratch.
Regulation caught up next. Colorado's AI Act sets compliance obligations for high-risk AI systems ahead of a January 2027 deadline, which we've covered in detail in our state AI law rundown, and it explicitly extends liability to deployers, not just developers, meaning the company that buys the AI tool carries exposure for how it's used. That single fact turns vendor vetting from a nice-to-have into something your legal counsel will ask about directly.
And the third factor is the quiet one: AI got embedded into tools you already pay for. Your CRM added an AI lead-scoring feature in a routine update. Your HR platform turned on a resume-screening model by default. Nobody re-ran a vendor security review because, technically, you didn't sign a new contract. That's the exposure most companies in the $5M to $50M range are sitting on right now without knowing it.
The AI vendor questionnaire: what to actually ask
A useful questionnaire has seven categories, and most vendors can answer the first three fluently and stumble on the last four. That stumble is the signal you're looking for.
Training data use: Does the vendor train, fine-tune, or improve its models using your inputs or outputs, by default or opt-out? Ask for the specific contract clause, not a marketing page.
Retention windows: How long is your data held after a request completes, and does that window differ for logs, embeddings, and cached responses?
Subprocessor disclosure: Which third-party model providers, hosting platforms, or inference APIs sit behind their product, and how often does that list change?
Access logging: Can they produce an audit trail of which employees or systems accessed your data, and for what purpose?
Output and IP ownership: Who owns content generated using your data as input? This matters more than most contracts admit.
Deletion SLAs: If you terminate the contract, what's the contractual timeline for full data deletion across all subprocessors, not just their primary systems?
Incident history: Have they had a data exposure or model behavior incident in the last 24 months, and what changed as a result?
PwC frames this differently than most compliance teams do, and it's worth borrowing their angle: proactive AI third-party risk management reduces time-to-contract, not just risk exposure. A vendor who has already answered these seven categories in writing closes faster than one your legal team has to interrogate line by line. Treat the checklist as a deal accelerator you hand to serious vendors early, not a gate you spring on them at the end.
What a SOC 2 report from an AI vendor should actually cover
A generic SOC 2 Type II report tells you almost nothing about AI-specific risk unless you know what to check for inside it. SOC 2 was built for traditional software controls, access management, change management, availability, and it doesn't natively address whether a model trains on customer data, whether outputs are logged and reviewable, or whether a subprocessor chain includes a foundation model provider with its own data policies.
When you get a SOC 2 report from an AI vendor, look at the scope section first. Does it explicitly name the AI or ML components as part of the audited system boundary, or does the report cover the surrounding infrastructure while treating the model itself as a black box outside scope? Many AI vendors hand over a SOC 2 report that was scoped before they added AI features, and the report simply doesn't speak to the part you're most worried about. Ask directly whether the trust services criteria included confidentiality controls specific to model inputs and outputs, not just the hosting environment.
Red flags that should stop the deal
Some answers are disqualifying on their own, regardless of price or feature fit.
The vendor can't tell you, in writing, whether your data trains their model. "We take privacy seriously" is not an answer to a yes-or-no question.
There's no audit right in the contract. If you can't verify their claims after signing, the questionnaire answers are marketing copy.
The subprocessor list is undisclosed or described as "may change without notice" with no notification clause.
There's no contractual data deletion guarantee with a defined timeline, only a policy statement on a webpage that can change unilaterally.
They can't produce evidence of a prior incident response, or they claim they've never had an incident. Every vendor at scale has had something go sideways; the ones who say otherwise either haven't been asked hard enough by anyone else, or aren't being straight with you.
These overlap heavily with what an insurance underwriter wants to see when pricing your own AI liability exposure, which we walked through in our AI liability insurance guide. If a vendor's answers wouldn't satisfy your own underwriter, they shouldn't satisfy your procurement team either.
When the answers justify building instead of buying
If a vendor fails two or more of the categories above, on a use case that touches regulated or sensitive data, the honest move is to stop negotiating and start scoping a narrower, self-hosted alternative. This isn't the default answer for every AI purchase. Most tools don't warrant it. But for a specific workflow, claims processing, patient intake, financial document handling, where the vendor can't answer basic questions about training data or subprocessor chains, building removes the entire conversation instead of managing around it.
This is the logic behind why Genta runs self-hosted open-source models on a client's own infrastructure for regulated or data-sensitive work, with zero data retention by design rather than by contract clause. It's not a broader claim that self-hosting beats every SaaS AI tool. It's a narrow one: for the subset of workflows where a vendor's answers to the seven-category checklist are unacceptable, owning the model and the infrastructure is often cheaper over three years than the compliance overhead of monitoring a vendor you don't fully trust. We've seen this pattern most clearly in medical-legal operations, where document intake and case assignment touch protected health information that made a self-hosted approach the lower-risk path from the start, not an afterthought bolted on after a vendor review failed.
How this plays out differently in finance and healthcare
In US financial services, examiners increasingly expect AI vendor risk to be documented as part of existing third-party risk management programs, which is why FS-ISAC built a workbook specifically for financial institutions rather than leaving them to adapt a generic template. If you're a fintech or a bank with under $50M in revenue and no dedicated GRC headcount, that workbook is a reasonable starting structure to adapt rather than build from scratch.
In healthcare, HIPAA's business associate agreement requirements mean an AI vendor touching protected health information needs a signed BAA before any data flows, and the BAA itself should reference the same training-data and retention questions from the checklist above, not just generic breach notification language. A BAA that doesn't mention model training is a BAA that hasn't caught up to what the vendor's product actually does.
For Singapore-based readers, the same logic maps onto MAS's technology risk management guidelines and PDPA's data protection obligations. MAS-regulated institutions are expected to extend outsourcing risk assessments to cover AI-specific data handling, and PDPA's cross-border transfer rules apply just as much to a US-based AI vendor's subprocessor chain as they do to any other cloud service. The categories on the checklist don't change; the regulator you're answering to does.
Vendor risk management as a software category is mature and, per Gartner's category overview, search interest in the broader term is actually declining as the market consolidates around a handful of platforms. AI vendor risk sits inside that category but hasn't been absorbed into it yet. Most TPRM platforms still ask generic security questions and treat "uses AI" as a checkbox rather than a distinct risk surface. That gap is exactly why a manual, AI-specific questionnaire still earns its place even if you already run a TPRM tool.
If you're working through a vendor decision like this and keep hitting answers you can't verify, that's usually the point where a Discovery-style diagnosis, mapping out what you actually need built versus bought, saves more time than another round of vendor calls. That's exactly the kind of conversation our enterprise AI work starts with, and we're happy to compare notes.
Frequently asked questions
What is an AI vendor risk assessment and how is it different from a normal vendor security review?
An AI vendor risk assessment evaluates whether a third-party AI product's data handling, model training practices, and subprocessor chain are safe enough to contract with. It differs from a standard vendor review because AI systems can retrain on your data and route it through changing third-party models, unlike static software with a fixed, diagrammable data flow.
What questions should I ask an AI vendor before signing a contract?
Ask whether they train models on your data, how long they retain inputs and outputs, which subprocessors and model providers sit behind their product, whether they log data access, who owns generated outputs, what their contractual data deletion timeline is, and whether they've had a data or model incident in the last 24 months.
Does my AI vendor train its models on our data?
It depends on the vendor and the contract terms, and you should never accept a vague answer here. Get it in writing, in the contract itself, not a privacy policy page that can change unilaterally. Many vendors train on customer data by default unless you explicitly opt out, so ask which default applies to your account.
What should a SOC 2 report from an AI vendor actually cover?
It should explicitly name AI or ML components inside the audited system boundary, and its confidentiality controls should address model inputs and outputs specifically, not just surrounding infrastructure. Many AI vendors hand over SOC 2 reports scoped before they added AI features, which means the report says nothing about the part of the product you're most concerned about.
What are red flags in an AI vendor's data retention or subprocessor policy?
Watch for no contractual deletion SLA, a subprocessor list that "may change without notice," no audit rights to verify claims after signing, and vague non-answers to whether your data trains their model. Any one of these should slow the deal down; two or more on a regulated workflow is a reason to consider building instead of buying.
Tell us where the manual work hurts
We’ll tell you straight whether AI can fix it, what it costs, and what it should return. Whatever we build, you own.
Tell us where the manual work hurts
We’ll tell you straight whether AI can fix it, what it costs, and what it should return. Whatever we build, you own.
Tell us where the manual work hurts
We’ll tell you straight whether AI can fix it, what it costs, and what it should return. Whatever we build, you own.
By
August 1, 2026
10 min read
What to Ask an AI Vendor Before You Sign the Contract



What AI vendor risk actually means (and why it's not the same review you already run)
An AI vendor risk assessment is the process of evaluating whether a third-party AI product handles your data, your model outputs, and your compliance obligations safely enough to sign a contract. It differs from a standard vendor security review in one core way: the thing you're buying keeps changing after you buy it. A normal SaaS tool has a fixed feature set and a static data flow you can diagram once. An AI vendor's model might retrain on customer inputs, route requests through a subprocessor chain that shifts quarterly, or update its behavior with no changelog you'd ever see.
That distinction is why the checklist your procurement team has used for the last decade doesn't cover it. You can ask a normal vendor "where is our data stored" and get a durable answer. Ask an AI vendor the same question and the honest answer is often "it depends on which model we're routing to this week." OneTrust's framing of this gap is useful: AI governance has to be folded into vendor assessment as its own category, not bolted onto the existing security questionnaire as three extra rows.
If you've been through governing AI agents you build in-house, this is the mirror image of that problem. You control the architecture of what you build. You control almost nothing about what you buy, except the terms you negotiate before you sign.
Why this became a board-level conversation in 2025 and 2026
Three things converged at once: certification standards matured, state law caught up, and AI features got embedded into tools nobody formally re-vetted. Search interest in "ai vendor risk assessment" is up roughly 600% quarter-over-quarter and year-over-year as of a July 2026 pull from DataForSEO Labs, and the related term "ai vendor risk" is up about 400% YoY over the same window. Volume is still small in absolute terms, but the trajectory tells you this went from a niche compliance concern to a live procurement question inside two years.
The certification landscape moved first. NIST published AI RMF 1.0 in January 2023, and it's now the default US reference for categorizing AI risk in a vendor context. ISO/IEC 42001, the first international AI management system standard, followed and is increasingly what enterprise buyers ask AI vendors to certify against, per isms.online's breakdown of the standard's vendor-risk provisions. Financial institutions got a concrete workbook to use: FS-ISAC's Generative AI Vendor Evaluation & Qualitative Risk Assessment guide, published for banks and credit unions in early 2024, structures exactly the kind of questionnaire most mid-market companies are now improvising from scratch.
Regulation caught up next. Colorado's AI Act sets compliance obligations for high-risk AI systems ahead of a January 2027 deadline, which we've covered in detail in our state AI law rundown, and it explicitly extends liability to deployers, not just developers, meaning the company that buys the AI tool carries exposure for how it's used. That single fact turns vendor vetting from a nice-to-have into something your legal counsel will ask about directly.
And the third factor is the quiet one: AI got embedded into tools you already pay for. Your CRM added an AI lead-scoring feature in a routine update. Your HR platform turned on a resume-screening model by default. Nobody re-ran a vendor security review because, technically, you didn't sign a new contract. That's the exposure most companies in the $5M to $50M range are sitting on right now without knowing it.
The AI vendor questionnaire: what to actually ask
A useful questionnaire has seven categories, and most vendors can answer the first three fluently and stumble on the last four. That stumble is the signal you're looking for.
Training data use: Does the vendor train, fine-tune, or improve its models using your inputs or outputs, by default or opt-out? Ask for the specific contract clause, not a marketing page.
Retention windows: How long is your data held after a request completes, and does that window differ for logs, embeddings, and cached responses?
Subprocessor disclosure: Which third-party model providers, hosting platforms, or inference APIs sit behind their product, and how often does that list change?
Access logging: Can they produce an audit trail of which employees or systems accessed your data, and for what purpose?
Output and IP ownership: Who owns content generated using your data as input? This matters more than most contracts admit.
Deletion SLAs: If you terminate the contract, what's the contractual timeline for full data deletion across all subprocessors, not just their primary systems?
Incident history: Have they had a data exposure or model behavior incident in the last 24 months, and what changed as a result?
PwC frames this differently than most compliance teams do, and it's worth borrowing their angle: proactive AI third-party risk management reduces time-to-contract, not just risk exposure. A vendor who has already answered these seven categories in writing closes faster than one your legal team has to interrogate line by line. Treat the checklist as a deal accelerator you hand to serious vendors early, not a gate you spring on them at the end.
What a SOC 2 report from an AI vendor should actually cover
A generic SOC 2 Type II report tells you almost nothing about AI-specific risk unless you know what to check for inside it. SOC 2 was built for traditional software controls, access management, change management, availability, and it doesn't natively address whether a model trains on customer data, whether outputs are logged and reviewable, or whether a subprocessor chain includes a foundation model provider with its own data policies.
When you get a SOC 2 report from an AI vendor, look at the scope section first. Does it explicitly name the AI or ML components as part of the audited system boundary, or does the report cover the surrounding infrastructure while treating the model itself as a black box outside scope? Many AI vendors hand over a SOC 2 report that was scoped before they added AI features, and the report simply doesn't speak to the part you're most worried about. Ask directly whether the trust services criteria included confidentiality controls specific to model inputs and outputs, not just the hosting environment.
Red flags that should stop the deal
Some answers are disqualifying on their own, regardless of price or feature fit.
The vendor can't tell you, in writing, whether your data trains their model. "We take privacy seriously" is not an answer to a yes-or-no question.
There's no audit right in the contract. If you can't verify their claims after signing, the questionnaire answers are marketing copy.
The subprocessor list is undisclosed or described as "may change without notice" with no notification clause.
There's no contractual data deletion guarantee with a defined timeline, only a policy statement on a webpage that can change unilaterally.
They can't produce evidence of a prior incident response, or they claim they've never had an incident. Every vendor at scale has had something go sideways; the ones who say otherwise either haven't been asked hard enough by anyone else, or aren't being straight with you.
These overlap heavily with what an insurance underwriter wants to see when pricing your own AI liability exposure, which we walked through in our AI liability insurance guide. If a vendor's answers wouldn't satisfy your own underwriter, they shouldn't satisfy your procurement team either.
When the answers justify building instead of buying
If a vendor fails two or more of the categories above, on a use case that touches regulated or sensitive data, the honest move is to stop negotiating and start scoping a narrower, self-hosted alternative. This isn't the default answer for every AI purchase. Most tools don't warrant it. But for a specific workflow, claims processing, patient intake, financial document handling, where the vendor can't answer basic questions about training data or subprocessor chains, building removes the entire conversation instead of managing around it.
This is the logic behind why Genta runs self-hosted open-source models on a client's own infrastructure for regulated or data-sensitive work, with zero data retention by design rather than by contract clause. It's not a broader claim that self-hosting beats every SaaS AI tool. It's a narrow one: for the subset of workflows where a vendor's answers to the seven-category checklist are unacceptable, owning the model and the infrastructure is often cheaper over three years than the compliance overhead of monitoring a vendor you don't fully trust. We've seen this pattern most clearly in medical-legal operations, where document intake and case assignment touch protected health information that made a self-hosted approach the lower-risk path from the start, not an afterthought bolted on after a vendor review failed.
How this plays out differently in finance and healthcare
In US financial services, examiners increasingly expect AI vendor risk to be documented as part of existing third-party risk management programs, which is why FS-ISAC built a workbook specifically for financial institutions rather than leaving them to adapt a generic template. If you're a fintech or a bank with under $50M in revenue and no dedicated GRC headcount, that workbook is a reasonable starting structure to adapt rather than build from scratch.
In healthcare, HIPAA's business associate agreement requirements mean an AI vendor touching protected health information needs a signed BAA before any data flows, and the BAA itself should reference the same training-data and retention questions from the checklist above, not just generic breach notification language. A BAA that doesn't mention model training is a BAA that hasn't caught up to what the vendor's product actually does.
For Singapore-based readers, the same logic maps onto MAS's technology risk management guidelines and PDPA's data protection obligations. MAS-regulated institutions are expected to extend outsourcing risk assessments to cover AI-specific data handling, and PDPA's cross-border transfer rules apply just as much to a US-based AI vendor's subprocessor chain as they do to any other cloud service. The categories on the checklist don't change; the regulator you're answering to does.
Vendor risk management as a software category is mature and, per Gartner's category overview, search interest in the broader term is actually declining as the market consolidates around a handful of platforms. AI vendor risk sits inside that category but hasn't been absorbed into it yet. Most TPRM platforms still ask generic security questions and treat "uses AI" as a checkbox rather than a distinct risk surface. That gap is exactly why a manual, AI-specific questionnaire still earns its place even if you already run a TPRM tool.
If you're working through a vendor decision like this and keep hitting answers you can't verify, that's usually the point where a Discovery-style diagnosis, mapping out what you actually need built versus bought, saves more time than another round of vendor calls. That's exactly the kind of conversation our enterprise AI work starts with, and we're happy to compare notes.
Frequently asked questions
What is an AI vendor risk assessment and how is it different from a normal vendor security review?
An AI vendor risk assessment evaluates whether a third-party AI product's data handling, model training practices, and subprocessor chain are safe enough to contract with. It differs from a standard vendor review because AI systems can retrain on your data and route it through changing third-party models, unlike static software with a fixed, diagrammable data flow.
What questions should I ask an AI vendor before signing a contract?
Ask whether they train models on your data, how long they retain inputs and outputs, which subprocessors and model providers sit behind their product, whether they log data access, who owns generated outputs, what their contractual data deletion timeline is, and whether they've had a data or model incident in the last 24 months.
Does my AI vendor train its models on our data?
It depends on the vendor and the contract terms, and you should never accept a vague answer here. Get it in writing, in the contract itself, not a privacy policy page that can change unilaterally. Many vendors train on customer data by default unless you explicitly opt out, so ask which default applies to your account.
What should a SOC 2 report from an AI vendor actually cover?
It should explicitly name AI or ML components inside the audited system boundary, and its confidentiality controls should address model inputs and outputs specifically, not just surrounding infrastructure. Many AI vendors hand over SOC 2 reports scoped before they added AI features, which means the report says nothing about the part of the product you're most concerned about.
What are red flags in an AI vendor's data retention or subprocessor policy?
Watch for no contractual deletion SLA, a subprocessor list that "may change without notice," no audit rights to verify claims after signing, and vague non-answers to whether your data trains their model. Any one of these should slow the deal down; two or more on a regulated workflow is a reason to consider building instead of buying.
Tell us where the manual work hurts
We’ll tell you straight whether AI can fix it, what it costs, and what it should return. Whatever we build, you own.
Tell us where the manual work hurts
We’ll tell you straight whether AI can fix it, what it costs, and what it should return. Whatever we build, you own.
Tell us where the manual work hurts
We’ll tell you straight whether AI can fix it, what it costs, and what it should return. Whatever we build, you own.
By
August 1, 2026
10 min read
What to Ask an AI Vendor Before You Sign the Contract



What AI vendor risk actually means (and why it's not the same review you already run)
An AI vendor risk assessment is the process of evaluating whether a third-party AI product handles your data, your model outputs, and your compliance obligations safely enough to sign a contract. It differs from a standard vendor security review in one core way: the thing you're buying keeps changing after you buy it. A normal SaaS tool has a fixed feature set and a static data flow you can diagram once. An AI vendor's model might retrain on customer inputs, route requests through a subprocessor chain that shifts quarterly, or update its behavior with no changelog you'd ever see.
That distinction is why the checklist your procurement team has used for the last decade doesn't cover it. You can ask a normal vendor "where is our data stored" and get a durable answer. Ask an AI vendor the same question and the honest answer is often "it depends on which model we're routing to this week." OneTrust's framing of this gap is useful: AI governance has to be folded into vendor assessment as its own category, not bolted onto the existing security questionnaire as three extra rows.
If you've been through governing AI agents you build in-house, this is the mirror image of that problem. You control the architecture of what you build. You control almost nothing about what you buy, except the terms you negotiate before you sign.
Why this became a board-level conversation in 2025 and 2026
Three things converged at once: certification standards matured, state law caught up, and AI features got embedded into tools nobody formally re-vetted. Search interest in "ai vendor risk assessment" is up roughly 600% quarter-over-quarter and year-over-year as of a July 2026 pull from DataForSEO Labs, and the related term "ai vendor risk" is up about 400% YoY over the same window. Volume is still small in absolute terms, but the trajectory tells you this went from a niche compliance concern to a live procurement question inside two years.
The certification landscape moved first. NIST published AI RMF 1.0 in January 2023, and it's now the default US reference for categorizing AI risk in a vendor context. ISO/IEC 42001, the first international AI management system standard, followed and is increasingly what enterprise buyers ask AI vendors to certify against, per isms.online's breakdown of the standard's vendor-risk provisions. Financial institutions got a concrete workbook to use: FS-ISAC's Generative AI Vendor Evaluation & Qualitative Risk Assessment guide, published for banks and credit unions in early 2024, structures exactly the kind of questionnaire most mid-market companies are now improvising from scratch.
Regulation caught up next. Colorado's AI Act sets compliance obligations for high-risk AI systems ahead of a January 2027 deadline, which we've covered in detail in our state AI law rundown, and it explicitly extends liability to deployers, not just developers, meaning the company that buys the AI tool carries exposure for how it's used. That single fact turns vendor vetting from a nice-to-have into something your legal counsel will ask about directly.
And the third factor is the quiet one: AI got embedded into tools you already pay for. Your CRM added an AI lead-scoring feature in a routine update. Your HR platform turned on a resume-screening model by default. Nobody re-ran a vendor security review because, technically, you didn't sign a new contract. That's the exposure most companies in the $5M to $50M range are sitting on right now without knowing it.
The AI vendor questionnaire: what to actually ask
A useful questionnaire has seven categories, and most vendors can answer the first three fluently and stumble on the last four. That stumble is the signal you're looking for.
Training data use: Does the vendor train, fine-tune, or improve its models using your inputs or outputs, by default or opt-out? Ask for the specific contract clause, not a marketing page.
Retention windows: How long is your data held after a request completes, and does that window differ for logs, embeddings, and cached responses?
Subprocessor disclosure: Which third-party model providers, hosting platforms, or inference APIs sit behind their product, and how often does that list change?
Access logging: Can they produce an audit trail of which employees or systems accessed your data, and for what purpose?
Output and IP ownership: Who owns content generated using your data as input? This matters more than most contracts admit.
Deletion SLAs: If you terminate the contract, what's the contractual timeline for full data deletion across all subprocessors, not just their primary systems?
Incident history: Have they had a data exposure or model behavior incident in the last 24 months, and what changed as a result?
PwC frames this differently than most compliance teams do, and it's worth borrowing their angle: proactive AI third-party risk management reduces time-to-contract, not just risk exposure. A vendor who has already answered these seven categories in writing closes faster than one your legal team has to interrogate line by line. Treat the checklist as a deal accelerator you hand to serious vendors early, not a gate you spring on them at the end.
What a SOC 2 report from an AI vendor should actually cover
A generic SOC 2 Type II report tells you almost nothing about AI-specific risk unless you know what to check for inside it. SOC 2 was built for traditional software controls, access management, change management, availability, and it doesn't natively address whether a model trains on customer data, whether outputs are logged and reviewable, or whether a subprocessor chain includes a foundation model provider with its own data policies.
When you get a SOC 2 report from an AI vendor, look at the scope section first. Does it explicitly name the AI or ML components as part of the audited system boundary, or does the report cover the surrounding infrastructure while treating the model itself as a black box outside scope? Many AI vendors hand over a SOC 2 report that was scoped before they added AI features, and the report simply doesn't speak to the part you're most worried about. Ask directly whether the trust services criteria included confidentiality controls specific to model inputs and outputs, not just the hosting environment.
Red flags that should stop the deal
Some answers are disqualifying on their own, regardless of price or feature fit.
The vendor can't tell you, in writing, whether your data trains their model. "We take privacy seriously" is not an answer to a yes-or-no question.
There's no audit right in the contract. If you can't verify their claims after signing, the questionnaire answers are marketing copy.
The subprocessor list is undisclosed or described as "may change without notice" with no notification clause.
There's no contractual data deletion guarantee with a defined timeline, only a policy statement on a webpage that can change unilaterally.
They can't produce evidence of a prior incident response, or they claim they've never had an incident. Every vendor at scale has had something go sideways; the ones who say otherwise either haven't been asked hard enough by anyone else, or aren't being straight with you.
These overlap heavily with what an insurance underwriter wants to see when pricing your own AI liability exposure, which we walked through in our AI liability insurance guide. If a vendor's answers wouldn't satisfy your own underwriter, they shouldn't satisfy your procurement team either.
When the answers justify building instead of buying
If a vendor fails two or more of the categories above, on a use case that touches regulated or sensitive data, the honest move is to stop negotiating and start scoping a narrower, self-hosted alternative. This isn't the default answer for every AI purchase. Most tools don't warrant it. But for a specific workflow, claims processing, patient intake, financial document handling, where the vendor can't answer basic questions about training data or subprocessor chains, building removes the entire conversation instead of managing around it.
This is the logic behind why Genta runs self-hosted open-source models on a client's own infrastructure for regulated or data-sensitive work, with zero data retention by design rather than by contract clause. It's not a broader claim that self-hosting beats every SaaS AI tool. It's a narrow one: for the subset of workflows where a vendor's answers to the seven-category checklist are unacceptable, owning the model and the infrastructure is often cheaper over three years than the compliance overhead of monitoring a vendor you don't fully trust. We've seen this pattern most clearly in medical-legal operations, where document intake and case assignment touch protected health information that made a self-hosted approach the lower-risk path from the start, not an afterthought bolted on after a vendor review failed.
How this plays out differently in finance and healthcare
In US financial services, examiners increasingly expect AI vendor risk to be documented as part of existing third-party risk management programs, which is why FS-ISAC built a workbook specifically for financial institutions rather than leaving them to adapt a generic template. If you're a fintech or a bank with under $50M in revenue and no dedicated GRC headcount, that workbook is a reasonable starting structure to adapt rather than build from scratch.
In healthcare, HIPAA's business associate agreement requirements mean an AI vendor touching protected health information needs a signed BAA before any data flows, and the BAA itself should reference the same training-data and retention questions from the checklist above, not just generic breach notification language. A BAA that doesn't mention model training is a BAA that hasn't caught up to what the vendor's product actually does.
For Singapore-based readers, the same logic maps onto MAS's technology risk management guidelines and PDPA's data protection obligations. MAS-regulated institutions are expected to extend outsourcing risk assessments to cover AI-specific data handling, and PDPA's cross-border transfer rules apply just as much to a US-based AI vendor's subprocessor chain as they do to any other cloud service. The categories on the checklist don't change; the regulator you're answering to does.
Vendor risk management as a software category is mature and, per Gartner's category overview, search interest in the broader term is actually declining as the market consolidates around a handful of platforms. AI vendor risk sits inside that category but hasn't been absorbed into it yet. Most TPRM platforms still ask generic security questions and treat "uses AI" as a checkbox rather than a distinct risk surface. That gap is exactly why a manual, AI-specific questionnaire still earns its place even if you already run a TPRM tool.
If you're working through a vendor decision like this and keep hitting answers you can't verify, that's usually the point where a Discovery-style diagnosis, mapping out what you actually need built versus bought, saves more time than another round of vendor calls. That's exactly the kind of conversation our enterprise AI work starts with, and we're happy to compare notes.
Frequently asked questions
What is an AI vendor risk assessment and how is it different from a normal vendor security review?
An AI vendor risk assessment evaluates whether a third-party AI product's data handling, model training practices, and subprocessor chain are safe enough to contract with. It differs from a standard vendor review because AI systems can retrain on your data and route it through changing third-party models, unlike static software with a fixed, diagrammable data flow.
What questions should I ask an AI vendor before signing a contract?
Ask whether they train models on your data, how long they retain inputs and outputs, which subprocessors and model providers sit behind their product, whether they log data access, who owns generated outputs, what their contractual data deletion timeline is, and whether they've had a data or model incident in the last 24 months.
Does my AI vendor train its models on our data?
It depends on the vendor and the contract terms, and you should never accept a vague answer here. Get it in writing, in the contract itself, not a privacy policy page that can change unilaterally. Many vendors train on customer data by default unless you explicitly opt out, so ask which default applies to your account.
What should a SOC 2 report from an AI vendor actually cover?
It should explicitly name AI or ML components inside the audited system boundary, and its confidentiality controls should address model inputs and outputs specifically, not just surrounding infrastructure. Many AI vendors hand over SOC 2 reports scoped before they added AI features, which means the report says nothing about the part of the product you're most concerned about.
What are red flags in an AI vendor's data retention or subprocessor policy?
Watch for no contractual deletion SLA, a subprocessor list that "may change without notice," no audit rights to verify claims after signing, and vague non-answers to whether your data trains their model. Any one of these should slow the deal down; two or more on a regulated workflow is a reason to consider building instead of buying.
Tell us where the manual work hurts
We’ll tell you straight whether AI can fix it, what it costs, and what it should return. Whatever we build, you own.
Tell us where the manual work hurts
We’ll tell you straight whether AI can fix it, what it costs, and what it should return. Whatever we build, you own.
Tell us where the manual work hurts
We’ll tell you straight whether AI can fix it, what it costs, and what it should return. Whatever we build, you own.