
How to Review Vendor AI Training Rights Before Signing
Review vendor AI training rights before you sign: search improve and enhance, define customer data, demand written opt-in, then ban, narrow or walk.
Key takeaway in 30 seconds
Knowing how to review vendor AI training rights before signing is a ban, narrow, or walk log, not a sales slide. Search improve, enhance, and aggregated. Write customer data to include prompts, files, outputs, and logs. Split service operation from model improvement. Demand written opt-in, not a console toggle. Then require derived-dataset deletion.
Sales said they don’t train. The paper may still grant it. Knowing how to review vendor AI training rights before signing is a 25-minute packet log: search improve / enhance / aggregated, inventory prompts and files, split service from training, demand written opt-in, then ban, narrow, or walk.
In August 2026, Leila — founder, 13-person UK SaaS — has Finance’s yes on a workflow vendor. The AE’s deck says we don’t train on your data. Friday she opens the MSA — a master services agreement: Provider may use submitted data to improve, develop, and enhance the service. A later paragraph lets them train on aggregated Customer Content, and that section survives expiry. Typical mistake: treating the slide as the clause. The hidden risk is Tuesday go-live on that grant.
Disclaimer: Checkory provides AI support, not legal advice. Consult a qualified lawyer for binding decisions.
Why does improve / enhance / aggregated mean training?
The word train may never appear. Improve, enhance, and develop the service often do the work. Freeze the packet — the exact file set you will sign — then search the MSA, order form, DPA — a data processing agreement — AUP, and any AI terms dated today. Log clause, file, and page.
For example, Common Paper CSA 2.1 §1.6 lets the vendor use Usage Data and Customer Content to train models, including third-party components, if aggregated — and §5.6 lets §1.6 survive. The 2026 SaaS Contract Benchmark (16,000 agreements to June 2026): no-training language rose from under 1% (2024) to 14% (2026). Most papers still do not ban. AI Policy Desk flags improve, develop, and enhance the service. Morgan Lewis: that phrase is being stretched into training rights.
- Do: search every file in the packet and log the hit.
- Do not: treat we don’t train, or a Data Protection heading, as the training clause.

What does customer data have to include?
Customer data is not the rows you type. Write the objects: data you import; prompts; uploaded files; model outputs; logs, embeddings, and eval sets; support transcripts. If the definition is only data submitted by Customer through the Service, outputs and logs are probably out. Mark that narrow.
Mark what the definition excludes: vendor analytics they claim; inferred scores; provider IP. You own your data is not a training ban. Tucker Ellis (August 2026) puts prompts, outputs, and logs in the vendor contract, not a privacy-policy paragraph. Other subscribe flags stay on the SaaS agreement red flags checklist — this page is the training-rights slice, not the nine.
- Do: write prompts, files, outputs, and logs into the Customer Data inventory.
- Do not: treat submitted by Customer as the whole set.

How do you split service operation from model improvement?
Service operation hosts, routes, encrypts, tickets, and abuse-scans this tenant. Model improvement trains, fine-tunes, evaluates, or benchmarks any model — the vendor’s or a third-party lab’s — that other customers can benefit from. Keep the first. Ban or written-opt-in the second.
In practice, Venable (5 June 2026): improve or enhance can authorise training, and de-identification is not a silver bullet. HowToContract: a ban that excepts aggregated data can swallow the restriction. If they train on personal data for their own models, name UK. ICO: a processor who processes for its own purposes becomes a controller. Walk Art. 28 on the data processing agreement checklist.
- Do: keep host / route / abuse-scan of this tenant; ban or written-opt-in any all-customer model.
- Do not: treat aggregated or de-identified as automatically service-only.
Aggregated is not the green lane
Anonymised or aggregated training can still carry confidential workflows. Treat it as narrow or walk, not a pass.
Is a console toggle enough?
A console toggle is plan-gated, reversible, and not the signed paper. Demand a cover-page or order-form sentence. A privacy-policy line or Slack we don’t train is not that sentence. The improve-the-service AI training opt-out you need is written and on the packet you will click.
Worked example: Atlassian’s 17 August 2026 change uses metadata and in-app data to improve AI for all customers. In-app data can be turned off; on Free / Standard / Premium, metadata cannot. That is a console, not a signed no-train sentence. Write on the order form: Provider will not use Customer Data — including prompts, files, outputs, logs — to train any model unless Customer gives prior written consent. Common Paper’s Language Library deletes §1.6 and replaces it with No AI Training — a starting point, not a stamp that the deal is done.
- Do: put written opt-in on the cover page or order form.
- Do not: schedule the opt-out for after Tuesday go-live.

What happens to derived datasets on exit?
Raw delete leaves embeddings, eval sets, and a surviving Machine Learning licence. Ask for wipe of derived datasets plus a written certificate. Deleting raw files later does not unwind a licence that survives expiry. If they keep derived copies indefinitely, mark walk.
CSA 2.1 default: delete Customer Content on request within 60 days — raw files, not a derived wipe — and §1.6 survives. Export first if you need a usable zip, then wipe production and derived copies — the SaaS data export and exit rights checklist walks that inventory. Ban withholding the wipe because a fee is disputed.
- Do: require derived-data wipe plus a written certificate, and kill any surviving ML licence.
- Do not: treat a 60-day raw delete as the training grant dying.
Ban / narrow / walk
| Track | Ban when | Narrow to | Walk |
|---|---|---|---|
| Improve / enhance | No-train covers prompts, files, outputs, logs | Written opt-in for a named purpose | No no-train sentence |
| Customer Data | Inventory matches the definition | Add prompts, outputs, logs | Submitted by Customer only |
| Licence split | Service operation only | Customer-specific model only | Aggregated treated as service |
| Consent | Cover-page written opt-in | Notice plus exit if they add AI | Toggle-only or privacy policy |
| Exit | Derived wipe plus certificate | Define aggregated corpora you can audit | §1.6-style survival |

Ban, narrow, or walk — which training rights can you keep?
Ban, narrow, or walk is the gate after you fill the log. Ban needs a cover-page no-train sentence plus derived-data delete. Narrow is written opt-in and a customer-specific model only. Walk when improve/enhance has no no-train sentence, or the ML clause survives.
Pause wording: Provider may use submitted data to improve, develop, and enhance the service. Usage Data and Customer Content may train models, including third-party components. Section 1.6 survives. That is a High flag — wording a human must verify before anyone signs — for counsel — a qualified lawyer. Success bar: a one-page log and one sentence that pauses signature. Workflow: packet → search improve/enhance/aggregated → inventory Customer Data → split service vs training → written opt-in → derived-data delete → ban/narrow/walk. Checkory can run a first-pass — a first machine pass — on the same file at document analysis. A named human still opens every High clause. If the pain is pasting this MSA into a consumer chatbot, stop and use how to keep a confidential contract private when using AI.
- Do: highlight one High sentence, then ban, narrow, or walk.
- Do not: tell Finance the paper is fine because the deck said we don’t train.
Vendor AI training-rights review
Freeze the packet
Lock MSA + order form + DPA + AUP + today’s AI terms. Search improve, enhance, aggregated, train.
Inventory Customer Data
List prompts, files, outputs, logs, embeddings. Mark submitted by Customer as narrow.
Split the two licences
Keep host / route / abuse-scan of this tenant. Ban or written-opt-in all-customer training.
Demand written opt-in
Cover page or order form. A console toggle is not that sentence.
Require derived-data delete
Embeddings, eval sets, plus a written certificate. Kill any surviving ML licence.
Decide ban / narrow / walk
Circle one High sentence. Ban, narrow, or walk.
Optional first-pass, then a human
Upload the same PDF/DOCX. A named human still opens every High clause.
Frequently asked questions
Is anonymised training safe?▼
Where does vendor AI training live — MSA, AUP, or privacy policy?▼
How is this different from pasting a contract into ChatGPT?▼
What if the vendor uses a third-party model lab?▼
Is an admin opt-out enough before Tuesday go-live?▼
Highlight Improve and Enhance
Upload the same packet. A human still reads every High sentence.
Start document analysisWhat to do next
Review the packet
Upload the version you will sign.
RelatedKeep a confidential contract private when using AI
Paste and privacy — a different door from this vendor paper.
RelatedSaaS agreement red flags checklist
Wider subscribe screen. Training is one row there.
RelatedData processing agreement checklist
If they train for their own models they may become a controller.
RelatedSaaS data export and exit rights checklist
Confirm a usable export, then wipe production and derived copies.
RelatedHow to Redline a Vendor Contract Before Signing
Open the sibling checklist after this screen.
Sources
Read also
Related guides





