I’m taking a break from the side-by-side experiments to address an item that came up during a discussion about tools that could be used for drafting a report using AI.

The question that often comes up about using an LLM to draft a report is: but if I upload a dataset to Claude or ChatGPT, how do I know it’s not being used to train the models?

But, after I answered this question, another question was raised about data retention and data processing that I think is worth addressing, as well. In my research into the differences between data retention and processing for most of the software we use vs LLMs, I learned a LOT. I went into this thinking, “We’ve dealt with software companies retaining our data and processing our data for ages now, why is it suddenly making people worried with LLMs?”

The short answer: LLMs complicate the whole retention and processing thing.

But just going to market research tech specifically built to process and analyze data doesn’t necessarily mean you’re safer than using the LLM directly.

Let’s unpack this.

Data retention and processing before LLMs

Data retention and processing are things that have been happening with software as a service for years. Your cell phone company, using the cloud version of Microsoft, Word, or Google Docs - any of those software as a service tools are and have been retaining your data and processing your data in some way.

However, those tools had a specific thing they were doing with your data. Each of those pieces of software had a defined purpose that you would use that software for. Also if you're using a cloud storage system, once you delete a file from that system, it's gone. You can't recover it.

Data retention and processing for LLMs

Large language models, however, operate differently.

First, a company could retain your data even after you have deleted it on your side. For example, in Google's policies, it says that if a conversation has been flagged for human review, even if you delete it, it remains flagged for human review. could still be reviewed by a human later.

In addition, your conversation doesn't necessarily get deleted by the large language model company. Sometimes it's still retained for a certain length of time even if you've deleted it on your side.

And when it comes to data processing, well, LLMs could be used to process your data in a wide variety of ways. You are the one that decides what that data processing looks like. And, at least for Microsoft Copilot, even if you’ve said “don’t use my data to train your models,” they state that “will not exclude your conversations from being used for other general product or system improvements nor from use for advertising, digital safety, security, and compliance purposes.”

…yeah.

So…what should I do?

This is all really focused for those who are using consumer subscriptions to do market research work. Consumer subscriptions don't carry the same terms and conditions that business subscriptions do. In essence, think of it as the difference between using your personal Google Drive for work versus using a company's Google Drive. Your use of the LLM should be based on the agreement that you have with the customer for whom you are doing the work.

Three questions can help with any tech being used for market research work:

  • what is this vendor allowed to do with the data, and who defines that?

  • how long is the data kept after I’m done, and can I make the data be deleted immediately?

  • which sub-processors or models sit underneath?

If you're working with non-identifiable data for a market research project, then turning the training off is a reasonable floor for addressing data privacy concerns. If, however, you are working with identifiable data, then you might want to take a look at using a purpose built market research technology tool, or using a commercial subscription to one of these LLMs.

But even if you're using a purpose-built market research technology tool, you still need to read the terms and conditions, because some might just be wrappers around a large language model. Others are very clear on who is responsible for the data processing and who is responsible for the data retention.

So, look in their terms and conditions and privacy policy (sometimes the answers are split between the two documents) for answers to:

  • who is responsible for what happens with the data at each stage of the processing

  • who does what with the data

  • is it easy for you to delete the data

  • how long is data stored

  • where is the data stored (since different locales have different policies around data security and data privacy).

Business plan vs consumer plan vs specialized restech

The answer to “should I use market research tech tools” (aka restech) follows a bit of a risk analysis on what you’re doing and what your contract states - at least for freelancers and solopreneurs.

Restech is best when you can see the data processors list and the data processing agreements in place between the tech company and the tools it might be using to work with the data AND:

  • contract requires the documentation of data security in place for the tools used to analyze data

  • you’re working with raw data that has PII, including recordings, transcripts with names and work places, etc.

  • you need to follow regulations (GDPR, HIPAA, etc.)

  • you need traceability and auditability of the work done and where the data went during the project

  • you’re dealing with a high volume of data (100 transcripts vs 10 transcripts; LLMs can’t reliably handle the 100).

A business subscription is fine when the customer has approved it AND:

  • the data is anonymized and you’re certain it can’t be re-identified from information in the data

  • the sensitivity is business confidentiality vs personal data

  • the volume of data being processed is small

  • nobody is requiring something stricter.

A consumer subscription is fine when the customer has approved it AND:

  • the data is anonymized and you’re certain it can’t be re-identified from information in the data

  • there’s nothing the customer would consider confidential in the data.

I know both Anthropic and OpenAI now have policies where the minimum monthly charge for a business subscription is double that of a consumer subscription. but it's worth the investment for an entrepreneur or a freelancer who might not have access to AI for the projects they’re doing and is being expected to use it on projects.

Recommended for you