A few weeks ago, I kicked off this series of research on LLMs by asking all four of the primary LLMs to draft a series of surveys. We started with a simple survey about cookie preference, then moved to survey testing ads, and ended with an advanced methodology - conjoint.
This week, I had every intention of creating a study designed specifically to measure drivers behind purchase intent, then creating synthetic data using the LLMs and asking all four to then analyze the data and execute a drivers analysis, factor analysis, and simple regression analysis. This post was going to compare how well all four did with the advanced data analysis.
What I found along the way, though, was an interesting dynamic between researcher and LLM in just creating a survey that was going to be used for a driver's analysis in the first place.
The Prompt
“I want to do a 15 minute survey about the drivers behind intent to purchase for back-to-school items among the drivers to test I want to include excitement about back-to-school anticipation about back-to-school the enjoyment of office supplies, prices of office supplies and school supplies, and brand recognition.
Suggest other drivers to be tested. I want to be able to generate data that can then be analyzed using factor analysis and regression analysis to truly determine what drives intent to purchase for back to school items.”
So, as always, I wanted to show you the prompt that was used for all four LLMs. Interestingly, for this particular exercise, I started with Claude running Opus 5. This is important simply because, after a lengthy explanation of what drivers it would add because mine were too limited, it added a list of items that, as it said, would determine “whether the data supports factor analysis and regression cleanly.”
Among these, of course, were the sample size and scale consistency, but it also included dependent variable and multicollinearity. I had a dependent variable (DV) already: intent to purchase back-to-school items. However, according to Claude, "‘Intent to purchase back-to-school items’ is close to universal among parents and will have almost no variance.”
Fair point.
And then it said it didn’t know the audience I wanted to target - parents of K-12 students, college students, or adults restocking their home offices. It ended the treatise with:
If you tell me the audience and the DV you want to model, I can draft the actual item battery with wording, scale, and a mapping of items to intended factors.
None of the other three LLMs went into the amount of detail that Claude did, nor did any of the LLMs point out that the dependent variable that I had originally listed was going to be fairly universal among parents and have almost no variance the way that Claude did.
Instead, all three of the other LLMs I simply responded to my initial prompt and listed additional drivers to consider and recommended analysis approaches for factor analysis, driver's analysis, and regression analysis to be successful.
Once I shifted the dependent variable to intent to purchase premium or brand label versus private label and specified the audience as parents of K-12 students, each of the LLMs shifted their approach. Each listed new drivers with new statements, and in some cases rearranged the order of the drivers to be listed, as well as a slightly different variation of the analysis recommended for the study.
The Comparison
Gemini Pro had the lightest weight study of all four LLMs. It had the fewest drivers and really did not seem as potentially well thought out as some of the others. You can find the Gemini conversation and output here.
Co-pilot came in third, but only after telling it to use a different dependent variable than what I had told it to originally, as well as to change the audience. You can find the Copilot conversation and output here.
For ChatGPT, I was using Luna Medium, which is the lightest of the 5.6 models. It did fairly well, though it did not include questions about the dependent variable like Claude did at the beginning, but it also created a 60-item study, which, for somebody taking this questionnaire, felt really long. You can find the ChatGPT conversation and output here.
Claude Opus 5 did the best in terms of having the most thorough study, but it was also really verbose in every explanation, and its study had 58 items by the end with two quality check (QC) questions. It was the only model that actually included QC questions, and it was the only model that had a really thorough explanation of the regression analysis. You can find the Claude conversation and output here.
But similar to ChatGPT, with 60 items in the study, I was worried about the actual length of interview (LOI). It estimated a 14-minute LOI, which I wasn't entirely certain would be the case for someone going through this and rating every single one of these 60 items on a scale of 1 to 7.
If you want to see a full side by side comparison of all four final studies, I asked Claude to do an analysis of them. You can find that side-by-side analysis here. I didn’t continue doing a full analysis of them only because I was struck by what I’m calling the real learning from this experiment.
The Real Learning
Now, let's talk about the real takeaway from this:
You would need to be a fairly well-versed quantitative researcher to know whether or not any of these four draft studies were any good.
The fact that Claude raised the question about the dependent variable and actually suggested that the initial dependent variable was no good and was going to be flat, surprised me in a good way. The fact that none of the other three LLMs didn't raise the same issue shows you that the prompt that goes into the LLM will define the quality of the output you get.
In fact, knowing that LLMs don't always give the same answer to the same prompt two times in a row, I decided to go back to Opus 5 with the same exact prompt that I had started with to see if it would come up with the same objection that it did the first time about the original dependent driver running the risk of being flat.
It did not.
I tried Claude Fable 5.1, and the same thing happened. No cautionary statement about the dependent variable potentially being flat. I tried ChatGPT 5.6 Terra and Sol, and neither of them brought up the objection that Opus 5 did on the first try with the initial prompt, either.
So, I got lucky with Opus 5 acting like a good second researcher on my first try with this prompt.
Expertise wins over just having a good prompt
Even with how good all of the LLMs are getting, expertise still wins no matter how good the prompt is. In this case, I was coming up with a prompt quickly to be able to do this experiment for this particular post. If this were an actual customer, I would need to know more about the drivers that mattered to the business and likely would have had a better dependent variable to start with.
When turning to AI for the draft of a study like this, you’re hoping that the LLM knows enough about:
office supply retail
back-to-school shopping
advanced analytics like drivers analyses, regression analyses, factor analyses, etc.
your customer’s business to know what drivers they regularly measure or that matter to the business.
It might know enough about some of those things, but, unless you’ve been building a good library of data for the LLM to draw from about the customer, previous studies, drivers the customer regularly measures, etc., it likely won’t know that last bullet point, which is the most important one for it to generate a strong draft.
Deciding What to Give to AI
That brings me to why it's important that we think about what tasks we hand to AI and what tasks we keep in our expertise. It's very easy to open any LLM and type in a prompt for it to create a study of any complexity. You will receive an answer.
Whether it's a good answer or not is where you need to exercise your expertise.
AI does well with operational work and execution of tasks. What you do well with is with the nuance about your customer, knowing exactly how to phrase certain questions, and knowing what to ask your customer to get the inputs you need to identify the correct methodology to answer their questions.
So, tempting as it may be to hand most things over, be judicious about what tasks you hand to AI. AI can take the role of pushing on your thinking to identify gaps, reviewing your drafts to point out weak spots, or checking that logic in the questionnaire is clean.
You keep the original thinking and strategizing.