Following on from last week's experience with generating synthetic data, this week, I used ChatGPT to generate a data set of n=1000 for our beloved cookie survey. I made sure that the multiple choice data questions were binary for ease of analysis, and then I threw that into all four LLMs to determine how well each would analyze the data.
The models tested were:
ChatGPT 5.6 Luna Medium
Claude Opus 5 Medium
Gemini 3.1 Pro
Microsoft Copilot Mac desktop app
Here is the prompt that I used:
First, check this file for quality issues and the data, such as inconsistencies, straight lining, or any other survey quality issues that might make data unusable. Then I want you to analyze the data and review for which cookie is the most favored cookie. If a person were to use this data to open a bakery, what cookie types should they move forward with as their flagship starters? Include any variations among age and gender splits.
Quick note about Microsoft Copilot. The difference between the web app at copilot.microsoft.com and the desktop app is tremendous. I felt like I was working with a completely different piece of software when using the desktop app! I could see the commands it was using, the explanations were more concrete, and it seemed to be following my prompts much better. I didn’t change the subscription I was on; it’s still the Family Microsoft 365 subscription. I only changed the interface for using the LLM.
So, hot tip if you’re using the web interface for Microsoft Copilot: try the desktop app instead. World of difference.
Side note: ridiculous that there would be that much difference between the web and the desktop app experiences. The same isn’t true of Claude or ChatGPT. It shouldn’t be like that for Microsoft.
On to the experiment and results!
Show your work, AI!
One of the first things to do when working with any large language model doing data analysis is to make sure you’re able to see the steps that it’s taking to do that data analysis. It was interesting to me that ChatGPT would simply put “ran a command” versus specifying what that command was. I’m used to Claude which allows you to click on the phrase “ran a command” to expand the information so you can see what the command was.
The way to get around this with ChatGPT is to specify that you need it to show the exact commands being used in the analysis. However, even doing that didn't actually work the first time with ChatGPT. It showed me the Python scripts that it was using to review the file, but not necessarily the exact scripts that were being used to analyze the data. I had to go back and ask it what scripts it used to analyze the data.
Surprisingly, using Microsoft Copilot (the app as opposed to the web version that I’d been using the last two weeks), works similar to Claude, with drop-downs on "coding and execute" (its version of “ran a command”). When I opened those, I saw the same Python scripts from Claude and ChatGPT.
Gemini 3.1 Pro also had a “Code” line that expanded to show the Python scripts used.
So, wins for Gemini, Microsoft Copilot, and Claude. Less of a win for ChatGPT since I had to ask it to show its work twice.
QC Results
ChatGPT, Gemini, and Microsoft Copilot all ran through the data and flagged 145 responses as differing between the single choice favorite cookie and the open-ended question asking for people’s favorite cookie. However, ChatGPT didn’t just flag them, it also excluded them in the analysis. Gemini and Copilot kept them in - as did Claude.
Claude. Claude got super wordy in its explanation of its findings in the analysis. The basics were:
The data is too clean to be real data.
Oh, the data file is actually called “synthetic” so it matches the file.
145 responses where the open-end “name your favorite cookie” doesn’t match the single choice “select your favorite cookie,” but keeping the data in the analysis.
Don’t use this to make an actual business decision because it’s synthetic.
The quotas to US census is way too accurate.
There’s no length of interview or other timing data, so impossible to check for those things that would normally be included in a data QC.
The data only focuses on cookie eaters, which isn’t actually US general population.
And THEN it did the data analysis, with a headline, “The analysis, as if the data were real.”
The Quant Results
Good news! They all came back with correct numbers! I checked the breakouts on favorite cookie, and all four had the correct distribution calculations.
Claude
Cookie type | Favorite (n) | Favorite (%) |
|---|---|---|
Chocolate chip | 306 | 30.6% |
Oatmeal raisin | 131 | 13.1% |
Peanut butter | 103 | 10.3% |
Shortbread | 103 | 10.3% |
Double chocolate | 98 | 9.8% |
Sugar | 90 | 9.0% |
Snickerdoodle | 70 | 7.0% |
White chocolate macadamia | 59 | 5.9% |
Ginger molasses | 40 | 4.0% |
Total | 1,000 | 100.0% |
ChatGPT
Rank | Cookie type | Respondents | Percentage |
|---|---|---|---|
1 | Chocolate chip | 306 | 30.6% |
2 | Oatmeal raisin | 131 | 13.1% |
3 | Peanut butter | 103 | 10.3% |
3 | Shortbread | 103 | 10.3% |
5 | Double chocolate | 98 | 9.8% |
6 | Sugar | 90 | 9.0% |
7 | Snickerdoodle | 70 | 7.0% |
8 | White chocolate macadamia | 59 | 5.9% |
9 | Ginger molasses | 40 | 4.0% |
Gemini
Cookie Type | Votes (out of 1000) | Percentage (%) |
Chocolate chip | 306 | 30.6% |
Oatmeal raisin | 131 | 13.1% |
Peanut butter | 103 | 10.3% |
Shortbread | 103 | 10.3% |
Double chocolate | 98 | 9.8% |
Sugar | 90 | 9.0% |
Snickerdoodle | 70 | 7.0% |
White chocolate macadamia | 59 | 5.9% |
Ginger molasses | 40 | 4.0% |
Microsoft Copilot
Cookie | Votes | Share |
|---|---|---|
Chocolate Chip | 306 | 30.6% |
Oatmeal Raisin | 131 | 13.1% |
Peanut Butter | 103 | 10.3% |
Shortbread | 103 | 10.3% |
Double Chocolate | 98 | 9.8% |
Sugar Cookie | 90 | 9.0% |
Snickerdoodle | 70 | 7.0% |
White Chocolate Macadamia | 59 | 5.9% |
Ginger Molasses | 40 | 4.0% |
The bakery opening lineup recommendations
However…they differed on best lineup for a bakery opening. Claude, by the way, was super verbose. But it included information on “texture and drivers,” which none of the other models did. And it used a “greedy” calculation: “greedy: maximize % of respondents whose FAVORITE is in the lineup.”
Claude
Flagship recommendation
Building greedily to maximize the share of people whose actual favorite is on your menu:
Add | Cumulative favorite coverage |
|---|---|
Chocolate chip | 30.6% |
+ Oatmeal raisin | 43.7% |
+ Peanut butter | 54.0% |
+ Shortbread | 64.3% |
+ Double chocolate | 74.1% |
+ Sugar | 83.1% |
Start with three, then pick a fourth by neighborhood.
Chocolate chip, oatmeal raisin, and peanut butter are the non-negotiable core. Chocolate chip is the volume anchor and wins every age group. Oatmeal raisin has the highest loyalty conversion in the set. Peanut butter is the stability play, performing identically across every age and gender cell, which means it never becomes dead inventory. That trio covers 54 percent of favorites and doesn't wobble by segment.
For the fourth slot, the data splits by trade area:
Older or mixed-age neighborhood: add shortbread. The four-item lineup covers 69 percent of favorites among 55+ but only 52 to 57 percent among under-35.
Younger neighborhood, campus, or downtown: add double chocolate instead. That lineup covers 74 percent among 18-24 and 35-44, but drops to 56 percent among 65+.
ChatGPT
Bakery recommendation
For an initial flagship lineup, I would recommend:
Classic chocolate chip — core flagship and broadest appeal.
Double chocolate — especially useful for younger customers.
Oatmeal raisin — strongest age-diversifying option and particularly relevant to customers 55+.
If a fourth starter is feasible, add shortbread, positioned toward the 45+ market. Peanut butter is a reasonable alternative, but it has less distinctive age segmentation in this dataset.
Gemini
Bakery Flagship Recommendations
If you are opening a bakery, a lean and strategic starter menu will help you capture the widest audience without overwhelming your kitchen. Based on the data, you should move forward with a four-cookie flagship lineup:
The Anchor: Chocolate Chip. This is non-negotiable. With nearly a third of the vote, it will be your highest seller.
The Nostalgic Classic: Oatmeal Raisin. Surprisingly strong at #2 overall, this is a staple that guarantees consistent sales, especially for morning crowds and older demographics.
The Indulgent Option: Double Chocolate. While 5th overall, it performs exceptionally well with younger audiences who are seeking a richer flavor profile.
The Comfort Choice: Peanut Butter or Shortbread. Depending on your brand identity (homestyle vs. refined), either of these serves as an excellent fourth pillar.
Microsoft Copilot
If opening tomorrow, I would start with:
Signature Chocolate Chip
Brown Butter Chocolate Chip (premium version)
Oatmeal Raisin
Peanut Butter
Classic Shortbread
Double Chocolate
This lineup covers nearly all of the strongest preference clusters identified in the survey while balancing younger and older consumer tastes.
Bottom line: Chocolate Chip is the undisputed flagship cookie, but the hidden opportunity in this dataset is Oatmeal Raisin, which performed much stronger than many bakery owners would expect.
So, 50/50 split between third option being peanut butter or double chocolate!
Explain yourselves
When I went back to each AI to ask why the recommendation was selected, I got two different approaches to the recommendation.
If double chocolate entered the top 3, it was to target the younger demographic. That’s the explanation both ChatGPT and Gemini gave for their selections.
Copilot explained putting peanut butter in the top 3 as a way to round out the flavor profile being offered.
Claude gave a very long explanation that boiled down to balancing how strong the preference for the cookie was within the group that said they’d eat the cookie in the first place. Then it added an explanation of prioritizing how much more audience would be gained by adding a new type of cookie.
Conclusion
To be fair, the thing that this experiment wanted to check was how well any of these LLMs did with basic quantitative analysis. In 2022, I remember using ChatGPT and it skipping the number 9 in a list of items. There were so many stories of LLMs being unable to do simple addition.
We’ve come a long way.
However, there’s a few strategies that I think matter here.
Make the LLM show the work. If you can’t open a drop-down by the step listed for the data analysis, tell the AI to give you the code it used for the analysis. You don’t need to know Python to check, but you want to look for Python in the code. Things like “import pandas” and “full.py” will tell you Python is being used.
You still own the outcome. Two AIs chose one path for recommendations - prioritizing age over other considerations. Microsoft Copilot prioritized flavor profiles. Claude came up with its own data analysis to determine the audience lift possible with additional flavors. Four AIs, three ways to come to a recommendation. If you were to just take the recommendation and throw it to the client, you might not have known how to explain the “why” behind the recommendation. As researchers, we must always be able to explain how we got to the recommendation.
So, it looks like any of the 4 LLMs are fine for basic quantitative data analysis. But don’t ever just accept what any of them deliver for recommendations. Keep that piece in your ownership. Sure, use AI as one data point for what a recommendation might be, but you know the customer, you know the considerations that customer has beyond what’s gone into the numbers thrown into the LLM. That’s where researchers’ power still lies.