Welcome to the start of a series of articles comparing how each of the major LLMs (Microsoft Copilot, Gemini, ChatGPT, and Claude) fare when given the same market research task.

In this series, I will be running a prompt through each LLM and comparing the results. I will post links to the results so that you can see for yourself what the outputs looked like. Each link will include the model used and the date the prompt was run, since these models change so quickly!

A note on models: I don’t have a Microsoft 365 business license, so I am using Microsoft Copilot that comes with a Microsoft 365 Family license. For Gemini, I’m using Gemini Pro that comes with Google Workspace Business Standard license - and I learned that the memory function for Gemini exists in a weird location after finishing with this experiment, so the results I received for this were based on it not knowing I work in market research. And ChatGPT was defaulted to ChatGPT Instant, which generally is one of the latest models. I have a paid subscription to ChatGPT and to Claude.

Experiment 1: Quantitative Survey Drafts

Because my background is heavily quantitative, I decided to start with a set of increasingly complicated quantitative surveys.

  1. A cookie survey just asking for the general public’s favorite cookie.

  2. An ad test.

  3. A conjoint study.

The prompts aren’t polished. I was speaking them into the computer, and I am terrible at speaking into the void. The prompts you see in this article are pasted exactly as they were entered into each LLM. I didn't change anything about the prompt between LLMs.

The prompt I used for the very first survey was for a very simple survey. This is a survey designed to measure people's favorite cookie. Here’s the prompt:

“Create a short questionnaire about cookies. The objective is to understand the general public's favorite type of cookie, choosing from the classic types of cookies.”

Nothing fancy.

Results (links open pdf files in a separate window):

Claude, Gemini, and Microsoft Co-Pilot all created very similar questionnaires that started with what's your favorite type of cookie, what do you like about the cookie, what kind of texture do you like and what do you like having with the cookie: tea, coffee, milk, etc.

They did have slightly different variations on the classic types of cookie. One had ginger/molasses in one category while others separated them, and one had macarons while another one didn't. So, slight differences there but generally speaking, the same overall.

I will note here that I did have issues with Claude. I was using Sonnet 5, and my first issue was that it took longer than I expected. I’ve noticed that for a lot of prompts with the most recent models. However, the bigger issue was that it output two different surveys at the same time.

The first was a downloadable file, and that survey looked good. It was pretty straightforward and what I would have expected. The other survey was in the chat itself, and it had an answer option for favorite type of cookie that listed “Snickerdoggle…Snickerdoodle,” as well as some other wonky items in the survey.

When I asked it, “Why is this survey so different from the survey that I can download in the right panel on the desktop app?” it came back with, basically, “What downloadable survey? I only created the survey you see in the chat.”

This is the first time that I have ever run into this with Claude. I’ve seen other reports of Claude doing weird things with the most current update.

So, be aware that Claude's latest versions of their models might fritz out on you, while also giving you decent answers, I guess? It was bizarre.

Model used: Sonnet 5; effort used: Medium

ChatGPT created a one question screener that asked how often do you eat cookies, terminating if the answer was, “Never.” ChatGPT was the only one that introduced the screener question.

Winner: ChatGPT

The Ad Test

Next, I decided to ask all of four to create a survey to test ads for an ad campaign. Here's the prompt.

“Create a survey that is going to be testing ads for an ad campaign for the upcoming holiday season. The customer is needing to test six different ads so that they can decide which two they are going to move forward with. The ads are 60 second spots as well as 30 second spots.

There's going to be six of each.”

Results:

Microsoft Copilot probably did the worst on this exercise just because its survey was the least fleshed out of all of them. It created a sequential monadic survey where respondents would watch three ads: one 60-second, one 30-second, and then one random.

It gave a decent set of metrics including emotional impact, message clarity, and length, followed by favorite, least favorite, and appropriate for the holidays. But it was fairly underwhelming.

For Gemini, I actually tested Flash, 3.6 Thinking, and 3.6 Pro because I’ve noticed that the three versions of Gemini will produce different outputs.

Flash separated the audience into the 60-second ad group and the 30-second ad group, but had both audiences watch all 6 ads and then answer appeal, brand recall, emotional impact, and key message.

Gemini 3.6 Thinking went for a monadic test and provided an analysis guide for selecting the top two ads with a weighted metric scheme:

  • Top-2-Box Purchase Intent (40% weight)

  • Correct Brand Linkage Score (30% weight)

  • Net Appeal Rating (20% weight)

  • Message Clarity Average (10% weight)

Gemini 3.6 Pro first had a survey with an intro, which was different from Flash and Thinking, but it also said that to combat survey fatigue, it would show everyone the 60-second ads and have them score those, and then just use the 30-second spots to check that message clarity was continuing through.

To me, that wasn't necessarily combating survey fatigue.

When I tested Pro a second time, it created a sequential monadic test. However, it kept the idea of testing the 60 second ad with the metrics and using the 30 second version to test message clarity. Also, interestingly, the second time that I used Gemini Pro, it left out the intro.

ChatGPT went straight for a monadic test, but it called for a balanced incomplete monadic test, where each person would see a subset of ads and the sample would be distributed evenly across all six concepts and lengths. GPT’s ad metrics included persuasion, relevance, differentiation, brand linkage, and overall appeal. It also gave a recommended scorecard and decision criteria based on the results of the study.

Create a scorecard for each concept containing:

  1. Overall appeal

  2. Consideration/persuasion lift

  3. Purchase motivation

  4. Brand linkage

  5. Message communication

  6. Relevance

  7. Holiday fit

  8. Differentiation

  9. Attention and memorability

  10. Comparative preference

ChatGPT

So, a much more robust ad measurement study than either Gemini or Copilot.

Claude put together a similar study to what ChatGPT did. I tested Sonnet 4.6 and Sonnet 5, and both of them created similar studies. Sonnet 4.6 also had decision frameworks and recommended decision criteria and included the different items that would be used in that decision framework rather than just going for a favorite ad.

Appendix A: A full rotation and cell assignment guide with the 12 ad ID scheme (AD1-60, AD1-30, etc.) and a table for your programmer.

Appendix B: Key metrics table mapping every question to a reporting format and a decision threshold, plus a recommendation framework for how to use the data to select the final two concepts.

Things I'd want to clarify with you or the client before programming: whether there's a specific product category that needs a screener question, the exact age and demographic targets, and whether forced video playback is feasible in the platform they're using (it matters a lot for validity with a 60-second spot).

Claude Sonnet 4.6

Claude’s latest Sonnet model, Sonnet 5, didn’t include a decision framework, but suggested confirming one with the customer.

Claude, ChatGPT, and Gemini all included open-ended questions you’d expect in an ad test.

Winners: ChatGPT and Claude Sonnet 4.6

The Conjoint

Okay, now for some real fun. Here is the prompt that I used for our last test of survey drafting among the four.

“I'm working with a customer who is trying to figure out the right combination of price and features for kitchen blenders that they want to sell the price points are a $30 $50 or a $100 blender. The buttons that could be included are pulse, ice crush, blend, liquefy, generic on/off, and speeds 1-5. Create a conjoint study that could be used to help this customer identify the price point and feature list that would work best for the audience of potential blender buyers.”

Results:

So, none of the four actually created a full draft study. I was honestly a bit disappointed by that. I was hoping for a fully programmable conjoint from one of these tools!

That aside, Microsoft Copilot and Gemini were hands down the worst of the four. Where they did well: they did come up with a choice-based conjoint and gave example tasks that you’d show people in the study. They also suggested no more than 10-12 tasks per person, and Copilot suggested 200-300 responses, Gemini 3.6 Pro suggested 300-500. Gemini 3.6 Thinking didn’t have a recommended audience size.

ChatGPT got closer. It went through a screener, feature definitions, how to set up the tasks, the recommendation of generating the tasks algorithmically and a recommendation to be strategic about the combinations used. It recommended some design restrictions, like what could a blender legitimately contain, some post-conjoint questions, and then some analysis guidance on estimating part-worth utility, creating attribute importance, willingness to pay, creating a market simulator, and optimization scenarios.

Pretty thorough, just not a fully fleshed out study, which is understandable.

Claude (Sonnet 5) was really interesting. The processing notes cracked me up, followed by, “This is squarely in your wheelhouse, so I put together a full CBC (choice-based conjoint) design rather than just a sketch.” It said of the 6 buttons, “That’s too many for a clean CBC design, and several combinations aren't realistic.” So, it recommended combining some into groups and testing as groups rather than individually. This reduced the possible number of combinations from 192 to 108.

Then it said, in essence, “But if your customer really wants to test every single possible combination, then you’d be better off with a menu-based conjoint.”

Instead of creating the robust study outline and analysis guidance that ChatGPT did, it gave some example tasks, and recommended 12 tasks per respondent, with 3 combinations and one “none,” just like the other 3 tools proposed. And then it gave a list of items that it didn't know such as the customer's target market, or if this was US versus multi-market (which would impact pricing). it even listed this, “Conjoint tells you what buyers will pay for a feature, not what it costs to include, and both numbers are needed to pick a final answer.”

So, interesting in the depth of response, but it left me feeling like there were more questions than answers, and a bit disappointed overall. I appreciated the questions posed to take into consideration, such as the fact that production costs should also play into decisions for selling goods, but it also felt a tad frustrating as a researcher. I suppose I was looking for “the research bundle” to react to that ChatGPT had provided.

Winner: ChatGPT

Runner up: Claude Sonnet 5 because it raised good questions

Conclusion

I should start by saying I don't have access to the business license version of Microsoft Copilot, and I don't know if the output from that version would be different from the version that comes with a family Microsoft 365 license. If it is, I would hope that it is more robust than what was given here.

If I were to rank the four large language models just from this experiment, it would look like this:

  1. ChatGPT and Claude

  2. Gemini

  3. Copilot.

Claude’s latest Sonnet model has me worried based on the experience with the super weird survey it generated, but that’s why we always check the output.

And speaking of always checking the output, you’ll see that none of these models created the perfect survey, at least not on the first try, and not even for something as simple as the cookie survey. Even ChatGPT with the screener question still had what I considered a slightly weird list of cookie options that I would edit before handing to someone for review.

Take that kind of work and scale it up to a conjoint, and you realize that a researcher still needs to know what good looks like. And I'm still convinced that the only way for any researcher to learn what good looks like is through the experience of drafting surveys themselves.

I have other experiments I’ll be running, like how each LLM does at generating crosstabs, how each does at checking crosstabs for mistakes, how each LLM does at generating discussion guides for an in-depth interview and a focus group, and how each does at quantitative and qualitative data analysis.

If there’s something else you’d like me to add to the list, reply to this email and let me know!