MercatIQ home

Structured buying

Spreadsheets, tabs and ChatGPT: what each one misses when comparing vendors

Three tools do most of the vendor research happening today. None of them was built for it, and it shows in different places.

Samuel Negru, Co-founder & CTO

Three scored bars labelled open tabs, chat answer and scored comparison, with scored comparison scoring highest.

Three tools, none of them built for this

Almost every vendor comparison happening today is assembled from the same three things: a pile of browser tabs, a spreadsheet, and increasingly a chat window.

None of them was designed for evaluating vendors. That is not a criticism. A browser is meant to show you pages, a spreadsheet is meant to hold numbers, and a general assistant is meant to answer whatever you happen to ask it. They are all good tools. They are just being asked to do a job that sits slightly outside what each was built for, and the gap shows up in a different place for each one.

Worth knowing where, because most teams use all three at once and assume the combination covers everything. It does not, and the holes overlap.

What a vendor comparison actually has to do

Before comparing the tools, it helps to be specific about the job. Five questions cover most of it.

Does it help you work out what should matter, rather than assuming you already know? Does it evaluate every option the same way? Can you change your mind about priorities without starting over? Can several people use the same thing? And is any of it still usable in six months?


Browser tabs

Spreadsheet

Chat assistant

Structured evaluation

Helps you work out what should matter

No

No

Only if you ask well

Yes, by default

Evaluates every option the same way

No

Only if you fill it in consistently

Not reliably

Yes

Changing priorities without starting over

Not applicable

Manual rework

Re-ask and hope

Weights update the result

Several people can use the same thing

No

Yes, with version problems

No

Yes

Still usable in six months

No

If someone maintained it

Chat history

A scored record with sources

Browser tabs find things and cannot hold them

Tabs are genuinely good at discovery. That is what a browser is for, and no structured tool replaces the moment where you follow a link and find something you did not know existed.

The problem is that a tab holds nothing. Nothing carries between them, nothing accumulates, and nothing is comparable. By the twelfth tab you are relying entirely on memory to hold what the third one said, and memory does not do this well. Order effects creep in too, because the option you read first sets the frame that every later option gets measured against, without anyone deciding that it should.

So the tabs are a research surface with no output. Which is why they always end up feeding a spreadsheet.

Spreadsheets hold things and cannot find them

A spreadsheet is the opposite. It is a good place to put a comparison and a bad place to produce one.

It does no research, so every cell arrives by hand from one of those tabs, which means the grid is only as current as the last hour someone spent maintaining it. It goes stale immediately and silently. Two options added in week three sit next to options researched in week one, and nothing in the file indicates that.

The subtler issue is weighting. Most comparison spreadsheets either treat every column as equally important, which is never true, or bury the weighting in a formula that one person wrote and nobody else reads. Either way, the most consequential judgement in the whole exercise is the one least visible to everyone looking at it.

None of this makes spreadsheets bad. It makes them a recording surface being asked to do research and reasoning.

A chat assistant answers fast and is not a system

This is the interesting one, because a chat assistant genuinely does some of this well. It summarises a category quickly, explains unfamiliar terms, and will happily draft a criteria list you can react to. For a first orientation into something you know nothing about, it is hard to beat.

What it is not is a tool built for this job, and that shows up in five places.

It answers the question you asked. If you do not know what should matter for this purchase, it will not work that out with you unless you already know enough to ask it to. Someone who has bought this category before can prompt their way to a good criteria set. Someone buying it for the first time, which is most people most of the time, gets a competent answer to a question that was not quite the right one.

Every conversation starts from nothing. Whatever nuance you established last time is gone. You rebuild the context, the constraints and the priorities each time, and some things cannot be rebuilt at all. You can get a table out of a chat window. You cannot get a comparison that five people will read the same way.

A thread is one person's. You can share the text, but you are sharing a transcript of someone else's session, not a thing the group can use. Everyone else is reading what you asked, in the order you asked it, which is not the same as looking at the same evidence.

Changing your mind means asking again. Move a priority and the answer regenerates. Nothing guarantees the new answer preserves the reasoning of the old one, and you have no straightforward way to see what changed and what did not.

It is built to end the conversation. A general assistant is optimised to give you something satisfactory now, which is exactly right for most questions and the wrong trade for a decision you will live with for three years. The Tow Center at Columbia tested eight AI search tools across 1,600 queries and found they consistently produced answers rather than declining when they did not have enough to go on. That study was about news attribution, not vendor research, so I would not stretch it further than the behaviour it documents. But that behaviour is the point: the default is to answer, not to stop.

Six months later, what you have is a conversation log. Not a scored comparison with sources attached to each score.

What is true today, and might not be forever

Two more things are worth separating out, because they are capability gaps rather than design differences.

The first is accuracy on product detail. Buyers we spoke to during research described getting feature lists that were wrong, products that had been confused with each other, and capabilities that did not exist. It is not universal and it is not every query. It is common enough that anyone building a comparison on top of it should be verifying rather than trusting.

The second is that the output is prose. You can ask for a table and get one. You cannot get a comparison built for looking at.

Both of these could improve. The reason not to plan around it: a general assistant has no particular commercial reason to specialise in buying comparison specifically. It is a small slice of what it is used for, and the gaps that matter to you are unlikely to be the gaps that get prioritised.

Where MercatIQ fits

MercatIQ exists because the job needs all five of those things at once and no combination of the three covers it.

A few guided questions help surface the criteria likely to matter for what you are buying, and then you take over: edit them, add your own, set how much each one counts. The options get researched and scored against those criteria, the same way for each one. Change a weight and the ranking updates rather than requiring you to start again. What comes out is a comparison with the reasoning and sources attached to each score, which is a thing a group can look at together and a thing that still means something when someone asks about it next year.

If you are currently running a decision across tabs, a spreadsheet and a chat window, tell us what you are evaluating and we will look at it with you.

Common questions

Can you use ChatGPT to compare vendors?

For orientation, yes. It is quick at explaining a category, listing the kinds of options that exist, and drafting criteria you can react to. What it does not give you is a consistent evaluation of every option against the same criteria, a way to change your priorities without asking again, something a group can use together, or a record that survives the conversation. Treat the output as a starting point to verify rather than a comparison to act on.

Is a spreadsheet good enough for vendor comparison?

A spreadsheet is a good place to hold a comparison and a poor place to build one. It does no research, so it is only as current as the last time someone updated it by hand, and it usually hides the weighting inside a formula or ignores weighting altogether. If you use one, make the weights their own visible row and put a date on the file.

What should a vendor comparison include?

The criteria, the weight on each one, every option scored against all of them, and the evidence behind each score. Options that lost should be in it, including the incumbent. A comparison that only explains the winner is a recommendation, and people treat recommendations as opinions rather than as findings.

Why do AI assistants get product features wrong?

They are answering from a mix of training data and whatever they retrieved at that moment, and they are optimised to give you an answer rather than to stop when the evidence is thin. Product pages also change constantly. The practical implication is not to avoid them, it is to verify anything you plan to make a decision on, particularly pricing, integrations and anything described as available.

How do you keep a vendor comparison useful after the decision?

Keep the criteria, the weights and the sources with the scores, not just the conclusion. Most of the value later comes from being able to answer why something was ruled out, and that is exactly the part people stop recording once they have picked a winner. A conclusion without its reasoning is impossible to revisit and impossible to reuse for the next decision.

Related articles

Aug 26, 2026

·

Team alignment

The analysis is the easy half. Getting six people from four departments to accept the same conclusion is where vendor decisions actually break down.

Three scored bars labelled loudest opinion, another meeting and agreed criteria, with agreed criteria scoring highest.

Aug 25, 2026

·

Why decisions stall

Buying groups spend months comparing options and still regret the purchase. The data says the failure happens earlier than anyone looks.

Three scored bars labelled more sources, more time and clear criteria, with clear criteria scoring highest.

Aug 19, 2026

·

Structured buying

Nine steps that turn scattered research into a scored comparison your team can defend, from defining scope to documenting the decision.

Three approaches compared as bars: gut feel is shortest, spreadsheet is longer, scored comparison runs full length and is highlighted in teal.

Not the decision you are trying to make?

Tell us what you are evaluating. We will show you how MercatIQ builds the criteria, runs the research, and scores the options.