Most tool comparisons measure the wrong things — feature counts, benchmark scores, and how impressive the demo looks. Here is a test that predicts whether you will still be using something in three months.
Test it on your work, not theirs
The single most important step. Take five tasks you actually did last week — including the awkward ones — and run them through.
Vendor examples are selected best cases. Your work has ambiguous inputs, missing fields, domain jargon and context that lives outside the document. That is the material that decides whether a tool is useful.
Include at least one task you know is hard. What a tool does at its limit is more informative than what it does in the middle.
Judge the failure, not just the success
Every tool succeeds on easy tasks. They differ on what happens at the edge, and this is the question most evaluations skip:
- Does it say "I don't have that", or produce something plausible and wrong?
- Does it show its sources so you can check, or assert?
- When it is uncertain, can you tell?
- Are errors legible, or does it just spin?
A tool that admits its limits can be used on unfamiliar work. A tool that fabricates confidently must be verified line by line, which usually cancels the time it saved.
The questions to ask before adopting
Data. Where is it processed? Is it retained? Is it used for training? Can you turn that off? If you handle anything confidential, get this in writing before you upload.
Export. Can you get your prompts, history and configuration out? Anything you cannot export is something you will rebuild if you leave.
Failure mode of the business. If the tool disappeared next month, what breaks? Depth of integration is a cost, not only a benefit.
Cost at real volume. Per-seat pricing and per-token pricing behave very differently as usage grows. Estimate at the volume you would reach if it worked well.
The decision
Adopt when a tool removes a specific, repeated friction you can name. "It seems powerful" is not that. "It saves me forty minutes every Monday on the report" is.
If you cannot name the friction after a week of real use, the honest conclusion is that you do not need it yet — and that is a legitimate result, not a failed evaluation.