The honest answer is that all three are good, they change every few months, and the differences that matter to you are probably not the ones in the benchmark charts.
What actually varies
Three things tend to matter more than raw capability. How much material you can paste in before the model loses the thread. How the model behaves when it does not know something. And how it writes by default, before you have given it a voice.
A practical way to decide
Rather than picking a winner, give each tool the jobs it suits. Run the same real task through all three, with the same prompt, and compare the outputs on work you actually care about. Ten minutes of that will tell you more than any comparison article, including this one.
Things worth checking for your own use
- How long a document can you paste before quality drops?
- Does it admit uncertainty, or does it fill gaps confidently?
- How much editing does the default writing style need?
- Does your organisation allow the data you want to paste into it?
The part people skip
Whichever you choose, the prompt matters more than the model. A well-specified prompt on a mid-tier model beats a vague prompt on the best one, almost every time. Spend your effort there before you spend it on tool selection.