Hyderabad: India is now the second-largest market on earth for AI chatbots.
ChatGPT alone crossed roughly 100 million weekly active users here, growing about 41 per cent year on year.
At the India AI Impact Summit earlier in 2026, OpenAI disclosed a detail that surprised even people in the industry: users aged 18 to 24 send close to half of all ChatGPT messages originating in India, and under-30s account for around 80 per cent of total usage.
A very large number of those messages are people asking a machine to write something. An email to a landlord. A leave application. A caption. A cover letter. A complaint to a municipal office.
So, the question is fair: of the three big assistants, which one writes best?
The honest answer is that the question has a cleaner answer than most articles admit, and a messier one than any of them admit. Both parts are below.
What you are actually choosing?
The three companies shipped so fast this summer that most comparison articles online, including ones published this month, are describing models that have already been replaced. Here is the real lineup.
Anthropic
Anthropic (Claude) released Claude Opus 5 on July 24, 2026. It carries a 1 million token context window, up to 128,000 tokens of output, and a knowledge cutoff of May 2026.
API pricing is 5 dollars per million input tokens and 25 dollars per million output. It is the default model on the Max plan and the strongest option on Pro. Above it sits Claude Fable 5, released June 9, which is the genuine frontier model at 10 dollars and 50 dollars per million.
Fable 5 has an unusual history: US Commerce Department export controls forced Anthropic to suspend access on June 12, and it only returned on July 1 after the controls lifted. Claude Sonnet 5 (June 30) is the cheaper workhorse, at 2 dollars and 10 dollars per million on introductory pricing until August 31, then 3 and 15.
ChatGPT
OpenAI (ChatGPT) made the GPT-5.6 family generally available on July 9.
Unusually, it is three separate models rather than three settings: Sol is the flagship, Terra the balanced middle, Luna the cheap high-volume tier. All three share a 1.05 million token context window and 128,000 token output. Sol costs 5 dollars and 30 dollars per million. On July 30, OpenAI cut Terra by 20 per cent and Luna by a striking 80 per cent, leaving Luna at 20 cents and 1.20 dollars. One catch buyers miss: requests above 272,000 input tokens are billed at double input and 1.5 times output, applied to the whole request.
Gemini
Google (Gemini) is the messiest picture. Gemini 3.7 Flash reached general availability this week, with Google’s own developer documentation updated on August 13. It runs a 1 million token context window, 64,000 tokens of output, tunable thinking levels, and introductory pricing of 75 cents and 3.75 dollars per million through December 31, after which it doubles. Note that almost every ‘August 2026 comparison’ article currently ranking on Google still describes 3.6 Flash as Google’s newest model. It is not, as of yesterday.
The flagship you actually pay a premium for remains Gemini 3.1 Pro at 2 dollars and 12 dollars per million, with the largest context window of the three at 2 million tokens. The long-awaited Gemini 3.5 Pro is still not out. Bloomberg reported on July 16 that it is months behind schedule and has fallen short of Google’s internal targets, particularly on coding.
What Indian readers actually pay
ChatGPT Plus is 1,999 rupees a month with GST included.
Google AI Pro is around 1,950 rupees.
Claude Pro sits near 2,000 to 2,400 rupees depending on billing.
ChatGPT and Gemini both support UPI and rupee billing. Free tiers on all three are usable for light writing work, and for many readers the free tier is the honest recommendation.
What the leaderboards say
On the public benchmarks that specifically measure writing, Claude models currently lead. EQ-Bench’s Creative Writing leaderboard, LiveBench Language, and the Arena text board all place Anthropic models at or near the top as of this month.
There is also one of the only genuine blind tests anyone has run.
Two newsletter writers put the three models through eight prompts with the labels stripped and the order randomised, and asked readers to vote. Around 134 people voted in the first round. Claude won four of the eight rounds. ChatGPT won exactly one, the business strategy prompt, at 53 per cent. Gemini never won a round outright and never came last either, which is its own kind of result.
That is the clean answer. Now the messy one.
Why you should not trust that clean answer
The most-cited writing benchmark uses a Claude model as its judge.
EQ-Bench’s Creative Writing test works by generating responses to 32 prompts, then having a judge model score them against a rubric. The benchmark’s own public source code specifies that for results comparable to the official leaderboard, the judge must be Claude Sonnet 4.6. The long-form writing leaderboard uses the same judge.
This matters because of a well-documented problem in AI research called self-preference bias.
In the foundational work behind these leaderboards, GPT-4 rated its own outputs about 10 per cent more favourably than human raters did. Claude-v1 favoured its own by 25 per cent.
A 2026 paper from researchers at Unbabel and Instituto Superior Técnico extended the finding: judges favour models from their own family, not merely their exact selves. Another June 2026 study across 20 mainstream models found that greater capability does not reduce this bias and is sometimes negatively correlated with it.
So the benchmark most often used to prove Claude writes best has a Claude model marking the papers. That does not make the result wrong. It makes it unproven. To EQ-Bench’s credit, they publish the judge model openly and flag the limitation themselves. The failure lies with the articles that quote the score and skip the method.
Private testing
The human preference leaderboard has its own problem.
In April 2025, researchers from Cohere Labs, Stanford, Princeton, AI2, Waterloo and the University of Washington published a 68-page audit of Chatbot Arena, now called Arena, analysing 2 million matchups across 243 models. They found that large labs can privately test many variants of a model and publish only the best-scoring one. In one case, a single provider tested 27 private variants before a public release. They also found that Arena voters reward style: bulleted lists and particular response lengths score better regardless of substance.
Arena disputed several figures in the paper and published a detailed response, which is worth reading alongside it. But the core point about private testing was not seriously contested.
Two leaderboards, same model, wildly different rank
The sharpest illustration: Claude Sonnet 5 currently sits around 13th on one writing benchmark and around 53rd on another. Same model, same month. If the leaderboards agreed, that could not happen.
What should you actually use?
Strip away the rankings and the practical picture is reasonably stable across independent hands-on testing, user discussion, and the blind test results.
Claude produces prose that needs the least editing to sound like a person wrote it. It holds a tone instruction across a long document without drifting, and it drops fewer requirements from complicated briefs. Its weakness is not quality but access: usage caps on the paid tiers are the single most common complaint from heavy users, and it lacks image generation.
ChatGPT is the most versatile and by a wide margin the largest ecosystem: image generation, voice, wider language support, more third-party integrations. Its writing is competent and well-structured but has a recognisable register that experienced editors spot immediately. It won the strategy prompt in the blind test, which fits a broader pattern of doing well on structured, analytical output.
Gemini is the research writer. Live Google Search access, Workspace integration inside Gmail and Docs, the biggest context window, and by far the cheapest tokens. It never topped a writing round in the blind test and never bombed one. For summarising long reference material, or for writing that must be anchored to current facts, the integration advantage is structural rather than stylistic.
The most useful framing is not which one wins. It is that Gemini finds the material, Claude writes the prose, and ChatGPT does everything adequately in one place. Most people who write for a living and can afford it use two.
How to test this yourself
If you want an answer for your own work rather than an average across strangers, the method matters more than the tool.
Use the same prompt, on the same day, in all three.
Turn web access on for all or off for all. Use default settings and a fresh chat with no memory or custom instructions. Run each prompt two or three times, because these models are non-deterministic and one run can mislead. Then strip the labels before you judge, and if you can, get someone else to rate them blind.
The single most useful measure for real work is not elegance. It is edit time: how many minutes from the AI’s draft to something you would actually send.
A note on reading AI comparisons
Search this question and page one is dominated by articles from companies selling AI writing tools, transcription services and AI platforms, all reaching conclusions that favour a purchase. Many recycle numbers without checking them.
One example currently circulating widely claims the United States has 745.87 million ChatGPT users. The population of the United States is about 340 million. The figure appears to be cumulative visits relabelled as users, and it has been syndicated across multiple sites without anyone stopping to compare it to a census.
That is a useful test to apply to anything in this genre, including this article. Check whether the numbers could physically be true. Check whether the model versions named are the current ones. Check who profits from the conclusion.
Disclosure: Research assistance for this article was provided using an AI assistant where that assistant’s own product family is under comparison. Claims have been sourced to independent benchmarks, peer-reviewed papers and primary vendor documentation, and the conflicts have been stated in the text.