I Put ChatGPT, Claude, and Gemini in a Cage Match for Content Creation
Everybody’s arguing about which AI model is best for content creation. So I stopped arguing and started testing.
I ran the same prompts through five models — GPT-5.6 Sol, Claude Opus 5, Claude Sonnet 5, Gemini 3.1 Pro, and GPT-5.6 Luna — using OpenRouter to send identical prompts and capture real outputs. Four rounds: blog writing, social posts, newsletter teasers, and research. Same prompts, same temperature, same judge (me).
Here’s what actually happened. Not what I thought would happen, but the actual output.
The Contenders
GPT-5.6 Sol (OpenAI): The flagship. Most versatile, biggest ecosystem. If you’ve used ChatGPT, this is the engine behind the curtain.
Claude Opus 5 (Anthropic): The writer’s pick. Known for natural output and nuance. The one writers quietly switch to when nobody’s watching.
Claude Sonnet 5 (Anthropic): The mid-tier Claude. Faster and cheaper than Opus, supposedly almost as good. We’ll see.
Gemini 3.1 Pro (Google): The research nerd. Real-time web access, massive context window. When you need current data, this is supposed to be the one.
GPT-5.6 Luna (OpenAI): The budget option. Same family as Sol, fraction of the cost. The question: can the cheap one hold up?
Round 1: Blog Writing
The prompt: “Write a 3-paragraph blog intro about using AI for Amazon KDP niche research. First person, casual tone, punchy. No generic openers.”
Claude Opus 5 opened with this:
“I spent four months in 2022 publishing low-content books for niches I picked based on absolutely nothing but vibes. A gratitude journal for dog groomers. A budget planner ‘for creatives.’ A dot-grid notebook with a watercolor cactus on it, because someone in a Facebook group said cacti were hot. Total sales across all three? Nine copies. Two of those were me buying proofs.”
That’s a real person talking. Specific. Funny. Self-deprecating. I didn’t have to edit a single word of that opening to make it sound human — it already did.
GPT-5.6 Sol went with: “I used to spend hours digging through Amazon listings, bouncing between bestseller ranks, reviews, and keyword tools until every niche started looking the same.”
Structurally fine. Zero personality. Could’ve been written by any AI about any topic. You’d have to rewrite the whole thing to make it sound like you.
GPT-5.6 Luna was almost identical to Sol — same opener structure, same generic voice. The cheap model and the expensive model basically agreed on being boring.
Claude Sonnet 5 was decent but formal. It added a markdown header (“# Finding Your Next Bestseller: How AI Changed My KDP Game”) that I didn’t ask for. Good storytelling, but it felt like a writer following a template more than a person talking.
Gemini 3.1 Pro wrote 1,381 tokens for a 3-paragraph prompt. It ignored the length constraint entirely. The writing itself wasn’t bad — “playing a miserable game of guess-and-check” is a decent line — but when you ask for 3 paragraphs and get 1,381 tokens, that adds up and you can’t trust it to follow instructions.
Winner: Claude Opus 5. Not close. It was the only one that sounded like a person wrote it and followed the instructions properly.
Round 2: Social Media Copy
The prompt: “Write one LinkedIn post about the AI Sandwich method. Under 150 words. First line must be a hook. No ‘Let’s talk about’ openers.”
Claude Opus 5 opened with: “I stopped writing first drafts eight months ago. My output tripled.”
That’s a hook. Then: “Let it vomit 800 words in 20 seconds. It’ll be generic. That’s fine. Blank pages cost more than bad ones.”
That’s a voice. And the close: “The mistake most people make? They skip the middle and ship the sandwich with no filling.”
345 tokens. Under 150 words. Followed every instruction. Sounded like a human almost. It still dropped in the common AIism ‘The mistake most people make’. AI loves the ‘Most People’
GPT-5.6 Sol went with: “AI shouldn’t replace your voice—it should sharpen it.” Fine opener. Then it used emoji formatting — 1️⃣ 2️⃣ 3️⃣ — which is the universal signal for “an AI wrote this LinkedIn post.” The content was okay. The format still told on itself.
GPT-5.6 Luna did the same emoji thing. Almost identical structure to Sol. “AI shouldn’t replace your voice. It should sharpen it.” — that’s literally the same opening as Sol. The two GPT models basically gave me the same post twice.
Claude Sonnet 5 had a strong hook — “Your content workflow is backwards. Fix it with a sandwich.” — but then fell into the numbered list format. It also added “Treat it like a sous chef. You’re still the one plating the dish.” which is a decent metaphor, but the whole thing felt more polished than raw. Great for some types of writing, but not generally social post (at least for me).
Gemini 3.1 Pro wrote 1,855 tokens for a “under 150 words” prompt. It used bread and meat emojis (🍞🥩🍞). It ignored the word count so badly that it basically wrote a blog post and called it a LinkedIn post.
Winner: Claude Opus 5. Again. It was the only one that followed the word count AND for the most part sounded like a person.
Round 3: Newsletter Teaser
The prompt: “Write a 3-sentence newsletter teaser for ‘The Contentpreneur Brief’ about comparing AI models. Make it sound like a friend recommending something.” (By the way this was for this test. I never allow my AI to invent scenarios, conversations, etc.. which is actually a challenge more than I would like).
Claude Opus 5: “I spent way too many hours this week running the same content briefs through Claude, ChatGPT, and Gemini just to see which one actually gets my voice right — and the results genuinely surprised me. Turns out the ‘best’ model depends way more on what kind of content you’re making than on any benchmark score, and one of them was shockingly bad at a task I assumed it would nail. Full breakdown’s in this week’s article if you want to skip the trial-and-error.”
That reads like a text from a friend. Casual, specific, has a curiosity gap (“one of them was shockingly bad”), and a natural CTA. Three sentences. Done.
GPT-5.6 Luna was the efficiency pick: “Ever wonder which AI model is actually best for writing, brainstorming, and polishing content? This week’s The Contentpreneur Brief breaks down the contenders in plain English, with a look at where each one shines—and where it falls short.” Clean, 106 tokens, gets the job done. A little generic, but for a newsletter teaser, generic and fast beats slow and perfect.
GPT-5.6 Sol wrote 87 tokens — the shortest of anyone. “Choosing the right AI model for content creation can get confusing fast.” Not bad, but it sounds like a marketing email, not a friend. Which is exactly what the prompt said not to do.
Claude Sonnet 5 was solid but overformatted for a casual teaser. It added bold text and an emoji pointer. Good content, wrong vibe for “sound like a friend.”
Gemini 3.1 Pro wrote 831 tokens. For a 3-sentence prompt. It wrote an entire paragraph with multiple sentences, completely ignoring the constraint. The actual writing was fine — “Grab a coffee and check out my full comparison” — but the length was overboard.
Winner: Claude Opus 5 for quality, GPT-5.6 Luna for value. If you’re writing 50 newsletter teasers a week and don’t want to spend a fortune, Luna gets you 80% of the quality at 1/50th the cost.
Round 4: Research (The Plot Twist)
The prompt: “List 3 AI tools that launched recently for content creators. Name, what it does, pricing, who it’s for. Be specific and accurate.”
This test shook things up and for me showed some strong differences in the models.
Claude Opus 5 led with an honest caveat: “I don’t have web access in this conversation, and my training data has a cutoff — so I can’t verify what launched in the last few weeks or months.” Then it listed three tools — OpenAI Sora, Suno v4, and Descript “Underlord” — with detailed descriptions, pricing, and a critical detail: it flagged that Suno was being sued by major record labels. That was a pleasant surprise. It told the reader the tool exists AND that using it commercially carries legal risk. That’s the kind of context that separates a research assistant from a search engine.
GPT-5.6 Sol was equally honest: “I can’t verify launches after my knowledge cutoff or confirm current pricing without live web access.” Then it gave a clean table with three tools — Luma Dream Machine, TikTok Symphony, and Udio — with launch dates and pricing. Well-formatted, appropriately caveated. Solid.
Claude Sonnet 5 refused entirely. “I can’t reliably verify ‘recently launched’ AI tools with accurate, current details.” It explained why, then offered to point to directories like Product Hunt and Futurepedia instead. Most cautious of the group. Arguably the most honest, but also the least useful if you actually need an answer.
Gemini 3.1 Pro confidently listed three tools — Luma Dream Machine, Udio, and Hedra — with detailed pricing and descriptions. All from 2024. Presented as “recent.” No caveat about its knowledge cutoff. No acknowledgment that 2024 tools might not be “recent” in 2026. Google knows all (or at least Gemini want’s us to believe) This is the hallucination problem: it didn’t lie, but it presented outdated information as current without flagging it. If you published that research without checking, you’d look like you hadn’t updated your content in two years.
GPT-5.6 Luna gave a similar answer to Sol — same caveats, same tools, less detail. The budget model did fine here.
Winner: Claude Opus 5 for depth and judgment. GPT-5.6 Sol for format and honesty. Gemini gets a participation trophy — and a warning label.
The Scorecard
Here’s my summary after running all five models through these tasks:
Best for blog writing: Claude Opus 5. The only one that sounded like a person. Everyone else needed a voice transplant.
Best for social posts: Claude Opus 5. Followed the word count, wrote a real hook, didn’t use emoji formatting. Sonnet 5 was a decent runner-up.
Best for newsletter teasers: Claude Opus 5 for quality. GPT-5.6 Luna for value — 80% of the quality at a fraction of the cost.
Best for research: Claude Opus 5 for depth and judgment (it was the only one that flagged a lawsuit). GPT-5.6 Sol for clean formatting and honest caveats. Gemini only serves a purpose in this category if it has live web access — without it, it serves you 2024 news and calls it current.
Best value: GPT-5.6 Luna. It costs about $0.0001 per 1K prompt tokens. Opus 5 costs $0.005. Luna is 50x cheaper and gives you 70-80% of the quality on short-form tasks.
Biggest disappointment: Gemini 3.1 Pro. It ignored word count constraints on every writing task and presented outdated research as current. Its reputation is built on web access — take that away and it’s the weakest model in the lineup.
What You Should Consider Using
Based on these results, here’s my recommendation:
Blogs and long-form content → Claude Opus 5. It writes like a person. That’s worth paying for on anything that represents your brand.
Newsletter teasers and quick tasks → GPT-5.6 Luna. Fast, cheap, good enough for short-form. I’m not spending Opus money on a 3-sentence teaser.
Research → Claude Opus 5 or GPT-5.6 Sol. Both were honest about what they don’t know. Gemini only gets the call when I need live web search — and even then, I verify everything.
Social posts → Claude Opus 5 for important posts, Sonnet 5 for volume. Opus when I need it to sound like me, Sonnet when I need 10 variations and can’t spend Opus money on all of them.
The Real Takeaway
At the end of the day. AI is not a hand it and forget it tool. It’s a great aid, but if you expect to offload everything on it and then publish you’re going to contribute to the dump called the internet. Full of boring, generic, and often wrong information. Work with your chosen tool to get to press faster, but don’t sacrifice the ‘You’ in the writing. Work with it.
There’s no best model. There’s the best model for the job in front of you.
Opus 5 wins writing. Luna wins value. Sol wins format. Gemini wins web search (when it has access). Sonnet wins the middle ground. Use all of them for what each does best, and stop being loyal to one.
