Is forcing GPT-5.6 Sol on every prompt actually better than letting it route to 5.5 Instant?
Follow on from the routing thread. Now that I know how to keep new chats on Sol, I am not sure I want to.
Roughly half my usage is "reformat this", "what is the flag for X", "summarise this paragraph". Instant answers those in about two seconds. Sol at medium effort takes noticeably longer to produce an answer I cannot tell apart. The other half is the stuff I actually care about, and there I do not want to be quietly routed down.
Before I pin it globally:
- Which classes of prompt does the higher effort tier actually change the answer on, in a way you could show someone? Interested in cases where you ran the same prompt at two effort levels, not in vibes.
- Does higher effort cost you anything on a paid plan beyond latency? I have not found a usage number I trust for the chat product and I would rather not learn it by hitting it in the middle of something.
- Has anyone found the per-project effort setting to be a better answer than a global default?
@stdlib_stef · 2h ago · 2 replies
Where the higher tier earns its latency, from a fortnight of doing exactly this comparison on my own work:
Where it is just slower: retrieval shaped questions, formatting, short self contained snippets, rewriting. If the answer is essentially lookup, effort buys you nothing.
The cheap way to settle it for your workload: keep ten prompts from real work, run each at both settings, save the outputs. Twenty minutes, and it is your workload rather than someone's benchmark table.
Reply
Report
@gramgrader_gus · 2h ago
And score them blind, with the labels stripped. If you know which one is the expensive tier you will find the improvement you paid for. I have caught myself doing it, which is why I now shuffle the file names before reading.
Reply
Report