Cartoon Nick is back, and this week he is annoyed. Not at you. At the way most people talk to AI. You know the move: ask a question, get an answer, ask a follow-up, get another answer, then 40 messages later you are wondering why the AI forgot what you were even doing.

Here is what is going on under the hood, and two habits that fix it.

AI at Work Tip #51: Edit your prompt instead of sending a follow-up

Every time you send a message to Claude, ChatGPT, or Gemini, the model re-reads the entire conversation from the top. Not just your new message. Everything. By message 30 you are paying for 50,000+ tokens of history just to ask "can you make that shorter?"

Fix: when the answer misses, do not reply. Go back to your original message, hit Edit, add the detail you forgot, and regenerate. The old answer gets replaced instead of stacked. Over a normal 10-message session this one habit cuts your token use by 80 to 90 percent, and if you are on a plan with usage limits, that is the difference between working all afternoon and hitting a wall at 2 PM.

Copy-paste version of a prompt that gets it right the first time:

You are helping me write a [type of content] for [business type] in [city].
Audience: [who reads it].
Tone: [3 adjectives].
Length: [word count or format].
Must include: [list].
Must avoid: [list].
Give me one finished version, not options.

AI at Work Tip #52: Start a fresh chat every 15 to 20 messages

Long conversations are not memory. They are a tax on every future message. When a chat gets long, the model gets slower, more expensive, and weirdly forgetful, because the important stuff is buried under 40 rounds of back-and-forth.

Fix: before you close a long chat, ask for a handoff summary, then paste it into a new chat. Zero context lost, massive savings.

Summarize everything we decided in this conversation as a handoff brief for a new chat. Include: the goal, decisions made, constraints, style rules, and what is still unfinished. Keep it under 300 words.

Then open a new chat and paste the brief with "Continue from this brief."

What's hot right now

Anthropic shipped Claude Fable 5.1 on September 1, and the headline for small business owners is not the model itself. It is that cached input pricing dropped 75 percent. Translation: the "edit, don't follow up" habit above now saves you even more, because repeated context is cheaper when it is cached and pointless when it is not. Google's Gemini 3.8 Flash landed the next day and is holding current pricing through December 31, then doubling on January 1. If you are building anything on the Gemini API, budget for that now.

Why a web guy cares about this

Because the same problem kills WordPress sites. Bloated history, never cleaned up, slowing everything down until it breaks. Whether it is your AI chats or your database, the fix is the same: keep it tight, summarize, start fresh. If your site feels like message 40 of a long chat, that is what the care plans at webguynick.com are for.

Frequently asked questions

Does editing a prompt really save money in ChatGPT or Claude?

Yes. Both re-send the full conversation with every message. Editing replaces the old branch instead of adding to it.

How long is too long for an AI chat?

Around 15 to 20 messages for most work. Ask for a handoff summary and start fresh.

author avatar
Web Guy Nick
X
Welcome to our website
Start a project →