A practical guide to getting more from every prompt – and making your AI subscription go further.
Most people who use Claude hit their usage limit far too quickly. They assume the plan is too small, so they upgrade, only to hit the wall again a week later. The real problem isn’t your subscription tier. It’s how you’re prompting. Every message you send to Claude is measured in tokens – roughly one token per word of input and output combined. The longer your conversation grows, the more tokens each new message costs, because Claude re-reads the entire conversation history every time you press enter. Understanding this one principle is the key to unlocking dramatically better value from your credits.
In this article, I’ll walk you through the habits, techniques, and prompt structures that I use daily, and that I teach to every client I work with. Some are small changes that take seconds. Others will reshape the way you work with AI entirely. All of them will extend your credits.
Understand What Actually Costs You Tokens
Before we look at solutions, let’s understand the problem. Every interaction with Claude has two token costs: input tokens (everything Claude reads, including your prompt, the conversation history, and any uploaded files) and output tokens (everything Claude writes back to you). As a conversation grows longer, the input cost of every subsequent message balloons, because Claude must re-process the entire thread.
That means message one in a new chat is cheap. But message thirty? Claude is re-reading twenty-nine previous exchanges before it even begins thinking about your new question. This is, by a significant margin, the biggest reason credits disappear.
KEY INSIGHT
File uploads are surprisingly expensive. A single PDF page can cost 1,500 to 3,000 tokens. Screenshots and images can be even more. A Word document or PowerPoint carries hidden metadata bloat you’ll never see. Wherever possible, extract just the relevant text and paste it in directly.
The Top 10 Prompt Hacks to Save Credits
1 – Start Fresh Conversations Often
This is the single highest-impact habit. Once a conversation reaches 15 to 20 messages, the accumulated context is costing you heavily on every exchange. Start a new chat for each distinct topic or task. If you need to carry context forward, ask Claude to write a brief summary at the end of the current chat, copy it, and paste it into the new one.
2 – Be Specific From the First Message
Vague prompts generate vague answers, which trigger follow-up corrections, which burn tokens on the bloated context of the failed first attempt. Front-load your requirements. Tell Claude the format you want, the length you need, and the audience you’re writing for, all in message one. One precise prompt is always cheaper than three rounds of refinement.
TOKEN-HUNGRY
“Can you help me write something about our new product? It’s a CRM tool. Maybe a blog post? Keep it professional.”
TOKEN-EFFICIENT
“Write a 400-word blog post announcing our new CRM tool, Relay. Audience: small business owners. Tone: confident and friendly. Include three benefits and a call to action.”
3 – Constrain the Output
When you don’t tell Claude how much to write, it defaults to being thorough, which means long, detailed responses packed with explanations and caveats you probably didn’t need. Add explicit constraints: “Reply in under 100 words,” “Numbered list only,” “No explanation, just the code,” or “Three bullet points maximum.” This cuts output tokens dramatically.
4 – Edit Your Prompt Instead of Sending Corrections
If Claude misunderstands your first message, resist the urge to type “No, I meant…” as a follow-up. That follow-up stacks on top of the full conversation history, and Claude re-reads all of it. Instead, use the edit message feature (click the pencil icon on your original message). This effectively re-sends a refined prompt without the cost of the failed exchange clogging the context window.
PRO TIP
The edit-and-resend approach is especially powerful in the first few messages of a chat. It’s essentially free do-overs, without the token penalty of accumulating failed attempts in the conversation history.
5 – Use the Right Model for the Job
Not every task needs the most powerful model. Claude offers several tiers: Haiku is fast and cheap, ideal for simple lookups, formatting, and quick answers. Sonnet handles the vast majority of daily work brilliantly. Reserve Opus for genuinely complex reasoning, multi-step analysis, or creative tasks where nuance matters. Choosing Sonnet over Opus for routine work can stretch your credits significantly.
6 – Don’t Upload Entire Files When a Snippet Will Do
Uploading a 30-page PDF when you only need information from page 7 is one of the most common, and most expensive, mistakes. Copy and paste the relevant section as plain text instead. If you must upload a document, tell Claude exactly where to look: “Refer only to section 3.2 on page 12.” This focuses its attention and reduces the processing overhead.
7 – Structure Your Prompts With Clear Sections
Claude responds well to structured input. Use labels like CONTEXT:, TASK:, FORMAT:, and CONSTRAINTS: to organise your prompt. This eliminates ambiguity, which means Claude spends fewer tokens on hedging, clarifying, or asking questions. Structured prompts consistently produce better results in fewer exchanges.
EXAMPLE PROMPT
ROLE: You are a marketing copywriter for a B2B SaaS company.
CONTEXT: We’re launching a new feature called “Smart Alerts” that sends automated notifications based on user behaviour.
TASK: Write an announcement email to existing customers.
FORMAT: Subject line + body text. Under 200 words total.
TONE: Professional, enthusiastic, not salesy.
8 – Use Projects and Custom Instructions
Claude’s Projects feature lets you set persistent context and custom instructions that apply to every conversation within that project. Instead of re-explaining your company, your tone of voice, or your preferred output format in every single chat, set it once in the project instructions. This saves tokens on every message because you’re not repeating boilerplate, and Claude still has full access to it.
9 – Ask for Outlines Before Full Drafts
When tackling a big piece of work, whether a report, a strategy document, or a presentation script, ask Claude for an outline first. Review it, tweak the structure, then ask it to flesh out the final version. This two-step approach is far cheaper than generating a 2,000-word draft, deciding the structure is wrong, and asking for a complete rewrite. The rewrite costs double because the failed draft is still sitting in the conversation.
10 – Monitor Your Usage and Plan Around It
Claude’s usage limits reset on a rolling window (roughly every five hours for paid plans). Check your usage in Settings > Usage and plan your most intensive work for the start of each window. Batch similar tasks together. It’s more efficient to do all your content drafting in one focused session than to scatter prompts throughout the day between unrelated conversations.
Bonus: The “Carry-Forward Summary” Technique
This is my favourite workflow for long-running projects. At the end of a productive session with Claude, before you close the chat, send this prompt:
EXAMPLE PROMPT
Write a concise session summary capturing: key decisions made, current status of the work, and the next steps to take. Format it as a handover note I can paste into a fresh chat to continue where we left off. Keep it under 300 words.
Copy the output. Start a new conversation. Paste it in as your opening message. You’ve just carried forward all the useful context at a fraction of the token cost, without dragging along dozens of messages of back-and-forth that Claude would otherwise re-read on every turn.
This single technique, used consistently, can reduce your token consumption by 40 to 60 per cent across multi-session projects.
The Golden Rule
Every tip above flows from one principle: treat your context window like a budget. Every word in the conversation, yours and Claude’s, has a cost that compounds with every subsequent message. The less noise in your thread, the further your credits go.
Be deliberate. Be specific. Start fresh often. And you’ll find that your current plan likely gives you far more capacity than you thought.
QUICK-REFERENCE CHEAT SHEET
New topic? Start a new conversation.
Wrong answer? Edit your prompt; don’t send a correction.
Simple task? Use Sonnet or Haiku, not Opus.
Big file? Paste the relevant snippet; don’t upload the whole thing.
Long project? Use the carry-forward summary at the end of each session.
Verbose output? Constrain it: “under 100 words” or “list only.”
If you’d like hands-on help optimising your team’s AI workflows, or if your business is curious about how to adopt Claude effectively, get in touch at trimontium.ai. We run workshops, one-to-one coaching, and bespoke integration projects designed to make AI practical, not theoretical.
Until next issue – prompt wisely.

