Back to Behind the Build
AI & FluencySep 30, 20265 min read

Is the Math Mathing?

A few weeks ago I was deep in a long chat, pasting in doc after doc, and it suddenly told me I'd hit my limit for the day. I hadn't asked that many questions. So I went looking for why.

If you've used an AI tool and wondered why it has limits, or why a bill looked bigger than you expected, this is for you. No tech background needed. We'll start at the very beginning and build up.

What is an LLM, really?

LLM stands for large language model. It's the engine behind tools like Claude and ChatGPT. Here's the simplest way to think about it: it's a very well read prediction machine. You give it text, and it predicts what text should come next, one small piece at a time. It's not looking things up in a database. It's writing, piece by piece, based on patterns it learned from a huge amount of reading.

Those small pieces matter, because they're what we're here to talk about.

What's a token?

A token is one of those pieces. Think of tokens like LEGO bricks. The model doesn't see a sentence. It sees a pile of bricks and snaps new ones on, one at a time.

A short common word like "deal" is one brick. A longer word like "unbelievable" gets split into a few. Spaces and punctuation count too. For English, 100 words is about 130 tokens. Code, numbers, and other languages take more bricks for the same idea.

"deal"deal1 brick"unbelievable"unbelievable3 bricksabout 130 tokens per 100 English words
Short words are one brick. Longer words get split into a few.

Why do we care? Because tokens are what the model counts. They're the unit of work, the unit of cost, and the unit behind your usage limits.

Two piles of bricks

Every time you use an AI, there are two piles. The stuff you send in is called input. The stuff it writes back is called output.

Output costs more than input, because writing takes more effort than reading. So every conversation is a little math problem: input bricks times the input price, plus output bricks times the output price. Prices are quoted per million tokens, so a single question costs a tiny fraction of a cent.

input: what you sendoutput: what it writescheaper per brickpricier per brick$$$
Every conversation is two piles, and the output pile costs more per brick.

I'm not listing exact prices here because they change often. If you want the real numbers, check the pricing page for the tool or model your team uses.

So if one question is basically free, where do the big numbers come from? Here's the part that surprises almost everyone.

The model has no memory

So how does a conversation feel continuous? Because every time you hit send, the app quietly sends the entire conversation again, from the very first message, plus your new one. The model re-reads the whole thing, then answers.

The same goes for anything you attach. Paste in a document, and that document rides along with every message after it.

The snowball

Let's make it real. Say each of your messages is about 100 tokens and each reply is about 300. Watch how much goes in on each turn:

  • Turn 1: 100 tokens sent in
  • Turn 2: 500 tokens sent in
  • Turn 3: 900 tokens sent in
  • Turn 4: 1,300 tokens sent in
  • Turn 5: 1,700 tokens sent in
100turn 1500turn 2900turn 31,300turn 41,700turn 5what you typedre-read history
You typed 500 tokens. The model read 4,500.

Five messages in, and 4,500 tokens went in. You only typed 500 of them. The rest was the model re-reading the conversation.

Now stretch it. A 5 message chat sends about 4,500 tokens in total. A 20 message chat sends about 78,000. Four times the messages, but seventeen times the input.

Is the math mathing? It is. And that's why long chats get expensive and hit limits fast. Cost doesn't climb in a straight line. It snowballs.

Why our brains miss this

Three quirks of the human brain make this easy to miss.

what our brains guesswhat actually happenswe'd guess 4010 people = 4520 people = 19051020people in the roomhandshakes
Five people means 10 handshakes. Twenty people means 190, not 40.
  • Linear thinking. We expect growth in a straight line. But think about a room where everyone shakes hands. Every new person has to meet everyone already there, so each arrival adds more handshakes than the last. Chats work the same way. Every new message gets read alongside everything before it.
  • The pain of paying. Handing over cash stings. Tapping a card barely registers. Tokens are even lighter than a tap. No receipt, no beep, so we don't feel the cost until we hit a limit.
  • The sunk cost trap. You stay an hour into a bad movie because you already bought the ticket. Same with a long chat. You're 40 messages in and it finally gets you, so starting over feels like losing it all. But that long chat is the priciest place to be.

Why this matters for us

Think about what we do every day across the GTM org: call transcripts, account research, RFP responses, demo prep. All of it involves pasting in big chunks of text and asking follow-up questions.

Picture someone pasting in a long call transcript and asking four questions about it. That transcript rides along all four times. Now picture a hundred people doing that several times a day. A tiny habit turns into a real line item.

What to do about it

A few habits that keep the numbers down:

  • Start a new chat when you change topics. Old context you don't need still costs you, on every reply.
  • Paste the section you need. Not the whole doc. Every brick you add gets re-read every time.
  • Be specific up front. A clear ask the first time skips the extra back and forth that grows the chat.
  • Tell the model the format you want. Table, email, or checklist, named up front, so you don't spend extra turns reshaping it.
  • Use a smaller, cheaper model for simple jobs. Save the bigger model for the work that actually needs it.

The one I swear by: when a chat starts getting long, I ask for a short summary of where we landed, then start fresh with just that summary.

The takeaway

You don't need to count tokens. You just need to know what spins the meter: the model forgets everything, so your whole chat gets re-read every time. Short chats, tight prompts, smart pasting. That covers most of it.

Here's my small ask: this week, try starting a new chat every time you switch topics. See if you notice the difference. The math will thank you.

Happy Learning,

KP

Keep reading