Every AI conversation eventually shifts to tokens. The word gets used as if everyone already knows what it means.
Maybe you have nodded and moved on. At this point, asking what a token is can feel a little like that meme: “And at this point I’m too afraid to ask.” I’ll admit that, even spending hours a week on this, I feel like the meme sometimes.
Bear with me. This is my longest article to date. Maybe you'll like it, but most importantly, I hope you find it useful. If you do, I will work on a part 2.
Why am I writing so much? Some AI products and company plans are replacing unlimited use with allowances, credits, or spending limits. Without understanding what consumes them, you can become so cautious that you never start, or run out before anything useful reaches the business. I want to explain the terms and help you control costs while continuing to deliver and innovate. GitHub now enforces premium-request allowances for Copilot, and Anthropic gives enterprise administrators controls over model access, effort, and spending.
Not understanding what a token is shouldn't stop you from creating, experimenting, and shipping things of value in your life and work.
We in HR like to understand the rules and the context before we make decisions. So let’s start with a few AI terms that technical people and AI companies use as if everyone already knows what they mean.
1. Token and tokenizer
A token is a small unit of information that an AI model processes. For text, it can be a whole word, part of a word, punctuation, or even a single character. When you upload a policy, spreadsheet, or slide deck, the words inside it can become text tokens. Images, files, tools, and request structure can also add to the input count. OpenAI explains how these inputs affect token counts.
Before the model can process your words, a tokenizer divides the text into those smaller pieces. A tokenizer does not divide words according to grammar or syllables, and two tokenizers may split the same sentence differently, making your token usage feel opaque.
For something more tangible, you can estimate that one token is about four characters, or roughly three-quarters of an English word. A 1,000-word policy might therefore use around 1,300 input tokens before your instructions, conversation history, or the model’s answer. That is an estimate, not a conversion formula. The exact count depends on the model, tokenizer, language, punctuation, and even the text surrounding a word. Yes, seriously.
Think of tokens as the units on a taxi meter. An economy car and a premium car can travel the same distance at different prices, just as different models can process tokens at different rates. The model may keep using tokens while reasoning through your request, like a meter running in traffic. It does not keep using tokens simply because the chat is open.
One important distinction: APIs publish explicit token prices. ChatGPT and Claude subscriptions may translate usage into allowances without showing a simple token bill. See the current OpenAI API models and prices and Anthropic pricing.

2. Model
A model is the trained system doing the work inside of, say, Claude and ChatGPT. It takes the tokens provided to it, identifies patterns and relationships, and generates the tokens that become its response.
Models vary in capability, speed, context window, reliability, token consumption, and cost. Some handle straightforward work quickly and cheaply. Others spend more time and tokens on complicated work.
Think about how you would staff a project. You would not hand a complicated workforce analysis to someone simply because their hourly rate is lowest. You choose someone with the level of capability that can complete the work accurately and reliably.
The most expensive model is not automatically the right choice. The cheapest model may not produce the cheapest completed work if it takes four failed attempts and another model to finish the job. A cheaper model can still produce quality work when matched to the right task.
Ruben Dominguez’s AI Corner breakdown of Martin Casado’s interview helped sharpen this tradeoff for me: token price matters, but so does the cost of completed work. You can also listen to the original interview.
3. Input, output, and reasoning
Tokens are consumed at different parts of your request. The three categories you need to understand are input, output, and reasoning.
Input is what the model receives for your request: what you type, existing instructions, relevant conversation history, documents, and results from tools such as web or file searches.
That input makes up the model’s context: what it has available when deciding how to respond. It does not automatically know about the document on your computer or the background you forgot to provide.
A context window is the maximum amount of information the model can work with at one time. Think of it as a workbench. A larger one holds more material, but irrelevant documents and a long conversation consume space and tokens without necessarily improving the result. The response also needs room.
Reasoning is the work between receiving your input and producing an answer. Some models use reasoning tokens to break down a problem, test answers, or work through several steps. You may not see that work, but it can still count toward usage.
Output is what the model sends back. Longer answers generally consume more output tokens, and providers may price them differently from input tokens. OpenAI’s token guide explains these categories.
Think again about assigning work to a consultant. The material you ask them to review is the input. What is on their desk is the context, and the desk can hold only so much at once. Their private analysis is the reasoning. The document, spreadsheet, or recommendation they return is the output. Providers measure those parts in tokens instead of billable hours.
This is why a short prompt does not always mean a cheap request. The model may also receive a long conversation, files, tool results, and standing instructions before reasoning and generating its answer.

Okay. So, now that we know what we are paying for, how do we use fewer tokens?
First, using fewer tokens cannot be the single goal. You can absolutely save tokens by choosing a cheaper model, giving it almost no context, and asking for a short answer, but the result may be unusable.
The goal is to spend enough tokens to complete valuable work without wasting them. Start with the finish in mind. What are you trying to finish? What does the model need to read? How complicated is the work? What would make the result useful, and who needs to review it?
If you cannot answer those questions, you are starting a taxi ride without knowing the destination. You may arrive somewhere, but you will spend more time and money finding out where you meant to go.
The opposite mistake is being so worried about the meter that you never take a ride. Companies can create this behavior when they impose tight usage limits without teaching people how to work within them. Employees either save their capacity for a perfect use case that never arrives or begin something ambitious and run out before producing anything the business can use.
Leadership then sees little finished work and decides AI is not improving productivity. That can mean tighter limits, less experimentation, and fewer useful results before employees have produced enough work to prove its value.
Cool, Mike. But what can I actually do by the end of this article to improve how I use AI?
Choose the right model, control unnecessary input and output, preserve useful work, and decide whether the result was worth the tokens.
Choose the least expensive model that can reliably complete the work
Many AI providers offer faster, less expensive models and more capable ones. Use the less expensive option for routine work, a reasoning model for multi-step work, and the most capable model where a mistake would be expensive. The names keep changing, so check the current OpenAI model catalog and Anthropic model guidance.
Depending on your plan and model, you may also be able to control effort. Low or medium effort can stretch your usage on routine work. Higher effort burns more tokens and is better saved for requests that need deeper reasoning.
Bottom line: even if model names change, start with the least expensive model that you reasonably believe can finish the work. If it fails twice for the same reason, stop paying it to fail and move up.

Control what the model reads and produces
Every old message, document, search result, and standing instruction can become input. Start a new chat when the old conversation is mostly irrelevant. Upload what the assignment needs, not everything that could conceivably matter.
Do not starve the model of context to save tokens. Missing facts create weak answers and repeated attempts. Give it the information that can change the result and leave out information that cannot.
The same applies to reasoning and output. You do not need the highest reasoning level to alphabetize a list. If you need a one-page manager guide or the five largest differences between two policies, ask for exactly that. You can always ask for more.
Shorter is not automatically better. The point is to avoid paying for work you will not read or use.
Test the approach and save the work
Guess, then check. If you have 500 employee comments to categorize, test the instructions on 20 first. If the categories are wrong, fix them before the model processes the other 480.
For longer work, save the research, decisions, outline, analysis, draft, or completed spreadsheet somewhere you control. Do not leave all of the value trapped inside one long chat.
When you move to another chat or model, bring over the smallest useful handoff: the goal, decisions already made, source material still needed, work completed, and next step.
Before you start the meter
Ask:
- What am I trying to finish?
- What information can change the answer and needs to be included?
- How costly would a weak answer or missed fact be?
- What is the least expensive model that can handle that risk?
- What output do I actually need?
- Where can I save useful work before moving to the next stage?
Afterward, ask whether you finished usable work, how much time it saved, whether it improved the work, how much correction it required, and whether you would use the same approach next time.
Before your next substantial request, answer the questions above. Then start the meter.
