Tokens are the small pieces a model reads and writes. A token is often a word or part of a word. Vendors bill by the token, so a long document in and a long answer out both cost money. Roughly, a few hundred English words are a few hundred tokens, not an exact word count.
This is the unit on the invoice. A chatbot that writes a page when the user needed a sentence costs several times more and feels slower. Cap the answer length for routine jobs. Keep instructions short. And remember that a pasted contract, a chat history, and the answer are all on the same bill.
Think of it this way: Tokens are to AI what minutes are to a phone plan. You are billed for every word in, every word out, and every system prompt. Volume is what drives the bill.
A team builds a reporting assistant that summarizes 10-page PDFs. At scale the per-report token cost makes the unit economics break. Redesigning the prompt to send only key sections cuts running cost by 70 percent.
A support bot was told to be thorough. It answered a password-reset question with eight paragraphs. Shortening the instruction to three steps cut the tokens per chat and the wait. Customers finished the reset. The monthly bill dropped without a model change.
Token counting belongs in every AI cost model from day one. Map your typical input and output sizes before you pick a model and before you price any product feature. You cannot avoid tokens; you can optimize them. Caching repeated system prompts, truncating unnecessary context, and routing simple tasks to smaller models are the main levers. The primary unit of ongoing AI cost.
RaftLabs points the model at your documents and your rules, then checks the answers against cases you already trust. That surrounding work is where these projects succeed or stall. The related work on our side is AI development.
This sits with the other building & tuning terms on the glossary. How a general model gets pointed at your documents, your tone, and your workflow. Worth reading next: Fine-tuning, Retrieval-Augmented Generation (RAG), and Prompt Engineering.