Building & tuning

What are tokens in AI?

Your AI running cost is essentially a token bill. Longer prompts and longer answers cost more, so cost control is a design decision, not an afterthought.

In plain terms

Tokens are the small chunks of text that AI models read and generate, roughly three-quarters of a word each, and the unit most AI usage is billed in.

Tokens are the small pieces a model reads and writes. A token is often a word or part of a word. Vendors bill by the token, so a long document in and a long answer out both cost money. Roughly, a few hundred English words are a few hundred tokens, not an exact word count.

This is the unit on the invoice. A chatbot that writes a page when the user needed a sentence costs several times more and feels slower. Cap the answer length for routine jobs. Keep instructions short. And remember that a pasted contract, a chat history, and the answer are all on the same bill.

Think of it this way: Tokens are to AI what minutes are to a phone plan. You are billed for every word in, every word out, and every system prompt. Volume is what drives the bill.

A team builds a reporting assistant that summarizes 10-page PDFs. At scale the per-report token cost makes the unit economics break. Redesigning the prompt to send only key sections cuts running cost by 70 percent.

A support bot was told to be thorough. It answered a password-reset question with eight paragraphs. Shortening the instruction to three steps cut the tokens per chat and the wait. Customers finished the reset. The monthly bill dropped without a model change.

Token counting belongs in every AI cost model from day one. Map your typical input and output sizes before you pick a model and before you price any product feature. You cannot avoid tokens; you can optimize them. Caching repeated system prompts, truncating unnecessary context, and routing simple tasks to smaller models are the main levers. The primary unit of ongoing AI cost.

RaftLabs points the model at your documents and your rules, then checks the answers against cases you already trust. That surrounding work is where these projects succeed or stall. The related work on our side is AI development.

This sits with the other building & tuning terms on the glossary. How a general model gets pointed at your documents, your tone, and your workflow. Worth reading next: Fine-tuning, Retrieval-Augmented Generation (RAG), and Prompt Engineering.

Common questions

You pay for tokens in and tokens out, on every request. A quiet pilot hides the number. Estimate a busy day: how many requests, how long the documents are, how long the answers are. Then ask the vendor's price per token and put a ceiling on runaway answers.
Close, not equal. Common words may be one token. Long or unusual words split into several. Numbers, code, and other languages convert differently. Use the vendor's own counter on a real document before you budget, rather than multiplying a word count.

Work with us

Tell us what's broken.

Tell us what's not working in your business. We'll find the real problem and tell you exactly what it would take to fix it.

  • Scope and cost agreed before work starts. No surprises. No obligation.
  • Working prototype within 3 weeks of kickoff.
  • Pay by milestone. You see progress before each invoice.
  • 60-day post-launch warranty. Bug fixes, UI tweaks, and deployment support. No retainer.
  • All conversations are NDA-protected.