Building & tuning

What is a context window in AI?

It sets a hard limit on how much a model can read in a single pass. Larger windows cost more per use, which is why retrieval is often smarter than pasting everything in.

In plain terms

The context window is the amount of text an AI model can consider at once, covering both your input and its response.

The context window is how much text the model can hold in mind for one request. That includes your instructions, the documents you attached, the conversation so far, and the answer it is writing. Past that limit, something is dropped, often the earlier rules.

Bigger is not automatically better. Sending a whole handbook for every question costs more, runs slower, and can bury the two pages that mattered. For a long document, send the pages that match the question. For a long chat, start a new one when the task changes. If a rule must hold, repeat it in the standing instruction, not only in the first message.

Think of it this way: The context window is the AI's short-term memory. It holds everything in front of it, but once the conversation or document exceeds its capacity, earlier content starts to fall off the edge.

A company building a contract review tool discovers the model misses a key clause on page 40 because the full contract exceeds the context window. They switch to chunked retrieval to solve it.

A lawyer pastes a 200-page contract and asks for the termination clause. The model answers from the early pages and misses a later amendment that never fit. The fix is to pull the clauses that mention termination, then ask. The model now sees the amendment.

Any time the AI reads a long document or a long conversation. The limit decides whether you send the whole document or only the pages that matter. Pasting an entire knowledge base into the context window to avoid building a retrieval layer is an anti-pattern. It is slower, more expensive, and less accurate than selective retrieval.

RaftLabs points the model at your documents and your rules, then checks the answers against cases you already trust. That surrounding work is where these projects succeed or stall. The related work on our side is LLM integration.

This sits with the other building & tuning terms on the glossary. How a general model gets pointed at your documents, your tone, and your workflow. Worth reading next: Fine-tuning, Retrieval-Augmented Generation (RAG), and Prompt Engineering.

Common questions

The system drops text, usually from the start of the request, or it refuses the request. Dropped text can include the rule that said do not give a price. You will not always get a warning. Keep requests to the pages that matter, and put hard rules in the standing instruction.
Pay for it when the job truly needs the whole document in one pass, such as comparing two long files. For most questions, a search that fetches the relevant pages is cheaper and more accurate. Test both on the same ten documents before you commit to the larger plan.

Work with us

Tell us what's broken.

Tell us what's not working in your business. We'll find the real problem and tell you exactly what it would take to fix it.

  • Scope and cost agreed before work starts. No surprises. No obligation.
  • Working prototype within 3 weeks of kickoff.
  • Pay by milestone. You see progress before each invoice.
  • 60-day post-launch warranty. Bug fixes, UI tweaks, and deployment support. No retainer.
  • All conversations are NDA-protected.