RAG Development Services

RAG development for answers that must point back to your source.

A fluent answer is not enough when the question concerns your policy, product, contract, or operating procedure. RaftLabs develops retrieval-augmented generation systems that index approved knowledge, preserve access boundaries, retrieve supporting passages, cite them, and refuse when the available evidence is insufficient.

Bring the problem, the current workflow, or the existing code. We reply with a practical next step within one business day.

Evidence and scope

4 to 8 weeks

First release

One use case, one or two sources, and a labelled question set.

$15K

Starting scope

Ingestion, retrieval, citations, evaluation, and handover.

Defined first

Release rule

Retrieval and answer quality measured on your questions.

Evidence · planning contextSee the work

The brief

Start with what is not working.

Good software decisions begin with the constraint, not a list of features or a preferred technology.

01

Does the assistant answer confidently when the supporting document is missing, outdated, or outside the user's permissions?

02

Can you tell whether a bad answer came from retrieval, source content, or the model itself?

Plain answer

RAG development lets a language model answer from approved private knowledge and cite the evidence it found. RaftLabs develops ingestion, search, permissions, refusal paths, and tests. A focused system for one use case and 1 to 2 sources starts at $15,000 and takes 4 to 8 weeks.

The answer cited a policy. The policy had expired six months earlier.

The assistant found a relevant paragraph, but the index contained two versions. It chose the older one, wrote a confident answer, and displayed a citation that made the mistake look trustworthy.

RAG is not accurate because it has a vector database. It becomes dependable when source versions, permissions, retrieval tests, citations, and refusal behaviour are designed as one system.

First scope

1 to 2
knowledge sources in a focused first release
Keep the evaluation boundary clear
4 to 8 weeks
typical focused delivery window
Depends on source and access complexity
$15K
starting development price
Fixed after source and question review

RAG earns its cost when answers must use private knowledge and show their evidence.

A plain model call is simpler when the necessary context is small, stable, and already supplied in the request.

A fit
01

Employees or customers repeatedly ask questions answered by approved documents, tickets, or records.

02

Knowledge changes often enough that citations, versions, and deletion behaviour matter.

03

Different users have different permission boundaries over the same source collection.

Not a fit
01

The task is mainly tone, style, or output-format control rather than knowledge retrieval.

02

The answer depends on judgment that no approved source contains.

03

The source material is contradictory and nobody can decide which version governs.

Scope

What a production RAG system needs

  • 01
    Ingestion with source identity
    Parsers preserve document, section, date, version, link, and access metadata. Update and deletion paths keep the index aligned with the approved source instead of creating an unmanaged copy.
  • 02
    Retrieval tuned on real questions
    Dense and keyword search, metadata filters, and reranking are chosen from the corpus and query set. Retrieval is evaluated before answer fluency can hide a missing passage.
  • 03
    Permission-aware context
    User identity and source permissions constrain results before text reaches the model. Tests cover allowed, denied, mixed-access, and revoked-document cases.
  • 04
    Citations, refusal, and monitoring
    Answers link to the supporting passage and version. When the evidence is weak or absent, the interface says so and offers the next useful step. Logs separate source, retrieval, generation, and cost failures.

Should you use RAG or fine-tuning?

RAG vs fine-tuning

RAGFine-tuning
Best fitPrivate or changing knowledge with source evidenceRepeatable behaviour, style, format, or a narrow task
Knowledge updateRe-index the changed or deleted sourcePrepare data, retrain, and evaluate another model version
CitationCan link the answer to retrieved passagesDoes not expose which training example supports a claim
Main failureWrong, missing, stale, or unauthorized retrievalInsufficient examples or behaviour that fails to generalize
First testCan search find the right evidence?Does training improve the task beyond prompting?

For broader model integration controls, see LLM integration services. AI knowledge management owns the employee-facing workflow, while AI semantic search owns search experiences that may not generate an answer.

How it works

From source inventory to evaluated answers

A labelled question set keeps retrieval and generation honest.

  1. Phase 1
    01

    Define questions and evidence

    Choose one workflow, approved sources, user groups, and a labelled set of questions with expected passages, unacceptable answers, and permission cases.

  2. Phase 2
    02

    Prove ingestion and retrieval

    Test parsing, chunking, metadata, versions, permissions, hybrid search, and reranking on representative content. Fix missing evidence before adding generation.

  3. Phase 3
    03

    Add answers and refusal paths

    Generate from the retrieved context, display source links, validate access, and decline or escalate when evidence is insufficient.

  4. Phase 4
    04

    Release and monitor quality

    Measure retrieval misses, unsupported statements, stale sources, latency, cost, and user escalations before expanding the corpus or audience.

Where RAG systems usually fail

A fluent answer hides weak retrieval
Score whether search found the expected evidence before judging the prose. The model cannot repair a missing source passage reliably.
Permissions are filtered after generation
Unauthorized content has already crossed the boundary once it reaches the model. Filter during retrieval and test revoked access.
Old versions remain searchable
Every connector needs update, supersession, and deletion behaviour. A citation to an expired policy is still a wrong answer.
A universal accuracy number replaces evaluation
Quality depends on the buyer's corpus, questions, and acceptance rules. We do not promise a benchmark percentage before testing those inputs.

Scope and price

A focused RAG system starts at $15,000.

Begin with one use case, one or two sources, a labelled question set, retrieval, citations, access controls, and a handover path.

Multi-domain systems grow in fixed phases after the first corpus clears its agreed retrieval and answer thresholds.

Starting investment

Starts at $15,000

A focused first release usually takes 4 to 8 weeks. Source quality, custom connectors, and permission complexity move the estimate most.

Retrieval measured separately

The evaluation shows whether the right evidence was found before generation quality is judged. That makes failures diagnosable.

Fixed-price phase

Once the sources, questions, permissions, acceptance rules, and handover are agreed, the phase price is locked in writing.

Useful next steps

More on RAG & knowledge management

RAG development questions

Retrieval-augmented generation retrieves relevant passages from an approved knowledge source before a language model answers. The response can cite those passages and reflect newer private information without retraining the model. The system still needs evaluation because retrieval can return irrelevant, incomplete, or unauthorized context.

Use RAG when answers must use private or frequently changing knowledge and show their supporting source. Fine-tuning fits repeatable behaviour, tone, format, or a narrow task when prompting has reached its limit. Some systems combine them, but retrieval should not be added when the required context already fits safely in the request.

We create a labelled set of representative questions, expected source passages, unacceptable answers, and permission cases. Retrieval is assessed separately from answer generation so the team can tell whether a failure came from the source, parser, index, search, reranker, prompt, or model.

Yes, when the source identity, document metadata, and application authorization model support it. Permission filters must be applied before retrieved text reaches the model, then tested with positive and negative cases. A RAG layer should never broaden access beyond the source system's approved policy.

A focused system for one use case and one or two sources starts around $15,000 and usually takes 4 to 8 weeks. Custom connectors, scanned content, several permission tiers, larger corpora, custom interfaces, and stricter evaluation requirements increase the scope.

Work with us

Bring the question your current search cannot answer.

Show us the approved sources, user groups, and evidence a correct answer should cite. We will scope one measurable retrieval workflow.

  • Scope and cost agreed before work starts. No surprises. No obligation.
  • Working prototype within 3 weeks of kickoff.
  • Pay by milestone. You see progress before each invoice.
  • 60-day post-launch warranty. Bug fixes, UI tweaks, and deployment support. No retainer.
  • All conversations are NDA-protected.