RAG as a Service

RAG as a service for knowledge that does not stay still.

A RAG system does not stay reliable after launch by itself. Sources change, permissions move, real questions reveal new misses, and model providers change. We make and operate one bounded retrieval workflow, then measure, correct, and document it with your team.

See our work

Bring one live or planned knowledge workflow. Leave with a build, hand over, manage, buy, simplify, or stop recommendation.

The brief

Start with what is not working.

Good software decisions begin with the constraint, not a list of features or a preferred technology.

01

The RAG prototype can answer a demo question, but nobody owns stale sources, failed retrievals, access changes, or the next model update.

02

Your team wants a managed retrieval service without giving up visibility, client-controlled assets, or a practical way to change providers later.

Plain answer

RAG as a service is an ongoing managed service for a retrieval-augmented generation workflow. It covers the source sync, permission-aware retrieval, evaluation, monitoring, failure review, changes, and incident handling needed after launch. Every RaftLabs project starts at $9,500 with one bounded workflow, written responsibilities, and a documented exit path.

What to remember

  • Choose RAG as a service when the knowledge, questions, permissions, or model behaviour will keep changing and an internal team cannot own every operating task.
  • Keep retrieval quality, answer support, access checks, source freshness, latency, and cost visible instead of reducing the service to a single uptime number.
  • Put client responsibilities, provider usage charges, artefact ownership, service hours, change control, and handover terms in writing before release.

Managed should describe a responsibility, not just the hosting model.

Start with the smallest arrangement that gives the workflow a credible owner after release.

ScenarioRecommendationReason
Users mainly need ranked documents or passagesImprove searchDo not add generated answers and model operations when people should inspect the source themselves.
A standard product fits the sources, permissions, controls, and exit pathBuy the productUse its maintained connectors and administration instead of commissioning a custom operating layer.
Your platform team can own the system after releaseBuild and hand overUse project-based RAG development, with evaluation and operating notes transferred to the internal team.
Knowledge, permissions, questions, or providers will keep changingUse RAG as a serviceGive source sync, evaluation, failures, changes, and incidents a named operating owner.
The source of truth and acceptable answer are still disputedFix governance firstA managed service cannot resolve missing ownership or contradictory policy through retrieval tuning.

Fit

Use a managed RAG service when the work continues after launch.

The service is a fit when ongoing source, quality, access, and provider changes matter more than simply keeping an endpoint online.

A fit
01

A repeated knowledge workflow has identifiable users, approved sources, reviewers, and a measurable cost when people search, verify, or correct it manually.

02

The system needs permission-aware retrieval, source updates, citations, evaluation, incident handling, and regular changes that no internal team can fully own.

03

The organisation wants an operating partner but still needs visibility into quality, usage cost, changes, client-controlled assets, and the path to handover.

Not a fit
01

The task is a deterministic lookup, calculation, or report that conventional software can return more reliably.

02

The source material is missing, contradictory, unowned, or updated outside any dependable system of record.

03

The organisation expects a provider to guarantee every answer or take responsibility for legal, clinical, financial, policy, or operational decisions.

If a simpler search product, direct model context, or a project handover is enough, the first conversation should make that visible before a managed service is proposed.

The index stayed online. The knowledge did not.

A policy is replaced on Monday. One connector misses the deletion event. The older version still ranks well because it contains the buyer's exact phrase. On Thursday, the assistant gives a polished answer and cites the expired page.

Nothing technically crashed. Uptime stayed green. The content owner did their job, but the update never reached the assistant. The fault could sit in ingestion, retrieval, or review, which is why a generic hosting promise would not catch it.

A managed RAG service needs an operating definition of healthy: the expected sources arrived, changed access is reflected, representative questions still find the right evidence, unsupported answers are contained, and someone can trace what changed when a result regresses.

Knowledge-base indexing specification and retrieval evaluation worksheet for a managed RAG workflow

Who owns what after launch?

ResponsibilityManaged software productRAG build and handoverRAG as a service
Source connectorsVendor maintains supported connectorsClient owns them after handoverNamed in the service boundary, including failed syncs and exceptions
Answer qualityPlatform supplies standard tools and metricsClient runs the evaluation processRaftLabs runs agreed checks; client experts settle meaning and policy
PermissionsLimited to the product's identity modelClient operates the implemented controlsAccess behaviour is monitored and reviewed within the agreed source and identity limits
Changes and incidentsHandled under the vendor's product termsClient owns release and incident responseSupport window, severity, containment, approval, rollback, and communication are written into the service
CostSubscription plus usage and add-onsProject cost plus client operations and provider usageStarting phase, managed fee, provider usage, licences, and change work are shown separately
ExitDepends on platform export and contract termsHandover is the planned end stateClient-controlled assets, export limits, documentation, and transition steps are agreed before operation

Managed scope

The work that keeps a retrieval workflow useful

Not every engagement needs every item. The proposal names the exact operating boundary, the client dependencies, and what counts as change work.

  • 01

    Source sync and reconciliation

    Watch approved sources, connector runs, parsing failures, updates, permission changes, archives, and deletions. Reconcile what the source says should exist against what the index can retrieve, then route content exceptions to the person who owns the source.
  • 02

    Retrieval and answer evaluation

    Maintain representative questions, expected sources, access cases, no-answer cases, and important failure groups. Score whether the right evidence was found before judging whether the response used that evidence correctly.
  • 03

    Feedback and correction handling

    Give users a useful way to report an answer, citation, source, or access problem. Trace the fault to the source, parser, metadata, sync, retrieval, reranking, prompt, or model, and turn important corrections into repeatable regression checks.
  • 04

    Access and security operation

    Test allowed and denied retrieval paths, review identity or role changes, protect secrets, limit sensitive logs, and define containment when the system exposes or appears to expose material to the wrong audience. The client remains responsible for identity policy and access approval.
  • 05

    Provider, model, and cost changes

    Measure important model, embedding, index, retrieval, or prompt changes against the same evaluation set before release. Track provider usage separately from the service fee so a quality improvement does not hide an unacceptable increase in cost or latency.
  • 06

    Release, incident, and exit records

    Keep the configuration, evaluation result, release note, known limitation, rollback route, and incident decision attached to each meaningful change. Maintain the code, data preparation, account map, and operating notes needed for a practical handover.

Managed does not mean outsourced judgement.

RaftLabs can operate the technical path from an approved source to a retrieved passage and a visible response. We can show which source was indexed, which passages were found, what the model received, what it returned, and where the system crossed an agreed threshold.

Your subject-matter owners still decide which source governs, what an acceptable answer means, who may see it, when a person must review it, and whether the result may shape a business decision. That division is useful. It prevents the engineering team from inventing policy and prevents the business team from being handed an opaque system it cannot challenge.

The operating agreement should name both sides of that boundary in plain language. If the client inputs are missing, a managed provider cannot quietly substitute its own judgement and still call the output reliable.

Operating model

Close one operating risk before widening the service.

Each phase produces a decision the buyer can inspect, not just activity inside a black box.

  1. 01
    Boundary

    Write the service boundary

    What exactly is being managed, for whom, and during which hours?

    Follow one repeated question from its approved source to the person who uses the answer. Record the current baseline, owners, access rules, assurance needs, provider constraints, and what the workflow may influence.

    Decision produced

    Named users, questions, sources, responsibilities, prohibited uses, support coverage, escalation, and client dependencies.

    Risk closed

    A buyer pays for managed RAG but discovers that the important failure sits outside the contract.
  2. 02
    Evidence

    Establish the baseline

    Can the current or proposed system be evaluated before the service begins?

    Test source intake, retrieval, answer support, citations, refusal, permissions, latency, and cost separately. Repair or simplify the path before attaching an operating promise to it.

    Decision produced

    A versioned question set, expected evidence, access cases, no-answer cases, measures, thresholds, and known limits.

    Risk closed

    The provider keeps tuning a system without a stable way to show whether it improved or regressed.
  3. 03
    Control

    Release to a bounded audience

    What happens when a source, answer, permission, or provider fails in real use?

    Start with the smallest group that can produce useful evidence. Compare live questions with the evaluation set, inspect misses, and separate a content problem from a technical failure.

    Decision produced

    Monitoring, feedback, review queues, severity, containment, communication, approval, release, and rollback procedures.

    Risk closed

    A wrong answer becomes a long-running operating problem because nobody can see, route, or contain it.
  4. 04
    Continuity

    Operate, change, and exit cleanly

    Can the service change without losing quality, cost control, or the ability to leave?

    Review failures and costs with the client owner, test meaningful changes before release, record decisions, and keep the assets another qualified team would need to take over.

    Decision produced

    Review rhythm, change approval, usage reporting, regression evidence, renewal terms, exports, documentation, and transition steps.

    Risk closed

    The workflow becomes too expensive or too dependent on one provider to change safely.

Evaluation needs more than one accuracy number

A retrieval workflow can fail before the model writes or after. A source may be missing, or the parser may damage a table. The index may hold an old version. A permission filter can remove the right passage. Even with good evidence, the model can still make an unsupported claim. One blended score hides which part needs attention.

The evaluation plan therefore separates retrieval from the generated response. It also groups results by the cases that matter to the buyer, such as exact identifiers, conflicting policies, different access roles, recently changed content, questions with no approved answer, or workflows that require a person to review the result. Amazon Bedrock's current evaluation guidance similarly distinguishes retrieve-only evaluation from retrieve-and-generate evaluation and includes measures for context relevance, coverage, correctness, faithfulness, and citations. Read the Amazon Bedrock evaluation documentation.

No universal threshold makes every RAG workflow safe. A support drafting assistant, an employee policy search tool, and a regulated research aid have different consequences, reviewers, and escalation requirements. The service agreement should state the measures and release decisions for the actual workflow rather than borrow an impressive percentage from another system.

Service contract

Six things to settle before calling RAG managed

Service boundary
Name the users, question class, sources, environments, components, support hours, incident severity, response targets, exclusions, and work treated as a new change.
Client responsibilities
Name the source, identity, security, policy, subject-matter, and commercial owners plus the decisions and access approvals only they can make.
Quality and change
Define the evaluation set, measures, review rhythm, correction path, release evidence, approval authority, known limitations, and rollback conditions.
Data and access
Document systems of record, regions, providers, subprocessors, secrets, permissions, retention, deletion, logs, review access, and what must never enter model context.
Commercial model
Separate the first phase, recurring service fee, provider usage, infrastructure, licences, support coverage, limits, and the events that change price.
Ownership and exit
Record client-controlled accounts, project-specific artefacts, third-party terms, exports, operating documents, notice, transition help, and the work required to change providers.

Starting point

Every managed RAG engagement starts at $9,500.

The first paid phase creates evidence and an operating boundary before you commit to continuing management. It may end with a managed service, a project handover, a product recommendation, a repair plan, or a decision not to proceed.

Price and timing depend on source condition, connectors, permissions, evaluation depth, the existing system, deployment, query and update volume, assurance needs, service hours, and the length of the operating period. The proposal names those assumptions before paid work begins.

What the first phase can be

  1. 01

    Managed-service readiness audit

    Trace one workflow, inspect the current architecture and operations, and decide what is ready to manage, repair, replace, or stop.

  2. 02

    Evaluation baseline

    Build representative questions, expected evidence, access cases, measures, thresholds, and a repeatable comparison for future changes.

  3. 03

    Takeover and repair plan

    Locate recurring failures, stabilise the important path, and define the service boundary for an existing RAG system.

  4. 04

    First operated workflow

    Release one bounded source-to-answer path with monitoring, review, escalation, change control, and handover records.

We recommend the smallest first phase that can produce a defensible manage, hand over, buy, simplify, or stop decision.

Starting investment

$9,500

Minimum first-phase scope. The continuing service fee, infrastructure, model and vector-store usage, third-party licences, and extended support are priced separately.

Price held for the phase

The agreed first-phase price does not change unless you approve a material change in scope.

Client-controlled assets

Project-specific code, evaluation cases, source accounts, and operating notes remain under client control where provider terms allow.

Written exit path

The proposal names exports, third-party limits, transition artefacts, notice, and handover responsibilities before operation begins.

Common questions

RAG as a service is an ongoing managed service for a retrieval-augmented generation workflow. The provider operates agreed parts of source ingestion, indexing, retrieval, permissions, evaluation, monitoring, incident response, changes, and support after release. The agreement should also name what the client still owns, including source authority, business policy, subject-matter review, user access, and decisions made from the output.

RAG development is a project to assess, make, repair, or hand over a retrieval-and-answer system. RAG as a service adds a defined operating period and continuing responsibilities after release. Use a project handover when your internal team can own source changes, evaluation, incidents, costs, and provider updates. Use the managed model when those tasks need a named outside owner and review rhythm.

No. A managed product provides a standard platform, connectors, administration, and commercial model. It can be the right answer when its source, permission, retrieval, deployment, evaluation, and exit compromises fit. A managed RAG engagement may configure such a product, add application-specific components, or operate a custom system. The first phase should compare those paths before assuming custom software is necessary.

The agreed service may cover source sync and reconciliation, parsing failures, index changes, retrieval and reranking, permission checks, evaluation runs, failed-query review, citations, refusal and escalation paths, model or provider changes, logs, alerts, cost review, releases, rollback, incident handling, and operating documentation. The exact boundary depends on the source systems, deployment, risk, and what your team already owns.

Yes, where the chosen services and account model allow it. We prefer client-controlled source, identity, cloud, model, vector-store, analytics, and monitoring accounts when practical. The proposal names any service that must remain vendor-managed, who can access it, where data is processed, how secrets are handled, what can be exported, and what is required to move the workload.

We define how each approved source signals a create, change, permission update, move, archive, or deletion. The managed service monitors sync jobs and failures, reconciles the source against the index, tests representative allowed and denied users, and records exceptions that need a source owner. This only works as well as the source API, identity data, and client governance allow, so unsupported or ambiguous cases are stated explicitly.

We keep a versioned set of representative questions and expected evidence, including access cases, conflicts, outdated material, exact terms, and questions the system should not answer. Retrieval is evaluated separately from the generated response. Depending on the workflow, monitoring can cover source freshness, retrieval relevance and coverage, answer support, citation quality, refusal, access, latency, cost, failure queues, and changes between releases.

The first task is to locate the failure. It may come from the source, parsing, metadata, sync, permissions, query handling, retrieval, reranking, context assembly, prompt, or model. The service agreement defines how users report a problem, what is logged, who reviews the meaning, how urgent cases are contained, when the system should refuse or escalate, and how a correction becomes a regression test.

No. Retrieval can give a language model relevant evidence, but the system can still find the wrong passage, miss a source, use an old version, misunderstand context, or make an unsupported inference. Citations, separate retrieval and answer evaluation, refusal rules, and human review reduce and expose risk. They do not make every output correct or transfer accountability for business decisions to the service provider.

Your team owns the business question, approved systems of record, source authority, user and role decisions, retention and policy, prohibited uses, acceptable failure, subject-matter judgement, and decisions made from the output. A managed service can operate the technical path and make failures inspectable. It cannot invent ground truth or decide which policy should govern when your organisation has not settled it.

Every RaftLabs project starts at $9,500. The first paid phase may be a managed-service readiness audit, evaluation baseline, takeover and repair plan, or first operated workflow. The proposal separates that work from continuing management, infrastructure, model and vector-store usage, third-party licences, support hours, and client duties. Ongoing cost depends on source volume and change rate, query volume, permissions, evaluation depth, deployment, assurance needs, and service coverage.

The exit path should be agreed before the service begins. We document client-controlled accounts, project-specific code, data preparation, evaluation cases, configuration, release and rollback notes, known limitations, exports, third-party licences, and the steps another team would need to operate the workflow. A provider change may still require migration work, but it should not require rediscovering how the system works from zero.

Work with us

Bring the RAG workflow your team does not want to operate by guesswork.

In a 30-minute call, we will help identify whether the sensible next move is a managed platform, a project handover, a service-readiness audit, a repair, a managed operating agreement, or no RAG at all.

  • One audience, question class, source boundary, and accountable client owner before expansion.
  • Retrieval, answer support, access, freshness, latency, cost, and failure handling evaluated separately.
  • Build cost, provider usage, managed fees, support coverage, client duties, and exit terms shown separately.
  • Client-controlled assets and operating notes retained for handover where provider terms allow.