AI Search and Semantic Search Development
AI search that understands meaning without losing exact matches.
We improve product, content, and internal search with the right mix of keyword, semantic, filtered, and reranked retrieval. Relevance is judged on the queries your users actually type, not on a polished demo.
Bring a query users keep reformulating and the result they should have found. Leave with a tune, buy, prove, rebuild, or stop recommendation.
The brief
Start with what is not working.
Good software decisions begin with the constraint, not a list of features or a preferred technology.
Users know the product, document, or record exists, but search returns nothing useful unless they remember the exact wording.
Adding semantic search improved some natural-language queries but pushed down product codes, names, filters, or other exact matches.
Plain answer
Semantic search retrieves results by meaning, while keyword search retrieves text that matches the words in a query. A dependable AI search system often combines both as hybrid search, then applies filters, business rules, permissions, and reranking. The right design depends on what users search, which exact terms must never be weakened, what a good result looks like, and whether the user needs ranked sources or a generated answer. Every RaftLabs project starts at $9,500 with a bounded first phase and written acceptance criteria.
What to remember
- Use semantic search when relevant content exists but users describe it differently from the words stored in the catalogue, document, or record.
- Keep keyword retrieval and structured filters when exact identifiers, names, categories, prices, dates, permissions, or business priorities must remain dependable.
- Evaluate search with representative queries and expected results before choosing embeddings, a vector database, reranking, or a complete platform rebuild.
Start with what the user should receive.
A better search box, a ranked source list, and a generated answer are different products. Choose the smallest result that helps the user finish the job.
| Scenario | Recommendation | Reason |
|---|---|---|
| Known item, exact identifier, or tightly filtered record | Tune keyword search and filters | Protect product codes, names, quoted phrases, categories, dates, prices, availability, and other exact constraints. |
| Users describe the right item in different words | Add semantic retrieval | Recover relevant products, documents, or passages when the query and stored text do not share the same vocabulary. |
| The search surface has exact and descriptive queries | Use hybrid search | Combine lexical precision and semantic recall, then rerank or apply business rules where the evidence supports it. |
| The user needs one answer supported by private sources | Use RAG | Retrieve permitted evidence, then generate a cited response with clarification, refusal, and separate answer evaluation. |
| The content is missing, stale, duplicated, or poorly structured | Repair the source first | A more sophisticated retrieval layer cannot return reliable information that the source does not contain or govern. |
Fit
Improve search when the right result exists but does not reliably surface.
The engagement is a fit when a search failure can be observed, judged, and connected to a real user or business consequence.
Users reformulate queries, abandon the search, browse manually, ask support, or use external search to find material that already exists.
The search surface mixes descriptive intent with exact names, codes, filters, permissions, freshness, availability, or business ranking rules.
Your team can provide representative content and someone who can judge why one result should rank above another.
The source content is not available, current, owned, or detailed enough to answer the user's need.
The real job is to summarize several sources, make a recommendation, or take an action rather than return ranked results.
The goal is to add a vector database or AI label without a query set, relevance owner, or measurable search problem.
If the current engine can be fixed with better fields, metadata, filters, synonyms, or ranking rules, the first recommendation can stop there.
The product code was exact. The result was only similar.
A buyer pastes a model number into a catalogue. A pure semantic search treats it like ordinary text and returns products with similar descriptions. The exact product appears lower, or not at all. On another query, a buyer writes what they need in plain language and keyword search returns nothing because the catalogue uses different words.
Neither query proves that keyword search is old or semantic search is better. They show two different relevance jobs on the same surface. Exact identifiers need lexical protection. Descriptive intent needs meaning. Price, availability, category, tenant, permission, and merchandising constraints may need structured rules before either result reaches the page.
That is why the useful starting artefact is a query and judgement set, not a vendor shortlist. It records what users type, what should appear, what must not appear, which constraints apply, and why the order matters. Architecture comes after that evidence.

An illustrative relevance worksheet. A real evaluation uses your query logs, searchable content, expected results, protected constraints, and reviewer decisions.
Keyword, semantic, hybrid, or RAG?
| Decision | Keyword search | Semantic search | Hybrid search | RAG |
|---|---|---|---|---|
| What it returns | Ranked results using lexical signals | Ranked results using meaning and similarity | Ranked results using lexical and semantic candidates | A generated answer based on retrieved evidence |
| Strongest use | Codes, names, exact phrases, filters, and predictable field boosts | Paraphrases, synonyms, concepts, and descriptive queries | Mixed query sets where exactness and meaning both matter | Questions that need synthesis across approved sources |
| Main risk | Relevant content is missed when the wording differs | Plausible but wrong similarities weaken exact or business-critical results | Fusion and reranking hide trade-offs if the query groups are not evaluated | The system retrieves weak evidence or writes beyond what the evidence supports |
| Buyer test | Can users name or filter the item precisely? | Does the right result exist under different wording? | Do both behaviours occur on the same search surface? | Does the user need a cited explanation rather than a source list? |
Current platform documentation reflects the same separation. Elastic describes hybrid search as combining lexical precision with vector similarity, while Microsoft documents semantic ranking as a later ranking step over an initial keyword or hybrid result set. Those are useful building blocks, not a reason to choose a platform before the search job is understood.
Production scope
The hard work is deciding what deserves to rank.
A first release can be narrow. It still needs the content, retrieval, ranking, interface, measurement, and operating decisions that keep search understandable.
- 01
Search behaviour and relevance model
Group real queries by job: known item, category, descriptive, comparative, filtered, misspelled, navigational, informational, and unsupported. Record expected results, relevance grades, protected exact behaviours, business rules, segments, and the consequence of a poor result. - 02
Content, catalogue, and index design
Map the fields, identifiers, variants, attributes, passages, metadata, language, dates, versions, ownership, freshness, and access data needed for retrieval. Expose missing, duplicated, stale, unparseable, or weakly described items before blaming the ranking model. If a separate vector index is justified, define stable source IDs, updates, deletions, re-embedding, reconciliation, versioning, rollback, and the team that will operate it. - 03
Lexical, semantic, filtered, and hybrid retrieval
Tune fields, analyzers, synonyms, typo handling, boosts, exact matches, filters, vector retrieval, fusion, and candidate depth around the query groups. Keep deterministic constraints outside similarity scoring when the rule must not become approximate. - 04
Reranking and business priorities
Rerank only a useful candidate set. Combine textual relevance with permitted signals such as availability, freshness, location, account context, quality, popularity, or merchandising without quietly turning relevance into whichever item the business wants to promote. - 05
Search interface and recovery states
Design query suggestions, filters, facets, snippets, highlights, result reasons, source links, spelling recovery, empty states, and reformulation paths around how people inspect results. The interface should help a user recover when confidence is low, not hide a weak result behind fluent copy. - 06
Evaluation, analytics, and search operations
Version the query set, expected results, configuration, and reviewer decisions. Monitor index health, zero-result and low-result queries, reformulation, result opens, conversion or completion signals, access failures, latency, cost, regressions, and the feedback path used to correct important misses.
A lower zero-result rate can still hide worse search.
Suppose the new system always returns something. The zero-result rate improves, but exact model numbers now surface similar products, unavailable stock ranks first, and a policy search favors an older page with more matching language. The dashboard is greener while the search experience is less dependable.
Search quality needs several views. Offline evaluation asks whether expected items appear high enough for representative queries. Segment checks show whether codes, names, descriptive phrases, filters, languages, and permission groups behave differently. Production behaviour adds reformulations, result opens, conversion or task completion, abandonment, and explicit feedback.
These signals still need judgement. A click may mean the result was useful, or that the user had no better option. A promoted item may convert well while answering the query poorly. We keep the ranking goal, commercial rule, and user behaviour visible as separate evidence so the team knows what it is optimizing.
What a credible relevance baseline contains
Keep the artefacts that let your team compare a change, challenge a result, and move platforms without rediscovering the search problem.
- 01
Representative query groups
Real successful and failed queries, exact identifiers, paraphrases, misspellings, filters, sensitive cases, unsupported requests, and the segments that matter commercially. - 02
Expected results and protected behaviour
Relevance grades, top-result expectations, acceptable alternatives, exclusions, exact-match guarantees, filters, access constraints, and the reason one item should outrank another. - 03
Versioned configuration and results
The content snapshot, index schema, retrieval settings, embedding and reranking choices, business rules, evaluation result, reviewer decision, latency, and cost attached to a meaningful change. - 04
Production feedback with context
Zero-result and low-result searches, reformulations, result opens, conversions or task completions, abandonment, manual corrections, and enough context to investigate without retaining unnecessary sensitive query data. - 05
Owners and correction path
Named owners for content, search relevance, permissions, commercial ranking rules, releases, incidents, feedback, and the decision to expand, simplify, buy, rebuild, or stop.
Proof
Relevant search product delivery

Musgrave receipt scanning software case study for SuperValu and Centra
Musgrave's post-campaign report records 1,062 users and 1,610 receipt submissions. At the reporting snapshot, 1,158 were AI approved, 409 AI rejected, 4 manually approved, 10 manually rejected, and 29 remained processing.

Eris Lifesciences gets 3,500 field employees logging 30 minutes daily on a training platform they actually use
Eris Lifesciences has 5,000+ pharmaceutical field reps who visit doctors to promote medicines. We built EMS Connect with a Learning Store, tests, gamification, and offline access. It reached 3,500+ daily active users averaging 30 minutes per session.
GrantHub demonstrates a real search-product constraint: users needed to narrow inconsistent records by country, industry, business type, funding stage, eligibility, and deadline while the underlying data kept changing. It is evidence of search, filtering, data modelling, freshness, and product delivery. It is not presented as proof of a semantic model or as a promise that another search surface will achieve the same result.
How it works
Close one relevance risk before opening the next.
Each phase ends with evidence and a decision. The next query group, source, or search surface is added only after the current one is understandable.
- 01Understand
Observe the search job
Who is searching, what are they trying to find, and what happens when the result is wrong or missing?
Review query logs where they exist, watch people search, inspect support language, and reproduce successful and failed journeys. Include the weird model number, vague description, misspelling, filtered query, stale item, and permission edge case that a polished demo avoids.
Decision produced
One search surface, user groups, query families, current path, expected results, exact constraints, filters, access, baseline friction, and accountable relevance owner.Risk closed
Replacing a search engine before separating a content problem, a query problem, a ranking problem, and an interface problem. - 02Measure
Build the relevance baseline
What should appear, in what order, and which behaviours must not regress?
Ask people who understand the domain to grade results and explain disagreements. Separate expected user relevance from commercial boosts and compliance rules. If nobody can judge the order, the project needs a decision owner before it needs a new model.
Decision produced
A versioned query and judgement set, content snapshot, protected exact cases, relevance grades, segment view, current scores, latency, and evidence gaps.Risk closed
Optimizing one headline metric while high-value queries, exact matches, small segments, or access boundaries quietly become worse. - 03Improve
Prove the smallest useful change
Can tuning, a managed feature, or a focused retrieval change solve the measured misses without creating a larger operating burden?
Start with the current engine and the simplest credible change. Compare lexical, semantic, hybrid, filtered, and reranked paths only where the evaluation shows a trade-off. Preserve exact and structured behaviour before adding approximation.
Decision produced
A compared approach, working proof or repair, evaluation result by query group, platform and cost trade-offs, interface behaviour, and a go, change, buy, build, or stop recommendation.Risk closed
Choosing embeddings, a vector database, or a reranker because the technology sounds current rather than because the result list improves. - 04Operate
Release an operated search surface
Can the team see and correct relevance as content, queries, permissions, inventory, and business priorities change?
Release to a bounded audience or traffic share when practical. Review failures by query group, rerun evaluation before material changes, document known compromises, and expand only when the first surface remains measurable and supportable.
Decision produced
A controlled release, index and query analytics, regression checks, feedback and incident paths, configuration history, operating owners, handover, and an evidence-based expansion gate.Risk closed
Shipping a one-time relevance improvement that becomes opaque after the catalogue, content, model, or ranking rules move.
Every project starts at $9,500.
The first paid phase is deliberately bounded. It should close the most expensive open search question before you commit to a platform migration or broad AI rollout.
The 30-minute tune, buy, prove, repair, rebuild, or stop conversation comes first and costs nothing. If ordinary search or a managed feature is enough, that can be the recommendation.
What the first phase can be
- 01
Search relevance audit
Trace representative misses through content, fields, analyzers, query handling, filters, exact matches, retrieval, ranking, interface, analytics, latency, and current operations.
- 02
Query and judgement baseline
Turn real queries into expected results, relevance grades, protected exact behaviours, segment checks, reviewer decisions, and a reusable benchmark for vendors or internal changes.
- 03
Platform and architecture proof
Compare tuning, a managed search feature, lexical, semantic, hybrid, filtered, and reranked retrieval on the same bounded content and query set.
- 04
One production search path
Deliver one search surface with prepared content, required filters and permissions, retrieval, ranking, recovery states, evaluation, analytics, documentation, and handover.
The costliest unanswered search question decides the first phase. Before it begins, you will know what result is included, how it will be judged, what remains outside scope, and what evidence would justify another investment.
Starting investment
$9,500
Minimum project scope. The search surface, users, query groups, content, constraints, evaluation, acceptance criteria, exclusions, ownership, price, and timing are written down before the phase starts.
Price held for the phase
Client-controlled assets
60-day launch warranty
Choose the next path from what the user must receive.
Ranked sources, generated answers, governed knowledge, and a focused technical proof solve different problems and create different operating responsibilities.
- 01
RAG Development
Use this when the product must retrieve permitted evidence and generate a cited answer rather than stop at a ranked source list.
- 02
AI Knowledge Management
Use this when source authority, ownership, permissions, versions, correction, and knowledge lifecycle are the main organisational constraint.
- 03
AI Proof of Concept
Test the riskiest query, content, retrieval, ranking, integration, or cost assumption before committing to a production search surface.
- 04
Enterprise search cost guide
Compare build and buy paths, search-platform scope, access complexity, and the cost drivers behind an enterprise search product.
Common questions
Semantic search finds content that is conceptually related to a query even when the wording is different. It normally represents the query and searchable content as embeddings, retrieves similar candidates, and may rerank them. That can help with synonyms, natural-language descriptions, related concepts, and long queries. It does not automatically understand your commercial priorities, permissions, current inventory, exact identifiers, or what your users consider a good result.
Keyword search is strong when the words matter: a product code, legal clause, person, error message, model number, or exact phrase. Semantic search is useful when meaning matters more than shared wording. Most real products need both behaviours. The question is not which method is more advanced. It is which signal should control each query group, field, filter, and position in the result list.
Hybrid search runs lexical and vector retrieval together and combines their candidates before optional reranking. It can preserve exact matches while recovering relevant results for paraphrases and descriptive queries. A production design may also use filters, field boosts, popularity, freshness, availability, permissions, merchandising rules, or domain-specific signals. The blend should be evaluated on real queries rather than fixed by a generic weight.
Search returns ranked products, documents, records, or passages for the user to inspect and choose. Retrieval-augmented generation, or RAG, retrieves evidence and gives it to a language model to compose an answer. Use search when source choice, comparison, browsing, or direct navigation matters. Use RAG when the product must synthesize an answer from approved evidence. Some experiences offer both, but each needs separate acceptance criteria.
It can if vector similarity replaces rather than complements lexical search and structured filtering. Exact identifiers, names, quoted phrases, categories, prices, dates, stock, geography, and access attributes often need deterministic handling. We identify those protected query and field behaviours, keep the relevant keyword or filter path, and test them as regressions when semantic retrieval or reranking changes.
We start with representative queries and expected results or relevance grades from people who understand the product or content. The evaluation can include exact-match success, useful results near the top, zero-result behaviour, filtered queries, protected access cases, segment performance, latency, and cost. Production signals such as reformulation, result clicks, source opens, conversions, abandonment, and explicit feedback add evidence, but none is a complete definition of relevance on its own.
Yes, when the source identity and application design expose dependable access data. Restricted candidates must be removed before their content is returned or passed to another model. The design must reflect changed and revoked access, tenant boundaries, group membership, and source-system rules. Positive and negative permission cases belong in the same regression set as relevance queries.
Yes. The first phase inspects the current search engine, index schema, query API, interface, content or catalogue, analytics, filters, business rules, traffic, and deployment constraints. The best change may be query tuning, synonym or field work, a managed feature, an additional vector index, reranking, or a replacement service behind the current interface. A visible rebuild is not assumed.
Choose the smallest platform that supports the required content, filters, ranking controls, vector and lexical retrieval, permissions, analytics, latency, deployment, operating model, and commercial constraints. Existing managed products are often the sensible answer. A custom layer is justified when the product needs specialised ingestion, ranking, identity, workflow, interface, or control that the selected platform cannot provide cleanly.
Not necessarily. PostgreSQL with pgvector, Elasticsearch, OpenSearch, Azure AI Search, and other platforms can add vector retrieval without creating a separate source of truth. A dedicated vector database may be justified when the measured workload needs different scale, filtering, tenancy, latency, hosting, or operational behaviour. We compare it against the existing platform using the same queries, content, filters, relevance judgements, update tests, latency, and cost before recommending another data system.
We need access to a representative searchable set, the current index or source structure, and examples of queries that succeed, fail, or return the wrong order. Someone on your team must be able to judge which results are useful, which exact behaviours must be preserved, and which business or access rules affect ranking. Perfect analytics are not required, but invented test queries are a weak substitute for real search behaviour.
Every RaftLabs project starts at $9,500. The first paid phase may be a search audit, relevance evaluation, platform comparison, architecture proof, targeted repair, or one bounded production search path. Timing depends on the current system, content condition, index size, query evidence, filters, permissions, interface, integration, traffic, latency, evaluation depth, and operating requirements. The result, acceptance criteria, exclusions, price, and schedule are written before work begins.
Project-specific code, index and ingestion configuration, query and relevance tests, analytics definitions, operating notes, and prepared data remain under client control. Client-controlled cloud, search, source, identity, and analytics accounts are used where practical and provider terms allow. Third-party platforms and open-source components retain their own licences. The handover names how the index changes, how relevance is tested, and who owns the next release.
Work with us
Bring the search query your users should not have to reword.
In a 30-minute call, we will help identify whether the sensible next move is content repair, search tuning, a managed feature, a relevance proof, a custom build, RAG, or no AI at all.
- One search surface, representative query groups, expected results, exact terms, filters, and business consequence before architecture.
- Lexical, semantic, filtered, hybrid, and reranked retrieval compared only where they solve a measured failure.
- Relevance, access, latency, analytics, cost, ownership, and operating responsibilities written before release.
- A 60-day warranty after release for defects in the agreed application scope.