AI Search and Semantic Search Development

AI search that understands meaning without losing exact matches.

We improve product, content, and internal search with the right mix of keyword, semantic, filtered, and reranked retrieval. Relevance is judged on the queries your users actually type, not on a polished demo.

See our work

Bring a query users keep reformulating and the result they should have found. Leave with a tune, buy, prove, rebuild, or stop recommendation.

The brief

Start with what is not working.

Good software decisions begin with the constraint, not a list of features or a preferred technology.

01

Users know the product, document, or record exists, but search returns nothing useful unless they remember the exact wording.

02

Adding semantic search improved some natural-language queries but pushed down product codes, names, filters, or other exact matches.

Plain answer

Semantic search retrieves results by meaning, while keyword search retrieves text that matches the words in a query. A dependable AI search system often combines both as hybrid search, then applies filters, business rules, permissions, and reranking. The right design depends on what users search, which exact terms must never be weakened, what a good result looks like, and whether the user needs ranked sources or a generated answer. Every RaftLabs project starts at $9,500 with a bounded first phase and written acceptance criteria.

What to remember

  • Use semantic search when relevant content exists but users describe it differently from the words stored in the catalogue, document, or record.
  • Keep keyword retrieval and structured filters when exact identifiers, names, categories, prices, dates, permissions, or business priorities must remain dependable.
  • Evaluate search with representative queries and expected results before choosing embeddings, a vector database, reranking, or a complete platform rebuild.

Start with what the user should receive.

A better search box, a ranked source list, and a generated answer are different products. Choose the smallest result that helps the user finish the job.

ScenarioRecommendationReason
Known item, exact identifier, or tightly filtered recordTune keyword search and filtersProtect product codes, names, quoted phrases, categories, dates, prices, availability, and other exact constraints.
Users describe the right item in different wordsAdd semantic retrievalRecover relevant products, documents, or passages when the query and stored text do not share the same vocabulary.
The search surface has exact and descriptive queriesUse hybrid searchCombine lexical precision and semantic recall, then rerank or apply business rules where the evidence supports it.
The user needs one answer supported by private sourcesUse RAGRetrieve permitted evidence, then generate a cited response with clarification, refusal, and separate answer evaluation.
The content is missing, stale, duplicated, or poorly structuredRepair the source firstA more sophisticated retrieval layer cannot return reliable information that the source does not contain or govern.

Fit

Improve search when the right result exists but does not reliably surface.

The engagement is a fit when a search failure can be observed, judged, and connected to a real user or business consequence.

A fit
01

Users reformulate queries, abandon the search, browse manually, ask support, or use external search to find material that already exists.

02

The search surface mixes descriptive intent with exact names, codes, filters, permissions, freshness, availability, or business ranking rules.

03

Your team can provide representative content and someone who can judge why one result should rank above another.

Not a fit
01

The source content is not available, current, owned, or detailed enough to answer the user's need.

02

The real job is to summarize several sources, make a recommendation, or take an action rather than return ranked results.

03

The goal is to add a vector database or AI label without a query set, relevance owner, or measurable search problem.

If the current engine can be fixed with better fields, metadata, filters, synonyms, or ranking rules, the first recommendation can stop there.

The product code was exact. The result was only similar.

A buyer pastes a model number into a catalogue. A pure semantic search treats it like ordinary text and returns products with similar descriptions. The exact product appears lower, or not at all. On another query, a buyer writes what they need in plain language and keyword search returns nothing because the catalogue uses different words.

Neither query proves that keyword search is old or semantic search is better. They show two different relevance jobs on the same surface. Exact identifiers need lexical protection. Descriptive intent needs meaning. Price, availability, category, tenant, permission, and merchandising constraints may need structured rules before either result reaches the page.

That is why the useful starting artefact is a query and judgement set, not a vendor shortlist. It records what users type, what should appear, what must not appear, which constraints apply, and why the order matters. Architecture comes after that evidence.

Illustrative search relevance worksheet comparing test queries with ranked results and reviewer judgements

An illustrative relevance worksheet. A real evaluation uses your query logs, searchable content, expected results, protected constraints, and reviewer decisions.

Keyword, semantic, hybrid, or RAG?

DecisionKeyword searchSemantic searchHybrid searchRAG
What it returnsRanked results using lexical signalsRanked results using meaning and similarityRanked results using lexical and semantic candidatesA generated answer based on retrieved evidence
Strongest useCodes, names, exact phrases, filters, and predictable field boostsParaphrases, synonyms, concepts, and descriptive queriesMixed query sets where exactness and meaning both matterQuestions that need synthesis across approved sources
Main riskRelevant content is missed when the wording differsPlausible but wrong similarities weaken exact or business-critical resultsFusion and reranking hide trade-offs if the query groups are not evaluatedThe system retrieves weak evidence or writes beyond what the evidence supports
Buyer testCan users name or filter the item precisely?Does the right result exist under different wording?Do both behaviours occur on the same search surface?Does the user need a cited explanation rather than a source list?

Current platform documentation reflects the same separation. Elastic describes hybrid search as combining lexical precision with vector similarity, while Microsoft documents semantic ranking as a later ranking step over an initial keyword or hybrid result set. Those are useful building blocks, not a reason to choose a platform before the search job is understood.

Production scope

The hard work is deciding what deserves to rank.

A first release can be narrow. It still needs the content, retrieval, ranking, interface, measurement, and operating decisions that keep search understandable.

  • 01

    Search behaviour and relevance model

    Group real queries by job: known item, category, descriptive, comparative, filtered, misspelled, navigational, informational, and unsupported. Record expected results, relevance grades, protected exact behaviours, business rules, segments, and the consequence of a poor result.
  • 02

    Content, catalogue, and index design

    Map the fields, identifiers, variants, attributes, passages, metadata, language, dates, versions, ownership, freshness, and access data needed for retrieval. Expose missing, duplicated, stale, unparseable, or weakly described items before blaming the ranking model. If a separate vector index is justified, define stable source IDs, updates, deletions, re-embedding, reconciliation, versioning, rollback, and the team that will operate it.
  • 03

    Lexical, semantic, filtered, and hybrid retrieval

    Tune fields, analyzers, synonyms, typo handling, boosts, exact matches, filters, vector retrieval, fusion, and candidate depth around the query groups. Keep deterministic constraints outside similarity scoring when the rule must not become approximate.
  • 04

    Reranking and business priorities

    Rerank only a useful candidate set. Combine textual relevance with permitted signals such as availability, freshness, location, account context, quality, popularity, or merchandising without quietly turning relevance into whichever item the business wants to promote.
  • 05

    Search interface and recovery states

    Design query suggestions, filters, facets, snippets, highlights, result reasons, source links, spelling recovery, empty states, and reformulation paths around how people inspect results. The interface should help a user recover when confidence is low, not hide a weak result behind fluent copy.
  • 06

    Evaluation, analytics, and search operations

    Version the query set, expected results, configuration, and reviewer decisions. Monitor index health, zero-result and low-result queries, reformulation, result opens, conversion or completion signals, access failures, latency, cost, regressions, and the feedback path used to correct important misses.

A lower zero-result rate can still hide worse search.

Suppose the new system always returns something. The zero-result rate improves, but exact model numbers now surface similar products, unavailable stock ranks first, and a policy search favors an older page with more matching language. The dashboard is greener while the search experience is less dependable.

Search quality needs several views. Offline evaluation asks whether expected items appear high enough for representative queries. Segment checks show whether codes, names, descriptive phrases, filters, languages, and permission groups behave differently. Production behaviour adds reformulations, result opens, conversion or task completion, abandonment, and explicit feedback.

These signals still need judgement. A click may mean the result was useful, or that the user had no better option. A promoted item may convert well while answering the query poorly. We keep the ranking goal, commercial rule, and user behaviour visible as separate evidence so the team knows what it is optimizing.

What a credible relevance baseline contains

Keep the artefacts that let your team compare a change, challenge a result, and move platforms without rediscovering the search problem.

  • 01

    Representative query groups

    Real successful and failed queries, exact identifiers, paraphrases, misspellings, filters, sensitive cases, unsupported requests, and the segments that matter commercially.
  • 02

    Expected results and protected behaviour

    Relevance grades, top-result expectations, acceptable alternatives, exclusions, exact-match guarantees, filters, access constraints, and the reason one item should outrank another.
  • 03

    Versioned configuration and results

    The content snapshot, index schema, retrieval settings, embedding and reranking choices, business rules, evaluation result, reviewer decision, latency, and cost attached to a meaningful change.
  • 04

    Production feedback with context

    Zero-result and low-result searches, reformulations, result opens, conversions or task completions, abandonment, manual corrections, and enough context to investigate without retaining unnecessary sensitive query data.
  • 05

    Owners and correction path

    Named owners for content, search relevance, permissions, commercial ranking rules, releases, incidents, feedback, and the decision to expand, simplify, buy, rebuild, or stop.

GrantHub demonstrates a real search-product constraint: users needed to narrow inconsistent records by country, industry, business type, funding stage, eligibility, and deadline while the underlying data kept changing. It is evidence of search, filtering, data modelling, freshness, and product delivery. It is not presented as proof of a semantic model or as a promise that another search surface will achieve the same result.

How it works

Close one relevance risk before opening the next.

Each phase ends with evidence and a decision. The next query group, source, or search surface is added only after the current one is understandable.

  1. 01
    Understand

    Observe the search job

    Who is searching, what are they trying to find, and what happens when the result is wrong or missing?

    Review query logs where they exist, watch people search, inspect support language, and reproduce successful and failed journeys. Include the weird model number, vague description, misspelling, filtered query, stale item, and permission edge case that a polished demo avoids.

    Decision produced

    One search surface, user groups, query families, current path, expected results, exact constraints, filters, access, baseline friction, and accountable relevance owner.

    Risk closed

    Replacing a search engine before separating a content problem, a query problem, a ranking problem, and an interface problem.
  2. 02
    Measure

    Build the relevance baseline

    What should appear, in what order, and which behaviours must not regress?

    Ask people who understand the domain to grade results and explain disagreements. Separate expected user relevance from commercial boosts and compliance rules. If nobody can judge the order, the project needs a decision owner before it needs a new model.

    Decision produced

    A versioned query and judgement set, content snapshot, protected exact cases, relevance grades, segment view, current scores, latency, and evidence gaps.

    Risk closed

    Optimizing one headline metric while high-value queries, exact matches, small segments, or access boundaries quietly become worse.
  3. 03
    Improve

    Prove the smallest useful change

    Can tuning, a managed feature, or a focused retrieval change solve the measured misses without creating a larger operating burden?

    Start with the current engine and the simplest credible change. Compare lexical, semantic, hybrid, filtered, and reranked paths only where the evaluation shows a trade-off. Preserve exact and structured behaviour before adding approximation.

    Decision produced

    A compared approach, working proof or repair, evaluation result by query group, platform and cost trade-offs, interface behaviour, and a go, change, buy, build, or stop recommendation.

    Risk closed

    Choosing embeddings, a vector database, or a reranker because the technology sounds current rather than because the result list improves.
  4. 04
    Operate

    Release an operated search surface

    Can the team see and correct relevance as content, queries, permissions, inventory, and business priorities change?

    Release to a bounded audience or traffic share when practical. Review failures by query group, rerun evaluation before material changes, document known compromises, and expand only when the first surface remains measurable and supportable.

    Decision produced

    A controlled release, index and query analytics, regression checks, feedback and incident paths, configuration history, operating owners, handover, and an evidence-based expansion gate.

    Risk closed

    Shipping a one-time relevance improvement that becomes opaque after the catalogue, content, model, or ranking rules move.

Every project starts at $9,500.

The first paid phase is deliberately bounded. It should close the most expensive open search question before you commit to a platform migration or broad AI rollout.

The 30-minute tune, buy, prove, repair, rebuild, or stop conversation comes first and costs nothing. If ordinary search or a managed feature is enough, that can be the recommendation.

What the first phase can be

  1. 01

    Search relevance audit

    Trace representative misses through content, fields, analyzers, query handling, filters, exact matches, retrieval, ranking, interface, analytics, latency, and current operations.

  2. 02

    Query and judgement baseline

    Turn real queries into expected results, relevance grades, protected exact behaviours, segment checks, reviewer decisions, and a reusable benchmark for vendors or internal changes.

  3. 03

    Platform and architecture proof

    Compare tuning, a managed search feature, lexical, semantic, hybrid, filtered, and reranked retrieval on the same bounded content and query set.

  4. 04

    One production search path

    Deliver one search surface with prepared content, required filters and permissions, retrieval, ranking, recovery states, evaluation, analytics, documentation, and handover.

The costliest unanswered search question decides the first phase. Before it begins, you will know what result is included, how it will be judged, what remains outside scope, and what evidence would justify another investment.

Starting investment

$9,500

Minimum project scope. The search surface, users, query groups, content, constraints, evaluation, acceptance criteria, exclusions, ownership, price, and timing are written down before the phase starts.

Price held for the phase

The agreed phase price does not move unless you approve a material change in scope.

Client-controlled assets

Project-specific code, prepared data, query and relevance tests, search accounts, analytics, and operating notes remain under client control where licences and security allow.

60-day launch warranty

Defects in the agreed application scope, release support, and small interface corrections are covered for 60 days after launch.

Common questions

Semantic search finds content that is conceptually related to a query even when the wording is different. It normally represents the query and searchable content as embeddings, retrieves similar candidates, and may rerank them. That can help with synonyms, natural-language descriptions, related concepts, and long queries. It does not automatically understand your commercial priorities, permissions, current inventory, exact identifiers, or what your users consider a good result.

Keyword search is strong when the words matter: a product code, legal clause, person, error message, model number, or exact phrase. Semantic search is useful when meaning matters more than shared wording. Most real products need both behaviours. The question is not which method is more advanced. It is which signal should control each query group, field, filter, and position in the result list.

Hybrid search runs lexical and vector retrieval together and combines their candidates before optional reranking. It can preserve exact matches while recovering relevant results for paraphrases and descriptive queries. A production design may also use filters, field boosts, popularity, freshness, availability, permissions, merchandising rules, or domain-specific signals. The blend should be evaluated on real queries rather than fixed by a generic weight.

Search returns ranked products, documents, records, or passages for the user to inspect and choose. Retrieval-augmented generation, or RAG, retrieves evidence and gives it to a language model to compose an answer. Use search when source choice, comparison, browsing, or direct navigation matters. Use RAG when the product must synthesize an answer from approved evidence. Some experiences offer both, but each needs separate acceptance criteria.

It can if vector similarity replaces rather than complements lexical search and structured filtering. Exact identifiers, names, quoted phrases, categories, prices, dates, stock, geography, and access attributes often need deterministic handling. We identify those protected query and field behaviours, keep the relevant keyword or filter path, and test them as regressions when semantic retrieval or reranking changes.

We start with representative queries and expected results or relevance grades from people who understand the product or content. The evaluation can include exact-match success, useful results near the top, zero-result behaviour, filtered queries, protected access cases, segment performance, latency, and cost. Production signals such as reformulation, result clicks, source opens, conversions, abandonment, and explicit feedback add evidence, but none is a complete definition of relevance on its own.

Yes, when the source identity and application design expose dependable access data. Restricted candidates must be removed before their content is returned or passed to another model. The design must reflect changed and revoked access, tenant boundaries, group membership, and source-system rules. Positive and negative permission cases belong in the same regression set as relevance queries.

Yes. The first phase inspects the current search engine, index schema, query API, interface, content or catalogue, analytics, filters, business rules, traffic, and deployment constraints. The best change may be query tuning, synonym or field work, a managed feature, an additional vector index, reranking, or a replacement service behind the current interface. A visible rebuild is not assumed.

Choose the smallest platform that supports the required content, filters, ranking controls, vector and lexical retrieval, permissions, analytics, latency, deployment, operating model, and commercial constraints. Existing managed products are often the sensible answer. A custom layer is justified when the product needs specialised ingestion, ranking, identity, workflow, interface, or control that the selected platform cannot provide cleanly.

Not necessarily. PostgreSQL with pgvector, Elasticsearch, OpenSearch, Azure AI Search, and other platforms can add vector retrieval without creating a separate source of truth. A dedicated vector database may be justified when the measured workload needs different scale, filtering, tenancy, latency, hosting, or operational behaviour. We compare it against the existing platform using the same queries, content, filters, relevance judgements, update tests, latency, and cost before recommending another data system.

We need access to a representative searchable set, the current index or source structure, and examples of queries that succeed, fail, or return the wrong order. Someone on your team must be able to judge which results are useful, which exact behaviours must be preserved, and which business or access rules affect ranking. Perfect analytics are not required, but invented test queries are a weak substitute for real search behaviour.

Every RaftLabs project starts at $9,500. The first paid phase may be a search audit, relevance evaluation, platform comparison, architecture proof, targeted repair, or one bounded production search path. Timing depends on the current system, content condition, index size, query evidence, filters, permissions, interface, integration, traffic, latency, evaluation depth, and operating requirements. The result, acceptance criteria, exclusions, price, and schedule are written before work begins.

Project-specific code, index and ingestion configuration, query and relevance tests, analytics definitions, operating notes, and prepared data remain under client control. Client-controlled cloud, search, source, identity, and analytics accounts are used where practical and provider terms allow. Third-party platforms and open-source components retain their own licences. The handover names how the index changes, how relevance is tested, and who owns the next release.

Work with us

Bring the search query your users should not have to reword.

In a 30-minute call, we will help identify whether the sensible next move is content repair, search tuning, a managed feature, a relevance proof, a custom build, RAG, or no AI at all.

  • One search surface, representative query groups, expected results, exact terms, filters, and business consequence before architecture.
  • Lexical, semantic, filtered, hybrid, and reranked retrieval compared only where they solve a measured failure.
  • Relevance, access, latency, analytics, cost, ownership, and operating responsibilities written before release.
  • A 60-day warranty after release for defects in the agreed application scope.