Why Your Agent Needs Both Keywords and Meaning
A user asks your support agent, "why is my card getting declined," and the agent returns nothing useful, because your knowledge base filed that answer under "payment authorization failures." The exact-match search wanted the word…
A user asks your support agent, "why is my card getting declined," and the agent returns nothing useful, because your knowledge base filed that answer under "payment authorization failures." The exact-match search wanted the word "declined." The semantic search matched three tangentially related billing docs and ranked the right one fourth. Either way, the customer sees a confident, wrong answer. This is the core failure of single-mode retrieval: keyword search misses paraphrase, and pure vector search misses the specific term that actually mattered, like a SKU, an error code, or a product name. This article argues that production agents need both lexical and semantic retrieval blended in a single ranked pass, and shows why running them as two separate systems (a vector database beside a search engine) is the expensive, drift-prone way to get there. In Sanity Context, both live inside one GROQ query against the Content Lake, so there is no second index to keep in sync.

Why does pure semantic search miss exact terms your users type?
Pure semantic search miss exact terms because embeddings compress meaning into a vector, and vectors are good at concepts but blunt about tokens. An error code like "ERR_492", a part number like "A17-Pro", or a version string like "v2.14.1" carries almost no semantic weight. To an embedding model, "ERR_492" and "ERR_493" sit close together in vector space even though one is the answer and the other is noise. The model learned that both are short alphanumeric error strings; it did not learn that a customer typing one of them needs a different document than a customer typing the other.
This is where teams that bet everything on a vector database get burned. Semantic retrieval shines when a user paraphrases ("my subscription won't renew" matching a doc titled "Failed recurring billing"), and falls down exactly when precision matters most. In regulated or technical domains, precision matters most constantly: legal clause numbers, dosage figures, API endpoint names, and internal product codenames are all high-stakes, low-semantic-signal tokens.
The naive fix is to lower the similarity threshold so more candidates come back, but that trades a miss for a flood of loosely related passages, which is worse for an agent because it now has more plausible wrong context to hallucinate from. What you actually want is a lexical signal running alongside the semantic one, so an exact token match can rescue the result that meaning-based ranking buried. Keyword matching is not a legacy technique you outgrow when you adopt embeddings. It is the half of retrieval that embeddings are structurally bad at.
Why does keyword-only search miss the answer when users paraphrase?
Keyword-only search miss paraphrased questions because lexical matching, whether classic BM25 or a simple term match, scores documents on the words they literally contain. A user who asks "why does my payment keep bouncing" will not match a knowledge base article that only ever says "declined transaction" and "authorization failure." The concepts are identical; the vocabulary does not overlap, so the score is zero, and a zero-scored document never reaches the agent.
Human support content and human questions rarely use the same words. Editors write in the product's formal vocabulary; customers write in whatever language their frustration produces at 11pm. That vocabulary gap is precisely what embeddings were built to close. A semantic match sees that "bouncing," "declined," and "authorization failure" occupy the same neighborhood of meaning, and surfaces the right doc even with zero shared tokens.
The standard workaround for keyword-only systems is synonym expansion: hand-maintained lists that teach the index that "bouncing" implies "declined." This works until it doesn't. Every new product, every regional phrasing, and every emerging slang term is a synonym list your team now owns and maintains forever, and the list is always behind. Synonym tables are a manual approximation of what an embedding model does automatically across your entire corpus. The lesson mirrors the previous section exactly: lexical search is strong where semantic search is weak, and weak where semantic search is strong. Neither mode is a superset of the other, which is why picking one is picking which class of question your agent will silently fail.
What is hybrid retrieval, and why does blending beat running two searches?
Hybrid retrieval is a single retrieval pass that scores each candidate document on both lexical overlap and semantic similarity, then blends those signals into one ranked list. It is not two separate searches whose results you union afterward. The distinction matters because a naive union throws away rank information: if a document is the number-three lexical hit and the number-two semantic hit, a blended score can promote it above a document that only ranked well in one mode, which is usually exactly the behavior you want.
Many teams build hybrid retrieval as an assembly project. A vector database (say, Pinecone) handles the semantic half, a search engine or SQL extension handles the lexical half, and application code fetches both result sets, normalizes their incomparable score scales, and merges them with something like reciprocal rank fusion. This works, but you now operate two data systems, keep two indexes consistent with your source content, and own the fusion logic as production code that has to be tuned and tested.
Blending inside one query removes the coordination problem. In Sanity Context, a single GROQ query calls text::semanticSimilarity() for the vector signal and match() for the lexical signal, then combines them with score() and boost() so the ranking is computed in one pass over the Content Lake. There is no second store to provision, no index to reconcile, and no fusion service to keep alive. The weighting between keyword precision and semantic recall becomes a tunable expression in the query itself, not glue code spread across two systems. That is the difference between hybrid retrieval as an architecture you maintain and hybrid retrieval as a query you write.
Why do separate vector and keyword indexes drift out of sync?
Separate vector and keyword indexes drift out of sync because every content change has to be applied to two systems that have no shared notion of truth. An editor updates a pricing doc in the CMS. The lexical index reindexes on one schedule; the embedding pipeline re-embeds on another, often a nightly batch. Between those two events, your agent can match the new wording lexically while retrieving the old meaning semantically, or vice versa, and it will do so with total confidence.
This drift is not a rare edge case; it is the steady state of any stack where embeddings live apart from content. The vector store does not know the source document changed unless something tells it. So teams build change-detection: webhooks, queues, dead-letter handling, and backfill jobs for the inevitable missed events. That plumbing is real engineering that produces no product value, and every gap in it is a window where the agent grounds itself in stale context.
The structural fix is to tie embeddings to the content itself rather than to a downstream copy. Sanity's dataset embeddings are computed against documents in the Content Lake, so when a document is published the embedding updates propagate within minutes, without a separate vector pipeline to babysit. The lexical and semantic signals read from the same content, so "reindex" is not a coordination event across two systems; it is a property of the store. This is one of the five differentiators between a Content Operating System and a legacy stack: legacy CMSes create silos while Sanity provides a shared foundation. When retrieval, embeddings, and editorial all read the same source, drift has nowhere to live.
How do you tune the balance between keyword precision and semantic recall?
You tune the balance between keyword precision and semantic recall by weighting the two signals per use case, because the right blend is not universal. A developer-docs agent answering "how do I call the batch endpoint" leans lexical: the exact endpoint name is the whole answer, and a semantic near-miss is useless. A support agent handling emotional, paraphrased complaints leans semantic: the customer will almost never use your internal term for their problem. One weighting cannot serve both, so the weighting has to be a parameter you control, not a vendor default you accept.
In an assembled stack, changing this balance means editing fusion code, redeploying the retrieval service, and re-running your evaluation set, because the merge logic sits in application code between two systems. That friction pushes teams to set the weighting once and never revisit it, which means the blend is tuned for the average query and wrong for the tails, and the tails are where agents embarrass you.
When the blend is an expression in the query, tuning is cheaper and testable. In Sanity Context, boost() lets you raise the weight of an exact match() on a title or SKU field while score() carries the semantic contribution, all in one GROQ statement you can version and evaluate. Because Studio and Content Releases let editors stage agent behavior the same way they stage a website, a retrieval-tuning change can be reviewed, previewed against real questions, and rolled forward with the same governance as a content change, rather than shipped as an opaque code deploy. Governance here is not bureaucracy; it is the record of why the agent answers the way it does, which is what evaluation and audit both depend on.
How do you evaluate whether your hybrid retrieval is actually working?
You evaluate hybrid retrieval by measuring whether the correct document reaches the agent's context window for a representative set of real questions, not by trusting that adding both signals must be better than one. The honest metric is retrieval recall at the cutoff you actually feed the model: if you pass the top five passages to the agent, the question is whether the answer-bearing passage is in that top five, across a labeled set that includes both exact-term queries and paraphrased ones.
Build the evaluation set from two piles deliberately. The first pile is high-precision queries: error codes, product names, version strings, the questions that punish a semantic-only system. The second is paraphrase queries: emotional, colloquial, vocabulary-mismatched questions that punish a lexical-only system. A healthy hybrid config should beat either single mode on the combined set while losing to neither mode on its home turf. If your blend underperforms pure lexical on the exact-term pile, your semantic weight is too high and is diluting precise matches.
This is where a governed content backend earns its place in evaluation. Because Sanity Context reads retrieval, embeddings, and agent instructions from the same Content Lake, you can point your evaluation harness at the exact content and query the agent will see in production, rather than a snapshot that has already drifted from the live index. Content Source Maps trace a retrieved passage back to its source document, so when a case fails you can see which document should have won and why the ranking buried it. Evaluation without that traceability tells you the agent was wrong; evaluation with it tells you where to fix the content or the weighting. That is the difference between a red test and an actionable one.
Hybrid retrieval: native blend vs assembled stack
| Feature | Sanity | Pinecone | Contentful | pgvector / Neon |
|---|---|---|---|---|
| Blending lexical + semantic | Native: text::semanticSimilarity() and match() combined with score() and boost() in one GROQ query, one ranked pass. | Sparse-dense hybrid is supported, but blending and any lexical richness beyond it is tuned in your application code. | No native blend; you pair the App Framework with an external search or vector service and merge results yourself. | Vector distance plus Postgres full-text search exist, but you write the SQL that fuses and normalizes the two scores. |
| Where embeddings live | Dataset embeddings computed against Content Lake documents, so the semantic signal reads the same source as everything else. | A dedicated vector store separate from your content system; you push vectors to it and keep them current. | External to the CMS; embeddings live in whatever vector service you bolt on via the App Framework. | In your Postgres tables, but populated by an embedding job you build and schedule yourself. |
| Freshness after an edit | Publishing a document propagates embedding updates within minutes, without a separate vector pipeline to maintain. | Depends on your sync: a webhook or batch job you own re-embeds and upserts after each content change. | Two systems to keep consistent; freshness is only as good as the pipeline you build between CMS and vector store. | You own the re-embed trigger; nightly batches are common, which opens a drift window between edits and vectors. |
| Governing agent instructions | Editors stage and review agent behavior in Studio with Content Releases, the same workflow used to stage the website. | Not in scope; prompt and instruction governance sits in your own tooling or code. | Content workflows exist for editorial, but agent instructions and retrieval config live outside them. | No content or governance layer; instruction management is entirely your application's responsibility. |
| Tracing a retrieved passage to source | Content Source Maps link a retrieved passage back to its source document for evaluation and audit. | Metadata you attach to vectors carries provenance; the mapping is only as complete as you make it. | Provenance depends on how you wired IDs across the CMS and the external search layer. | You join back to source rows via your own IDs; traceability is whatever your schema records. |
| Systems to operate | One content backend serves editorial, lexical retrieval, semantic retrieval, and the MCP endpoint agents query. | A vector database plus a separate content source plus fusion code between them. | A CMS plus an external search or vector service plus the integration glue you maintain. | A Postgres database plus an embedding pipeline plus the application layer that fuses lexical and vector scores. |