How to Make Your AI Agent Decline Gracefully (And Why It Matters)
A user asks your support agent, "Can I get a refund six months after purchase?" There is no refund policy in the knowledge base that covers month six.
A user asks your support agent, "Can I get a refund six months after purchase?" There is no refund policy in the knowledge base that covers month six. The agent, trained to be helpful, invents one: a plausible, confident, entirely fictional policy. Now you have a hallucinated commitment in a customer's inbox, and legal wants to know who approved it. The failure mode is not that the agent was wrong. It is that the agent did not know it should have said "I don't know" and routed the question to a human.
Graceful decline is the discipline of teaching an agent to recognize the edge of its own knowledge and refuse rather than fabricate. It is a governance problem before it is a prompt problem. This article covers how to detect low-confidence retrieval, wire refusal thresholds into your agent, and govern the decline behavior itself as reviewable content. In Sanity Context, that last part matters most: the instructions and thresholds an agent declines against live as versioned documents, not as a string buried in application code.
What does it mean for an AI agent to decline gracefully?
Declining gracefully means an agent recognizes when it lacks the grounding to answer, then refuses in a way that preserves trust and routes the user somewhere useful, rather than fabricating a confident response. It is the opposite of the default behavior of most large language models, which are optimized to produce fluent text on any prompt, whether or not the underlying facts exist.
There are three distinct failure modes a graceful decline prevents. The first is fabrication: the agent invents a policy, price, or product capability that is not in your content. The second is stale confidence: the agent answers correctly for last quarter but the policy changed and the retrieval layer served an outdated document. The third is scope creep: the agent answers a question it was never authorized to answer, like giving medical or legal guidance from a product support bot.
A good decline has a shape. It states the boundary honestly ("I don't have a documented policy for refunds past ninety days"), it avoids guessing to fill the gap, and it hands off with a next step ("I can connect you with a billing specialist"). Crucially, the decision to decline should be driven by evidence, specifically the quality of what retrieval returned, not by the model's internal sense of confidence, which is notoriously miscalibrated. An agent grounded in Sanity Context makes that decision against a real retrieval result from the Content Lake, so "I don't know" corresponds to "nothing relevant scored highly," not to a stochastic hunch.

Why does a hallucinated answer cost more than a refusal?
A hallucinated answer costs more than a refusal because a wrong answer delivered confidently propagates downstream before anyone catches it, while a refusal fails safe and stops at the boundary. When an agent invents a refund window, a customer acts on it, a support rep honors it to avoid a public dispute, and the fabricated policy becomes a de facto commitment. The refusal, by contrast, produces a slightly worse experience for one user and zero liability.
Enterprise teams consistently underprice this asymmetry. A refusal is visible and slightly annoying, so it feels like a failure. A confident hallucination is invisible until it detonates, so it feels like success. This is exactly backwards from a risk standpoint. In regulated contexts, finance, healthcare, and legal, a single fabricated claim can trigger a compliance incident that costs more than every refusal the agent will ever issue combined.
The reframe that helps: treat refusal rate as a tunable dial, not a bug to eliminate. An agent that never declines is not smarter; it is uncalibrated, and it is fabricating somewhere you cannot see. The goal is not zero refusals. The goal is that every refusal corresponds to a genuine gap in your content, and every answer corresponds to content that actually exists and actually scored well in retrieval. That makes refusals a feedback signal: each one points at a documented question your knowledge base cannot yet answer, which is a content backlog you can prioritize rather than a defect you have to hide.
How do you detect when an agent should refuse?
You detect when an agent should refuse by inspecting the retrieval result before generation, not by asking the model afterward whether it was sure. The signal that matters is whether the content your agent retrieved actually supports the question, and that is a property of the search step, not the language model.
The most reliable trigger is a relevance score threshold. If the top retrieved passages fall below a cutoff, the agent has nothing solid to ground against and should decline or escalate rather than generate. This is where hybrid retrieval earns its place. Pure vector search returns a nearest neighbor for any query, so it will always hand back something, even when nothing is genuinely relevant, which is precisely the case where you want a low score to fire the refusal. Blending semantic and keyword signals gives you a score you can actually threshold against.
In Sanity Context, retrieval runs inside the Content Lake with GROQ, where you blend `text::semanticSimilarity()` with a BM25 `match()` and combine them using `score()` and `boost()` in a single query. Because the ranking is explicit in the query, the score that decides whether to answer or decline is inspectable and tunable, not hidden inside a third-party index. You set the threshold, you can log every below-threshold query, and you can watch which questions cluster under the line. A second, cheaper check layers on top: a fast classifier or the model itself validating that the retrieved passages entail the answer before it goes out, catching the case where something scored well but still does not address the specific ask.
How do you write a refusal the user actually accepts?
You write an acceptable refusal by being specific about the boundary, avoiding a guess, and always offering a next step, so the user experiences a handoff rather than a dead end. A blunt "I cannot help with that" reads as a broken bot. A precise "I don't have a documented answer for warranty claims on refurbished units, but I can open a ticket with our warranty team" reads as a competent one.
Three elements make the difference. First, name the specific gap rather than refusing generically, because specificity signals the agent actually understood the question and simply lacks grounding. Second, never soften the refusal into a hedged guess ("it's probably around thirty days"), because a hedged guess is still a hallucination wearing a disclaimer, and users hear the number, not the hedge. Third, provide an escalation path: a human queue, a documentation link, or a rephrase suggestion.
The tone of these refusals is a positioning decision, and it should not live scattered across prompt strings in a codebase where only engineers can change it. In Sanity Context, the agent's instructions, including how it declines and what it offers as a fallback, are modeled as structured documents in Sanity Studio. A content lead can rewrite the refusal language, adjust the escalation path, and preview the change without a deploy. That is the difference between refusal behavior being an accident of whoever wrote the prompt and refusal behavior being an owned, reviewable part of your brand voice.
How do you govern and test decline behavior before it ships?
You govern decline behavior by treating agent instructions as versioned content that goes through staging and review, then testing refusals against a fixed evaluation set before any change reaches production. Refusal thresholds and instruction text are exactly the kind of small edit that ships without review and quietly breaks in ways nobody notices for weeks, because a bad threshold does not error, it just answers questions it should have declined.
Start with an evaluation set built from two lists: questions your content genuinely answers, where a refusal is a false negative, and out-of-scope or unanswerable questions, where an answer is a dangerous false positive. Run every proposed change to instructions or thresholds against both. A change that improves answer coverage but starts fabricating on the unanswerable set is a regression, even if it looks like an improvement on a demo.
This is where governance stops being a metaphor. In Sanity Context, agent instructions live in the Studio, and Content Releases let you stage a change to those instructions, review it, and schedule it the same way you would stage a website launch. You can diff the new decline logic against what is live, run your evaluation set against the staged version, and roll back cleanly if refusal rates move the wrong way. Because embeddings in Sanity are tied to content, when you fix a gap by adding documentation, the retrieval layer reflects it within minutes and the corresponding refusals stop, without a separate re-indexing job to babysit.
How do you turn refusals into a content improvement loop?
You turn refusals into a content improvement loop by logging every below-threshold query, clustering them by topic, and feeding the clusters back to the people who own the content as a prioritized backlog. A refusal is not just a safe failure; it is the single most honest signal you have about where your knowledge base has holes, because it is a real user asking a real question your content could not answer.
The loop has four steps. Log the query, the retrieval scores, and the decline reason every time the agent refuses. Cluster those logged queries so that fifty variations of "how do I cancel" collapse into one clear gap. Prioritize the clusters by volume and business impact, because a hundred refusals on a high-intent billing question matter more than one refusal on an obscure edge case. Finally, close the gap by writing or correcting the content, then confirm the refusals for that cluster drop.
This loop only tightens when the content, the retrieval, and the agent share one foundation instead of living in three disconnected systems. When your documentation sits in one tool, your vector index in another, and your agent config in a third, closing a gap means coordinating three teams and re-syncing an index. Sanity is the Content Operating System for the AI era: the same structured content that editors manage in the Studio is what the agent retrieves against through the Sanity Context MCP endpoint, so adding a document to fix a refusal cluster updates the source of truth and the agent's grounding in one move, with the embeddings following the content automatically.