The Loop That Doesn't Know When to Stop: Adaptive and Agentic RAG in Production

RAG Architecture Series · Part 5

The Loop That Doesn't Know When to Stop

Adaptive RAG makes one decision before retrieval. Agentic RAG makes a decision after every step, and keeps making them until something tells it to stop. The loop was never the hard part. The stopping condition is.

Reviewed August 2026 / Developer track / 15 min read

Before you build any of this: most teams who tell me they need an agent need a router.

A router is a classifier and three code paths. You can reason about it, you can unit test it, you can put a number on what it costs before you ship it. An agent is a control flow you handed to a language model, which means your p99 latency and your monthly bill are now both functions of how confused that model gets on inputs you haven't seen yet.

That's not an argument against agentic RAG. Part of this article is about the cases where the loop genuinely earns its keep, and they're real. It's an argument about ordering. Route first. Measure what's left failing. Then decide whether the thing that's still broken is broken because your pipeline only got one shot at it, because that's the only failure the loop actually fixes.

Everything below assumes you already have hybrid retrieval and a reranker in place from Part 2, and some flavor of retrieval grading from Part 3. If you don't, this article is premature. Adding a loop on top of retrieval that was never good enough gets you the same bad chunks, several times, at four times the cost.

Terminology

Adaptive and agentic are not synonyms

Adaptive RAG classifies the query once, up front, and picks a path: answer directly with no retrieval, do a single-shot retrieve and generate, or escalate to something iterative.3 One decision, made before any work happens. After that the control flow is fixed, even though what happens inside it still varies: retrieval returns what it returns, generation is its own kind of nondeterministic. What you get is a bounded shape. You know which path a query took and roughly what that path costs, before it runs.

Agentic RAG puts the model in charge of control flow. It decides whether to retrieve, what to retrieve, whether what came back was any good, whether to reformulate and go again, whether to call a different tool entirely. Every one of those is a decision made at runtime with no fixed bound unless you put one there yourself.

The distinction matters because the vocabulary has collapsed. "We went agentic" usually means somebody wrapped a while loop around hybrid search, and nobody costed it. Meanwhile the actual adaptive part, the routing, is the piece that pays for itself fastest and gets skipped, because it isn't interesting to talk about at standup.

There's a connection back to Part 3 worth naming. Corrective RAG's grader is a judgment call made after retrieval, on the chunks.4 A router is the same species of judgment moved upstream, made on the query alone. That makes it dramatically cheaper, since you're scoring twenty tokens instead of twenty chunks, and dramatically less informed, since you're guessing what retrieval will find before you've looked. Those two properties are the whole tradeoff. Route on what's obvious from the query. Grade on what actually came back. Teams that try to do all the deciding in one place end up with a router that's slow or a grader that fires on everything.

Aside · tool opinion

Every major framework ships a ReAct-style loop as the starter template. The quickstart, the tutorial, the thing you get from create_agent(). Which means the most expensive architecture in the entire design space is the one you arrive at by not making a decision.

To be fair, the frameworks do put a crude stop in there: LangGraph raises after 25 steps by default,1 the OpenAI Agents SDK stops at 10 turns and makes you explicitly ask for unlimited.2 Those are loop-safety catches, not budgets. Nothing in either default knows what a query costs you, and 25 steps of retrieval plus reranking plus generation is not a number anyone chose with your workload in mind. So the default isn't uncapped, it's capped on the wrong axis, which in practice reads the same on the invoice.

Implementation

One decision, made cheaply, on the query alone

Three ways to do it, in increasing order of cost and decreasing order of how much I like them.

Embedding classifier

Embed the query, run a small trained classifier or nearest-centroid match over labeled examples, pick the path. Single digit milliseconds after the embedding call you were probably making anyway. Needs labeled queries, which is the actual work.

Small dedicated model

A fine-tuned classifier or a small instruct model with a tight prompt. Tens of milliseconds. Better on phrasing it hasn't literally seen.

LLM router

Ask the big model which path to take. Easy to build, easy to change, and it adds a full model round trip to the front of every single query, including the 60 percent that were trivial. This is the one everybody starts with and the one nobody profiles.

That last point is the one I'd underline. The router runs on 100 percent of traffic. Whatever latency it adds, it adds to your floor, and it adds it to the queries that were supposed to be the fast ones. A 400ms LLM router in front of a 600ms single-shot path didn't optimize anything, it just moved where the time goes.

PATHS = ("direct", "single_shot", "iterative")

ESCALATE = {"direct":      "single_shot",
            "single_shot": "iterative",
            "iterative":   "iterative"}   # already at the top

def route(query: str, threshold: float = 0.62) -> str:
    scores = classifier.predict_proba(embed(query))   # dict[path, float]
    best, conf = max(scores.items(), key=lambda kv: kv[1])

    if conf < threshold:
        return ESCALATE[best]         # unsure: one level up, never down

    if best == "direct" and conf < 0.85:
        return "single_shot"          # asymmetric: skipping retrieval is
                                      # the expensive mistake to get wrong
    return best

Routing on the query alone. The tie-break is the whole design.

When the router is unsure, route up. Escalate relative to the predicted path, not to the middle: a low-confidence iterative prediction is not evidence that the query is easy.

The asymmetry is the reason. Routing a hard query to the cheap path produces a confident, fluent, wrong answer, and nothing in your logs marks it as a failure. Routing an easy query to the expensive path produces a correct answer that cost too much and took too long, and shows up as a line item you can go find later. One of those failure modes is visible. The other one is the insurer bot from Part 1, citing an exclusion that doesn't exist, with nothing in the trace to say why.

Same logic applies to the direct path specifically. "Answer with no retrieval" is the highest-variance decision in the whole system, because it's the only one where a wrong route means the model is working purely from parametric memory on a question your users assumed was grounded in your documents. I hold that path to a higher confidence bar than the others. It's not elegant. It works.

FIG.01 — ADAPTIVE ROUTING — ONE DECISION, THREE PATHS BUDGET · TOKENS · WALL CLOCK · PASSES QUERY ROUTER classifier decides on the query alone DIRECT ANSWER no retrieval · parametric memory only 0.2x HYBRID + RERANK GRADER ANSWER grades chunks, not the query 1x HYBRID + RERANK GRADER ANSWER REFORMULATE, GO AGAIN 2x to 12x
Fig. 01 The router decides once, on the query. The grader decides again, later, with better information. Only the third path can spend without a ceiling, which is why the budget is drawn around it and not around the system.

Where the loop actually earns it

There are queries a single retrieval pass cannot answer, no matter how good your reranker is, because the information needed to write the second query is only available after you've read the results of the first one.

The clean version of this is a join key you don't have yet. "What are the termination terms for the vendor that took over our old logistics contract" is that shape: the vendor's name isn't in the query, it's inside the novation agreement, and until you've retrieved and read that document you cannot write the query that finds the contract. Same for anything routed through an alias, a successor entity, a ticket that references another ticket. Step two is genuinely not knowable at step one.

Worth separating from a case that looks similar and isn't. "Which of our vendor contracts have termination clauses that conflict with the new data retention policy" reads like multi-hop, but both halves are named in the query and you can go get them independently, in parallel if you want. What it needs is broad or exhaustive retrieval plus a comparison step, not a discovered second query. A loop is one way to do that. So is a wider retrieval budget and a map-reduce pass, and that's usually cheaper and easier to reason about. If your "we need an agent" case is really this shape, check the cheaper thing first.

That's the case for the loop. It's real, I've seen it pay off, and I'd build it. What I'd note is how narrow it is compared to how often the loop gets deployed. In the systems I've looked at, this shape is a minority of traffic, frequently a small one. Which is exactly why it belongs behind a router instead of in front of everything.

Budgets

Your loop needs three caps, and none of them is max_iterations

Here is where the money goes.

A retry cap is not a budget. max_iterations=8 tells you how many times the loop may go around, and tells you nothing about what those eight iterations cost, because iteration four might retrieve 40 chunks and iteration five might call three tools. Two queries that both hit the cap can differ by an order of magnitude in spend.

What you want is three caps, checked every pass:

  1. Token budget per query. Cumulative, prompt plus completion, across every model call in the loop including the grader. This bounds model token consumption, which is usually the biggest line, but it is not the same as bounding spend. Tokens don't price a paid search API, a vector store that charges per query, an OCR pass, or the fact that your grader and your generator are different models at different rates. For a real ceiling, count weighted model cost plus paid tool calls in one budget, and give the expensive tools their own per-query call limit on top.
  2. Wall-clock budget per query. Because the loop can be cheap and still blow your p99 through tool latency you don't control.
  3. Iteration cap. Still useful, as a backstop against pathological cycles where each pass is individually tiny.

When a budget is exhausted, the loop must terminate into a degraded answer, not a retry. Answer from what you have and say what you couldn't verify, or hand off to a human. What it must never do is treat exhaustion as a reason to try once more with different phrasing, which is the shape of every runaway I've seen.

The arithmetic is worth doing once, because it explains why the dashboard looks fine right up until it doesn't. Take a loop with continuation probability p, per-pass cost T, and depth cap d:

E[tokens] = T_route + T · Σi=1..d pi-1

with p = 0.4, T = 4,000, d = 12:
  Σ pi-1 ≈ 1.667
  E[tokens] ≈ 6,700    (excluding the router call)
  worst case = 48,000  (about 7x the mean, and it is a reachable path)

The mean is a liar here. With p at 0.4 most queries stop after one or two passes and your average cost per query looks great. The bill is not paid by the average. It's paid by the queries where p is not 0.4, which is to say the ambiguous ones, the ones where retrieval keeps returning plausible-but-wrong chunks and the grader keeps saying no, again, and the loop obligingly walks all the way to the cap. Those queries cluster. They are not randomly distributed across your traffic. They're all the same kind of query.

FIG.02 — LOOP DEPTH — WHERE THE SPEND ACTUALLY LIVES PASS 1 PASS 2 PASS 3 PASS 4 PASS 5+ 16% OF QUERIES SHARE OF QUERIES 60% 24% 10% SHARE OF TOKENS 36% 28% 17% 9% PASS 3+ · 36% OF TOKENS p = 0.4 · 4,000 tokens per pass · cap 12 · geometric model, not a benchmark
Fig. 02 Same distribution, two views. The deep passes are a sixth of your traffic and better than a third of your token spend, and in practice they are not spread evenly: they are one segment of queries, arriving together.

Which is the composite cost anecdote from Part 1, and I'll keep the honest framing on it: this is a composite of a few documented cases and public postmortems, not one incident I stood next to. Uncapped retries, up to twelve passes, roughly $40K in a month, concentrated almost entirely in ambiguous part-number lookups on a returns bot. The point isn't the number. The point is that the queries that trigger the deep loop are correlated, so the tail isn't a tail, it's a segment, and it will find you the moment that segment shows up in volume.

Aside · still unresolved

Still bothering me: I have a suspicion that most Adaptive RAG papers benchmark against the same three or four public datasets, and that routing accuracy on HotpotQA tells you approximately nothing about routing accuracy on a messy internal knowledge base where half the queries are ticket numbers and the other half are three words long.

I said this in Part 1 and I still haven't gone and checked it properly. I keep meaning to pull the actual eval sections and count. If somebody's already done that, I'd like to see it.

PathAdded p50p95CostCharacteristic failure
System prompt only 000.2x Answers from parametric memory, sounds grounded
Single-shot hybrid + rerank baselinebaseline1x Misses when the query needs a second hop
Corrective loop, capped at 2 +1 pass on ~30%+2 passes1.4x to 1.8x Grader too strict, retries on good chunks
Full agentic, uncapped smallworkload-dependent,
no hard ceiling
2x to 12x+ Correlated deep loops on one query segment

Numbers are shape, not benchmark. Measure yours. But note the p95 column, because that's the one the demo never shows you.

Opinion

Agents demo well because demos are five queries long

Five queries is not enough to see variance. It is barely enough to see a mean. Every agentic RAG demo I've watched, including ones I've given, runs a handful of hand-picked questions where the loop does something visibly clever, converges in two or three passes, and everybody in the room watches the reasoning trace scroll by and finds it delightful. It is delightful. It's also a sample size that cannot possibly surface the failure mode that will actually hurt you, which lives at p99 and is invisible until you're doing thousands of queries a day.

So the demo selects for the architecture with the worst tail, and then the tail arrives in week three, and everyone is confused.

If you're evaluating an agentic setup, mine your logs for the fifty ugliest real queries you have, the ambiguous ones, the ones that are just an ID, the ones with a typo in the only word that mattered, and run those. Not the ones you'd put in a deck. The loop's behavior on the clean queries is not information. You already knew it could do those.

I don't think this is a case of anyone being dishonest. It's just that the thing that makes agentic RAG compelling to watch and the thing that makes it dangerous to operate are the same property: it does something different every time.

Correction

I validated a router against the wrong queries

This one is recent enough that it still annoys me.

I built a router for a client, three paths, the whole structure above. To validate it, I used the query set the team already had, which was the set they'd been using for manual testing since the project started. Roughly a hundred queries. Well-formed, complete sentences, unambiguous. The router hit something like 94 percent on them, which felt great, and we shipped it.

The problem is that their real users don't type like that. Real traffic was fragments. Bare IDs. Two words and a question mark. Copy-pasted error strings with the timestamp still attached. And on that distribution the router was systematically routing down, because short unadorned queries look like simple queries to a classifier trained on tidy examples, so a bare part number that needed a full retrieval pass went to the direct path and got answered from the model's own memory.

Which failed silently. That's the part I want to sit with. Nothing errored. Latency looked better. Cost per query went down. Every dashboard we had said the router was working, and it had been quietly degrading answer quality for a few weeks before somebody in support noticed the bot confidently making up a spec.

The lesson I actually took wasn't "test on real queries," which everyone already knows and says. It was that I'd built a system whose main failure mode improves every metric I was watching. I should have seen that when I designed it, and I write the property down explicitly now, for every routing change: if this decision is wrong, what does it look like in the dashboard? If the answer is "better," I don't ship it without a human-labeled sample.

So if you were one of the people getting confident wrong answers from that thing in the spring, that's mine, sorry.

Aside · half finished

Related thing I started and abandoned: I wanted to know whether a small fine-tuned classifier actually beats an LLM router once you weight for p95 rather than raw accuracy. My hunch is that the classifier wins by a lot on the metric that matters, and that everyone uses the LLM router because it's four lines of code and accuracy is the number that gets reported.

I got as far as instrumenting both on one project and then the project changed shape and I never went back for the data. Might be a real result in there. Might be that the difference disappears once you're behind a cache. I don't know, and I'm not going to pretend I do.

Production

Log the route, or you can't debug anything

If you log only the query and the final answer, you cannot distinguish a bad answer caused by bad retrieval from a bad answer caused by a bad route. Those have completely different fixes and you will spend a week tuning your reranker to solve a routing bug. I have done this.

Per query, minimum:

  • routing decision, the confidence score, and the path actually taken (they diverge more than you'd like)
  • number of loop passes and where it terminated: converged, iteration cap, token budget, wall clock
  • cumulative tokens and cumulative wall clock
  • every query reformulation the loop generated, in order
  • retrieved chunk IDs per pass, and the grader's verdict per pass

That termination-reason field is the highest-value thing in the list and it's the one nobody logs. The distribution of termination reasons is your architecture's actual health metric. If the majority of your deep loops are ending on budget exhaustion rather than convergence, the loop isn't solving the problem, it's just spending your money on the way to the same answer, and you should route those queries somewhere else instead.

And then, once you have that logged: a spreadsheet. Query, route taken, pass count, termination reason, final answer, one column for a human to write yes or no. Still the best debugging tool for a broken RAG pipeline, still beats the observability platform that's been in onboarding for six weeks, still faintly embarrassing to recommend in 2026. I stand by it.

Evaluation

Routing accuracy is its own metric

Short section, because Part 7 is the whole story.

Routing accuracy and answer quality are different measurements and need different labels. You need a set of queries where a human has said "this one needed the loop" or "this one didn't," which is a separate labeling job from "was this answer correct." Without it, a routing regression is indistinguishable from a retrieval regression in your aggregate scores, and you'll chase the wrong one.

Also worth measuring separately: how often the loop's second pass actually improved on the first. Not whether the final answer was good. Whether the extra pass changed anything. If that number is low, your loop is a very expensive way of running retrieval twice, and the fix is upstream in Part 2, not here.

Aside · the running admission

I do not have a golden dataset for my own side project. I have been meaning to build one for about nine months now, which is up from eight months when I mentioned this in Part 1, so at least the number is honest. I'm aware of the shape of this. I'm writing the evaluation article anyway.

Here's what I keep running into, and I'd like to be wrong about it.

Every agentic RAG system I've looked at closely in production has, after tuning, converged to a loop that effectively runs one or two passes and then stops. Sometimes that's an explicit cap somebody added after an incident. More often it's implicit: the grader is tuned strict enough that pass three basically never fires, or the budget is set so that it can't. Either way the deployed reality is a corrective loop wearing an agent's clothes, and it's usually working fine, and nobody wants to say so because "we built a two-step retry" doesn't sound like the thing that got funded.

So: is there a system out there where the loop genuinely goes deeper than that, on a meaningful share of real traffic, and where the depth is measurably worth what it costs? Not a benchmark. Not a demo. Production, with the termination-reason distribution to back it up. I'd like to see it, because if it exists my priors here are too narrow, and if it doesn't then a lot of teams are paying agent prices for corrective behavior.

If you've got one, tell me what the numbers look like. Comment, DM, whatever's easiest.

Notes
  1. LangGraph raises GraphRecursionError once a graph exceeds the maximum number of steps; recursion_limit defaults to 25 and is raised by passing a higher value in the invocation config. docs.langchain.com back
  2. The OpenAI Agents SDK raises MaxTurnsExceeded when a run exceeds max_turns, which defaults to 10; passing max_turns=None disables the limit entirely. openai.github.io back
  3. Jeong, Baek, Cho, Hwang and Park, "Adaptive-RAG: Learning to Adapt Retrieval-Augmented Large Language Models through Question Complexity," NAACL 2024. The three-way split used here (no retrieval, single-step, iterative) and the small trained classifier that picks between them both come from this paper. arXiv:2403.14403 back
  4. Yan, Gu, Zhu and Ling, "Corrective Retrieval Augmented Generation," 2024. The lightweight retrieval evaluator that scores retrieved documents after the fact, discussed at length in Part 3. arXiv:2401.15884 back

Framework defaults in notes 1 and 2 were checked in August 2026 and will drift. Check the current docs before you rely on either number.

  • 01 · Choosing an architecture
  • 02 · Hybrid & reranking
  • 02a · Addendum: the middle ground
  • 03 · Corrective & Self-RAG
  • 04 · GraphRAG in practice
  • 05 · Adaptive & agentic
  • 06 · Multimodal
  • 07 · Evaluation & golden datasets

I publish a new article in this series every week: mostly what breaks in production, not what looked good in the demo. Part 6 is Multimodal RAG, where the retrieval problem stops being about text and the chunking assumptions we've spent five articles refining quietly stop applying. Follow along if you want it when it lands; I post here first, before anywhere else.