I Finally Built the Golden Dataset: Evaluation for RAG in Production

RAG Architecture Series · Part 7

I Finally Built the Golden Dataset

Six articles ended with some version of “measure it.” This one shows how, starting with the dataset I spent most of the series admitting I didn’t have.

Reviewed September 2026 / Developer track / 17 min read

In Part 1 I admitted I didn’t have a golden dataset for my own side project. I’d “been meaning to build one for about eight months.” That was July. Add the summer and we’re well past ten.

It exists now. 212 questions, each with the evidence that should come back, a reference answer, and a label saying what kind of question it is. It took four evenings and one rainy weekend, which is a little embarrassing given how long I talked about it. And on the first run it told me five things I believed about my pipeline were wrong.

This is the last part of the series, and it’s the one every other part quietly depended on. Every “it depends,” every “tune this for your corpus,” every “check whether it actually helps” pointed here. You can’t check anything without something to check against.

Gut check before we start, same spirit as Part 1: sit down and write 20 questions your system absolutely must answer correctly, plus the document each answer lives in. If you can’t do that in an hour, you don’t have an evaluation problem yet. You have a scoping problem.

Nobody, including you, has decided what the system is for. Fix that first. Everything below assumes you can.

01 · The confession, paid off

The project, and what the dataset found

The side project is called Filing Desk. It answers questions over SEC filings: 10-Ks, 10-Qs and 8-Ks for 30 US companies over five years, half large tech companies, half banks. That’s about 2,100 filings and roughly 60,000 chunks after parsing. A handful of people I know in equity research and credit use it for the tedious part of their job: finding the paragraph, the footnote, or the table where a company actually said the thing.

I picked those two sectors on purpose. Tech companies and banks report in completely different shapes. A bank’s “revenue” isn’t even the same concept as a software company’s. Mix them and you get a corpus that breaks a RAG pipeline in almost every way this series covered: exact identifiers, tables that parsers mangle, documents that amend other documents, and questions that sound answerable but aren’t.

The pipeline is what this series recommended. Hybrid retrieval with BM25 and embeddings, a cross-encoder reranker on top (Part 2), a retrieval grader with a fallback path (Part 3). I thought it worked well. I’d vibes-tested it for months. Nobody was complaining much.

CategoryEntriesWhat it tests
Exact identifiers41Tickers, form types, fiscal periods
Narrative44What the company says: risk factors, MD&A
Numeric lookup38Values that live in financial tables
Comparison27Multi-hop: across companies, across years
Temporal23Latest filing, amendments, restated figures
Ambiguous23Should trigger a clarifying question
Unanswerable16Not disclosed, or forward looking

Sourcing: 138 came from real query logs, anonymized, but I kept the phrasing, abbreviations included. 51 I wrote by hand to cover gaps. 23 are synthetic, the survivors out of 120 I generated. Why so few made it is section 04.

What the first run showed

The reranker was hurting one category while helping overall. Recall@5 across all questions went up six points with the reranker on. On numeric lookups it went down, from 0.79 to 0.71. Financial tables get flattened into pipe-separated rows during parsing, and the cross-encoder scores those chunks low because they don’t read like prose answers. The overall average hid it completely. I would never have found this by poking at the chat window.

Exact identifiers were still the weakest category, even with hybrid search. Recall@5 of 0.64, against 0.86 for narrative questions. Part 1 was about my embeddings-only mistake; adding BM25 fixed the obvious half of it. What neither half fixed was fiscal periods. People type “FY24,” “fiscal 2024,” and “2024” and mean the same thing. Except they don’t, because several of the tech companies have fiscal years that end in June or September, so “fiscal 2024” and calendar 2024 are different filings. And “Q3” means something different for almost every company in the set.

Amendments and restatements lost to the originals. On 11 of 23 temporal questions, the top chunk came from the wrong version. Sometimes the original 10-K beat the 10-K/A. More often it was sneakier: every 10-K repeats prior-year figures as comparatives, and when a company restates, the old number is still sitting in last year’s filing, perfectly relevant, perfectly wrong. Nothing in the ranking knew what “superseded” means.

The fallback path was a liability here. The retrieval grader from Part 3 was routing about 19% of queries to web search. For a general-knowledge corpus that’s fine. For this one, the web results were mostly news articles and data aggregators quoting adjusted, non-GAAP figures, presented next to GAAP numbers as if they were the same thing.

And the numbers were wrong even when retrieval was right. This is the one that changed the architecture. On numeric questions where the correct chunk was in the top 5, roughly one answer in five was still wrong. Prior-year column instead of current year. Thousands read as millions. A segment subtotal reported as the total. The retriever did its job and the model misread the table anyway. I’d been proud of how well the pipeline “handled tables.” It didn’t. It just sounded like it did.

And the one I least wanted to see: of the 16 unanswerable questions, Filing Desk answered 9 confidently. Not hallucinated from nothing. Stitched together from the adjacent quarter, or from guidance language that reads like a forecast if you squint. In a research memo, that’s the insurer bot from Part 1 with a ticker symbol attached.

What I changed, and whether it worked, is section 06. First, the machinery.

Aside · Why it took almost a year

The honest reason is that labeling is boring and there’s nothing to demo. I can spend an evening adding a reranker and have something to show people. An evening of labeling produces a spreadsheet that’s 30 rows longer.

I kept telling myself I was waiting for a quiet week. I think the truth is less flattering. Whether that’s a personal failing or the actual reason most teams don’t have one either, I don’t know. Probably both. It would explain a lot of the vendor pitches I get.

02 · Two scores, not one

Score retrieval and generation separately

If you’ve read the rest of the series you know the claim: retrieval is the bottleneck, not generation. The practical consequence is that one end-to-end “answer quality” score isn’t enough. When it drops, you don’t know which half broke, and you’ll end up swapping the model because that’s the easiest knob to turn.

Yes, the table-misreading finding above is a case where generation was the problem. That’s the point of scoring separately. Without two numbers I’d have blamed the retriever and tuned it for a week.

Retrieval asks: did the right evidence come back, and how high? Three metrics cover most of it.

  • Recall@k. Of the evidence that should come back, what fraction is in the top k. Pick k to match what you actually pass to the model. If the prompt gets 5 chunks, recall@20 is a vanity number.
  • MRR (mean reciprocal rank). How high the first correct chunk lands. Good when one chunk is enough. It lies when an answer needs two chunks from two filings, which is exactly the comparison category.
  • nDCG. Rewards putting the most relevant chunks first, if you have graded relevance labels. Worth it only if you’ll actually label graded relevance. Most people won’t. I didn’t, for most entries.

Generation asks: given what was retrieved, is the answer right?

  • Faithfulness, or groundedness: is every claim supported by the retrieved context?
  • Correctness: does it match the reference answer?
  • Refusal correctness: does it say “I don’t know” when it should, and only then?

Faithfulness and correctness are not the same thing, and the gap between them is exactly where my 9 confident wrong answers lived. They were faithful. Every claim was backed by a retrieved chunk. The chunks were just from the wrong quarter. A faithfulness score on its own gives those answers full marks.

FIG.01 · TWO STAGE EVALUATIONRETRIEVAL STAGEGENERATION STAGEQUERYplus as_of dateRETRIEVERhybrid, rerankedRANKED IDStop k equals 5GENERATORANSWERRETRIEVAL SCORErecall@5partial creditmrrrank of first hitcompletebinary, all of itmodel not involvedGENERATION SCOREfaithful to the context?matches the reference?refused when it should?binary answers only
FIG.01Retrieval is scored against the expected evidence before the model is involved at all. Generation is scored separately, against the reference answer and the context it was actually given.
03 · Anatomy

What an entry looks like, and why I trust it

My first schema had a flat list of expected chunks. It fell apart on the comparison questions within a day.

Take a real one from the logs: how did one of the banks’ net interest margin change from 2022 to 2023? The FY2023 10-K answers it alone, because the MD&A table shows both years side by side. The FY2022 and FY2023 10-Ks together also answer it. With a flat list I had two bad options: require both filings, which punishes the retriever for finding the better single chunk, or accept either filing alone, which rewards half an answer.

So each entry now lists evidence paths. A path is a set of required pieces, and each piece can be satisfied by any of several interchangeable chunks. The score is the best path.

{"id": "fd-0183",
 "query": "bank-04 NIM 22 vs 23, did it expand",
 "category": "comparison",
 "answerable": true,
 "as_of": "2024-03-31",
 "evidence_paths": [
   [["bank-04/10-K/fy2023#mdna-nii-table"]],
   [["bank-04/10-K/fy2022#mdna-nii-table"],
    ["bank-04/10-K/fy2023#mdna-nii-table",
     "bank-04/10-K/fy2023#mdna-nim-text"]]
 ],
 "reference_answer": "NIM expanded by roughly 40 bp year over year;
   higher asset yields outpaced rising deposit costs.",
 "source": "log",
 "reviewed_by": 2,
 "added": "2026-08-14"}

One entry. JSONL, versioned in the repo with the indexing code.

And the scorer. Note that as_of goes into retrieval, not just into the file.

def path_recall(path, top):
    # each piece in a path is a set of interchangeable chunks
    return sum(1 for piece in path if top & set(piece)) / len(path)

def recall_at_k(paths, retrieved, k):
    if not paths:
        return None   # unanswerable: scored on the generation side
    top = set(retrieved[:k])
    return max(path_recall(p, top) for p in paths)

def reciprocal_rank(paths, retrieved):
    wanted = {c for p in paths for piece in p for c in piece}
    for rank, chunk_id in enumerate(retrieved, start=1):
        if chunk_id in wanted:
            return 1 / rank
    return 0.0

def evaluate(golden, retrieve, k=5):
    by_cat = {}
    for e in golden:
        # the retriever only sees filings EDGAR accepted on or before as_of,
        # so "most recent quarter" means what it meant on that date
        retrieved = retrieve(e["query"], as_of=e["as_of"])
        r = recall_at_k(e["evidence_paths"], retrieved, k)
        if r is None:
            continue
        rr = reciprocal_rank(e["evidence_paths"], retrieved)
        complete = 1.0 if r == 1.0 else 0.0   # every piece of one path
        by_cat.setdefault(e["category"], []).append((r, rr, complete))
    for cat, rows in sorted(by_cat.items()):
        n = len(rows)
        print(f"{cat:18} n={n:3}  recall@{k}={sum(x[0] for x in rows)/n:.2f}"
              f"  mrr={sum(x[1] for x in rows)/n:.2f}"
              f"  complete={sum(x[2] for x in rows)/n:.2f}")

Per category always. The average is where the reranker problem hid.

Those last two columns are not interchangeable, and I’ve watched myself confuse them. Mean recall gives partial credit: a comparison question that needs two filings and retrieves one scores 0.5. complete is the binary, every piece of at least one path present, which is the only version the user experiences. Six questions moving from 0.5 to 1.0 raise the complete-evidence rate by three points and the mean recall by one and a half. Same change, two different stories. The statistical aside at the end of this article is entirely about the binary one, and I’ll say so again there.

The as_of cutoff matters more than it looks. “What was the most recent quarterly revenue” has a different correct answer every three months. If the date only lives in the dataset and never reaches the retriever, two copies of that question with different dates get identical retrieval, and one of them is guaranteed to be scored wrong. The index stores each filing’s EDGAR acceptance date, and the retriever filters on it. Storing the date without enforcing it is decoration.

Two more things matter more than the code: report per category, always. And seed categories on purpose, because ambiguous, temporal and unanswerable questions almost never show up if you only sample popular queries.

Now the uncomfortable part: why call any of this golden?

The questions came from real logs, fine. But the evidence and reference answers came from me, and I’m exactly the person most likely to label whatever my own pipeline tends to find. Three rules keep that in check.

  1. Label from the corpus, not from the pipeline. For each entry I searched the filings directly, full-text over the raw documents plus the XBRL data for numbers, and listed every chunk that independently answers the question. I never looked at what Filing Desk retrieved while labeling. Entries ended up with 2.3 acceptable chunks on average, which says a lot about how unfair a single expected chunk would have been.
  2. “Not disclosed” needs a higher bar than “I didn’t find it.” An unanswerable label requires a full-text search of every relevant filing for that company and period, an empty XBRL lookup, and a second person agreeing. I originally had 18 unanswerable entries. Two turned out to be disclosed in footnotes I’d skipped. Those two became numeric entries, which is how the category ended up at 16. Remember this bar. It comes back in section 06, because the running system cannot clear it.
  3. Someone else checks my judgment. Two of the people who use Filing Desk, one equity analyst and one credit analyst, independently labeled a random 50 entries. We disagreed on 7. Four were my mistakes: two wrong fiscal years, one equally valid 10-Q chunk I’d missed, one restated figure. The other three were questions where all three of us had a defensible reading, so they moved to the ambiguous category. The reviewed_by field records how many people have signed off on an entry.

Same logic on the grading side. A correctness judge that rejects valid answers is as bad as a wrong label, so I fed it 20 correct answers in different formats (“$4.2 billion,” “4,213 million,” “up about 40 bp”) and checked that it accepted them. It rejected one, a rounding case, and the judge prompt now says explicitly what counts as equivalent.

Aside · The spreadsheet, again

In Part 1 I said the best debugging tool for a bad RAG pipeline is a spreadsheet with the query, the retrieved chunks, and a human yes/no column. That’s literally how this started. The JSONL came later, when I wanted it in CI.

For the first 80 entries it was a sheet with conditional formatting and a column called “wtf” for things I didn’t know how to label yet. I tried a proper annotation tool for about an hour. It wanted me to define a labeling schema before I’d seen enough data to know what the schema should be. Given that my schema changed completely on day two, I stand by that one.

04 · The fight

Synthetic golden datasets are grading your own homework

The popular shortcut: take your chunks, have an LLM write a question for each one, and that chunk becomes the expected answer. Minutes instead of evenings. Plenty of eval tooling will do this for you out of the box, and I get the appeal. I tried it.

300 synthetic questions, generated from randomly sampled chunks. Recall@5: 0.95. On the 138 real log questions, same pipeline, same day: 0.80.

The reason isn’t subtle. A question written while looking at a chunk borrows the chunk’s vocabulary. The filing says “subscription and support revenue,” so the question says “subscription and support revenue.” An actual analyst types “recurring rev” or “sub rev growth y/y.” The filing says “provision for credit losses,” the analyst types “reserve build.” The synthetic set tests whether your retriever can find a chunk when you hand it its own words back. Of course it can. That’s the easiest retrieval task there is.

Then it gets worse. The same setups usually score answers with an LLM judge, often from the same model family. The model writes the question, the model’s phrasing makes retrieval easy, the model grades the answer. The loop is closed, and the 95% at the end says almost nothing about your users.

There’s a coverage problem too. Each synthetic question comes from one chunk, so they’re almost never unanswerable, ambiguous or comparative. They cover precisely the categories you’re already good at and skip the ones that hurt.

I’m not saying never use them. I kept 23. My rules:

  1. Seed from real logs first. Always.
  2. Generate synthetic only to fill a category gap you can name.
  3. Rewrite each one the way a real user would phrase it, or have someone who talks to users do it.
  4. Throw away anything answerable by string-matching the source chunk.

Of 120 generated for gaps, 23 survived. That ratio is the point.

If your eval dashboard is green and built mostly on synthetic questions, I’d bet your users are having a worse time than the dashboard says. Happy to be proven wrong on that one, but I haven’t seen it yet.

FIG.02 · THE CLOSED LOOPCHUNKfrom your indexLLM WRITESTHE QUESTIONRETRIEVALfinds its own wordsLLM JUDGEsame model familySCORE 0.95means littleEVERY ARROW ISTHE SAME MODELREAL TRAFFICanalyst types “sub rev y/y”retrieval, no shared words0.80no loop, no flattery
FIG.02The model writes the question, its own phrasing makes retrieval easy, and it grades the result. Real traffic shares no vocabulary with the chunk, which is where the 15 points go.
05 · Judges

LLM-as-judge, on a leash

At any real volume you need an LLM judge for the generation side. Fine. Calibrate it before you trust it.

What I did: labeled 100 answers by hand, ran the judge on the same 100, compared. Overall agreement was 83%, which sounds acceptable. Per category it wasn’t. The judge marked 7 of the 9 confident wrong answers as good, because it was checking faithfulness against the retrieved context, and the answers were perfectly faithful to the wrong quarter.

The fix was boring. Give the judge the reference answer, the answerable flag and the as_of date. Ask two separate questions, both binary: is every claim supported by the context? Does the answer match the reference? Disagreement between those two is itself a useful signal.

The known biases are well documented and worth designing around. Judges tend to prefer longer answers. In pairwise comparisons they favor whichever answer comes first, so swap the order and run both. And they can favor output that sounds like their own model family. Re-check calibration every time you change the judge model, not just once.

One more thing, and it’s the insurer bot from Part 1 again. Nobody could trace which chunk the invented policy exclusion came from, because nothing was logged. Evaluation has the same dependency. If production only logs the question and the answer, you can never compute a single retrieval metric on real traffic. Log the retrieved chunk IDs and their scores for every query. It’s cheap, and there’s no way to add it retroactively. In finance it’s also the first thing a compliance team will ask for, so you might as well have it before they do.

Aside · On 1-to-10 scales

Pet peeve. Judge prompts that ask for a quality score from 1 to 10. What’s the difference between a 6 and a 7? Nobody knows, including the model. The scores cluster at 7 and 8, the average moves by 0.2, and somebody puts it on a slide as progress.

Binary questions look less impressive and are far more useful. “Does the answer use the correct fiscal year, yes or no.” I’d take ten of those over one 7.4.

06 · The rerun

What happened after the fixes

Here’s what I changed.

  • Fiscal periods normalized from filing metadata at index time, parsed out of the query, filtered before ranking.
  • Numeric questions routed to a plain lookup against the XBRL financial data EDGAR publishes. RAG handles the text around the number, not the number.
  • Web fallback disabled for anything numeric or company-specific.
  • Table chunks exempted from the reranker and merged back by their retrieval score.

Supersession recorded per item, not per document. This is the one I got wrong twice. My first version flagged an amended filing as superseding the original, full stop, and ranking demoted the original. It immediately started burying good MD&A chunks, because a 10-K/A that exists only to add the Part III information (directors, compensation, ownership) doesn’t touch the financial statements at all. The original stays authoritative for everything the amendment doesn’t reach. So the index now records which items a filing amends, and the original is demoted only for those items. A blanket flag is a worse bug than the one it fixes, because instead of ranking stale evidence too high it hides valid evidence, and nothing in the output tells you it happened.

A stricter refusal instruction, and the retrieval grader’s “insufficient evidence” path now ends in “I could not find this in the filings I searched,” not “not disclosed.” That distinction is the whole point of rule 02 above. Labeling something as not disclosed took me a full-text search of every relevant filing, an empty XBRL lookup, and a second person agreeing. A grader that has seen the top 20 chunks of one query has cleared none of that. It knows retrieval came up short, which is not the same as the corpus coming up empty, and a system that says “not disclosed” on a retrieval miss is wrong in exactly the insurer-bot way, just pointed in the other direction. It also reads worse to an analyst, who knows perfectly well the company discloses that.

And the rerun. The middle column is the same 212 questions I’d just spent two weeks tuning against, so treat it as optimistic. The right column is 60 new questions from the logs of the month after the fixes, labeled before I looked at what the pipeline did with them.

MetricFirst run (212)After fixes, same setFresh sample (60)
Mean recall@5, all answerable0.790.880.83
Complete evidence retrieved0.710.840.78
Mean recall@5, exact identifiers0.640.900.85 (n=13)
Numeric answers correct24 of 3836 of 3811 of 12
Temporal: right version first12 of 2320 of 236 of 8
Unanswerable answered anyway9 of 163 of 162 of 6
Queries sent to web fallback19%4%5%

Row one is the partial-credit mean, row two the binary. They move together here, which is the pleasant case. When they don’t, the binary is the one I believe, because half the evidence for a comparison question is not half an answer.

The gap between the middle and right columns is the part I’d want a client to look at. Five points of mean recall went away the moment the questions were new. Some of that is sample size, the fresh categories are tiny. Some of it, I’m fairly sure, is that I’d tuned to my own test.

What’s still broken: the remaining numeric misses are custom metrics companies report outside the standard XBRL tags, where the lookup has nothing and falls back to reading tables. The temporal misses are restatements that happen quietly inside a later filing’s comparatives, with no amendment to flag and no item to attach a supersession to. And 2 of 6 fresh unanswerable questions still got a confident answer. Both were forward-looking questions phrased as if they were about the past. That one I don’t have a fix for yet.

Which brings up overfitting. Every time I check a change against the 212 and keep it because the number went up, I’m selecting on that set. Do that for two weeks and the set stops measuring “does this work” and starts measuring “does this work on these 212 questions.” The CI gate is still the right tool for catching regressions. It’s the wrong tool for claiming improvement.

So there are now two sets with two jobs. The main set, 231 entries and growing, is the regression suite: I tune against it freely and it’s allowed to be a collection of old failures. The fresh sample is 50 to 60 new questions from the most recent month’s logs, labeled blind, used once to report how the system does on current traffic, and never tuned against. Next month it’s replaced, and its failures get added to the main set. If I only had time for one, I’d keep the fresh sample. It’s the only number I’d put in front of someone else.

07 · In the workflow

Keeping it alive

A golden dataset in a folder is a snapshot. With filings it rots on a schedule: every quarter, new 10-Qs land and “latest” means something else.

  • Every retrieval change triggers a run. Chunking, embedding model, reranker, grader threshold, prompt. About six minutes on Filing Desk, and the CI job fails if any category’s complete-evidence rate drops by more than a set margin. Any category, not the average.
  • Every production failure becomes an entry. That’s where the 19 entries since the first run came from.
  • Version the dataset with the index. When filings get re-parsed, chunk IDs go stale. I learned this when a parser update shifted section anchors and recall dropped to 0.3 overnight. The pipeline was fine. The labels pointed at chunks that no longer existed. There’s now a pre-check that every expected chunk exists in the current index before scoring starts.
  • Re-check time-sensitive entries every filing season. Anything with “latest,” “most recent” or “current” in it gets reviewed after each quarter’s filings arrive.
FIG.03 · TWO SETS, TWO JOBSLANE 01 · MEASURE, NEVER TUNELANE 02 · GUARD, TUNE FREELYPRODUCTIONlogged chunk idsFRESH SAMPLE50 to 60, labeled blindREPORTED ONCEthe client numberNEXT MONTHsample replacedFILINGSEASONre-check “latest”MAIN SET231 entriesversioned with indexCI GATEper categorycomplete evidenceSHIPnext changeits failures, and only thoseback into production, where the next failures come from
FIG.03The main set remembers old failures and may be tuned against. The fresh sample measures current traffic once, then retires into the main set.
Aside · How many is enough?

The question I get most, and one I got wrong in my head for a long time. Everything below is about the complete-evidence rate, the binary from the scorer: every piece of at least one evidence path in the top 5, yes or no. Not the partial-credit mean recall in the table above. The two move differently, and six questions going from 0.5 to 1.0 shift the binary by three points and the mean by one and a half, so a test run on one of them says nothing about the other.

At 200 entries and a rate around 0.8, the standard error on a single score is about 0.028, a 95% interval of roughly plus or minus 5.5 points. I used to read that as “a 3-point change is noise.” That’s the wrong test. Both runs use the same questions, so what matters is the paired result: which questions flipped, and in which direction.

Count the flips. If a change fixes 6 questions and breaks none, that’s a 3-point gain with an exact McNemar p-value of about 0.03. Same 3 points from fixing 16 and breaking 10: p around 0.33, and I wouldn’t trust it. The net number is identical. The flip counts are the whole story, so the CI report now prints them. If you’d rather test the mean recall directly, McNemar doesn’t apply, because those aren’t binary outcomes; you’re looking at the per-question differences and want a signed-rank test over the ones that moved.

Per category, with 20 to 40 entries, you rarely get enough flips to conclude anything, which is part of why the fresh sample exists. What I still haven’t worked out is how to account for the dozens of comparisons I’ve run against the same main set. Every test I run makes the next p-value a bit less honest. I know the textbook corrections exist. I haven’t decided which one I actually believe applies to a CI gate.

Seven parts. Looking back, a lot of the advice in this series came down to “it depends on your corpus.” The golden dataset is how you find out what it depends on. It’s the most useful few hundred lines in the whole project, and I built them last.

What I haven’t figured out: when does an old failure stop deserving a place in the main set? The fresh sample tells me what’s true now. The main set remembers what went wrong before, and at some point it turns into a museum of old bugs, weighted toward whatever broke most often last year. Prune it? Cap each category? Retire entries once they’ve passed for four straight filing seasons? I don’t have a rule I trust yet. If you do, I want to hear it.

  • 01 · Choosing an architecture
  • 02 · Hybrid & reranking
  • 02a · Addendum: the middle ground
  • 03 · Corrective & Self-RAG
  • 04 · GraphRAG in practice
  • 05 · Adaptive & agentic
  • 06 · Multimodal
  • 07 · Evaluation & golden datasets

That’s the series. Seven parts, mostly about what breaks in production, not what looked good in the demo. Something new is already in the works: a different angle, same rule. Stay tuned; I post here first, before anywhere else.

And if you’ve built a golden dataset: tell me how many entries, how long it took, and the first thing it told you that you didn’t want to hear. Comment, DM, whatever.