Latino Canon

Evals

How we check the AI in Latino Canon, and what we find.

Latino Canon uses AI in three places: search that understands half-remembered plots and Spanish, the short “why it matters” note on each title, and a server that lets AI assistants like Claude and ChatGPT use the canon.

AI gets things wrong. It can miss the film you meant, or say something about a film that no source supports - and for a canon about who gets represented and how, a confident mistake is worse than none. So instead of asking you to trust it, we test it after every change and publish the results here: what works, what doesn't yet, and what we're fixing next.

What we measure

Each number has a target, so you can tell at a glance whether it's good. They're our own targets: there's no industry standard for these, because a score depends on how hard the test is - finding a film by its exact title is far easier than from a half-remembered plot in Spanish.

About 83 of every 100 test searches find the right title in the top 5. Genre, era or kind searches are the weakest (64% in the top 5), like “Mexican family stories from the 90s”.

Recall@5 is the share of our 97 test searches whose right answer lands in the top 5 results - for example “dos estafadores venden estampillas falsas a un coleccionista” → Nine Queens. MRR is how high it ranks: 1 for first place, ½ for second.

Search quality over time

Hybrid recall@5 across the last 15 retrieval runs, one after every api deploy. A run fails the deploy gate when it drops more than 0.03 below the last passing run.

0.750.800.850.90Sep 28Oct 10.829gate failed

no change vs previous run+0.000 vs last passing run (gate limit −0.03)

Where search is weakest

Recall@5 by search type in the latest run. Weakest first.

  • Genre, era or kind
    0.639 · MRR 0.68
  • By person
    0.800 · MRR 0.68
  • By plot
    0.822 · MRR 0.75
  • In Spanish
    0.861 · MRR 0.68
  • Exact title
    1.000 · MRR 1.00
How these evals work

Retrieval runs the 97-query golden set (known titles, people, half-remembered plots, facets, Spanish) against the live api after every deploy and weekly. Recall@5 is the share of each query's correct titles in the top 5; MRR is how high the first one ranks.

Groundedness gives a 70B LLM judge each blurb and the sources it was written from, and scores the share of its claims a source supports. Runs are started by hand from eval-groundedness.yml. The same judge gates new blurbs at ingest: only a fully supported, properly cited blurb is approved without an editor.

Tool selection connects to the live MCP server like any client, reads its instructions and tool list, and gives them to Llama 3.3 70B with each test request (some carry an earlier turn, as in "tell me more about the second one"). Some requests are written to trip a model up - a film set in the 1940s is not a 1940s film - and other models are scored on the same requests. A request passes when the model calls the right tool, or none for an off-topic question, with valid arguments, the expected title id or filter, and no search filter the person didn't ask for - filters exclude, so a guessed one can hide the answer. Runs are started by hand from eval-mcp-tools.yml, after a change to a tool's wording.

Runs marked invalid stay on the record. Their judge was given each source's reference (a title slug and a director's name) instead of its text, so it never saw the synopsis, and it ran at a non-zero temperature. Scores only compare within one judge version.

Unscored in the latest run: 10-years-of-taking-the-sky-by-storm-1991 (The judge's reply wasn't valid JSON, so the blurb went unscored rather than failed.)

Glossary
Golden set
Our test searches, each with the title(s) a good search should return, checked by hand against the live catalog.
Search types
How the test searches are grouped: exact title ("Selena"; engineers call it known-item), by person (a director or actor), by plot (a half-remembered story), genre, era or kind ("Mexican family stories from the 90s"; called facet), and in Spanish.
Recall@5
The share of test searches whose right answer appears in the top 5 results.
MRR
Mean reciprocal rank: how high the first right answer ranks - 1 for first place, ½ for second, ⅓ for third - averaged over all searches.
Hybrid search
Two searches combined: keyword matching (BM25), which finds exact words, and meaning matching (embeddings), which finds related ideas even in other words or languages.
Deploy gate
The rule that re-runs the tests after every code change and flags it if recall@5 drops more than 0.03 below the last passing run.
Baseline
The run a new one is compared against. Changing the golden set starts a new baseline, because scores only compare on the same tests.
Groundedness
The share of a note's claims that its cited sources (the synopsis, credits, awards) actually support.
LLM judge
An AI model (Llama 3.3 70B) that reads a note and its sources and lists any claim no source backs. Its version is frozen, so scores stay comparable.
Approved by the judge
A new note the judge found fully supported and properly cited, so it's shown without waiting for a person.
Held for an editor
A note with at least one unsupported claim. It isn't shown until an editor approves it, or a rewrite passes.
Tool selection (MCP)
Whether an AI assistant connected to our MCP server picks the right tool for a request - search, open a title, more like this, recommend - with the right details, or rightly uses none.
Invalid run
A past run we found was measuring the wrong thing. It stays on the record, labeled, rather than being deleted.
Run log (24 runs)
RunTypeJudgenFailedScore
Oct 1, 2026, 12:02 AMretrieval—9700.829
Sep 30, 2026, 11:23 PMretrieval—9700.829
Sep 30, 2026, 7:56 PMretrieval—9700.829
Sep 30, 2026, 6:04 PMretrieval—9700.829
Sep 29, 2026, 7:05 PMretrieval—9700.832
Sep 29, 2026, 6:27 PMretrieval—9700.829
Sep 29, 2026, 5:51 PMretrieval—9700.829
Sep 29, 2026, 6:26 AMretrieval—9700.839
Sep 29, 2026, 3:48 AMretrieval—9700.808
Sep 29, 2026, 2:35 AMretrieval—9700.853
Sep 29, 2026, 2:32 AMretrieval—9700.853
Sep 29, 2026, 12:31 AMretrieval—9700.817
Sep 28, 2026, 7:34 PMretrieval—9700.818
Sep 28, 2026, 6:44 PMretrieval—9700.818
Sep 28, 2026, 5:18 PMretrieval—9700.832
Sep 27, 2026, 3:17 AMgroundednessv320510.685
Sep 24, 2026, 6:55 AMgroundednessv321110.682
Sep 24, 2026, 6:35 AMgroundednessv221110.769
Sep 24, 2026, 6:33 AMgroundednessv221200.777
Sep 24, 2026, 6:03 AMgroundednessv121200.431invalid
Sep 22, 2026, 5:48 PMgroundednessv121200.525invalid
Sep 20, 2026, 6:40 AMgroundednessv121600.486invalid
Sep 17, 2026, 4:11 AMgroundednessv118700.487invalid
Sep 17, 2026, 4:03 AMgroundednessv118700.496invalid