Evals
How we check the AI in Latino Canon, and what we find.
Latino Canon uses AI in three places: search that understands half-remembered plots and Spanish, the short “why it matters” note on each title, and a server that lets AI assistants like Claude and ChatGPT use the canon.
AI gets things wrong. It can miss the film you meant, or say something about a film that no source supports - and for a canon about who gets represented and how, a confident mistake is worse than none. So instead of asking you to trust it, we test it after every change and publish the results here: what works, what doesn't yet, and what we're fixing next.
- Automatic. The search tests run after every code change; a drop of more than 0.03 flags it.
- Nothing hidden. Weak results stay on the page - they're the list of what to fix next.
- Comparable. The AI judge's version is frozen, so a score today compares with one last month.
- On the record. Runs later found to measure the wrong thing stay here, labeled invalid, not deleted.
What we measure
Each number has a target, so you can tell at a glance whether it's good. They're our own targets: there's no industry standard for these, because a score depends on how hard the test is - finding a film by its exact title is far easier than from a half-remembered plot in Spanish.
A model handled 95% of 41 test requests exactly right.
The canon is also an MCP server that Claude, ChatGPT or Cursor can connect to. The server calls no model - the assistant brings its own - so its tool names and descriptions are all an assistant has to go on. This test gives a model only those, plus a request, and checks what it calls.
AI assistants using the canon
Of 41 test requests, the share where the model did exactly the right thing.
Picked the right tool - or rightly used none for an off-topic question.
When the tool was right: valid arguments, the right title id or filter, and no filter nobody asked for.
Written to trip a model up: a film set in the 1940s (not made then), Spanish follow-ups, a title named in an off-topic question.
- Look something up“What has Gael García Bernal been in?”100% of 11
- Search with a filter“Horror movies from Argentina”100% of 7
- Open a result“Tell me more about the second one”100% of 6
- More like this“More like Coco, please”100% of 5
- Recommend“Something uplifting by a Latina director”100% of 6
- Off-topic: no tool“What's the capital of Peru?”67% of 6
What it got wrong (2)
hard-none-translate-title(hard) - expected no tool, called search_titles({"query":"Y tu mamá también English translation"}). called search_titles for a request no tool covershard-none-pronounce(hard) - expected no tool, called search_titles({"query":"Alejandro González Iñárritu pronunciation"}). called search_titles for a request no tool covers
Latest run 2026-09-30 · Llama 3.3 70B · MCP server v1.7.0
How these evals work
Retrieval runs the 97-query golden set (known titles, people, half-remembered plots, facets, Spanish) against the live api after every deploy and weekly. Recall@5 is the share of each query's correct titles in the top 5; MRR is how high the first one ranks.
Groundedness gives a 70B LLM judge each blurb and the sources it was written from, and scores the share of its claims a source supports. Runs are started by hand from eval-groundedness.yml. The same judge gates new blurbs at ingest: only a fully supported, properly cited blurb is approved without an editor.
Tool selection connects to the live MCP server like any client, reads its instructions and tool list, and gives them to Llama 3.3 70B with each test request (some carry an earlier turn, as in "tell me more about the second one"). Some requests are written to trip a model up - a film set in the 1940s is not a 1940s film - and other models are scored on the same requests. A request passes when the model calls the right tool, or none for an off-topic question, with valid arguments, the expected title id or filter, and no search filter the person didn't ask for - filters exclude, so a guessed one can hide the answer. Runs are started by hand from eval-mcp-tools.yml, after a change to a tool's wording.
Runs marked invalid stay on the record. Their judge was given each source's reference (a title slug and a director's name) instead of its text, so it never saw the synopsis, and it ran at a non-zero temperature. Scores only compare within one judge version.
Unscored in the latest run: 10-years-of-taking-the-sky-by-storm-1991 (The judge's reply wasn't valid JSON, so the blurb went unscored rather than failed.)
Glossary
- Golden set
- Our test searches, each with the title(s) a good search should return, checked by hand against the live catalog.
- Search types
- How the test searches are grouped: exact title ("Selena"; engineers call it known-item), by person (a director or actor), by plot (a half-remembered story), genre, era or kind ("Mexican family stories from the 90s"; called facet), and in Spanish.
- Recall@5
- The share of test searches whose right answer appears in the top 5 results.
- MRR
- Mean reciprocal rank: how high the first right answer ranks - 1 for first place, ½ for second, ⅓ for third - averaged over all searches.
- Hybrid search
- Two searches combined: keyword matching (BM25), which finds exact words, and meaning matching (embeddings), which finds related ideas even in other words or languages.
- Deploy gate
- The rule that re-runs the tests after every code change and flags it if recall@5 drops more than 0.03 below the last passing run.
- Baseline
- The run a new one is compared against. Changing the golden set starts a new baseline, because scores only compare on the same tests.
- Groundedness
- The share of a note's claims that its cited sources (the synopsis, credits, awards) actually support.
- LLM judge
- An AI model (Llama 3.3 70B) that reads a note and its sources and lists any claim no source backs. Its version is frozen, so scores stay comparable.
- Approved by the judge
- A new note the judge found fully supported and properly cited, so it's shown without waiting for a person.
- Held for an editor
- A note with at least one unsupported claim. It isn't shown until an editor approves it, or a rewrite passes.
- Tool selection (MCP)
- Whether an AI assistant connected to our MCP server picks the right tool for a request - search, open a title, more like this, recommend - with the right details, or rightly uses none.
- Invalid run
- A past run we found was measuring the wrong thing. It stays on the record, labeled, rather than being deleted.
Run log (24 runs)
| Run | Type | Judge | n | Failed | Score |
|---|---|---|---|---|---|
| Oct 1, 2026, 12:02 AM | retrieval | — | 97 | 0 | 0.829 |
| Sep 30, 2026, 11:23 PM | retrieval | — | 97 | 0 | 0.829 |
| Sep 30, 2026, 7:56 PM | retrieval | — | 97 | 0 | 0.829 |
| Sep 30, 2026, 6:04 PM | retrieval | — | 97 | 0 | 0.829 |
| Sep 29, 2026, 7:05 PM | retrieval | — | 97 | 0 | 0.832 |
| Sep 29, 2026, 6:27 PM | retrieval | — | 97 | 0 | 0.829 |
| Sep 29, 2026, 5:51 PM | retrieval | — | 97 | 0 | 0.829 |
| Sep 29, 2026, 6:26 AM | retrieval | — | 97 | 0 | 0.839 |
| Sep 29, 2026, 3:48 AM | retrieval | — | 97 | 0 | 0.808 |
| Sep 29, 2026, 2:35 AM | retrieval | — | 97 | 0 | 0.853 |
| Sep 29, 2026, 2:32 AM | retrieval | — | 97 | 0 | 0.853 |
| Sep 29, 2026, 12:31 AM | retrieval | — | 97 | 0 | 0.817 |
| Sep 28, 2026, 7:34 PM | retrieval | — | 97 | 0 | 0.818 |
| Sep 28, 2026, 6:44 PM | retrieval | — | 97 | 0 | 0.818 |
| Sep 28, 2026, 5:18 PM | retrieval | — | 97 | 0 | 0.832 |
| Sep 27, 2026, 3:17 AM | groundedness | v3 | 205 | 1 | 0.685 |
| Sep 24, 2026, 6:55 AM | groundedness | v3 | 211 | 1 | 0.682 |
| Sep 24, 2026, 6:35 AM | groundedness | v2 | 211 | 1 | 0.769 |
| Sep 24, 2026, 6:33 AM | groundedness | v2 | 212 | 0 | 0.777 |
| Sep 24, 2026, 6:03 AM | groundedness | v1 | 212 | 0 | 0.431invalid |
| Sep 22, 2026, 5:48 PM | groundedness | v1 | 212 | 0 | 0.525invalid |
| Sep 20, 2026, 6:40 AM | groundedness | v1 | 216 | 0 | 0.486invalid |
| Sep 17, 2026, 4:11 AM | groundedness | v1 | 187 | 0 | 0.487invalid |
| Sep 17, 2026, 4:03 AM | groundedness | v1 | 187 | 0 | 0.496invalid |