Evals
How we check the AI in Latino Canon, and what we find.
Latino Canon uses AI in three places: search that understands half-remembered plots and Spanish, the short “why it matters” note on each title, and a server that lets AI assistants like Claude and ChatGPT use the canon.
AI gets things wrong. It can miss the film you meant, or say something about a film that no source supports - and for a canon about who gets represented and how, a confident mistake is worse than none. So instead of asking you to trust it, we test it after every change and publish the results here: what works, what doesn't yet, and what we're fixing next.
- Automatic. The search tests run after every code change; a drop of more than 0.03 flags it.
- Nothing hidden. Weak results stay on the page - they're the list of what to fix next.
- Comparable. The AI judge's version is frozen, so a score today compares with one last month.
- On the record. Runs later found to measure the wrong thing stay here, labeled invalid, not deleted.
What we measure
Each number has a target, so you can tell at a glance whether it's good. They're our own targets: there's no industry standard for these, because a score depends on how hard the test is - finding a film by its exact title is far easier than from a half-remembered plot in Spanish.
On average, 69% of a note's claims are backed by its sources. The most common problem: a significance claim no source makes (115 of 205 notes). +0.003 vs previous run
Groundedness: an AI judge checks each note, claim by claim, against the sources it was written from (synopsis, credits, awards). 1.0 means every claim is backed.
Why notes fail
In the latest run, 131 of 205 notes had a claim the judge couldn't match to a source, grouped by what went wrong.
- A significance claim no source makes115“It matters as a portrayal of counter-culture” The blurb prompt no longer asks for one, and the ingest gate holds any blurb that still makes one.
- A credit the judge couldn't match4“It was directed by Jon M. Chu.” Credit sources now name the work ("Heli (2013) is directed by…"), so the judge can match them.
- An interpretation beyond the synopsis6“It reimagines cinema and art.” Held for an editor: "It explores…" can be fair reading, but no source states it.
- A detail the sources don't state6“It's a comedic tale of lost jewelry” Held for an editor. The OMDb awards line now gives blurbs more to draw on.
Titles to fix first
The lowest-scoring notes and the claim the judge flagged. A flag means no source states the claim, not that it's false.
42nd Street2025score 0.50“It matters as a portrayal of counter-culture”
A 3 Minute Hug2018score 0.50“It matters as a portrayal of border life.”
A Better Life2011score 0.50“It matters as it portrays struggles of Indigenous Latino Americans”
Show 3 more
A Day Without a Mexican2004score 0.50“It matters because it highlights the economic impact of Mexican workers on the state”
A Dog's Will2000score 0.50“It matters as it showcases their journey to paradise, judged by Christ, the Devil, and the Virgin Mary.”
A Fantastic Woman2017score 0.50“It matters as it portrays her mourning process under scrutiny.”
How these evals work
Retrieval runs the 97-query golden set (known titles, people, half-remembered plots, facets, Spanish) against the live api after every deploy and weekly. Recall@5 is the share of each query's correct titles in the top 5; MRR is how high the first one ranks.
Groundedness gives a 70B LLM judge each blurb and the sources it was written from, and scores the share of its claims a source supports. Runs are started by hand from eval-groundedness.yml. The same judge gates new blurbs at ingest: only a fully supported, properly cited blurb is approved without an editor.
Tool selection connects to the live MCP server like any client, reads its instructions and tool list, and gives them to Llama 3.3 70B with each test request (some carry an earlier turn, as in "tell me more about the second one"). Some requests are written to trip a model up - a film set in the 1940s is not a 1940s film - and other models are scored on the same requests. A request passes when the model calls the right tool, or none for an off-topic question, with valid arguments, the expected title id or filter, and no search filter the person didn't ask for - filters exclude, so a guessed one can hide the answer. Runs are started by hand from eval-mcp-tools.yml, after a change to a tool's wording.
Runs marked invalid stay on the record. Their judge was given each source's reference (a title slug and a director's name) instead of its text, so it never saw the synopsis, and it ran at a non-zero temperature. Scores only compare within one judge version.
Unscored in the latest run: 10-years-of-taking-the-sky-by-storm-1991 (The judge's reply wasn't valid JSON, so the blurb went unscored rather than failed.)
Glossary
- Golden set
- Our test searches, each with the title(s) a good search should return, checked by hand against the live catalog.
- Search types
- How the test searches are grouped: exact title ("Selena"; engineers call it known-item), by person (a director or actor), by plot (a half-remembered story), genre, era or kind ("Mexican family stories from the 90s"; called facet), and in Spanish.
- Recall@5
- The share of test searches whose right answer appears in the top 5 results.
- MRR
- Mean reciprocal rank: how high the first right answer ranks - 1 for first place, ½ for second, ⅓ for third - averaged over all searches.
- Hybrid search
- Two searches combined: keyword matching (BM25), which finds exact words, and meaning matching (embeddings), which finds related ideas even in other words or languages.
- Deploy gate
- The rule that re-runs the tests after every code change and flags it if recall@5 drops more than 0.03 below the last passing run.
- Baseline
- The run a new one is compared against. Changing the golden set starts a new baseline, because scores only compare on the same tests.
- Groundedness
- The share of a note's claims that its cited sources (the synopsis, credits, awards) actually support.
- LLM judge
- An AI model (Llama 3.3 70B) that reads a note and its sources and lists any claim no source backs. Its version is frozen, so scores stay comparable.
- Approved by the judge
- A new note the judge found fully supported and properly cited, so it's shown without waiting for a person.
- Held for an editor
- A note with at least one unsupported claim. It isn't shown until an editor approves it, or a rewrite passes.
- Tool selection (MCP)
- Whether an AI assistant connected to our MCP server picks the right tool for a request - search, open a title, more like this, recommend - with the right details, or rightly uses none.
- Invalid run
- A past run we found was measuring the wrong thing. It stays on the record, labeled, rather than being deleted.
Run log (24 runs)
| Run | Type | Judge | n | Failed | Score |
|---|---|---|---|---|---|
| Oct 1, 2026, 12:02 AM | retrieval | — | 97 | 0 | 0.829 |
| Sep 30, 2026, 11:23 PM | retrieval | — | 97 | 0 | 0.829 |
| Sep 30, 2026, 7:56 PM | retrieval | — | 97 | 0 | 0.829 |
| Sep 30, 2026, 6:04 PM | retrieval | — | 97 | 0 | 0.829 |
| Sep 29, 2026, 7:05 PM | retrieval | — | 97 | 0 | 0.832 |
| Sep 29, 2026, 6:27 PM | retrieval | — | 97 | 0 | 0.829 |
| Sep 29, 2026, 5:51 PM | retrieval | — | 97 | 0 | 0.829 |
| Sep 29, 2026, 6:26 AM | retrieval | — | 97 | 0 | 0.839 |
| Sep 29, 2026, 3:48 AM | retrieval | — | 97 | 0 | 0.808 |
| Sep 29, 2026, 2:35 AM | retrieval | — | 97 | 0 | 0.853 |
| Sep 29, 2026, 2:32 AM | retrieval | — | 97 | 0 | 0.853 |
| Sep 29, 2026, 12:31 AM | retrieval | — | 97 | 0 | 0.817 |
| Sep 28, 2026, 7:34 PM | retrieval | — | 97 | 0 | 0.818 |
| Sep 28, 2026, 6:44 PM | retrieval | — | 97 | 0 | 0.818 |
| Sep 28, 2026, 5:18 PM | retrieval | — | 97 | 0 | 0.832 |
| Sep 27, 2026, 3:17 AM | groundedness | v3 | 205 | 1 | 0.685 |
| Sep 24, 2026, 6:55 AM | groundedness | v3 | 211 | 1 | 0.682 |
| Sep 24, 2026, 6:35 AM | groundedness | v2 | 211 | 1 | 0.769 |
| Sep 24, 2026, 6:33 AM | groundedness | v2 | 212 | 0 | 0.777 |
| Sep 24, 2026, 6:03 AM | groundedness | v1 | 212 | 0 | 0.431invalid |
| Sep 22, 2026, 5:48 PM | groundedness | v1 | 212 | 0 | 0.525invalid |
| Sep 20, 2026, 6:40 AM | groundedness | v1 | 216 | 0 | 0.486invalid |
| Sep 17, 2026, 4:11 AM | groundedness | v1 | 187 | 0 | 0.487invalid |
| Sep 17, 2026, 4:03 AM | groundedness | v1 | 187 | 0 | 0.496invalid |