Latino Canon

Evals

How we check the AI in Latino Canon, and what we find.

Latino Canon uses AI in three places: search that understands half-remembered plots and Spanish, the short “why it matters” note on each title, and a server that lets AI assistants like Claude and ChatGPT use the canon.

AI gets things wrong. It can miss the film you meant, or say something about a film that no source supports - and for a canon about who gets represented and how, a confident mistake is worse than none. So instead of asking you to trust it, we test it after every change and publish the results here: what works, what doesn't yet, and what we're fixing next.

What we measure

Each number has a target, so you can tell at a glance whether it's good. They're our own targets: there's no industry standard for these, because a score depends on how hard the test is - finding a film by its exact title is far easier than from a half-remembered plot in Spanish.

A model handled 95% of 41 test requests exactly right.

The canon is also an MCP server that Claude, ChatGPT or Cursor can connect to. The server calls no model - the assistant brings its own - so its tool names and descriptions are all an assistant has to go on. This test gives a model only those, plus a request, and checks what it calls.

AI assistants using the canon

Pass rate
95%

Of 41 test requests, the share where the model did exactly the right thing.

Right tool
95%

Picked the right tool - or rightly used none for an off-topic question.

Right arguments
100%

When the tool was right: valid arguments, the right title id or filter, and no filter nobody asked for.

Hard requests
82%

Written to trip a model up: a film set in the 1940s (not made then), Spanish follow-ups, a title named in an off-topic question.

  • Look something up100% of 11
  • Search with a filter100% of 7
  • Open a result100% of 6
  • More like this100% of 5
  • Recommend100% of 6
  • Off-topic: no tool67% of 6
What it got wrong (2)
  • hard-none-translate-title (hard) - expected no tool, called search_titles({"query":"Y tu mamá también English translation"}). called search_titles for a request no tool covers
  • hard-none-pronounce (hard) - expected no tool, called search_titles({"query":"Alejandro González Iñárritu pronunciation"}). called search_titles for a request no tool covers

Latest run 2026-09-30 · Llama 3.3 70B · MCP server v1.7.0

How these evals work

Retrieval runs the 97-query golden set (known titles, people, half-remembered plots, facets, Spanish) against the live api after every deploy and weekly. Recall@5 is the share of each query's correct titles in the top 5; MRR is how high the first one ranks.

Groundedness gives a 70B LLM judge each blurb and the sources it was written from, and scores the share of its claims a source supports. Runs are started by hand from eval-groundedness.yml. The same judge gates new blurbs at ingest: only a fully supported, properly cited blurb is approved without an editor.

Tool selection connects to the live MCP server like any client, reads its instructions and tool list, and gives them to Llama 3.3 70B with each test request (some carry an earlier turn, as in "tell me more about the second one"). Some requests are written to trip a model up - a film set in the 1940s is not a 1940s film - and other models are scored on the same requests. A request passes when the model calls the right tool, or none for an off-topic question, with valid arguments, the expected title id or filter, and no search filter the person didn't ask for - filters exclude, so a guessed one can hide the answer. Runs are started by hand from eval-mcp-tools.yml, after a change to a tool's wording.

Runs marked invalid stay on the record. Their judge was given each source's reference (a title slug and a director's name) instead of its text, so it never saw the synopsis, and it ran at a non-zero temperature. Scores only compare within one judge version.

Unscored in the latest run: 10-years-of-taking-the-sky-by-storm-1991 (The judge's reply wasn't valid JSON, so the blurb went unscored rather than failed.)

Glossary
Golden set
Our test searches, each with the title(s) a good search should return, checked by hand against the live catalog.
Search types
How the test searches are grouped: exact title ("Selena"; engineers call it known-item), by person (a director or actor), by plot (a half-remembered story), genre, era or kind ("Mexican family stories from the 90s"; called facet), and in Spanish.
Recall@5
The share of test searches whose right answer appears in the top 5 results.
MRR
Mean reciprocal rank: how high the first right answer ranks - 1 for first place, ½ for second, ⅓ for third - averaged over all searches.
Hybrid search
Two searches combined: keyword matching (BM25), which finds exact words, and meaning matching (embeddings), which finds related ideas even in other words or languages.
Deploy gate
The rule that re-runs the tests after every code change and flags it if recall@5 drops more than 0.03 below the last passing run.
Baseline
The run a new one is compared against. Changing the golden set starts a new baseline, because scores only compare on the same tests.
Groundedness
The share of a note's claims that its cited sources (the synopsis, credits, awards) actually support.
LLM judge
An AI model (Llama 3.3 70B) that reads a note and its sources and lists any claim no source backs. Its version is frozen, so scores stay comparable.
Approved by the judge
A new note the judge found fully supported and properly cited, so it's shown without waiting for a person.
Held for an editor
A note with at least one unsupported claim. It isn't shown until an editor approves it, or a rewrite passes.
Tool selection (MCP)
Whether an AI assistant connected to our MCP server picks the right tool for a request - search, open a title, more like this, recommend - with the right details, or rightly uses none.
Invalid run
A past run we found was measuring the wrong thing. It stays on the record, labeled, rather than being deleted.
Run log (24 runs)
RunTypeJudgenFailedScore
Oct 1, 2026, 12:02 AMretrieval—9700.829
Sep 30, 2026, 11:23 PMretrieval—9700.829
Sep 30, 2026, 7:56 PMretrieval—9700.829
Sep 30, 2026, 6:04 PMretrieval—9700.829
Sep 29, 2026, 7:05 PMretrieval—9700.832
Sep 29, 2026, 6:27 PMretrieval—9700.829
Sep 29, 2026, 5:51 PMretrieval—9700.829
Sep 29, 2026, 6:26 AMretrieval—9700.839
Sep 29, 2026, 3:48 AMretrieval—9700.808
Sep 29, 2026, 2:35 AMretrieval—9700.853
Sep 29, 2026, 2:32 AMretrieval—9700.853
Sep 29, 2026, 12:31 AMretrieval—9700.817
Sep 28, 2026, 7:34 PMretrieval—9700.818
Sep 28, 2026, 6:44 PMretrieval—9700.818
Sep 28, 2026, 5:18 PMretrieval—9700.832
Sep 27, 2026, 3:17 AMgroundednessv320510.685
Sep 24, 2026, 6:55 AMgroundednessv321110.682
Sep 24, 2026, 6:35 AMgroundednessv221110.769
Sep 24, 2026, 6:33 AMgroundednessv221200.777
Sep 24, 2026, 6:03 AMgroundednessv121200.431invalid
Sep 22, 2026, 5:48 PMgroundednessv121200.525invalid
Sep 20, 2026, 6:40 AMgroundednessv121600.486invalid
Sep 17, 2026, 4:11 AMgroundednessv118700.487invalid
Sep 17, 2026, 4:03 AMgroundednessv118700.496invalid