span-01 vs mercury-decide: same score, opposite failures
span-01 vs mercury-decide: same score, opposite failures Last time I tested a "decision model" — a model that takes a plain-language question about a text and answers with a probability — as a gate for keeping Japanese narration free of English words. That article is here: Is regex enough? I tested span-01 on mixed-language text Code and measured data: / span01-eval span01-eval Evaluating prompt-defined "decision models" (span-01-lite, mercury-decide) as a language gate — against a plain regex and a generic chat model — with every number recomputed from saved JSON artifacts. English | 日本語 Can
This is an AI-generated summary. ShortSingh links to the original source for the complete article.



Discussion (0)
Log in to join the discussion and vote.
Log in