Domain-specific benchmarks built from verified sources reveal that LLM performance gaps aren't just about model size—reasoning capabilities, self-preference bias, and access to reference materials dramatically change accuracy, with some models gaining 33+ percentage points when given context.
OenoBench is a wine-domain benchmark with 3,266 multiple-choice questions built from verified facts extracted from government registries and peer-reviewed sources.