You can evaluate LLMs on subjective tasks by checking if their responses follow theoretical predictions and internal consistency, rather than comparing to a ground truth—a framework borrowed from decades of economics research.
This paper adapts validity frameworks from economics to evaluate language models on subjective questions without ground truth answers. Using stated-preference survey methods, the authors test whether LLM responses follow economic theory predictions (like downward-sloping demand curves), demonstrating that validity testing can assess model coherence even when correct answers don't exist.