Instead of using a single LLM to judge AI readiness, using multiple specialized LLMs with a conservative aggregation rule produces more reliable and honest assessments of project maturity.
This paper creates AIRL, a unified 9-level AI readiness scale combining three existing frameworks, and RAIL, an automated classifier using multiple specialized LLMs to assess AI project maturity. The system evaluates projects across six dimensions (data, specifications, expertise, algorithms) and prevents overestimation through a conservative review process.