There is no honest universal number. A receipt-extraction test with four stable layouts asks a different sampling question from a support assistant that must search thousands of changing documents. A prediction model for a rare event needs a different evaluation from a voice agent handling common appointment calls.
Start by listing the input families that could change the result. For documents, that may include layout, image quality, language, missing pages, handwritten fields, and older formats. For retrieval, it may include source type, permission level, conflicting guidance, date, and questions the system should refuse. For prediction, it may include time periods, relevant user or customer groups, rare outcomes, and shifts between training and future use.
The evaluation set should contain enough examples from each important group to expose material failure patterns. It should also preserve a final holdout that is not used to adjust prompts, rules, or models. If the sample is too small to estimate a claim, the handover should say so plainly and limit the decision to what the evidence can support.